7,165 Matching Annotations
  1. Last 7 days
    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version.

      As discussed before, the authors employ a wide range of techniques (FOS IHC, FP for fine scale PVN OXT population dynamics, behavioural analysis, core and surface temperature tracking, physiological recordings to assess AAV specificity, optogenetic activation of PVN OXT neurons, and projection tracing) to address a clear question. The outcomes of these techniques seem to drive the same conclusion that PVN OXT neurons signal transitions from rest to arousal (behavioural and thermogenic) in a state-dependent manner:

      - FOS data identifies PVN OXT population activity following behavioural onset

      - Ca activity in these cells peaks at behavioural and thermogenic state transitions

      - Rump temperature and BAT activity increase at state transition points

      - Optogenetic stimulation of these cells recapitulates the thermogenic effects seen during physiological state transitions (in low body temperature animals) with a trending increase in physical activity

      Despite the inconclusive IHC results when validating the specificity of their AAV, the virgin female/ lactation experiment is convincing that they are specifically targeting PVN OXT neurons. The rationale for this experiment is clearer in the revised manuscript.

      Generally, in terms of the revised manuscript, the authors give strong responses to reviewer comments, either incorporating feedback, or giving clear explanations for the choices they made in the original manuscript. The revised manuscript is clearer about the question the authors aim to address, the reasons for their choice of experiments, and the limitations of the techniques used.

      We thank the reviewer for the close attention to the manuscript, the response to reviewers, and the revision, all of which have improved the manuscript.

      Criticisms:

      I appreciate and agree with the authors' point that this manuscript is more fundamental than simply social basis oxytocin neuron function. This is point is well made by their data, and in the revised text. However, I still believe more behavioural analysis would be welcome to any reader.

      They partly justify the lack of behavioural analysis in Figure 6 with the problem of "animal merging" on the SGBS images. However, in Figure 6C, they confirm that, in solo conditions, the SGBS readings are consistent with core body temperature readings. So why not stick to core body temperature, opto stimulate and analyse the social behaviour with DLC (with normal video recordings)?

      This is a good suggestion. Because we find that quiescent huddling (paired) bouts were associated with stronger body temperature regulation compared to solo quiescence and other behavioral states, and because PVNOT peak probability and frequency were higher in the paired compared to solo context, these experiments are warranted. We made the following edits to the discussion:

      “Future experiments should attempt to disentangle the effects of PVNOT light stimulation on social vs. non-social aspects of these behavioral state transitions; of particular interest would be to examine how light stimulation affects the duration and thermoregulatory control of social huddling.”

      The lactation validation still seems out of place in manuscript order. It is a very valuable validation, but it feels more like supplementary data for Figure 1. I feel the authors wanted it as a main figure because of how much work it must have been. In that case, it still makes more sense to include it in Figure 1.

      The purpose of the lactation experiment arose from the inadequacy of using histology to test whether AAV-transfected cells were oxytocin-immunoreactive. Because we observed intense oxytocin immunoreactivity in the fibres lining the ventricle, and less reactivity in the cell bodies than what we would have predicted from the Oxytocin-Cre-dependent AAV, we turned to the known physiological relationship between oxytocin-positive neurons and lactation. As such, this study is not associated with Figure 1, which demonstrates our initial, coarse-grained findings relating FOS activity in the PVN and in oxytocin-positive neurons during social thermoregulation.

      To your point, it typically does make sense to have the cellular validation “up front” as supporting or background information that enables the downstream experiments. However, what gives this data credibility as a standalone figure is the novel finding that PVNOT neurons display burst-like patterns of activity outside the context of lactation. Previous discussions with experts in the field, along with a review of the literature, unexpectedly led us to the observation that the burst-like patterns we observed during the transition from rest to wake and thermogenesis in virgin females represents a new aspect of oxytocin neuron physiology. Because we wanted to directly compare the new virgin female activity pattern (i.e., Figure 2) with the known lactation activity pattern, we decided it made the most sense to combine the validation aspect with the novel aspect into a standalone figure.

      Though their lactation experiment validates that they are targeting PVN OXT neurons, their optogenetic stimulation protocol may not be specifically inducing OXT release from these cells. PVN OXT neurons co-release glutamate but can also release glutamate independently of OXT following lower frequency tonic stimulation. OXT release from PVN neurons requires pulsatile stimulation at a higher frequency (Leithead et al., 2021; Piñol et al., 2014; Lincoln & Wakerley, 1975). In this paper, the authors use a low stimulation frequency (10Hz) and continuous pulse train (20s) to optogenetically manipulate the target PVN population which may bias the cells towards glutamate release over OXT. Therefore, though they find evidence that PVN OXT neurons are involved in driving the transition between states in their other experiments, their optogenetic stimulation may not necessarily involve OXT release/signalling. It may be valuable to separate this out to identify the signalling molecule underlying this behavioural/ thermogenic transition. This could be done by using an opto protocol that recapitulates physiological OXT release.

      The authors do however mention that isolating the specific contribution of OXT signalling compared to other co-transmitted molecules was not the aim of this study, so this is not an essential question for this manuscript.

      Thank you for this thoughtful point. We agree our optogenetic stimulation experiment should be interpreted as activation of PVNOT neurons rather than as selective evidence for oxytocin release or oxytocin signaling. PVNOT neurons can co-release glutamate (an idea we had also briefly touched upon in the Limitations and caveats section), and the stimulation pattern/frequency may influence the relative engagement of fast glutamatergic transmission versus peptide release. We agree the lactation literature, including Lincoln et al., highlights the importance of high-frequency pulsatile activity for oxytocin release, and that Piñol et al. provide evidence that PVNOT-linked glutamatergic transmission can interact with oxytocin-receptor-dependent modulation of downstream synapses–so thanks for pointing these out.

      We made revisions to support our protocol and now acknowledge this important aspect of the neuronal physiology. In Results, we now explain why we selected 10Hz: this frequency was grounded in the study by Fukushima et al. (2022), where 10Hz stimulation of PVNOT terminals in the rMR elicit thermogenic responses and 10Hz stimulation of PVNOT somata produce thermogenesis that’s dependent on oxytocin receptors in rMR.

      In the Limitations section, we now cite these three references to include broader context around stimulation frequency and differential release. We emphasize that our optogenetic data demonstrate sufficiency of PVNOT neuron activation, but do not establish whether the downstream thermogenic and behavioral effects are mediated by oxytocin, glutamate, or both. We note that resolving this issue will require future experiments using stimulation-pattern comparisons together with receptor-targeted pharmacology or genetic loss-of-function approaches.

      References

      Leithead, A. B., Tasker, J. G., & Harony-Nicolas, H. (2021). The interplay between glutamatergic circuits and oxytocin neurons in the hypothalamus and its relevance to neurodevelopmental disorders. Journal of neuroendocrinology, 33(12), e13061. https://doi.org/10.1111/jne.13061

      Lincoln, D. W., & Wakerley, J. B. (1975). Factors governing the periodic activation of supraoptic and paraventricular neurosecretory cells during suckling in the rat. The Journal of physiology, 250(2), 443-461. https://doi.org/10.1113/jphysiol.1975.sp011064

      Piñol, R. A., Jameson, H., Popratiloff, A., Lee, N. H., & Mendelowitz, D. (2014). Visualization of oxytocin release that mediates paired pulse facilitation in hypothalamic pathways to brainstem autonomic neurons. PloS one, 9(11), e112138. https://doi.org/10.1371/journal.pone.0112138

      A loss of function experiment to test for sufficiency would be a nice addition to further confirm their claims, but the authors mention that there were technical limitations to their attempts at inhibiting PVN OXT neurons. I appreciate the authors declaring that the DREADDs attempt suffered from unfortunate confounds. But for optogenetic attempts, I don't think they need a closed-loop system to get some useful results. They still can shine the light at "random" moments (that will correspond to random body temperatures) and then separate the data per body temperature.

      We thank the reviewer for this constructive suggestion. Such an experiment would strengthen our claims and complement the optogenetic activation (Fig. 6). Reviewer 3 brought up a similar concern.

      Building directly on the reviewer’s proposal, we now describe a loss-of-function experiment as an important next step. Optogenetic inhibition of PVNOT neurons can be delivered at pseudo-random times across light and rest phase. Because animals spend extended periods at rest during this phase, a substantial fraction will fall within established rest bouts, which can then be analyzed and stratified by body temperature, as the reviewer notes. The prediction is that silencing PVNOT neurons during rest should prolong the average duration of rest bouts and delay the onset of activity and thermogenesis, relative to matched unstimulated bouts.This provides a direct test of whether PVNOT activity is necessary for the transition from rest to activity. We have revised the Limitations and caveats section to describe this experiment.

      “Third, although we show that PVNOT neurons are sufficient to drive thermogenic and behavioral transitions (Fig. 6), we did not perform acute loss-of-function experiments. Such experiments are warranted because decreases in baseline PVNOT calcium activity were associated with transitions toward the onset of quiescence (Fig. 3I-L), suggesting this system may bidirectionally regulate thermo-behavioural state. A tractable next step would be to optogenetically inhibit PVNOT neurons during established rest bouts, delivered at pseudo-random times across the light and rest phase and analyzed post hoc by behavioral state and body temperature; we predict that silencing during rest would prolong the average duration of rest bouts and delay the onset of activity and thermogenesis. Pairing the inhibition with selective oxytocin antagonist (such as L-368,899), would further test whether the thermogenic and autonomic components of these transitions are oxytocin receptor dependent rather than driven by glutamate released by the same neurons.”

      Lastly, the mention of Raam et al. 2026 is insufficient. The authors just mention it regarding the potential differences with males, to be explored in future experiments. Even if not using males in the current study doesn't affect the stated conclusions, the fact that they chose females because "their thermo-behavioural states were readily discernible" is a considerable bias. Testing males in this very study might be out of scope, but more discussion is warranted.

      We thank the reviewer for this point. We agree that our decision to study females deserves fuller treatment, and we have expanded the Limitations and caveats section accordingly.

      We want to be clear about the rationale, because it was methodological rather than an assumption of sex specificity. Our previous study on behavioral thermoregulation in mice (Landen et al., 2024) showed that, during the light/rest phase, females–but not males–display clearly rhythmic episodes of rest and activity that align with transitions between thermoregulatory states, and are therefore well suited to the analyses that form the core of this study. This choice does constrain the generality of our findings to females, but it does not affect the validity of the conclusions we draw, all of which concern PVNOT neurons in females.

      At the same time, we agree that whether these mechanisms extend to males is a substantive open question and we now say so explicitly. A direct comparison in males, while beyond the scope of the present study, is an important next step, and the recently defined neural basis of collective thermoregulatory huddling (Raam et al. 2026) offers a useful framework for that work. We have modified the Discussion/Limitations and caveats as follows:

      “We focused on females for a practical reason: during the light and rest phase, females show clear, rhythmic bouts of rest and activity, which makes transitions between thermoregulatory states readily discernible and well suited to the analyses around each state transition used here (Landen et al., 2024). This choice constrains the generality of our conclusions, which pertain specifically to females. Because oxytocin signaling can differ between sexes (https://doi.org/10.1016/j.yfrne.2015.04.003), and because the neural control of thermoregulatory behavior may not be identical in males, whether the PVNOT dynamics we describe operate similarly in males remains an open question. Testing males directly was beyond the scope of the present study, but it is an important next step, particularly as the neural basis of collective thermoregulatory huddling has recently begun to be defined (Raam et al. 2026).”

      Reviewer #2 (Public review):

      Summary:

      This is a very interesting study from Vandendoren and colleagues examining the role of PVN oxytocin neurons during thermoregulatory behaviors, in particular during thermoregulatory huddling. The findings are important and have implications for the thermoregulation field as well as the social/naturalistic behavior field. The findings are compelling and use a combination of state-of-the-art tools (photometry, optogenetics, automated behavior tracking, thermal imaging, and core body temperature measurement), often in combination with each other, to produce a rigorous and high-dimensional dataset.

      Comments on revised version.

      I appreciate the effort the authors have put into addressing all of my questions, and I have no remaining concerns.

      Thanks for the comments; they have greatly improved the manuscript.

      Reviewer #3 (Public review):

      Summary:

      This study investigates how the activity of hypothalamic paraventricular oxytocin (PVNOT) neurons relates to physiological states in female mice, with a particular focus on behavioral states and thermogenic sympathetic activity. To address this question, the authors combined automated video-based behavioral classification with calcium imaging of PVNOT neuron activity. Sympathetic thermogenesis was inferred from surface temperature changes measured by infrared thermography, and the authors have made their custom analysis scripts available. The authors report that strong, pulsatile activation of PVNOT neurons was "occasionally" observed immediately before transitions from resting to active states. This observation suggests that PVNOT neuronal activity may facilitate the transition from rest to activity. This phenomenon was observed in both pair-housed and individually housed animals. Taken together, these findings raise the possibility that the oxytocinergic system contributes to naturalistic behavior transitions even in the absence of social interactions. However, concerns regarding the selectivity of GCaMP expression in oxytocin-expressing neurons call into question the validity of the recorded PVNOT neuronal activity.

      Strengths:

      The oxytocinergic neural system is believed to subserve a wide range of physiological functions. Elucidating these roles requires monitoring PVNOT neuronal activity under diverse behavioral contexts, as well as manipulating this activity to establish causal relationships. In this study, the authors present a technically sound experimental framework that integrates behavioral tracking in both individually and group-housed mice with the monitoring and manipulation of PVNOT neuron activity. This setup represents a valuable methodological resource for researchers investigating the physiological functions of oxytocin.

      Thanks for the comments. We are encouraged to hear this framework will open new doors in understanding how the oxytocin system regulates behavior and energy homeostasis.

      Weaknesses:

      (1) Immunohistochemical validation of selective GCaMP expression in oxytocin-expressing neurons showed that only 24-51% of GCaMP-positive neurons expressed oxytocin. As an alternative approach, the authors demonstrate that GCaMP-expressing PVN neurons in virgin females exhibit calcium peaks during rest-wake transitions with kinetics similar to those observed in PVNOT neurons during early lactation. However, this comparison is based solely on population-level peak profiles and does not provide direct evidence for cell-type specificity of GCaMP expression in oxytocin neurons. This limitation substantially undermines the validity of the optical calcium imaging data. In situ hybridization targeting oxytocin mRNA, rather than immunohistochemistry, may provide a more reliable assessment of expression specificity.

      We view our data as showing strong evidence that the recorded neurons include, but may not be limited to, PVNOT neurons for the following two reasons: (1) as the reviewer notes, our longitudinal experiment shows conservation in the physiological and biophysical profile of these neurons in females that went from virgins to parturition and lactation, and (2) as described in Discussion/PVNOT neurons in context of arousal and peptidergic PVN cell-types, non-OT cell-types in the PVN do not show this pulsatile busting profile.

      In the “Discussion/Thermal tracking and validation of PVNOT recording specificity” section we had stated “We note that the animals were perfused at ~ZT4–8, before we were aware that somatic OT immunoreactivity in PVN neurons reaches a daily low during the early light phase [56]”. We now add to this the idea, suggested by the reviewer, that “In situ hybridization targeting oxytocin mRNA, rather than immunohistochemistry, may provide a more reliable assessment of expression specificity.”

      (2) Although the authors' interpretation is generally consistent with the data presented, their main conclusions rely heavily on observational findings. Moreover, optogenetic stimulation of PVNOT neurons failed to robustly recapitulate behavioral state transitions (Figs. 6D and S5B). Further interventional experiments will be necessary to more rigorously test the authors' interpretation and to establish mechanistic insight into the causal relationship between PVNOT activity and rest-to-active transitions. In particular, loss-of-function approaches targeting the PVNOT system, such as OXTR antagonism, inhibitory DREADDs, or cell-type-specific ablation, will be essential to determine whether perturbation of this system alters behavioral state transitions These points should be addressed in future studies.

      Reviewer 1 brought up a similar concern. We have added to the Discussion/Limitations and caveats to address this.

      “Third, although we show that PVNOT neurons are sufficient to drive thermogenic and behavioral transitions (Fig. 6), we did not perform acute loss-of-function experiments. Such experiments are warranted because decreases in baseline PVNOT calcium activity were associated with transitions toward the onset of quiescence (Fig. 3I-L), suggesting this system may bidirectionally regulate thermo-behavioural state. A tractable next step would be to optogenetically inhibit PVNOT neurons during established rest bouts, delivered at pseudo-random times across the light and rest phase and analyzed post hoc by behavioral state and body temperature; we predict that silencing during rest would prolong the average duration of rest bouts and delay the onset of activity and thermogenesis. Pairing the inhibition with selective oxytocin antagonist (such as L-368,899), would further test whether the thermogenic and autonomic components of these transitions are oxytocin receptor dependent rather than driven by glutamate released by the same neurons.”

      Note: as described in the previous response to reviewers, we have tried inhibitory DREADDs in this system and have concluded that it is of little value because delivering DREADD ligand requires handing the animals for an IP injection—a procedure that disrupts sleep/rest and induces stress hyperthermia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors have answered our criticisms and can proceed as they chose. This is an important paper, and it is the author's choice whether to develop their research here or in a subsequent paper.

      Thank you.

      Reviewer #2 (Recommendations for the authors):

      I thank the authors for citing my pre-print, as suggested by Reviewer 1. The paper has now been published and the authors may like to cite the published version (doi.org/10.1038/s41593-026-02224-0).

      Thank you.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors now interpret their results as indicating that PVNOT activity biases the system toward state transition (from rest to active), rather than acting as a deterministic trigger. This interpretation is reasonable. However, the wording "PVNOT peaks (or neurons) predict transitions to behavioral arousal and thermogenesis" may be misleading. If arousal and thermogenesis occur in more than 80% of cases following PVNOT peaks, then such peaks could reasonably be described as "being predicted". Otherwise, the terminology should be revised for clarity.

      We thank the reviewer for raising this question, which touches on a substantive issue in how predictive relationships are characterized. We agree that "predicts" can misleadingly imply a high positive predictive value: i.e., that a large fraction of peaks are followed by transitions.

      This is not the claim we intend, nor is it the appropriate statistical criterion. A variable is predictive when it shifts the conditional probability (or, here, the conditional distribution) of the outcome relative to its base rate — the criterion underlying likelihood ratios, relative risk, and signal-detection measures — rather than when it exceeds an absolute occurrence threshold such as 80%. By this standard, a peak can be informative even if transitions do not follow the majority of peaks, provided transitions are substantially more likely (or thermogenically warmer) when a peak precedes them than when one does not.

      Our data support precisely this. The logistic regression shows peaks are much more probable immediately before rest offset than at other transitions or at baseline, and our new analysis shows that transitions preceded by peaks carry significantly larger post-offset Tb increases than those without. We are not claiming peaks act as a deterministic trigger, and we agree with the reviewer that they are not present before every transition.

      To keep our language aligned with these results, we have revised the wording to avoid "predict" where it could imply high hit-rate determinism, replacing it with comparative phrasing. Accordingly, we have revised the terminology throughout the manuscript: where a claim concerns timing, we now state that peaks “precede” transitions. We have removed “predict”/”predictive” from the section heading, figure legend, introduction and results as follows.

      “Then, we discovered that PVNOT calcium dynamics during huddling were associated with increased likelihood of transitions to body warming and arousal.”

      “PVNOT neuronal activity precedes transitions towards thermogenesis and behavioral arousal in social and non-social contexts.”

      Fig. 3 legend title: “PVNOT peaks are associated with increased likelihood of thermogenic rest-to-active transitions.”

      “Thus, PVNOT peaks are at least five-fold more likely to occur near the offset of quiescence/quiescent compared to onset, and signal an increase in physical activity—a correlate of behavioral arousal 53 and a means of increasing metabolic rate and Tb [26]”

      “Thus, for nesting and active huddling, PVNOT peaks are two- to three- fold more likely to occur at bout onset than offset.” Dropping flagged word here lol.

      “Together these results suggest that elevated PVNOT activity dynamics precede the offset of two rest states (quiescence and quiescent huddling) by approximately 100 seconds, and the onset of two post-quiescence active states (nesting and active huddling) by around 20 seconds, in solo and paired mice respectively.”

      “Moreover, PVNOT peaks aligned with the low point of a U-shaped body temperature profile: on average, Tb decreased before, and increased after, the time of the calcium peak in both solo and paired conditions (Fig. 3O,R). Together, these results suggest that PVN<sup>OT</sup> peaks occur during a low Tb trough and mark a subsequent rise in Tb.”

      (2) Regarding the 400-sec latency of BAT surface temperature increases following optogenetic stimulation, the authors now attribute this delay to slow peptidergic transmission. However, the authors should consider prior findings showing that BAT temperature increased immediately following optogenetic stimulation of PVN→rMR oxytocin neurons in anesthetized rats (Fukushima et al., 2022).

      My hunch is that doing this in anesthetized rats gives a stronger signal to noise… not sure if I can back that up though.

      At the least we can add a sentence that says “rMR oxytocin neurons immediately increases BAT temperature, while infusion of OXT or NMDA in the rMR results in BAT temperature increases after approximately one minute…” (see Fig. 3,4,5).

      We thank the reviewer for redirecting us to Fukushima et al. (2022). We note, however, that in that study the fast-responding variable was BAT sympathetic nerve activity, whereas the BAT temperature itself rose over several minutes following both optogenetic stimulation (their Fig. 4F, quantified at 5 and 10 minutes) and focal rMR infusion of oxytocin or NDMA (their Fig. 5, multiminute traces). This thermal timescale is comparable to the one we observe.

      The remaining difference could reflect methodological differences: we stimulated PVNOT somata rather than rMR terminals, measured intrascapular surface rather than BAT temperature directly, and recorded in awake, freely behaving animals (rather than anesthetized animals) in which competing thermoeffector and behavioral processes are active. Consistent with a methodological basis for the delay, focal infusion of oxytocin or NMDA into the rMR in that study increased BAT temperature over roughly a minute (Fukushima et al., 2022). Slow, diffuse peptidergic neuromodulation may further contribute, oxytocin is released from large dense-core vesicles and can act over extended time scales (Ludwig and Leng, 2006; Parmaksiz and Kim, 2025; Qian et al., 2023), although our data cannot isolate this mechanism from the factors above or from fast glutamatergic co-transmission that likely accompanies PVNOT activation (Hrabovszky and Liposits, 2008).

      (3) In the previous review, clarification was requested regarding the rationale and histological basis for intravenous FluoroGold injection. While the authors have now added methodological details, they should also incorporate the following explanatory text (previously provided in their rebuttal) into the manuscript for readers unfamiliar with PVN histological analyses:

      "Intravenous injection of FluoroGold (FG) was used to histologically differentiate between magnocellular and parvicellular oxytocin neurons in the PVN. Because the posterior pituitary is located outside the blood-brain barrier, i.v. FG is selectively taken up by terminals of magnocellular neurons and retrogradely transported to their cell bodies. This allows us to infer the neuroanatomical identity (magno- vs. parvicellular) of the PVNOT neurons of interest."

      We thank the reviewer for this suggestion. We have added the explanatory text to the results subsection, “PVN<sup>OT</sup> cellular projections to the rMR”. The text now reads: “rMR cell types in mice, we used FluoroGold (FG to disambiguate magno- vs. parvocellular PVN<sup>OT</sup> projections [67] (Fig. S6A-C). Because the posterior pituitary is located outside the blood-brain barrier, intravenous FG is selectively taken up by terminals of magnocellular neurons and retrogradely transported to their cell bodies. This allows us to infer the neuroanatomical identity (magno- vs. parvicellular) of the PVN<sup>OT</sup> neurons of interest.”

    1. Author response:

      Reviewer #1 (Public review):

      The manuscript from Zhu et al. identifies microbial riboflavin-derived MR1 ligands as potent pharmacological activators of human MAIT cells and provides evidence that MR1 ligand stimulation can enhance MAIT-mediated tumor killing across multiple solid tumor models. The study is conceptually interesting and supported by a broad combination of human primary samples, tumor cell lines, 3D models, SC transcriptomics, and xenograft experiments. Overall, the data largely support the central conclusion that MR1 ligand stimulation can strongly activate human MAIT cells and enhance anti-tumor cytotoxicity. However, the broader conclusions concerning endogenous MAIT mobilization, tumor specificity, and translational potential are not yet fully supported by the current data and should either be moderated or addressed with additional experiments.

      We thank the reviewer for the positive feedback. We will address all comments and suggestions point by point.

      Comments:

      (1) The authors use one-way ANOVA throughout the manuscript, but this may not be appropriate for some analyses, particularly when multiple experimental factors are present and their interaction effects need to be considered. For example, Figure 3f appears to involve multiple factors, for which a two-way ANOVA may be more appropriate. Similar issues may apply to other panels.

      We thank the reviewer for this valuable comment. We will carefully review the statistical analyses and revise the tests as appropriate, including the use of two-way ANOVA where multiple experimental factors are present.

      (2) In Figure 3f, the authors show data from patients #1 and #2 and state that the experiment is representative of three experiments. What does the reported "n=4" represent in this figure?

      We thank the reviewer for this valuable comment. We will revise the figure legend to clearly define what the reported n = 4 represents.

      (3) There appears to be a discrepancy between Figure 3f and Supplementary Figure 3b. The two panels appear to use the same treatment conditions and the same label, and both appear to use patient #1 samples, yet the reported values are different. Please clarify the experimental design and explain the reason for this discrepancy.

      In addition, the gating strategy used to define live tumor cells should be clearly described in the figure legend and/or Methods. The authors define "live tumor cells" as MR1/5-OP-RU tetramer-CD45- cells. However, in primary liver tumor samples, the CD45-/tetramer- population may contain other non-hematopoietic cells, such as fibroblasts, and therefore may not exclusively represent tumor cells. The authors should clarify whether additional tumor-specific markers or other criteria were used. The gating strategies for the relevant flow cytometry experiments should be provided in the Supplementary figures.

      We thank the reviewer for this valuable comment. Figure 3f (patient #2) and Supplementary Figure 3b (patient #1) were generated using samples from different patients. We will clarify this in the revised manuscript and provide the relevant gating strategies in the Supplementary Information.

      (4) I have some concerns regarding the claims of "selective activation of anti-tumor inflammatory pathways rather than generalized cytokine release" and "avoiding induction of tumor-supportive mediators." The authors show that MAIT cells stimulated with 5-OP-RU can substantially reduce tumor cell viability. Therefore, the cellular composition of the co-culture is likely to change considerably during the assay, which may affect the absolute levels of cytokines and other soluble mediators detected. For example, reduced tumor cell numbers could lead to lower production of tumor-derived factors such as VEGF, potentially confounding the interpretation that these mediators are not induced by MAIT activation. The authors should consider whether cytokine measurements have been normalized to viable cell numbers or otherwise account for differences in tumor cell abundance.

      We thank the reviewer for this important comment. We agree that differences in tumor cell abundance may affect cytokine measurements. We will moderate our claims accordingly and acknowledge this limitation in the revised manuscript.

      (5) The in vivo tumor models may show substantial variability between independent experiments. Rather than presenting a single representative experiment, the authors should consider showing pooled data from all independent experiments, with the total number of mice clearly indicated.

      We thank the reviewer for this valuable comment. We will provide pooled data from all independent in vivo experiments and clearly indicate the total number of mice.

      (6) Why did the authors use an MR1-overexpressing tumor cell line for the in vivo studies rather than the parental cells with endogenous MR1 expression, together with MR1-KO cells as a negative control? The authors demonstrate that MR1 is detectable across multiple tumor cell lines and that endogenous MR1 expression is sufficient to support MAIT-mediated killing in vitro. Moreover, MR1 overexpression substantially enhances tumor cell susceptibility to MAIT-mediated killing. Therefore, it is unclear whether the strong therapeutic efficacy observed in vivo reflects physiologically relevant MR1 expression or is driven by artificially elevated MR1 expression. An in vivo comparison using parental and MR1-KO tumor cells would substantially strengthen the translational relevance and establish whether the therapeutic effect can be achieved at endogenous levels of MR1.

      We thank the reviewer for this important comment. We agree that comparison with endogenous MR1 expression would strengthen the translational relevance of our findings. We will include new in vivo experiment comparing parental tumor cells.

      (7) How is tumor specificity of MAIT achieved? The authors propose that MAIT-cell activation by MR1 ligands provides an antigen-independent approach for tumor targeting. However, MR1 is broadly expressed and is not tumor specific. While the relative sparing of T and B cells in Figure 7B provides some evidence of cell-type selectivity, this does not establish tumor versus normal tissue specificity. It remains unclear whether activated MAIT cells can discriminate tumor cells from other normal MR1-expressing cells and tissues. This raises an important question regarding the potential systemic toxicity of MAIT cells activated by systemic administration of 5-OP-RU. In particular, could other MR1-expressing cells be targeted when a large number of MAIT cells are simultaneously activated? The authors should consider assessing systemic toxicity in vivo, for example by examining serum ALT/AST levels and tissue pathology, and/or by evaluating the effects of MAIT + 5-OP-RU in tumor-free animals. At least, the potential specificity and safety limitations of systemic MR1 agonism should be discussed.

      We thank the reviewer for this important comment. To further evaluate the potential safety concerns associated with systemic MR1 ligand stimulation, we will include a new experiment assessing the effects of MAIT cells plus 5-OP-RU in tumor-free animals. We will also discuss the potential specificity and safety limitations of systemic MR1 agonism in the revised manuscript.

      Reviewer #2 (Public review):

      The manuscript by Zhu et al. describes MAIT cell activation by riboflavin metabolites presented by MR1. The authors provide solid evidence for this activation and anti-cancer functional consequence using an array of selected cell lines, primary ex vivo and engineered xenograft models. Broadly, the results are thorough and well controlled, and provide a highly informative insight into the metabolite-MAIT-cancer cell interactions. However, the majority of this work is undertaken using models that preferentially express key targets, and whilst still useful, the (current) broader implications of this research are overstated. Additionally, the suggested MAIT modulation of the tumor microenvironment requires clarification.

      We thank the reviewer for the positive feedback. We will address all comments and suggestions point by point.

      Major Comments:

      (1) In Figures 2b-d, the authors suggest microbial metabolite stimulation of PBMC cultures increased MAIT cell frequency up to 60%. Whilst their flow data is compelling, the frequency of one population can be influenced by changes in other populations. A form of absolute or relative-to-total count should be used.

      We thank the reviewer for this valuable comment. We will provide absolute cell counts and/or normalized data to more accurately assess changes in MAIT cell frequency.

      (2) The statements regarding cytokine induction in Figure 4e are too strong; many of those inflammatory cytokines are not automatically and consistently tumour-suppressive. The line 299 '...were not induced' may just reflect death of tumor cells. It would be useful to include tumour cell-only controls in Figure 4.

      We thank the reviewer for this valuable comment. We agree that the statements regarding cytokine induction should be interpreted more cautiously. We will revise the relevant claims.

      (3) Figure 7 is interesting, but the authors' conclusion that MAIT+5-OP-RU controls the tumor microenvironment is not robustly supported by their evidence.

      (a) It is not clear how CD14+ cells established a sustained suppressive environment.

      We thank the reviewer for this valuable comment. We will include additional experiments to further characterize the contribution of CD14+ cells to the observed suppressive environment.

      (b) It is not clear how the peritoneal addition of microbial metabolites 'significantly enhanced MAIT-mediated tumor control'. The authors show that the addition of 5-OP-RU reduced the number of GFP-expressing tumour cells present in peritoneal lavage fluid. There is limited evidence to suggest this occurs through MAIT cells or MR1 in this figure.

      We thank the reviewer for this valuable comment. We will include additional T-cell and T-cell + 5-OP-RU control groups to further determine the contribution of MAIT cells to the observed tumor control.

      (c) It is difficult to draw conclusions from peritoneal lavage flow when some experimental groups received cells IP, but then all groups were equally assessed for key populations, and all data are presented as frequencies. The authors should use absolute counts (or similar) to appropriately show changes in cell populations to account for varying total/live/cd45+ cell compartments.

      We thank the reviewer for this valuable comment. We will provide absolute cell counts, in addition to frequencies, to account for differences in total and viable CD45+ cell numbers.

      (d) It would be necessary at a minimum to include 5-OP-RU-only controls, and ideally include MR1 blocking or the cancer line with MR1 removed. Alongside this, the authors should substantially reduce the strength of their statements on microbial metabolite-MAIT suppression of the tumor microenvironment.

      We thank the reviewer for this valuable comment. We will include additional T-cell and T-cell + 5-OP-RU control groups and will substantially moderate our statements regarding microbial metabolite-mediated modulation of the tumor microenvironment.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This Review Article provides a thorough overview of whole-brain activity changes induced by brain stimulation and summarizes the current state of the field. However, it lacks integration across spatial and mechanistic scales, which limits the reader's ability to understand how the different findings relate to one another. In addition, several key concepts are not explained in sufficient depth for non-expert readers. The manuscript would benefit from the development of a cohesive conceptual framework to more clearly synthesize the existing literature.

      Thank you for the positive assessment. We fully agree, and as suggested we have added a new conclusion paragraph that outlines a synthesis of the paper and suggests a conceptual framework :

      “In this paper, we have reviewed aspects of neuronal responsiveness, from the microscale level of neurons and circuits, the mesoscale level of single brain areas, and the macroscale level of the whole brain. At the microscale, it is apparent that the circuit operating in an asynchronous mode displays the highest responsiveness, as seen in brain slices (D’Andola et al., 2018). The underlying mechanism is that the high levels of synaptic « noise » in asynchronous states set neurons in a high responsive mode, as seen in models of single neurons (Ho & Destexhe, 2000). This higher responsiveness is confirmed at mesoscale, and can be seen for example with Utah-array recordings comparing wake and anesthesia (Dwarakanath et al., 2025). Similarly, propagating waves occur in the asynchronous state in awake monkey (Muller et al., 2014), and sensory inputs evoke more propagating patterns (and higher PCI) in wakefulness with asynchronous states compared to slow-wave states of anesthesia in mice (Montagni et al., 2024). At the whole-brain scale, experiments also find that evoked responses are more complex and propagating compared to slow-wave states (Massimini et al., 2005), a situation which models can reproduce (Goldman et al., 2023; Sacha et al., 2025). Other measures, such as fluidity (Breyton et al., 2024) and reversibility (Camassa et al 2024) also point to the same conclusion. Collectively, these results show that asynchronous and irregular activity states set neurons in a high responsive mode, which in turn impacts network behavior and favors the propagation of activity as mesoscale propagating waves, or macroscale activity patterns that propagate across brain regions. It is therefore not surprising that the best correlate of conscious states is the asynchronous activity (Koch et al., 2016).”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper is a comprehensive review of perturbation studies and the state-dependence of the brain's response to perturbation at the circuit, mesoscale, and macroscale levels.

      Strengths:

      The strengths of the paper are the thorough description of many perturbation studies at different levels of organization, and the integration of both experimental and modeling studies. The review clearly communicates the need to consider (1) brain or local-population state, and (2) multiple levels of organization, in order to understand perturbation responses. Another major strength is the ability for the reader to reproduce figures using the EBRAINS platform.

      Weaknesses:

      Two major points of improvement should be resolved with the review, in order to make it useful for a broad audience.

      The first is that the review does not include a significant integration across scales, and as a result, reads like three separate (though comprehensive) reviews. Currently, the only integration across the scales is in the brief conclusion paragraph. I would recommend adding an additional section, in which the overarching picture is discussed. (i.e. a unifying view of state dependence, and what is learned by considering across scales). This need not be too long, but it should be longer than a single conclusion paragraph.

      Thank you for the positive assessment. We fully agree with the excellent suggestion of adding a concluding paragraph where the conceptual framework and overarching picture are presented. Please see the new conclusion paragraph that we added to the paper (copied above in the reply to Editors).

      The second major weakness is that there is a lack of clarity on many points throughout, which is needed for the reader to fully understand the results described.

      See our answer to the specific comments below for the list of unclear points.

      Reviewer #2 (Public review):

      Summary:

      In this review article, the authors discuss the whole-brain activity changes induced by brain stimulation. They review the literature on how these activity changes depend on the cognitive state of the brain and divide the results by the scale of the change being induced, from microscale changes across small groups of neurons, up to macroscale changes across the entire brain. Finally, they describe attempts to model these changes using computational models.

      Strengths:

      The review provides an overview of the results within this subfield of neuroscience, and the authors are able to discuss a lot of prior results. The framing of the changes in neuronal activity in terms of computational changes is also a helpful approach.

      Weaknesses:

      However, the authors are not able to contextualize these results within a single framework, i.e. explaining from first principles how different aspects of stimulus-induced changes interact to generate functional changes in the brain, and how different changes - at distinct spatiotemporal scales - combine to form larger effects. This is a significant weakness in generating a review of the literature, since the authors do not provide a cohesive conceptual framework on which to frame the results. Similarly, the authors do not explain how their different computational models fit together, and how one can get a singular computational understanding of the distinct mechanisms of brain activity changes due to stimulation under different brain states, by combining the results derived from each separate model.

      Thank you for the positive assessment. This is an excellent suggestion, actually also requested by Reviewer 1. We have now added a new conclusion paragraph where the conceptual framework is explained (copied above in the reply to Editors).

      Major Comments:

      (1) The authors have written this review as if it were intended for an audience who is already familiar with the topics. For example, they introduce concepts like complexity, spiral vs planar waves, without much explanation.

      Thank you for this helpful comment. We agree that the Introduction should be more accessible to readers who are less familiar with these concepts, and we have therefore revised the text to provide a clearer definition of complexity and a more explicit explanation of propagating wave patterns.

      Specifically, we now clarify that complexity can be understood as the richness of the set of accessible states of a system, which in our context can be related to the diversity of slowwave propagation modes and to the high-entropy, desynchronized activity of the awake brain. We also expanded the description of propagating slow waves to distinguish planar from spiral waves and to explain how their relative prevalence changes with anesthesia depth.

      In the Introduction we replaced the sentence “New methods … at various scales” with “New methods for characterizing the complexity of network dynamics and their response patterns have emerged, particularly recently (Krohn et al., 2023; Wolf et al., 2018), and are presented here at various scales. Here, complexity is associated with the set of accessible states of a system (Parisi, 2006). In the present context, this notion can be linked to the diversity of slow-wave propagation modes and to the richness (i.e., the entropy) of perturbation-evoked responses in brain activity.”

      While in Results (p. 13) the sentence “Spontaneous slow waves … administered (Huang et al., 2010).” has been expanded in “Spontaneous slow waves can also display propagating patterns, as shown in anesthetized mice (Huang et al., 2010; Mohajerani et al., 2010; Pazienti et al., 2022; Stroh et al., 2013). These patterns may take the form of planar waves, which travel across the cortex along a relatively regular front, or spiral waves, which rotate around a central core and therefore produce a more complex spatiotemporal organization. Under relatively deep anesthesia, spiral waves occur more frequently than planar waves, whereas the opposite imbalance is observed as anesthesia is lightened (Huang et al., 2010).”

      (2) Regarding complexity, the authors present a quantification termed PCI. However, in the associated box, they state that PCI could be implemented in a number of different ways, using analogous metrics (which are, nonetheless, not identical). Yet the authors simply claim that all these metrics are sufficiently similar to be grouped together as "PCI". The authors do not provide much intuition about this, and they also don't present any other potential quantifications. This makes any interpretation of their results strongly dependent on your understanding of the concept of PCI. It would be helpful to present some other, analogous metric to demonstrate that the results that the authors are focusing on are not somehow tied to the specific computational structure of the PCI metric.

      Thank you for pointing out to this inconsistency. We agree that the rationale for focusing on perturbational complexity was not sufficiently introduced in the original version of the manuscript.

      Broadly speaking, complexity measures used in consciousness research can be divided into two major classes. The first includes observational measures, which are computed from spontaneous ongoing activity and quantify statistical dependencies within neural time series. The second includes perturbational measures, which quantify the deterministic causal interactions revealed by a controlled perturbation of the system and their spatiotemporal propagation across the network (see Sarasso et al., 2021).

      The primary aim of the present Review was to discuss how complexity changes across spatial and temporal scales in response to perturbations. For this reason, we focused on perturbational complexity measures and, in particular, on the Perturbational Complexity Index (PCI), which remains one of the most widely adopted and validated approaches in this category.

      As described in Box 2, different implementations of PCI have been proposed. The two most established versions are PCI based on Lempel–Ziv complexity (PCI^LZ) and PCI based on state transitions in principal component space (PCI^ST). Although these implementations differ algorithmically, they were developed to operationalize the same theoretical construct and have been shown to correlate strongly when applied to the same datasets (Comolatti et al., 2019). For this reason, throughout the Review we use the term “PCI” as an umbrella label encompassing these related perturbational complexity measures.

      Importantly, all complexity measures discussed in the studies reviewed here belong to this broader class of perturbational approaches. While adaptations of the original algorithms are often required when dealing with different recording modalities and spatial scales, these modifications mainly concern preprocessing and signal representation rather than the underlying theoretical construct being quantified.

      To clarify this point, we have revised the Introduction to explicitly motivate our focus on perturbational complexity, to distinguish perturbational from observational complexity measures, and to explain why different PCI implementations can be discussed within a common conceptual framework. We believe that these additions make the rationale of the Review substantially clearer and reduce the impression that the conclusions depend on a specific implementation of PCI.

      (3) The authors divide the review into sections organized by the spatial extent of the effects that they are exploring (e.g. from microscale to macroscale). However, they don't bring together these insights into a cohesive structure - for example, by providing potential explanations of the macroscale effects by using the microscale changes.

      We agree – and this is now the focus of the newly-added conceptual-framework conclusion paragraph.

      (4) The authors completely ignore any aspect of cell-type specificity in their review, despite the known importance of specific cell types at the microcircuit scale. This makes it difficult to map their results onto the true biological system.

      We agree that cell-type specificity could be made more explicit. The revised manuscript now clarifies that several models already include cell-type specificity. For example, the AdEx-based models distinguish excitatory regular-spiking or pyramidal populations with adaptation from inhibitory fast-spiking populations without adaptation. This differentiation is not made with other models like leaky or quadratic integrate-and-fire models. At the mesoscale, mean-field approaches can be derived for different structures, such as cortex, thalamus, hippocampus, striatum, or cerebellum, and can incorporate the experimentally experimentally observed firing properties of relevant cell classes.

      (5) The authors introduce several different computational models, such as the Hopf model, the AdEx model, and the MPR model. However, they do not provide the reader with a conceptual understanding of the structure of each of these models (except through potentially more complex terminology, e.g. the Hopf model is a "phenomenological StuartLandau nonlinear oscillator"). Additionally, though they present the results of each simulation, they don't provide the reader with intuition about how these models compare against each other, and how best to interpret results derived from each model.

      Very good question, and the answer is not easy. If the goal is to capture large-scale phenomena with models as simple as possible, then Stuart-Landau, Hopf, or Jahnsen-Rit may be appropriate. This approach is rather top-down. But if the goal is to assess how microscopic changes (synaptic receptors for example) affect large-scale brain activity, then we need a bottom-up approach, where mean-field models are derived. We can better explain this.

      We agree that while the technical definitions of the whole-brain models (Hopf, AdEx, MPR) were provided, a clear conceptual framework comparing their underlying structures, specific trade-offs, and interpretation guidelines was missing. We have substantially revised the "Macroscale" section on Page 22 to provide immediate intuition regarding what each model represents structurally (e.g., macroscopic phenomenology vs. microscopic biological realism). We emphasized the structural assumptions of each of them, as well as the explicit utility in interpreting brain responsiveness. This ensures readers understand exactly why a researcher would choose one model over another depending on the mechanistic question at hand.

      (6) In several cases, the authors make statements that they appear to believe to be completely straightforward (and require no justification), but that do not appear so to the reader. For example, they mention: "In wakefulness and REM sleep, ..., the membrane potential is depolarized and close to the spike threshold, which explains why neurons respond more reliably and with less response variability compared with slow-wave sleep". However, this statement is not obvious to the reader and requires explanation (for example, in a system that is close to balance, bringing cells closer to the firing threshold can result in increased response jitter).

      We agree that the original statement was an over-simplification of a complex situation. We have revised it to avoid suggesting that depolarization alone monotonically increases reliability. The relevant mechanism is the combination of depolarization, desynchronized high-conductance synaptic input, balanced fluctuations, and reduced tendency to enter long silent Down states. In this regime, weak inputs are more likely to be converted into spikes and propagate through the network. However, too high conductance, excessive noise, can shunt inputs, enhance jitter, or saturate the network. We are now more explanatory.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As stated in the public review, there is a lack of clarity on many points throughout, which is needed for the reader to fully understand the results described.

      Points needing clarification:

      (1) sPCI (slice Perturbational complexity index) - is this different from other PCIs in box 2? Regardless, the metric and its interpretation should be briefly explained in the main text.

      We use sPCI to refer to the PCI measure adapted for application to cortical brain slices (D’Andola et al., 2017; see Box 2). It relies on the same core algorithm as PCI, namely the Lempel–Ziv complexity of the spatiotemporal pattern of significant responses (Casali et al., 2013) but differs in the preprocessing steps required for slice recordings. We now added this clarification to the text.

      (2) Page 9 "by decreasing fast inhibition but also enhancing it" - What does that mean? More info about the model is needed.

      Thanks for raising this point, since this sentence was indeed confusing. We have revised it now.

      (3) Page 9 "balance between segregation and integration, a crucial ingredient on which sPCI relies" - How is this balance seen in the figure? All I see is sPCI and blockage of GABA.

      The comment is correct, and this mention of segregation and integration has now been eliminated.

      (4) Figure 3D, Page 11 "two different desynchronized (AI) states in a network of AdEx neurons"- What are the two different states? Why is the response different?

      The different AI states correspond to different synaptic strength parameters, we added this precision in the text.

      (5) Figure 3B - "Bifurcation diagram showing the different activity regimes displayed by spiking neuron network." Which model? Multiple are cited. This is a general issue throughout where multiple models are mentioned in the text, and it's unclear which is shown in the figure.

      We agree with the Reviewer's helpful remark. We have revised the manuscript to explicitly state the types of models depicted in the different panels of Figure 3. Corresponding details have also been incorporated into the relevant text in the Results section (previously pages 10–12)."

      A few editorial issues:

      (1) The text in many of the figure panels was too small to read. This is a significant issue that must be addressed.

      We will fix this at the next round, can you please let us know which figures are not visible?

      (2) I recommend reading through for writing flow. E.g. In the first paragraph of the introduction, there are two sentences that start with "importantly, ..." in a row.

      Thanks for noting this – this is now fixed.

      (3) Figure 1E - How does the color on the left relate to the right? What is the y-axis?

      The colour code corresponds to the latency of activation (light blue, 0 ms; red, 300 ms). The Y-axes is the global mean field power (voltage). It has now been included in the figure caption.

      (4) Figure 1C - What is the stimulus?

      The triangle corresponds to the electrical stimulation of the homotopic area 18 of the contralateral hemisphere. This information is now included in the figure legend.

    1. Author response:

      Reviewer 1 is concerned that our astrocyte enriched cultures have significant contamination of microglia or other myeloid cells, OPCs (and related cells) and neurons. Further, they assert that purified astrocytes do not express TNF.

      That astrocytes can’t make TNF directly contradicts our previous paper (Heir, et al., JNeurosci, 2024) showing that the TNF driving homeostatic plasticity is generated by astrocytes. The reviewer seems to want to dispute that paper, which is not really the topic of the current paper (which covers the regulation of TNF production, not whether particular cell types make TNF). The Nedergaard group also saw TNF release from human astrocytes (Wang, et al., 2006). The papers cited by the reviewer (all from the same group) rely on RNAseq data, which has limited depth and cannot distinguish if something is not expressed or simply below threshold. Further, as these datasets were generated from astrocytes isolated from brain (which has normal levels of activity), the astrocytic TNF expression would be very low. Plenty of data supports that astrocytes can express TNF when stimulated (by LPS or other activators), and our previous paper shows that this is also true when neuronal activity is blocked (or absent). But at baseline, astrocyte TNF is quite low and likely undetectable as assayed in those papers.

      Here we are using highly purified astrocyte cultures. The reviewer is perhaps unfamiliar with the type of cultures we are using. Given that we use mechanical disruption to remove neurons, followed after 1-2 weeks by shaking to remove microglia, and finally cell passaging, all before experiments, it is surprising that the reviewer thinks there could be neuronal contamination. Neurons cannot survive that procedure, and we do not observe them by morphology or immunostaining, nor see neuronal markers by qPCR.  The microglial contamination is also minimal, as noted in the manuscript, with qPCR for microglia markers is at noise levels (Iba1 Ct value of 34.6), while GFAP shows robust expression (Ct of 16.8; >100,00 fold more than Iba1). But it is possible, if unlikely, that some small number of microglia are making a lot of TNF. However, treating our cultures with the microglia toxin LME (used in Heir, et al., 2024) did not alter our results, further suggesting microglia are not contributing here. Other contaminating cell types (in the OPC lineage, for example) are also possible. However, the majority of cells in our astrocyte-enriched culture are positive for TNF by immunostaining (done while blocking protein export, to prevent any release of TNF). This makes it highly probably that astrocytes are producing TNF (and this production is regulated by g-protein signaling). To verify this, we will show TNF protein in cells co-labeled with astrocyte markers in our upcoming revision of the paper. This will definitively identify astrocytes as producing TNF in these rat cultures. With the human iPSC-derived astrocytes, microglial contamination is not possible (this requires a completely different differentiation protocol). We agree the ALDH1L1 labeling is not as expected, but it is unclear if this is an antibody issue or mis-localized protein. However, the cells also label with S100beta and GFAP, making the astrocyte identity the most likely option by far. We have additional qPCR data showing expression of ALDH1L1 by these cells, in addition to the other astrocyte markers (which will also be added to the revision). The in vivo situation is more complex, and we can’t exclude that astrocyte-DREADD signaling here indirectly alters TNF production in other cells. However, given the direct regulation of astrocyte TNF production in culture, the simplest explanation is that the same is occurring in vivo.

      Reviewer 2 was concerned that the Gs data was indirect and the use of pharmacological approaches with microglia. As for Gs signaling, it is a bit unclear what the reviewer is suggesting as an alternative hypothesis. We activate the Gs-coupled beta-adrenergic receptor to reduce TNF levels and get the same effect by activating adenylyl cyclase, the canonical downstream pathway from Gs-coupled receptors. While it is possible that beta-adrenergic receptors could have alternate coupling or that Gs activation acts on additional pathways, it seems odd to argue that Gs would not be working through adenylyl cyclase activation when activating the cyclase yields the same response. Certainly the most parsimonious explanation is that Gs-couple receptors act through adenylyl cyclase to reduce TNF production.

      As for the use of pharmacology with microglia, this was the more expedient solution to the difficulty of using AAV virus on microglia. Gathering the necessary Cre and conditional DREADD lines was an impractical solution in terms of time and resources. However, the pharmacology of these receptors is well characterized, as is the g-protein coupling. Given that the results are identical to the results from more specific manipulations in astrocytes, it seems reasonable to conclude that there is a common pattern of GPCR regulation of TNF production. The criticism that non-canonical pathways can be activated by these receptors seems equally true for the DREADDs, as these are just GPCRs with mutated binding sites. If anything, the forskolin experiment is the most specific, yet the reviewer dislikes this approach. The overall consistency of the responses, whether due to DREADD activation, native receptors or direct activation of adenylyl cyclase, is the strongest argument.

      This reviewer was also concerned about the limits of in vivo experiments. As noted above, we agree that the in vivo situation is less controlled and indirect effects are possible. However, since the direct action on astrocytes in a defined culture system is identical to what we observe in vivo, the most likely explanation is that the GPCR is having the same effect on TNF production, rather than leading to an unknown secondary signaling which then alters TNF production in microglia (or other cell types).

      The remaining concerns about sample size, statistics, etc will be fully addressed in an upcoming revision. All reported n’s are biological replicates. The iPSCs were generated from 3 distinct unrelated individuals.

    1. Author response:

      We would like to thank the editor and reviewers for their constructive and thoughtful feedback. We appreciate the reviewers' assessment that our work addresses a fundamentally important research question through a novel approach. We are also glad the reviewers found our data to be rigorously analysed, and that they valued our focus on the whole cortex rather than localised regions or electrodes. We are encouraged by the overall assessment of our work and welcome the suggestions for improving the manuscript. Below, we summarise how we plan to address the reviewers' comments in our revision:

      Analyses

      - We will include quantitative results for Fig. 1E that describe the distribution of the optimal alpha values across sensors and participants.

      - For Figures 2 and 3 we will include analyses of individual participants in the Appendix.

      - We will provide more detailed descriptions on how the timing of peak saccade curvature relates to saccade onset and to the identified optimal alpha value.

      Presentation of Methods and Results

      - We will phrase our claims and conclusions more carefully and nuanced throughout, ensuring direct coverage by the data and analyses.

      - We will revise the currently complex sections of the Methods and Results to improve clarity and readability.

      - We will be more explicit about how the data for scene onset were selected.

      Revision of the Discussion

      - We will extend the Discussion section to address possible mechanisms linking the timing of peak saccade curvature and ERF initiation. We will also provide a more thorough discussion of existing and more recent literature on the topic.

      - We will emphasise the main takeaway of the study: the observation that saccade-related processes are more important to the M100 than previously thought, and, reversely, that this component may be less directly related to fixation-locked responses. We will also present our observation of peak saccade curvature as a starting point for future research, as it was not intended as conclusive mechanistic insight into how and why this process relates to early cortical responses.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript seeks to make use of information about Ct values from PCR testing of mosquito pools for West Nile virus infection to make inferences about mosquito prevalence and West Nile risk. It does so through analysis of empirical data and simulated data with a realistic agent-based model.

      Strengths:

      This work is conceptually innovative for mosquito-borne viruses, building on ideas developed primarily during work on SARS-CoV-2. Exploring this topic is worthwhile regardless of the outcome. The use of data, testing in multiple labs, and the complementarity of modeling and empirical data analysis are all strengths of the approach.

      Weaknesses:

      Some of the primary weaknesses include a dependence of the results on relatively narrow model assumptions, and a lack of compelling improvement over existing methods. None of these weaknesses are fatal flaws; they are modest weaknesses that limit the potential of or excitement about the method.

      Thank you for the comment

      Reviewer #2 (Public review):

      Summary:

      The authors extend their previous population-based Ct-value framework for inferring community epidemic trajectories from human infections to vector infections, using mosquitoes as vectors for West Nile virus. They use agent-based modelling to distinguish virus-positive detections arising from non-active infection states from those reflecting active infections, and then apply this framework to mosquito surveillance data from Colorado and Texas.

      Overall, this is a well-designed and carefully evaluated study. The manuscript proposes a feasible and potentially valuable framework for vector infection surveillance. The findings are supported by both mechanistic agent-based simulations and applications to real-world mosquito surveillance data, which strengthens the biological plausibility and practical relevance of the proposed approach.

      Strengths:

      A major strength of the study is its clear methodological extension from human infection surveillance to vector infection surveillance. The agent-based modelling framework provides a useful basis for distinguishing active infections from virus-positive detections that may reflect non-active infection states. The application to surveillance data from two different geographic settings further supports the feasibility of the framework. Overall, the study is carefully designed, and the model schematic and main analyses are generally clear.

      Weaknesses:

      (1) It would be helpful if the authors could provide plots showing variation across locations and over time. This would further support the claim made in the paragraph at lines 101-107.

      Thank you for the comment. Our supplementary Material Figures S5 and S6 already included these visualisations. However, we note that these were not referenced in the manuscript. We have now referenced these within the lines:

      “First, the variation we observe is consistent across five trapping seasons and two states (Figures S5 and S6).”

      (2) Figure 2: The model schematic is clear in terms of workflow, but it would benefit from more information on model parameterization. In particular, it would be helpful to clarify which parameters or migration rates were estimated from the data and which were assumed based on prior literature.

      Thank you for the comment. All the parameters are used from the literature and recorded in the Supplementary Material. However, we have now added a note in the caption of Figure 2, referencing the Supplementary Material as below:

      “Overall structure of the agent-based model (parameters were derived from the literature; see Supplementary Material S2, S3 and S4)”

      (3) Figure 4: I wonder whether the authors examined how changes in the proportion of mosquitoes with static viral-kinetics trajectories would affect the observed bimodal distribution. Relatedly, it would be useful to know whether there is a threshold proportion at which the method becomes less able to distinguish active from static viral-kinetics patterns.

      Thank you for the comment. We have conducted this in analysis and have already included the relevant figures in the Supplementary Material, In particular, Figures S14 (in Section S7) and S22. We have referenced Figure S14 where we discuss the proportion of mosquitoes with static viral-kinetics trajectories that would affect the observed bimodal distribution. However, we had not included a reference to Figure S22, where we illustrate the proportions at which the method becomes less able to distinguish active from static viral-kinetics patterns. We have now included this reference in the same line.

      “We found that the simulated pooled Ct values aligned well with the observed data when the percentage viral load inherited from birds was 100% and the probability of a productive or non-productive infection in the mosquitoes was 0.5, capturing the bimodal distribution of low Ct values (from productively infected mosquitoes) and high Ct values (from non-productively infected mosquitoes) (Figure 4 (B), (C) & (D); Supplementary Material S7.1, Figure S14 for Ct distributions of viral inheritance probability vs. model change probability and Figure S22 for the accuracy and confidence-interval coverage across different productive infection proportions).”

      Conclusion:

      Overall, the evidence is reasonably strong for demonstrating the feasibility and biological plausibility of the proposed framework. Some conclusions would be further strengthened by additional sensitivity analyses on key assumptions, especially the proportion of static viral-kinetics trajectories and spatial-temporal heterogeneity across surveillance sites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for th e authors):

      (1) ll 70-73 - There are a lot of ideas in this sentence. It would be useful to support this with a schematic figure or something like that, which illustrates the conceptual predictions made here. Such a step is necessary given the novelty of what is being explored here.

      This portion of the introduction has been simplified to better introduce the core observations that formed our hypothesis:

      “The substantial variation in viral quantities observed during cross-sectional entomological surveillance suggests more complex vector/virus interactions, and the precedence set by population SARS-CoV-2 testing in humans suggests that Ct value data for WNV in mosquitoes could inform metrics of disease risk to humans. However, there are substantial differences in the epidemiology and biology of WNV infection in mosquitoes with SARS-CoV-2 in humans.”

      We have chosen not to include an additional schematic figure, as the core idea (population viral loads reflect the convolution of infection incidence and within-host viral kinetics) is illustrated in the referenced literature, and the main message of the current manuscript is the explore how this phenomenon is observed in arbovirus vector surveillance.

      (2) ll 134-137 - Mosquito species is another factor that could result in wide variation in Ct values due to differences in vector competence and infection kinetics. Given that 2-4 mosquito species are present in these pools with unknown frequencies, this seems like a potentially major source of unexplained variation.

      Importantly, Culex pipiens and Culex tarsalis mosquitoes are separated prior to testing for WNV. We have clarified in the legend for Figure 1 that Ct values are presented from pools of either Culex pipiens/restuans/salinarus or Culex tarsalis.

      Additionally, we have added the following text to the Materials and Methods section:

      “For identification purposes, Cx. Pipiens species mosquitoes are not separated from the Cx. Salinarius or Cx. Restuans, which are nearly identical morphological. However, Cx. Pipiens is far more abundant than either Cx. Salinarius or Cx. Restuans in Nebraska.”

      We also show in Figure S22, S23, and S24 that we observe similar variation in Ct values across mosquito species, location, and epi week, and thus we do not think that differences between species play a major impact in our findings.

      (3) ll 173-175 - I believe that this is a consequence of the trapping method. Could this please be spelt out a bit more?

      Indeed, all of the data in this manuscript were derived from mosquitoes collected in CDC Light Traps that attract host-seeking mosquitoes (i.e mosquitoes looking for a bloodmeal). We are not considering vertical transmission in our model as it has been reported to occur infrequently in laboratory studies. Therefore, WNV-positive mosquitoes collected in CDC Light Traps have been exposed to WNV through a previous blood meal from a bird. We have clarified the text to include this explanation:

      “The mosquito pool Ct value data in this study come from specimens collected using CDC Light Traps that are baited with CO2, specifically targeting host-seeking mosquitoes. Vertical transmission is not factored into our model, thus, for a WNV-positive mosquito to be captured in the pool, it must have already obtained one blood meal from an infected bird and be seeking its next blood meal, which introduces a delay between infection and being captured.”

      (4) ll 177-178 - Doesn't the temporal trend in Ct values primarily reflect temporal changes in mosquito infection prevalence?

      Thank you for the comment. We agree that the temporal changes in mosquito infection prevalence is the main factor influencing the distribution of the Ct values in pools, as the time-since-infection distribution of trapped mosquitoes does not vary sufficiently to lead to trapping mosquitoes at very different points in their viral kinetics trajectory. Our intended point from this sentence was, given the prevalence and pool size, the variation in the viral load of the infected mosquitoes does not vary in time as all infected mosquito are trapped after they have reached a constant high-viral load level. We have now revised this sentence to reflect this.

      “As the infected mosquitoes progress from increasing viral load to a high set-point viral load, temporal trends in pooled Ct values primarily reflect time-varying infection prevalence and the number of infected mosquitoes in each pool. The remaining non-temporal variation in pooled Ct values reflect individual-level variation in mosquito set-point viral loads.”

      (5) Section 2.2 - It would seem that the assumed viral kinetics in birds would be important to this line of reasoning, given that that determines initial viral load ingested by mosquitoes. I am unclear on what was assumed in the model regarding viral kinetics in birds.

      Thank you for the comment. We have discussed the viral kinetics of the birds in detail in Section 5.3 and Supplementary Material S3. However, we agree that we have not explicitly mentioned this in Section 2.2. Therefore, we have added a reference to these sections in the following paragraph:

      “The model assumes that the mosquito's initial viral load is proportional to the infector bird's viral load (see Section 5.3 and Supplementary Material S3 for further details on the bird viral kinetics model).”

      (6) ll 194-202 - Whilst you have shown that this hypothesis leads to predictions that are consistent with the data, this is a relatively narrow hypothesis, and others are neither discussed nor refuted.

      There are two features of the data which we discuss. First, the substantial variation in Ct values across pools. This is described in detail in Section 2.1. The second observation is the bimodal pattern, which L192-202 refers to. While we agree that we have not modelled alternative hypotheses, our point is that the distribution of pooled Ct values is bimodal, and capturing some mosquitoes with very low viral loads is the most plausible explanation for the mode at high Ct values. However, we contend that this is actually a fairly broad hypothesis, as there are many plausible mechanisms generating mosquito infections with low viral loads, which we already discuss (discussion section beginning “This could be explained by a variety of factors…”). No changes have been made to the manuscript.

      (7) Section 2.3, first paragraph - The problem with this approach is that these simulations depend on a number of assumptions and parameter settings that are not estimated as part of the model fitting process. Thus, the model is very narrow and contingent on these narrow and not compellingly justified assumptions.

      Thank you for the comment. While we agree with the reviewer that this is a potential limitation of our study, we have discussed this in detail in the discussion. As mentioned in the manuscript “the main objective of this study was not to formally fit the multi-scale agent-based model to the data, but rather to understand how individual-level viral kinetics in mosquitoes are reflected in pooled surveillance data, and to demonstrate the use of pooled Ct values in estimating WNV infection prevalence”, we believe the assumptions and model are sufficient to address the research objectives. Furthermore, the fact that simulated Ct value distributions from the ABM can be used directly with the prevalence estimation method to give similar estimates to the existing PooledInfRate package supports the validity of our assumptions, though we agree that this does not necessarily mean all of our assumptions are correct, nor that our model is generalisable to other settings. No changes have been made to the manuscript.

      (8) Section 2.3, second paragraph - So the newly proposed method using Ct values does no better than the existing method using binary data?

      Thank you for the comment. We agree with the reviewer that our method and the existing PooledInfRate package perform similarly at the estimated prevalence levels of WNV. However, the Ct-based method, as we have discussed and shown, is robust at all prevalence levels where the binary-only method fails, and our method can distinguish the productive and non-productive prevalence.. Thus, while the prevalence estimates are similar under both methods for the current dataset, the novelty lies in the ability to reconstruct prevalence using the data in an entirely different way, and the proof-of-concept for how Ct values may harbour more biological information than treating pools as positive/negative. We believe that these points are sufficiently discussed throughout the manuscript. We have not made any changes to the manuscript.

      (9) ll 252-254 - This may only be true because the simulation model and the inference model are identical. If the inference model were misspecified (due, for example, to incorrect assumptions about kinetics, etc), this result would likely weaken.

      Thank you for the comment. The difference in robustness between the binary-only method and Ct-based method is not a feature of the method, but rather of how the data is used. At higher prevalence, all pools are likely to have at least one positive mosquito in them, and thus all pools will be positive, removing all information to discriminate between different prevalence levels. In contrast, the Ct-based method is able to still discriminate between prevalence levels even when all of the pools are positive, as there is still information based on whether the positive pools have low or high Ct values. No changes have been made to the manuscript.

      (10) ll 272-273 - Again, this is highly dependent on built-in model assumptions.

      Thank you for the comment. We agree that the performance of a model can depend on the underlying assumptions and the structure of the model. This is true in general for any model-based inference technique (see White, 1982, for example). Therefore, our simulations, results and interpretations are intended to be evaluated under the model structures and underlying assumptions we have used throughout the manuscript. However, to be explicit, we have now added this line at the end of the paragraph that included the sentence.

      “These results are based on the model structure and the underlying assumptions we used and they may be affected by model misspecification, including incorrect assumptions (see White, 1982, for example).”

      (11) ll 273-281 - Can this be done with pooled data only, or does it require individual mosquito Ct values? The latter would seem to be less practical to obtain in real-world applications.

      Thank you for the comment. As we have cited the related work for SARS-CoV-2, in theory, these methods are applicable when individual Ct values are present. Both pooled data and individual-level data will work, but using pooled data requires the pooling and dilution process to be modelled explicitly. However, as the reviewer mentions, for mosquito surveillance, this is not a practical approach as mosquitoes are always pooled prior to testing to reduce effort and costs. No changes have been made to the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) The supplementary figures do not appear to be ordered according to their first mention in the manuscript, which makes them somewhat harder to follow.

      Thank you for the helpful comment. We have now made sufficient changes to the Supplementary Material and updated the references in the manuscript. Where possible, the supplementary figures and sections are now numbered and presented in the order of their first mention in the manuscript.

      (2) Lines 71-73: This sentence is somewhat vague, and I was not fully clear on the intended message. The authors may wish to revise it for clarity.

      Thank you for the comment. Similar comments have been made by Reviewer #1. We have revised this sentence for clarity.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study investigates low-affinity Ca2+ binding by WT calreticulin and mutant calreticulin associated with type I myeloproliferative neoplasms, as well as the impact on Ca2+ fluxes in suspension cultures of megakaryocyte-like cells in vitro in response to ER Ca2+ ATPase inhibitors that deplete endoplasmic reticulum (ER) Ca2+ store and open plasma membrane Ca2+ channels through STIM1-Orai interactions. The results are important in that they show that Ca2+ binding by calreticulin and store-operated Ca2+ entry are not fundamentally impacted by the type I deletion mutation in calreticulin, which rules out a direct effect of the calreticulin mutation on its own low-affinity Ca2+ binding and any broad impact on ER Ca2+ regulation. The strength of the data and methods used ranges from solid to convincing, although the use of suspension-based flow cytometric assays to investigate ER Ca2+ levels and Ca2+ entry can be challenged. High-affinity Ca2+ binding sites could be further considered, and possible confounding effects of Abl kinase activity in the megakaryocyte-like cell lines could be offset.

      The authors thank the editors and the reviewers for the summary, comments and many helpful suggestions. In the revised manuscript, we have used fluorimetry for more precise ER calcium measurements (new Figure 5), clarified the high-affinity calcium binding site concern based on our previous work, and addressed possible effects of BCR-ABL translocation kinase activity upon cellular calcium signaling using the drug Imatinib (new supplements to Figure 6 and 7).

      Public Reviews:

      Reviewer #1 (Public review):

      The researchers conducted their study using advanced techniques. They found almost no difference in calcium binding between the two proteins and observed no impact on calcium signaling, specifically store-operated calcium entry (SOCE). The study also noted an increase in ER luminal calcium-binding chaperone proteins. Surprisingly, the authors selected flow cytometry as a technique for measurements of ER luminal calcium. Considering the limitations of this approach, it would be better to use alternative approaches.

      Thank you for this suggestion. We have undertaken fluorimetry-based ER calcium measurements, which are shown in a new Figure 5. These also indicate similar ER calcium levels in CRT-KO HEK293T cells, compared to those reconstituted with wild-type CRT and CRT<sub>Del52</sub>.

      This is particularly important as previous reports, using cells from MPN patients, indicate reduced ER luminal calcium and effects on SOCE (Blood, 2020). This issue matters because earlier research with MPN patient cells reported reduced ER luminal calcium levels and altered SOCE (Blood, 2020). How do the authors explain the difference between their results and previous findings about lower ER luminal calcium and changed SOCE in MPN patient cells expressing CRTDel52?

      We thank the reviewer for asking for these clarifications. We have revised the discussion to address some of these points and also clarify the findings of the referenced study (Di Buduo et al., 2020) which did not directly measure ER calcium levels. We also discuss findings from a related study with cultured megakaryocytes from patients that indicated different effects of type I vs type II mutations (Pietra et al., 2016). In the absence of engineered controls, sample-to-sample heterogeneities in primary cells make it difficult to attribute any measured differences as direct effects of CRT mutations. Different from these experiments, by using purified proteins and ITC, our studies show that the Del52 mutant has calcium-binding characteristics resembling that of the wild-type protein. Additionally, through genetic manipulations in cell lines, our studies directly address the effects of calreticulin KO and its Del52 mutation upon ER luminal and cytosolic calcium levels, and cellular SOCE signals. We did not measure significant differences in any of these parameters between the KO cells and those reconstituted with wild-type calreticulin or the Del52 mutant. As noted by the editors, these results show that Ca2+ binding by calreticulin and SOCE in a cell are not fundamentally impacted by the type I deletion mutation.

      Other studies have found that unfolded protein responses are activated in MPN cells with CRTDel52 calreticulin (see Blood, 2021), and increased UPR could account for higher levels of some ER-resident calcium-binding proteins observed here.

      These points are addressed in the discussion. Either protein misfolding in cells with wild-type calreticulin deficiency or the sensing of cellular calcium perturbations could induce the expression of ER calcium-binding proteins in calreticulin-deficient cells, although we favor the latter model for the reason specified in the discussion. Regardless of the precise mechanisms underlying the expression changes in calcium-binding proteins, the upregulated factors are predicted to compensate for calreticulin deficiency and contribute to the maintenance of the overall cellular calcium homeostasis.

      Overall, it remains unclear how this work improves our understanding of MPN or clarifies calreticulin's role in MPN pathophysiology.

      Multiple studies referenced in the manuscript have suggested links between altered calcium signaling/binding by CRT mutants and MPN pathogenesis. Our studies indicate that ER and cytosolic calcium levels and SOCE are not directly impacted by the MPN type I CALR mutation, points noted in the abstract and discussion. Thus, calcium signaling may not play a specific role in MPN CALR mutant pathology via suggested mechanisms. We are confident that readers will find these results important for better understanding the role of calreticulin type I mutations in MPN.

      Reviewer #1 (Recommendations for the authors):

      This study aimed to express, purify, and evaluate low-affinity calcium binding by a calreticulin deletion mutant (CRTdel52) that is linked to myeloproliferative neoplasms (MPN). The researchers performed cell imaging, flow cytometry, and isothermal titration calorimetry to compare calcium binding between wild-type calreticulin and CRTDel52. They assessed cytosolic calcium levels and store-operated calcium entry (SOCE) in HET293T cells (CRT knocked-out background) and in megakaryoblastic MEG-1 cells. Additionally, they examined changes in the abundance of endoplasmic reticulum (ER) resident calcium-binding proteins in cells expressing either wild-type or mutant CRT.

      This study is well executed but lacks clear relevance to MPN, and it is not clear how this work advances our knowledge of calreticulin biology. No differences were found in calcium binding between wild-type calreticulin and CRTDel52, nor was SOCE impacted. They noticed, however, a compensatory increase in the abundance of some ER resident calcium-binding proteins. The lack of any significant changes in calcium behavior between wild type and CRTDel52 is not surprising based on the known amino acid sequence of calreticulin and calreticulin mutant and based on our knowledge about CRT calcium binding in general. Consequently, it is not clear how this work advances our understanding of the pathophysiology of MPN. Calcium may not play a critical role in the MPN pathology; instead, CRTDel52 secretion and receptor signaling appear more central. Further research should address how these findings relate specifically to MPN and cell biology, in general.

      Our current studies demonstrate increased expression of other calcium-binding proteins in the context of heterozygous MPN type I CALR mutations (Figure 8C) or conditions resembling homozygous MPN type I CALR mutations (Figures 8D-8G). These results, together with findings of maintained ER and cytosolic calcium levels and SOCE signals (Figures 4-7 and Figure 6, supplemental Figure 2 and Figure 7, supplemental Figure 1), indicate that altered calcium binding/signaling by Del52 does not directly contribute to MPN pathology.

      What is the biological or pathophysiological relevance of the CRTDel52-KDEL construct?

      The KDEL sequence is important for the ER retention of CRT (Sonnichsen et al., 1994), and its addition was expected to at least partially remedy the ER retention defect of CRT<sub>Del52</sub>. This point is clarified in the revised results section.

      The rise in ER calcium-binding proteins is noteworthy but anticipated, given likely genetic changes from UPR pathway activation in these cells. Is this relevant to MPN?

      We suggest that increased expression of other calcium-binding proteins in the context of heterozygous MPN type I CALR mutations (Figure 8C) or conditions resembling homozygous MPN type I CALR mutations (Figures 8D-8G) would contribute to the maintenance of the cell’s calcium signaling capacity.

      How do the authors explain the difference between their results and previous findings about lower ER luminal calcium and changed SOCE in MPN patient cells expressing CRTDel52?

      The findings related to SOCE are addressed in the points discussed above and in the revised discussion. Related to ER luminal calcium, the study by Ibarra et al. (Ibarra et al., 2022) reported that CRT<sub>Del52</sub> overexpressed in U2OS cells (expressing endogenous CRT) had reduced ER calcium levels compared to the same cells expressing WT CRT or CRT<sub>Ins5</sub>. Those measurements did not use a ratiometric ER calcium probe, and additionally it is possible that the expression of compensatory calcium-binding proteins is more muted in cells expressing endogenous wild-type CRT.

      Other studies have found that unfolded protein responses are activated in MPN cells with CRTDel52 calreticulin (see Blood, 2021), and increased UPR could account for higher levels of some ER-resident calcium-binding proteins observed here.

      We agree that increased UPR could account for higher levels of some ER-resident calcium-binding proteins. As noted in the revised discussion, regardless of the precise mechanisms underlying the expression changes in calcium-binding proteins, the upregulated factors are predicted to compensate for calreticulin deficiency and contribute to the maintenance of the overall cellular calcium homeostasis.

      It is not clear why flow cytometry was a choice of technique for measurements of ER calcium. Pacific Blue was detected at 405 nm excitation and 452-455 nm emission, while unbound probe signals appeared in the AmCyan channel (405 nm excitation/498 nm emission). It is unnecessary to mention fluorochrome labels (like Pacific Blue or AmCyan) for channels that are not in use. Only the channels actually utilized need to be specified: The calcium-bound GEM-CEPIA1er probe's signal was collected using the Pacific Blue channel (excitation at 405 nm, emission at 452-455 nm), while the unbound probe's signal was detected using the AmCyan channel (excitation at 405 nm, emission at 498 nm). Since it's not possible to monitor emission only at a specific wavelength with a Fortessa, could this be a different channel that is being recorded?

      Additionally, due to the similar spectra and potential for bleed-through between Pacific blue and Amcyan, compensation is likely necessary. Therefore, you should include single-stained control cells containing only one probe for proper reporting. Additionally, only the "GEM-CEPIA1er probe" is displayed, while the second probe is referred to solely as "unbound".

      A single genetically encoded GEM-CEPIA1er probe (Suzuki et al., 2014) was used for measuring both the bound and unbound signals. The GEM-CEPIA1er probe was excited with the 405 nm violet laser. In the methods section of the revised manuscript, the GEM-CEPIA1er probe wording is included for describing both the bound and unbound signal collections. Additionally, we have undertaken new spectrofluorimetric experiments (new Figure 5), which allow for the distinct emission peaks to be recorded corresponding to the Ca<sup>2+</sup>-bound and Ca<sup>2+</sup>-unbound signals. Similar results were obtained as reported for the flow cytometry-based experiments.

      Increased expression of wild-type calreticulin compared to parental cells should impact on ER calcium content and dynamics in back-transfected HEK293T-KO or MEG-1 cells. Direct ER calcium measurements in HEK293 cells with various calreticulin constructs would significantly strengthen this presentation.

      Our experiments were structured to compare calcium signaling in cells expressing only wild-type CRT or CRT<sub>Del52</sub> (resembling homozygous type I MPN CALR mutations) compared to CRT-KO cells. The over-expression of CRT in the reconstituted cells compared to endogenous expression level is a limitation of our study which we have acknowledged in the revised manuscript discussion. Understanding the effects of over-expression of wild-type CRT vs the CRT<sub>Del52</sub> mutant upon ER and cytosolic calcium signals and SOCE is interesting, but beyond the scope of the present study.

      The authors should examine the immunolocalization of CRTDel52 and wild-type protein in HEK293 cells.

      Previous published studies from another lab showed that CRT<sub>Del52</sub> is secreted from HEK cells and that the addition of a KDEL sequence to CRT<sub>Del52</sub> reduces secretion and induces its increased intracellular accumulation (Arshad and Cresswell, 2018). This point is noted in the revised results section and the reference is cited. This appears to be the general theme in primary cells and cell lines. Previous studies and our own prior published studies have shown that CRT<sub>Del52</sub> (but not wild-type CRT) is detectable in the media of cell lines and patient serum as well on the cell surface of primary cells and cell lines (Kaur et al., 2024, Venkatesan et al., 2021, Pecquet et al., 2023).

      SDS-PAGE of purified proteins is overloaded, and chromatograms show extra peaks or shoulders, making protein quality assessment uncertain.

      Representative chromatograms, peaks corresponding to protein monomers used for ITC analyses and the relevant gels are clarified in the revised manuscript. In new analyses since the original submission, intact protein mass spectrometry was undertaken for CRT<sub>Del52</sub>. The results indicate a 35-42 amino acid truncation in different preparations. The truncated proteins would still include acidic residues (between 340–351) previously implicated in low-affinity calcium binding by murine CRT that are shared between wild-type and CRT<sub>Del52</sub>. This new information is now included in the revised results section.

      Analysis of SOCE in calreticulin-deficient cells and cells reconstituted with calreticulin or overexpressing the protein has already been reported (PMID12324449).

      The indicated reference and additional related papers examining effects of CRT deficiency and overexpression on cellular calcium signaling (Arnaudeau et al., 2002, Bastianutto et al., 1995, Mery et al., 1996, Nakamura et al., 2001) are cited in the revised manuscript.

      Reviewer #2 (Public review):

      Tagoe and colleagues present a thorough analysis of the calcium (Ca2+) binding capacity of calreticulin (CRT), an endoplasmic reticulum (ER) Ca2+-buffer protein, using a mutant version (CRT del52) found in myeloproliferative neoplasms (MPNs). The authors use purified human CRT protein variants, CRT-KO cell lines, and an MPN cell line to elucidate the differing Ca2+ dynamics, both on the level of the protein and on cell-wide Ca2+-governed processes. In sum, the authors provide new insights into CRT that can be applied to both normal and malignant cell biology.

      First, the authors purify CRT protein and perform isothermal titration calorimetry to quantify the Ca2+ binding capacity of CRT. They use full-length human CRT, CRT del52, and two truncations of CRT (1-339 and 1-351, the former of which should lead to the entire loss of low-affinity Ca2+ binding). While CRT del52 has previously been shown to lead to a decrease in Ca2+ binding affinity in other models, the ITC data show that this is retained in CRT del52.

      Next, the authors utilize a CRT-KO cell line with subsequent addition of CRT protein variants to validate these findings with flow cytometric analysis. Cells were transfected with a ratiometric ER Ca2+ probe, and fluorescence indicates that CRT del52 is unable to restore basal ER Ca2+ levels to the same extent as CRT wild-type. To translate these findings to MPNs, the authors perform CRT-KO in a megakaryocytic cell line, where reconstitution with either CRT variant did not cause a difference in cytosolic calcium levels. The authors further test store-operated calcium entry (SOCE), an important process for maintaining ER Ca2+ levels, in these cells, and find that CRT-KO cells have lower SOCE activity, and that this can be slightly recovered with CRT addition.

      Finally, the authors ask whether other effects of CRT-KO/reconstitution can affect the cellular Ca2+ signaling pathway and levels. RNASeq analysis revealed that CRT-KO leads to an increase in various chaperone protein expressions, and that reconstitution with CRT del52 is unable to reduce expression to the same extent as reconstitution with CRT wildtype.

      Strengths:

      The authors provide new insights into CRT that can be applied to both normal and malignant cell biology.

      We thank the reviewer for the recognition that this study is important for our understanding of both normal and malignant cell biology.

      Weaknesses:

      (1) The authors should consider discussing the high-affinity Ca2+ binding site more in the introduction. Can they show a proof-of-concept experiment that validates that incubation of recombinant CRT reduces the function of that high-affinity Ca2+ binding site?

      In a previous study (Wijeyesakere et al., 2011), we showed that at a starting calcium concentration of 0 mM and with CaCl<sub>2</sub> injections to a final concentration of 70-80 mM the measured K<sub>D</sub> value was 16.6 mM for calcium binding to wild type murine calreticulin, (which has ~95% sequence identity with human calreticulin), corresponding to the high-affinity site. On the other hand, at a starting calcium concentration of 50-100 mM and CaCl<sub>2</sub> injections to a final concentration of 700-850 mM, the measured K<sub>D</sub> value for calcium binding to wild-type murine calreticulin was 590 mM (corresponding to the low-affinity sites). We did not observe the high-affinity sites when the starting calcium concentration was 50 mM and calcium injections were at 33 mM each; similar conditions are used in the present study. These points are clarified in the revised manuscript in the results section.

      (2) For Figure 2B, do you have an explanation for why the purified proteins run higher than predicted (48-52kDa) - are these proteins still tagged with pGB1?

      Yes, the purified proteins shown in Figure 2B retained a GB1 tag. This point is clarified in the revised methods.

      (3) The MEG-01 cell line has the BCR:ABL1 translocation, while CRT mutations are strictly found in BCR:ABL1 negative MPNs. Could these experiments be repeated in these cells treated with imatinib to decrease these effects, or see if basal MEG-01 Ca2+ levels/activity are changed with or without imatinib?

      Thank you for this important point. We have assessed cytosolic calcium levels in MEG-01 cells that were treated or not treated with imatinib in new Figure 6, supplemental Figures 1 and 2 and Figure 7, supplemental Figure 1) and show that the prior results hold in imatinib-treated cells.

      References

      ARNAUDEAU, S., FRIEDEN, M., NAKAMURA, K., CASTELBOU, C., MICHALAK, M. & DEMAUREX, N. 2002. Calreticulin differentially modulates calcium uptake and release in the endoplasmic reticulum and mitochondria. J Biol Chem, 277, 46696-705.

      ARSHAD, N. & CRESSWELL, P. 2018. Tumor-associated calreticulin variants functionally compromise the peptide loading complex and impair its recruitment of MHC-I. J Biol Chem, 293, 9555-9569.

      BASTIANUTTO, C., CLEMENTI, E., CODAZZI, F., PODINI, P., DE GIORGI, F., RIZZUTO, R., MELDOLESI, J. & POZZAN, T. 1995. Overexpression of calreticulin increases the Ca2+ capacity of rapidly exchanging Ca2+ stores and reveals aspects of their lumenal microenvironment and function. J Cell Biol, 130, 847-55.

      DI BUDUO, C. A., ABBONANTE, V., MARTY, C., MOCCIA, F., RUMI, E., PIETRA, D., SOPRANO, P. M., LIM, D., CATTANEO, D., IURLO, A., GIANELLI, U., BAROSI, G., ROSTI, V., PLO, I., CAZZOLA, M. & BALDUINI, A. 2020. Defective interaction of mutant calreticulin and SOCE in megakaryocytes from patients with myeloproliferative neoplasms. Blood, 135, 133-144.

      IBARRA, J., ELBANNA, Y. A., KURYLOWICZ, K., CIBODDO, M., GREENBAUM, H. S., ARELLANO, N. S., RODRIGUEZ, D., EVERS, M., BOCK-HUGHES, A., LIU, C., SMITH, Q., LUTZE, J., BAUMEISTER, J., KALMER, M., OLSCHOK, K., NICHOLSON, B., SILVA, D., MAXWELL, L., DOWGIELEWICZ, J., RUMI, E., PIETRA, D., CASETTI, I. C., CATRICALA, S., KOSCHMIEDER, S., GURBUXANI, S., SCHNEIDER, R. K., OAKES, S. A. & ELF, S. E. 2022. Type I but Not Type II Calreticulin Mutations Activate the IRE1alpha/XBP1 Pathway of the Unfolded Protein Response to Drive Myeloproliferative Neoplasms. Blood Cancer Discov, 3, 298-315.

      KAUR, A., VENKATESAN, A., KANDARPA, M., TALPAZ, M. & RAGHAVAN, M. 2024. Lysosomal degradation targets mutant calreticulin and the thrombopoietin receptor in myeloproliferative neoplasms. Blood Adv, 8, 3372-3387.

      MERY, L., MESAELI, N., MICHALAK, M., OPAS, M., LEW, D. P. & KRAUSE, K. H. 1996. Overexpression of calreticulin increases intracellular Ca2+ storage and decreases store-operated Ca2+ influx. J Biol Chem, 271, 9332-9.

      NAKAMURA, K., ZUPPINI, A., ARNAUDEAU, S., LYNCH, J., AHSAN, I., KRAUSE, R., PAPP, S., DE SMEDT, H., PARYS, J. B., MULLER-ESTERL, W., LEW, D. P., KRAUSE, K. H., DEMAUREX, N., OPAS, M. & MICHALAK, M. 2001. Functional specialization of calreticulin domains. J Cell Biol, 154, 961-72.

      PECQUET, C., PAPADOPOULOS, N., BALLIGAND, T., CHACHOUA, I., TISSERAND, A., VERTENOEIL, G., NEDELEC, A., VERTOMMEN, D., ROY, A., MARTY, C., NIVARTHI, H., DEFOUR, J. P., EL-KHOURY, M., HUG, E., MAJOROS, A., XU, E., ZAGRIJTSCHUK, O., FERTIG, T. E., MARTA, D. S., GISSLINGER, H., GISSLINGER, B., SCHALLING, M., CASETTI, I., RUMI, E., PIETRA, D., CAVALLONI, C., ARCAINI, L., CAZZOLA, M., KOMATSU, N., KIHARA, Y., SUNAMI, Y., EDAHIRO, Y., ARAKI, M., LESYK, R., BUXHOFER-AUSCH, V., HEIBL, S., PASQUIER, F., HAVELANGE, V., PLO, I., VAINCHENKER, W., KRALOVICS, R. & CONSTANTINESCU, S. N. 2023. Secreted mutant calreticulins as rogue cytokines in myeloproliferative neoplasms. Blood, 141, 917-929.

      PIETRA, D., RUMI, E., FERRETTI, V. V., DI BUDUO, C. A., MILANESI, C., CAVALLONI, C., SANT'ANTONIO, E., ABBONANTE, V., MOCCIA, F., CASETTI, I. C., BELLINI, M., RENNA, M. C., RONCORONI, E., FUGAZZA, E., ASTORI, C., BOVERI, E., ROSTI, V., BAROSI, G., BALDUINI, A. & CAZZOLA, M. 2016. Differential clinical effects of different mutation subtypes in CALR-mutant myeloproliferative neoplasms. Leukemia, 30, 431-8.

      SONNICHSEN, B., FULLEKRUG, J., NGUYEN VAN, P., DIEKMANN, W., ROBINSON, D. G. & MIESKES, G. 1994. Retention and retrieval: both mechanisms cooperate to maintain calreticulin in the endoplasmic reticulum. J Cell Sci, 107 (Pt 10), 2705-17.

      SUZUKI, J., KANEMARU, K., ISHII, K., OHKURA, M., OKUBO, Y. & IINO, M. 2014. Imaging intraorganellar Ca2+ at subcellular resolution using CEPIA. Nat Commun, 5, 4153.

      VENKATESAN, A., GENG, J., KANDARPA, M., WIJEYESAKERE, S. J., BHIDE, A., TALPAZ, M., POGOZHEVA, I. D. & RAGHAVAN, M. 2021. Mechanism of mutant calreticulin-mediated activation of the thrombopoietin receptor in cancers. J Cell Biol, 220, e202009179.

      WIJEYESAKERE, S. J., GAFNI, A. A. & RAGHAVAN, M. 2011. Calreticulin is a thermostable protein with distinct structural responses to different divalent cation environments. J Biol Chem, 286, 8771-85.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a multi-modal study of associative learning and memory in humans, that combines scalp EEG, pupillometry and behavioral analysis to explore the construct of mnemonic prediction errors (MPEs), in terms of their relationship to attention and cognitive control. Across two pooled studies, participants performed associative memory tasks in which they learned the relationship between a cue word (action verb) and subsequent picture (animate or inanimate) with a strong vs. weak (4 or 1 repetitions) encoding manipulation. At test, participants were encouraged to generate a prediction following the cue word to determine whether the subsequently presented picture was a match or mismatch.

      The timecourse of pupillary responses during match decisions were decomposed using temporal principal components analysis, which identified 6 distinct and overlapping processes. Some of the components (PC3/PC4) exhibited sensitivity to both the strength and mismatch conditions, as well as behavior (both RT and accuracy) and retrieval success on the subsequent trial. Furthermore, relationships were also observed between pupillary responses (specifically for PC4) and both frontal theta and posterior alpha power measures obtained from scalp EEG in Experiment 2, as well as for frontal theta and subsequent learning from mismatch stimuli (assessed using subsequent memory findings from a surprise recognition test). The authors suggest the findings indicate that MPEs elicit changes in attention, arousal and cognitive control which impact subsequent learning.

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPEs.

      Weaknesses:

      The technical proficiency and complexity of the study and analysis also presents a clear limitation and challenge for interpretation. It is likely that readers, even those that are quite knowledgeable about the methods, constructs, and questions being addressed will often struggle (as this reviewer did) to keep the large set of findings in mind and gain understanding of how they all fit together.

      Indeed, it seems like there many threads running together in the paper which make it challenging to find the through-line of the key findings. The authors do address some of the key questions motivating the paper in the Introduction, but the results are somewhat ambiguous with regard to the primary question of the study as to whether the detection of MPEs leads to interaction among cognitive control, attention, and arousal. To their credit, the authors tackle this question through both cross-correlation and formal mediation analyses, and summarize these in diagrammatic figures (Figure 3, Figure 6). Yet it is not resolved whether the results represent a clear answer pointing to independence, or rather a lack of statistical power, or ill-resolved formulation of the mediational relationship. In particular, the cross-correlation suggests that posterior alpha suppression in response to MPEs does precede frontal theta, yet this indirect relationship does not explain the variation in trial-by-trial RTs on mismatches. This suggests a potential model misspecification.

      In addition to the primary interaction issue mentioned above (between cognitive control, attention & arousal), the Introduction lays out a number of claims:

      (1) That pupil size will be more sensitive to strong than weak MPEs.

      (2) That MPE-linked increases in attention (indexed with posterior alpha suppression) and arousal (indexed with pupil size) will be linked to learning.

      (3) That MPE learning will vary as a function of prediction strength.

      Given the focus on learning, it is somewhat surprising that learning is not included in the mediation models. As the authors indicate in the Discussion, the use of trial-by-trial RT variation to drive the mediation model might be problematic, given that the RTs are sensitive to a range of factors beyond mnemonic prediction strength and also are under competing pressures (longer for mismatches than matches, due to surprise-linked slowing, but also faster following stronger rather than weaker mnemonic predictions). Thus, an alternative possibility might be to use trial-by-trial recognition of mismatches as the outcome variable in mediation models rather than trial-by-trial RT as the independent variable.

      A large component of the results (Sections 2 and 3) is devoted to analyses of cue-linked pupil and EEG processes that putatively reflect mnemonic predictions (i.e., occurring before picture probes are presented and match/mismatch detection, i.e., MPEs occur). Yet these Results and the subsequent pupillary PCA components (PC1 and PC5) that are elicited are not well-integrated with the primary themes of the paper or the causal hypotheses. One finding that does seem to figure prominently (in that it is mentioned in Abstract, Introduction & Discussion) relates to the amount of attention allocated to the mnemonic prediction generation. Yet this finding is not well emphasized in the Results themselves. Possibly it refers to the negative relationship between posterior alpha during memory retrieval and the magnitude of pupillary PC3 component, described in Section 3. But it was quite challenging to identify amongst the wealth of results described in this Section as well as the others. More generally, the large amount of findings described across all four lengthy Results sections makes it challenging for readers to discern what are the key ones that the authors would like to highlight.

      It is recommended that the authors do another pass through the paper to better highlight the most critical findings that they want to emphasize or which are most interpretable from a mechanistic and causal flow perspective and then de-emphasize or move other findings to the Supplemental Materials. Although the authors are to be commended for such a rigorous and comprehensive set of analyses, there are so many of them and findings, that the key points get buried and the reader needs to struggle potentially unnecessarily to identify the key take-away points.

      We thank Reviewer 1 for the helpful feedback on how the manuscript can be further strengthened. We recognize that the rich set of findings can overwhelm the reader, resulting in difficulty discerning the main take aways about the effects of mnemonic prediction errors. We particularly appreciate Reviewer 1’s encouragement to restructure the manuscript so as to focus on the findings reported in Sections 1 and 4 of the original revision; the current revision now focuses on these key observations.

      As part of this restructuring, Reviewer 1 also proposed moving the content from Sections 2 and 3 of the original revision to the Supplement. We agree with the Reviewer that the questions addressed in these sections on retrieval-related processes are not the main focus of the paper, but that they are informative in their own right. To avoid their getting lost in the Supplement, we decided that these results would be better served in a separate manuscript and thus we have removed them entirely. 

      We acknowledge in the revised manuscript that the mediation and cross-correlation analyses were exploratory and that these specific analyses may not be well powered in the current experiments. With respect to Reviewer 1’s concerns about the specification of the mediation models, we were motivated to test whether MPEs trigger an increase in cognitive control that in turn, triggers an increase in attention and/or arousal (Fig. 4a); as such, we designed the model to assess whether, on strong MPE trials, frontal theta mediates the relationship between prediction strength and attention/arousal. As noted in the manuscript and raised by Reviewer 1, mismatch RT here is an imperfect measure of trial-level prediction strength. Future experiments that selectively elicit strong MPEs and have a more controlled measure of trial-level prediction strength may be better equipped to address these questions about interactions between control, attention, and arousal. We agree with Reviewer 1 that models assessing subsequent memory as an outcome would be desirable. However, given that (a) we did not find strong evidence for interactions at the time of a strong MPE and (b) we only observed a relationship between frontal theta and subsequent memory (but not posterior alpha or pupil), subsequent memory mediation models do not appear to be well justified. Altogether, these findings illuminate open avenues for future research.

      Reviewer #2 (Public Review):

      Summary:

      The authors studied cognitive control and attention in response to mnemonic prediction errors (MPEs): situations in which the external reality violates internal memory-based predictions. The behavioral task first established strong versus weak predictions, and then either confirmed or violated these predictions. The authors examined markers of cognitive control (frontal theta) and attention (posterior alpha suppression, pupil response) while strong and weak predictions were confirmed or violated. They found increased cognitive control (frontal theta) for strong MPEs, which correlated with subsequent memory. Markers of attention (alpha suppression, pupil response) also accompanied strong MPEs but did not correlate with subsequent memory.

      Pupil response was investigated using an interesting approach that decomposes the response into different components, finding that different components respond earlier or later and show different correlations with MPEs and their strength. The authors also investigated how EEG, reaction time, and pupil responses correlated with one another, providing further insight into the mechanism underlying the response to MPEs. Together, the study points toward multiple control and attention mechanisms involved in MPE response and memory.

      Strengths:

      The study has a clear behavioral paradigm with multiple measures — behavioral, EEG, and pupillometry — that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

      Weaknesses:

      The methods are rigorous, and the data support the claims. The weaknesses are minor and are offered here as avenues for future research.

      (4) The relationships the authors find between brain measures and pupil components were largely not specific to mismatches/matches. Thus, the specificity of this relationship is untested.

      (5) The results with subsequent memory are important and address a major gap in the field that largely did not relate neural effects of MPE to subsequent memory. However, one major limitation of the study is that the authors did not test memory for matches. I understand the logic of avoiding testing matches. Because matches were repeated more times in the study, it’s not a fair comparison and could change participants’ overall criterion for old/new decisions. Future research could address this, e.g., by testing weak matches or potentially using a between-subject design.

      We appreciate Reviewer 2’s helpful feedback during the review process and encouraging comments on the strengths of the manuscript. We agree and note in the revision that it would be illuminating for future studies to contrast memory for events that violate and confirm mnemonic predictions.

      Comments on revised version

      The authors addressed all my concerns. I appreciate the authors’ thoughtful and detailed response.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It was challenging to read the paper with key findings happening at the time of the MPE presented first, and then to go “backwards in time” to examine process that occurred at preceding time periods (i.e., prior to probe presentation), and then again forward in time to examine learning related processes.

      (1) Restructuring the Results

      In this regard two distinct recommendations are made:

      - Move Sections 2 and 3 to Supplemental Materials, to maintain the focus on the key findings related to detection of MPEs and their effects on subsequent learning.

      - An alternative structure would be to present the cue-locked findings first, which relate to mnemonic predictions, then to those that occur when those predictions get violated, and finally to the learning process that occur following MPEs and which can be detected with subsequent recognition tests.

      We thank the Reviewer for this encouragement to restructure the manuscript. To address concerns regarding the density of the paper and the cohesiveness of the findings, we removed the content that was in Sections 2 and 3 of the original revision. We will publish those results in a separate paper with additional analyses to more directly address previously raised questions regarding the mechanisms indexed by frontal theta and posterior alpha at the time of retrieval.  

      (2) Integration of the summary figures

      At the minimum, a recommendation would be to better integrate Figure 6, and maybe various versions of Figure 3, earlier into the text, preferably even in the Introduction, and then repeatedly refer to them throughout the results. For example, in Figure 6, linking the leftmost panel of the figure to Sections 2 and 3 is critical, and the righthand panel to Sections 1 with the rightmost part related to learning explicitly linked to Section 4.

      We moved the summary figure up to now be Figure 1; we reference this figure in the Introduction and throughout the manuscript; and we additionally make reference in the figure to the association between frontal theta and PC3 at the time of a strong MPE as well as the cross-correlation outcome. We hope these modifications further aid the reader in identifying the main findings of the manuscript.

      (3) Specification of the mediation models

      Additionally, for the mediation models it is quite unclear why mismatch RT is treated as the index of MPE magnitude, as this seems to be where the problem may lie in model fitting. Why not think of this as an outcome variable (since would seem to be a causal outcome of the underlying functional processes elicited by mismatch detection)?

      We thank the reviewer for these thoughtful comments. Our goal in designing the mediation models in Fig. 4a-c was to test our hypothesis that strong MPEs trigger an increase in cognitive control that, in turn, triggers an increase in attention and/or arousal; thus, attention/arousal should be the outcome and cognitive control should be the mediator. Given that increases in attention and cognitive control were selectively observed for strong and not weak MPEs, the models were restricted to strong MPEs. We therefore needed a trial-level measure of MPE magnitude to determine whether stronger MPEs in the strong condition elicit greater attention/arousal by engaging more cognitive control. As such, we decided to use mismatch RT as a proxy measure of MPE magnitude; we acknowledge in the text that this measure is imperfect. While a model with mismatch RT as the outcome variable is possible, the relative timing of the attention/arousal effects (which largely occur following responses) would render interpretation to be more challenging. Had stronger evidence of indirect effects emerged in Figs. 4b-c, we could have conducted model comparison with RT as an outcome. We acknowledge in the text that these analyses were exploratory and characterization of potential indirect effects will require more data and a more precise measure of MPE magnitude.

      Alternatively, examining trial-by-trial recognition memory of mismatches as the relevant outcome variable would also seem to capture the functional process of interest. In this regard, have the authors examined whether trial-by-trial RT on mismatches predicts subsequent recognition of these items? If this direct relationship holds, it could be a target for mediation analyses in itself.

      We agree with the reviewer that in theory, a mediation model predicting subsequent memory would be a desirable test of an integrated model of the mechanisms underlying MPE-driven learning. However, the mediation analyses conducted to address the functional relationships at the time of a prediction error (Fig. 4) are not well-powered to begin with; this limitation is raised in the Results and the Discussion. A mediation model predicting subsequent memory would be similarly underpowered; given that there is not strong evidence for indirect effects at the time of a strong MPE, and that only frontal theta – and not posterior alpha nor pupil – predicts subsequent memory, we think that such a mediation model is not well justified to include. Such a model should be more directly tested by well-powered designs in future studies.

      (4) Hippocampal theta in the Introduction

      The Introduction discusses hippocampal theta as well as frontal theta yet also makes clear that the former is not really well-detected or analyzed using scalp EEG. Consequently, a recommendation would be to remove this paragraph from the Introduction, since it can be misleading and a “red herring” for the reader, and instead only bring up this point in the Discussion section, as a pointer to the need for future research using methods that may be more sensitive to hippocampal interactions with PFC regions.

      We appreciate this point and moved discussion of the hippocampus from the Introduction to the Discussion.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This manuscript is very interesting and timely. By introducing the critical effects of desolvation barriers and solvent (water)-separated minima into the implicit-solvent potentials (of mean force, PMFs) for coarse-grained molecular dynamics simulations of biomolecular liquid-liquid phase separation (LLPS), this work fills a gap that should be apparent to researchers of protein folding in the past couple of decades but has so far escaped deserved attention such that these basic features of aqueous solvation have seldom, though not never, been invoked in recent studies of biomolecular condensates. Although the present paper deals almost exclusively with homopolymers, this work can be a foundation for the future development of a new, more physical coarse-grained interaction scheme for simulating amino acid sequence-dependent effects, which I presume is the authors' ongoing or next endeavor. The results presented in this manuscript are highly valuable.

      We thank the reviewer for all these positive comments.

      However, there is room for improvement in the authors' description of (i) the broader impact of effects of desolvation barrier and solvent-separated minimum in the thermodynamics of biomolecular condensates, especially with regard to the ramifications on hydrostatic pressure-dependent effects; (ii) the physical implication of using a 20-parameter hydropathy scale rather than a 210-parameter pairwise amino acid interaction scheme; and (iii) temperature-dependent effects, including the authors' discussion of "enthalpic" and "entropic" contributions. In all these aspects, the authors' discussion should be put in a more comprehensive context of the existing literature. At a few other places, the description of the methods and results should be clarified as well. Accordingly, the authors should revise the manuscript to address the following items thoroughly within the revised manuscript (not merely in the response letter) with the additional references mentioned below included in the revised discussion:

      (1) In several places, e.g., on line 77 (p.2), the authors appear to suggest that "implicit-solvent representation" is the origin of the deficiency in commonly utilized coarse-grained potentials that this study is aiming to rectify. But desolvation barriers and solvent-separated minima are also features of implicit-solvent representations; they are just features that should be incorporated in more accurate implicit-solvent potentials. This point is stated quite clearly and accurately in the Abstract (p.1) but not consistently in the rest of the text. The authors should check the entire text carefully to ensure that a coherent, accurate perspective is presented.

      We thank the reviewer for pointing out this important issue. We agree that implicit-solvent representation itself is not the origin of the deficiency. Our intention is to incorporate desolvation-inspired effective terms within an implicit-solvent coarse-grained framework, because many commonly used implicit-solvent potentials do not directly account for the desolvation features, such as the desolvation barrier and the solvent-separated potential well.

      We have revised the Abstract, Introduction, Results, and Discussion to make this distinction consistent throughout the manuscript. The revised text now emphasizes that the model remains an implicit-solvent CG model, but contains additional effective terms inspired by desolvation features observed in all-atom PMFs.

      Corresponding changes:

      (1) (page 1, lines 17–20) The Abstract identifies the model as an implicit-solvent CG model with added desolvation terms:

      "Here, guided by all-atom simulations and experimental measurements, we develop a desolvation-aware implicit-solvent CG model by incorporating residue-level desolvation terms directly into the pairwise energy function and apply it to investigate LLPS of intrinsically disordered proteins."

      (2) (page 2, lines 80–81) The Introduction retains the implicit-solvent description of existing residue-level CG models:

      "Despite these advances, most residue-level CG models rely on implicit solvent representations, in which individual water molecules are not explicitly represented."

      (3) (page 2, lines 87–89) The specific limitation is identified as the absence of a direct account of the multi-step desolvation process:

      "More importantly, conventional implicit-solvent CG models used for LLPS usually do not directly account for the multi-step desolvation process that accompanies the transition from a dilute solution to a dense condensate."

      (4) (page 4, lines 172–174) The added terms are described as part of a desolvation-inspired effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (5) (page 15, lines 521–522) The Discussion restates the implicit-solvent model framework:

      "To address this challenge, we developed a desolvation-aware implicit-solvent CG framework that incorporates desolvation barrier and solvent-separated terms into the pairwise potential."

      (2) In the discussion of the importance of desolvation barriers and solvent-separated minima in the Introduction (pp.1-3), connections should be drawn to recent works that utilize these PMF features to rationalize hydrostatic pressure (P)-modulated effects on biomolecular LLPS, including the P-dependent reentrant phase separation of alpha elastin; see Cinar et al. (2019) Chem Eur J 25:13049 (https://chemistryeurope.onlinelibrary.wiley.com/doi/full/10.1002/chem.201902210) and references therein, especially discussions around Figures 10, 11 & 13 in this reference.

      We thank the reviewer for bringing this literature to our attention. We agree that pressure-modulated LLPS provides important context for the physical relevance of desolvation barriers and solvent-separated minima. We have therefore expanded the Introduction and Discussion to connect our model to prior work on hydrostaticpressure-dependent condensate behavior, including pressure-dependent reentrant phase separation of alpha-elastin.

      Corresponding changes:

      (1) (page 3, lines 93–95) The Introduction connects the PMF features to hydrostaticpressure-dependent LLPS:

      "Related studies on hydrostatic pressure effects have further suggested that desolvation barriers and solvent-separated minima can help rationalise pressure-modulated LLPS behaviors, including the pressure-dependent reentrant phase separation of α-elastin Cinar et al. (2019, 2018)."

      (2) (page 15, lines 533–537) The Discussion states the pressure-dependent implication conservatively:

      "These findings may also provide a useful physical basis for future studies of pressure-dependent condensate behavior, as pressure-induced changes in hydration, solvent-separated states, and desolvation barriers have been proposed to contribute to pressure-modulated and reentrant LLPS Dias and Chan, 2014); Cinar et al. (2019, 2018)."

      (3) In the lower panels of Figures 2D, E (p.5), what do the differently colored small circles in the double-minimum free energy profiles represent? Does the color shading have the same meaning as that in the upper panels? If so, what do the positions of the circles on the free energy profile represent? The authors should clarify this.

      We thank the reviewer for identifying this ambiguity. The small circles in the lower panels of Figures 2D and 2E are qualitative schematic representations of residue-pair configurations along the effective pair-potential profile. Their blue and green colors distinguish the barrier-variation and solvent-separated-well cases, respectively; they are not a quantitative scale and do not encode temperature or population magnitude. The positions of the circles indicate the direct-contact, barrierregion, or solvent-separated regions, while the density of circles schematically represents the population of configurations.

      We have clarified this interpretation in the Figure 2 caption and aligned the Results text with the redistribution among direct-contact, barrier-region, and solvent-separated states.

      Corresponding changes:

      (1) (page 6, Figure 2D and E lower panels) The schematics distinguish low and high ε_b or ε_ss and use the density and position of the circles to depict populations in the direct-contact, barrier-region, and solvent-separated regions; the blue and green colors distinguish the two parameter families and are not a quantitative scale.

      (2) (page 6, Figure 2 caption) The caption defines the population encoding used in the lower panels:

      "The lower panels schematically illustrate how changes in ε<sub>b</sub> and ε<sub>ss</sub> alter the distribution of residue-pair configurations. The small circles indicate schematic populations of residue-pair configurations along the potential profile, with denser circles representing a higher population."

      (3) (page 5, lines 210–213) The Results text connects the schematics to redistribution among the three residue-pair states:

      "These opposing effects suggest that the desolvation potential regulates macroscopic phase behavior by redistributing residue-pair configurations between direct-contact, barrier-region, and solvent-separated states (lower panels of Figure 2D and E)."

      (4) The discussion regarding entropy and enthalpy around Figure 2 is quite confusing as it stands. What do the authors mean exactly by the association of entropy or enthalpy with the desolvation barrier of the solvent-separated minimum? Are they referring to conformational entropy?

      We thank the reviewer for pointing out this ambiguity. We agree that our original wording around entropy and enthalpy could be misleading, because it might imply a rigorous thermodynamic decomposition of the PMF. In the revised manuscript, we have therefore clarified that the effect of the desolvation barrier refers to an entropyrelated configurational restriction of residue-pair configurations near the barrier region, rather than the overall conformational entropy of the entire chain. We also replaced the previous "enthalpic stabilization" wording with "effective free-energy stabilization" to avoid implying that the solvent-separated minimum is treated as a purely enthalpic contribution.

      Corresponding changes:

      (1) (page 5, lines 199–201) The barrier effect is described in terms of the sampled residue-pair population:

      "Analysis of residue-residue radial distribution functions showed that higher ε<sub>b</sub> suppresses the population of configurations near the barrier region (Figure 2—figure Supplement 1C)."

      (2) (page 5, lines 201–202) The entropy-related statement is restricted to configurational sampling near the barrier:

      "This reduction in the statistical weight of barrier-region configurations can be interpreted as an entropy-related configurational restriction and thus disfavors phase separation."

      (3) (page 5, lines 209–210) The solvent-separated minimum is described as an effective free-energy contribution:

      "This solvent-separated minimum provides effective free-energy stabilization for water-mediated configurations and thereby promotes phase separation."

      (4) (page 6, Figure 2D and E lower panels) The schematic headings are "Barrier-mediated Restriction" and "Solvent-separated Stabilization", avoiding a strict entropy-enthalpy decomposition.

      (5) Do the authors assume that the PMF (effective implicit-solvent potential) is a purely enthalpic term? It appears to be the authors' assumption. If so, the assumption has to be stated clearly in their discussion of "entropy" vs "enthalpy" around Figure 2.

      We thank the reviewer for raising this important point. We do not assume that the PMF obtained from all-atom simulations is a purely enthalpic term. The PMF is a free-energy profile that contains enthalpic and entropic contributions. The current manuscript defines the PMF as −k<sub>B</sub> T lnP(r), uses the all-atom PMFs to motivate a nonbonded effective coarse-grained potential, and describes the solvent-separated minimum as providing effective free-energy stabilization. We do not perform a rigorous enthalpy-entropy decomposition, and the revised wording avoids implying such a decomposition.

      Corresponding changes:

      (1) (page 16, lines 594–595) The Methods define the PMF as a free-energy profile obtained from the radial probability density:

      "The potential of mean force (PMF) was computed as PMF(r) = −k<sub>B</sub>T ln P(r), where P(r) is the radial probability density obtained from the production trajectory."

      (2) (page 5, Figure 1 caption) The CG interaction is labeled as an effective potential rather than as an enthalpic PMF decomposition:

      "Pairwise effective potential incorporating desolvation-inspired terms. Different curves correspond to different desolvation parameters."

      (3) (page 4, lines 172–174) The parameters are described as shaping an effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (4) (page 5, lines 209–210) The solvent-separated contribution is described using free-energy language:

      "This solvent-separated minimum provides effective free-energy stabilization for water-mediated configurations and thereby promotes phase separation."

      (6) Closely related to points 3-5 above, it should be stated clearly that the "temperature" used in the authors' simulations does not represent experimental temperature if the authors are using purely enthalpic effective potentials because PMFs are in fact temperature-dependent. This clarification is necessary to avoid misunderstanding. In this regard, it should be noted that temperature-dependent effective interactions have been used for modeling biomolecular condensates in analytical theory (Lin, Song, Forman-Kay & Chan, J Mol Liq 2017, already in the citation list) as well as in coarse-grained molecular dynamics simulations [Dignon et al. (2019) ACS Cent Sci 5:821-830 (https://pubs.acs.org/doi/10.1021/acscentsci.9b00102); Chakravarti & Joseph (2025) Protein Sci 34:e70284 (https://onlinelibrary.wiley.com/doi/10.1002/pro.70284)]. The latter two studies, not cited currently, are particularly relevant and thus should be cited because the authors may wish to incorporate temperature-dependent features in their ongoing or future effort in constructing a more comprehensive coarse-grained interaction scheme for biomolecular LLPS simulation.

      We agree with the reviewer that the simulation temperature should be interpreted carefully. In the present simulations, the effective potential is temperature-independent within each chosen parameter set. Therefore, the reduced temperature primarily serves as a model temperature controlling the relative strength of thermal fluctuations, rather than as a direct experimental temperature. We have clarified this point in the revised manuscript and added relevant references on temperature-dependent effective interactions, which represent an important direction for future model development.

      Corresponding changes:

      (1) (page 4, lines 177–179) The manuscript states that the absolute simulation temperature is not an experimental temperature:

      "Because the effective interaction parameters used in the model are temperature-independent, the absolute simulation temperature should not be directly interpreted as an experimental temperature."

      (2) (page 4, lines 182–184) The reduced temperature is identified as a model temperature:

      "Accordingly, T<sup>*</sup> should be interpreted primarily as a model temperature that controls the relative strength of thermal fluctuations, rather than as having a direct quantitative correspondence with experimental temperature."

      (3) (page 14, lines 481–485) The FUS LC temperature comparison is framed cautiously:

      "It is worth noting that the residue-level coarse-grained models used here employ temperature-independent effective interaction parameters. As a result, the temperature values reported here cannot be interpreted as quantitatively equivalent to experimental temperatures, particularly when they deviate substantially from room-temperature conditions."

      (4) (page 15, lines 563–567; continues on page 16, lines 568–569) The Discussion identifies temperature-dependent effective interactions and a corresponding future extension:

      "In addition, effective interactions themselves can be temperature-dependent, as demonstrated in analytical theories and coarse-grained simulations of biomolecular condensates Lin et al. (2017); Dignon et al. (2019); Chakravarti and Joseph (2025). Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      (7) In tackling "entropy" vs "enthalpy", it should be noted that the temperature dependence of the effective interactions entails an entropic contribution (which is itself temperature dependent) in addition to conformational entropy. As for the effective potential with desolvation barrier and solvent-separated minimum, it should be noted that the decomposition into entropic and enthalpic contributions at the direct contact, desolvation barrier, and solvent-separated minimum can be dramatically different, see, e.g., MaCallum et al. (2007) PNAS 104:6206-6210 (https://www.pnas.org/doi/full/10.1073/pnas.0605859104) and references therein.

      We thank the reviewer for this important clarification. We agree that temperature-dependent effective interactions can contain entropic contributions beyond conformational entropy and that the balance of enthalpic and entropic contributions may differ among the direct-contact minimum, desolvation barrier, and solvent-separated minimum. The present model does not decompose the PMF into temperature-dependent enthalpic and entropic components; accordingly, we have avoided assigning those components to individual PMF features. The Discussion cites explicit-solvent PMF analyses when noting residue-pair and temperature dependence and identifies temperature-dependent desolvation parameters as an important future extension.

      Corresponding changes:

      (1) (page 15, lines 561–563) The Discussion cites explicit-solvent PMF work when noting residue-pair and temperature dependence:

      "Explicit-solvent PMF analyses have shown that desolvation barrier heights and solvent-separated minima can differ substantially among residue pairs and may also exhibit temperature dependence Cinar et al. (2019); Debiec et al. (2014); MacCallum et al. (2007)."

      (2) (page 15, lines 563–566) The manuscript states that effective interactions can themselves depend on temperature:

      "In addition, effective interactions themselves can be temperature dependent, as demonstrated in analytical theories and coarse-grained simulations of biomolecular condensates Lin et al. (2017); Dignon et al. (2019); Chakravarti and Joseph (2025)."

      (3) (page 15, lines 566–567; continues on page 16, lines 568–569) Temperature-dependent desolvation parameters are identified as a future model extension:

      "Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      (8) P.7, line 340: The proportionality relation follows directly from the standard FloryHuggins result T_c = T chi(T)/chi_c, thus the proportionality constant is exactly 1/chi_c. Is this the standard relation that the authors are invoking here? The authors should clarify this.

      We thank the reviewer for pointing out the missing intermediate steps. Yes, the relation we invoked is based on the standard Flory-Huggins critical condition. In the revised manuscript, we have expanded the derivation to explicitly show how the critical condition chi(T_c) = chi_c leads to the relation between chi(T_sim) - chi_c and the normalized thermal distance (T_c - T_sim)/T_sim.

      We also revised the wording to avoid presenting this as a universal law. The relation is now presented as a simulation-supported trend within the present model, rationalized by a simplified linear-response assumption between Delta R_g and the excess interaction strength.

      Corresponding changes:

      (1) (page 7, lines 263–264) The critical-condition substitution is now shown explicitly:

      "At the critical point, χ(T<sub>c</sub>) = χ<sub>c</sub>, which gives ε<sub>eff</sub> = k<sub>B</sub>T<sub>c</sub>χ<sub>c</sub>. Substituting this relation into the expression for χ(T<sub>sim</sub>) yields χ(T<sub>sim</sub>) = χ<sub>c</sub>T<sub>c</sub>/T<sub>sim</sub>."

      (2) (page 7, line 265) The resulting relation is written as Equation (2):

      "χ(T<sub>sim</sub>) − χ<sub>c</sub> = χ<sub>c</sub> (T<sub>c</sub> − T<sub>sim</sub>)/T<sub>sim</sub>."

      (3) (page 7, lines 266–267) The fixed-chain-length assumption is stated explicitly:

      "For systems with the same chain length, χ<sub>c</sub> is a fixed constant. Thus, the deviation from the critical interaction parameter is directly related to the rescaled thermal distance (T<sub>c</sub> − T<sub>sim</sub>)/T<sub>sim</sub>."

      (9) The study on dynamic consequences on pp.8-11 is interesting, but clarifications are necessary:

      (i) The vertical schematic in Figure 4A should be explained in detail in its entirety. As it stands, no explanation is provided either in the figure caption or in the text. In particular, what does "elasticity driven" refer to?

      (ii) The top snapshot in Figure 4A is labeled t_sim = 0 ns. Does it mean that the snapshot shown is the only chain configuration that the authors used to start the simulation, and that the snapshot does NOT represent the result of any time evolution, no matter how short the duration is? However, if that is the case, why is this snapshot identified with spinodal decomposition if it is not the product of a time evolution from a more homogeneous configuration?

      (iii) Related to (ii) - do the rectangular boxes shown represent the entire simulation box or just part of the box containing the polymer chains? One would imagine that if the top snapshot represents spinodal decomposition, the simulation would have been started at a more uniform distribution a short time prior? Why is this not the case?

      (iv) What precisely do the small yellow beads and black-colored springs in the zoomin image of Figure 4E represent?

      We thank the reviewer for all these inspiring comments and questions. We agree that the original Figure 4 schematic did not sufficiently explain the sequence of dynamical events and the meaning of several graphical elements. We have therefore revised both the Figure 4 caption and the Results text to make the schematic self-contained and to clarify how it relates to the quantitative analyses in Figure 4F and G.

      First, we replaced the phrase "elasticity driven" with a more precise description of "viscoelastic resistance". In the revised text, interfacial tension is described as favoring domain fusion thermodynamically, whereas transient inter-chain network connectivity within dense domains generates viscoelastic resistance to the deformation required for coalescence. This revision avoids implying that elasticity is the driving force and instead identifies it as a resistance that delays domain fusion kinetically during the plateau regime.

      Second, we clarified the meaning of t_sim = 0 ns and its relation to spinodal decomposition. The system was equilibrated at a supercritical temperature to obtain a homogeneous one-phase state and was then instantaneously quenched to the target temperature. The label t_sim = 0 ns denotes the first snapshot immediately after the quench. It is not the only initial configuration used in all simulations; the reported kinetic metrics were averaged over six independent slab simulation replicas.

      The t_sim = 0 ns snapshot is therefore described as a homogeneous but thermodynamically unstable post-quench state. Spinodal decomposition refers to the subsequent amplification of the initial density fluctuations after the quench, including the development of interconnected density fluctuations within 1-2 ns, rather than to a preceding evolution represented by the t_sim = 0 ns snapshot.

      Third, we clarified that the rectangular snapshots in Figure 4A show the entire simulation box viewed along the z-axis. The subsequent snapshots show how post-quench density fluctuations grow and reorganize into dense domains during spinodal decomposition.

      Finally, we clarified the symbols in the zoom-in schematic of Figure 4E. Yellow beads now denote residues involved in transient inter-chain contacts, and black springs denote schematic network connections formed by these contacts. These elements are meant to illustrate transient network connectivity and are not additional simulated particles or force-field terms.

      Corresponding changes:

      (1) (page 10, Figure 4A) The vertical schematic now labels the progression as "Thermodynamic instability", "Kinetic arrest (viscoelastic resistance)", "Domain coarsening (interfacial-tension dominated)", and "Dynamic equilibrium (chain self-diffusion)".

      (2) (page 10, Figure 4E) The plateau schematic labels the competing effects as "Interfacial Tension" and "Transient network resistance".

      (3) (page 10, Figure 4 caption) The caption defines the snapshots and the vertical schematic:

      "Upper snapshots show the simulation box along the z-axis at t<sub>sim</sub> = 0, 10, and 500 ns. The vertical schematic summarizes the dynamical progression described in the main text, from the post-quench spinodal instability to kinetic arrest, domain coarsening, and dynamic equilibrium."

      (4) (page 11, lines 368–370) The first recorded time point after the quench is defined explicitly:

      "Here, t<sub>sim</sub> = 0 ns denotes the first snapshot immediately after the temperature quench, corresponding to a homogeneous but thermodynamically unstable nonequilibrium state."

      (5) (page 10, Figure 4 caption) The yellow beads and black springs are defined:

      "In the zoom-in view, yellow beads denote residues involved in transient inter-chain contacts, and black springs denote schematic network connections formed by these contacts."

      (6) (page 12, lines 401–404) The Results explain the physical meaning of the transient network:

      "In the zoom-in schematic in Figure 4E, this transient network is represented by connections between residues involved in inter-chain contacts, illustrating how multivalent interactions can resist domain deformation during the plateau regime."

      (10) In discussing dynamic effects, it is useful to draw connections to related works on the effect of chain flexibility on "aging" of condensate [Biswas & Potoyan (2024) PRX 45:9222-9245 (https://journals.aps.org/prxlife/abstract/10.1103/PRXLife.2.023011)] and characterization of viscoelasticity in simulations of biomolecular condensates [Tejedor et al. (2023) J Phys Chem B 127:4441-4459 (https://pubs.acs.org/doi/10.1021/acs.jpcb.3c01292)], as the effects of desolvation can be explored further based on these prior works.

      We thank the reviewer for these important references. We have added connections to simulation studies of condensate viscoelasticity and aging. The revised manuscript now places our dynamic results in the context of transient network connectivity, chain flexibility, sticker lifetime, desolvation-associated rigidification, and viscoelastic or aging-like material behavior.

      We present these connections conservatively as relevant context and as future directions for extending the current model, rather than claiming a new universal dynamic mechanism.

      Corresponding changes:

      (1) (page 12, lines 397–399) The dynamics section cites simulation-based rheological analyses of condensate viscoelasticity:

      "Similar viscoelastic effects have recently been quantified in molecular simulations of biomolecular condensates using rheological analyses of time-dependent material properties Tejedor et al. (2023)."

      (2) (page 12, lines 399–401) The manuscript connects condensate aging to chain flexibility, sticker lifetime, and desolvation-associated rigidification:

      "Molecular simulations of condensate aging have further highlighted the roles of chain flexibility, sticker lifetime, and desolvation-associated rigidification in promoting more solid-like states Biswas and Potoyan (2024)."

      (3) (page 12, lines 423–427) The kinetic interpretation is connected to viscoelastic andaging-like behavior:

      "The sensitivity of kinetic arrest and coarsening dynamics to desolvation parameters underscores the importance of incorporating desolvation features into coarse-grained potentials for more physically plausible molecular simulations of LLPS, especially when connecting microscopic interaction lifetimes to emergent viscoelastic or ageing-like material behavior."

      (4) (page 16, lines 569–572) The Discussion identifies simulation-based rheological analysis as a future direction:

      "An additional direction would be to combine these potentials with simulation-based rheological analyses to quantify how desolvation reshapes condensate viscoelasticity, aging-like maturation, and long-time material relaxation Tejedor et al. (2023); Biswas and Potoyan (2024)."

      (11) Much of the present study is based on the original HPS formulation of Dignon et al. (2018). In this regard and also in anticipation of future development of improved interaction schemes, several issues should be stated and discussed, even if briefly:

      (i) The original HPS model has a basic shortcoming in accounting for the relative interaction strengths of, among others, arginine vs lysine residues [Das et al. (2020) PNAS 117:28795-28805 (https://www.pnas.org/doi/10.1073/pnas.2008122117)].

      (ii) Compared to 210-parameter pairwise interaction schemes, such as KH in Dignon et al. (2018) and Joseph et al. (2021), the 20-parameter interaction scheme is likely too restrictive to account for pairwise amino acid residue interactions [Wessén et al. (2022) J Phys Chem B 45:9222-9245 (https://pubs.acs.org/doi/10.1021/acs.jpcb.2c06181)].

      (iii) The height of the desolvation barrier may vary significantly for different amino acid residue pairs, see, e.g., Figure 11 of Cinar et al. (2019) mentioned above (and references therein). The authors should discuss these nuances in the revised version. They may also wish to take them into consideration in future investigations.

      We thank the reviewer for the suggestion to clarify these limitations. We have revised the Discussion to acknowledge explicitly the limitations of the 20-parameter hydropathy-scale representation relative to more flexible 210-parameter pairwise interaction schemes for describing amino-acid-pair interactions. We have also added discussion emphasizing that future desolvation-aware models should incorporate residue-pair-specific parameters for the desolvation barrier and solvent-separated potential well.

      Corresponding changes:

      (1) (page 13, lines 445–446) The scope of the averaged baseline parameterization is stated explicitly:

      "This uniform parameterization captures the generic desolvation features of the PMFs but does not resolve residue-pair-specific variations in desolvation energetics."

      (2) (page 15, lines 551–554) The Discussion identifies the limitation of the 20-parameter HPS representation, including Arg/Lys interactions:

      "In particular, HPS-type models use a 20-parameter hydropathy-scale representation, which is useful for capturing generic IDP phase behavior but is not flexible enough to resolve residue-pair-specific chemical effects, such as the distinct interaction patterns of arginine and lysine residues Das et al. (2020)."

      (3) (page 15, lines 554–557) The greater flexibility of 210-parameter pairwise schemes is described:

      "More general 210-parameter pairwise interaction schemes, such as KH-type and related residue-pair-specific models, provide greater flexibility for encoding amino acid-pair preferences and capturing sequence-specific interaction heterogeneity Dignon et al. (2018b); Joseph et al. (2021); Wessén et al. (2022)."

      (4) (page 15, lines 557–561) The limitation of using one averaged desolvation parameter set is stated:

      "Second, the present desolvation model employs a single set of averaged parameters (α<sub>b</sub>, α<sub>ss</sub>) for all residue pairs. While this simplification is effective for isolating the generic physical consequences of desolvation, it has limitations in describing the pair-specific variations in the desolvation barrier and the solvent-separated minimum."

      (5) (page 15, lines 566–567; continues on page 16, lines 568–569) Residue-specific and temperature-dependent parameters are identified as a future extension:

      "Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in the molecular simulation of biomolecular condensates. Most residue-level coarse-grained models used for IDP phase separation employ implicit solvent and represent effective interactions through relatively simple pairwise potentials. While these models have been very useful, they usually do not explicitly distinguish direct contacts from solvent-separated interactions, nor do they include an energetic barrier associated with water removal. This manuscript attempts to address that limitation by introducing desolvation-inspired terms into coarse-grained models and examining their consequences for phase behavior, chain conformations, dense-phase packing, and dynamics. Strengths:

      The central idea is physically well motivated. Using a simple homopolymer model, the authors show that increasing the desolvation barrier suppresses phase separation, whereas stabilizing solvent-separated contacts enhances phase separation. They further show that solvent-separated interactions can reduce densephase over-compaction, which is a meaningful result given the known challenges in obtaining both accurate single-chain dimensions and realistic dense-phase properties from the same coarse-grained model. The finding that desolvation-like terms can reshape dense-phase packing without simply rescaling the overall interaction strength is interesting and could be useful for future model development. I also found the attempt to connect conformational changes across dilute and dense phases with thermal distance from the critical point to be intriguing. The dynamic analysis, including the FRAP-like simulations and the discussion of kinetic arrest during coarsening, adds another useful dimension to the work.

      Weaknesses:

      At the same time, there are several places where the manuscript would benefit from more careful framing. First, the desolvation terms are still effective coarse-grained parameters rather than a direct representation of water molecules. The language sometimes gives the impression that desolvation is being treated explicitly, whereas the model introduces desolvation-inspired effective interactions into an implicitsolvent framework.

      We thank the reviewer for the positive assessment and constructive suggestions. We agree that the desolvation terms should be described as effective coarse-grained parameters rather than explicit water molecules. We have revised the manuscript to describe the model as a desolvation-aware implicit-solvent coarse-grained framework with desolvation-inspired effective interaction terms.

      Corresponding changes:

      (1) (page 1, lines 17–20) The Abstract identifies the model as an implicit-solvent CG model:

      "Here, guided by all-atom simulations and experimental measurements, we develop a desolvation-aware implicit-solvent CG model by incorporating residue-level desolvation terms directly into the pairwise energy function and apply it to investigate LLPS of intrinsically disordered proteins."

      (2) (page 4, lines 151–153) The Results describe the added contributions as desolvation-related effective terms:

      "These observations underscore the importance of incorporating desolvation-related effective terms and exploring the effects of different desolvation strengths on the thermodynamics and kinetics of protein LLPS."

      (3) (page 4, lines 172–174) The pair interaction is described as a desolvation-inspired effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (4) (page 15, lines 541–543) The Discussion emphasizes that the framework retains water-mediated features within an implicit-solvent representation:

      "By retaining key water-mediated features while preserving the computational efficiency of implicit-solvent representations, this framework provides a mechanistic means to decouple overall phase-separation propensity from condensed-phase packing."

      Second, the conformational analysis is interesting, but the broader context of prior work on dilute-to-dense phase conformational reorganization of IDPs could be more clearly discussed. This would help clarify what is new in the present work, whether it is the conformational change itself, its dependence on desolvation terms, or the proposed scaling with distance from the critical point.

      We thank the reviewer for this suggestion. We agree that the conformational change itself should be placed in the context of prior work. The contribution of the present analysis is not simply the observation that IDP conformations can reorganize upon condensation. Rather, we examine how desolvation-inspired effective terms modulate dilute- and dense-phase conformations and how the conformational change correlates with thermal distance from the critical point within the present model.

      We have revised the Results section discussing Figure 3 to cite prior work and to state the interpretation of ΔR_g more clearly.

      Corresponding changes:

      (1) (page 6, lines 230–232) Prior work on conformational reorganization upon condensation is cited:

      "Previous studies have shown that IDP condensation can reorganize chain conformations by redistributing the balance between intra-chain and inter-chain interactions Wei et al. (2017); Hazra and Levy (2021); Tesei et al. (2021); von Bülow et al. (2025)."

      (2) (page 7, lines 241–242) The phase-dependent conformational response to the desolvation parameters is introduced:

      "In addition to the difference between dilute- and dense-phase conformations, varying the desolvation parameters further reveals a phase-dependent conformational response (Figure 3A–C)."

      (3) (page 7, lines 243–244) The stronger response in the dilute phase is stated directly:

      "Increasing ε<sub>b</sub> or decreasing ε<sub>ss</sub> shifts the dilute-phase R<sub>g</sub> distributions toward larger values, whereas the dense-phase R<sub>g</sub> remains comparatively insensitive to these parameter changes."

      (4) (page 7, lines 247–249) The source of the desolvation-dependent variation in ΔR_g is identified:

      "As a result, the desolvation-dependent variation in ΔR<sub>g</sub> = R<sub>g</sub><sup>dense</sup> − R<sub>g</sub><sup>dilute</sup> arises predominantly from the conformational changes of isolated chains in the dilute phase."

      (5) (page 7, lines 254–256) The observed relationship is presented as an approximate trend in the simulated systems:

      "Notably, data from the simulated systems approximately follow a common trend, revealing a strong correlation between the magnitude of conformational change and the thermal distance to the phase transition point (R<sup>2</sup> = 0.942, Figure 3D)."

      Third, the dynamic results are potentially useful, but the manuscript should more clearly articulate what is nontrivial beyond the expected slowing of local rearrangements by an added barrier in the potential.

      Overall, I think this is a useful and potentially important contribution.

      We thank the reviewer for this constructive comment and the positive overall assessment. We have revised the dynamics section to clarify that the nontrivial result lies in the competition between two effects: although the desolvation barrier directly slows local rearrangements, its reduction of dense-phase packing can reverse the net mobility trend at fixed temperature. At matched thermodynamic quench depth, the intrinsic slowing associated with energy-landscape roughness becomes evident. We also clarified that desolvation modulates transient kinetic arrest and domain-scale coarsening, not only local rearrangements.

      Corresponding changes:

      (1) (page 11, lines 353–357) The fixed-temperature and matched-quench-depth analyses are summarized as opposing contributions:

      "Together, the fixed-temperature and renormalized analyses in Figure 4C and D reveal two distinct and opposing contributions of desolvation to condensate dynamics. At fixed temperature, increasing ε<sub>b</sub> loosens dense-phase packing and thereby increases the measured diffusion coefficient, whereas at matched thermodynamic quench depth, the same parameter change suppresses chain mobility by roughening the microscopic energy landscape and slowing local rearrangements."

      (2) (page 11, lines 358–361) The multiscale interpretation is stated explicitly:

      "Condensate dynamics therefore emerge from a balance between density-regulated mobility and energy-landscape-regulated mobility, with macroscopic packing determining the dominant trend and microscopic barrier roughness imposing an additional kinetic modulation. This interplay highlights how desolvation reshapes condensate dynamics across multiple physical scales."

      (3) (page 12, lines 421–423) The dynamics section distinguishes the result from simple local slowing:

      "This picture shows that desolvation does more than slow down local chain rearrangements through an added barrier. It also regulates the balance between fluctuation growth, transient arrest, and domain coarsening, thereby shaping the evolution of phase-separated domains."

      Reviewer #2 (Recommendations for the authors):

      (1) The model is physically motivated and useful, but I would encourage the authors to be more precise in describing the added terms as desolvation-inspired effective interactions rather than explicit desolvation.

      We thank the reviewer for the comment and suggestion. We have revised the manuscript accordingly and describe the added terms as desolvation-inspired effective interactions within an implicit-solvent CG framework throughout the Abstract, Results, and Discussion. More detailed changes are provided in our response to the first point raised in Reviewer #2's Public Review.

      (2) The desolvation barrier is introduced as part of the equilibrium pair potential, and therefore it is expected to affect not only kinetics but also the phase boundary through changes in the configurational partition function. The manuscript would benefit from clarifying this point, since the term "barrier" may otherwise suggest a primarily kinetic role. In particular, the authors should explain whether the observed shift in T_c reflects a change in the effective pair attraction, for example, through the integrated Boltzmann weight or second virial coefficient, rather than only an entropic penalty associated with restricted configurations.

      We thank the reviewer for this important point. We agree that the desolvation barrier is part of the equilibrium pair potential and therefore affects the phase boundary through the Boltzmann-weighted sampling of residue-pair configurations, not only through kinetic slowing.

      Following this recommendation, we added a bead-level second virial coefficient analysis based on the effective pair potential. This analysis provides a pair-potentiallevel measure of the integrated effective attraction and clarifies why increasing the barrier lowers T_c, whereas stabilizing the solvent-separated minimum raises T_c.

      Corresponding changes:

      (1) (page 5, lines 203–205) The equilibrium role of the barrier is stated explicitly:

      "At the pair-potential level, the desolvation barrier modifies the equilibrium Boltzmann weight and thereby alters the integrated effective attraction, as quantified by the bead-level second virial coefficient B<sub>2</sub>."

      (2) (page 5, lines 205–207) The barrier-dependent second virial coefficient is connected to the shift in critical temperature:

      "Specifically, increasing ε<sub>b</sub> makes B<sub>2</sub>/σ<sup>3</sup> larger (Figure 2—figure Supplement 1G), indicating a weaker integrated effective attraction and providing a thermodynamic basis for the lower T<sub>c</sub><sup>*</sup>."

      (3) (page 5, lines 207–209) The solvent-separated-well trend is linked to a smaller second virial coefficient:

      "By contrast, deepening the solvent-separated well ε<sub>ss</sub> elevates T<sub>c</sub><sup>*</sup> (Figure 2E), which is associated with the enhanced population of solvent-separated configurations and a smaller B<sub>2</sub>/σ<sup>3</sup> (Figure 2—figure Supplement 1D, H)."

      (4) (page 18, lines 662–665) The Methods specify the integration range and the quantity used to compare integrated effective attraction:

      "The upper limit of integration r<sub>c</sub> is set as 3σ, which is sufficiently large to capture the full range of interactions while ensuring numerical convergence. The reduced value B<sub>2</sub>/σ<sup>3</sup> was used to compare the integrated effective attraction under different desolvation parameters."

      (5) (Figure 2—figure supplement 1G, H) The new panels report the integrated effective attraction:

      "(G, H) Bead-level second virial coefficient (B<sub>2</sub>/σ<sup>3</sup>) calculated from the effective pair potential under varying ε<sub>b</sub> at fixed ε<sub>ss</sub> = 0.02 kcal/mol (G) and varying ε<sub>ss</sub> at fixed ε<sub>b</sub> = 3.12 cal/mol (H)."

      (3) The conformational analysis in Figure 3 is interesting and potentially important. It would help to better place this result in the context of prior work showing dilute-todense phase conformational reorganization of IDPs, and to clarify what is new here beyond that broader observation.

      We thank the reviewer for the comment and suggestion. We have revised the Results section discussing Figure 3 to place dilute-to-dense conformational reorganization of IDPs in the context of previous studies and then to emphasize the specific contribution of the present work.

      The revised text clarifies that, within the present model, desolvation-inspired interactions mainly regulate chain conformations in the dilute phase, whereas dense-phase conformations remain comparatively insensitive. Detailed changes are provided in our response to the second point raised in Reviewer #2's Public Review.

      (4) The proposed scaling between ΔR_g and distance from the critical point is intriguing, but the argument relies on simplifying assumptions. I would present this more as an empirical scaling supported by a plausible theoretical argument rather than a general result.

      We thank the reviewer for this helpful suggestion. We agree that the correlation between Delta R_g and the distance from the critical point relies on simplifying assumptions and should not be presented as a general law. In the revised manuscript, we have softened the interpretation and now present this relationship as an empirical correlation supported by a simplified Flory-Huggins-based theoretical argument.

      Corresponding changes:

      (1) (page 7, lines 254–256) The relationship is described as an approximate trend in the simulated systems:

      "Notably, data from the simulated systems approximately follow a common trend, revealing a strong correlation between the magnitude of conformational change and the thermal distance to the phase transition point (R<sup>2</sup> = 0.942, Figure 3D)."

      (2) (page 7, lines 256–258) The interpretation is limited to an association with thermal distance from the critical point:

      "This result suggests that the conformational response to phase separation is closely associated with how far the system resides thermally from the critical point."

      (3) (page 7, lines 263–264) The critical-condition derivation is shown explicitly:

      "At the critical point, χ(T<sub>c</sub>) = χ<sub>c</sub>, which gives ε<sub>eff</sub> = k<sub>B</sub>T<sub>c</sub>χ<sub>c</sub>. Substituting this relation into the expression for χ(T<sub>sim</sub>) yields χ(T<sub>sim</sub>) = χ<sub>c</sub>T<sub>c</sub>/T<sub>sim</sub>."

      (4) (page 7, lines 268–272) The structural relation is explicitly introduced as a first-order linear-response approximation:

      "The thermodynamic driving force χ(T<sub>sim</sub>) – χ<sub>c</sub> can then be related to the structural observable ΔR<sub>g</sub>. Since ΔR<sub>g</sub> captures the structural transition from an intrachain-interaction-dominated state in the dilute phase to an interchain-interaction-dominated state in the dense phase, we assume, as a first-order approximation, that this conformational shift responds approximately linearly to the excess interaction strength, expressed as ΔR<sub>g</sub> ∝ [χ(T<sub>sim</sub>) − χ<sub>c</sub>]."

      (5) (page 8, lines 285–287) The unscaled relation is labeled as an empirical scaling approximation:

      "Although the complete relation in Equation (3) contains an additional T<sub>sim</sub> factor, the unscaled quantities remain strongly correlated over the simulated range. We therefore use T<sub>c</sub> − T<sub>sim</sub> ∝ ΔR<sub>g</sub> as an empirical scaling approximation."

      (5) The dynamics section would benefit from a statement of what is nontrivial, since a desolvation barrier is expected to slow local rearrangements.

      We thank the reviewer for the comment and suggestions. As described above, we have revised the dynamics section to clarify what is nontrivial beyond the expected slowing of local rearrangements by an added barrier. The revised text emphasizes that desolvation affects condensate dynamics through competing effects of macroscopic packing and microscopic energy-landscape roughness, and that it also regulates transient kinetic arrest and domain-scale coarsening. More detailed changes are provided in our response to the third point raised in Reviewer #2's Public Review.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the editor and the reviewers for their time and efforts to evaluate our manuscript. We have taken into account all the comments and revised the manuscript accordingly which has considerably strengthened the message.

      Several parts of the manuscript, have been extensively rewritten to add explanations and clarify our hypothesis and claims, this also led us to add four new references.

      In addition, we have added the following new figures:

      - New part of figure 2. Figure 2E shows a platelet in a constrained clot after 4h of retraction with the fibrin cage around the platelet center still present. The actin staining of the platelet shows that radial actin fibers are present in each bulb extending to the platelet center. This observation supports our hypothesis that in each bulb an individual cytoskeletal swirling could take place resulting in the accumulation of fibrin fibers at the base of each bulb.

      - Modification of figure 3, to include the criteria used to define four categories of platelets and associated fibers in the 2D fiber retraction assay (new Fig. 3C).

      - New figure 13, illustrating the quantification of fibrin fibre compaction mediated by platelets in the 2D fiber retraction assay and the rotational movement of a fiber mass (video 9).

      - New supplementary figure 1, showing the result of a new model simulation in the absence of cytoskeleton swirling. Under this condition the fibrin fiber does not loop around the platelet bulb.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper reports a previously unrecognized mechanism by which platelets compact fibrin fibers during clot retraction. Rather than simply pulling on fibers, the authors propose that platelets generate swirling motions that wind and loop fibrin into dense structures.

      While the results are intriguing, the underlying physical mechanism remains unexplained. In particular, it is unclear how platelets generate swirling motion capable of inducing fibrin coiling, especially when suspended in 3d fibrin mesh. This raises concerns about the conclusions.

      The reviewer is right, it is difficult to imagine how platelets in a 3D fibrin mesh can accumulate fibers at the base of their extensions to form a cage-like fiber organisation around the center of the platelets. We therefore developed the 2D fibre-retraction assay, which we believe provides important insight for the coiled fiber accumulations above spread platelets in the 2D situation but also provides a framework for interpreting similar processes that may occur within a 3D clot. In response, we have placed greater emphasis on clarifying and strengthening the comparison between the potential mechanistic aspects in the 2D and 3D assays, in order to better support our proposed model (see Results, section: "Platelets, spread on a 2D surface, organize fibers above them", last paragraph). In addition, the Ideas and Speculations section of the discussion has been extensively rewritten to provide more detailed explanations about the potential mechanism leading to fibrin fiber accumulations around platelet bulbs in a 3D fibrin mesh.

      Also, does fibrin have inherent chirality or structural asymmetry that could promote coiling independently of platelet activity?

      Yes, double-stranded fibrin protofibrils have a helical twist [1]. Furthermore, a clot formed in the absence of platelets and other cellular components shows intrinsic tensile forces [2]. However, we show that inhibition of actomyosin actions prevents fibrin fiber accumulation in the 2D fibre-retraction assay providing evidence that platelet actions are necessary to observe the coiled fibers above spread platelets. This has been accentuated in the revised version and three references have been added.

      Furthermore, platelet retraction typically involves platelet aggregation rather than isolated cells, and it is unclear how fibrin coiling would proceed in clustered platelets.

      Under the in vitro fiber retraction conditions used in our study (constrained or unconstrained clots or even in the 2D assay) individual platelets are homogenously distributed within the forming clot or on the coverslip. Therefore, there are no big platelet aggregates or clusters of platelets under our experimental conditions and the results can only demonstrate how individual platelets act on fibrin fibers. This point has been emphasized in the revised version (Discussion, third paragraph).

      Reviewer #2 (Public review):

      Summary:

      Grichine et al. investigate platelet-mediated fibrin compaction using human donor platelets and propose a novel mechanistic model in which platelets generate contractile forces and wind fibrin fibers into compact coiled structures. Using a combination of 2D spread assays, 3D clot imaging via expansion microscopy, live-cell imaging, and computational modelling, the authors present evidence of cage-like fibrin architectures, coiled-fibre morphologies, and platelet centred "rosette" structures present during fibre compaction. They further suggest that actomyosin-driven cytoskeletal dynamics, potentially involving rotational or swirling motion, underlie this proposed winding mechanism, analogous to DNA looping and compaction. The study addresses an important and longstanding question in thrombosis and hemostasis and offers a conceptually novel perspective on clot compaction.

      Strengths:

      The integration of multiple imaging modalities is a notable strength of this paper. In particular, the 2D fiber-retraction assay provides a useful model for understanding the spatio-temporal dynamics of platelet-mediated fibrin compaction, which can be applied to other systems and may yield detailed mechanistic insights into biological processes. The live-imaging approaches are particularly well executed and offer valuable dynamic insight.

      Weaknesses:

      The primary weakness of this paper lies in its descriptive nature and its reliance on correlative rather than causal evidence. Several interpretations are not uniquely supported by the data presented. For example, the categorisation of fibrin accumulation in 2D assays as "fiber winding" and "fibre compaction" remains descriptive without establishing winding as a mechanism.

      When introducing the 2D fiber-retraction assay (figure 3) in the revised version, we now only mention the terms fiber accumulation and compaction to better align with the level of evidence, since wound-up fibers cannot be distinguished in this figure. The criteria to establish the four categories of platelets and associated fibers in the 2D fiber retraction assay have now been included in figure 3C.

      Nevertheless, coiled fibers above spread platelets are clearly visible in figure 4 and 8 and dynamic fiber rotations or winding-up are observed in figure 12 and video 9. These observations have been presented more cautiously, as indicative rather than definitive evidence of a winding mechanism.

      Alternative mechanisms, such as circular bundling, stacked fibers under tension, or fibrin crosslinking-induced aggregation, are neither excluded nor investigated.

      For fibrin fiber bundling, staggered or crosslinked protofilaments no platelet actions are necessary as described previously [2,3]. Since we observed a clear difference between +/- blebbistatin conditions in the 2D fiber-retraction assay, the fiber compaction we observe depends on platelet actions. Consequently, we consider these alternative mechanisms unlikely based on our data. This has been stated explicitly in the results section and discussion and three references have been added.

      Although the authors present compelling live imaging, establishing winding as a dynamic phenotype would require quantitative analyses, such as measuring angular velocities and coiling rates.

      We have incorporated quantitative measurements (new figure 13) about platelet mediated fibrin fiber compaction and angular rotation velocities to complement the observations obtained from live imaging. It is important to note, however, that angular velocities and coiling rates are likely influenced by the number of fiber–fiber contacts present at the time coiling occurs. Specifically, an increased number of contacts is expected to elevate tension within the network, thereby modulating the forces generated by platelets and, consequently, affecting both velocity and coiling dynamics.

      The use of a second fluorophore-labelled fibrin population could further strengthen evidence for rotational dynamics.

      These live videos are quite difficult to acquire because of the following reasons:

      - Small platelet size

      - Heterogeneity of platelets within the population (10 d half-life, old platelets may not be able to compact fibers efficiently).

      - The speed of the process and the time needed to adjust parameters for image acquisition, necessitates an arbitrary choice of the acquisition window and only one acquisition (90 min) per sample preparation is possible.

      - Furthermore, the laser-induced illumination can perturb the observed processes. We therefore use high-spatial-resolution 3D confocal time-lapse imaging, performed in photon-counting mode with very low laser excitation.

      For these reasons, the use of additional markers would be technically challenging and could perturb the delicate equilibrium and dynamics of the process under investigation.

      Similarly, the inference of rotational contractility or actomyosin "swirling", based on chiral actin organisation and blebbistatin treatment, is not sufficiently supported to conclude that platelets actively wind or loop fibrin fibers.

      Importantly, in the 2D fiber-retraction assay, we do not propose that the rotational actomyosin activity leads to a contractility of the platelets which would allow fiber retraction. Rather, we suggest that cytoskeletal actomyosin swirling (as demonstrated for nucleated cells by Bershadsky's team) can induce rotational dragging of extracellular bound fibrin fibers around the pseudonucleus of spread platelets thereby promoting accumulation of fibrin fibers (shown in figure 12C, video 9, third panel). Consistent with this interpretation, inhibition of myosin by blebbistatin prevents the accumulation of fibrin fibers above spread platelets in the 2D fibre retraction assay (Fig. 3).

      The mathematical model, while complementary and well-constructed, relies on multiple assumptions and lacks predictive validation.

      We thank the reviewer for this insightful comment and acknowledge that the proposed model relies on several important assumptions. In our view, the most significant assumption is that integrin molecules undergo rotational downstream motion as a consequence of their coupling to the swirling cytoskeleton. To assess the necessity and impact of this assumption, we provide an additional simulation performed in absence of the cytoskeletal swirling. Under this condition the fibrin fibers are not looped around the platelet bulb (this result has been added as supplementary figure 1). This analyses also provides further validation of the proposed model and underlying mechanism. At the same time, it is important to emphasize that the primary purpose of the model was to examine whether the hypothetical swirling dynamics of the cytoskeleton, together with the associated receptors, could in principle reproduce the experimentally observed fibrin organization.

      Appraisal:

      While the authors successfully document intriguing fibrin architectures and provide a compelling descriptive framework, they do not fully demonstrate a mechanistic model of active fibrin winding by platelets. The conclusions regarding platelet-driven winding and rotational dynamics are not sufficiently supported by direct or quantitative evidence. To substantiate these claims, the study would benefit from experiments that directly link platelet dynamics to fibrin organisation, including coordinated measurements of platelet motion and fibre rearrangement. As it stands, the results are suggestive but do not definitively support the proposed mechanism.

      Discussion and Impact:

      Despite these limitations, the study addresses an important question in thrombosis and hemostasis and introduces a potentially impactful conceptual framework for understanding clot compaction. The imaging approaches and datasets presented will be valuable to the community, particularly for researchers interested in platelet mechanics and fibrin organisation. However, the overall impact will depend on whether the proposed mechanism can be more rigorously validated. In its current form, the study presents an interesting and thought-provoking model, but would benefit from either stronger experimental support for the proposed mechanisms or a more cautious interpretation of the findings.

      We agree that the proposed mechanism requires further validation. In the revised version we have added a new result (figure 2E) showing that radial actin filaments are present in each bulb of a platelet in a constrained clot, supporting the possibility that rotational cytoskeletal movements could take place in individual bulbs. In a new figure 13, we have also quantified fiber compaction and the angular velocity of a rotating fibrin mass observed in video 9. Furthermore, in the revised manuscript, we present a more cautious and explicitly hypothesis-driven interpretation of the mechanism. We hope that the publication of our observations will be of interest to researchers in the field of thrombosis and clot mechanics who possess the specialized tools and expertise necessary to rigorously evaluate and either substantiate or refute the proposed mechanistic model.

      Reviewer #3 (Public review):

      Summary:

      This work aims to understand the mechanisms that platelets use to interact with and compact fibrin fibers during clot formation. This is an important process during wound healing, and recent work has demonstrated that platelets play a critical role in generating the force required to drive the accumulation of fibrin. The authors argue that current models are insufficient to account for the observed reduction in clot volume and propose that platelets actively 'wind up' these fibers by undergoing myosin-dependent rotation. While interesting, the experiments performed by the authors do not directly test this mechanism, and further evidence is required to support their claims.

      We do not "propose that platelets actively 'wind up' these fibers by undergoing myosin-independent rotation" of the whole platelet, but rather of the cytoskeleton winding-up extracellular fibrin fibers attached to integrin receptors.

      Weaknesses:

      (1) The motivation to switch from the system used in Figures 1 and 2 to the '2D fiber-retraction assay' is not clear. While the authors state that this system has 'reduced complexity', the differences between these assays appear to disrupt the 'cage-like' organization of fibrin around platelets shown in Figures 1 and 2 (compare images in Figure 2 with those in Figure 4). An indepth comparison of two methods is needed to support the conclusions from the 2D system.

      We agree that the cage-like fibrin organization around platelets is disrupted in the 2D fibre-retraction assay when platelets are completely spread on the coverslip before they have encountered fibrin fibers (Fig. 4). This has been explicitly stated in the revised version. However, some platelets in the 2D fiber-retraction assay form the same number of extensions as platelets in a 3D clot (Fig. 9 A, B) and are not completely spread on the glass surface. For these platelets a cage-like fibrin organisation is retained under the 2D conditions (Fig. 5 and 6). Nevertheless, the fiber density at the base of the bulbs is higher in the 2D assay than under the constrained 3D clot retraction conditions (Fig. 1C and Fig. 2), probably because in the 2D condition the fibers are less constrained and readily available for compaction.

      Furthermore, the change in plasma volume (Figure 2 vs Figure 7) should also be tested - the authors state that this increases fibrin fiber formation, but this is not quantified or demonstrated in the figures. Notably, this appears to change the morphology of the fibrin fibers shown (comparing Figure 2 and Figure 7).

      We thank the reviewer for raising this point. We would like to clarify that Figure 2 and Figure 7 correspond to two distinct experimental setups: the constrained clot retraction assay (Figure 2) and the 2D fiber-retraction assay (Figure 7). As such, they are not directly comparable. We understand, however, that the reviewer is likely referring to the apparent differences between Figures 3–6 (lower plasma volume, higher fiber density) and Figures 7–8 (higher plasma volume, lower apparent fiber density).

      The reduced number of visible fibers in the latter condition is not solely a consequence of plasma volume per se, but rather results from the formation of a labile fibrin gel at higher plasma concentrations, which is lost during the fixation and aspiration steps. This effect was initially observed across samples from two donors with differing plasma fibrinogen levels. In one case, an unusually low fibrinogen concentration allowed the addition of higher plasma volumes without inducing gel formation. In contrast, in the other sample, a more typical fibrinogen level resulted in gel formation under the same conditions.

      Importantly, we performed all experiments using matched donor plasma and platelets. As a result, the precise fibrinogen concentration could not be determined prior to experimentation. Nonetheless, post hoc measurements confirmed that fibrinogen levels in most donor samples fell within the normal physiological range, which allowed us to always use the same plasma volumes for low and high plasma concentrations (4ul/ml PBS and 7 ul/ml PBS, respectively) except for one donor as mentioned above.

      (2) It is unclear how the classification of platelets as 'fiber-winding' versus 'fiber compaction' differs in Figure 2. The criteria used for these classifications should be stated. Further, it seems premature to characterize fibers as wound without having established this earlier in the manuscript.

      The reviewer probably refers to figure 3 and he is right; it is premature to mention fiber winding at this stage of the results section (see our response to reviewer #2). In the revised version, we have modified figure 3 to include the criteria used to classify the platelets into four different categories (Fig. 3C).

      (3) Is the 'gearwheel' different from the 'cage' of fibrin fibers? They appear similar, but it is difficult to distinguish between them with only qualitative descriptions of these phenotypes.

      The "gearwheel" is observed for completely spread platelets in the 2D fiber-retraction assay and a figure illustrating our hypothetical speculations to compare the 2D gearwheel with the 3D clot situation is presented in the discussion under the "Ideas and Speculations" paragraph (now Fig. 14). We have given a more comprehensive explanation of the proposed mechanism in the revised version.

      (4) The quantification of platelet extensions in Figure 9 is confusing. While those in 9A are clear, those in 9B are not. For instance, what is the difference between #7 and #8 in the middle panel of 9B? It does not seem like #8 is labeling an extension.

      For the platelet shown in the middle panel of Figure 9B, the extensions cannot be clearly distinguished in the MIP (Maximum Intensity Projection) image because extension #8 is positioned above extension #7 and is therefore superimposed in the projection. However, the two extensions can be differentiated when examining the 3D image stack (Video 4, upper panel). As indicated in the figure legend, the number of extensions was determined manually by scrolling through the z-stack image sequence. In the revised version, we will also define the abbreviation “MIP” as Maximum Intensity Projection.

      (5) It is unclear what the modeling accomplishes, as there is no comparison between the results of these simulations and their experiments.

      We thank the reviewer for this valuable concern. We chose not to combine the experimental fibrin organization and the modeling results within the same figure panel, as the resulting image would be too complex and difficult to interpret. We have, however, added a supplementary figure 3 showing the results of a new simulation in the absence of cytoskeletal swirling. Under these conditions no winding of the fibrin fiber around the platelet bulb can be observed. It is also important to emphasize that the comparison between the model and the experimental data was intended to be primarily qualitative rather than quantitative.

      (6) The data presented in Figure 12 provides the most direct support for their mechanism, but falls short of directly testing their claims. These experiments should be repeated to include blebbistatin to test the contribution of myosin and include quantitative rather than qualitative comparisons of these experiments.

      As mentioned already above, these live videos are quite tricky to acquire because of the following reasons: - small platelet size

      - Heterogeneity of platelets within the population (10 d half-life, old platelets may not be able to compact fibers efficiently).

      - The speed of the process and the time required to optimize imaging parameters, necessitate the selection of an arbitrary acquisition window. Consequently, only a single acquisition of approximately 90 min can be performed per sample preparation, with no guarantee that relevant platelet-fibrin interactions can be acquired in the acquisition window.

      - Furthermore, after blood donation, the first sample is usually ready to be acquired around 3 pm, acquisition time 90 min. At least 10 successful acquisitions per condition would be required to ensure statistical robustness, but maximal 4 can be acquired per donor, because platelet samples start to deteriorate within twelve hours after blood donation.

      Taken together, the intrinsic heterogeneity of the platelet population, the low likelihood of capturing informative events, and the limited availability of suitable imaging resources at our institute render a robust and quantitative comparison between conditions with and without blebbistatin extremely challenging, if not impractical, within a reasonable timeframe.

      In accordance with the reviewer's request, we have added a new figure 13 to the revised version, presenting quantitative data on the platelet-mediated fibre compactions and the speed of angular fibrin rotations observed in video 9.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Throughout the manuscript, it is difficult to map the data presented in the figures to the text in the results section. Often, many subpanels are referred to collectively (for example, 'Fig 4 AE and animation, Video 3' on line 150), and the reader is left to piece together how this data fits into the statements in the results section. More guidance from the authors would help to understand the connection between these data and their conclusions.

      In the revised version, we have provided clearer explanations to make it easier to understand the conclusions drawn from the data. Concerning the indication "Fig 4 A-E and animation, Video 3" just means that platelets shown in panels A-E of figure 4 can also be visualized in the animation video 3. We have also put an effort to clearly indicate which figure part is presented in the associated video.

      There are also many figures that contain redundant information. The authors should consider revising these figures and including some of these repeated images as supplemental figures.

      As noted by the reviewers, our study provides predominantly qualitative observations essentially because it is not obvious to choose parameters which would be pertinent and could be quantified accurately using expansion microscopy. A quantitative analysis would allow to show the quantification and a representative image to describe the phenotypes of platelet-mediated fibre organisations. Without a quantitative analysis, we consider it more appropriate to provide multiple examples, enabling the reader to assess the consistency as well as the variability across repeated observations.

      Additional References

      (1) Jansen KA, Zhmurov A, Vos BE, et al. Molecular packing structure of fibrin fibers resolved by X-ray scattering and molecular modeling. Soft Matter. 2020;16(35):8272-8283.

      (2) Spiewak R, Gosselin A, Merinov D, et al. Biomechanical origins of inherent tension in fibrin networks. J Mech Behav Biomed Mater. 2022;133:105328.

      (3) Ramanujam RK, Lavi Y, Poole LG, Bassani JL, Tutwiler V. Understanding blood clot mechanical stability: the role of factor XIIIa-mediated fibrin crosslinking in rupture resistance. Res Pract Thromb Haemost. 2025;9(4):102871.

      (4) Gaertner F, Ahmad Z, Rosenberger G, et al. Migrating Platelets Are Mechano-scavengers that Collect and Bundle Bacteria. Cell. 2017;171(6):1368-1382 e1323.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public review:

      Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure. To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to neural changes across sessions. Repeated retrieval can induce memory updating (reconsolidation) or extinction, the latter involving the formation of a new association between the CS+ and safety. Although memory updating may have contributed to the ensemble reorganization observed here, several aspects of our findings argue against extinction as the primary explanation for the observed neural changes.

      First, we observed substantial neuronal ensemble turnover beginning with the first retrieval session. This early turnover is consistent with previous observations in the prefrontal cortex [1, 2] and with growing evidence that cortical memory representations remain dynamic throughout systems consolidation [3, 4]. Longitudinal studies have shown that neurons are continuously recruited into and removed from cortical memory ensembles while memory expression remains stable [1-4].

      Second, we calculated discrimination ratios to quantify discrimination of each tone relative to the CS+ across retrieval sessions (Figure S1). These analyses showed that discrimination increased, rather than decreased, over successive retrieval sessions, a pattern inconsistent with the behavioral profile expected if repeated testing had induced extinction.

      Finally, one of the most novel findings of our study is that ensemble turnover does not affect all neuronal populations equally. The graded neurons identified by our clustering analysis maintained their identity and functional organization across retrieval sessions, and their activity was better explained by tone threat value than by freezing behavior (Figure 8). This selective stability indicates that ensemble reorganization is not a uniform process but instead preferentially affects specific neuronal subpopulations while preserving a stable threat-value generalization gradient. Thus, although repeated retrieval may contribute to ongoing ensemble reorganization, our results demonstrate that this process is selective and largely spares the neuronal subpopulations that encode graded threat-value representations.

      Accordingly, we have revised the Discussion to explicitly acknowledge these points as follows:

      “The ensemble turnover observed here is consistent with previous studies demonstrating dynamic reorganization of cortical activity patterns over time [1-3, 5]. Such reorganization has been proposed to provide flexibility by allowing new information to be incorporated into existing cortical representations while preserving stable behavioral performance [4, 6]. Several mechanisms could contribute to this turnover, including systems consolidation, retrieval-induced reconsolidation or memory updating, and repeated nonreinforced stimulus exposure [4, 6-8]. Although our experiments cannot distinguish between the first two possibilities, the behavioral data argue against extinction as the primary explanation. Extinction is generally associated with the formation of new CS+-safety associations [9], whereas discrimination ratios increased across retrieval sessions, indicating that animals progressively improved their discrimination between threat-associated and safe stimuli rather than acquiring generalized safety responses. This pattern is consistent with previous work showing that discrimination learning sharpens stimulus representations and narrows behavioral generalization gradients [10-13]. Importantly, turnover was not uniform across the population. Graded neurons retained remarkably consistent response profiles across retrieval sessions, and their activity remained more strongly associated with learned threat value than with freezing behavior. These observations indicate that stable components of the population code can coexist with extensive reorganization of surrounding neuronal ensembles.” Pg. 19

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      This is an important point, which was also highlighted by the other reviewers. To directly address this concern, we implemented the generalized linear model (GLM) analysis suggested by Reviewer 3. We modeled the neuronal activity time series using both tone identity and freezing behavior as simultaneous predictors. Because tone identity was fixed across trials whereas freezing varied from trial to trial, the GLM allowed us to dissociate their independent contributions to neuronal activity.

      As described in the original submission, freezing was estimated from the miniscope's onboard inertial measurement unit (IMU), which measures body acceleration along three axes. Rather than classifying freezing using a fixed threshold, we estimated the continuous probability of freezing from the accelerometer signal using a Gaussian mixture model. This probabilistic estimate was incorporated directly into the GLM together with tone identity, providing a conservative test of whether neuronal activity was better explained by freezing behavior or by the auditory stimulus.

      We applied the GLM both to all sound-responsive neurons contributing to the population response curves (Figure 4) and to the graded and frequency-selective neuronal subpopulations identified by our clustering analysis (Figure 8). Across both experimental groups and all analyses, the median regression coefficients (β) associated with tone identity were consistently larger than those associated with freezing, indicating that tone identity contributed more strongly to neuronal activity. Moreover, tone coefficients exhibited graded monotonic profiles that closely tracked the learned threat value of each tone, with graded neurons showing the strongest gradients (Figures 4a, 8a, and 8e). Consistent with previous reports [14, 15] freezing accounted for a modest but significant component of PL activity. However, only 6–8% of graded neurons were classified as freezing-dominant, indicating that for the vast majority of these neurons, tone identity was the stronger predictor. Together, these findings demonstrate that the graded representation of learned threat value persists after accounting for freezing behavior, supporting our conclusion that PL activity reflects learned threat value rather than merely the behavioral expression of fear.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      We thank the reviewer for this suggestion. We now include measures of registration quality in the resubmission. Specifically, we calculated shifts in centroid distances, proportion of ROIs retained across all sessions, and representative examples of matched imaging fields over time (Fig, S3).

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      We corrected correspondence between text and Figure.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      We clarified the labelling of the Figure 2a and call the graphs “activity-plots”.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      For Figure 3, we maintained the previous one-way ANOVAs assessing changes in AUC per day to be able to note significance on the Figure panels. However, we added a three-way mixed-effects analysis, using group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30) as variables, with frequency and day of testing as repeated measures. The results were described as follows (statistical details Table S1):

      “To determine how AUC varied across groups over time, we performed a three-way mixed-effects ANOVA with group (CS+15, CS+3, and no shock), frequency (3, 7, 11, and 15 kHz), and time (test days 1, 15, and 30) as factors, with repeated measures on frequency and time. For positive responder neurons, the analysis revealed significant main effects of group (p < 0.001) and time (p < 0.05), as well as a significant group × frequency interaction (p < 0.001), whereas the time × frequency and group × time × frequency interactions were not significant (p > 0.05; Table S2a). Tukey-corrected post hoc comparisons showed that, in the CS+15 group, AUC differed between all frequency pairs except 11 and 15 kHz (p < 0.05). In the CS+3 group, the AUC at 3 kHz differed from those at 7, 11, and 15 kHz (p < 0.05), whereas no significant frequency differences were observed in the no-shock controls (p > 0.05). For negative responder neurons, the only significant effect was a time × frequency interaction (p < 0.01). Tukey-corrected simple-effects analyses revealed that, on day 30, the AUC at 15 kHz differed from those at 3, 7, and 11 kHz (p < 0.05; Table S2b). Because this pattern was observed across all experimental groups, including the no-shock controls, it is unlikely to reflect associative learning. These results indicate that although the AUC exhibited modest changes over time, these changes were not group-specific and therefore do not support learning-dependent alterations in neuronal responses. Together, these results show that despite substantial neuronal turnover, PL population responses encode generalization gradients, closely matching behavioral expression.” Pg. 9

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term inference may overstate the cognitive processes engaged by the current task. Accordingly, we revised the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli. The new GLM analyses further support this interpretation by demonstrating that, in the conditioned groups, neuronal activity at both the population and single-neuron levels is explained substantially better by tone identity than by freezing behavior (Figures 4 and 8). We therefore retained the term threat value, as our results indicate that PL activity primarily reflects learned threat value rather than simply the expression of freezing behavior, but removed inference.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

      We corrected the language and replaced valence for “threat value”

      Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      We thank the reviewer for this thoughtful comment. We agree that our original wording may have implied that turnover was a spontaneous process. Repeated retrieval provides opportunities for updating the learned contingencies associated with both the conditioned and generalization stimuli, and therefore changes in ensemble composition across sessions need not arise independently of experience. We also agree that the stability of graded neuronal representations may be related to the animal's certainty about the learned contingencies. However, in our data the graded neuronal population remained remarkably stable across retrieval sessions, whereas changes occurred primarily within the dynamic, frequency-selective neuronal populations. This suggests that stable ensembles preserve representations of learned threat value while updating is concentrated in a distinct neuronal subpopulation. We have now incorporated these ideas into the Discussion.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      We thank the reviewer for these specific criticisms. We revised the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.

      (a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      (a) We hypothesized that the PL generates representations of learned threat value that support threat generalization and discrimination, and that these representations emerge from the coordinated activity of stable and dynamic neuronal subnetworks, preserving consistent relationships among stimuli despite ongoing cellular turnover.

      (b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      (b) The summary statement was rewritten: " Together, these findings provide a neural framework for understanding how the PL supports adaptive threat generalization and discrimination.” pg. 4

      (c) 'In CS<sup>+</sup> 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS<sup>+</sup>15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      We adopted the reviewer's suggested rewording: " In CS<sup>+</sup> 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting learned contingency value across testing days" pg. 9

      We will systematically review the entire manuscript to ensure consistency with this revised framing.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we followed Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using the time series activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis allowed us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we concluded that tone identity was a stronger predictor of neuronal activity than freezing. These results are summarized in Figures 4 for all cells contributing to population responses and Figure 8 for the main neuron types identified in the clustering analysis (frequency-selective and graded neurons). All details of this extensive new analysis are shown in red in the revised resubmission.

      In addition, we revised the text to remove the terminology of “learned and inferred valence” throughout the manuscript.

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      We thank the reviewer for this helpful comment. We agree that our use of the term valence in describing the no-shock controls was imprecise. Because none of the tones was associated with reinforcement in this group, there was no learned valence that could modulate neuronal activity. Our intention was simply to convey that, although both positive and negative sound-responsive neurons were present, the population responses did not vary systematically across tone frequencies. We have revised this section accordingly.

      We also agree that our original wording overstated the interpretation of the graded population responses. Our data do not demonstrate that associative learning is required for sound responsiveness itself; rather, they show that associative learning is required for the emergence of graded population responses that distinguish tones according to their learned threat value. We have revised the text to make this distinction explicit.

      Finally, we agree that our previous references to "learning and inference" were not justified by the behavioral paradigm. We have removed this language throughout the manuscript and now describe the findings more directly as graded representations of learned threat value that closely parallel the observed behavioral generalization gradients.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      We thank the reviewer for this thoughtful comment. We agree that repeated retrieval is an inherent limitation of longitudinal memory studies and that repeated non-reinforced presentations of the tones provide opportunities for memory updating. Accordingly, we have revised the Discussion to explicitly acknowledge that repeated retrieval may contribute to the ensemble reorganization observed across sessions through memory updating or reconsolidation processes (Discussion, pg. 19).

      We also agree that the progressive sharpening of the behavioral generalization gradients across retrieval sessions is consistent with memory updating. Both the behavioral data (increased discrimination ratios) and the neuronal data (progressively sharper population generalization gradients among neurons active after conditioning) indicate that the memory representation became more precise over time. We now discuss this possibility explicitly in the revised Discussion. We also agree that fear generalization often broadens with time; however, this is not universal. Under discriminative conditioning paradigms, repeated retrieval can instead produce progressively narrower generalization gradients [11]. We have revised the Discussion to clarify this distinction and added the appropriate references (pg. 19).

      While the reviewer's interpretation is therefore plausible, we do not believe it fully accounts for our observations. If repeated non-reinforced presentations were the sole driver of the observed neuronal changes, one might expect a more uniform reorganization across the neuronal populations engaged by the task. Instead, the reorganization was highly selective. Neurons encoding graded threat value remained remarkably stable across retrieval sessions, whereas neuronal turnover occurred primarily within the frequency-selective subpopulations. Thus, although repeated retrieval may update the memory representation, the neuronal substrate supporting graded threat-value coding is largely preserved while refinement occurs within a distinct neuronal subpopulation.

      Moreover, we observed substantial neuronal turnover beginning with the first retrieval session, consistent with previous longitudinal studies showing that cortical memory ensembles remain dynamic despite stable memory [1-4]. This early emergence of turnover suggests that repeated testing alone is unlikely to account for the continuous population dynamics observed throughout the experiment.

      Rather than viewing these findings as evidence exclusively for either memory updating or systems consolidation, we believe they are more consistent with current models proposing that these processes occur in parallel. Several influential frameworks argue that memories are continuously modified through retrieval while simultaneously undergoing systems-level reorganization [4, 6, 16, 17].We have therefore revised the Discussion to interpret the longitudinal changes more conservatively as reflecting the combined influence of retrieval-dependent memory updating and systems-level reorganization.

      In summary, we have revised the manuscript to better acknowledge the contribution of repeated retrieval while emphasizing what we believe is the principal finding of our study: despite substantial turnover within the overall ensemble, the neuronal population encoding graded threat value remained remarkably stable, whereas refinement occurred primarily within dynamic frequency-selective neuronal populations.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS<sup>+</sup>. In CS<sup>+</sup>15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      We agree that, because population similarity is highest for the 3/3, 15/15, and 15/11 tone pairs and freezing is also greatest for these same stimuli, the neural data could, in principle, reflect a correlate of behavioral expression rather than an independent representation of learned threat value.

      To directly address this possibility, we implemented a generalized linear model (GLM) to dissociate the contributions of tone identity and freezing behavior to neuronal activity. Across all analyses, tone identity consistently explained substantially more variance in neuronal activity than freezing behavior. Importantly, this finding held not only for the full population of sound-responsive neurons used to generate the population similarity analyses (Figure 4), but also for both the stable graded neurons and the dynamic tone-selective neuronal populations identified by our clustering analysis (Figure 8). Thus, although freezing behavior contributes modestly to PL activity, it cannot account for the enhanced similarity of population vectors across stimulus presentations or the graded population responses that form the basis of our conclusions.

      In addition, the temporal dynamics of the population vector similarity analysis are not entirely consistent with the interpretation that PL activity simply reflects the expression of freezing behavior. Population vector similarity peaked during the first 5 seconds following tone onset, whereas freezing occurred intermittently throughout the tone presentations. Although this temporal relationship does not establish causality, it is consistent with the interpretation that PL activity reflects the learned threat value associated with each tone rather than merely tracking the magnitude of freezing.

      Finally, we have revised the manuscript to more clearly acknowledge the correlational nature of these analyses. Specifically, we now state that population vector similarity is associated with, rather than determines, the degree of threat generalization.

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      We made this correction. (“These findings indicate that population-level similarity at stimulus onset scales with behavioral threat generalization”. pg. 13)

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      We agree that the phrase "integrated learned valence" is unnecessarily opaque and we replaced it with more precise language “Our previous analyses demonstrated that threat-value generalization gradients are represented at the population level. However, these findings do not reveal how these representations arise. Specifically, the observed population gradients could emerge either from the pooled activity of frequency-selective neurons that respond to individual tones or from neuronal subnetworks that integrate information across tones to encode their learned threat-value.” (Pg. 13)

      (8) Another example of what has been a common theme in this review:

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      We thank the reviewer for pointing out that this section was unclear. We agree that our original wording was imprecise and could be interpreted as implying cognitive processes that were not directly tested in the present study. Accordingly, we have revised the terminology throughout the manuscript. We no longer refer to "inferred emotional valence" or "core memory content" and instead describe these neurons more specifically as exhibiting graded representations of learned threat value.

      This is not the interpretation we intended. To determine whether these neurons primarily reflected defensive behavior rather than learned stimulus value, we implemented a Generalized Linear Model (GLM) that dissociates the contributions of tone identity and freezing behavior to neuronal activity. Across the entire neuronal population, as well as within the stable graded and dynamic tone-selective neuronal subpopulations, tone identity consistently explained substantially more variance than freezing behavior (Figures 4 and 8). Furthermore, after accounting for freezing, the regression coefficients of the graded neurons continued to follow the learned threat value of the tones, exhibiting opposite monotonic gradients in the CS+15 and CS+3 groups. If these neurons simply tracked defensive behavior irrespective of the stimulus presented, this relationship would not be expected to persist after accounting for freezing. We therefore conclude that the activity of this stable neuronal subpopulation is better explained by graded representations of learned threat value than by defensive behavior alone, and we have revised the manuscript accordingly.

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      We thank the reviewer for highlighting that this section was unclear. We agree that the original phrasing was insufficiently precise. Our intention was to convey that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. This observation led us to hypothesize that neurons recruited into the active population after conditioning— likely dynamic, frequency-selective neurons—also contribute to these graded population responses through modulation of their firing rates.

      The reviewer correctly notes that the phrase "shaped by associative processes" was too vague. By this we meant that the firing properties of these neurons are modified by the animal's associative history, including both the original conditioning experience and any retrieval-dependent updating that may occur during subsequent test sessions. We have revised the manuscript to make this interpretation explicit rather than leaving it to the reader to infer.

      To test this hypothesis, the GLM analysis we implemented dissociated the contributions of tone identity and freezing behavior to neuronal activity. After accounting for freezing, tone identity (i.e., learned threat value) remained a significant predictor of neuronal responses. Importantly, this was also true for the dynamic, frequency-selective neurons (Fig. 8e–f), indicating that these neurons contribute to population-level representations of learned threat value through firing-rate modulation rather than simply reflecting defensive behavior.

      To clarify our interpretation, we have rewritten the relevant section as follows:

      "Graded clusters encode generalization gradients but constitute only a subset of the active neuronal population. Nevertheless, population-level representations, which incorporate all active neurons, remain robust and accurately preserve these gradients. This observation led us to hypothesize that neurons recruited over time (e.g., dynamic, frequency-selective cells) also contribute to threat-value representations. Consistent with findings in the hippocampus showing that neurons can encode task contingencies through firing-rate modulation despite responding selectively to a single location (Gagliardi et al., 2024; Huxter et al., 2003; Sanders et al., 2019), we tested whether dynamic, frequency-selective clusters exhibited firing-rate differences proportional to learned threat value." (page 15)

      Regarding the reviewer's suggestion that the characteristics of the newly recruited neurons may reflect learning during repeated non-reinforced test sessions, we agree that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population, and we now explicitly acknowledge this possibility in the Discussion. However, we do not believe that our findings are fully explained by repeated non-reinforced retrieval alone. First, no-shock control animals underwent the same repeated testing but failed to develop graded neuronal representations, indicating that repeated exposure in the absence of associative learning is insufficient to account for the observed changes. Second, both behavioral discrimination and the corresponding population-level neural gradients became progressively sharper over time, consistent with refinement of learned threat representations rather than an effect of repeated testing alone, which must lead to extinction.

      In summary, we thank the reviewer for highlighting both the ambiguity of our original wording and an important alternative interpretation. In response, we have clarified the text to explicitly define what we mean by associative processes, added a GLM analysis demonstrating that the newly recruited neurons encode learned threat value beyond freezing behavior, and revised the Discussion to acknowledge that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population.

      (10) The following points all relate to the Discussion and reiterate many of the points above.

      (a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      We modified the language as stated in the prior points.

      (b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      We incorporated the new GLM analysis to address this point and conclusions.

      (c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      We agree that the term "graded representational axis" was insufficiently defined and could be interpreted in multiple ways. Because this terminology was not essential to our conclusions, we have removed it and instead describe the observed phenomenon as a graded population representation of learned threat value across tone frequencies. We also removed the statement suggesting that recurrent connectivity provides a stable scaffold for these representations, as this mechanistic interpretation is not directly supported by our data.

      We also agree that some sections of the manuscript overstated the scope of our conclusions and have revised the wording accordingly. Our study uses neuronal activity recorded during memory retrieval after learning, an approach widely used in studies of systems consolidation to infer how learned information is represented within neural populations. Accordingly, we have revised the manuscript to explicitly state that our findings pertain to neural representations of learned threat value during memory retrieval rather than the broader content of emotional memories.

      Finally, we agree that our data are correlational and do not establish the causal role of the neuronal representations we identify. Throughout the manuscript, we now refer more precisely to population- and single-neuron correlates of learned threat value during memory retrieval following auditory fear conditioning.

      (d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

      We thank the reviewer for this thoughtful comment. Our interpretation that these neurons reflect preexisting sensory-driven properties of PL cortex is based on two observations. First, tone-selective neuronal clusters were present in both conditioned and no-shock control animals, consistent with previous reports of sensory responsiveness in PL cortex [18, 19]. Second, these responses were already present during the first retrieval session, when the intermediate frequencies were presented for the first time. Thus, they cannot be explained by repeated exposure to those tones across subsequent test sessions.

      We therefore interpret the frequency-selective response properties as pre-existing features of PL circuitry that are present independently of conditioning. In contrast, associative learning modifies the firing activity of these neurons, allowing them to contribute to graded representations of learned threat value. This interpretation is supported by our GLM analysis, which showed that, after accounting for freezing, tone identity significantly predicted the activity of frequency-selective neurons in conditioned animals but not in no-shock controls. Thus, while the frequency-selective response properties are present independently of learning, associative learning modifies how these neurons encode learned threat value. We have revised the manuscript to clarify this distinction.

      Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded neuronal responses may partially reflect freezing-related activity rather than representations of learned threat value. In the revised manuscript, we acknowledge that previous studies have identified PL neurons whose activity tracks freezing independently of stimulus identity or associative content. To directly address this possibility, we implemented the reviewer's suggestion by fitting a generalized linear model (GLM) to the neuronal activity time series derived from the Ca<sup>2+</sup> signals, using tone identity and freezing behavior as predictors. Because tone identity is fixed across trials, whereas freezing varies both during tone presentation and across trials (see below our answer to the Recommendations to Authors), this approach allowed us to dissociate their respective contributions to neuronal activity. We are grateful for this excellent suggestion, which has substantially strengthened both the manuscript and the conclusions that can be drawn from our data. The new analyses are summarized in Figures 4 and 8.

      In the points below we summarize the new findings.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

      We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such as combinations of optogenetic and two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.

      We examined inter-individual variability in freezing relative to the proportion of graded cells but the number of mice used in this study (CS+3= 5 and CS+15=7) did not give us enough power to reach significance.

      Therefore, we modified the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the correlational nature of our conclusions.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Many sections of the paper should be rewritten along the lines that I have suggested in my public review and below.

      Minor Comments:

      (1) INTRO - 'This broad accessibility reduces spatial specificity and increases learning variability...'.

      Broad accessibility of what, exactly? And how does the 'broad accessibility' reduce spatial specificity and increase learning variability? That is, I do not understand what the terms 'reduced spatial specificity' and 'increased learning variability' refer to at this point in the first paragraph...

      We rewrote the introduction and discussion to address the points raised by the reviewer.

      (1) INTRO - 'The prelimbic cortex (PL) contributes to the expression (Burgos-Robles et al., 2009; SierraMercado et al., 2011; Sotres-Bayon & Quirk, 2010) and the proper discrimination and generalization of threat memories (Rosas-Vidal et al., 2025; Stujenske et al., 2022).'

      What is achieved by calling it 'proper' discrimination and generalization? Can't one simply say that the PL contributes to the expression, discrimination, and generalization of threat memories?

      This was corrected.

      (3) INTRO - '... and that population-level firing rate similarity across stimulus presentations determines threat generalization'.

      Or, alternatively, that generalization of conditioned freezing responses from the tone CS to variants along the dimension of Hz values is reflected in systematic changes in the firing rate of PL neuronal ensembles; when the test stimulus is similar to the conditioned stimulus, the two elicit similar behavioural responses and evoke a similar population-level firing rate in the respective PL neuronal ensembles.

      We have revised the Introduction as stated above. However, as discussed in our detailed responses, freezing behavior cannot fully account for the observed patterns of PL activity.

      (4) METHODS - 'Memory retrieval was tested on days 1, 15, and 30 after conditioning to probe early, long-term, and remote memory (Bontempi et al., 1996). During retrieval, mice were tested in a novel context with the CS<sup>+</sup>, CS1<sup>-</sup>, and two intermediate frequencies (7 and 11 kHz), presented in semirandom order, with each tone repeated three times (Fig. 1a).'

      Why was testing conducted in a different context than that of conditioning? This is likely to result in an underestimation of generalization to the different tones...

      In tone fear conditioning, it is always customary to test in a different context to dissociate conditioning to the context vs conditioning to the tones, which usually take place simultaneously in the same context [20]. Therefore, testing generalization in a novel context gives the correct estimate of generalization to the tones in the absence of contextual conditioning confounds. Please note that while overall freezing levels may be lower in a novel context due to the absence of contextual conditioning, the relative generalization gradient across tones — which is what your study measures — is unlikely to be systematically distorted by context change.

      (5) RESULTS - 'No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that freezing reflected associative learning.'

      The inference doesn't follow from the result described. Was there more freezing among animals in the shocked groups compared to those in the no-shock group? I presume so - my point is that this comparison is the one that most directly speaks to the presence or absence of associative learning.

      Experimental animals exhibited not only higher overall freezing but also graded freezing responses across tone frequencies. It is important to note that no-shock controls did not display this pattern, ruling out the possibility that the different frequencies themselves elicited graded behavioral responses. To clarify this point, we revised the sentence as follows: "No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that the graded freezing patterns resulted from associative learning rather than the acoustic properties of the tones." (Pg. 6)

      (6) 'Across animals and sessions, we identified distinct neuronal populations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound (Fig. 2b)...'

      To be clear, do you mean to say that there were distinct neuronal populations that consistently [i.e., across all three sessions] increased their responses to the tones [positive modulation], decreased their responses to the tones [negative modulation], showed variable responses to the tones [mixed responses], and did not respond to tones [not modulated]?

      The sentence refers to neuronal populations identified within each recording session based on their responses to the tones, not to neurons that maintained the same response profile across all three sessions. We have revised the text to make this distinction explicit. The only stable patterns across sessions were observed in graded neurons that were stable across retrieval.

      “Across animals, we identified distinct neuronal subpopulations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound in each session (Fig. 2b)” Pg. 7

      (7) What does 'active' mean in relation to Figure 2? Does this refer to neurons that displayed either positive responses, negative responses, and/or mixed responses? In the text, it is stated that 'Sound responder neurons were classified using a test that detected modulation based on magnitude relative to baseline variability, allowing reliable identification of both transient and sustained responses while remaining robust to noise...'

      I can't work out if this is the same classification criteria used for the determination of positive modulation, negative modulation, and mixed responding.

      We thank the reviewer for pointing out this ambiguity. In Figure 2, the term "active" referred to neurons that exhibited significant sound-evoked modulation and were subsequently classified as showing positive, negative, or mixed responses. Thus, active and sound-responsive refer to the same population of neurons. We removed the word active to avoid confusion.

      The reference to transient and sustained responses describes the temporal profile of the calcium signals rather than separate response categories. Some neurons exhibited brief calcium transients that rose and decayed rapidly, whereas others displayed sustained activity throughout the tone presentation. The sound-response detection algorithm was designed to reliably identify both temporal response profiles. We have revised the manuscript to make these definitions explicit. We modified the sentence as follows: “Sound-responsive neurons were identified using a statistical test that detected activity modulation relative to baseline variability, allowing reliable identification of responses while remaining robust to noise. This approach was effective for neurons exhibiting either brief calcium transients that rose and decayed rapidly or sustained activity throughout the tone presentation.” Pg. 7-8

      (8) 'A moderate proportion of neurons was present across all retrieval sessions, with no differences between groups (p > 0.05).'

      Do you mean to say that 'A moderate proportion of neurons was ACTIVE across all retrieval sessions, with no differences between groups (p > 0.05)'?

      We replaced the word present and replaced it with “active”. Pg. 8

      (9) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'If neurons encoding graded responses carry core mnemonic information, they should exhibit enhanced stability over time. To test this hypothesis, we quantified the proportion of registered neurons that retained their cluster identity across at least two retrieval sessions and compared these values to a shuffled null distribution (10,000 iterations), with multiple comparisons controlled using the BenjaminiHochberg procedure.'

      What does the comparison to the shuffled null distribution tell us exactly? I accept that some neurons were stable positive responders across at least two sessions. The comparison to the shuffled null distribution creates a false impression about the robustness of this stability or the 'enhanced stability over time'.

      Our intention in comparing the observed stability to a shuffled null distribution was to evaluate whether the proportion of neurons retaining cluster identity exceeded chance levels expected from random assignment. The shuffled distribution therefore provides a statistical baseline against which the observed degree of stability can be evaluated. We agree, however, that the wording “enhanced stability over time” may be confusing regarding this finding. We rephrased this paragraph to clarify that a subset of neurons retained cluster identity across all retrieval sessions at levels greater than expected by chance as follows:

      “These data demonstrate that graded clusters remain consistently active at levels exceeding chance, preserving their cellular identity and providing a stable representation of learned contingencies and generalization gradients.” Pg. 15.

      (10) ABSTRACT. The abstract states that, 'Stimulus-evoked population similarity scaled precisely with behavioral generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing, indicating that population-level similarity predicts inferred threat.'

      I believe that the sentence could be rewritten as, 'Stimulus-evoked population similarity reflected the degree of generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing.'

      We revised the text according the reviewer’s suggestion; however, we had to shorten the sentence due to word limits. “Population similarity tracked behavioral generalization, whereas consistent population states emerged only for shock-associated or highly generalized tones.” Pg. 2

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The ensembles with graded activation in proportion to stimulus valence are described at various points in the manuscript as "maintaining the emotional 'gist'", "preserving core components of the memory trace", and "preserving core components of the memory trace". This conclusion is premature because there is an alternative interpretation. The graded response ensembles would also be consistent with coding for the freezing behavior itself, irrespective of the specific memory or stimulus association that drives it. An ensemble that encodes a behavior in this way would not be considered mnemonic, just as motor neurons in the spinal cord are not, even if they may fire during a conditioned response. Indeed, previous work has identified neurons in the prelimbic cortex that encode freezing independently from the stimuli that signal an aversive outcome (e.g., Kyriazi, Headley, and Pare 2020; Casanova, Pouget, ..., Vetere 2024).

      There are two ways the authors can address this point.

      (a) Fit a generalized linear model to the time series of inferred spiking activity from the Ca2+ signal and include stimuli and freezing as predictors. Since freezing behavior is inconsistent across trials, while stimulus presence is fixed, they can be disassociated. If, after accounting for freezing, responsiveness neurons still show a graded coding of stimuli that agrees with inferred aversiveness, this would strengthen their claim that they have identified an ensemble that corresponds with mnemonic or salience aspects of the stimuli.

      (b) Conduct no further analysis but cover the issue in the discussion as a limitation to their study and to dampen some of the language throughout the manuscript that implies that a memory trace has been identified.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an important alternative interpretation is that graded-response ensembles could reflect freezing-related activity rather than representations of learned threat value. To directly address this possibility, we implemented a Generalized Linear Model (GLM) analysis, as suggested by the reviewer. The GLM was fitted to the activity of every sound-responsive neuron included in the population analyses and simultaneously incorporated tone identity and continuous freezing probability (derived probabilistically from miniscope acceleration) as predictors, allowing us to quantify their independent contributions to neuronal activity.

      We want to note that freezing was quantified from the miniscope's inertial measurement unit (IMU) using a two-component Gaussian mixture model applied to the log-transformed body-acceleration signal. Rather than classifying freezing with a binary threshold, we used the posterior probability of the low-movement state as a continuous freezing regressor. This approach captures graded variations in immobility and provides a more conservative test of tone encoding, because it accounts for more behaviour-related variance than a binary classifier, making it more difficult to detect an independent contribution of tone identity.

      We applied the GLM both to all sound-responsive neurons contributing to the population response curves and separately to the identified frequency-selective and graded neuronal subpopulations. Across all analyses, tone identity consistently explained neuronal activity better than freezing. Furthermore, the freezing-corrected tone β coefficients scaled with learned threat value, with graded neurons exhibiting the strongest monotonic gradients, indicating that they provide the most robust representation of learned threat value. These findings demonstrate that the graded coding of learned threat value persists after accounting for freezing behavior and therefore cannot be explained simply by the behavioral expression of fear. The new analyses are presented in Figures 4 and 8. Notably, although freezing-dominant neurons were present in both the tone-selective and graded populations, they represented only a small fraction of each group and were least prevalent among graded neurons (6–8%), further supporting the conclusion that graded neurons primarily encode learned threat value.

      In addition, we revised the manuscript to avoid language implying that these neuronal populations constitute a mnemonic trace. Instead, we consistently describe them as encoding learned threat value, a more accurate interpretation that is directly supported by the new GLM analyses.

      (2) The title makes a seemingly causal claim by using the term 'arise', "Learned and inferred valence arise from interactions between stable and dynamic subnetworks". While it is true that the authors show that both stable and dynamic ensembles encode valence, they do not demonstrate that the behavioral expression of valence depends on these codes, nor their interaction. Experimentally testing this is beyond the scope of this study (holographic two-photon stimulation of transient and stable ensembles?), but they may be able to get closer to it by examining inter-individual variability. The authors could measure the proportion of neurons in each subject that participate in the stable (graded responding) and dynamic (stimulus-specific) ensembles, and see if they predict individual differences in the expression of freezing behavior or its generalization. Indeed, this correlation may change across testing days.

      We agree that the term “arise” in the title may imply a stronger causal relationship than is directly supported by the present data. We modified the title in the resubmission as follows: “Complementary stable and dynamic prelimbic ensembles encode learned threat value underlying generalization and discrimination”

      The new LGM analysis confirms that a large proportion of neural activity can be predicted by tone threat value; therefore, we think this title fully captures our findings.

      We also appreciate the reviewer's suggestion to examine inter-individual variability. In the revised manuscript, we tested whether the proportion of graded neurons correlated with freezing behavior. However, the limited number of experimental animals in each experimental group provided insufficient statistical power to reliably assess this relationship. Accordingly, we revised the manuscript to clarify that our conclusions are based on correlational observations rather than causal inferences.

      In summary, we revised the title and related language throughout the manuscript to avoid implying causal mechanisms beyond the scope of the current experiments.

      Minor points:

      (1) I was surprised by the absence of an ensemble in the No-shock group that responded uniformly to all stimuli. Can the authors confirm this?

      Yes, we confirm this finding. It was unexpected to us as well. We would like to clarify, however, that some control neurons may have responded to more than one frequency, but these responses were too infrequent or too weak to be classified as a distinct graded neuronal population by our clustering algorithm. Thus, while broadly responsive neurons may have been present in the control group, they did not form a robust, identifiable ensemble comparable to that observed after fear conditioning.

      (2) Several different approaches were used to analyze the same Ca2+ responses to stimuli across testing days. These were the "Sound responder classification", "Average stimulus-aligned trace procedure", "Population similarity over time across tone pairs", and the construction of "Stimulus response vectors". These feature differing alignment/binning/interpolation, normalization, and response quantification procedures, and it is unclear why they cannot all be in agreement, at least when it comes to alignment and normalization.

      We thank the reviewer for this careful reading of our Methods. All analyses were performed on the same underlying calcium imaging dataset, but they were designed to address different aspects of the data and therefore required different preprocessing steps. The analyses share a common initial pipeline leading to the calcium traces (all z-scored across the session). Differences in subsequent processing (e.g., use of ΔF/F versus CASCADE-deconvolved activity, normalization, baseline correction, temporal binning, and interpolation) were introduced only when required by the specific analysis method.

      To make this clearer, we have substantially revised the Methods. We added a new overview of preprocessing section that summarizes the common preprocessing pipeline and explicitly distinguishes the shared steps from those that are analysis-specific. We also included a summary table describing the input signal (ΔF/F or CASCADE-deconvolved activity), normalization procedure, and temporal processing used for each analysis. Finally, the individual Methods sections were revised to eliminate redundancies and more clearly describe the steps to avoid confusion. We hope these revisions make the rationale for the different preprocessing procedures and the overall analytical workflow more transparent (Pg. 22-23)

      (3) In the methods section "Window-wise response quantification" the Ca2+ signal was baselinesubtracted and divided by the standard deviation in the baseline across trials, but that data was already presumably z-normalized to the baseline of each trial ("Data alignment and normalization"). This second step of normalization seems excessive. Why is it not sufficient to just take the average peri-stimulus response across the z-normalized trials from the "Data alignment and normalization" section? This is simpler and would capture the effect size of the response relative to baseline.

      We thank the reviewer for this careful observation.

      The two operations are also not the same normalization applied twice; they standardize different sources of variability. During the alignment step, each trial is z-scored relative to its own baseline by dividing by the standard deviation of that trial's baseline across time. This places all trials on a common within-trial scale before averaging. In the window-wise step, the trial-averaged, baseline-subtracted response is expressed relative to the standard deviation of the per-trial baseline levels across trials—a distinct quantity that reflects trial-to-trial baseline stability rather than within-trial fluctuations. The purpose of this second term was to down-weight windows in cells with unstable baselines across trials, and it entered the analysis only as a significance criterion; the magnitude threshold defining a sound responder was applied to the trial-averaged baseline-relative response itself. We have revised the Methods to clarify the distinct roles of these two normalization steps.

      To further address this concern, we re-ran the sound-responder classification after removing the second (between-trial) normalization step, so that responder detection depended only on the per-trial baselinerelative response magnitude and its temporal persistence. Across all cells, tones, and sessions (n = 89,504 cell–tone–session classifications), the two procedures agreed on 95.3% of labels. The small fraction of cells whose labels changed were almost exclusively those lying immediately at the detection threshold: 86.6% of changes involved cells moving into or out of the "modulated" category, whereas direct reversals between excitatory and inhibitory classification occurred in only 8 of 89,504 cases (0.009%). Consistent with the between-trial standard deviation being a less stable quantity when few trials are available, label changes were approximately twice as frequent in the three-trial retrieval sessions (5.2%) as in the ten-trial conditioning sessions (2.1%). Overall responder proportions changed only minimally (positive responders +2.2%, negative responders +4.3%), and all population-level findings—including the graded threat-value gradient across tones, its absence in no-shock controls, and its persistence after controlling for freezing in the , as analysis—were unaffected. These analyses demonstrate that our conclusions are robust to this methodological choice.

      (4) It would increase confidence in the tracking of neurons across days if the authors showed some example images of neurons tracked across days.

      We added an example in the Supplement. Additionally, we now provide measures of registration quality (Fig. S3)

      (5) Table S2 is a bit confusing. I take it that Graded A/B were only for CS15, and Graded C/D/E were only for CS3. Also, Common B1-4 were the cells with positive responses to individual stimuli, and Common C1-4 were the cells with negative responses to stimuli. If this is the case, it should be explained in the figure legend (or even better, clusters should be named and numbered consistently in all figures.

      Thank you for pointing this out, we corrected the Table to indicate which test corresponds to which figure and cluster, specifying which ones were positive or negative modulated. Please note old Table 2 is now Table 5

      (6) The term network and subnetworks implies some connectivity between neurons, but in this study, it is used to refer to the ensembles of cells activated in a similar manner. Since connectivity is never assessed, it would be better if the authors stuck to the terms ensembles or populations.

      We changed the wording and now use ensembles or populations

      (7) The legend for Figure S3 has the text 'eded', which seems to be a typo.

      We corrected this typo.

      References

      (1) Kitamura, T., et al., Engrams and circuits crucial for systems consolidation of a memory. Science, 2017. 356(6333): p. 73–78.

      (2) DeNardo, L.A., et al., Temporal evolution of cortical ensembles promoting remote memory retrieval. Nat Neurosci, 2019. 22(3): p. 460–469.

      (3) Tome, D.F., et al., Dynamic and selective engrams emerge with memory consolidation. Nat Neurosci, 2024. 27(3): p. 561–572.

      (4) Mau, W., M.E. Hasselmo, and D.J. Cai, The brain in motion: How ensemble fluidity drives memory-updating and flexibility. Elife, 2020. 9.

      (5) Gallego, J.A., et al., Long-term stability of cortical population dynamics underlying consistent behavior. Nat Neurosci, 2020. 23(2): p. 260–270.

      (6) Zaki, Y. and D.J. Cai, Memory engram stability and flexibility. Neuropsychopharmacology, 2024. 50(1): p. 285–293.

      (7) Lacagnina, A.F., et al., Distinct hippocampal engrams control extinction and relapse of fear memory. Nat Neurosci, 2019. 22(5): p. 753–761.

      (8) Sangha, S., Plasticity of Fear and Safety Neurons of the Amygdala in Response to Fear Extinction. Front Behav Neurosci, 2015. 9: p. 354.

      (9) Bouton, M.E., S. Maren, and G.P. McNally, Behavioral and Neurobiological Mechanisms of Pavlovian and Instrumental Extinction Learning. Physiol Rev, 2021. 101(2): p. 611–681.

      (10) Jenkins, H.M. and R.H. Harrison, Effect of discrimination training on auditory generalization. J Exp Psychol, 1960. 59: p. 246–53.

      (11) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.

      (12) Herzog, K., et al., Reducing Generalization of Conditioned Fear: Beneficial Impact of Fear Relevance and Feedback in Discrimination Training. Front Psychol, 2021. 12: p. 665711.

      (13) Lommen, M.J.J., et al., Training discrimination diminishes maladaptive avoidance of innocuous stimuli in a fear conditioning paradigm. PLoS One, 2017. 12(10): p. e0184485.

      (14) Casanova, J.P., et al., Threat-dependent scaling of prelimbic dynamics to enhance fear representation. Neuron, 2024. 112(14): p. 2304–2314 e6.

      (15) Kyriazi, P., D.B. Headley, and D. Pare, Different Multidimensional Representations across the Amygdalo-Prefrontal Network during an Approach-Avoidance Task. Neuron, 2020. 107(4): p. 717–730 e5.

      (16) McKenzie, S. and H. Eichenbaum, Consolidation and reconsolidation: two lives of memories? Neuron, 2011. 71(2): p. 224–33.

      (17) Winocur, G. and M. Moscovitch, Memory transformation and systems consolidation. J Int Neuropsychol Soc, 2011. 17(5): p. 766–80.

      (18) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.

      (19) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.

      (20) Phillips, R.G. and J.E. LeDoux, Differential contribution of amygdala and hippocampus to cued and contextual fear conditioning. Behav Neurosci, 1992. 106(2): p. 274–85.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      We have addressed the outstanding points made by the reviewers and provide a detailed description of the additional analyses performed & key results below. We have also updated the manuscript to reflect these additional results, and to contain a more detailed consideration of alternative plausible models.

      We also note that we have corrected one figure panel (Fig.2 panel B, Exp. 2 only), where we identified a small bug in the visualisation code whereby the data of either one or two participants was not correctly plotted in some conditions. This makes no difference to the reported effects.

      Please note that the reviewers acknowledged that your introduction now more broadly refers to the previous work from various groups on motor beta lateralisation (MBL).

      (1) Evaluating the correlation between CPP and MBL, which is key for supporting the claim that CPP is feeding MBL. If, as you are alluding to in your rebuttal, single-trial estimates of CPP are too noisy, trials could be binned based on CPP.

      As requested, we now provide additional analyses binning the data by CPP amplitudes, for the high-low coherence conditions at P1. Full details are provided below. In both experiments we find that, for a given coherence, greater CPP amplitudes at P1 correlate with stronger motor beta lateralisation.

      (2) Examining the possibility of down-weighting (in line with previous studies) compared to your current bounded integration description. Specifically, does the your model predict a bi-modal CPP-P2 distribution that is not evident in the data?

      We have now fit 3 additional models investigating alternative mechanisms that might account for the behavioural results. In particular, we have explored 4 different ways in which flexible weighting of the second pulse might account for both the behavioural and neural data. A full account of the results is provided below, and has been included in the manuscript. In sum, we find that a model which directly and uniformly downweighs evidence from the second pulse (as opposed to indirectly through little or no distance remaining to bound, as in our model) can account for the behavioural data well, but it cannot recapitulate CPP-P2 results unless an accumulation-terminating bound is also included in the model, and the additional complexity of a model with these two free parameters is not supported by model comparison. An alternative model where P2 is downweighed as an inverse function of P1 strength (i.e., stronger downweighing for P1-high coherence pulses) could recapitulate both the behavioural and neural data, but again, this was not favoured by model comparison metrics that account for complexity in the current dataset. We have added a piece on this in the discussion, noting how previous studies mentioned by the reviewer such as Cheadle et al., 2014 and Glickman et al., 2022 find a consistency bias where later evidence is boosted when it agrees with the earlier evidence, opposite to the dampening suggested by the model here, but that a key distinction in our task is that P2 always agreed with P1, so that a dampening might be plausible if subjects tend to withdraw some of their attention from the confirmatory P2 based on the strength of P1.

      Regarding the CPP-P2 distribution, our original bounded model does indeed predict a bimodal CPP-P2 distribution with a peak at 0 arising from the early termination trials, which does not appear in our data (see Fig. S19 and related reply below). However, that EEG noise precludes the detection of any such bimodality in single trial amplitude distributions is demonstrated by the fact a bimodal distribution is strongly predicted for CPP-P1 amplitudes due to the two coherences, most strongly in fact for the unbounded model since there would be nothing to cap the higher-coherence, yet no trace of such bimodality is evident there either, due to EEG noise.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper characterises the physiological and computational underpinnings of the accumulation of intermittent glimpses of sensory evidence, with a focus on the centroparietal positivity and motor beta lateralization. The main finding is that the centroparietal positivity builds up during evidence accumulation but falls back to baseline during gaps, while motor beta lateralization maintains a continuous a sustained representation throughout the gap and until response.

      Strengths:

      - Elegant combination of electroencephalography and computational modelling.

      - Innovative task design, including parametric manipulation of gap duration.

      - The authors describe results of two separate experiments, with very similar results, in effect providing an internal replication.

      Weaknesses:

      - A direct characterization of how the centroparietal positivity and motor beta lateralization interact is missing, which limits the novelty. In their reply to reviewers, the authors argue that the signal-to-noise ratio of EEG signals is insufficient for such analyses at the single-trial level. If so, a binned or trial-averaged approach could still be attempted.

      As requested, we have now performed an additional analysis binning trials according to single-trial CPP-P1 amplitudes. To this aim, we sorted trials according to P1 coherence, and median-split them within condition according to the CPP-P1 amplitudes integrated in a time window around the grand-averaged peak [0.4 to 0.6s] after pulse onset, on the same subset of electrodes as in the manuscript. We then plotted motor beta lateralisation (MBL) as the difference in [Contra - Ipsi] hemispheres. Stronger negativities thus indicate stronger lateralisation towards the correct response. In all 4 cases, (both experiments and both coherence levels), higher CPP amplitudes were associated with stronger lateralisation from 0.5s post-pulse onwards (Author response image 1).

      Author response image 1.

      MBL (bottom) traces aligned to P1 onset (time = 0), median split by CPP amplitude [0.4-0.6s] post pulse onset, within P1 coherence condition. Trials with stronger CPPP1 potentials were linked to stronger MBL lateralisation toward the correct response.

      - An exhaustive characterisation of sensors and frequency bands is also missing. In their reply to reviewers, the authors suggest that this would detract from their hypothesis-driven focus. I disagree: the main hypothesis and figures could remain centred on the centroparietal positivity and motor beta lateralization, with a more comprehensive mapping of sensors and frequencies placed in supplementary material. Since the purpose of the paper is to examine EEG-based decision signals in a novel behavioural context, a broader characterisation of the underlying EEG landscape would seem appropriate.

      To broaden our characterisation, we have now included an additional supplementary figure that describes another distinct, relevant EEG signal. Fig. S12 shows the lateralised readiness potential (LRP), a lateralised motor preparation signal that has long been used as an index of relative motor preparation with high temporal resolution (Eimer, 1998; Kelly & O’Connell, 2013; Vidal et al., 2015).The LRP is typically computed as the difference in voltage between [IpsiContra] lateral motor electrodes with respect to eventual response, and it captures the fact that the contralateral motor cortex exhibits more pronounced negative ramps than the ipsilateral one immediately preceding action execution. The EEG landscape characterised in our paper thus comprises four distinct signals that are all functionally relevant to the task, including occipital alpha power, relevant for attention & temporal expectation encoding, which was included both in the main manuscript (Fig. 2) and the supplement (Figs. S9, S11). Given our already extensive supplementary material (18 figures) focused on our main research questions, we feel that a full, hypothesis-free exploration across the dimensions of frequency, space (sensors) and time, considering that there are 60 experimental conditions among which differences may be tested for (2 directions x 3 gaps x 4 coherence pairings in exp 1, plus 2 directions x 4 gaps x 4 coherence pairings in exp 2, plus single-pulse trials), would render the supplemental materials excessive in volume. Again, the data will be shared publicly for future exploration of these many dimensions.

      Reviewer #2 (Public review):

      Summary:

      This manuscript examines decision-making in a context where the information for the decision is not continuous, but separated by a short temporal gap. The authors use a standard motion direction discrimination task over two discrete dot motion pulses (but unlike previous experiments, fill the gaps in evidence with 0-coherence random dot motion of differently coloured dots). Previous studies using this task (Kiani et al., 2013; Tohidi-Moghaddam et al., 2019; Azizi et al., 2021; 2023) or other discrete sample stimuli (Cheadle et al., 2014; Wyart et al., 2015; Golmohamadian et al., 2025) have shown decision-makers to integrate evidence from multiple samples (although with some flexible weighting on each sample). In this experiment, decision-makers tended not to use the second motion pulse for their decision. This allows the separation of neural signatures of momentary decision-evidence samples from the accumulated decision-evidence. In this context, classic electroencephalography signatures of accumulated decision-evidence (central-parietal positivity) are shown to reflect the momentary decision-evidence samples.

      Strengths:

      The authors present an excellent analysis of the data in support of their findings. In terms of proportion correct, participants show poorer performance than predicted if assuming both evidence samples were integrated perfectly. A regression analysis suggested a weaker weight on the second pulse, and in line with this, the authors show an effect of the order of pulse strength that is reversed compared to previous studies: A stronger second pulse resulted in worse performance than a stronger first pulse (this is in line with the visual condition reported in Golmohamadian et al., 2025). The authors also show smaller changes in electrophysiological signatures of decision-making (central parietal positivity, and lateralised motor beta power) in response to the second pulse. The authors describe these findings with a computational model which allows for early decision-commitment, meaning the second pulse is ignored on the majority of trials. The model-predicted electrophysiological components describe the data well. In particular, this analysis of model-predicted electrophysiology is impressive in providing simple and clear predictions for understanding the data.

      Weaknesses:

      Some readers may be left questioning why behaviour in this experiment is so different from previous experiments which use almost exactly the same design (Kiani et al., 2013; TohidiMoghaddam et al., 2019; Azizi et al., 2021; 2023). Overall performance in this experiment was much worse than previous experiments: Participants achieved ~85% correct following 400 ms of 33 - 45% coherent motion. In previous work, performance was ~90% correct following 240ms of 12.8% coherent motion. A second weakness is that, while the authors present a model which describes the data based on pre-mature decision-commitment, they do not examine explanations from the existing literature, that evidence is flexibly weighted, and do not provide any analyses which could be used to compare these descriptions. While their model can describe the data in this manuscript, it cannot explain the data from previous experiments showing a stronger weight on the second pulse.

      The revised version of the manuscript includes a detailed discussion about possible reasons why our stimulus characteristics, task design & experimental protocol may have led to the observed behavioural results (lines 605 onwards). Furthermore, we have now included an extended model comparison as a supplementary note which examines alternative models that could account for the observed data and an additional discussion section that links it to the existing literature.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have responded to each of the comments in the previous review. The manuscript introduction and discussion have been substantially improved, and now more adequately address the previous literature. Limited improvements were made to the analysis, although the authors acknowledged why the suggested improvements from the reviewers were unlikely to be successful, but did not attempt to address the comments using other methods.

      One common theme to both reviews was that, although the model broadly describes the data, it is not fully tested, and alternative descriptions are not fully considered.

      We thank the reviewer for their careful consideration of our data & their reply. We agree it is important to formally test alternative descriptions, most particularly those involving downweighting of the processing of P2, and we have now done so. If participants were simply downweighting P2 by implementing generally smaller drift rates regardless of P1, we would expect the CPP-P2 to exhibit the same classical pattern as in P1, with higher amplitudes following high-coherence P2. The interaction pattern we observe, whereby CPP-P2 amplitudes are systematically lower following high-coherence P1, within each P2 coherence, can only be explained by some dependency between the processes occurring at P1, and those following at P2. What if, as the reviewer suggested originally, “additional evidence from the second pulse was down-weighted according to certainty following the first pulse?” We thus consider this also.

      To formally test this, we have now fit four further models and compare both the fit quality to accuracy data and the predicted EEG results to our original implementation. All models were fit to the grand-averaged accuracy data for all conditions, using 10K simulations, and for both experiments separately (as was done for the bounded model presented in the manuscript). For EEG simulations in unbounded models, we assumed that the CPP signal fell back down to zero upon the dots turning blue, as for the simulations in the main manuscript. Fig. S15 illustrates the fit quality as measured by means of Bayesian Information Criterion (BIC), and we go into more detail on each model in turn below.

      “(1) Unbounded model + P1-independent P2 drift rate reweighting (DR<sub>P2</sub>)

      First, we fit a model with no bound in which the drift rate for P2 was estimated by reweighting (positive or negative) the P1 drift rate via an additional free scaling parameter w. This model thus had the same complexity as our original one (k = 3 free parameters): two drift rate parameters for high and low coherence of P1 (d<sub>high</sub>, d<sub>low</sub>), plus a scaling parameter, w, dictating the strength of P2 relative to same-coherence P1. The drift rate for P2 (d<sub>P2</sub>) was computed as the d*w, where d equals d<sub>high</sub> or d<sub>low</sub> depending on P2 coherence, multiplied by the scaling parameter w, so that if w < 1, P2 drift rates would decrease compared to the same coherence in P1. This model implements the reviewers’ suggestion that systematic downweighting of the second pulse might also account for the results (while also allowing for upweighting (w > 1) for the sake of flexibility).

      This first model (DR<sub>P21</sub>) yielded a similar fit quality to the behavioural data compared to our original bounded model (Bnd), as measured by BIC (Fig. S15). The new model broadly recapitulated the key behavioural results, including order effects, (Fig. S16,A) and could also recapitulate the generally lower CPP-P2 amplitudes (Fig. S16,B). However, this model failed to capture the key coherence-based pattern observed in the CPP-P2 data. Namely, while our EEG results showed that CPP-P2 in trials following P1-low coherence pulses reached overall higher amplitudes than that in trials following P1-high coherence pulses (see manuscript Fig. 3), this model’s simulations predicted that CPP-P2 should scale only with P2 coherence, showing higher amplitudes and steeper build-up rates for P2-high trials, regardless of P1 coherence (Fig. S16C). This is at odds with our empirical results.”

      “(2) Unbounded model + P1-dependent P2 Drift rate reweighting (invDR<sub>P2</sub>)

      Next, we tested a model in which P2 downweighting could depend on P1 strength. That is, we made the P2 drift rate scaling parameter w inversely proportional to P1 coherence so that P2 drift rate d<sub>P2</sub> = d*w/d<sub>P1</sub>, where d equals d<sub>high</sub> or d<sub>low</sub> depending on P2 coherence, and d<sub>P1</sub> indicates the preceding P1 coherence. This implements a kind of certainty weighting, whereby evidence following a strong P1 is more strongly dampened than evidence following a weak P1. This model could recapitulate the key behavioural findings (Fig. S17A), and also qualitatively captured the CPP-P2 effects (i.e. P1-low trials reaching overall higher amplitudes than P1-high trials, Fig. S17C), although the magnitude of this effect was substantially smaller than predicted by the original simple bounded model with no drift rate modulations. However, BICs indicated that this model provided an overall worse fit to the accuracy data compared to the original bounded model (Fig. S15).”

      (3) Bounded models + P2 drift rate reweighting (Bnd + DR<sub>P2</sub>, Bnd+ invDR<sub>P2</sub>)

      Finally, we investigated how well models with both a bound and either of the two P2 drift rate scaling methods we investigated above (uniform downweighting, P1-dependent downweighting) could capture the data, thus effectively testing two extensions of our original implementation that allowed for flexible reweighting of P2.

      The systematic reweighting model with a bound (Bnd + DR<sub>P2</sub>) could recapitulate all key behavioural and EEG findings (Fig. S18), but the additional complexity of the model was not supported by BIC (Fig. S15). Crucially, the reason that this model could recapitulate the CPPP2 results was still the presence of a bound, although we note that the proportion of trials that were predicted to terminate early was reduced in this model compared to the original bounded model presented in the manuscript (c.f. Fig. 4). Yet, this relatively small fraction of trials where accumulation ended early meant that 1) in some trials no accumulation was allowed to occur at all during P2, and 2) where it occurred, the DV was closer to the bound following P1-high coherence trials, thus needing to accumulate less further evidence before reaching a bound and yielding smaller CPP-P2 amplitudes overall in those trials. The inclusion of a bound was thus key to allow a model with systematic P2 downweighting to account for both behavioural and EEG data.

      The model with P1-based scaling of P2 drift rates with a bound could also recapitulate all key behavioural and EEG findings (Fig. S18), but again, the additional complexity of the model was not supported by BICs (Fig. S15).

      “Conclusion

      This extended modelling exercise suggests that 1) a simple bounded model is favoured by model comparison, 2) an alternative model of equal complexity which includes a P1dependent systematic downweighting of P2 rather than a bound can produce qualitatively similar results, the common feature of both viable models being the push-pull relationship between P1 and P2, and 3) more complex models including both a bound and P2 modulations can also account for both the behavioural and EEG results, but the additional complexity is not supported by the current data. While it is possible that some flexible weight modulations occur, these are not sufficiently influential to justify its inclusion in the model. Future work would nevertheless be warranted to explore this possibility in more detail using tailored task paradigms. For the scope of this paper, we have maintained the bounded account in the main manuscript as it is the one supported by the model comparison in the current dataset, but we have also included the alternative P1-based downweighting account as a supplementary figure, along with some additional discussion.”

      In response to my previous comment 3, the authors show their model predicts that there should be no CPP-P2 if the bound is reached before P2, otherwise CPP-P2 is similar to CPPP1 (Figure R2). The argument in the manuscript is that the lower CPP-P2 is because of this bound. The distribution of CPP-P2 amplitudes should therefore have higher variance than CPP-P1 amplitudes, and one might even predict a second mode in the distribution, around 0 amplitude (those trials that terminated before P2). The authors do not show this.

      In Figure R3, it looks like the data have been normalised independently for CPP-P1 and CPPP2 (since the means are approximately the same); normalisation also prevents a comparison of the variance. However, it is apparent that there is no bimodality in the CPP-P2 distribution - were there substantially more trials with 0 CPP-P2 amplitude than CPP-P1? Is the EEG data actually more consistent with a model that systematically downweights P2?

      In the previous Figure R3, data were normalised across CPP-P1 and CPP-P2, not separately. We plot the non-normalised values here (Author response image 2), for comparison, along with median, variance and skewness values for CPP-P1 and CPP-P2. Additionally, we attach the single-participant plots at the bottom of this document (Author response image 3)

      We reanalysed the non-normalised data, excluding outliers (defined as values exceeding the mean +/- 3 times the standard deviation, computed for each pulse & for each participant separately). We found, in both experiments, lower median amplitudes (Exp. 1: t(21) = 2.68, p = 0.013; Exp. 2: t(20) = 4.37, p < 0.001), higher variances (Exp. 1: t(21) = 0.59, p = 0.55; Exp. 2: t(20) = 2.33, p = 0.03) and more positive skewness (Exp. 1: t(21) = 1.79, p = 0.08; skewness: 0.027 vs. 0.69; Exp. 2: t(20) = 3.85, p <0.001; skewness: 0.027 vs. 0.214) in CPP-P2 compared to CPP-P1, although variance and skewness effects were only significant in Exp. 2.

      The reviewer argued in the previous review as well as here that increased variance would be predicted by the model – this is correct, but we believe that the mere presence of higher CPPP2 variances in our empirical data does not, on its own, necessarily support the model. That is because this increased variance could be explained by other factors, such as the increased EEG signal complexity of data at P2 compared to P1. This increased complexity naturally arises from overlapping potentials from CPP-P1 (which can be corrected for, but will increase data noise and thus variability nonetheless), as well as the various gap durations across conditions, which would also affect pre-P2 dynamics. Thus, while our data (partially - in Exp. 2 only) support the reviewer’s interpretation, we would be cautious in using the observation of increased variability as evidence for or against our model given the considerations above.

      Author response image 2.

      A. Empirical CPP–P1 and CPP-P2 amplitude [450-550ms post-pulse] distributions, pooled across coherences. Data were not normalised within-participant. Data were baselined 100ms before pulse onset prior to CPP-P2 amplitude extraction. In both experiments, CPPP2 amplitudes had a lower median (vertical line) amplitude and higher variance than CPP-P1. B. CPP amplitude simulations based on the original bounded model. A high number of trials where CPP-P2 amplitude should equal 0 due to early terminations. C. CPP amplitude simulations based on the inverse weighting model (invDR<sub>P2</sub>). The model did not predict any trials with zero CPP-P2 amplitude because early terminations were not allowed. Rather, the mean distribution shifted towards lower predicted amplitudes, because P2 was downweighted proportionally to P1 coherence.

      The reviewer asks: “Were there substantially more trials with 0 CPP-P2 amplitude than CPPP1?”. Given the noisiness of single-trial EEG data due to high-frequency artifacts, and/or spurious signal drifts (as illustrated by the raw CPP value distributions in Author response image 2A above), it is not be possible to directly detect trials on which CPP = 0. Instead, this must be inferred by other means. If the CPP is in reality at 0 on a larger proportion of trials, then on average, this should manifest as a higher fraction of trials with lower amplitudes, resulting in a more positively skewed distribution of CPP-P2 amplitudes compared to CPP-P1. As reported above, in both experiments we find that skewness is higher in CPP-P2 than in CPP-P1, in line with this hypothesis.

      Regarding the question “Is the EEG data actually more consistent with a model that systematically downweights P2?”, we point to the additional modelling we conducted in response to the comment above. To recapitulate, we find that a model that systematically downweights P2 could account for behavioural findings, but could only recapitulate the key EEG CPP-P2 patterns if an accumulation-ending bound was also included in the model. Instead, a model where P2 is downweighted as a function of P1 strength could qualitatively capture both behavioural and EEG findings, but was not strongly supported by goodness of fit measures in this dataset. We note, however, that the latter model does not predict a bimodal CPP-P2 distribution, but rather a shifted mean and more positive skewness for CPP-P2 trials (Author response image 2C; skewness: CPP-P1 = 1.08; CPP-P2 = 1.24). Thus, in that respect, it does appear to provide a better qualitative recapitulation of the single-trial CPP-P2 data. However, given the extent of EEG noise, the bimodal underlying distribution of our bounded model would also translate to a unimodal, skewed distribution as observed, so this does not provide a strong basis for adjudication. Underscoring this, it is noteworthy that an unbounded model in fact predicts a more separated bimodal distribution for P1 than a bounded model, yet, again, with EEG noise, we are not able to identify any such bimodality in the empirical P1 amplitude distribution.

      Minor:

      In the discussion, the authors write "We also used a narrower range of coherences than the previous studies, which possibly lends itself to calibrating a bound to achieve acceptable accuracy while saving cognitive effort." (Page 22). Perhaps this should be reworded. The range of coherence in this study was ~26-44% in Exp 1 (a difference of 18%) in previous experiments the coherence was 3.2-12.8% (a difference of ~10%). The ranges of performance were similar.

      We agree that the phrasing could be improved. We meant to say that we used only two coherences (high-low) with less than a twofold difference between them, instead of multiple levels of evidence strength in previous studies (e.g. 0,3.2,6.4,12.8) – we have clarified this.

      The authors mention in their rebuttal "However, in contrast to previous studies, we did not include any feedback on a trial-by- trial basis, instead only providing feedback at the end of each block indicating the average accuracy." Actually, Kiani et al., 2013 also only gave feedback at the end of each block. This is also implied in the discussion. I suggest this be removed as the common feedback in Kiani et al., 2013 suggests this cannot explain the difference.

      The methods in Kiani et al. 2013 state “At the end of motion stimulus, a 400–1000 ms delay period (truncated exponential) was imposed before the Go signal, disappearance of the fixation point, was presented. The subject was required to report the net direction of motion within 1 s after the Go signal by pressing a left or right key. Distinctive auditory feedback was delivered for correct and error responses. On trials with 0% coherence, the type of feedback was chosen randomly.“ We understand this means feedback was provided after every trial. If this interpretation is wrong, we would like to kindly ask the reviewer to point us to the relevant methods section so that we can correct the manuscript.

      Author response image 3.

      Individual CPP-P1 (blue) and CPP-P2 (orange), for both experiments (non-z-scored). Vertical lines indicate median CPP amplitudes for each pulse, respectively.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work by Antonnen et al. was triggered by claims of auditory-mediated effects on altricial avian embryos, which were published without any direct evidence that the relevant parental vocalizations were actually heard. I agree with Anttonen et al. that, based on the available evidence about avian auditory development, those claims are highly speculative and therefore necessitate more direct experimental verification.

      Attonen et al. have embarked on a comprehensive series of experiments to:

      (1) Better characterize acoustically the relevant parental vocalizations (heat whistles; in a separate preprint, not reviewed here)

      (2) Characterize the auditory sensitivity of zebra finches at various stages of their posthatching development. Despite the long-standing importance of the zebra finch as a songbird model in neuroethology of learned vocalizations, the auditory development of the species has not been studied so far.

      (3) Explore an alternative hypothesis of how the parental vocalizations might be perceived.

      The principal method used here is the non-invasive recording of ABR (auditory brainstem response), a standard neurophysiological method in auditory research. The click-evoked ABR provides a quick and objective assessment of basic hearing sensitivity that does not require animal training. Weaknesses of the technique include its limited frequency specificity and low signal-to-noise ratio. The authors are experienced with ABR measurements and well aware of those issues. ABR responses in zebra finches are shown to gradually appear during the first week posthatching and to mature in subsequent weeks, consistent with the auditory development in other altricial bird species studied previously. When matching the acoustic properties of parental heat whistles and auditory sensitivities, hearing of the parental heat whistles by zebra finch hatchlings was convincingly excluded. Although not directly measured, this also convincingly extrapolates to zebra finch embryos. Finally, the authors tested the hypothesis that parental heat whistles could induce perceptible vibrations of the egg and thus stimulate the embryo via a different modality. The method used here was laser doppler vibrometry, an appropriate, state-of-the-art technique that the authors also have proven experience with. The induced vibrations were shown to be several orders of magnitude below known vibrotactile sensitivities in mammals and birds. Thus, although zebra finch vibrotactile thresholds were not obtained directly, the hypothesis of vibrotactile perception of parental heat whistles by zebra finch embryos could also be rejected convincingly.

      In summary, even when considering some weaknesses of the techniques (which the authors are aware of), the conclusions of the paper are well supported: Auditory and/or vibration perception of parental heat whistles can be excluded as an explanation for previous reports of developmental programming for high ambient temperatures. As a constructive suggestion towards resolving the apparent paradox, the authors recommend repeating some of the crucial, previous playback experiments at lower sound levels that better match the natural parental vocalizations.

      (R1. A1) We thank the reviewer for their time and effort to thoroughly review our paper and for the positive comments on our manuscript. In the revised manuscript we have addressed the concerns that you have raised in the Joint recommendations above (Pages 1-4).

      Reviewer #2 (Public review):

      This study by Anttonen, Christensen-Dalsgaard, and Elemans describes the development of hearing thresholds in an altricial songbird species, the zebra finch. The results are very clear and along what might have been expected for altricial birds: at hatch (2 days post-hatch), the chicks are functionally deaf. Auditory evoked activity in the form of auditory brainstem responses (ABR) can start to be detected at 4 days post-hatch, but only at very loud sound levels. The study also shows that ABR response matures rapidly and reaches adult-like properties around 25 days post-hatch. The functional development of the auditory system is also frequency dependent, with a low-to-high frequency time course. All experiments are very well performed. The careful study throughout development and with the use of multiple time-points early in development is important to further ensure that the negative results found right after hatching are not the result of the experimental manipulation. The results themselves could be classified as somewhat descriptive, but, as the authors point out, they are particularly relevant and timely. Since 2016, there have been a series of studies published in high-profile journals that have presumably shown the importance of prenatal acoustic communication in altricial birds, mostly in zebra finches. This early acoustic communication would serve various adaptive functions. Although acoustic communication between embryos in the egg and parents has been shown in precocial birds (and crocodiles), finding an important function for prenatal communication in altricial birds came as a surprise. Unfortunately, none of those studies performed a careful assessment of the chicks' hearing abilities. This is done here, and the results are clear: zebra finches at 2 and 6 days post-hatch are functionally deaf. Since it is highly improbable that the hearing in the egg is more developed than at birth, one can only conclude that zebra finches in the egg (or at birth) cannot hear the heat whistles. The paper also ruled out the detection on egg vibrations as an alternative path. The prior literature will have to be corrected, or further studies conducted to solve the discrepancies. For this purpose, the "companion" paper on bioRxiv that studies the bioacoustical properties of heat calls from the same group will be particularly useful. Researchers from different groups will be able to precisely compare their stimuli.

      Beyond the quality of the experiments, I also found that the paper was very well written. The introduction was particularly clear and complete (yet concise).

      Weaknesses:

      My only minor criticism is that the authors do not discuss potential differences between behavioral audiograms and ABRs. Optimally, one would need to repeat the work of Okanoya and Dooling with your setup and using the same calibration. The ~20dB difference might be real, or it might be due to SPL measured with different instruments, at different distances, etc. Either way, you could add a sentence in the discussion that states that even with the 20 dB difference in audiogram heat whistles would not be detected during the early days post-hatch. But adding a (novel) behavioral assay in young birds could further resolve the issue.

      (R2. A0) We thank the reviewer for their time and effort to thoroughly review our paper, and for the positive comments on our manuscript.

      In our revision, we have added a new figure (Fig 5) and three new paragraphs (Lines 387-422) in the discussion to compare all published ABR and behavioral audiograms and the differences between and among these datasets. Our adult data is consistent with the reported findings from four other labs despite differences in stimulus design, setups and genetic background of the animals, providing strong support for our findings.

      Furthermore we have more clearly presented our argument why we think juvenile and embryos cannot detect heat whistles. For clarification, we have added a new figure (Fig 4) and four new paragraphs (Lines 271-355) in the discussion.

      We agree with the reviewer that more data is needed on the development of hearing in songbirds and zebra finches especially; both anatomical data and functional, such as innervation. We emphasize this need in our discussion (Lines 352-355 and Lines 405-411).

      More Minor Points:

      (1) As mentioned in the main text, the duration of pips (from pips to bursts) affects the effective bandwidth of the stimulus. I believe that the authors could give an estimate of this effective bandwidth, given what is known from bird auditory filters. I think that this estimate could be useful to compare to the effective bandwidth of the heat-call, which can now also be estimated.

      (R2. A1) Please see answer A3.4 under Joint Recommendations.

      (2) Figure 5b. Label the green and pink areas as song and heat-call spectrum. Also note that in the legend the authors say: "Green and red areas display the frequency windows related to the best hearing sensitivity of zebra finches and to heat calls, respectively". I don't think this is what they meant. I agree that 1-4 kHz is the best frequency sensitivity of zebra finches, but they probably meant green == "song frequency spectrum" and pink == "heat call spectrum". In either case, the figure and the legend need clarification.

      (R2. A2) We thank the reviewer for pointing out these issues. We have changed the figure and legend accordingly. In the meantime, we published a paper measuring the in vivo source levels of the heat whistles (Anttonen et al., Current Biology 2025), and we have carefully gone through this manuscript to incorporate those findings and adjust the text accordingly. We have therefore adjusted the analysis in Fig 3B to correct for the narrower frequency distribution of the heat whistles.

      (3) Figure 5c. Here also, I would change the song and heat-call labels to "song spectrum", "heat call spectrum". The authors would not want readers to think that they used song and heat calls in these experiments (maybe next time?). For the same reason, maybe in 5a you could add a cartoon of the oscillogram of a frequency sweep next to your speaker.

      (R2. A3) We thank the reviewer for pointing out these issues We mention the frequency sweep in the legend for panel A, but decided against including this in the figure to prevent too much clutter. We have changed the figure = legend to (new text underlined):

      Legend Fig 3. “A Setup used to measure sound-induced vibrations of eggs. A 94 dB, 0.25-to-10 kHz frequency sweep was played at the eggs to determine the vibration transfer function.”

      (4) Methods. In the description of the stimulus, the authors describe "5ms long tone bursts", but these are the tone pips in the main part of the manuscript. Use the same terms.

      (R2. A4) Thank you for catching this, we have changed this into “5ms long tone pips".

      Reviewer #3 (Public review):

      Summary

      Following recent findings that exposure to natural sounds and anthropogenic noise before hatching affects development and fitness in an altricial songbird, this study attempts to estimate the hearing capacities of zebra finch nestlings and the perception of high frequencies in that species. It also tries to estimate whether airborne sound can make zebra finch eggs vibrate, although this is not relevant to the question.

      Strength

      That prenatal sounds can affect the development of altricial birds clearly challenges the long-held assumption that altricial avian embryos cannot hear. However, there is currently no data to support that expectation. Investigating the development of hearing in songbirds is therefore important, even though technically challenging. More broadly, there is accumulating evidence that some bird species use sounds beyond their known hearing range (especially towards high frequencies), which also calls for a reassessment of avian auditory perception.

      Weaknesses

      Rather than following validated protocols, the study presents many experimental flaws and two major methodological mistakes (see below), which invalidate all results on responses to frequencyspecific tones in nestlings and those on vibration transmission to eggs, as well as largely underestimating hearing sensitivity. Accordingly, the study fails to detect a response in the majority of individuals tested with tones, including adults, and the results are overall inconsistent with previous studies in songbirds. The text throughout the preprint is also highly inaccurate, often presenting only part of the evidence or misrepresenting previous findings (both qualitatively and quantitatively; some examples are given below), which alters the conclusions.

      Conclusion and impact

      The conclusion from this study is not supported by the evidence. Even if the experiment had been performed correctly, there are well-recognised limitations and challenges of the method that likely explain the lack of response. The preprint fails to acknowledge that the method is well-known for largely underestimating hearing threshold (by 20-40dB in animals) and that it may not be suitable for a 1-gram hatchling. Unlike what is claimed throughout, including in the title, the failure to detect hearing sensitivity in this study does not invalidate all previous findings documenting the impacts of prenatal sound and noise on songbird development. The limitations of the approach and of this study are a much more parsimonious explanation. The incorrect results and interpretations, and the flawed representation of current knowledge, mean that this preprint regrettably creates more confusion than it advances the field.

      (R3. A0) We thank the reviewer for their detailed and critical assessment. We agree that establishing auditory sensitivity in very young altricial birds is technically challenging and that careful interpretation of ABR data is essential. We also appreciate the reviewer’s recognition of the importance of obtaining direct physiological data on auditory development.

      However, we respectfully disagree with the reviewer’s central claim that our methodology is flawed or that our conclusions are unsupported. Many of the concerns raised reflect misunderstandings of ABR methodology, selective interpretation of the literature, or assumptions that are not supported by empirical evidence. Below, we address the main points in turn.

      As also stated in our Provisional response, the reviewer’s critique can be distilled into four main arguments:

      (1) ABR cannot be reliably measured in very small animals.

      (2) Our stimulus design (especially 25 ms tone bursts) invalidates frequency-specific results.

      (3) ABR thresholds should be corrected to behavioral thresholds, which would alter conclusions.

      (4) Our findings are inconsistent with prior studies in songbirds.

      We address each of these below before responding point-by-point.

      (1) Suitability of ABR in small animals.

      Reviewer claim: ABR may not be suitable for very small hatchlings.

      This claim is not supported by existing evidence. ABR measures summed neural activity, and signal amplitude depends in part on the distance between neural tissue and recording electrodes. In smaller animals, this distance is reduced, which can increase signal amplitude and improve signal-to-noise ratio.

      Consistent with this, ABR has been successfully recorded in animals substantially smaller than zebra finch hatchlings, including zebrafish (Jørgensen et al., 2012), 10 mm froglets (Goutte et al., 2017) and 5 mm salamanders (Capshaw et al., 2020). It is in fact much more surprising the technique still provides robust signals even in extremely large animals such as Minke whales, where the distance between electrodes and brain is on the decimeter scale (Houser et al., 2024). We have extensive experience of recording ABRs in such small systems.

      Thus, there is no principled reason why ABR would be an invalid method to study auditory sensitivity in zebra finch hatchlings.

      (2) Stimulus design and tone duration

      Reviewer claim: Use of 25 ms tone bursts invalidates frequency-specific results.

      We agree that stimulus duration affects frequency specificity and ABR detectability. However, the reviewer’s assertion that there is a single “correct protocol” (≤5 ms) is inaccurate. In avian ABR studies, stimulus duration varies depending on experimental goals.

      Our choice of 25 ms tone bursts was intentional and necessary to accurately represent low frequencies (down to 250 Hz), ensuring sufficient cycles per stimulus in the plateau segment of 15 ms and minimizing spectral splatter (see auditory brainstem response design considerations discussed in Lauridsen et al., 2021) and our responses below.

      Key clarifications:

      a) Click-evoked ABRs form the basis of our conclusions about onset of hearing, not tone bursts.

      b) Tone bursts were used primarily to assess frequency-dependent maturation, not detect earliest sensitivity.

      c) We explicitly demonstrate that:

      - 25 ms bursts yield higher thresholds (lower sensitivity)

      - 5 ms pips yield lower thresholds and align with published ABR audiograms

      We have now:

      - Further clarified the rationale of stimulus design in the methods (Line 520-530) and added a section in the discussion (Lines 397-411).

      - Included additional comparison between burst and pip datasets (Lines 387-395).

      - Clarified that conclusions about early hearing do not depend on tone-burst data (Fig 4 and Lines 271-294).

      (3) ABR vs behavioral thresholds

      Reviewer claim: Failure to correct ABR thresholds (20–40 dB) invalidates conclusions.

      We agree that ABR thresholds typically overestimate behavioral thresholds. However, we disagree that this invalidates our conclusions.

      Importantly:

      a) We do not replace measured ABR data with corrected values, as this would be methodologically inappropriate.

      b) Instead, we:

      - Present measured ABR thresholds transparently (Fig 1-3)

      - Compare them directly to published behavioral audiograms (Fig 5)

      - Explicitly discuss the expected offset (Lines 413-422)

      In the revised manuscript we:

      a) Add a new figure (Fig. 5) compiling all published ABR and behavioral audiograms

      b) Show that:

      - ABR and behavioral audiograms have similar shapes (Fig 5, new discussion Lines 387-422)

      - Offsets are typically ~20 dB (Line 413-422)

      Crucially, even under conservative corrections:

      - Early hatchlings remain far less sensitive than adults (>54 dB SPL) to clicks.

      - Heat whistle levels remain at or below detection limits even in adults (new figure Fig 4)

      - The developmental gap (>50 dB between adults and 2 DPH hatchlings) remains decisive.

      Thus, incorporating ABR–behavioral differences does not change the central conclusion.

      (4) Consistency with prior literature in developing songbirds.

      Reviewer claim: Results contradict previous studies in developing songbirds.

      We respectfully disagree. The cited studies fall into three categories:

      (1) Behavioral studies (e.g., alarm-call responses)

      (2) Gene expression studies (e.g., ZENK activation)

      (3) Different species with different developmental trajectories

      None of these directly measure auditory sensitivity thresholds in zebra finch embryos or hatchlings.

      We emphasize:

      - Behavioral responses do not provide threshold measurements

      - Neural activation (e.g., ZENK) does not demonstrate functional perception thresholds

      - Cross-species comparisons must consider differences in developmental timing.

      We have now expanded the Discussion to explicitly address these studies and clarify how they relate to our findings (Lines 357-367). Our data are consistent with what is known about the physiology of auditory development in all birds studied so far.

      (5) Final statement

      We have revised the manuscript extensively to:

      - Clarify methodology and experimental design

      - Expand discussion of ABR limitations

      - Incorporate additional literature and comparisons

      - Correct inconsistencies in reporting

      We maintain that our central conclusions—that early zebra finch hatchlings lack detectable auditory brainstem responses and are unlikely to perceive parental heat calls at natural levels—is robust and supported by the data. We go into more detail in our point-by-point rebuttal below.

      Detailed assessment

      For brevity, only some references are included below as examples, using, when possible, those cited in the preprint (DOI is provided otherwise). A full review of all the studies supporting the points below is beyond the scope of this assessment.

      (A) Hearing experiment

      The study uses the Auditory Brainstem Response (ABR), which measures minute electrical signals transmitted to the surface of the skull from the auditory nerve and nuclei in the brainstem. ABR is widely used, especially in humans, because it is non-invasive. However, ABR is also a lot less sensitive than other methods, and requires very specific experimental precautions to reliably detect a response, especially in extremely small animals and with high-frequency sounds, as here.

      (1) Results on nestling frequency sensitivity are invalid, for failing to follow correct protocols:

      (R3. A1.1) We disagree that our protocol is invalid. There is no universal ABR protocol standard in birds. Our approach is consistent with established principles of stimulus design and is validated by:

      - Robust click-evoked responses

      - Consistent developmental trajectories

      - Agreement between pip-based ABR and behavioral audiograms

      We now clarify this explicitly. See response R3. A0 (Stimulus design and tone duration).

      The results on frequency testing in nestlings are invalid, since what might serve as a positive control did not work: in adults, no response was detected in a majority of individuals, at the core of their hearing range, with loud 95dB sounds (Figure S1), when testing frequency sensitivity with "tone burst".

      This is mostly because the study used a stimulation duration 5 times larger than the norm. It used 25ms tone bursts, when all published avian studies (in altricial or precocial birds) used stimulation of 5ms or less (when using subdermal electrodes as here; e.g., cited: Brittan-Powell et al 2004; not cited: Brittan-Powell et al 2002 (doi: 10.1121/1.1494807), Henry & Lucas 2008 (doi: 10.1016/j.anbehav.2008.08.003)). Long stimulations do not make sense and are indeed known to interfere with the detection of an ABR response, especially at high frequencies, as, for example, explicitly tested and stated in Lauridsen et al 2021 (cited).

      (R3. A1.2) ABR with long-duration stimuli were shown previously to work perfectly well in birds and do not interfere with the detection of an ABR response. Longer stimuli have been used, e.g. in the following bird papers:

      - Amin et al., J Neurophysiol 2007 (cited) on zebra finch ABR: 20 ms tone bursts

      - Korneeva et al. (2006), evoked responses from field L in flycatcher (cited by the reviewer): 20 ms

      - Saunders et al. (1973) (cited): 60 ms tone bursts.

      - Larsen ON, Wahlberg M, Christensen-Dalsgaard (2020) Amphibious hearing in a diving bird, the great cormorant (Phalacrocorax carbo sinensis), J Exp Biol, doi:10.1242/jeb.217265: 25 ms tone bursts

      Human ABR has been measured even with long-duration speech signals (duration 40 ms and longer). See for example Binkhamis et al., Ear and Hearing 40: 659-670, 2019.

      Furthermore, the reviewer unfortunately misunderstood some aspects in Lauridsen et al., which tested a specifical method exploiting neural phase locking, and showed that this method has a low-frequency bias because neural phase locking decreases at high frequencies.

      Also, Lauridsen et al clearly show the reason for using 25 ms bursts: because we aimed to measure frequencies down to 250 Hz, we need to have a sufficient number of cycles (3) in the plateau segment to represent the frequency adequately (Lauridsen et al. fig 1). A 15 ms plateau contains 3 cycles, plus 5 ms rise/fall time (1 cycle) equals 25 ms. We have clarified this in our methods (L520-525) and added a section in the discussion (L397-411).

      Thus, long-duration stimulations make sense and have been used before successfully. See also response R3.A0 (Stimulus design and tone duration).

      Adult response was then re-tested with a correct 5ms tone duration ("tone-pip"), which showed that, for the few individuals that responded to 25ms tones, thresholds were abnormally high (c.a. by 30dB; Figure 2C).

      Yet, no nestlings were retested with a correct protocol. There is therefore no valid data to support any conclusion on nestling frequency hearing. Under these circumstances, the fact that some nestlings showed a response to 25ms tones from day 8 would argue against them having very low sensitivity to sound.

      (R3. A1.3) Please see answer A3.1 under Joint Recommendations and R3.A0 (Stimulus design and tone duration).

      (2) Responses to clicks underestimate hearing onset by several days:

      Without any valid nestling responses to tones (see # 1), establishing the onset of hearing is not possible based on responses to clicks only, since responses to clicks occur at least 4 days after responses to tones during development (Saunders et al, 1973). Here, 60% of 4-day-old individuals responding to clicks means most would have responded to tones at and before 2 days post-hatch, had the experiment been done correctly.

      (R3. A2.1) We disagree that clicks necessarily underestimate onset.

      Clicks are broadband stimuli that:

      - Recruit large neural populations

      - Are commonly used to detect early auditory responses

      The cited delay between tone and click responses reflects stimulus energy differences, not an inherent limitation of clicks. We have clarified this in the revision and softened language to refer to “no detectable ABR response” rather than absolute deafness.

      The report that Saunders could only see responses to clicks later than to tones only reflects that he used click amplitudes that were insufficiently high. The ABR responses reported were to extremely intense tones (110 dB SPL) of long duration (60 ms).

      If Saunders had used clicks (duration 60 µs) with comparable sound energy, they would have had been very difficult to produce. He should have used clicks with an amplitude 1000 times (60 dB) higher than the tones to produce the same sound energy. This would have been clicks at 170 dB SPL, equivalent to the sound at the mouth of a medium-sized military cannon. Applying this pressure would not be a recommendable method for hearing assessment, but instead lead to irreversible hearing damage.

      In budgerigars, hearing onset occurs before 5 days post hatch, since responses to both clicks and tones were detectable at the first age tested at 5dph (Brittan-Powell et al, 2004).

      (R3. A2.2) This is not how we interpret the cited paper. They state that ‘Responses were first obtained from 1-week-old at high stimulation, and their click responses (Fig. 1) show no wave 1 peak at 6 days post-hatch. Also, their conclusion (p 3101) states that ‘budgerigars probably cannot hear at hatching’.

      (3) Experimental parameters chosen lower ABR detectability, specifically in younger birds: Very fast stimulus repetition rate inhibits the ABR response, especially in young:

      (a) The stimulus presentation rate (25 stim/ sec) is 6 times faster than zebra finch heat-calls, and 5 to 25 times faster than most previous studies in young birds (e.g., cited: Saunders et al 1973, 1974: 1 stim/sec or less; Katayama 1985: 3.3 clicks/sec; Brittan-Powell et al 2004: 4 stim/sec).

      Faster rates saturate the neurons and accordingly are known to decrease ABR amplitude and increase ABR latency, especially in younger animals with an immature nervous system.

      In birds, this occurs especially in the range from 5 to 30 stim/sec (e.g., cited: Saunder et al 1973, Brittan-Powell et al 2004). Values here with 25 rather than 1-4 stim/min are therefore underestimating true sensitivity.

      (R3. A3a) Please see answer A3.3 under Joint Recommendations.

      (b) Averaging over only 400 measures is insufficient to reliably detect weak ABR signals: The study uses 2 to 3 times fewer measures per stimulation type than the recommended value of 1,000 (e.g., Brittan-Powell et al 2002, 2024; Henry & Lucas 2008). This specifically affects the detection of weak signals, as in small hatchlings with tiny brains (adult zebra finches are 12-14g).

      (R3. A3b) Please see answer A3.2 under Joint Recommendations.

      (c) Body temperature is not specified and strongly affects the ABR:

      Controlling the body temperature of hatchlings of 1-4 grams (with a temperature probe under a 5mm-wide wing) would be very challenging. Low body temperature entirely eliminates the ABR, and even slight deviance from optimal temperature strongly increases wave latency and decreases wave amplitude (e.g., cited: Katayama 1985).

      (R3. A3c) Please see answer A2 under Joint Recommendations.

      (d) Other essential information is missing on parameters known to affect the ABR: This includes i) the weight of the animals,

      (R3. A3d-i) These important details have now been added in Table S6.

      (ii) whether and how the response signal was amplified and filtered,

      (R3. A3d-ii) Signal was amplified 500 times (74 dB). We have included these important details to the methods (Line 508).

      (iii) how the automatised S/N>2 criteria compared to visual assessment for wave detection,

      (R3. A3d-iii) There is no universally accepted/fitting protocol for performing ABR recordings in various animals, we decided to perform both visual and automated criteria detection of thresholds. The automated criterion in our experience is more strict approach than visual detection of thresholds, because using an automated criteria for threshold detection will remove potential experimenter bias from the results. We added this in our methods (Lines 583-585).

      (iv) what measures were taken to allow the correct placement of electrodes on hatchlings less than 5 grams.

      (R3. A3d-iv) We have placed electrodes in much smaller animals than 5 grams, and the common landmarks (ear opening, midline of skull) could easily be identified in the hatchlings.

      (4) Results in adults largely underestimate sensitivity at high frequencies, and are not the correct reference point:

      (a) Thresholds measured here at high frequencies for adults (using the correct stimulus duration, only done on adults) are 10-30dB higher than in all 3 other published ABR studies in adult zebra finches (cited: Zevin et al 2004; Amin et al 2007; not cited: Noirot et al 2011 (10.1121/1.3578452)), for both 4 and 6 kHz tone pips.

      (b) The underlying assumption used throughout the preprint that hearing must be adult-like to be functional in nestlings does not make sense. Slower and smaller neural responses are characteristic of immature systems, but it does not mean signals are not being perceived.

      (R3. A4) We acknowledge variation across studies and now include a comprehensive comparison of all zebra finch audiograms (new Fig. 5 and discussion Lines 387-422).

      Importantly:

      - Our pip-based audiogram aligns with previous ABR studies

      - Differences in high-frequency sensitivity likely reflect methodological variation or population differences.

      Our conclusions rely on relative developmental changes, not absolute thresholds.

      (5) Failure to account for ABR underestimation leads to false conclusions:

      (a) Whether the ABR method is suitable to assess hearing in very small hatchlings is unknown. No previous avian study has used ABR before 5 days post-hatch, and all have used larger bird species than the zebra finch.

      (R3. A5a) As stated above (R3.A0), small animals should give better signals, and we have been able to measure ABR in much smaller animals previously.

      (b) Even when performed correctly on large enough animals, the ABR systematically underestimates actual auditory sensitivity by 20-40 dB, especially at high frequencies, compared to behavioural responses (e.g., none cited: Brittan-Powell et al 2002, Henry & Lucas 2008, Noirot et al 2011). Against common practice, the preprint fails to account for this, leading to wrong interpretations.

      (R3. A5a) See our answer to R3.A0 above.

      For example, in Figure 1G (comparing to heat call levels), actual hearing thresholds would be 3040dB below those displayed. In addition, the "heat whistle" level displayed here (from the same authors) is 15dB lower than their second measure that they do not mention, and than measures obtained by others (unpublished data). When these two corrections are made - or even just the first one - the conclusion that heat-call sound levels are below the zebra finch hearing threshold does not hold.

      (R3. A5a) Our conclusion that heat whistles are unlikely to be perceived does not rely on a single dataset or method, but on the convergence of three independent constraints: (i) signal amplitude, (ii) adult auditory sensitivity, and (iii) developmental immaturity of the auditory system.

      First, heat whistles are low-amplitude signals. Our in vivo measurements show levels of ~33 dB re 20 µPa at 10 cm and ~14 dB at 1 m (Anttonen et al., 2025, Curr Biol). Even allowing for uncertainty in near-field estimation, a conservative upper bound at very close range (<5 cm) is ~40 dB SPL.

      Second, the most sensitive available measure—behavioral audiograms—places adult zebra finch thresholds at ~40 dB SPL at ~6 kHz (Okanoya and Dooling, 1987, J Comp Psychol), increasing steeply toward higher frequencies. Thus, even under optimal conditions, heat whistles fall at or below the detection threshold of adults, and only potentially at very close range.

      Third, auditory sensitivity in early development is substantially reduced. Our ABR data show a ≥40–60 dB decrease in sensitivity in hatchlings relative to adults for click stimuli, which provide the most favorable conditions for eliciting responses. Because frequency-specific sensitivity develops later, thresholds at 6–8 kHz are expected to be even higher in hatchlings and embryos.

      Taken together, these constraints define a narrow and unfavorable detection window: a low-amplitude, high-frequency signal positioned at the edge of adult sensitivity, combined with a large developmental decrease in auditory sensitivity. Under these conditions, it is unlikely that heat whistles are detectable by hatchlings or embryos.

      Importantly, this conclusion does not depend on precise correction factors between ABR and behavioral thresholds. Even when considering the most sensitive behavioral data and conservative estimates of sound level, the signal remains at or below the limits of detection in adults, and far below expected sensitivity in early developmental stages.

      We have included this argument more clearly in our discussion (Line 271-367), illustrated by new Fig. 5.

      (c) Rather than making appropriate corrections, the preprint uses a reference in humans (L180), where ABR is measured using a much more powerful method (multi-array EEG) than in animals, and from a larger brain. The shift of "10-20dB" obtained in humans is not applicable to animals.

      (R3. A5c) Again our conclusions rely on relative developmental changes, not absolute thresholds. The clinical practice in humans to measure ABR is with 4 electrodes, not a multi-electrode EEG array.

      Animal studies where ABR audiograms have been compared directly to psychophysical audiograms show differences of around 20 dB. For example in Brittain powell et al 2002, audiogram comparisons between behavioral and ABR in budgerigars were made within the same lab, same animal population and by the same people, leading to 20 dB difference. Our discussion includes a new paragraph on this topic (Lines 413-422).

      (6) Results are inconsistent with previous findings in developing songbirds:

      (R3. A6) We now explicitly discuss all cited studies. Key points:

      - Early behavioral responses do not imply high-frequency sensitivity

      - Studies in other species do not directly translate to zebra finches

      - None of the cited work provides direct measures of auditory thresholds in embryos

      As expected from all of the above, results and conclusions in the preprint are inconsistent with findings in other songbirds, which, using other methods, show for example, auditory sensitivity in: a) zebra finch embryos, in response to song vs silence (not cited: Rivera et al 2018, doi: 10.1097/WNR.0000000000001187)

      (R3. A6a) We thank the reviewer for pointing out this study. We agree that the question of auditory responsiveness in embryos is important, and that a range of approaches have been used to address it. However, the study cited (Rivera et al., 2018) does not directly measure auditory sensitivity, but instead infers auditory processing from differences between treatment groups exposed to different acoustic conditions. As such, it is not directly comparable to physiological measures of hearing sensitivity, such as ABR or behavioral thresholds.

      In addition, interpretation of these results is complicated by limited characterization of the acoustic environment and differences in experimental handling between groups, which may introduce confounding factors unrelated to auditory perception. Given these considerations, and because our study focuses specifically on quantifying auditory sensitivity using established physiological methods, we have chosen not to include a detailed discussion of this work.

      (b) flycatcher hatchlings at 2-3d post hatch (first age tested), across a wide range of frequencies (0.3 to 5kHz), at low to moderate sound levels (45-65dB) (cited: Aleksandrov and Dmitrieva 1992, not cited: Korneeva et al 2006 (10.1134/S0022093006060056)).

      (R3. A6b) Korneeva et al. (2006) and Aleksandrov and Dmitrieva (1992) report evoked responses in very young flycatcher hatchlings across a broad frequency range. Notably, these measurements were obtained using more invasive recording approaches (e.g., implanted electrodes in Field L) in unanesthetized birds, which are known to yield lower thresholds compared to far-field ABR recordings under anesthesia. These methodological differences likely account for part of the higher sensitivity reported.

      Importantly, even in flycatchers, auditory sensitivity shows substantial postnatal improvement: thresholds decrease by up to ~40 dB over the first days after hatching, and the upper frequency limit expands from ~4 to ~7 kHz. Thus, while absolute sensitivity may differ across species and methods, the overall developmental trajectory—gradual improvement in sensitivity and progressive extension toward higher frequencies—is consistent with our findings and with broader patterns reported in songbirds.

      We have added this paper in our discussion (Lines 362-365).

      (c) songbird nestlings at 2-6d post hatch, which discriminate and behaviourally respond to relevant parental calls or even complex songs. This level of discrimination requires good hearing across frequencies (e.g., not cited: Korneeva et al 2006; Schroeder & Podos 2023 (doi: 10.1016/j.anbehav.2023.06.015)).

      (R3. A6c) The species mentioned are different species from our study species. In the Pied flycatchers (Korneeva et al. 2006) experimental conditions were different: recordings were made from unanesthetized nestlings with implanted electrodes directly in the brain (field L), so likely with better SNR. The audiograms show a 10 dB SPL threshold after day 11, so the species may be considerably more sensitive than the zebra finch. The swamp sparrows in the Schroeder and Podos (2023) behavioral study were exposed for 4 days starting at 4-7 days post-hatch, so the study does not address embryonal hearing.

      (d) zebra finch nestlings at 13d post-hatch, which show adult-like processing of songs in the auditory cortex (CNM) (Schroeder & Remage-Healey 2021, doi: 10.1002/dneu.22802).

      (R3. A6d) This study does not conflict with our data. Even though sensitivity is lower at 10 days than in adults, cortical processing could still be ‘adult-like’.

      (e) zebra finch juveniles, which are able to perceive and learn song syllables at 5-7kHz (fundamental frequency) with very similar acoustic properties to heat calls, and also produced during inspiration (Goller & Daley 2001, doi: 10.1098/rspb.2001.1805).

      (R3. A6e) This result is not in conflict with our data. First, the onset of song learning occurs earliest at 20 DPH as discussed in the paper and our work demonstrates that click-evoked ABR thresholds are adult-like at 20 DPH. In the cited paper, the tutoring experiments were initiated at 35 DPH so the auditory system of the studied juveniles is mature.

      Second, even though Goller & Daley 2001 do not report the source level of the specific syllable or the playback sound pressure levels, the source level of the inspiratory notes is comparable to other syllables, and thus around ~67 dB and ~34 dB louder than heat whistles.

      NONE of these results - which contradict results and claims in the preprint - are mentioned.

      Instead, the preprint focuses on very slow-developing species (parrots and owls), which take 2-4 times longer than songbirds to fledge (cited: Brittan-Powell et al 2004; Köppl & Nickel 2007; Kraemer et al 2017).

      (R3. A6f) We have included papers in our discussion that reflect the known data (to our best knowledge) on the developmental neurophysiology and neuroanatomy of the auditory system and not proxies thereof.

      (7) Results in figures are misreported in the text, and conclusions in the abstract and headers are not supported by the data:

      For example:

      (a) The data on Figure 1E shows that at 4 days old, 8 out of 13 nestlings (60%) responded to clicks, but the text says only 5/13 responded (L89).

      (R3. A7a1) We apologize for this typo. Corrected.

      When 60% (4dph) and 90% (6dph) of individuals responded, the correct term would be that "most animals", rather than "some animals" responded (L89).

      (R3. A7a2) We have rephrased this sentence into “observable in most animals during” as suggested.

      Saying that ABR to loud sound appeared "in the majority only after one week" (L93) is also incorrect, given the data.

      (R3. A7a3) We have rephrased this sentence into: “Thus, sounds at loud, yet physiologically relevant SPLs do not evoke ABRs in the first days after hatching, but do so in all animals at 8 DPH.” (Lines 97-99).

      It follows that the title of the paragraph is also erroneous.

      (R3. A7a4) The paragraph title supports our conclusions and we will keep it.

      (b) The hearing threshold is underestimated by 40dB at 6 and 8Kz on Fig 2C, not by "10-20dB" as reported in the text (L178).

      (R3. A7b) We have changed the title of this section and moved the last sentence to the discussion to remove the focus on heat whistles. We added a paragraph in the discussion to specifically address the difference between ABR and behaviorally measured audiograms (Lines 413-422).

      (B) Egg vibration experiment

      (8) Using airborne sound to vibrate eggs is biologically irrelevant:

      (R3. A8.1) We agree that parental contact could influence vibration transmission.

      However, (1) prior studies assume airborne sound transmission, and (2) our experiment tests this assumption directly. We now clarify this scope (Lines 328-342 and Figure 4) and discuss contact-based transmission as a potential future direction.

      The measurement of airborne sound levels to vibrate eggs misunderstands bone conduction hearing and is not biologically meaningful: zebra finch parents are in direct contact with the eggs when producing heat calls during incubation, not hovering in front of the nest. This misunderstanding affects all extrapolations from this study to findings in studies on prenatal communication.

      (R3. A8.2) The definition of bone conduction is the response to sound that is not mediated by a functional middle ear, but through the skull. In the earlier study, the eggs were stimulated by sound from a headphone, so that is the reason for using the same stimulation here. See also joint response A3.5 above.

      (C) Misrepresentation of current knowledge

      (9) Values from published papers are misreported, which reverses the conclusions:

      (R3. A9) We thank the reviewer for identifying inconsistencies and have:

      - Corrected heat whistle frequency ranges consistently through our paper

      - Added a comprehensive comparison figure gathering all available audiograms (Fig 5)

      - Expanded discussion of high-frequency hearing.

      These revisions do not alter our conclusions.

      Most critical examples:

      (a) Preprint: "Zebra finch most sensitive hearing range of 1-to-4 kHz (Amin et al., 2007; Okanoya and Dooling, 1987; Yeh et al., 2023)" (L173).

      Actual values in the studies cited are:

      1-to-7kHz, in Amin et al 2007 (threshold [=50dB with ABR] is the same at 7kHz and 1KHz).

      1-to-6 kHz, in Okanoya and Dooling (the threshold [=30dB with behaviour] is actually lower at 6kHz than at 1KHz).

      1-to-7kHz, in Yeh et al (threshold [=35-38dB with behaviour] is the same at 7kHz and 1KHz).

      (R3. A9a.1) In this sentence presenting our results (“sensitive hearing range of 1-to-4 kHz”) we originally wrote that “these are consistent with the following papers (Amin et al., 2007; Okanoya and Dooling, 1987; Yeh et al., 2023)". This latter part was left out during the writing process. This explains the different numbers. We apologize for this mistake.

      To avoid confusion, in our revision we have placed all ABR curves together into new Fig 5 and have included a new paragraph to discuss the differences (Lines 387-422).

      Note that zebra finch nestlings' begging calls peaking at 6kHz (Elie & Theunissen 2015, doi: 10.1007/s10071-015-0933-6), would fall 2kHz above the parents' best hearing range if it were only up to 4kHz.

      (R3. A9a.2) Of course that is possible. However this representation is incorrect because begging calls are harmonic sounds with a fundamental frequency around 500 Hz and formant at 6 kHz. Begging calls thus contain lots of energy at frequencies below 6 kHz, while the heat whistles do not. The peak frequency of heat whistles is also their lowest frequency component.

      (b) The preprint incorrectly states throughout (e.g., L139, L163, L248) that heat-calls are 7-10kHz, when the actual value is 6-10kHz in the paper cited (Katsis et al, 2018).

      (R3. A9b) The authors in Katsis et al. 2018 provided a range of 6-10 kHz estimated from the spectrogram without any further specification of methods. In another manuscript, we have quantified the heat whistle frequency (Anttonen et al Curr Biol https://doi.org/10.1016/j.cub.2025.08.054) to be 6.8 ± 0.6 kHz. We have changed this accordingly throughout our manuscript.

      (c) Using the correct values from these studies, and heat-calls at 45 dB SPL (as measured by others (unpublished data), or as measured by the authors themselves, but which is not reported here (Anttonen et al 2025), the correct conclusion is that heat calls fall within the known zebra finch hearing range.

      (R3. A9c) Please see our answer R3.A5a. We have included this argument more clearly in our discussion (Lines 271-355), illustrated by new Fig. 4.

      (10) Published evidence towards high-frequency hearing, including in early development, is systematically omitted:

      (a) Other studies showing birds use high frequencies above the known avian hearing range are ignored. This includes oilbirds (7-23kHz; Brinklov et al 2017; by 1 of the preprint authors, doi: 10.1098/rsos.170255) and hummingbirds (10-20kHz; Duque et al 2020, doi: 10.1126/sciadv.abb9393), and in a lesser extreme, zebra finches' inspiratory song syllables at 57kHz (Goller & Dalley, 2001).

      (R3. A10a) We agree that some bird species produce or use acoustic signals extending into high frequencies. However, signal production is not evidence of perceptual sensitivity. Many animals, including birds and mammals, produce signals that contain harmonic or broadband components extending beyond their most sensitive hearing range without implying functional detection at those frequencies.

      The cited examples (oilbirds, hummingbirds, inspiratory song syllables in zebra finches) concern signal production or ecological specializations in different species, not measured auditory sensitivity in zebra finches, and particularly not during early development. As such, they do not provide evidence that zebra finches—adults or embryos—can detect low-amplitude, narrowband signals in the 6–7 kHz range.

      Our study explicitly addresses auditory sensitivity using physiological measurements, which is the relevant metric for evaluating detectability.

      (b) The discussion of anatomical development (L228-241) completely omits the well-known fact that the avian basilar papilla develops from high to low frequencies (i.e., base to apex), which - as many have pointed out - is opposite to the low-to-high development of sensitivity (e.g., cited: Cohen & Fermin 1978; Caus Capdevila et al 2021).

      (R3. A10b) We agree that the avian basilar papilla develops from base to apex (high to low frequency). We have now added a sentence in the Discussion to acknowledge this (Lines 406411).

      Importantly, morphological development does not directly translate to functional sensitivity. Functional hearing depends critically on factors such as hair cell innervation, synaptic maturation, and central auditory processing, which are known to develop over time.

      Our data show a low-to-high frequency progression in functional sensitivity, consistent with previous physiological studies. This apparent mismatch between anatomical gradients and functional onset has been noted in other systems and likely reflects the later maturation of neural encoding rather than hair cell differentiation per se. We now clarify this distinction in the revised manuscript (Lines 406-411).

      (c) High frequency hearing in songbirds at hatching is several orders of magnitude better than in chickens and ducks at the same age, even though songbirds are altricial (e.g., at 4kHz, flycatcher: 47dB, chicken-duck: 90dB; at 5kHz, flycatcher: 65dB, chicken-duck: 115dB; Korneeva et al 2006, Saunders et al 1974). That is because Galliformes are low-frequency specialists, according to both anatomical and ecological evidence, with calls peaking at 0.8 to 1.2kHz rather than 2-6kHz in songbirds. It is incorrect to conclude that altricial embryos cannot perceive high frequencies because low-frequency specialist precocial birds do not (L250;261).

      (R3. A10c) We agree that species differ in their auditory ecology and frequency specialization, and we do not claim that all altricial birds share identical developmental trajectories.

      However, the cited comparisons involve different species, methodologies, and developmental timelines, which limits their direct comparability. In particular:

      Developmental staging is not directly comparable across species using days post-hatch alone.

      - Different methods (e.g., invasive recordings vs. ABR vs behavioural assays) yield systematically different thresholds.

      - Ecological specialization (e.g., low-frequency vs. broadband species) influences adult audiograms and likely developmental trajectories.

      We have revised the Discussion to explicitly acknowledge these limitations and to avoid overgeneralization across species. Importantly, our conclusions are based on within-species comparisons (adult vs. hatchling zebra finches) combined with measured signal levels of heat whistles. These constraints are sufficient to evaluate detectability without relying on cross-species extrapolation.

      (11) Incorrect statements do not reflect findings from the references cited For example:

      (a) "in altricial bird species hearing typically starts after hatching" (L12, in abstract), "with little to no functional hearing during embryonic stages (Woolley, 2017)." (L33).

      There is no evidence, in any species, to support these statements. This is only a - commonly repeated - assumption, not actually based on any data. On the contrary, the extremely limited evidence to date shows the opposite, with zebra finch embryos showing ZENK activation in the auditory cortex in response to song playback (Rivera et al, 2018, not cited).

      The book chapter cited (Woolley 2017) acknowledges this lack of evidence, and, in the context of song learning, provides as only references (prior to 2018), 2 studies showing that songbirds do not develop a normal song if the song tutor is removed before 10d post-hatch. That nestlings cannot memorise (to later reproduce) complex signals heard before d10 does not mean that they are deaf to any sound before day 10.

      Studies showing hearing in young songbird nestlings (see point 6 above) also contradict these statements.

      (R3. A11a) We agree that the precise onset of hearing in altricial embryos is not well established. We have therefore revised the wording in the Abstract and Introduction to avoid categorical statements and instead reflect the limited available evidence (Lines 13-16 and 33-37).

      Our data provide direct physiological measurements showing extremely low sensitivity immediately after hatching, which constrains the likelihood of functional hearing in earlier embryonic stages.

      Regarding the cited ZENK study, we note that immediate early gene expression indicates neural activation but does not provide a measure of auditory sensitivity or detection thresholds. As such, it cannot be directly compared to physiological or behavioral measures of hearing.

      (b) "Zebra finch embryos supposedly are epigenetically guided to adapt to high temperatures by their parents high-frequency "heat calls" " (L36 and L135).

      This is an extremely vague and meaningless description of these results, which cannot be assessed by readers, even though these results are presented as a major justification for the present study. Rather than giving an interpretation of what "supposedly" may occur, it would be appropriate to simply synthesize the empirical evidence provided in these papers. They showed that embryonic exposure to heat-calls, as opposed to control contact calls, alters a suite of physiological and behavioural traits in nestlings, including how growth and cellular physiology respond to high temperatures. This also leads to carry-over effects on song learning and reproductive fitness in adulthood.

      (R3. A11b) We thank the reviewer for raising this point. In the revised manuscript, we have replaced the previous phrasing with a more precise and neutral summary of what these studies report, namely that embryonic exposure to heat-call playbacks has been associated with differences in physiological and behavioral traits.

      Our study, however, addresses a distinct question—whether such acoustic signals are detectable by embryos given known constraints on signal amplitude and auditory sensitivity. The cited studies do not directly quantify auditory perception or the physical sound environment experienced by embryos. As a result, they do not provide a direct test of the sensory mechanism required for acoustic communication. A detailed evaluation of experimental design and interpretation in those studies is beyond the scope of the present manuscript, and we therefore limit our discussion to assessing the biophysical and physiological plausibility of the proposed mechanism.

      (c) "The acoustic communication in precocial mallard ducks depends specifically on the lowfrequency auditory sensitivity of the embryo (Gottlieb, 1975)" (L253)

      The study cited (Gottlieb, 1975) demonstrates exactly the opposite of this statement: it shows that duckling embryos, not only perceive high frequency sounds (relative to the species frequency range), but also NEED this exposure to display normal audition and behaviour post-hatch. Specifically, it shows that duckling embryos deprived of exposure to their own high-frequency calls (at 2 kHz), failed to identify maternal calls post-hatch because of their abnormal insensitivity to higher frequencies, which was later confirmed by directly testing their auditory perception of tones (Dimitrieva & Gottlieb, 1994).

      (R3. A11c) We thank the reviewer for this clarification and have revised the relevant text. Our intention was to highlight that embryonic auditory experience can shape postnatal behavior, not to imply strict low-frequency limitation. Therefore we already included the actual frequency in the original sentence. We have removed the non-descriptive term “low-frequency” (Lines 330-332).

      (12) Considering all of the mistakes and distortions highlighted above, it would be very premature to conclude, based on these results and statements, that altricial avian embryos are not sensitive to sound. This study provides no actual scientific ground to support this conclusion.

      (R3. A12) We respectfully disagree with the reviewer’s conclusion.

      Our study does not make a general claim that altricial embryos are incapable of perceiving sound. Rather, we evaluate a specific hypothesis: whether zebra finch embryos and hatchlings can detect sound and parental heat whistles.

      Our conclusions are based on the convergence of:

      (1) Measured low sound pressure levels of heat whistles,

      (2) Established adult auditory thresholds (behavioral data),

      (3) A large developmental decrease in auditory sensitivity demonstrated by our ABR measurements.

      Even under conservative assumptions, these constraints place heat whistles at or below adult detection thresholds and far below expected sensitivity in hatchlings and embryos.

      Thus, our conclusion is not based on absence of evidence, but on quantitative constraints that make detection unlikely under biologically realistic conditions.

      Recommendations for the authors:

      Joint recommendations:

      In response to the joint recommendations, we have:

      - Expanded methodological transparency (temperature, electrode setup, stimulus parameters),

      - Added new data (Fig S3) and figures (Fig 4 and 5),

      - Clarified ABR limitations and interpretation,

      - Strengthened the separation between measured results and interpretation,

      - Reframed conclusions to avoid overstatement.

      These revisions leave the central two conclusions unchanged: 1) zebra finch hatchlings and embryos are functionally deaf, and 2) under biologically realistic conditions, heat whistles are unlikely to be detectable by zebra finch hatchlings or embryos.

      (A) Reviewers 1 and 2:

      Much of the reviewer discourse revolved around providing clarifications of methodology for measuring the ABR and caveats for interpretation. There was near consensus with reviewers 1 and 2 on issues related to the ABR, which should be addressed.

      We appreciate the reviewers’ consensus that the main conclusions are supported, while requesting clarification of methodological details and interpretation of ABR measurements.

      (1) Please address all of the issues raised by reviewers 1 and 2 above.

      (A1) All points raised by Reviewers 1 and 2 have been addressed in detail in our point-by-point rebuttal below. In addition, we have revised the manuscript to improve clarity, added new figures (Fig. 4, 5), and substantially expanded the Discussion with eight new paragraphs to better contextualize our findings.

      (2) Please also

      - clarify all aspects of experimental details of the ABR that were missing, including temperature control (estimate body and ambient temperatures during ABR recordings,

      - please address the possibility of hypothermia of hatchlings that could have reduced ABR responses,

      - and potential local head cooling due to surgical exposure and its likely effect on highfrequency response depression).

      (A2) In our revision, we have expanded the Methods section (L479-485) and added the following new data:

      Body and ambient temperature/hypothermia

      We have now included the body temperatures during ABR recordings in new table S6. These data show that:

      (1) Body temperature was stable throughout recordings,

      (2) Temperatures were within the physiological range,

      (3) Conditions were consistent across all age groups.

      Importantly, even the youngest hatchlings maintained stable temperatures and showed no indication of hypothermia. Therefore, differences in ABR responses cannot be attributed to temperature effects.

      Potential cooling due to surgical exposure

      This concern does not apply to our experiments. We used subdermal needle electrodes, which do not require surgical exposure. Therefore, no local cooling of the head occurred, and no tissue exposure could affect high-frequency sensitivity. We have added a clarifying sentence in the Methods section to explicitly state this (Line 501-503).

      (B) Reviewer 3 also had additional requests for clarification that should also be addressed:

      (3.1) Stimulus duration too long: The study used 25 ms tone bursts instead of the standard {less than or equal to} 5 ms "pips." Could this prevent reliable ABR detection, especially at high frequencies?

      (A3.1) We agree that stimulus duration affects ABR characteristics and now clarify our rationale in the manuscript.

      - The 25 ms tone bursts were deliberately chosen to ensure sufficient cycle representation at low frequencies (down to 250 Hz) and to avoid frequency splatter.

      - Using a constant duration across frequencies ensures comparable stimulus energy.

      Importantly:

      - The 25 ms data yield audiogram shapes consistent with both click responses (Fig 2C) and published behavioral data (new Fig 5).

      - To address potential high-frequency limitations, we included a dataset using 5 ms tone pips, which produced thresholds consistent with published ABR studies (new Fig 5).

      Thus, both stimulus types support the same conclusion: a gradual maturation of hearing sensitivity from low to high frequencies. We have expanded the Discussion with three paragraphs to clarify these methodological trade-offs (Lines 387-422).

      (3.2) Were 400 sweeps enough averaging? Might a signal appear at 1000 or more?

      (A3.2) Signal-to-noise ratio improves with the square root of the number of averages. Increasing from 400 to 1000 sweeps would therefore reduce thresholds by at most ~4 dB. This magnitude is small relative to the >54 dB developmental differences observed, and the large gap between signal levels and detection thresholds. Thus, increasing sweep number would not alter the conclusions. We now clarify this explicitly in the Methods (L530).

      (3.3) Was the repetition rate too high? How does the stimulus presentation affect the ABR? Might a signal have emerged with 1-4 per second?

      (A3.3) We have clarified stimulus presentation rates in the revised manuscript:

      - Clicks were presented at 25 Hz. Control measurements (now included as Supplementary Fig. S3) show no effect of this rate on ABR amplitude or threshold.

      - Tone bursts and pips were presented at ~3 Hz, consistent with commonly used rates that avoid neural adaptation. We apologize for leaving this out in our original submission.

      We now explicitly describe these parameters and their rationale in the Methods (Lines 543-552).

      (3.4) If possible, provide an estimate of the effective bandwidth of the tone pips and compare it with the bandwidth of the parental heat-whistles.

      (A3.4) We agree that stimulus bandwidth differs between tone pips and heat whistles, and that broader signals may stimulate multiple auditory filters. Shorter stimuli (e.g., 5 ms pips) have broader bandwidth and may stimulate multiple filters—particularly at low frequencies—potentially lowering thresholds, whereas longer stimuli (25 ms bursts) are more frequency-specific and may yield higher thresholds. At higher frequencies (including the heat whistle range), this effect is expected to be smaller.

      However, quantitative correction is currently not possible due to a lack of species-specific data on auditory tuning curves in zebra finches. The only available avian data (budgerigar; Saunders et al., 1979) suggest auditory filter bandwidths (Q10 dB, i.e., the bandwidth 10 dB below the peak divided by peak frequency) of ~1.4 at low frequencies and ~1000 Hz at higher frequencies, but how multifilter stimulation affects thresholds is unknown and likely species- and frequency-dependent.

      Given these uncertainties, direct comparison between tone stimuli and heat whistles requires strong assumptions. We therefore suggest that future studies should measure responses to natural heat whistles directly.

      (3.5) Egg Vibration Experiment. Address the possibility that if a parent were physically lying on top of an egg and generated a heat call, parental body vibration could significantly communicate some perceptual vibrotactile signal to the egg. Reviewer 3 raised the possibility that the experiments in this paper tested the extent to which an auditory input can vibrate the egg - what if a vocalizing bird was on the egg?

      (A3.5) We agree that embryos may receive multiple types of sensory input from parents, including direct mechanical cues.

      However, our experiment specifically tests the hypothesis proposed in prior work: that airborne sound (heat whistles) induces egg vibrations sufficient for perception. Our findings show that airborne sound-induced vibrations are orders of magnitude below known vibrotactile sensitivity thresholds.

      Regarding parental contact:

      - Heat whistles are produced by an aerodynamic whistle mechanism, not tissue vibration (Anttonen et al., Curr Biol 2025), meaning most respiratory energy is radiated as sound rather than dissipated as heat/vibration in the body.

      - A parent sitting on the egg would attenuate airborne sound transmission, not amplify it.

      We now clarify in the Discussion that other cues (e.g., respiration, direct contact, temperature) may exist and need to be included in new experiments (Lines 349-352). Even so, these are distinct mechanisms and were not the hypothesis tested in prior playback studies.

      In our paper, we will not add a detailed discussion of these prior papers as this is outside the scope of this paper. Instead, we added a paragraph what would be a constructive way forward (Lines 349-355). Hopefully somebody in the community will have the good fortune to secure research funding to continue this benchmarking work.

      (4) Finally, all reviewers agreed that some more context on the ABR and its relationship to functional hearing could be provided, with less direct focus on the heat-call experiments.

      (A4) We agree and in the revision discussion have compiled all published ABR and behavioral audiograms (Fig. 5) and added new paragraphs on functional hearing (Lines 261-269), and ABR vs behavioural audiograms (Lines 413-422).

      Furthermore, to remove focus on the heat whistles, we have moved all heat-whistle-specific interpretation out of the Results into a single, focused Discussion section (Line 271-355).

      The behavioral studies mentioned below show auditory responses (e.g., begging suppression). However, these behaviors are tested between ~5–10 days post-hatch (consistent with our findings), in different species, and do not provide quantitative sensitivity thresholds, nor do they address detectability of low-amplitude, high-frequency signals like heat whistles.

      In our revision we added a new discussion paragraph including these behavioral studies (Lines 357-367).

      For example, there are ample cases in the literature of altricial birds exhibiting behavioral evidence of auditory sensitivity by reducing begging calls in response to parental alarm calls:

      Platzen & Magrath (2004) - Playback of parental alarm calls nearly abolished nestling non-begging calls and reduced begging in scrubwrens. Proc. R. Soc. B 271:1271-1276.

      Different species: scrubwrens. Playback age: 5-, 8- and 11-DPH nestlings.

      Magrath, Haff, Horn & Leonard (2006) - Review and experiments on the developmental shift to silence/freeze after aerial alarm calls as chicks become fledglings; documents nestling quieting to alarms. Proc. R. Soc. B 273:2335-2341.

      Different species: scrubwrens. Playback age: 7-9 DPH nestlings, and 2- 4 days after fledging.

      Magrath, Pitcher & Dalziell (2007) - Nestlings respond to the sound of a predator's footsteps and parental food/alarm calls; includes begging suppression following predator sounds. Anim. Behav. 74:1117-1129.

      Different species: scrubwrens. Playback age: 8 DPH nestlings.

      Haff & Magrath (2012) - Nestlings suppress calling after heterospecific alarm calls (when acoustically similar to conspecific alarms), indicating generalized auditory danger recognition. Anim. Behav. 84:e.g., 495-505 (article).

      Different species: scrubwrens. Playback age: 5-6 and 10-11DPH nestlings. They show that 10-11 days old suppress calling while 5-6 days old do not.

      Barati & McDonald (2017) - Noisy miner nestlings suppress begging after conspecific alarm calls and some heterospecific cues; stronger/longer suppression for terrestrial-predator alarms. Sci. Rep. 7:9563.

      Different species: Noisy miner (Manorina melanocephala ). Playback age: 14 DPH. Nestlings started to vocalise at 5 DPH.

      Suzuki (2011) - In Paridae, parental alarm calls encode predator type; prior work (cited within) shows young of altricial species suppress vocalizations to alarms. Curr. Biol. 21:15-20.

      Different species: great tits. Playback age: 17 DPH.

      Can you please contextualize the present results about the timing of auditory development with the above body of work with respect to the timing of alarm call-induced begging call suppression?

      In our revision we have added a new paragraph in the discussion on these papers (Lines 357-367), and highlight the need for comparative work on hearing development in different species (Lines 352-355 and Lines 405-411).

      (5) Strictly speaking, a flat ABR does not equal deafness - at the extreme, an average of 10,000 trials may pull out a minuscule signal. Thus, the more rigorous path would be, in the results section, to ensure that statements summarize the data as they are, representing an absence of a neural signal.

      Save the interpretation of what this may mean for the discussion, and provide alongside this interpretation the necessary caveats related to temperature, rendition rate, averaging, etc.

      Clarify the conditions where a flat ABR demonstrates or fails to demonstrate immature deafness.

      Expand clarification for how the known 20-40 dB difference between ABR and behavioral thresholds can exist if a flat ABR can be interpreted as deafness.

      Consider refraining from concluding deafness from a flat ABR. Discuss that behavioral, single-unit, or alternative physiological assays might detect responses below the ABR threshold. If such cases exist, cite.

      (A5) We thought about this considerably before starting our measurements. What constitutes the absence of a signal? Even with intracellular recordings of all but one of the auditory neurons, the last one could still contain a signal and theoretically transmit information to the nervous system. We agree with the reviewers that absence of an ABR response should not be equated with absolute deafness. We have revised the manuscript accordingly and removed all statements implying “deafness” from the Results. The Results now strictly report presence or absence of detectable ABR responses.

      However, in both clinical and comparative contexts, absence of ABR responses at high SPLs (e.g., 90–95 dB) is widely interpreted as functionally non-responsive hearing. The developmental shift we observe (>54 dB) is far larger than typical ABR–behavioral offsets (20–40 dB). In the Discussion, we have added a new paragraph arguing that we think that the term functional deafness is reasonable here (Lines 261-269).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary.

      In this manuscript, the authors investigate the mechanisms underlying macrostructure formation in a freshwater filamentous cyanobacterium strain, F. draycotensis, focusing on how its ability to aggregate and form these structures depends on the physical properties of the filaments. Using experimental observations, they demonstrate that the cyanobacterium actively captures and surrounds particles, a process driven primarily by gliding motility.

      To explain these physical dynamics, the authors present a 3D model indicating that particle collection relies on filament length, as well as a specific mechanical response, namely, filament buckling and the subsequent formation of loops of bundles of filaments. While the authors have previously documented the buckling and looping characteristics of this strain, this study provides new insight by demonstrating that these physical phenomena are essential for particle capture and collection.

      Strengths:

      This manuscript benefits from a rigorous and detailed quantitative analysis of video recordings, which clearly documents the motility, buckling behaviour, and particle collection dynamics of the filaments.

      The authors effectively validate their hypothesis by using a naturally shorter filamentous strain, which fails to collect particles, suggesting that filament length is indeed a critical parameter.

      To further confirm the length dependency within the same species, the authors experimentally generated shorter filaments of F. draycotensis. The fact that these shortened filaments also lose the capacity to collect particles provides strong evidence supporting their proposed mechanism.

      We thank the reviewer for the accurate summary of our work and for their identified strengths of the study.

      Weaknesses:

      There is a conceptual concern. The authors linked the specific physical properties of this strain to evolutionary data, highlighting that the studied lineages diverged approximately two billion years ago. This creates a misleading impression that particle collection via flexible looping filaments is a recent evolutionary adaptation. However, particle collection has been observed in other cyanobacteria, such as Trichodesmium, which features short, rigid filaments. Therefore, the term ”emerging” does not seem appropriate for the title and text. The capacity to collect particles in the studied strain F. draycotensis appears to be primarily a function of physical characteristics (filament length and flexibility) rather than evolutionary age. Any cyanobacterial strain possessing similar physical properties is likely to exhibit comparable behaviour, rendering the evolutionary timeframe largely irrelevant to the core mechanism.

      We would like to first clarify our use of the term “emergent”. It seems that the referee took this in an evolutionary context, whereas we are using this term in the context of its use in systems dynamics, and referring to: “a complex entity displaying behaviors that its components do not have on their own, and emerge only when they interact in a wider whole”. Here, particle collection and dynamic aggregate formation “emerges” from the buckling and interaction of many filaments.

      With regards to the evolution of particle collection behavior, our comment on the evolutionary distance between F. draycotensis and Pseudoanabena sp. was meant to highlight the point that particle collection seems to be a function of physical characteristics and motility: Despite a large evolutionary distance, and possibly many biological differences, a physics-based argument is capturing the difference between the particle collection ability of these two organisms. Thus, we are in agreement here with the reviewer. We did not intend to make any arguments about “evolutionary age” of the particle collection behavior.

      We see that the short, evolutionary comment in the Introduction has confused the reviewer and potentially is confusing to other readers too. We will therefore remove this evolutionary comment from the Introduction section of the revised manuscript and make the point in more detail in the Discussion section.

      In addition, the phylogenetic tree presented in Figure S5 does not reflect the current consensus on cyanobacterial evolution and systematics and does not align with modern phylogenomic frameworks (see, for example, Strunecky et al., 2023 https://doi.org/10.1111/jpy.13304). There is also no such order Cyanobacteriales, which has been mentioned in a few older publications but is clearly outdated.

      We thank the reviewer for this comment, as it has made us realise that we never explained our choice of taxonomic framework in the manuscript, and perhaps this is the source of the confusion.

      The order Cyanobacteriales does exist: it is the order-level name applied in the Genome Taxonomy Database (GTDB) [5, 6], currently the most comprehensive and actively curated genome-based taxonomy of prokaryotes. GTDB classifies taxonomic groups algorithmically, as monophyletic groups in a concatenated marker-protein phylogeny with ranks normalised by relative evolutionary divergence. This has produced a number of re-groupings and new names relative to the older, morphology-derived classifications; many of these have since been formally proposed under the International Code of Nomenclature of Prokaryotes and the SeqCode [2], and are progressively being adopted by the NCBI. The placement of Cyanobacteriales, and of the other orders shown in Figure S5, can be inspected directly on the GTDB “Taxonomy Tree” (see here for the orders within the class Cyanobacteriia).

      We would also like to note that we do not see our tree and the framework of Strunecky et al. as being in conflict. Strunecky et al. constructed their phylogenomic backbone using GTDB-Tk and the same 120-marker concatenated alignment that the GTDB itself uses. What differs between the two schemes is therefore not the underlying phylogeny but the nomenclature applied to the resulting clades: Strunecky et al. work within the botanical tradition and combine the phylogenomic tree with phenotypic characteristics, thereby proposing ten new orders and fifteen new families, whereas GTDB assigns rank boundaries purely by evolutionary divergence and so draws broader order limits. In practice, the GTDB order Cyanobacteriales spans several of the families (e.g. Oscillatoriales and Coleofasciculales) and orders (e.g. Chroococcales and Nostocales), that are proposed within the Strunecky et al. work. Our reason for adopting the GTDB nomenclature is for practical reasons specific to this study. F. draycotensis is a recently described organism [3] that is not included in Strunecky et al. and has no placement in their tree. In GTDB it falls within a family-level lineage (placeholder name JAAUUE01) inside the Cyanobacteriales, with the sequenced members of the Coleofasciculaceae as its closest relatives. We could not have assigned it to one of the Strunecky orders without inventing a placement. The same applies to some of the other, recent metagenomically described cyanobacteria [10], which similarly have no assigned names in the literature. GTDB, by contrast, provides a reproducible, algorithmic assignment for all of these genomes, and is now widely used for this reason in genome- and metagenome-based studies of cyanobacteria (e.g. [1]). We therefore used it consistently throughout.

      Finally, with regards to the reviewer’s point about the tree itself, we would like to note that Figure S5 was intended only to convey the evolutionary distance between F. draycotensis and Pseudanabaena sp., and it was built from a modest set of six concatenated ribosomal protein markers using an approximate maximum-likelihood method with SH-like local support values. This is considerably less rigorous than the 120-marker RAxML and Bayesian analysis of Strunecky et al., and we agree that a stronger tree may be preferable. For the revised manuscript we are recomputing the tree from a substantially larger set of concatenated single-copy marker genes, using IQTREE with model selection and non-parametric bootstrap support. We would note, however, that the specific conclusion drawn from this figure — that the two strains we use for our experimental work, namely F. draycotensis and Pseudoanabena sp. belong to deeply divergent cyanobacterial lineages — is supported by the deep backbone of the cyanobacterial tree, which is stable across marker sets and inference methods, and is equally supported by the tree of Strunecky et al.

      We will make these points clearer in the Methods and Discussion sections of the revised manuscript, as well as the Figure S5 legend.

      Another concern is that the authors nearly completely ignore the role of type IV pili in the gliding motility of cyanobacteria, including filamentous strains. For a long time, there was a misconception that the gliding motility of cyanobacteria was due to slime protrusion. Slime plays a role in this process. However, several studies have shown that filamentous strains also use type IV pili to glide on surfaces. The authors should discuss this and include it in their model.

      The reviewer is correct that we did not include molecular details of gliding motility in our biophysical model. They are also correct to point out that pili and slime biosynthesis genes are shown to be involved in gliding motility [8]. It is, however, still unclear how these factors interact to produce mechanical gliding forces that can result in filament rotation (observed only in some filamentous cyanobacteria), filament reversal, as well as decoordination during such reversals, which we have previously shown in F. draycotensis [9]. Therefore, we have chosen to keep the biophysical model at a coarse-grained, phenomenological level. Instead of explicitly modelling the detailed molecular mechanisms behind force generation, we model only the minimum necessary physical forces and torques needed to reproduce the observed rotation and translation of the filament during gliding under de-coordinated conditions. This model is able to reproduce the experimentally observed buckling and twisting of filaments, and is therefore sufficient and useful to achieve a coarse-grained understanding of mechanical forces and their relationship to buckling, twisting and entanglement, which are the main processes we focus on here. As molecular details behind force generation in rotating, filamentous cyanobacteria become available, more detailed physical models can be constructed. We also note, in this context, that the two filamentous cyanobacteria we compare both encode the type IV pilus machinery, so the presence of a pilus motor does not by itself distinguish a particle-collecting from a non-collecting strain (see our response to the reviewer’s next point).

      We will make these points clearer in the Methods and Discussion sections of the revised manuscript.

      In addition, the authors concluded that gliding motility is responsible for particle collection by Fluctiforma draycotensis. Although I believe that their conclusion is correct, there might be several limitations to the experiments which allow for other reasons to be considered. Their conclusions were based on the use of a non-motile strain and an unspecified community without the motile Fluctiforma draycotensis strain. The problem I see here is that it is not clear why this strain is not motile; it could be because of the lack of type IV pili, mutations which alter their functionality, defects in slime secretion, any other mutation (e.g. in chemoreceptors), cellular structure, metabolism, or combinations of these. Furthermore, it is possible that the community changes its composition and behaviour when it lives without the cyanobacterium with a rich carbon source (glucose) or with a non-motile cyanobacterium which may not secrete slime or, for example, a signalling component which controls behaviour of the bacteria in the community. For that reason, the authors should be more cautious with their conclusion that solely motility behaviour of Fluctiforma draycotensis is responsible for particle collection. Additional factors might be responsible for these effects.

      Our conclusion that gliding motility is the main factor underpinning particle collection is based on several observations.

      Firstly, on the macroscopic scale we present several control experiments where we did not observe particle collection: (i) in the community featuring a non-motile F. draycotensis, and with mostly the same other bacterial species as the community featuring the motile F. draycotensis, (ii) in a bacterial community derived from the original F. draycotensis community but lacking any cyanobacteria, (iii) in the original community with physically shortened F. draycotensis, and (iv) in another cyanobacterial community featuring different bacteria and a naturally shorter, filamentous gliding cyanobacteria Pseudanabaena sp. A straightforward, parsimonious explanation that satisfies all these observations is that particle collection is underpinned by physical characteristics of gliding filamentous cyanobacteria.

      Secondly and more directly, in time-lapse microscopy imaging we repeatedly observe clusters of beads being moved by gliding filaments, and thereby being collected into larger clusters. Thus, whilst factors such as slime secretion also contribute, the primary mechanism driving the observed particle motion seems to be that particles stick to filaments and are carried around with them as they glide. We cannot rule out a contribution of pili to bead attachment and transport. We note, however, that both cyanobacteria compared here encode the type IV pilus machinery. In a homology survey of the two genomes, Pseudanabaena sp. and F. draycotensis both carry orthologues of the core T4P components — the assembly ATPase PilB, the retraction ATPase PilT, the inner-membrane platform protein PilC, the prepilin peptidase PilD, and the alignment-complex proteins PilM and PilF — together with the hormogonium-associated hmpD, hmpF and hmpG. Pseudanabaena sp. is therefore not pilus-deficient, and it does glide, yet it does not collect particles. The difference between the two organisms consequently cannot be attributed to the presence or absence of the pilus motor, which we would argue supports the physical argument we make here. Consistent with this, we have not identified mutations in pilus-related genes in the mutant, non-motile F. draycotensis.

      We are currently in the process of preparing another manuscript describing the mutations that led to motility loss in the non-motile F. draycotensis, as well as the proteins that are differentially expressed in the motile and non-motile F. draycotensis. These analyses will shed more light on the molecular mechanisms abolishing motility and how they might be influencing particle collection.

      In the revised manuscript, we will make these points clearer in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      The authors studied aggregation, buckling, and particle collection by the filamentous cyanobacterium Fluctiforma draycotensis, as well as by the filamentous Pseudanabaena sp. (order Pseudoanabenales). They performed a range of experiments, from imaging individual gliding filaments to multiple-day experiments showing the formation of large aggregates around a particle formed from a precipitate. They also developed a model of buckling filaments to argue that the ability of elastic filaments to collect particles and form macrostructures is confined to a part of the filament phase space in terms of length and flexibility, meaning that gliding combined with certain filament length and flexibility naturally reproduces the observations.

      Strengths:

      This is an impressive study that uses multiple tools to connect macrostructure formation with filaments’ gliding motility and buckling. It adds an important perspective on the biological and physical factors at play in the emergence of aggregates.

      We thank the reviewer for the accurate summary of our work and highlighting the strengths of the study.

      Weaknesses:

      The authors ignore the possibility that filament behavior plays an important role in the emergence of the observed patterns. Cyanobacteria have been shown to control their gliding motility (Pfreundt et al Science 2023; Kurjahn et al Nature Comm 2024), and their molecular motors are known to be regulated by chemotaxislike signaling pathways (Risser ARM 2025). As far as I know, how the coordination between the pulling agents along an individual filament works is actively debated, but there seems to be little doubt that it exists.

      To illustrate this point better, note that the aggregation observed by the authors is consistent with the length-dependent ability of filaments to coordinate gliding (I’m not saying this is how it works in Fluctiforma draycotensis; I’m saying it’s consistent). Suppose the coordination requires sufficiently long filaments, which could be the case when signaling molecules travel along the filament, propagating information about when individual pulling agents should reverse. In such a model, short filaments act randomly because they fail to coordinate gliding by the time they glide off nascent aggregates, whereas longer filaments can perform informed reversals because they have more time for coordination. Such behavior then explains the lack of aggregation in Pseudanabaena sp. (via behavior, not lack of stiffness). Note that Trichodesmium is stiff; its filaments do not buckle, yet Trichodesmium forms organized aggregates via tightly controlled motility. Note also that, as the authors report, since Pseudanabaena sp. is both shorter and faster, its filaments have relatively (to the time needed to glide the filaments’ length) little time to coordinate reversals. In my opinion, whether the observed patterns passively emerge from gliding and buckling or result from active behavior remains an open question.

      We appreciate the comment by the reviewer. We certainly agree that behavioral responses exist in filamentous cyanobacteria and will interplay with the physical aspects to produce exciting, complex dynamics. Besides the exemplar ideas that the reviewer provides, there can be many other scenarios involving behavioral responses, such as responses to light and to quorum sensing molecules or photosynthesis-generated radicals. For example, in F. draycotensis we have observed photo-responses at the aggregate level, which we are are currently studying. Photoresponses are also observed in Trichodesmium aggregates [7]. In general, a full understanding of the interaction of the biological (i.e. behavioral) and the physical aspects will require several future studies.

      In the current study, however, we focus on characterising the physical aspects of gliding motility alone, combined with experimental observations. We believe that this approach is important to establish a form of “null expectation” from the physics of gliding, elastic filaments alone. Currently, the molecular mechanisms responsible for coordinating the reversal behaviour of multiple filaments are still unclear, so it is difficult to experimentally demonstrate behavioural contributions to aggregate formation, e.g. via experiments where such behaviour is switched off. In the meantime, simulations such as those presented here allow us to test more precisely the potential role of activity, coordinated reversals and the elastic properties of the filament. In future it will be interesting to scale up the presented model to include multiple interacting filaments, and to systematically test the respective roles of active coordination behaviour for one individual filament (reversals) and for multiple interacting filaments (where contacts modulate activity), as well as the physical properties (length and flexibility). Such modelling studies can then identify if a ‘purely physical’ model can or cannot generate realistic aggregates, and pinpoint whether additional coordination mechanisms are needed to regulate aggregation. By testing the combination of different physical and biological coordination mechanisms, it would then help to indicate how much of a role is played by various potential active coordination behaviours.

      We will bring out this point more clearly in the Discussion section of the revised manuscript.

      I also have a small suggestion regarding this statement on model novelty:

      The essential novelty of this model is that the filament itself is active and out of equilibrium, and additionally, the forces and torques are applied locally along its centreline, and not at its extremities as in previous steady-state mechanical studies of elastic, twistable filaments such as DNA [31-33] (see Methods and SI).

      This statement needs to be revised as it ignores a substantial body of work on self-organization of active filaments: (R. E. Isele-Holder, J. Elgeti, G. Gompper, Soft Matter 2015; Pfreudnt et al, Science 2023; Faluweki et al PRL 2023; Kurjahn et al Nature Comm 2024).

      We agree with the reviewer that there is a significant literature on active filaments, some of which we have already cited and will now discuss in more details, as well as adding and discussing the suggested additional references. Our statement on “model novelty” refers to the analysis of buckling instabilities of biological filaments, and in particular DNA, due to a combination of forces and torques. To our knowledge, this has only be studied explicitely by [4], and only in the local (resistive force theory) limit. The elastohydrodynamic simulations coupled to local active forces and torques, as we implemented here, are therefore novel and will expand the analysis of both microbial filaments and other biological polymers. We will clarify these points in the Methods and Discussion sections of the revised manuscript.

      Last point: the authors often say that their observations are reproducible (’...reproducibly forms macroscopic granules...’). What is meant? Different experiments on different days, different aliquots?

      The “replicability” statement was in reference to different experiments started on different days using cultures obtained from serial transfer experiments, as well as cultures re-initiated from cyrostocks. This point will be made clear in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The authors report and characterize the formation of aggregate microstructures by the motile filamentous cyanobacterium Fluctiforma draycotensis, which exhibits gliding motility accompanied by rotation along the long axis while excreting EPS. In experiments with motile F. draycotensis cultures, they observed the formation of granular structures composed of cyanobacteria and other material (iron, polystyrene beads, etc.), with macrostructures on the scale of 1mm within 24 hours. The structures were motile at speeds comparable to that of the cyanobacteria filaments, resulting in their growth through coalescence over time. Notably, such macrostructures were absent in nonmotile F. draycotensis, pointing to the role of filament motility in their formation. Through experiments examining the micro-scale dynamics, inert material such as small polystyrene beads was found to be transported by the gliding, buckling, and plectoneme dynamics of the filaments, pointing to the underlying mechanism by which particles are collected into larger-scale microgranule structures.

      To interrogate the properties that drive the cyanobacteria filament buckling, plectoneme formation, and entanglement, the authors develop a mechanical model for filaments as nearly inextensible, slender bodies with resistance to twisting and bending under active gliding forces and torques and responding to fluid flows and surface adhesion. They derive expressions for the thresholds for buckling and twisting instabilities, which are additionally demonstrated and interrogated through simulation via the Immersed Boundary Method. Most importantly, bending and plectoneme formation only occur with sufficiently long filaments, and the threshold is shorter for bending than for plectoneme formation. Experimental observations with wild-type filaments agree with the model-predicted thresholds. The authors perform additional experiments with shorter filaments below both thresholds, including the filamentous bacterium Pseudanabaena, which fail to collect particles (though can in principle form macrostructures).

      Strengths:

      This work appears to be novel (notably, the discovery and characterization of the particle collection behavior of a filamentous cyanobacterium) and has interesting implications for both naturally observed cyanobacterial macrostructures as well as the controllable parameters in engineering them. The experimental and modeling work is well motivated, contributing to the broader understanding of macrostructure formation and material aggregation through active filament dynamics (not exclusive to cyanobacteria), as well as the underlying physical properties governing important filamentous cyanobacterium dynamics. As such, I would expect the results of this paper to be of broad interest to both biophysicists and microbiologists. Generally, the manuscript is well written with clear, compelling figures that illustrate the important conclusions of this study.

      We thank the reviewer for the accurate summary of our work and recognising the broad relevance of the study.

      Weaknesses:

      In the section on “Shorter gliding filaments cannot collect particles nor form granule macrostructures”, the filamentous cyanobacteria considered “all” fall below the predicted thresholds for bending and twisting. The “long” F. draycotensis are 60 microns in length, notably less than the 120 and 320 micron thresholds derived in the previous section as well as the lengths of filaments considered in Figure 3D, yet these “long” 60 micron filaments form macrostructures. How can this be understood in the context of the model predictions? Is the nature of the macrostructures in Figure 4B, the microscale parameters, or the collection of particles somehow different than those with filaments an order of magnitude longer in earlier parts of the paper? The paper would be stronger if these sorts of questions were addressed in the text and/or with supplementary figures.

      We thank the reviewer for this point. Indeed as we mention in the text, the ‘long’ population has a mean length of 60 micron. However, as shown in the length distribution plot in Fig 4A, the maximum filament lengths observed in these populations (within the samples used for microscopy) are 560 microns for the long filaments, versus 240 microns for the short filaments. Thus, we expect the long population to contain multiple filaments that can buckle and a few that can form plectonemes, whilst the short population might have some buckling filaments and none that form plectonemes. We stress that Fig 4A only shows the length distribution for what we believe to be a representative sample taken from the long and short populations, not the full data from the entire population.

      We will revise the main text to include the maximum filament lengths of the two populations as well as the mean values. We will also add lines to Fig 4A to indicate the buckling and plectoneme threshold lengths from the analytical estimate for the F. draycotensis filaments (same values as in Fig 3), to make it clear that the long population contains more buckling/plectoneming filaments than the short population.

      References:

      (1) E. S. Cameron et al. “Diversity and specificity of molecular functions in cyanobacterial symbionts”. In: Sci Rep 14.1 (2024), p. 18658. issn: 2045-2322 (Electronic) 2045-2322 (Linking). doi: 10.1038/s41598-024-69215-8. url: https: //www.ncbi.nlm.nih.gov/pubmed/39134591.

      (2) M. Chuvochina et al. “Proposal of names for 329 higher rank taxa defined in the Genome Taxonomy Database under two prokaryotic codes”. In: FEMS Microbiol Lett 370 (2023). issn: 1574-6968 (Electronic) 0378-1097 (Print) 0378-1097 (Linking). doi: 10.1093/femsle/fnad071. url: https://www.ncbi.nlm.nih.gov/ pubmed/37480240.

      (3) S. J. N. Duxbury et al. “Niche formation and metabolic interactions contribute to stable diversity in a spatially structured cyanobacterial community”. In: ISME J (2025). issn: 1751-7370 (Electronic) 1751-7362 (Linking). doi: 10.1093/ismejo/ wraf126. url: https://www.ncbi.nlm.nih.gov/pubmed/40577531.

      (4) Raymond E. Goldstein, Thomas R. Powers, and Chris H. Wiggins. “Viscous Nonlinear Dynamics of Twist and Writhe”. In: Physical Review Letters 80.23 (June 1998), pp. 5232–5235. issn: 1079-7114. doi: 10.1103/physrevlett.80.5232.

      (5) D. H. Parks et al. “A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life”. In: Nat Biotechnol 36.10 (2018), pp. 996–1004. issn: 1546-1696 (Electronic) 1087-0156 (Linking). doi: 10.1038/ nbt.4229. url: https://www.ncbi.nlm.nih.gov/pubmed/30148503.

      (6) D. H. Parks et al. “GTDB release 10: a complete and systematic taxonomy for 715 230 bacterial and 17 245 archaeal genomes”. In: Nucleic Acids Res 54.D1 (2026), pp. D743–D754. issn: 1362-4962 (Electronic) 0305-1048 (Print) 03051048 (Linking). doi: 10.1093/nar/gkaf1040. url: https://www.ncbi.nlm.nih. gov/pubmed/41123020.

      (7) U. Pfreundt et al. “Controlled motility in the cyanobacterium Trichodesmium regulates aggregate architecture”. In: Science 380.6647 (2023), pp. 830–835. issn: 1095-9203 (Electronic) 0036-8075 (Linking). doi: 10.1126/science.adf2753.

      (8) Douglas D Risser. “Motility in Filamentous Cyanobacteria”. In: Annual Review of Microbiology 79 (2025).

      (9) Jerko Rosko et al. “Cellular coordination underpins rapid reversals in gliding filamentous cyanobacteria and its loss results in plectonemes”. In: eLife 13 (2025), RP100768.

      (10) A. Scarampi et al. “Enrichment of convergent metabolic functions in microbial communities through imposed and emergent environmental niches”. In: bioRxiv (2026). doi: 10.64898/2026.02.11.705344.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      The study aimed to assess the associations between meteorological drivers and influenza is important although not new. The authors used 6 years of surveillance data and deep learning models, combining distributed lag non-linear models (DLNM) with Bayesian-optimized LSTM neural networks for predictive modeling. The key interest in this area is to explore the subtropical locations, where influenza is less common and circulates year-round. The authors further claimed that such an association could be able to provide an early warning in the community.

      Strengths:

      Study design based on a prospective cohort to analyse the data for retrospective outcomes.

      We would like to express our sincere and heartfelt gratitude to all of you for your exceptionally thorough, constructive, and intellectually rigorous evaluation of our manuscript. The breadth and depth of the feedback we have received reflect a high standard of scientific scrutiny that we deeply respect and appreciate.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 Mpro enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the Mpro monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how Mpro inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      None. The requested mutagenesis data have been provided in the revised manuscript, and all of my previous concerns have been satisfactorily addressed.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      None. The overall quality of the manuscript has been substantially improved. The authors have added supportive mutagenesis data in the revised manuscript to validate the proposed role of C300. All of my concerns have been adequately addressed.

      We appreciate the reviewer’s recognition of the improvements made in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents a sophisticated investigation into the mechanisms by which different inhibitor classes affect the SARS-CoV-2 main protease (Mpro), a pivotal antiviral drug target. This study reveals that effective inhibition can be achieved by modulating the stabilization of the essential dimeric state. It also indicates the dimer interface could be a druggable allosteric site, which may offer a strategy for developing broad-spectrum anticoronaviral agents.

      Strengths:

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      Comments on revised version:

      The authors have very nicely addressed most of the previous comments raised. But one comment remains to be clarified relating to original point 5 and the authors' response:

      "We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296-306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation."

      Two questions remain for the statements in line 376-380. First, while it is stated "residues 296-304 in the C-terminal region of Mpro were more flexible upon ebselen binding", the segment of 296-306 is shown Figure 4c. Second, the HDX change for this segment upon ebselen binding is very subtle in the figure (in contrast to the significant HDX change of the same segment in the protein upon PF-07321332 binding), thus making the strong conclusion that "This suggests that ebselen targeting C300 may induce structural changes in the C-terminal helical segment, weakening key hydrogen bonds at the dimer interface and ultimately inhibiting activity" not convincing. The reviewer would suggest the authors either delete this conclusion or largely tone it down.

      We thank the reviewer for the recognition of our efforts and agree with the reviewer’s suggestion. We have corrected “residues 296–304” to “residues 296–306” in Line 377 and removed the statement “This suggests that ebselen targeting C300 may induce structural changes in the C-terminal helical segment, weakening key hydrogen bonds at the dimer interface and ultimately inhibiting activity.”, as suggested.

    1. Author response:

      We thank the reviewers for their careful and constructive evaluation of our study. We are encouraged that the reviewers recognized the significance of DSB-induced genomic amplification (DIGA) and the evidence implicating DNA-end processing and recombination-associated DNA synthesis in this response.

      To our knowledge, this study provides the first description of DIGA as a large-scale increase in genomic DNA content following the induction of DSBs in cancer cells and represents an initial effort to define factors that regulate this phenomenon. The present work shows that DIGA can be induced by several sources of DSBs, involves de novo DNA synthesis, is genetically distinguishable from canonical CDT1-dependent origin re-licensing, is regulated by pathways controlling DNA-end protection and resection, and requires RAD51, RAD52, POLD3, and POLD4.

      At the same time, we agree that many important questions remain regarding the physical organization and genomic distribution of the additional DNA, the sites from which synthesis originates, the length and number of synthesis tracts, and the full determinants that render some cancer cells more susceptible to DIGA than others. We view these as important questions that arise from the initial characterization of this previously unrecognized phenotype and that will require substantial additional investigation.

      Reviewer #1:

      We agree that p53 status alone does not explain the considerable variation in DIGA observed among the cancer cell lines examined. Our experiments using isogenic HCT116 cells identify p53 as one factor capable of limiting DIGA, while the genetic studies implicate DNA-end protection and resection pathways as additional determinants. The present data, however, do not establish which of these or other pathways account for the differences among individual cancer cell lines. Defining the molecular basis for this variability will require systematic comparison of DIGA-prone and DIGA-resistant cells.

      The reviewer also raises an important question regarding the temporal relationship between DIGA and normal S-phase DNA replication. Our conclusion that the increase in DNA content involves de novo synthesis within the same cell-cycle interval is supported by BrdU incorporation in cells with >4N DNA content, the persistence of DIGA when progression through mitosis is blocked by nocodazole, and the detection of newly synthesized DNA in synchronized irradiated cells. These experiments do not, however, define precisely when DIGA-associated synthesis begins relative to normal S-phase replication. More detailed time-resolved analysis will be required to establish this relationship.

      We also agree that the present findings support a BIR-like mechanism rather than providing a complete physical reconstruction of classical BIR. The dependence of DIGA on DNA-end resection, RAD51, RAD52, POLD3, and POLD4 provides genetic evidence for recombination-associated DNA synthesis with features of BIR. Direct determination of synthesis-tract architecture, template usage, and genomic distribution will be required to define the underlying synthesis mechanism more completely.

      Reviewer #2:

      We agree that direct characterization of the additional DNA represents an important next step in understanding DIGA. The current study demonstrates a substantial increase in cellular DNA content, de novo DNA synthesis within the >4N population, and dependence on factors involved in DNA-end resection, strand invasion, and BIR-associated synthesis. These experiments do not determine which genomic regions are amplified or the length of individual synthesis tracts. Genomic analysis of cells undergoing DIGA should help determine whether the observed increase in DNA content reflects numerous amplification events, extensive synthesis from a subset of sites, or a different organization of the additional DNA.

      We appreciate the reviewer highlighting the study by Costantino et al. (2014), which provided important evidence that BIR-associated repair of damaged replication forks can generate segmental genomic duplications in human cells. The genomic alterations characterized in that study, however, arose under a substantially different experimental setting. Costantino et al. induced replication stress through cyclin E overexpression and analyzed copy-number alterations accumulated over a three-week period in clonally derived cells. Among these alterations, amplifications smaller than 200 kb were reduced following depletion of POLD3 or POLD4, leading the authors to propose that this subset of segmental duplications may represent BIR events, whereas larger amplifications and deletions could involve other repair mechanisms.

      DIGA, however, differs from these previously described alterations in several readily observable respects. DIGA develops over approximately one to three days following acute induction of DSBs by IR or AsiSI and produces increases in total cellular DNA content sufficiently large to be detected directly by flow cytometry. Thus, the two phenomena differ in their mode of induction, kinetics, and scale. At the same time, the involvement of POLD3 and other recombination-associated factors in both settings raises the possibility that they share aspects of the underlying DNA-synthesis machinery. The present data do not establish whether the additional DNA in DIGA consists of numerous segmental duplications, substantially longer synthesis products, or another genomic configuration.

      We also agree that our experiments do not directly demonstrate that DIGA-associated DNA synthesis initiates precisely at individual DSB sites. The ability of AsiSI-generated DSBs to induce DIGA, together with its dependence on DNA-end resection, RAD51, RAD52, POLD3, and POLD4, links the phenomenon closely to DSB processing. Direct mapping of newly synthesized DNA relative to defined DSBs will ultimately be required to determine where DIGA-associated synthesis originates.

      The reviewer asks whether MLN4924-induced re-replication and DIGA have been examined simultaneously. We have not examined this combination. Our distinction between these processes instead rests on their different genetic requirements. In particular, depletion of CDT1 strongly suppresses MLN4924-induced re-replication but does not suppress IR-induced DIGA, and DIGA is stimulated while rereplication is inhibited by the depletion of SET8. These observations argue against canonical CDT1-dependent origin re-licensing as the mechanism underlying DIGA, although they do not exclude more complex interactions between replication and DSB-associated DNA synthesis.

      We agree that the effects of XRCC4, XLF, and LIG4 are mechanistically intriguing and not yet fully understood. The present experiments establish that loss of XRCC4 or XLF, and to a lesser extent LIG4, suppresses DIGA, whereas loss or inhibition of DNA-PKcs enhances it. Stabilization of broken DNA ends by XRCC4/XLF is one possible interpretation, but effects on end resection, DSB persistence, repair-pathway choice, or other functions of these proteins could also contribute. The opposing effects of different components of the NHEJ machinery therefore identify an important mechanistic question that remains to be resolved.

      Finally, we agree that the correlation between DIGA and radiation sensitivity across the melanoma cell-line panel does not by itself demonstrate causality. The data establish an association between the propensity to undergo DIGA and sensitivity to IR. Because these cell lines differ in multiple additional properties that may influence the radiation response, matched models in which DIGA can be selectively altered will be important for determining the extent to which DIGA itself contributes to radiation-induced loss of proliferative capacity.

      Reviewer #3:

      We agree that the magnitude of the increase in DNA content is one of the most interesting unresolved features of DIGA. Previous analyses of BIR-associated synthesis at defined mammalian lesions have generally described synthesis events considerably smaller than the total increase in DNA content observed here. Our experiments do not establish the length or number of individual synthesis events responsible for DIGA. We therefore use the term BIR-like to describe the genetic requirements of the process rather than to imply that each DSB gives rise to a single exceptionally long BIR tract. Determining how many genomic sites participate and how much DNA is synthesized at individual sites will be important for understanding how the large increase in total DNA content is generated.

      Regarding the AsiSI BrdU experiments, it is important to note that BrdU was provided as a one-hour pulse immediately before harvesting. BrdU signal at 48, 72, or 96 hours therefore reports DNA synthesis occurring during that particular one-hour interval and does not measure the cumulative DNA synthesis that preceded the measurement. Consequently, relatively modest BrdU incorporation in cells that have already accumulated high DNA content does not indicate that the preceding increase occurred independently of DNA synthesis. Conversely, these experiments alone do not define the physical mechanism by which the additional DNA accumulated.

      We agree that the requirement for XRCC4, XLF, and LIG4 is unexpected under a simple model of BIR. As noted above, the present experiments establish this genetic relationship but do not define its molecular basis. The differential effects of DNA-PKcs and downstream NHEJ factors suggest that individual NHEJ components may influence DIGA through functions that are not adequately represented by viewing the pathway simply as a linear ligation reaction.

      Finally, we agree that the similarity between the flow-cytometric profiles produced by IR and MLN4924, as well as their common sensitivity to aphidicolin, does not by itself distinguish the underlying mechanisms. The distinction in the current study instead derives from their different genetic requirements, particularly the dependence of MLN4924-induced re-replication on CDT1 compared with the lack of such a requirement for DIGA, together with the differential effects of factors involved in DNA-end protection, resection, and DSB repair. These observations argue that DIGA is not simply the consequence of canonical origin re-licensing (i.e., rereplication or endoreduplication), while leaving open the possibility that additional replication-associated mechanisms contribute to the phenotype.

      In summary, we appreciate the reviewers highlighting several important mechanistic questions raised by our findings. We view these questions as natural extensions of the initial discovery and characterization of DIGA. The present study identifies a large-scale DSB-associated increase in genomic DNA content and establishes important roles for DNA-end protection, resection, strand invasion, and BIR-associated factors in regulating this response. Determining the genomic architecture of the additional DNA, the sites and molecular intermediates from which synthesis originates, and the cellular determinants of DIGA susceptibility will be important goals for future studies.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dong et al. present an in-depth analysis of mutant phenotypes of the Rab GTPases Rab5, Rab7, and Rab11 in Drosophila second-order olfactory neuron development. These three Rab GTPases are amongst the bestcharacterized Rab GTPases in eukaryotes and have been associated with major roles in early endosomes, late endosomes, and recycling endosomes, respectively. All three have been investigated in Drosophila neurons before; however, this study provides the most detailed characterization and comparison of mutant phenotypes for axonal and dendritic development of fly projection neurons to date. In addition, the authors provide excellent high-resolution data on the distribution of each of the three Rabs in developmental analyses.

      Strengths:

      The strength of the work lies in the detailed characterization and comparison of the different Rab mutants on projection neuron development, with clear differences for the three Rabs and by inference for the early, late, and recycling endosomal functions executed by each.

      Weaknesses:

      Some weakness derives from the fact that Rab5, Rab7, and Rab11 are, as acknowledged by the authors, somewhat pleiotropic, and their actual roles in projection neuron development are not addressed beyond the characterization of (mostly adult) mutant phenotypes and developmental expression.

      We would like to thank Reviewer #1 for their appreciation of our characterization of distinct Rab mutants.

      Reviewer #2 (Public review):

      Summary:

      This study by Dong et al. characterizes the roles of highly-expressed Rab GTPases Rab5, Rab7, and Rab11 in the development and wiring of olfactory projection neurons in Drosophila. This convincing descriptive study provides complementary approaches to Rab expression and localization profiling, conventional dominantnegative mutants, and clonal loss-of-function mutants to address the roles of different endosomal trafficking pathways across circuit development. They show distinct distributions and phenotypes for different Rabs. Overall, the study sets the stage for future mechanistic studies in this well-defined central neuron.

      Strengths:

      Beautiful imaging in central neurons demonstrates differential roles of 3 key Rab proteins in neuronal morphogenesis, as well as interesting patterns of subcellular endosome distribution. These descriptions will be critical for future mechanistic studies. The cell biology is well-written and explanatory, very accessible to a wide audience without sacrificing technical accuracy.

      Weaknesses:

      The Drosophila manipulations require more explanation in the main text to reach a wide audience.

      We appreciate Reviewer #2’s analysis of our work and thank them for their suggestions to improve the clarity of our manuscript.

      Reviewer #3 (Public review):

      Summary:

      The authors aimed at a comprehensive phenotypic characterization of the roles of all Rab proteins expressed in PN neurons in the developing Drosophila olfactory system. Important data are shown for a number of these Rabs with small/no phenotypes (in the Supplements) as well as the main endosomal Rabs, Rab5, 7, and 11 in the main figures.

      Strengths:

      The mosaic analysis is a great strength, allowing visualization of small clones or single neuron morphologies. This also allows some assessment of the cell autonomy of the observed phenotypes. The impact of the work lies in the comprehensiveness of the experiments. The rescue experiments are a strength.

      Weaknesses:

      The main weakness is that the experiments do not address the mechanisms that are affected by the loss of these Rab proteins, especially in terms of the most significant cargos. The insights thus do not extend far beyond what is already known from other work in many systems.

      We thank this reviewer for their feedback and appreciation of our genetic manipulations.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Consensus suggestions after discussion of all three reviewers:

      All three reviewers agree that the morphological and phenotypic analysis of the fly olfactory neurons is a strength of the manuscript. The shared perceived weakness is that the experiments do not address the mechanisms that are affected by the loss of these Rab proteins, especially in terms of the most significant cargos; the findings are in line with a large body of literature.

      The three reviewers feel that the manuscript could be strengthened greatly by adding data on an actual cargo (cell surface proteins?) and a more detailed analysis of the actual developmental origin (what happens when during axon and dendrite development) with respect to sorting of such cargo in the neurons they analyzed.

      We appreciate the time and effort of all three of our reviewers and share their interest in both identifying Rab-regulated cargos as well as determining the developmental origins of the Rab phenotypes. We have added three additional main figures (new Figure 4, Figure 8, and Figure 9), two supplemental figures (Figure 1—figure supplement 1 and 2), two additional supplemental tables (Table S2 and S3), and five additional panels (in Figure 3 H–L) of mutant developmental phenotype analysis.

      Regarding cargos, we also share the reviewers’ desire to identify cargos regulated by each Rab and made attempts to do so but were ultimately unable to achieve this goal. The main obstacles to this were: (1) it is not known which cell-surface proteins are most robustly endocytosed in PNs; without this knowledge it is difficult to identify candidates whose localization would predominantly reflect endosomal rather than plasma membrane distribution, making it challenging to detect changes in compartment-specific localization upon Rab perturbation; (2) reagents to evaluate cell-surface proteins in PNs are not cell-type-specific making it difficult to evaluate changes in their distribution in PNs; (3) tagged overexpressed proteins are either unavailable or expressed at levels too high to sensitively detect changes in their distribution. We have elaborated on each of these points below and feel that cargo identification, while an important future direction, is beyond the scope of the present study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      There are a number of experiments and ideas that the authors might consider to further improve on this work.

      (1) The idea, introduced by the authors in the introduction, that Rab-mediated recycling of cell surface proteins back to the membrane versus degradation is, of course, excellent and interesting. It is less clear how this applies to the present study. The functions of Rab5, Rab7, and Rab11 are so widespread, potentially affecting many signaling roles resulting in primary or secondary effects on membrane and even cytoskeletal regulation, that it remains unclear whether the mutant phenotypes are related to the recycling or degradation of cell surface receptors. To link the idea to experimentation, the authors have an excellent opportunity in their system to look at endogenously tagged cell surface proteins (or at least one example), many of which the Luo lab has characterized in these neurons, to minimally correlate cell surface protein defects to the observed developmental defects.

      We understand this critique and share this reviewer’s interest in identifying the specific cargos regulated by each Rab during development. We attempted to use antibodies to evaluate changes in cell-surface protein localization in response to disrupting individual Rabs but were unable to reliably distinguish shifts in association with specific endosomal compartment as many available antibodies label cell-surface proteins expressed in antennal lobe cells beyond projection neurons (such as olfactory receptor neurons, glia, or local interneurons) which complicates analyses.

      Additionally, although we have, in other work, generated multiple 'flp-on' tags for PN cell-surface proteins, these cannot be used in combination with the MARCM system, as it relies on a heat-shock-inducible flp to label singlePN clones. Heat shock would simultaneously induce tag expression in other cells expressing the tagged gene, preventing PN-type-specific detection. This incompatibility thus prevents us from simultaneously perturbing individual Rabs and tracking corresponding changes in surface-protein localization with single-cell resolution.

      Moreover, for proteins that are not highly endocytosed, it is difficult to separate plasma-membrane from endosomal localization, and we currently do not know which cell-surface proteins are most robustly endocytosed in PNs. Thus, while we share the reviewer’s interest in identifying candidate cargos, technological limitations make it difficult to achieve this goal within the scope of the current study.

      (2) The mutant phenotypes are mostly characterized based on adult outcomes. Maybe a little more can be learned about when and how Rab5, Rab7, or Rab11 function is locally required by characterizing the developmental processes that lead to, e.g., aberrant dendritic development in Rab5 and Rab11.

      We also feel that charting the developmental origins of Rab mutant phenotypes is important. Prior to mid-pupal stage (around 48 hours after puparium formation), glomeruli in the antennal lobe have not yet assumed their stereotyped positions, which complicates analyses and interpretation; thus, many of our analyses are conducted at the adult stage. For Rab11 mutants we did perform many developmental analyses to evaluate the origins of the axonal development (Figure 6—figure supplement 1) and dendrite elaboration phenotypes (Figure 5 J–L) we observed at the adult stage. We realize that the developing axonal analyses were in supplemental material where they could be missed. We have moved these data to the main figures (Figures 8 and 9) and emphasized these analyses. Further, we extended our Rab5 mutant analyses to evaluate developmental phenotypes (Figure 3H– L and Figure 4). We believe that these new analyses have strengthened the manuscript.

      (3) Regarding the subcellular localization analyses: a collection of endogenously tagged Rabs in Drosophila has been generated by Dunst et al. (2015), which is surprisingly not cited. Maybe the authors could consider looking at the endogenous localization of Rabs using the tagged version in parallel to their overexpressed tagged versions.

      We have now cited and discussed this paper (line 82) and thank the reviewer for pointing out this omission. We previously attempted to evaluate these endogenously tagged Rab proteins in PNs; however, since PN dendrites project into the antennal lobe, a dense neuropil region containing PN dendrites, ORN axons, glial processes, and neurites of local interneurons, we are unable to resolve individual Rab puncta from cytosolic (non-vesicle associated Rabs) or evaluate Rab localization in a cell-type-specific manner. For this reason, we focused on evaluating the localization of tagged Rab proteins from UAS-transgenes using a MARCM rescue strategy. We directly addressed this in the text (starting on line 83).

      (4) It is maybe not entirely surprising that Rab5 and Rab11 have the strongest phenotypes, as these have been implicated in early and recycling endosomal processes in basically all eukaryotic cells with major implications for signaling throughout development and function, often causing cell death (and in the case of Rab5, tumorigenic phenotypes in flies). By contrast, the Drosophila brain can develop in the absence of Rab7 (Cherry et al., 2013; also the reference for the Rab7 null mutant, not Chan et al., 2011). A key concern in any developing fly cell rendered mutant using clonal analysis is the perdurance of RNA or protein (ultimately even maternal contribution), which could be addressed by discussion or experimentally.

      We thank the reviewer for pointing out our citation error, which we have now corrected.

      However, we note that Cherry et al. (2013) found that loss of Rab7 causes pupal lethality at stages prior to completion of 50–80% of development and can also cause embryonic lethality when maternal Rab7 contribution is blocked. This indicates that the whole organism cannot fully develop in the absence of Rab7. And while Cherry and colleagues did evaluate overall brain morphology in Rab7 mutant pupae, they did not look at the development of individual cell types. So, it is still unclear how loss of this GTPase affects the development of individual central nervous system neurons.

      Thus, to understand whether Rab7 has functions in PN development, we used the QMARCM system to perform Rab7 LOF analyses in PN clones. While we did not observe any phenotypes in single-cell MARCM clones (Figure 6), we did see mild defects in neuroblast clones (Figure 6—figure supplement 1A–C). Since single-cell MARCM clones are more susceptible to RNA/protein perdurance, we further evaluated Rab7 function by expressing a Rab7 dominant-negative transgene in DL1-PNs using a DL1-specific GAL4 driver (Figure 6—figure supplement 2), which circumvents potential perdurance issues mentioned by this reviewer. Importantly, this same transgene produces dendrite targeting defects when expressed in all PNs (Figure 1I), confirming its efficacy. However, no phenotypes were observed when expression was restricted to DL1-PNs, suggesting that Rab7 may not be required in DL1-PNs for their dendrite targeting. Given that both Rab7 mutant neuroblast clones and pan-PN expression of Rab7 dominant negative causes PN dendrite targeting defects we conclude that Rab7 is nonautonomously required for dendrite targeting of DL1-PNs.

      We have softened our language with regards to the Rab7 analysis and have emphasized, and strengthened, our previous discussion of these points beginning on line 268 in the results section and on line 427 of the discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 1B, it would be useful to show the circuit over multiple developmental stages, rather than just in its final form.

      We have added this to Figure 1; it is now panel C. Thank you for this suggestion.

      (2) Expression analysis of Rabs in Figure 1C - how do these levels and ratios compare to the whole brain? Whole body?

      Unfortunately, we are unable to evaluate how Rab expression in PNs compares to all other cells in the brain as there is no sequencing data available for this organ at this time point. We did compare the expression of endosomal Rabs between PNs and their presynaptic targets, ORNs. We found that many Rabs displayed similar expression patterns between these two cell types during development. We have added a new paragraph on this, beginning on line 100 and we added two additional supplemental figures (Figure 1—figure supplement 1 and 2).

      (3) The authors should include at least a few sentences comparing the current approach and results to previous comprehensive Rab protein expression analysis, for example, in PMID 17409086, 22000105, 22844416, and 33666175.

      Thank you for pointing out this omission, we have amended it beginning on line 82.

      (4) For the non-Drosophila reader (for example, a cell biologist working on endosomal traffic in cultured neurons), the paper is less accessible. Some examples:

      We thank this reviewer for their suggestions for ways to clarify our work for the non-Drosophila reader. We have addressed each of their points.

      (a) The severity difference between Rab5, Rab11, Rab7 and Rab4, Rab 21, Rab35 isn't immediately obvious from the images to someone who doesn't work with this system - does the brightness of the ectopic growths indicate the number of ectopically grown processes? It might help to have half a sentence to make this difference more accessible for the readers who aren't familiar with this system.

      We have clarified this beginning on line 112.

      (b) Figures 1D-E require more extensive description of the experimental setup with orthogonal expression systems than is provided briefly in the cartoon, figure legend, and supplement. For example, it should be noted what white vs blue represents in the marked glomeruli.

      We have clarified this point beginning on line 114.

      (c) There should be at least one sentence introducing what is marked and what it means when MARCM clones are first shown in Figure 2B, in addition to the supplemental figure.

      We have added a detailed explanation of MARCM on line 155.

      (d) It's not clear to a non-expert what the meaning is of no innervation of non-adPN glomeruli in wild-type in Figure 2E. This requires a sentence of explanation.

      We have added additional details about this on line 159 and 169.

      (e) Can a control image be shown for the experiment in 3B?

      We have added an additional set of control images in Figure 3B on the left.

      (5) The experiment measuring axonal projection to the lateral horn in Rab5 clones in Figure 3 J-L is underpowered (n=3 for mutant). While this may be due to the frequency of an overall projection defect as shown in Figure 3B, it makes it difficult to assess the robustness of the terminal phenotype. Further, for clarity, similar measurements (e.g., bouton width) should be aligned vertically between E-G and J-L.

      We have performed additional analyses on Rab5 axons in the lateral horn and added them to a new Figure 4. The n’s are now n=10 for controls and n=7 for Rab5 mutants. Additionally, we have aligned similar measurements in the figure panels and standardized the axes of each graph so that it is easier to compare between developmental stages.

      (6) The argument that cell-type-specific phenotypes are due to distinct cargoes is weak. The same cargo could have different functions or signaling properties in different cell types (e.g., "Taken together, the distinct branching phenotypes observed in the mushroom body versus lateral horn suggest that Rab5 may regulate the trafficking of a distinct set of cargos in each axonal compartment"). Similarly, this argument is just one of many possibilities, as the effect could be quite indirect (eg via mis-regulated signal transduction): "Yet, the terminal boutons in Rab5 mutants were nearly 2-fold larger than those of controls (Figure 3G), suggesting Rab5 regulates the trafficking of cell-surface proteins that normally restrain bouton growth " and "Page 12 "Rab7mediated degradation does not have a major role in regulating axon or dendrite development" - should be softened since Rab7 may easily play an important but redundant role.

      We have made all of these changes and removed references to trafficking of specific CSPs.

      (7) Statistics need to be added to: Figure 3B, Figure 5D-E, Figure 2 - figure supplement 1C, Figure 4 - figure supplement 1E, F, Figure 5 - figure supplement 1A, E, Figure 6 - figure supplement 1K.

      We thank this reviewer for pointing out this omission, we have added statistical measurements to our graphs.

      As to not visually overwhelm readers with statistical measurements on already dense graphs (such as Figure 7B), we have added two supplemental tables (Table S2 and S3) that display all the results of all of the comparisons performed in the statistical tests.

      We have cited this table in the figure legends and in-text figure references.

      In addition to the methods, we have also added the exact statistical test and post-hoc tests (when applicable) to the figure legends.

      (8) The BSDC identifier for the UAS-Rab11-mCherry stock may be incorrect.

      It appears that some of the values in the ‘Identifiers’ column of the Key Resources table shifted downward. We have fixed this and appreciate that the reviewer pointed this out.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Our revision includes:

      (1) The generated data from long-read whole-genome sequencing of 1000 Genomes Project samples, including FASTQ files, SV calls, and the imputation panel, are now openly available via ENA and OpnMe. The imputed structural variant data have been submitted to UK Biobank for release through the UK Biobank Research Analysis Platform, subject to UK Biobank release procedures. SV-WAS summary statistics have been made available via OpnMe.

      (2) Clarification of analyses and methods, addition of two new Supplementary Figures, and correction of minor issues throughout the manuscript.

      (3) A significantly expanded Discussion to address the reviewers’ comments and better contextualise our methods and results.

      eLife Assessment

      This fundamental work significantly enhances our understanding of how structural variants influence human phenotypes. The conclusion is convincingly supported by rigorous analyses of long-read sequencing data. If the raw data are made publicly available, these high-quality datasets and findings will further advance our knowledge of genetic variation in the human population.

      We thank the editors for this positive assessment of our work. The raw long-read sequencing data (FASTQ files) can now be accessed through the European Nucleotide Archive (ENA) under accession number PRJEB89727, as part of a larger collection of 1019 sequenced probands from the 1000 Genomes Project (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The generated imputation panel and the structural variant calls, based on the 888 probands used in the present manuscript, remain freely available for download at https://opnme.com/genomiclens. We have now added the summary statistics of 32 SV-wide association studies to the same resource. In addition, we have submitted the imputed SV genotypes for UK Biobank participants to the UK Biobank; once processed by UK Biobank, these genotypes will be released via the UK Biobank Research Analysis Platform (RAP).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors sequenced 888 individuals from the 1000 Genomes Project using the Oxford Nanopore long-read sequencing method to achieve highly sensitive, genome-wide detection of structural variants (SVs) at the population level. They conducted solid benchmarking of SV calling and systematically characterized the identified SVs. While short-read sequencing methods, including those used in the 1000 Genomes Project, have been widely applied, they exhibit high accuracy in detecting single nucleotide variants (SNVs) and small insertions and deletions but have limited sensitivity for SV detection. This study significantly enhances SV detection capabilities, establishing it as a valuable resource for human genetic research. Furthermore, the authors constructed an SV imputation panel using the generated data and imputed SVs in 488,130 individuals from the UK Biobank. They then conducted a proof-of-principle genome-wide association study (GWAS) analysis based on the imputed SVs and selected traits within the UK Biobank. Their findings demonstrate that incorporating SV-GWAS analysis provides additional insights beyond conventional GWAS frameworks focusing on SNVs, particularly in improving fine mapping.

      The authors constructed a high-sensitivity reference panel of genome-wide SVs at the population level, addressing a critical gap in the field of human genetics. This resource is expected to significantly advance research in human genetics. They demonstrated the imputation of SVs in individuals from the UK Biobank using this panel and conducted a proof-of-concept SV-based GWAS. Their findings highlight a novel and effective strategy for integrating SVs into GWAS, which will facilitate the analysis of human genetic data from the UK Biobank and other datasets. Their conclusions are supported by comprehensive analyses.

      We thank the reviewer for highlighting the value of our SV imputation reference panel.

      Weaknesses:

      (1) Although the authors employ state-of-the-art analytical approaches for the identification of SVs, the overall accuracy remains suboptimal, as indicated by an F1 score of 74.0%, particularly in tandem repeat regions. To enhance accuracy, it would be beneficial to explore alternative SV detection methods or develop novel approaches. Given the value of the reference panel and the fact that improved SV accuracy would lead to more precise SV imputation and GWAS results, investing effort in methodological refinement is highly encouraged.

      Accurate SV calling remains an active area of research and is beyond the scope of the present study. Tandem repeat regions are particularly challenging for standardised SV detection. We believe that achieving a benchmark for NA12878 of F1 = 74% on a genome-wide level and, notably, F1 = 91% when excluding longer tandem repeats, represents strong performance. This result is especially convincing when considering that our benchmarking compared the SV calls to data generated using a different sequencing technology and processed using different bioinformatics pipelines.

      (2) From the Methods section, it appears that the authors employed Beagle for both imputation and the UK Biobank imputation.

      (a) It would be better to explicitly clarify this in the Results section and provide a detailed description of the corresponding procedures and parameters in the Methods section for both analyses, as this represents a key aspect of the study.

      We thank the reviewer for these suggestions. Accordingly, we added the clarification to the Methods section that the leave-one-out imputation used exactly the same pipeline and settings as the UK Biobank imputation (page 14, section “Leave-one-out imputation performance”):

      “We excluded one individual from the panel and imputed SVs for this individual using the panel of the remaining 887 samples, applying exactly the same pipeline and settings as those later used for SV imputation into UK Biobank (see below).”

      (b) Additionally, Beagle is not specifically designed for SV imputation, the imputation quality of SVs is generally lower than that of SNVs. Exploring strategies to improve SV imputation, such as developing a novel method with reference panel data, may enhance performance.

      As stated in our manuscript (page 4), we believe that, in our study, the imputation quality of SVs is lower than of that of SNVs primarily because of a) the greater difficulty of SV calling compared to SNV genotype calling and b) the heterogeneity in SV representation across samples. Once SVs are encoded as bi-allelic markers in the reference panel, they can be imputed using the same LD/haplotype-based framework as any other variants. Accordingly, improving SV imputation is likely to benefit most from more accurate upstream SV calling and genotyping (e.g., through more robust multi-sample calling and harmonised variant representations) and not so much from improved or SV-specific imputation methods. While improved imputation is an important research direction, it is beyond the scope of the present manuscript.

      (c) It is also important to assess how this reduced imputation quality may influence GWAS results. For instance, it would be useful to examine whether associated SVs exhibit higher imputation quality and whether SVs with lower quality are less likely to achieve significant association signals. In addition, the lower imputation quality observed for INV, DUP, and BND variants (Figure 3) may be due to their greater lengths (Figure 2). It is better to investigate the relationship between SV length and imputation quality.

      We agree that imputation quality can influence GWAS results. For example, for the FEV1/FVC phenotype, SVs with INFO > 0.9 are almost twice as likely to reach genome-wide significance (p < 5e-8) compared with SVs with 0.7 < INFO < 0.9 (odds ratio 1.95; Fisher’s exact test p-value 2.5e-5). This is consistent with the intuitive notion (applicable to any variants, not only to SVs) that greater uncertainty in the imputed genotypes dilutes association signals and therefore reduces power. For a detailed discussion of the relationship between allele frequency, imputation accuracy, and GWAS association results, see Zhang et al., Human Molecular Genetics 31(1):146–155 (2022), https://doi.org/10.1093/hmg/ddab203

      We have now investigated the relationship between SV length and imputation quality (the new Supplementary Figure 6). The results suggest that the observed association between imputation quality and SV size is primarily driven by the SV-size–dependent minor allele frequency in the imputation panel.

      (3) All examples presented in the manuscript focus on SVs that overlap with genes. It may also be valuable to investigate SVs that do not overlap with genes but intersect with enhancer regions. SVs can contribute to disease by altering regulatory elements, such as enhancers, which play a crucial role in gene expression. Including such analyses would further demonstrate the utility of SV-GWAS and provide deeper insights into the functional impact of SVs.

      We agree with the reviewer that examining SVs intersecting with enhancer regions could be an interesting direction for future studies, as it would provide additional insights into regulatory mechanisms and disease associations. However, in the present proof-of-principle study, we prefer focusing on SVs overlapping with genes and have now highlighted this in additional detail in the revised manuscript (Discussion, page 7):

      “In the present proof-of-principle study, we focused on SVs overlapping with the coding sequence of genes. In future applications of our SV imputation panel, more refined gene mapping approaches could be employed, e.g., including SVs overlapping enhancer regions or epigenetic marks. Such an enhanced mapping would increase the number of identified associated genes and thus provide additional insights into regulatory mechanisms and disease biology.”

      (4) The data availability link currently provides only a VCF file ("sniffles2_joint_sv_calls.vcf.gz") containing the identified SVs.

      (a) It would be beneficial for the authors to make all raw sequencing data (FASTQ files) and key processed datasets (such as alignment results and merged SV and SNV files) available. Providing these resources would enable other researchers to develop improved SV detection and imputation methods or conduct further genetic analyses.

      Thank you for emphasising the importance of data sharing, which we agree with.

      The Data Availability section of the manuscript already includes a link to https://opnme.com/genomiclens, where we made both the SV calls and the full and reduced SV imputation panel files freely available. We have now added SV summary statistics from 32 SVwide association studies to the same resource. Based on the reviewer’s request, we now also reference the ENA repository project PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727), where the raw FASTQ files are available for download, in the manuscript.

      We have appended the Data Availability statement on page 22 of the revised manuscript as follows:

      “Raw SV calls, the long-read sequencing-based SV imputation panel, and the SV summary statistics from 32 SV-wide association studies are available through the OpnMe initiative of Boehringer Ingelheim GmbH (https://opnme.com/genomiclens). The raw long-read sequencing data (FASTQ files) for the 1000 Genomes Project samples included in this study are accessible via the European Nucleotide Archive under accession number PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The dataset analysed here constitutes a subset of this broader collection.”

      (b) Furthermore, establishing a dedicated website for data access, along with a genome browser for SV visualization, could significantly enhance the impact and accessibility of the study. Additionally, all code, particularly the SV imputation pipeline accompanied by a detailed tutorial, should be deposited in a public repository such as GitHub. This would support researchers in imputing SVs and conducting SV-GWAS on their own datasets.

      The Methods section provides a full and detailed description of the imputation pipeline and parameters in the section “Preprocessing and imputation of SVs into UK Biobank” on page 14 of the revised manuscript. Our data processing simply consisted of running standard bioinformatics tools with the parameters exactly as described in the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) In the Results section, Figure 3b is mentioned before 3a, and it is better to switch them in the Figure.

      Thank you for highlighting this fact. We acknowledge that typically the sequence of sections matches exactly between text and figures. However, in this specific case, we would prefer to deviate from the norm: In our opinion, Figure 3 is easier to interpret in its current sequence. At the same time, the text flows more logically in its current sequence, describing 3b before 3a. We would therefore prefer to stick to the current order, even if it means that Fig. 3b is described before 3a in the text.

      (2) Page 10, "Figure 1e" -> "Figure 2e".

      Thank you, we corrected this issue.

      (3) Page 14, "Leave-one out" -> "Leave-one-out".

      Thank you, we corrected this mistake.

      (4) It is better not to use abbreviations in the subheadings, especially "UKB" (page 3).

      Thank you, we changed the acronym ‘UKB’ to ‘UK Biobank’ in all subheadings.

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to develop a novel and efficient method for SV detection, utilizing data from the 1000 Genomes Project (1KGP) for modeling and calibration. This method was subsequently validated using UK population data and applied to identify structural variants associated with specific disease phenotypes.

      Strengths:

      Third-generation single-molecule sequencing data offers several advantages over traditional high-throughput sequencing methods, particularly due to its long-read lengths, which provide valuable insights into significant forms of genomic variation. The authors have developed an efficient method for detecting structural variations and optimizing the utilization of genomic data. We hope that this method will continue to be refined, enabling researchers to more effectively leverage long-read data, high-throughput data, or even a synergistic combination of both.

      Weaknesses:

      Although this research contributes to our ability to more effectively utilize long-length and high-throughput data, there are some key issues that need to be addressed in terms of analyzing the specific results as well as writing the article.

      Reviewer #2 (Recommendations for the authors):

      (1) How to discuss the lower detection rate of structural variations (SVs) in East Asian populations, it is worth considering whether the authors' training dataset, which may have been based on raw data with insufficient representation of East Asian individuals, could have introduced a bias favoring other populations. This potential bias might arise from the relatively limited data available for Asian ancestry. Alternatively, the observed differences could also be influenced by the role of natural selection, which may have shaped the genomic landscape of East Asian populations in distinct ways. Further investigation is needed to clarify these possibilities.

      Thank you for raising this important point. Although an interesting research direction, a detailed investigation of the factors affecting SV detection rates is beyond the scope of the present study. However, we do not think that the lower detection rate in East Asians is due to an underrepresentation of Asian ancestry in our dataset. To explain this to all readers, we have added the following explanation to page 7 of the Discussion:

      “In this context, we observed that the number of SVs detected per individual differed between superpopulations. We identified the highest average number of SVs in individuals of African descent and a slightly lower average in East Asians, compared to the other superpopulations. While we included a higher number of African ancestry individuals, the number of East Asian individuals included in our reference panel was comparable to the number of individuals from other, non-African ancestries. In fact, it was even larger than the number of European ancestry individuals (AFR n=241, SAS n=171; EAS n=168; EUR n=164; AMR n=144). Therefore, we do not expect a major bias from underrepresentation of any superpopulation in the training dataset. It is well established that African ancestry is more diverse than is the case for other superpopulations [32, 33] and previous studies indicate that East Asian populations tend to exhibit slightly lower genetic diversity compared to European populations [34], which is consistent with the lower observed SV counts per genome.”

      (2) The authors did not present the results of the detection of CNV.

      Copy number variations (CNVs) are considered a subclass of structural variants. In our analysis, we detected deletions and duplications, which represent the most common forms of CNVs. However, we did not specifically investigate high copy-number SVs, as these are often larger than what can be reliably detected using long-read sequencing. Large-scale CNVs are typically identified in biobank studies through analysis of intensity data from genotyping microarrays using tools like PennCNV, and there is extensive literature supporting the use of this microarray approach in UK Biobank and other genotyped cohorts, see for example Aguirre et al.: Phenomewide Burden of Copy-Number Variation in the UK Biobank. Am J Hum Genet. 2019, Aug 1;105(2):373-383. doi:10.1016/j.ajhg.2019.07.001.

      (3) Multiple testing correction is essential for ensuring the validity of large-scale structural variation (SV) association analyses. It is strongly recommended that the statistical methods and correction strategies employed, such as Bonferroni correction or false discovery rate (FDR) control, be explicitly detailed to enhance the transparency and reliability of the findings.

      For genome-wide SV association analyses, we applied the commonly used genome-wide significance threshold of 5e-8, which is standard in genome-wide studies. Given that these were exploratory proof-of-principle analyses illustrating use cases for SV analyses, we decided not to correct on top of that for multiple testing for the number of traits (32) tested. For the pQTL analyses, we further adjusted this threshold using a Bonferroni-type correction based on the number of proteins tested (1,463), to account for the increased number of multiple comparisons.

      We have now added a more detailed description of this multiple testing procedure to the Methods subsection “SV-wide association studies in UK Biobank” on page 16 of the revised manuscript:

      “In the exploratory SV-WAS, we used the standard threshold for genome-wide significance of p < 5×10<sup>-8</sup>. For the pQTL analyses, we applied Bonferroni correction for multiple testing on top of that genome-wide threshold, correcting for the number of tested protein levels (n=1463): p < 5×10<sup>-8</sup>/ 1463 = 3.4×10<sup>-11</sup>.”

      (4) The study primarily relied on data from the 1000 Genomes Project (1KGP) and the UK Biobank; however, the UK Biobank cohort is predominantly composed of individuals of European ancestry, which may restrict the generalizability of the research findings to other populations.

      Our reference panel was constructed to cover multiple ancestries, enabling imputation for diverse populations. Thus, our imputation panel can be applied to biobanks around the world and is freely available for this purpose. As a proof of principle, we have demonstrated the feasibility and performance of SV imputation in UK Biobank as an example of a broadly accessible cohort. We are looking forward to biobanks from diverse ancestries downloading our imputation panel and applying it to their populations.

      (5) Although the study employed long-read sequencing technology, the validation of structural variation (SV) detection accuracy predominantly relied on internal data, such as 'leave-one-out' validation. To further strengthen the reliability of the SV detection methods, it is recommended to incorporate additional external independent datasets for validation.

      The leave-one-out procedure in our study was used to validate the imputation performance, not the accuracy of SV detection. To assess SV calling accuracy, we performed extensive benchmarking against external SV call datasets derived from PacBio long-read sequencing and Illumina short-read sequencing. These details are provided under the subheading ‘Structural variant calling and benchmarking’ in the Results section on page 2 of the manuscript.

      (6) Some of the SVs mentioned in the study overlap with disease association loci in the GWAS Catalog, but functional annotation and exploration of the biological mechanisms of these SVs are more limited. It is suggested that LD can be added to analyse whether there are SNP that are highly linked to them to further explore their functions.

      We thank the reviewer for this suggestion. We have actually conducted an analysis addressing exactly this question: We performed conditional association analyses of the SV signals with nearby short variants (SNPs and InDels) at the SV locus. Such a conditional analysis addresses whether the observed SV association is influenced by LD-correlated SNPs or not. The results of this analysis are reported in Supplementary Tables 16 and 17. These tables include both the conditional analysis results and the LD between each SV and the variant at the locus with the second-highest evidence for an association.

      Researchers interested in exploring the biological significance of the SV-WAS results in more detail can now download the full SV-WAS summary statistics from https://opnme.com/genomiclens.

      (7) The discussion section could be further expanded to explore the role of SV in complex diseases and its potential application in precision medicine. For example, it could discuss how SV information can be integrated into existing GWAS frameworks to enhance the accuracy of disease risk prediction.

      Thank you for the suggestion, we have now added the following sentences to the discussion (page 7/8):

      “Structural variants can influence complex disease biology through either the disruption of coding sequence or an altered regulation of gene expression. Such effects may not be well captured by short variants alone. Incorporating SVs into GWAS and follow-up analyses would thus provide more accurate disease risk prediction, uncover underlying pathomechanisms by highlighting actionable pathways and targets, and support precision medicine by providing biomarkers for patient stratification.”

      (8) The geographic labeling of certain samples in Figure 2 appears to contain inaccuracies. For instance, the CDX sample, which represents the Dai population from Xishuangbanna in China's Yunnan Province, is currently mislabeled as originating from China's Inner Mongolia. This discrepancy should be corrected to ensure the accuracy of the data representation.

      We apologise for the misunderstanding. The geographic map in Figure 2a serves as an illustrative mapping of the samples to countries. It is intended to provide readers with an overview of population coverage, rather than to indicate the precise geographic origins of individual populations. The populations CDX, CHB, and CHS are displayed within the outline of China in alphabetical order, without any intention to indicate their exact geographic origin. We changed the respective figure caption to make this clear (page 19 of the revised manuscript):

      “Map of the 888 samples from the 1000 Genomes project, mapping the samples to countries and not indicating detailed geographical origins of populations.”

      Reviewer #3 (Public review):

      This study successfully identified genetic loci associated with various traits by generating large-scale long-read sequencing data from a diverse set of samples. This study is significant because it not only produces large-scale long-read genome sequencing data but also demonstrates its application in actual genetics research. Given its potential utility in various fields, this study is expected to make a valuable contribution to the academic community and to this journal. However, there are several critical aspects that could be improved. Below are specific comments for consideration.

      Strengths:

      Producing high-quality, large-scale variant datasets and imputation datasets

      Weaknesses:

      (1) Data availability

      Currently, it appears that only the Genomic Lens SV Panel is available on the webpage described in the Data Availability section. It is unclear whether the authors intend to release the raw sequencing data. Since the study utilized samples from the 1000 Genomes Project, there should be no restriction on making the data publicly accessible. Given this, would the authors consider making the raw sequencing reads publicly available? If so, NCBI SRA or EBI ENA would be the most appropriate repositories for data deposition. I strongly encourage the authors to consider public data release. Additionally, accessing the Genomic Lens SV Panel data does not seem straightforward. The manuscript should provide a more detailed description of how researchers can access and utilize these data. In my opinion, the best approach would be to upload the variant data (VCF files) to a public database such as the European Variation Archive (EVA) hosted by EBI.

      I strongly request that the authors publicly deposit the variant data. At a minimum:

      (a) The joint genotype data for all 888 samples from the 1000 Genomes Project must be publicly available.

      Thank you for emphasising the importance of data sharing, which we agree with.

      The Data Availability section of the manuscript already includes a link to https://opnme.com/genomiclens, where we make both the SV calls and the full and reduced SV imputation panels (provided as multi-sample VCF files) freely available. Based on the reviewer’s request, we now also reference the ENA repository project PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727), where the raw FASTQ files are available for download.

      We have appended the Data Availability statement on page 22 of the revised manuscript as follows:

      “Raw SV calls, the long-read sequencing-based SV imputation panel, and the SV summary statistics from 32 SV-wide association studies are available through the OpnMe initiative of Boehringer Ingelheim GmbH (https://opnme.com/genomiclens). The raw long-read sequencing data (FASTQ files) for the 1000 Genomes Project samples included in this study are accessible via the European Nucleotide Archive under accession number PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The dataset analysed here constitutes a subset of this broader collection.”

      (b) For the UK Biobank samples, at least allele frequency data should be disclosed.

      Supplementary Table 5 includes the allele frequencies of the SVs imputed into UK Biobank.

      (c) Since eLife has a well-established data-sharing policy, compliance with these guidelines is essential for publication in this journal.

      By sharing the FASTQ files, the SV calls, the SV imputation panels, the SV summary statistics, and (once processed by UK Biobank) the genotypes of SVs imputed into UK Biobank, we are providing all SV data generated in our study.

      (2) Long-read sequencing data quality

      While the manuscript presents N50 read length and mean or median read base quality for each sample in a table, it would be highly beneficial to visualize these data in figures as well. A violin plot or similar visualization summarizing these distributions would significantly improve data presentation.

      Notably, the base quality of ONT long-read sequencing data appears lower than expected. This may be attributed to the use of pore version 9.4.1, but the unexpectedly low base quality still warrants attention. It would be helpful to include a small figure within Figure 2 to illustrate this point. A visual representation of read length distribution and base quality distribution would strengthen the manuscript.

      We thank the reviewer for this suggestion. We have now included two violin plots (the new Supplementary Figure 1) to the revised manuscript, summarising a) the N50 read length per sequencing run and b) the median read quality per sequencing run. These plots provide a clearer visualisation of the underlying distributions. We do not consider the ONT base quality to be low. Importantly, structural variant detection is generally robust to modest variations of per-base quality. Therefore, we do not expect the observed base quality levels to significantly affect SV calling in this study.

      (3) Variant detection precision, recall, and F1 score

      This study focuses on insertions and deletions (indels) {greater than or equal to}50 bp, but it remains unclear how well variants <50 bp are detected. I am particularly interested in the precision, recall, and F1 score for variants between 5-49 bp.

      While ONT base quality is relatively low, single-base variants are challenging to analyze, but variants {greater than or equal to}5 bp should still be detectable as their read accuracy is still approximately 90%, making analysis feasible. Given that Sniffles supports the detection of variants as small as 1 bp, I strongly encourage the authors to conduct an additional analysis.

      A simple two-category classification (e.g., 5-49 bp and {greater than or equal to}50 bp) should suffice. Additionally, a comparative analysis with HiFi and short-read sequencing data would be highly valuable. If possible, I strongly recommend that all detected variants {greater than or equal to}5 bp be made publicly available as VCF files.

      Because short InDels are available from high-coverage Illumina sequencing data generated for the same individuals (i.e., the data referred to as the NYGC dataset in our manuscript), we decided against calling such short variants from our lower coverage Oxford Nanopore data and thus concentrated our efforts on reliably calling longer variants covering at least 50 bp, consistent with the conventional definition of structural variants.

      (4) Assembly-based methods

      Given the low read accuracy and low sequencing depth in this dataset, it is understandable that genome assembly is challenging. However, the latest high-quality human genome datasets-such as those produced by the Human Pangenome Reference Consortium (HPRC)demonstrate that assembly-based approaches provide significant advantages, particularly for resolving complex and long structural variants.

      Since HPRC data also utilize 1000 Genomes Project samples, it would be highly informative to compare the accuracy of ONT sequencing in this study with HPRC's assembly-based genome data. The recent publication on 47 HPRC samples provides a valuable reference for such a comparison. Given its relevance, the authors should consider providing a comparative analysis with HPRC data.

      The aim of the present study was to generate an SV reference panel that enables SV imputation for large biobanks. Detailed assessments of ONT sequencing quality in general and comparisons to other sequencing efforts and technologies are out of scope for the present manuscript. We invite the scientific community to use the FASTQ files provided at ENA for conducting such detailed assessments in follow-up studies.

    1. Author response:

      The following is the authors’ response to the original reviews.

      General comments:

      You will see that many of the reviewers’ comments overlap. From our discussion with them, we agree that several of these comments should be addressed in this study, particularly comments related to the interpretation of the effect of drugs acting on the cellular cytoskeleton (reviewers #1 and #2). We also agree that the comparison of isogenic cell lines such as the mcf10a series or the 4T1 series should address some of the concerns regarding the interpretation of the mechanical fingerprint (reviewer #3). Also, certain methodological aspects should be easily clarified (reviewers #1 and #3).

      We also agreed that other comments may be more difficult to address in the context of this study. This is the case for comments related to establishing a link between different mechanical signatures and different cellular functions/outcomes (Reviewers #1 and #3). One could test whether migration or proliferation is altered by changing the mechanical fingerprint, or you could simply discuss these aspects by carefully reviewing the literature to corroborate mechanical signatures with known cellular phenotypes (e.g. migration speed, adhesion, cell size...). This is also the case for comments on the influence of other cellular parameters such as molecular crowding and energy metabolism, which could be left for future work or where you could use a low dose of cycloheximide (below the level of deleterious effects) to address the effect of cytoplasmic proteins (reviewer #3).

      We thank the editor for providing this helpful overview of the requested revisions. We have carefully addressed these points throughout the revised manuscript. The only difficulty was to establish the isogenic cell lines as requested. It took us over 18 months to find a source of these cells in Europe, and since then we are trying hard, but not successful to get these cells stably growing in the condition necessary for the optical tweezers experiments. As we have now spent more than 2 years on this without success, we decided to resubmit the paper without this part to not further delay this manuscript. The additional experiments and revisions have substantially strengthened the manuscript. Especially, the addition of Latrunculin A as suggested was an excellent request, as now the results regarding actin depolymerization and mechanical properties are in excellent agreement with the expected effects, as Latrunculin A is much more efficient in depolymerizing actin than cytochalasin B. The major changes are summarized below, followed by a detailed point-by-point response to all reviewer comments.

      General changes

      (1) Repeated measurements on wild-type HeLa cells.

      (2) Repeated all Cytochalasin B and Nocodazole experiments and increased the number of analyzed cells to approximately 60 per condition.

      (3) Performed additional experiments using Latrunculin A and combined Latrunculin A + Nocodazole treatment.

      (4) Performed immunostainings for all cytoskeletal perturbation conditions (WT, Cytochalasin B, Latrunculin A, Nocodazole, Cytochalasin B + Nocodazole, and Latrunculin A + Nocodazole).

      (5) Refined the rheological analysis procedure and expanded the methodological description.

      (6) Revised the manuscript text throughout and expanded the discussion of limitations and biological interpretation.

      Public Reviews:

      Reviewer #1 (Public Review):

      A limit of the paper is that the biological mechanisms by which intracellular mechanics is modulated (e.g. among cell types) remains unexplored and only briefly discussed. Yet this limit is greatly offset by the rigor of the approach.

      We thank the reviewer for this positive assessment and agree that a more extensive discussion of the biological mechanisms underlying the observed mechanical fingerprints strengthens the manuscript. We have substantially expanded the Discussion and Conclusion sections to address potential contributions of cytoskeletal organization, intracellular transport, molecular crowding, and metabolic state. In addition, we now discuss the relationship between the identified mechanical phase space and known cellular phenotypes where appropriate, while explicitly outlining the limitations of the current study and the need for future investigations linking intracellular mechanics to cellular function.

      Reviewer #2 (Public Review):

      The most difficult part of the method is the part with actin polymerization inhibition with cytochalasin B. The data shows that viscoelastic parameters as well as active energy parameters are unaffected by cytochalasin B. It is reasonable to expect that elasticity will reduce and fluidity will increase upon application of such a drug. The stiffness-reducing effect was observed only when CB was used with nocodazole most likely because of phagocytosis of the bead, which is governed by microtubule. The use of other actin-depolymerizing drugs such as latrunculin A would be needed to test actin’s role in mechanical fingerprints. If actin’s role is only explained by accompanying microtubule inhibition, it is not a convenient system to directly test the mechano-adaptation process.

      We thank the reviewer for this important suggestion. To strengthen the interpretation of the actin perturbation experiments, we repeated the Cytochalasin B measurements with an increased number of cells and performed additional experiments using Latrunculin A, a mechanistically distinct and more potent actin-depolymerizing compound. Together with complementary immunostaining experiments, these additional data reveal distinct contributions of the two major cytoskeletal systems to the intracellular mechanical fingerprint. Whereas actin depolymerization primarily affects intracellular stiffness and fluidity, microtubule depolymerization has the strongest effect on intracellular activity while also contributing to cellular softening. These additional experiments provide a substantially clearer interpretation of the respective roles of actin filaments and microtubules in shaping the intracellular mechanical fingerprint.

      Depolymerization of MT with nocodazole did not reduce the solid-like property A. Adding discussion and comparison with other papers in the literature using nocodazole will be helpful in understanding why.

      We thank the reviewer for this suggestion. We have expanded the discussion and now compare our observations with previous AFM studies investigating Nocodazole treatment. While AFM measurements of cortical mechanics often report little change or even increased stiffness after microtubule depolymerization, our intracellular measurements reveal pronounced softening and strongly reduced intracellular activity. We now discuss that this difference likely reflects the distinct intracellular mechanical compartment probed by intracellular microrheology compared with cortical AFM measurements.

      Overall, the usefulness of the concept of mechanical fingerprints and comparisons with other cell mechanics studies (from other groups) will make this manuscript stronger.

      We thank the reviewer for this suggestion. Throughout the revised manuscript we have strengthened the comparison of the mechanical fingerprint with previous literature. In particular, we now discuss the cytoskeletal perturbation experiments in the context of published AFM studies, compare the observed mechanical differences between cell types with previous measurements where available, and expand the discussion of the biological interpretation and limitations of the proposed mechanical fingerprint.

      Reviewer #3 (Public Review):

      The importance of the mechanical fingerprint is diluted due to some missing controls needed for biological relevance.

      We thank the reviewer for raising this important point. To strengthen the biological interpretation of the mechanical fingerprint, we performed substantial additional experiments, including repeated cytoskeletal perturbation measurements with increased sample sizes, additional Latrunculin A experiments, and complementary immunostaining analyses. We also expanded the discussion to address the influence of factors beyond the cytoskeleton, including molecular crowding and metabolic state, and explored possible relationships between the proposed mechanical phase space and cellular phenotypes. While we agree that future studies using well-controlled isogenic model systems will be required to establish direct links between intracellular mechanics and biological function, we believe that the additional experiments and expanded discussion substantially strengthen the biological relevance of the present study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      A caveat of the general methodology, which is partially acknowledged in the MS is that beads are endocytosed and likely end up in specific lysosomal compartments. Therefore, it is not clear whether the mechanical fingerprint fully represent the material properties of bulk cytoplasm, and not something more specific to lysosomal organelles. For instance, lysosome motion may be largely driven by motors moving along MT cytoskeletal track, and the extracted effective energy may as such not fully represent the crowding and effective active temperature of the cytoplasm. This limit certainly affect the interpretation of the results in other cell types, in which membrane trafficking and cytoskeletal organization may vary largely. I believe it would be very important to outline this limitation of the work and discuss it in light of the results obtained throughout.

      We thank the reviewer for pointing out this important limitation, which was not sufficiently addressed in the original manuscript. We have now acknowledged this issue throughout the manuscript and added a limitation section to the conclusion to clarify that our findings specifically relate to internalized objects surrounded by a membrane and therefore primarily reflect the properties of membrane-bound organelles in the 1 µm size regime, rather than the bulk cytoplasm as a whole.

      We consider this focus on membrane-enclosed intracellular objects to be biologically relevant and interesting in its own right. Alternative approaches for introducing tracer particles, such as microinjection or particle guns, are generally more invasive and less reproducible. We therefore deliberately focused on phagocytosed beads as a minimally perturbative and robust experimental system in this study.

      The evolution of the mechanics in Hela Cells using cytoskeletal drugs in interesting, but I was confused by the fact that authors interpret the effect of cytochalasin solely on the cortex. As they are probing intracellular rheology, variations (or lack thereof) may rather reflect bulk F-actin networks? Also the compensation mechanism is interested, but it would need to be strengthened by immunostaining for instance, to support the claim, that microtubule depolymerization enhances F-actin networks.

      We thank the reviewer for this important comment. To elaborate on the effect of cytoskeletal filaments, we extended our analysis by repeating the experiments, increasing the number of samples, and investigating the effect of an additional drug, Latrunculin A. Additionally, we conducted immunostaining with subsequent confocal imaging to deepen our understanding of the effect of the respective drugs. The additional experiments reveal that actin and microtubules contribute differently to the fingerprint. Actin depolymerization primarily affects intracellular stiffness and fluidity, whereas microtubule depolymerization has the strongest effect on both mechanics and intracellular activity. Combined perturbation produces the largest overall effect. Based on these additional data, we no longer invoke the compensation mechanism proposed in the original manuscript. While interactions between the actin and microtubule cytoskeleton have been reported previously, our immunostaining experiments do not provide evidence for a compensatory increase in actin organization following microtubule depolymerization. We have therefore removed this interpretation from the revised manuscript and replaced it with a discussion based on the newly acquired perturbation and imaging data.

      The final figure using principal component analysis is very interesting, but it would be important to link this to phenotypic signatures of the different cells. Could the authors try to link resistance, fluidity and activity to the different functions/behavior of cells? For instance, some of these cells are migratory but some may move much faster than others, and it would be very interesting to correlate the degree of activity or fluidity with speed of migration, or cell shape/size/contractile state for example.

      Indeed, this is an important point. Establishing direct links between the mechanical fingerprint and functional cellular properties such as migration, contractility, proliferation, or morphology would substantially strengthen the biological interpretation of the identified phase space. We carefully considered this suggestion and explored several approaches to relate the measured mechanical parameters to cellular phenotype. However, obtaining directly comparable quantitative functional data across all investigated cell types proved challenging. Parameters such as migration speed, adhesion, and contractility depend strongly on experimental conditions, including substrate properties, assay design, and culture conditions, making literature values difficult to compare across studies. To address the reviewer’s concern, we expanded the discussion and incorporated comparisons to available literature where appropriate. For example, previous studies have reported higher migration rates for HeLa cells compared with MCF7 cells, which is qualitatively consistent with the higher intracellular activity observed in HeLa cells. However, due to the limited comparability and availability of quantitative functional data across the investigated cell types, we refrained from performing a formal correlation analysis. In addition, we grouped the investigated cell lines according to several broad phenotypic classifications, including epithelial/mesenchymal character, cancer status, metastatic potential, and migratory potential, and examined their distribution within the proposed phase space. While this exploratory analysis provides additional biological context, it did not reveal robust relationships that could support definitive conclusions regarding structure–function relationships. We therefore agree with the reviewer that establishing direct links between intracellular mechanical fingerprints and cellular function represents an important next step. To this end, future studies will combine intracellular rheological measurements with independently quantified functional assays, ideally in well-controlled isogenic model systems.

      Reviewer #2 (Recommendations For The Authors):

      The study needs more thorough validation against known technology (such as AFM) or literature, e.g., rheological change upon the same drugs used in the current study.

      We thank the reviewer for this suggestion. We have expanded the discussion of the cytoskeletal perturbation experiments and now compare our observations to previous AFM studies and related literature on cytoskeletal mechanics. Consistent with AFM measurements of cortical mechanics, actin depolymerization using Cytochalasin B or Latrunculin A resulted in a reduction of cellular stiffness. In contrast, microtubule depolymerization produced effects that differ from many AFM studies, which report either no change or an increase in cortical stiffness following Nocodazole treatment. We now explicitly discuss that this discrepancy likely reflects the different mechanical compartments probed by the two techniques. AFM predominantly measures the actin-rich cell cortex, whereas our intracellular microrheology measurements probe the mechanical environment experienced by membrane-bound intracellular particles. We therefore interpret the differing response to microtubule depolymerization as evidence that intracellular active mechanics and cortical mechanics can be influenced by distinct physical mechanisms. These comparisons have been incorporated into the Results and Discussion sections of the revised manuscript.

      Page 8: Citation to Fig. 3a is missing before mentioning Fig. 3b.

      We revised the manuscript to ensure that all references are given in an appropriate order.

      Proper uses of hyphens are recommended to avoid confusion. For example, ’a yet not understood change’ can be written as ’ a yet-not-understood change’.

      We thank the reviewer for this suggestion. We carefully revised the manuscript to improve the use of hyphenation and compound modifiers throughout the text. The specific example highlighted by the reviewer, as well as similar constructions, have been corrected to improve readability and avoid ambiguity.

      Reviewer #3 (Recommendations For The Authors):

      As it reads, sinusoidal waves are applied sequentially from 1- 1024Hz. Please clarify if amplitude is the same for each frequency, also how many frequencies are used? On this point, due to perturbations due to alterations in pre-stress, are the orders of frequencies randomized?

      We thank the reviewer for pointing out this ambiguity. We have revised the manuscript to provide a more detailed description of the active microrheology protocol. Specifically, we now state that all measurements were performed using a constant trapping-laser oscillation amplitude of 200 nm and that the applied frequencies were 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, and 1024 Hz. The frequencies were applied sequentially in increasing order and were not randomized. This information has now been added to the manuscript.

      How many beads are probed in a given cell?

      We thank the reviewer for this question. We have clarified this point in the Methods section and now explicitly state that only a single phagocytosed probe particle was analyzed per cell. Of course, many different cells, and hence beads, have been analyzed per cell type.

      Is the graph in 1 c G’, G” per cell or average of many cells?

      We thank the reviewer for pointing out this ambiguity. In the original version of the manuscript, Figure 1b showed data from a representative cell, whereas Figure 1c displayed an average over multiple cells. To avoid confusion, we revised Figure 1 and now show representative data from a single measurement throughout the analysis workflow (Figure 1c,e,f).

      Figure 1e is quite nice, however, is there an equivalent performed in a nonlinear ECM such as collagen for comparison, in a similar vein can the equivalent be calculated for cells with/without treatment with low doses of cycloheximide to reduce protein synthesis? Yes, cytoskeletal elements are important for cell mechanics, but cytoplasm crowding is often an overlooked factor.

      We thank the reviewer for this important suggestion, and we are glad that the reviewer likes figure 1e. Regarding non-linear ECM, we have not done such experiments using optical tweezers. Collagen is a highly heterogeneous material and using the small deformations that we can obtain using the optical tweezers, our access to the non-linear contributions is rather limited.

      However, we agree that factors beyond the cytoskeleton, including molecular crowding and protein content, can make important contributions to intracellular mechanics. While investigating these effects experimentally, for example through cycloheximide treatment, would be highly interesting, such studies were beyond the scope of the present work.

      The primary focus of this study was to establish and validate a mechanical fingerprint for intracellular active microrheology and to investigate how this fingerprint responds to perturbations of the cytoskeleton. The additional experiments performed during revision therefore concentrated on strengthening the interpretation of the cytoskeletal contributions.

      At the same time, we agree that molecular crowding represents an important alternative mechanism influencing intracellular mechanics. We have therefore expanded the Discussion and Conclusion sections to explicitly acknowledge this limitation and now cite recent studies demonstrating strong effects of molecular crowding on intracellular rheology (Umeda et al,. 2023, Ebata et al., 2023). We further discuss that, besides cytoskeletal organization, metabolic state, intracellular transport, and molecular crowding are likely contributors to the observed mechanical fingerprint.

      The biggest issue is the interpretation of the different factors as each of these cells have different energetic needs. The comparison between cancer cells with different aggressiveness, immune and epithelial cells. For example, some types of cancer cells will be dominated by glycolysis vs oxphos, which will influence both the cytoplasmic and nuclear mechanics? It would be useful to carefully assess factors not restricted to

      (a) Cytoskeleton

      (b) Protein synthesis

      (c) Metabolic state

      For similar lines and/or cells where there are lineages that are either more metastatic in cancer, normal counterpart or drug resistant in an effort to link the fingerprint to a biological output. Specifically, is migration, proliferation, survival correlated with the measurements. The reviewer is sensitive to the technical difficulties of the experiments. However, the interpretation and importance of the mechanical fingerprinting requires additional work as mentioned above.

      We thank the reviewer for this thoughtful comment. We agree that intracellular mechanics is likely influenced by a broad range of biological factors beyond the cytoskeleton, including metabolic state, molecular crowding, intracellular transport processes, and protein synthesis. We also agree that the biological significance of the mechanical fingerprint would be strengthened by establishing direct links to functional cellular outputs such as migration, proliferation, or survival. To address the first point, we have expanded the Discussion and Conclusion sections of the manuscript to explicitly acknowledge that the observed fingerprint is unlikely to be determined solely by cytoskeletal organization. In particular, we now discuss the potential contributions of metabolic state, intracellular transport, and molecular crowding, and cite recent studies demonstrating the importance of these factors for intracellular mechanics. To address the second point, we explored several strategies to relate the measured mechanical fingerprints to cellular phenotype. We expanded the discussion of available literature, including examples where mechanical properties and migratory behavior appear qualitatively consistent. In addition, we grouped the investigated cell lines according to broad biological characteristics, including epithelial/mesenchymal character, cancer status, metastatic potential, and migratory potential, and examined their distribution within the proposed phase space. While this exploratory analysis provides additional biological context, it did not reveal robust relationships that would support definitive conclusions regarding structure–function relationships. We therefore agree that establishing direct links between intracellular mechanics and cellular function represents an important next step. Such studies will require quantitative functional assays performed under controlled and directly comparable conditions, ideally using well-defined isogenic model systems. We now discuss these limitations and future directions explicitly in the revised manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Reviewing Editor, the Senior Editor, and the three reviewers for their careful and constructive assessment of our manuscript. We were encouraged that the reviewers found the question timely and novel, the experimental design thoughtful and well-replicated, and the analyses diverse and informative. The reviewers also raised a number of valuable concerns, which clustered around three themes: (i) the framing of host-mediated selection as microbiome “engineering” versus a proof of concept; (ii) the interpretive challenges introduced by microbial dispersal and the resulting limits on the sterile-inoculated controls; and (iii) requests for clearer methodological detail and additional context from the recent literature. We have revised the manuscript to address these points through clearer framing, expanded discussion, and fuller methodological detail. Consistent with the nature of this long-term experiment, our revisions strengthen the interpretation and presentation of the existing dataset rather than adding new experiments.

      eLife Assessment

      The study has also shortcomings in that the rescuing effect is not benchmarked against healthy well-watered plants, the sterilized controls do not add much information, and the dispersal between inocula confounds the interpretation of the results… the presentation would overall benefit from more extensive consideration of recent developments in the field.

      We appreciate this balanced summary and have revised the manuscript accordingly. We have reframed the abstract and Introduction to present the study explicitly as a proof of concept rather than a completed engineering effort (ll. 27–31; ll. 96–101); we now address the well-watered benchmarking limitation and the limits of the sterile-inoculated controls directly in the Discussion (ll. 543–551); we discuss dispersal and its confounding effect on interpretation head-on, including an alternative hypothesis (ll. 546–551); and we have incorporated the recent studies suggested by Reviewer 3 (ll. 206–209, 442, 476–478). Each change is detailed in the point-by-point responses below.

      Reviewer #1 (Public Review):

      Weaknesses:

      The findings demonstrate the efficacy of host-mediated microbiome selection, but the engineering part for enhancing rice performance under drought-stress conditions has not been provided. The proposed mechanisms rely on correlations but not direct experimental proofs.

      We agree, and we have adopted this framing throughout. Our study demonstrates host-mediated selection as a discovery framework rather than a completed engineering pipeline, and we now say so explicitly: the abstract has been reframed (ll. 27–31) and a statement added at the end of the Introduction (ll. 96–101) clarifying that the work reproducibly enriches beneficial taxa and functions and yields simplified candidate communities, but does not yet benchmark those communities against single-isolate inoculants or test them in the field or against a resident native microbiome. We likewise agree that the functional inferences from our metagenome-assembled genomes (MAGs) are correlative; we now state this explicitly in the Methods and Discussion (ll. 776–778) and note that establishing causal roles for individual taxa or genes will require targeted isolation and genetic manipulation.

      Reviewer #1 (Recommendations For The Authors):

      The experimental design… could benefit from more detailed explanations. For instance, what are the criteria for choosing these soils and how are they relevant to rice growth phenotype? Also, the word ‘generation’ is misleading as it implies the use of seed-to-seed experiments… It would also be good to explain why the authors chose 6 generations for rice fields and 4 generations for deserts and serpentine seep. Importantly, the contribution of the rice seed microbiome… has not been considered and is also missing from… the discussion.

      We have addressed each part of this comment. Soil selection criteria: the Results section “Source inocula bacterial diversity” describes our rationale — we screened nine field soils in a pilot experiment, then selected the three that both supported rice growth and had negligible taxonomic overlap (providing three distinct starting points), with a stated per-soil expectation (rice-adapted, drought-adapted, and high-diversity). We are happy to expand this further if the reviewer feels additional detail is needed. “Generation”: we now define this term as a single 40-day selection cycle rather than a seed-to-seed generation (l. 109). Six vs. four generations: we explain in the Results (l. 309) that, having observed convergence of microbiome composition across soil treatments by the fourth selection generation, we concentrated resources on Rice Field and carried it through two additional cycles. Seed microbiome: we have added a note to the Discussion (ll. 444–446) that, although seeds were surface-sterilized before each generation, a residual contribution of seed-borne endophytes common to all treatments cannot be excluded.

      The authors stated that microbiomes were not selected for propagation into future generations in control lines. In this case, have the authors tested if the control LI microbiome in SG1 through SG6 did or did not significantly change in all the soil types?

      We have clarified the role of the live-inoculated (LI) lines in the text. LI lines were well-watered controls that were re-inoculated each generation with unsterilized selection-line material; they were included to identify drought-enriched taxa (by contrast with the droughted selection lines) and to test whether drought-optimized microbiomes were deleterious under well-watered conditions — not as an independently propagated selection line. Because LI communities were re-derived from selection-line inocula each generation, their composition necessarily tracked the changes occurring in the selection lines; this is the basis of the SL-versus-LI differential-abundance analysis (Figure 7B, Supplemental Figure 9). We note that comprehensive, temporally resolved 16S sequencing was performed for Rice Field, so we are appropriately cautious about extending LI comparisons across every soil type, and we have tempered our conclusions from the control lines accordingly (ll. 543–551).

      In Figure 3B (rice field), the tolerance in terms of AUC NDVI contrastingly increases to the biomass values in SG5 and SG6… the [NDVI] does not seem to be a good measure… It would be interesting to analyze these data sets under normal conditions… include representative pictures of all the ‘generations’… It will also be important to include the LI control data in Figures 3B and 3C.

      We appreciate these suggestions and respond to each. Our metric is biomass-adjusted AUC NDVI, which we use precisely to separate drought performance from plant size; NDVI itself was validated against shoot water content in preliminary experiments (R = 0.98; Supplemental Figure 2E), so we are confident it is an appropriate, validated proxy for drought status. We have substantially expanded the Methods to explain this adjustment and why the metric can diverge from raw biomass (ll. 666–669). Regarding the specific additions requested: analyzing the well-watered plants as a phenotypic dataset, adding representative images for every generation, and plotting LI data in Figure 3B/C would each require new analyses or figures that are outside the scope of this revision; moreover, LI plants were never droughted and therefore have no drought-response score comparable to the SL and SI lines, so they cannot be placed on the same axes. Representative images contrasting the first and last selection generations are already provided in Figure 3A. We have, however, added an explicit acknowledgement that our design does not quantify the absolute magnitude of drought rescue relative to well-watered performance (ll. 551–552).

      It is less clear how sterile soils acquired environmental taxa over time. Was this a seepage of microbes from inoculated samples to the calcinated clay, possibly via the water irrigation system? In this regard, four Venn diagrams representing all the generations… would be relevant.

      Each plant was grown in an individual container with its own separate water reservoir (Supplemental Figure 4), so shared irrigation was not a route of transfer; the most likely routes are airborne movement and handling within the growth chamber, together with within-treatment shuffling of plants. Our dispersal analysis (Figure 5) already traces the origins of taxa in each treatment, and Supplemental Figure 6 quantifies the ASVs shared among treatments over generations; we have added explicit criteria for these origin assignments (ll. 248–253). We therefore prefer to retain the existing Figure 5 / Supplemental Figure 6 presentation rather than add four separate Venn diagrams, which would convey the same information less quantitatively, but we are glad to reconsider if the editor feels a Venn representation would help readers.

      What is the logic behind the so-called ‘immigrating taxa’ in this study?

      “Immigrating” (dispersed) taxa are those that appear in a treatment despite not being attributable to that treatment’s own starting material — i.e., ASVs not detected in that treatment’s field soil or enrichment-generation inoculum, which must therefore have arrived by dispersal from other treatments or from the growth-chamber environment. We have made this definition explicit in the text (ll. 248–253).

      The decrease in alpha-diversity in subsequent generations… should be thoroughly discussed. Have authors tried to culture these few remaining taxa? If yes… tested for their individual drought tolerance supported by physiological assays… If no, is the microbiome of SG6 (and associated functions) ideal or sufficient to create drought tolerance in field conditions?

      We have expanded the discussion of the diversity decline. In addition to niche filtering along the soil-to-root gradient and dilution-to-extinction (already discussed), we now note that DNA-based profiling cannot distinguish metabolically active cells from relic DNA or dormant/non-viable cells, so part of the apparent collapse in diversity may reflect enrichment for the taxa that were active in the original inoculum (ll. 206–209). We agree that culturing the remaining taxa and characterizing them with physiological assays (e.g., water potential, water-use efficiency, stomatal conductance) is a valuable next step; these experiments are outside the scope of the present study, which we have now framed explicitly as a proof of concept, and we identify field validation of selected communities as a key open question (ll. 96–101).

      The result that Ideonella was identified as the dominant taxa in all selection conditions is highly interesting… This… should have been followed up for isolating the strains and performing direct tests to test their importance for conveying drought stress.

      We agree that isolating and directly testing dominant taxa such as Ideonella is the logical next step, and we now emphasize that a central value of host-mediated selection is that it yields simplified communities from which such taxa can be more readily isolated (ll. 27–31). These isolation and functional-validation experiments are beyond the scope of the current study and we have framed them as future directions rather than undertaking them here.

      The MAGs shown in Figure 8 have apparently ‘been assigned to ASVs…’. These data are not shown anywhere… the MAG data only give correlations but not direct genetic proofs of the biological functions of the identified genes.

      We have expanded the Methods to describe how each MAG was matched to an ASV (by closest taxonomic assignment and by concordance of relative abundance across samples), and we now state explicitly that these assignments are approximate and that the functional inferences drawn from them are correlative rather than definitive (ll. 776–778). We would be glad to add a supplemental table listing the MAG-to-ASV assignments if the reviewer or editor would find it useful; because it reports assignments already in hand, it requires no new analysis.

      Reviewer #2 (Public Review):

      Strengths:

      I think this study examines an important and exciting topic in the area of plant microbiomes. I predict the findings of the experiments will inform a wide audience of researchers attempting similar studies and be helpful in their designs.

      We thank the reviewer for recognizing the novelty of this complex experiment as well as the effort we put into designing it. Like the reviewer, we hope that this manuscript can serve a wide audience and help inform subsequent experiments in this new topic area.

      Weaknesses:

      Although the controls were well designed, the dispersal of the microbiomes erased the utility of the sterile inoculated (SI) controls… the SI lines acquired microbes from the experiment and never appeared to significantly deviate from the SL plants. The dispersal of the microbes… also minimizes any conclusions that can be made about the different starting inocula and how prone to selection they may be.

      We agree that microbial dispersal confounded our ability to use the sterile-inoculated (SI) plants to account for batch variation between generations. By maintaining each plant as a spatially discrete unit (individual pots and watering reservoirs), we had originally intended SI plants simply to acquire a similar consortium of environmental microbiota each generation. Truly axenic SI plants would have been better suited to this purpose, but would have severely limited the number of replicates and replicate selection lines we could include. We have now addressed this limitation directly in the Discussion (ll. 543–551): we state that the SI lines cannot be treated as static, microbe-free baselines, that the batch-to-batch variation they were meant to capture is only partially controlled, and that dispersal limits the strength of the conclusions we can draw about differences between starting inocula. This shared trajectory of selection and intended control lines has been observed in other host-mediated selection studies but rarely discussed in detail, and we now foreground it as a lesson for experimental design.

      Reviewer #2 (Recommendations For The Authors):

      My first concern is the framing of the approach… the authors never show that this approach has better efficacy than single-isolate inoculates… The phase of the research is still proof of concept, understandably, but these caveats should be mentioned/addressed head-on in the Introduction and Discussion.

      We agree and have made these caveats explicit rather than implicit. The abstract now frames the work as identifying candidate taxa and communities rather than delivering a finished engineering solution (ll. 27–31); the Introduction now states plainly that this is a proof of concept that does not benchmark the passaged communities against single-isolate inoculants or evaluate them in the field or against a resident native microbiome (ll. 96–101); and the Discussion reiterates these limitations (ll. 543–551).

      I disagree with the authors that the selected microbiota better approximate field conditions (line 55) - because… the diversity of the microbiome is drastically reduced… it is likely that exclusion of taxa is just as important as the passaging of bacterial members to see the desired effect.

      We take this point and have revised the sentence at (former) line 55 accordingly (now l. 57): we now say that community-level screening more closely approximates field complexity than single-isolate screens only at the outset, and we no longer imply that the selected (diversity-reduced) communities better approximate the field. We agree that taxon exclusion may be as important as enrichment; this is consistent with our balance analyses, in which the denominator groups comprise taxa negatively associated with phenotype (Figure 7), and with the diminishing returns we observe as diversity collapses. We have also added an explicit sentence to the Discussion (l. 488) stating that the exclusion of detrimental taxa may be as important as the enrichment of beneficial ones, and that a microbiome’s finite membership may contribute to the diminishing returns of selection we observe over generations.

      (1) It is unclear what the reason (or methodology) for correcting NDVI by biomass. Much of the findings hinge on corrected NDVI values, so a more thorough explanation of the correction method… would benefit the reader.

      We have substantially expanded this explanation in the Methods (ll. 666–669). We now state that biomass and AUC NDVI were anti-correlated (Supplemental Figure 12) and that we adjusted for plant size by taking the residuals of a linear regression of AUC NDVI on shoot dry-weight biomass, using these biomass-adjusted values as our measure of drought performance so that selection would reflect drought tolerance rather than plant size alone.

      (2) Are data for panels B and C of Figure 3 scaled?… how can one have a negative area under the curve if all the NDVI values are positive? For panel B, the representative plant images are much larger than 0.8 grams.

      This is a helpful catch, and the confusion stems from our terse original description. The values plotted are the biomass-adjusted AUC NDVI (regression residuals), which are centered on zero by construction; negative values therefore indicate poorer-than-expected drought performance for a plant of a given size and do not reflect negative raw NDVI or a negative raw area under the curve. We now explain this explicitly (ll. 666–669). In panel C, shoot biomass is plotted as dry weight in grams; the representative plant images in panel A are qualitative illustrations and are not scaled to the biomass axis. We will make the axis labels and legend state the units and the residual nature of the adjusted metric explicitly (noted in our accompanying figure-revision guide).

      (3) The dispersal analysis… What are the criteria for classifying ASVs as specific to an input source? Was it that they were observed in all samples of field soil, i.e. was a prevalence threshold implemented? Could they be observed in any other soil at a smaller threshold?

      We have added the criteria explicitly (ll. 248–253). An ASV was attributed to a given soil treatment if it was detected (present/absent) in that treatment’s field-soil or enrichment-generation inoculum samples; ASVs detected in none of the field soils or source inocula were designated environmental in origin (“unk/env”), and ASVs meeting the criterion for more than one treatment were assigned to each. Assignments were thus based on detection in the source samples rather than on an abundance-prevalence threshold within later generations.

      This reviewer finds the results around [inoculum source] inconclusive… the serpentine seep microbiome appears to provide more benefit from the first round of selection than any other soil… The slope of improvement… is different between soils, but mainly because the serpentine microbiomes start out conveying greater benefits than the other soils.

      We agree the Serpentine Seep result is not clear-cut. The Discussion already presents inoculum provenance as one of several factors shaping the outcome rather than a decisive one, and we have now added an explicit acknowledgement that Serpentine Seep conferred comparatively large benefits in the earliest cycles before plateauing, so its weaker response to continued selection may reflect an early approach to a performance ceiling rather than an inherently poorer substrate for selection (l. 423). We have tempered our “source matters” language accordingly.

      Have the authors assessed the biomass and ndvi of the well-watered plants?… showing this data would allow the reader to assess the degree to which the microbiomes are rescuing the plant… and… whether tradeoffs exist… under fully watered conditions.

      We have added an explicit statement that our design does not pair each droughted line with a well-watered readout of the same phenotype, so we refrain from estimating the absolute magnitude of drought rescue (ll. 551–552). We note, however, that shoot biomass increased in parallel with drought performance across selection generations (Figure 3C), which provides no evidence that selection for drought tolerance came at a cost to growth under our conditions. Collecting matched well-watered phenotypes to quantify effect size and trade-offs is a worthwhile aim for future work but would constitute a new analysis beyond this revision.

      How can the authors exclude the possibility that environmental microbes pre-existing in the growth chamber taxonomically overlap with the field soil-specific microbes?… the alternative hypothesis should be mentioned… A clearer representation of the ASVs categorized as source soil-specific in Figure 5… would be useful and how many of these ASVs make up the bar plots.

      We now state this alternative hypothesis explicitly: because dispersed taxa came to dominate all treatments, we cannot fully exclude that taxa shared across treatments were recruited from a common growth-chamber pool rather than dispersing directly between soils (ll. 546–549). We note that the two processes are difficult to distinguish retrospectively, but that the bias of each SI line toward its own treatment’s native diversity (Figure 5) is more consistent with genuine cross-treatment dispersal. Regarding the figure, the number of ASVs underlying each origin category is available in Supplemental Figure 6; we describe in the accompanying figure-revision guide how the Figure 5 legend can be clarified to state the assignment criteria and the ASV counts.

      The sterile inoculated plants were a nice control in theory, but I question their utility… A contrast that should be made is the microbiomes of only SI plants. It is striking that sterilized controls assemble and retain more microbes from the unsterilized starting inoculum. I would expect everything to be acquired from dispersal.

      We agree, and we have foregrounded this in the Discussion (ll. 543–551). As the reviewer notes, SI communities were biased toward their own treatment’s native diversity rather than being assembled entirely from dispersal (Figure 5) — an informative observation, but one that also demonstrates why the SI lines cannot serve as the clean, microbe-free baseline we had intended. We now treat this as a key design lesson and note that a fully isolated (e.g., gnotobiotic) control would be required to separate these effects in future experiments (l. 560).

      Reviewer #3 (Public Review):

      Weaknesses:

      Sterile/non-inoculated calcined clay also tends to enrich similar microbes… In a future experiment, the work would benefit from including a truly sterile control… the reader may get to wonder whether these efforts are necessary at all… This is discussed across the paper but not directly addressed and I think the manuscript would benefit from a clear argument for or against this idea.

      We thank the reviewer for this insightful point and have made our argument explicit rather than leaving it implicit. First, we agree a fully isolated, truly sterile control would strengthen future iterations of this design; the manuscript notes that gnotobiotic plants would be the ideal (if costly) means of achieving this (l. 560). Second, on whether selection is necessary if plants recruit beneficial microbes from the environment: the phenotypic gains seen even in the sterile-inoculated lines do not indicate that selection was superfluous, but rather that those plants recruited from a metacommunity that was itself being optimized by selection in the neighboring selection lines each generation. In other words, environmental acquisition propagated the benefits of selection across the shared growth-chamber environment rather than replacing it. We have clarified this reasoning in the Discussion (ll. 543–551).

      Reviewer #3 (Recommendations For The Authors):

      It is mentioned multiple times… that host genotype is the driver of the microbiota selection… However, this is not the case [multiple lines] and therefore I don’t find that surprising that there is a convergence of the microbiota across soils and selection rounds.

      We agree and have added text making this explicit: all plants were a single, near-isogenic rice genotype, and because host genotype is itself a strong filter on microbiome composition, the use of one genotype — together with shared environmental conditions and selection criteria — makes convergence across lines an expected rather than a surprising outcome (ll. 438–441). We have softened language that could be read as attributing selection to host-genotype variation.

      Another possibility… is that those microbes that are found in the later generations are actually the ones that were active/alive in the initial inoculum. It is not possible to rule out that most of the sequenced microbes in the input were not actually dead. Similar observations were made… in Duran et al. 2022. New Phytol.

      We have added this possibility to the manuscript, noting that DNA-based profiling cannot distinguish metabolically active cells from relic DNA or dormant/non-viable cells, so part of the apparent diversity decline may reflect enrichment for the subset of taxa that were active in the original inoculum, with reference to the transplantation work the reviewer cites (Durán et al. 2022; ll. 206–209).

      In the shotgun data, was there any observation of other microbes present (fungi, virus)? Did they follow the same trends as the bacterial communities?… I think addressing this will be very interesting and very novel.

      We agree this is an interesting question. Our shotgun workflow was designed and assembled specifically to recover high-quality bacterial and archaeal MAGs, and a rigorous cross-kingdom analysis (fungi, viruses) would require dedicated, eukaryote- and virus-specific assembly, binning, and reference databases — a substantial new analysis that lies outside the scope of this revision. We therefore flag cross-kingdom community dynamics as a promising direction for future work rather than presenting a new analysis here.

      Any interesting overlap with the results found in Karasov et al. 2022 (biorxiv)?

      We have added a comparison to drought-driven selection on host-associated microbiomes in Arabidopsis (Karasov et al. 2022) at the relevant point in the Discussion (l. 442).

      In Liu et al., 2024 Nat. Comms, the authors found Devosia as an interesting candidate for disease suppression (to add to the discussion?).

      Added — we now note that Devosia, one of the lesser-known genera enriched in our experiment, has recently been highlighted as a candidate mediator of disease suppression in the rhizosphere (Liu et al. 2024; l. 476).

      Lipids as a signal for host-microbe interaction: Rich et al., 2021 Science.

      Added — in the functional-enrichment discussion we now cite lipids as increasingly recognized central signaling molecules in host–microbe symbioses (Rich et al. 2021; l. 478).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Dwulet et al. combined experimental and modeling approaches to investigate how correlated spontaneous activity in the mouse's primary visual (V1) and primary somatosensory (S1) areas drives the development of multisensory integration in area RL. Notably, they focused on early developmental stages, before sensory experience occurs. Consistent with previous experimental findings, the authors first demonstrated that spontaneous activity becomes more sparse across development in all three areas, as measured by event amplitude, event duration, and participation ratio. Using a linear mixed model analysis to compare the maturation of this spontaneous activity, they found evidence that S1 matured the fastest. The authors then presented experimental evidence suggesting that these spontaneous events were moderately correlated both spatially and temporally.

      They hypothesized that activity-dependent mechanisms use these correlations to establish connectivity across these regions. To test this hypothesis, the authors modeled a feedforward network with connections from S1 to RL and from V1 to RL, where the strength of connections depended on a Hebbian term for potentiation and a heterosynaptic term for depression. By investigating different levels of V1-S1 correlations, they found that moderate levels of correlation led to the significant development of topographically organized connectivity while maintaining a mix of bimodal and unimodal cells in RL. Additionally, when simulating a network with a more mature S1, they observed that topographical maps improved not only between S1 and RL but also between V1 and RL. Finally, the authors use linear regression to suggest that the mixture of bimodal and unimodal cells in RL is optimal for encoding the maximum amount of information from both V1 and S1.

      However, there are significant gaps between the experimental data and the modeling setup, which weaken the paper's conclusions. Additionally, some key details are omitted, making it difficult to fully assess their analysis and interpret some of their figures.

      (1) Some of the statistical measures and techniques in Figure 1 could benefit from clearer definitions. While the thresholds for activation (peak with at least 5% dF/F0) and events (20% of recorded cells activated simultaneously) are provided, event duration and participation rate are not clearly defined. Based on this definition of event alone, it is unclear why the minimum participation rate in Figure 1F is not 20%. Additionally, the conclusion that S1 matures earlier than RL and V1 could be strengthened by including a direct comparison between S1 and RL, as the current analysis only compares these areas to V1.

      We thank the reviewer for this comment. We have now updated the Methods to include the event duration as time above half max, participation rate as % of cells out of total in that region active during an event. Also, the threshold of 20% recorded cells to identify an event was incorrectly stated, in fact the threshold was 5% consistent with what the reviewer observed in Figure 1F. This error has been corrected throughout the Methods. We chose 5% because spontaneous activity significantly sparsifies over development, with events involving far fewer cells, as previously shown by multiple studies (Golshani et al., 2009; Rochefort et al., 2009; Gribizis et al., 2019; Leighton et al., 2021; Murakami et al., 2022; Chini et al., 2022; reviewed in Lakhera et al., 2024).

      For the linear mixed model (LMM) analysis, we used V1 as a reference just for convenience, but this has no influence on the results. We now added a direct comparison using each area as reference in the LMMs. Several Supplementary Tables (S1-3) now show these results with coefficient estimates and stars showing statistical significance and are mentioned in the legend of Figure 1 and the main text.

      (2) The wide-field experiments in Figure 2 could be expanded to support the feedforward modeling assumptions. Currently, the spatial and temporal correlations presented leave open the possibility that these spontaneous events are traveling waves propagating from V1 to RL to S1 (or vice versa). This scenario would suggest a different connectivity scheme for the model. Clarifying this point with additional data analysis, specifically including temporal correlations involving RL, could provide stronger support for the model's assumptions.

      We agree with the reviewer that the correlation analyses shown in Figure 2 do not differentiate between two possibilities: activity that travels smoothly from one cortical area to another, thereby correlating correlations between these areas, versus activity that is spatially confined to individual areas but occurs near-synchronously across those areas. To address this point, we have revised Figure 2 in two ways.

      First, we added examples of spontaneous activity showing near-synchronous but spatially distinct activation of sub-areas in V1, RL and S1 (new Figure 2D). These examples show that localized activity can remain confined to individual sensory cortical areas and RL, while occurring at similar times across areas. Thus, the observed correlations are not simply due to single large events spreading continuously across the entire imaged field.

      Second, we added a lagged cross-correlation analysis between V1 and S1 activity (new Figure 2G). This analysis shows that the correlation between V1 and S1 peaks close to zero lag and decays for both positive and negative lags. This argues against a stereotyped travelling-wave-like propagation from V1 to S1 or from S1 to V1 with a fixed delay. The cross-correlation curves show a mild asymmetry, with somewhat higher correlations when S1 precedes V1. However, because the dominant peak is centered near zero lag, we interpret the data primarily as evidence for near-synchronous, spatially structured coactivity across sensory areas, rather than fixed directional propagation.

      Together, these two analyses support the modeling abstraction that V1 and S1 provide temporally correlated, spatially structured inputs to RL. We have added the new activity examples and the lagged cross-correlation analysis to Figure 2 and revised the Results accordingly. Although these analyses do not exclude all forms of propagating activity, they argue against the specific concern that the correlations are dominated by stereotyped traveling waves passing sequentially through V1, RL, and S1.

      (3) The functional correlation map in Figure 2D appears contradictory to the authors' modeling assumption that inputs are correlated spatially in V1 and S1. While V1 seed points align topographically with RL, this organization breaks down when extended into S1. In contrast, and in support of the modeling assumption, Figure 2E shows clearer topography across all three regions. A discussion of this discrepancy would be helpful, as it's a key conclusion of the figure. Additionally, it is unclear when this data was collected during development. Clarifying the developmental stage and analyzing how this map changes over time could strengthen the results.

      We thank the reviewer for pointing out this ambiguity. In the original version, the functional correlation maps were generated using separate seed locations in V1 and S1, and the interpretation relied heavily on thresholded RGB maps in which each pixel was assigned to the color channel with the strongest correlation. This representation made it difficult to directly compare the V1- and S1-seeded maps and may have given the impression that topographic organization was preserved in one direction but not the other.

      We have therefore revised the analysis and presentation of Figure 2. Instead of using separate seeds in V1 and S1, we now use common seed locations in RL and compute the correlations of these RL seeds with activity across the imaged cortical field. This allows us to ask directly whether different RL locations are associated with spatially distinct regions in both V1 and S1. We now show both the raw correlation maps, in which the RGB channels reflect the correlation values for the three RL seeds (new Figure 2E), and the thresholded/maximum-channel representation, in which each pixel is assigned to the strongest of the three color channels (new Figure 2F). The raw correlation maps make the correlation structure visible without relying solely on thresholding, whereas the thresholded representation highlights the spatial ordering of the strongest correlations.

      With this revised analysis, the topographic relationship across V1, RL, and S1 is clearer and no longer depends on comparing separate V1- and S1-seeded maps. We also clarified in the figure legend how the RGB maps are computed and how thresholded pixels are represented.

      The reviewer also asked about the developmental stage and progression of this phenomenon. The example shown in Figure 2 was recorded at PN9, and we now state this explicitly. In addition, we added examples from PN9–PN13 in Supplementary Figure S1, showing that similar functional correlation-map structure is present across the developmental period analyzed here. This is consistent with previous work showing that retinotopy-like patterns in higher visual areas can be recovered from functional-connectivity analysis of spontaneous activity before eye opening (Murakami et al., 2022), and with recent work showing that retinotopy-like and somatotopy-like patterns of ongoing activity, together with their rough topographic correspondence in RL, are already present before eye opening at PN10–11 (Matsumoto, Murakami & Ohki, 2025).

      (4) The modeling of spontaneous events with fixed amplitude and duration seems inconsistent with the experimental data in Figure 1, which shows variability in these parameters. This is particularly confusing in Figure 4, where S1 maturation is modeled as a stronger topographical alignment with RL, but the experimental data defines maturation based on amplitude, duration, and event rates. Justifying these modeling choices or adapting the model to reflect experimental variability would create a better connection between the theory and data.

      We agree with the reviewer that the original presentation did not sufficiently distinguish between the experimentally measured maturation of spontaneous activity and the way S1 maturation was implemented in the model. In the experiments (Figure 1), earlier maturation of S1 was reflected by lower event amplitudes, shorter durations, and higher event rates. In contrast, the original model explored the effect of a stronger or more spatially refined S1-to-RL projection (Figure 4). This modeling choice was motivated by pilot anatomical data suggesting that projections from S1 to RL become more elaborate earlier than projections from V1 to RL at comparable developmental ages. We include examples of these pilot data (Author response image 1), but we have not included them in the manuscript because the dataset is preliminary and does not yet allow for a sufficiently complete quantitative analysis.

      Author response image 1.

      Projections from V1 and S1 to RL at different developmental ages. Pilot anatomical data suggest that the S1 projection to RL becomes more elaborate and mature earlier than the V1 projection.

      To address the reviewer’s concern more directly, we have now extended the model to incorporate differences in the spontaneous activity patterns of V1 and S1, including the lower amplitude and higher frequency of S1 events. We then examined how these activity differences interact with different levels of initial connectivity bias between the primary sensory cortices and RL (Supplementary Figure S2). We also quantified the resulting topography, map alignment, and fraction of bimodal RL neurons as a function of the S1 bias and included these additional plots in Figure 4 (panels C-E).

      This analysis shows that incorporating the more mature S1-like activity patterns alone was not sufficient to generate the appropriate topographic and aligned maps. Rather, the model still required an initial connectivity bias, together with an appropriate level and structure of correlated activity. This is consistent with the results shown in Figure 3B,E,G–I and discussed in our response to Reviewer 2, point 3, where we show that the initial bias does not by itself determine the final map structure, but instead interacts with the level of V1–S1 correlation. We have added the new analysis to Supplementary Figure S2 and revised the text to clarify the interpretation. Rather than presenting the stronger S1 bias as a direct consequence of the more mature S1 activity dynamics revealed through the differences in amplitude, duration, and event rate, we now frame it as a model prediction: earlier S1 maturation may need to be accompanied by, or act through, a more advanced anatomical or functional S1-to-RL projection, whose refinement still depends on the temporal and spatial structure of spontaneous activity.

      The results suggest that differences in spontaneous activity dynamics and differences in projection maturity may act together during the emergence of topographically aligned multisensory maps, with neither component alone being sufficient to determine the final organization. Future experiments will be needed to establish whether such an S1-to-RL connectivity bias is present systematically, to quantify its developmental progression, and to disentangle the relative contributions of more mature spontaneous activity dynamics and more mature connectivity.

      (5) Several important details of the mathematical model are missing or unclear, partly due to typos. The Results section mentions the general framework of the input correlation matrix (e.g., "S1 and V1 neurons were driven by a combination of events, independent and shared in each V1 and S1" and "each independent event activated a randomly chosen, contiguous set of neurons"), but the specifics are not fully explained. Additionally, the caption of Figure 5 refers to a non-linear transfer function (a sigmoid), but these details are not provided in the Methods section, which instead suggests a linear model was used. A careful review of the main text and Methods section would help ensure that all the necessary details are included and that the story is both complete and accurate.

      We thank the reviewer for pointing out these missing details and inconsistencies. We have carefully revised the Results, figure captions, and Methods to make the model description more complete and internally consistent.

      First, we clarified how spontaneous input events were generated. Specifically, V1 and S1 activity was constructed from independent events in each area and shared events across the two areas. These event streams were generated using Poisson processes, with the rates chosen such that the total event rate was matched across simulations while varying the fraction of shared versus independent events. We also clarified that each event activated a spatially contiguous group of neurons, thereby implementing local spatial correlations within each primary sensory area, while shared events activated corresponding topographic locations in V1 and S1.

      Second, in the Methods we clarified the use of the nonlinear transfer function in the decoding analysis shown in Figure 5. The simulated RL activity was transformed with a sigmoid nonlinearity before performing the regression analysis, and we have now added the corresponding equation (15) to the Methods.

      Third, we clarified the distinction between the numerical decoding analysis and the analytical calculation of the optimal weight matrix. The decoding analysis uses the nonlinear transformation described above, whereas the analytical calculation uses a linearized version of the model to obtain a tractable closed-form solution. We now state this explicitly in the Methods to avoid the impression that two inconsistent models were used.

      Finally, we corrected several typographical errors and checked that the Results, Methods, and figure captions use consistent terminology for the input generation, correlation structure, and decoding analysis.

      (6) While Figure 5 supports the paper's conclusion that a mixture of unimodal and bimodal neurons in RL optimizes information encoding, the authors missed an opportunity to strengthen the connection between the model and experimental data. Specifically, they could apply this reconstruction method to the experimental data and examine how RL's ability to reconstruct V1/S1 activity changes across development. Their model predicts that this performance would improve over time, and if this trend is observed in the experimental data, it would provide strong validation that these feedforward connections are developing in line with the model's predictions.

      We agree with the reviewer that applying the reconstruction analysis directly to the experimental data would provide an important additional test of the model. However, the current experimental datasets are not well suited for this analysis. The two-photon recordings used to characterize spontaneous activity in V1, S1, and RL were acquired sequentially rather than simultaneously, and therefore cannot be used to reconstruct V1/S1 activity from RL activity. In principle, a related analysis could be attempted using the wide-field recordings, which are simultaneous across cortical areas. However, these data have lower spatial resolution, include movement-related variability, and do not provide cellular-resolution measurements of RL activity. We explored this possibility, but the resulting reconstructions were not sufficiently reliable or interpretable to include in the manuscript.

      We now state this explicitly as a limitation in the Discussion and identify simultaneous multiarea recordings at cellular resolution as an important future test of the model. Such experiments would make it possible to determine whether the ability of RL activity to reconstruct V1/S1 activity improves across development, as predicted by the model.

      Reviewer #2 (Public review):

      The authors aim to investigate the role of spontaneous activity in shaping the development of multisensory integration in the brain, specifically focusing on the connections between primary visual and somatosensory sensory areas (V1 and S1) and a higher-order cortical area rostrolateral to V1 (RL). They seek to understand how spontaneous activity guides the formation of aligned topographic maps and the emergence of bimodal neurons in RL.

      First, the authors found that spontaneous activity in all three areas sparsifies over time, but S1 exhibits more mature patterns earlier than V1 and RL. They claimed that correlated activity among neighboring regions of these areas during development carries topographic information. These data were used to implement a computational model that employed Hebbian rules of synaptic plasticity. The model indicated that correlated spontaneous activity can generate topographic connectivity between S1/V1 and RL and bimodal neurons in RL. The model suggested that the more mature spontaneous activity in S1 can guide map alignment between V1 and RL. In addition, the model also suggested that a mixture of bimodal and unimodal neurons in RL is optimal for decoding information from V1 and S1.

      While the data presented in the manuscript is promising and provides preliminary insights into the role of spontaneous activity in multisensory integration, it would be beneficial to strengthen the experimental foundation regarding the correlation between V1, S1, and RL. Incorporating more rigorous spatio-temporal analyses of spontaneous activity could enhance the robustness of these findings.

      Here are some important concerns:

      (1) The analysis of how spatial topography influences activity correlations in Figure 2 has several issues.

      (1a) While squares in V1 and S1 covered a small area of these sensory areas, the correlated territories in RL covered the entire area of RL. The topographic map in V1 continues caudally, so where is the rest of the map in RL? Something similar applies to the relationship between S1 and RL.

      We thank the reviewer for pointing out this ambiguity. In the original version, the functional correlation maps were generated using separate seed locations in V1 and S1, and the interpretation relied heavily on thresholded RGB maps in which each pixel was assigned to the color channel with the strongest correlation. This made it difficult to directly compare the V1- and S1-seeded maps and could give the impression that the correlation structure extended differently across RL depending on the chosen seed area.

      We have therefore revised the analysis and presentation of Figure 2. Instead of using separate seeds in V1 and S1, we now use common seed locations in RL and compute the correlation of each RL seed with activity across the imaged cortical field. This allows us to ask more directly whether different locations in RL are associated with spatially distinct regions in both V1 and S1. We now show both the raw correlation maps, in which the RGB channels reflect the correlation values for the three RL seeds (new Figure 2E), and the thresholded/maximum channel representation, in which each pixel is assigned to the strongest of the three color channels (new Figure 2F). The raw correlation maps make the correlation structure visible without relying solely on thresholding, whereas the maximum-channel representation highlights the spatial ordering of the strongest correlations.

      With this revised analysis, the topographic relationship across V1, RL, and S1 is clearer and no longer depends on comparing separate V1- and S1-seeded maps. We also clarified in the figure legend and Methods how the RGB maps are computed, how the maximum-channel maps are generated, and how thresholded pixels are represented. In addition, we added Supplementary Figure S1 to show further functional-correlation-map examples across PN9, PN10, and PN13 recordings, with seed locations in V1, S1, or RL as indicated in each panel.

      (1b) It is essential to know how areas were drawn. High precision is required.

      Consistent delineation of cortical areas is absolutely essential for interpreting the functional correlation maps. We have therefore expanded the Methods to describe how cortical areas were delineated from the wide-field recordings. Briefly, recordings were acquired in a field of view defined relative to lambda and the midline, and cortical-area outlines were assigned using published reference maps together with the spatial organization of spontaneous activity patterns and functional correlation maps. This approach follows the procedure we previously validated for developmental wide-field recordings (Leighton et al., 2021).

      To make this transparent, we added Supplementary Figure S3, which illustrates how the reference-map-based outlines were overlaid on the imaging field of view and how functional correlation maps and individual network events helped identify the boundaries of V1 and neighboring areas. We also clarified this in the Methods.

      (1c) It is not clear if correlated activity means different events in sync or large events that cover 2 or all 3 cortical areas of interest. The figure points to the second option, which contradicts the size of events at these stages, mainly in the oldest mice analyzed here.

      The reviewer asks whether the correlations reflect spatially confined events occurring near-synchronously in different cortical areas, or instead large events spanning V1, RL, and S1. To clarify this point, we revised Figure 2 to show representative activity traces and individual frames from the wide-field recordings (new Figure 2B–D). These examples show that activity can be localized to distinct subregions within V1, RL, and S1 while occurring at similar times across areas. Thus, the observed correlations are not well explained by single large events spreading continuously across the entire imaged field.

      We have revised the Results and Figure 2 to make this clearer. In addition, the lagged cross-correlation analysis in Figure 2G shows that V1–S1 correlations peak near zero lag and decay for both positive and negative lags, arguing against a stereotyped travelling-wave-like propagation between the two primary sensory cortices as the dominant explanation for the observed correlations.

      (1d) It is fundamental to know in detail and provide examples of how the detection of events was performed. For instance, could the dispersion of light from an event in V1 close to RL cause the detection of activity in RL?

      The reviewer asks how events were detected in the wide-field recordings and whether light dispersion could lead to false-positive correlations between neighboring areas. We have clarified this point in the Methods. For the functional correlation analyses shown in Figure 2, we did not perform event detection. Instead, the correlation maps were computed from the continuous fluorescence time courses by calculating Pearson correlations between seed region activity and the activity of every pixel in the field of view. Thus, the functional correlation maps do not depend on detecting or assigning individual events.

      To address the concern about whether correlations could reflect light spread from large events rather than genuine co-activity across areas, we revised Figure 2 to include representative activity traces and individual frames from the wide-field recordings. These examples show that activity can be spatially confined to distinct subregions in V1, RL, and S1 while occurring at similar times across areas. This argues against the interpretation that the correlations are simply caused by a single event spreading continuously across the imaged field or by light dispersion from one area into another. We have also described the area delineation procedure in more detail in the Methods and added Supplementary Figure S3 to illustrate how activity patterns and functional correlation maps were used to assign outlines of distinct cortical areas.

      Although wide-field imaging cannot completely exclude minor contributions from light scattering near area borders, the spatially localized activation patterns and the topographically ordered correlation maps support the interpretation that the correlations reflect genuine nearsynchronous co-activity across V1, RL, and S1.

      (2) For the correlations among V1, S1, and RL, it is crucial to have a consistent method to delineate the borders of cortical areas. The authors mention in one sentence that areas were drawn according to a reference map. More details are needed to convince the reader that the borders are accurate, especially because their shape and position change with age.

      As described in our response to point 1b, we have expanded the Methods to clarify how cortical-area borders were delineated in the wide-field recordings. Briefly, recordings were acquired in a field of view defined relative to lambda and the midline, and cortical area outlines were assigned using published reference maps together with the spatial organization of spontaneous activity patterns and functional correlation maps. We also added Supplementary Figure S3, which illustrates how the outlines based on reference maps were overlaid on the imaging field of view and how functional correlation maps and individual network events helped identify the boundaries of V1 and neighboring areas. This makes the delineation procedure more transparent across animals and developmental ages.

      (3) The results from the model seem to be based on the initial bias in connectivity between neighboring cells from the different areas. Then, it seems straightforward that implementing correlated activity with Hebbian and synaptic depression rules will force the strengthening of connections between spatially close cells. Despite this apparent predisposition of the model towards a defined outcome, the flaws in the experimental data used prevent a rigorous interpretation of the computational model.

      We understand the reviewer’s concern that the initial topographic bias could predispose the model toward the emergence of topographic maps. However, the model results show that this bias is not by itself sufficient to determine the final organization (Figure 3B,E,G– I). When V1–S1 correlations are weak, many RL neurons decouple from the primary sensory inputs, resulting in poor topography and few bimodal neurons (Figure 3E,G–I). Conversely, when V1–S1 correlations are very strong, the two input maps become highly aligned, but topography is degraded because many RL neurons receive similar visual and somatosensory inputs at the same topographic location, thereby overriding the initial topographic bias (Figure 3E,G,H). Thus, the initial bias does not simply determine the final map structure. Rather, appropriate topography, map alignment, and the emergence of a mixture of unimodal and bimodal neurons require an intermediate level of correlated activity.

      We have revised the manuscript to make this interpretation clearer. We also strengthened the experimental basis for the activity structure used in the model by revising Figure 2 and the corresponding Results and Methods. The revised analyses now show near-synchronous but spatially distinct activation of V1, RL, and S1, a lagged cross-correlation analysis arguing against stereotyped travelling-wave-like propagation between V1 and S1, and functional correlation maps computed from common RL seed locations. Together, these additions clarify the spatial and temporal structure of the spontaneous activity used to motivate the model.

      Finally, as described in our response to Reviewer 1, point 4, we have extended the model to test the role of the initial bias more directly in combination with experimentally measured differences in V1 and S1 activity patterns. In this analysis, we incorporated these activity differences and examined how they interact with different levels of initial connectivity bias (Supplementary Figure S2). These simulations show that more mature S1-like activity patterns alone are not sufficient to generate the appropriate topographic and aligned maps, and that an initial connectivity bias is required. At the same time, consistent with Figure 3, this bias does not by itself determine the final organization; the outcome also depends on the temporal correlation structure of V1 and S1 activity.

      We agree that the initial topographic bias remains an important modeling assumption, consistent with the idea that coarse activity-independent mechanisms provide an initial scaffold for later activity-dependent refinement. We now present the model accordingly: not as showing that correlated activity alone creates topography from an entirely unstructured circuit, but as showing how structured spontaneous activity can refine an initially coarse topographic scaffold to produce aligned multisensory maps and a mixture of unimodal and bimodal RL neurons.

      (4) In the Introduction, the authors nicely and briefly explain the role of primary and higher order sensory cortices in information processing. They also explain how spontaneous activity during development helps to build these circuits by refining connections or establishing hierarchies. They continue explaining the relevance of aligning different topographic maps to allow multisensory integration. Then they provide some examples of sites of multisensory integration. This provides a general context for the data presented in the Results section; however, and importantly, there is no specific introduction of why they are interested in RL and its interaction with V1 and S1. The authors should introduce the RL area and explain why it is an interesting site for multisensory processing.

      We thank the reviewer for pointing this out. We have revised the Introduction to make the rationale for focusing on RL more explicit. Specifically, we now introduce RL as a higher-order cortical area located between V1 and S1 that receives topographically organized input from both primary sensory cortices and contains overlapping visual and tactile representations. We also clarify that RL is a particularly relevant area for studying multisensory map alignment because corresponding locations in visual and whisker space can converge onto the same RL neurons, including bimodal neurons. Finally, we expanded the Introduction to explain that RL has been implicated in visually guided tactile behaviors and cross-modal generalization, making it an appropriate model system for studying how aligned multisensory representations emerge during development.

      (5) The results shown in Figure 1 corroborate published data from Golshani et al, Rochefort et al, Murakami et al. While the reproduction of data is more than welcome, the authors should specify which part of the data is completely new and acknowledge clearly the rest as corroboration of previous data. The sentence "As described in previous experiments ..." partially acknowledges this fact but is not clear enough. In addition, the transition between this part of the manuscript and the next data is not smooth. Data seems to be used to feed the model so perhaps the organization of the manuscript leaves room for improvement.

      We thank the reviewer for pointing this out. We have therefore revised the Results to clarify that the developmental sparsification of spontaneous activity in V1 is consistent with previous work, including Portera-Cailliau, Konnerth, Hanganu-Opatz, Crair and Ohki labs as well as our own (Siegel et al. 2021) and that similar developmental trends in S1 and RL corroborate and extend these observations across the sensory and higher-order cortical areas analyzed here.

      We also clarified what is new in the present analysis. Specifically, our contribution is not simply to reproduce previously described developmental sparsification, but to compare V1, S1, and RL within the same experimental and statistical framework, revealing that S1 exhibits more mature activity features earlier than V1 and RL. We also revised the transition to the next section to make clearer how these measurements motivate the subsequent analysis of temporally and spatially correlated spontaneous activity between V1, S1, and RL.

      Reviewer #3 (Public review):

      Summary:

      The study by Dwulet et al. explores how the development of spontaneous neural activity in primary sensory cortices influences the co-alignment of multiple sensory modalities in higher order brain areas (HOAs). To address this question, they focus on connectivity between the primary visual (V1) and somatosensory (S1) cortices and an associative cortical area (RL) in mice. The authors combine experimental (wide-field and two-photon calcium imaging) and computational approaches to show that spontaneous activity matures at a different pace across these brain regions. Their data indicate that S1 develops more rapidly than V1, which is possibly beneficial for RL's integration of visual and somatosensory inputs through correlated spontaneous activity. Using a computational model, they demonstrate that a moderate correlation between V1 and S1 activity can optimally guide the formation of bimodal neurons in RL, which are crucial for maximizing the decodability of multisensory stimuli. This finding highlights the role of correlated spontaneous activity in primary sensory cortices in establishing co-aligned topographic multimodal sensory representations in downstream circuits.

      Strengths:

      The manuscript is well written and it provides strong enough evidence to support the main claim of the authors. The insights on the role of correlated activity on instructing co-aligned multisensory maps in HOAs are not trivial and are an important advancement for the field.

      Weaknesses:

      In the opinion of this reviewer, the study has no major weaknesses. A drawback of the work is that none of the predictions of the computational modeling have been corroborated through mechanistic experimental manipulations of early brain activity.

      We thank the reviewer for their positive assessment of the manuscript and for highlighting the importance of the model predictions. We agree that a direct mechanistic perturbation of early spontaneous activity would provide an important future test of the model. Such experiments could, for example, perturb the temporal correlation structure between V1 and S1 during the relevant developmental window and then test whether this affects the alignment of V1/S1 maps in RL and the emergence of bimodal RL neurons.

      In the present study, we focused on identifying candidate features of spontaneous activity that could instruct multisensory map alignment and testing their sufficiency in a computational model. We now explicitly acknowledge in the Discussion that causal perturbations of early spontaneous activity will be needed to validate the model predictions experimentally. We believe this provides an important direction for future work while preserving the main conclusion of the current study: that structured, moderately correlated spontaneous activity provides a plausible developmental mechanism for refining aligned multisensory representations in higher-order cortex.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Additional comments/suggestions for the figures:

      (1) In Figure 1D-G, some of the dots lie almost directly on top of each other, essentially "hiding" certain data points. Using different shapes for each of the three regions might help alleviate this issue and make the data more visually distinct.

      We thank the reviewer for this suggestion. We have revised Figure 1D-G so that the three cortical regions are shown with different marker shapes. This should make overlapping data points easier to distinguish and clarify that each point corresponds to the average value for one animal and cortical region at the indicated postnatal age.

      (2) In Figure 2D-E, RGB color values are used to represent the highest correlation coefficient across the three seeded areas. It would be more informative if these also depicted the magnitude of the correlations, possibly through a color gradient. Additionally, the black regions in these panels are not currently defined and should be clarified.

      We have revised the functional correlation-map analysis and its presentation in Figure 2, as suggested by the reviewer. In the revised figure, the main correlation-map panels now use three seed locations in RL and show the resulting correlations across the imaged cortical field. We present the maps in two complementary ways. First, the raw RGB correlation map shows the correlation values for all three seed locations, with the intensity of each color channel reflecting the magnitude of the corresponding Pearson correlation coefficient (new Fig. 2E). Second, the maximum-channel representation assigns each pixel to the seed location with the strongest correlation, while still preserving correlation strength through pixel intensity (new Fig. 2F).

      We have also added color scales to relate pixel intensity to correlation magnitude and clarified that black pixels in the maximum-channel representation correspond to pixels below the correlation threshold used for visualization. The figure legend and Methods now describe how the RGB maps and maximum-channel maps were computed. Finally, we added Supplementary Figure S1 with additional examples from PN9, PN10, and PN13 recordings, with seed locations in V1, S1, or RL as indicated in each panel. This illustrates that they are quite similar across the ages investigated here.

      (3) I found Figure 3F a bit difficult to interpret without referring to the Methods section for the definitions of Topography and Alignment. Since these definitions are relatively short and essential for understanding all the modeling figures, I suggest moving them into the main text where they are first introduced.

      The definitions of Topography and Alignment have been added to the text where they are introduced.

      (4) In Figures 3-5, it is unclear what causes the variability in the model’s responses, as there are two potential sources of randomness: the initial random connectivity matrix and the correlated inputs driving the system. Are either of these fixed? For example, is the distribution of dots along the y-axis in Figure 3G-H, which corresponds to zero correlation between V1 and S1, driven by variability in the initial connectivity matrix, the random timing of input events, or a combination of both? If it’s a combination, it would be interesting to tease this effect apart by fixing one form of randomness and recreating these plots.

      In the original simulations in Figures 3–5, neither source of randomness was fixed across runs: each point corresponds to an independent developmental realization with a newly sampled initial connectivity matrix and a newly sampled sequence of spontaneous input events. The initial connectivity was random but weakly biased toward matched topographic location, while spontaneous activity consisted of stochastic independent and shared events activating randomly chosen contiguous groups of neurons (as explained in the main text and Methods). Thus, for example, the spread of points at zero V1–S1 correlation in Figures 3G–H reflects a combination of variability in the initial connectivity and variability in the independent V1 and S1 event histories. At zero correlation, no shared V1–S1 events are present, so this spread does not reflect variability in correlated shared events, but rather run-to-run differences in the two independently refined maps.

      We have clarified this point in the text and figure legend. We agree that fixing one source of randomness while varying the other would be an interesting additional analysis to decompose the relative contribution of initial wiring versus input history. However, the goal of the present simulations was to characterize the ensemble of possible developmental outcomes when both initial connectivity and spontaneous activity vary, as expected biologically.

      This interpretation is also consistent with the earlier two-layer model from developmental refinements from retina/thalamus to V1 (Wosniack et al., eLife 2021) on which our model builds, where final receptive fields emerge from the interaction between weak biased initial connectivity and stochastic structured spontaneous activity. In the current three-layer extension, the same principle applies to two converging projections, from V1 to RL and from S1 to RL. The initial topographic bias constrains the possible map structure, while the spatiotemporal statistics of V1 and S1 activity determine whether the two maps remain separate, align, or collapse into overly bimodal representations.

      (5) The specific parameter values used to create the panels in the modeling figures (Figures 3 and 4) should be made clearer, at least in the figure captions. For example, in Figure 3E, the exact values for the “weak,” “medium,” and “strong” correlations should be provided. Additionally, Figure 4 does not mention the strength of the correlated input considered, which should be specified as well.

      The values for the weak, medium and strong correlations have been added to the figure caption of Figure 3. The input correlation for Figure 4 is also now specified in the figure caption.

      (6) There is an odd vertical line in Figure 3I that doesn’t appear to be discussed or defined. Its purpose should be clarified, or the line should be removed if it is unintentional.

      This line has been removed.

      (7) There is a typo in the caption for Figure 3. Panel 'K' should be panel 'J'.

      This typo has been corrected.

      (8) In the text, the authors write "With these connectivity refinements, the generated activity in RL became sparser in terms of amplitude and participation rate (Figure 3J)." While this appears to be the case for this single example, it is difficult to confirm without zooming in on the panel. These quantities should be computed across multiple instances, and a summary plot should be provided to support this statement.

      The experimentally measured developmental sparsification of RL activity is quantified (independent of the model) in Figure 1D–F.

      We see how the original wording placed too much weight on the illustrative example in Figure 3J. We have revised the text to clarify that Figure 3J shows a representative simulation illustrating how RL activity changes as V1/S1-to-RL connectivity refines, rather than a separate population-level quantification across model instances.

      At the same time, this example is not meant to introduce a new, unsupported mechanism. The model used here is an extension of our previous two-layer model of developmental refinements between retina/thalamus and V1, in which spontaneous activity refined feedforward receptive fields from thalamus to V1. In that study, we specifically quantified how receptive field refinement led to sparsification of cortical activity in V1 over development, including reduced event amplitude, reduced event size/participation, and reduced pairwise correlations (Wosniack et al., 2021). Thus, the example shown in Figure 3J is consistent with a mechanism that has already been systematically characterized in the simpler two-layer setting.

      In the present manuscript, the central modeling results concern the emergence of topography, alignment, and the balance of unimodal and bimodal RL neurons. We therefore have softened the corresponding statement and explicitly refer to Figure 3J as an illustrative example.

      (9) Figure 5C is a bit difficult to interpret. The corresponding text states, "However, when activity across V1 and S1 is moderately correlated, having some unimodal RL neurons can achieve a higher total maximum fraction of variance for both V1 and S1 compared to the purely bimodal case (Figure 5C)", from which I infer that these dots represent networks resulting from "moderate correlations." However, the exact range of correlations considered should be mentioned in the text or figure caption. Additionally, I find it unusual that some networks with close to 0% bimodal cells perform quite well in reconstructing both S1 and V1. Many data points overlap, but I notice quite a few pale dots in the upper right of the plot. I believe this should be addressed in the main text.

      We thank the reviewer for this helpful comment. We have added the correlation values used for the simulations in Figure 5C to the figure caption and clarified the interpretation in the Results. The high reconstruction performance for some networks with relatively few bimodal cells arises because, when V1 and S1 activity are not perfectly correlated, unimodal RL neurons can provide unambiguous information about activity in one sensory area. In contrast, a purely bimodal population can make it more difficult to distinguish whether one or both primary sensory cortices were active. Thus, for moderately correlated inputs, a mixture of unimodal and bimodal RL neurons can reconstruct both sensory areas better than a population composed entirely of bimodal neurons. We have revised the main text to make this point explicit.

      (10) The network schematics in Figures 3A and 5A could be improved to better illustrate the network setup using a similar approach as the one used by this research group in Wosniack et al. (2021). Adding arrowheads to the lines from V1/S1 to RL would clarify that these are purely feedforward inputs. It would also be helpful to depict that V1 and S1 are driven by correlated events that are spatially structured.

      We thank the reviewer for this helpful suggestion. We have revised the schematics in Figures 3A and 5A to make the feedforward nature of the model clearer by adding arrowheads to the projections from V1 and S1 to RL. We have also clarified the depiction and description of the input activity. Specifically, Figure 3C shows the spontaneous events driving V1 and S1 in the model, including shared events that are both temporally correlated and spatially structured across corresponding topographic locations in the two primary sensory areas. These shared events activate matched contiguous groups of neurons in V1 and S1, while independent events activate randomly chosen contiguous groups within each area. We have clarified this point in the Results and Methods.

      General comments regarding the text (including typos):

      (1) In Statistical analysis, "In wide-field calcium imaging (we re-analyzed data from [46] (Figure 1))..." should be referencing Figure 2.

      Typo fixed.

      (2) Right before Table 1, the authors mention that they ran the simulations for 500,000 milliseconds, which is 500 seconds. This doesn't seem long enough for the weights to approach their steady-state values given the inter-event interval. Since the example simulations in Figure 3 are 1,000 seconds long, I'm guessing this is a typo.

      Typo fixed. Indeed the simulations in Figure 3 were 1,000 ms (1 s) long.

      (3) The specific time step used for the simulations should be specified. Currently, the text only mentions "sufficiently small time steps".

      We have now specified the simulation time step in the Methods.

      (4) In the Rate-based network model section, you write "These biased weights decay with a Gaussian profile with increasing distance (Figure 3)), with amplitude a and spread s", but Table 1 denotes these parameters differently.

      We have corrected the notation so that the parameter names are consistent between the Methods and Table 1.

      (5) Currently, all differential equations are written as 1/tau*df/dt. Based on the units of your time constants (seconds), I believe these equations should be tau*df/dt.

      We have corrected the differential-equation notation.

      (6) Equations 5-6 and 8 should be differential equations.

      We have corrected these equations so that they are written as differential equations. These mistakes happened because we changed formats between from Word to Latex.

      (7) The expectation in Equation 8 is not clearly defined and I would think here that the W_ij's should be within expectations. In the next paragraph, the authors specify that they are interested in a specific case of W_ij's, but this condition has not been introduced yet.

      We thank the reviewer for pointing out this ambiguity. We have revised the text around Equation 8 to define the expectation more clearly and to introduce the specific steady-state connectivity configuration before it is used. Because the expectation is taken over the input activity statistics at steady state, the weights are fixed quantities in this calculation. Including W_ij inside the expectation would therefore not change the result, but we have revised the notation and explanatory text to make this clearer.

      (8) The expectation in Equation 8 is not clearly defined, and I believe that the W_ij’s should be included within the expectations (in the following paragraph, the authors mention that they are interested in a specific case of W_ij’s, but this condition has not yet been introduced).

      This comment is the same as the one above. Please see the point above for the reply.

      (9) At the start of "Optimal weight matrix for correlated input populations", you write that the vector X is M x 1. If that is the case X'X would be a 1x1 matrix. I'm not sure if you meant to write X as 1 x M or to examine XX'.

      We thank the reviewer for pointing out this dimensional inconsistency. We have corrected the notation in the Methods. The concatenated input vector X=[v; s] has size M x 1, so the relevant input covariance matrix is X X^T not X^T X. This covariance matrix has size M x M, as required for the eigenvector analysis. We revised the corresponding equations and explanatory text accordingly.

      (10) Equation 11 has an s_i on the right-hand side that should be a \mu_s.

      Typo fixed.

      Reviewer #2 (Recommendations for the authors):

      Some sentences may require more scientific rigor. For instance: "We found that activity between the visual and the somatosensory cortex is often, but not always, temporally synchronized.

      We have revised the Results to state the quantitative observations more explicitly. Specifically for this example, we now report that the average activity in V1 and S1 across PN9PN12 animals showed a range of Pearson correlation coefficients with a mean of approximately 0.5. We also describe the examples in Figure 2B-D as near-synchronous but spatially distinct activation of subregions in V1, RL, and S1, and we use the lagged cross-correlation analysis in Figure 2G to support the conclusion that V1-S1 correlations peak near zero lag rather than reflecting stereotyped propagation with a fixed delay.

      Reviewer #3 (Recommendations for the authors):

      Minor suggestions on how to improve some specific aspects of the manuscript.

      Introduction:

      (1) What do the authors mean when they write "Higher-order areas (HOAs) situated between primary sensory areas"? This sentence might need some editing.

      We have revised the sentence to clarify that we are referring to higher-order cortical areas that receive and combine inputs from multiple primary sensory areas. We now also state explicitly that some of these areas, including RL, are anatomically positioned between the primary sensory cortices whose inputs they integrate.

      (2) In later portions of the manuscript, it becomes clear what the authors mean when they write “whereby sensory neurons converge onto higher-order cortex while preserving space”, but I think that it would be beneficial if this statement would be better explained also in the introduction.

      This has been clarified in the introduction. Specifically, we now clarify that topographic convergence means that neurons representing corresponding regions of sensory space in different primary sensory areas can project to overlapping or nearby locations in higher-order cortex. In the case of RL, this means that visual and tactile representations with corresponding spatial organization can converge onto RL neurons, including bimodal neurons.

      (3) Could the authors provide some more information about RL and the rationale as to why it was chosen as the HOA that they investigated in the study?

      We have expanded the Introduction to make the rationale for focusing on RL more explicit. We now introduce RL as a higher-order cortical area located between V1 and S1 that receives topographically organized input from both primary sensory cortices. We also explain that RL contains overlapping visual and tactile representations, including bimodal neurons, and that corresponding locations in visual and whisker space can converge in RL. In addition, we now note that RL has been implicated in visually guided tactile behavior and cross-modal generalization. These anatomical and functional properties make RL a particularly suitable model system for studying how aligned multisensory representations emerge.

      Results:

      (1) "RL was found to slightly lag behind V1 and S1". On what evidence is this statement based upon? As far as I can understand, there are no significant differences between V1 and RL besides amplitudes being higher in RL, which I don't think can be univocally interpreted as a sign that RL lags behind V1 in the developmental profile.

      The evidence for a delayed RL maturation relative to V1 and S1 is limited and comes from the pattern of coefficient estimates in the linear mixed models, now shown in Supplementary Tables S1-S3, rather than from a robust difference across all measured activity features. We have therefore revised the Results to state more conservatively that RL and V1 develop more similarly during the second postnatal week, while S1 shows more mature activity features earlier in development. The full linear mixed-model comparisons using V1, S1, and RL as reference areas are provided in Supplementary Tables S1-S3.

      (2) Figure 1H is very hard to read.

      (a) The slopes and the intercepts have values that differ by orders of magnitude, so the slopes get squeezed and become invisible. Further, the different parts of the plots (e.g. the one of amplitude and duration) are almost overlapping, which is a bit confusing. Slopes and intercepts should also have different units of measure (see Equation 3), so I wonder how they can lie on the same axis. Can the authors try to plot the data in a manner that is easier to visually inspect?

      (b) Including the "reference" (V1) intercept in H is also a bit misleading, as one might intuitively interpret it as a difference between V1 and other brain areas. Perhaps the overall differences between brain areas (regardless of age) might be best represented in a plot without age on the x-axis (only brain area). Alternatively, one might point them out directly on the plots in DG.

      (c) In D-G, what do the individual dots represent? The legend states N=10 animals, but I only see ~6 dots per plot.

      We thank the reviewer for these helpful points. We have revised the caption of Figure 1H and added Supplementary Tables S1-S3, which provide the full linear mixed-model estimates for each choice of reference area. These tables report the intercepts, slopes, interaction terms, confidence intervals, and significance levels in a format that avoids placing quantities with different units and scales on the same visual axis.

      For the caption of Figure 1H: The V1 value corresponds to the model intercept at PN8, whereas the age coefficient corresponds to the slope for V1. The S1, RL, Age: S1, and Age: RL terms represent differences relative to this reference model. To avoid the impression that the V1 intercept represents a difference between areas, we now explicitly state that the coefficients in Figure 1H are interpreted relative to V1 at PN8, and that the complete comparisons using S1 and RL as reference areas are provided in Supplementary Tables S2 and S3.

      Finally, we clarified that the individual points in Figure 1D–G represent animal-level averages for each cortical area at the indicated age. The value N = 9 refers to the total number of animals included across the dataset, not to the number of animals at each postnatal age. Because recordings were distributed across ages and some points overlap visually, fewer points are visible in individual panels than the total N.

      (3) Figure 2B-C: at which lag does this correlation peak? Is it at 0ms? Or does one brain area precede/follow the other one?

      We thank the reviewer for this comment. We have revised Figure 2 to include a lagged V1–S1 cross-correlation analysis. The V1–S1 correlation peaks close to zero lag and decreases for both positive and negative lags, indicating that the dominant temporal relationship is near-synchronous rather than consistent with fixed-delay propagation from one primary sensory cortex to the other. The curves show a mild asymmetry, with somewhat stronger correlations when S1 precedes V1, but because the dominant peak is near zero lag, we interpret the data primarily as evidence for near-synchronous, spatially structured coactivity across areas rather than stereotyped travelling-wave propagation. We have added this interpretation to the Results and clarified the temporal-lag convention in the Figure 2 legend.

      (4) Figure 2D-E: in the methods section the authors report that "The actual color of each pixel represents the highest coefficient of correlation value across the three channels." I think that this important information should be included in the main text or the legend of the figure.

      We have changed Fig. 2 now to clarify the quantification of the functional correlation maps and also added the information requested by the reviewer to the figure legend.

      (5) Figure 3D: I think that it would be beneficial if the authors would highlight directly in the figure that those connectivity matrices are between V1/S1 and RL.

      This information has been added to the figure.

      (6) Figure 3I: does the vertical line correspond to the "critical amount of temporal correlation" (eq. 2)? If so, could the authors provide this information in the figure or the figure legend?

      This line was unintentional and has been removed.

      (7) It would be nice if the data that was generated for this study (and the data that has already been published and was used to generate Figure 2) would be made publicly available on an open-access repository.

      We agree that open data sharing is important. We have made the code used for the model and figure generation available in the repository listed in the Data and Code Availability section. At present, we are not able to deposit the complete raw imaging datasets in an open repository because the wide-field and two-photon imaging files are very large, amounting to multiple terabytes, and we do not currently have a sustainable hosting solution for these raw data. We will share data upon request, and we will deposit the raw imaging datasets in an appropriate open repository if a feasible long-term hosting solution becomes available.

      References

      M. Chini, T. Pfeffer, and I. Hanganu-Opatz. An increase of inhibition drives the developmental decorrelation of neural activity. eLife, 11:e78811, 2022.

      P. Golshani, J. T. Gonçalves, S. Khoshkhoo, R. Mostany, S. Smirnakis, and C. PorteraCailliau. Internally mediated developmental desynchronization of neocortical network activity. Journal of Neuroscience, 29(35):10890–10899, 2009.

      A. Gribizis, X. Ge, T. L. Daigle, J. B. Ackman, H. Zeng, D. Lee, and M. C. Crair. Visual cortex gains independence from peripheral drive before eye opening. Neuron, 104(4):711–723.e3, 2019.

      S. Lakhera, E. Herbert, and J. Gjorgjieva. Modeling the emergence of circuit organization and function during development. Cold Spring Harbor Perspectives in Biology, 17(2):a041511, 2025.

      A. H. Leighton, J. E. Cheyne, G. J. Houwen, P. P. Maldonado, F. De Winter, C. N. Levelt, and C. Lohmann. Somatostatin interneurons restrict cell recruitment to retinally driven spontaneous activity in the developing cortex. Cell Reports, 36(1):109316, 2021.

      H. Matsumoto, T. Murakami, and K. Ohki. Topographic correspondence between retinotopic and whisker somatosensory map in mouse higher visual area and its development. Frontiers in Neural Circuits, 19:1552130, 2025.

      T. Murakami, T. Matsui, M. Uemura, and K. Ohki. Modular strategy for development of the hierarchical visual network in mice. Nature, 608:578–585, 2022.

      N. L. Rochefort, O. Garaschuk, R.-I. Milos, M. Narushima, N. Marandi, B. Pichler, Y. Kovalchuk, and A. Konnerth. Sparsification of neuronal activity in the visual cortex at eyeopening. Proceedings of the National Academy of Sciences of the United States of America, 106(35):15049–15054, 2009.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study builds on earlier work showing that early-life odor exposure can trigger glial-mediated pruning of specific olfactory neuron terminals in Drosophila. Moving from indirect to direct functional imaging, the authors show that pruning during a narrow developmental window leads to long-lasting suppression of odor responses in one neuron type (Or42a) but not another (Or43b). The combination of calcium and voltage imaging with connectomic analysis is a strength, though the voltage imaging results are less straightforward to interpret and may not reflect synaptic output changes alone.

      Strengths:

      Biologically, one of the main strengths of this work is the direct comparison between two odor-responsive OSN types that differ in their long-term adaptation to early-life odor exposure. While Or42a OSNs undergo pruning and remain persistently suppressed into late adulthood, Or43b OSNs, which also respond to the same odor, show little lasting change. This contrast not only underscores the cell-type specificity of critical-period plasticity but also points to a potential role of inhibitory network architecture in determining susceptibility. The persistence of the Or42a suppression well beyond the developmental window provides compelling evidence that early glia-mediated pruning can imprint a stable, life-long functional state on selected sensory channels. By situating these functional outcomes within the context of detailed connectomic data, the study offers a framework for linking structural connectivity to long-term sensory coding stability or vulnerability.

      Weaknesses:

      The narrative begins with the absence of changes in PN dendrites and axons. While this establishes specificity, it is a relatively weak starting point compared to the novel OSN functional results.

      We agree that switching the order of Figures 1 and 2 recontextualizes the negative PN morphology findings to make their significance more clear, especially with the addition of PN odour-evoked activity data (see Figures 2A, B of the revised manuscript).

      Calcium imaging with GCaMP, though widely used, is an indirect measure of synaptic function, and reduced signals could reflect changes in non-synaptic calcium influx as well as release probability. The interpretation of the voltage imaging results is also unclear: if suppression were solely due to impaired synaptic release, one might expect action potential-evoked voltage signals to remain unchanged. The reported changes raise the possibility of deficits in action potential initiation or propagation, which would shift the mechanistic explanation.

      Although it is true that non-synaptic Ca<sup>2+</sup<> influx could contribute to odour-evoked signals in OSN axon terminals, it seems likely to be a relatively small contribution when compared to Ca<sup>2+</sup> influx via voltage-gated Ca<sup>2+</sup> channels at the active zone. Given the observation that synaptic markers are eliminated during this form of critical period plasticity and remain decreased even after OSNs regrow their terminals days later (consistent with our observed continued decrease in odour-evoked responses), the most parsimonious explanation is that we are seeing a reduction in synaptic Ca<sup>2+</sup> influx. We cannot dismiss the possibility that there is a decreased voltage signal arising from fewer action potentials being elicited by the odour stimulation. However, the reduction in voltage signal must arise at least in part from the observed reduction in Ca<sup>2+</sup> influx. We have therefore provided additional text to this effect in the results section.

      The difference between Or42a and Or43b OSNs is attributed to varying inhibitory input densities from connectome data, but this remains speculative without functional tests such as manipulating GABA receptor expression in OSNs. In Or43b, there is essentially no strong phenotype, making it premature to ascribe the absence of suppression solely to inhibitory connectivity.

      We have tempered our conclusions to posit additional mechanisms that could explain the more mild pruning that occurs for Or43b OSNs. While the pruning phenotype for Or43b OSNs is not as strong as Or42a, it is not absent. To further explore the contribution of inhibition as a candidate mechanism underlying differences in susceptibility of Or42a and Or43b to this form of critical period plasticity we compared the relative impact of knocking down expression of GABA-A receptor (called “rdl”) in Or42a and Or43b OSNs. Consistent with the degree of pruning being regulated inhibition, knocking down expression of rdl enhanced pruning for both Or42a and Or43b OSNs. However, because the magnitude of the enhancement was similar between both OSN types, we agree with the reviewer that inhibitory connectivity cannot be the sole mechanism that explains the difference and have therefore tempered our language appropriately.

      Finally, the study does not connect circuit-level changes to behavioral outcomes; assays of odor-guided attraction or discrimination could place the findings in an organismal context.

      We agree that behavioral assays will be a critical component for understanding the functional consequences of this form of critical period plasticity. However, the goal of this study was to extend our prior work to determine the longevity and selectivity of the critical period pruning. Behavioral assays testing the consequences of this form of early life plasticity will be a component of future studies.

      Some introduction material overlaps with the authors' 2024 paper, and the novelty of the present study could be signposted more clearly.

      We have included text to highlight the novelty of the present study.

      Reviewer #2 (Public review):

      Recent work from the authors identified the synaptic changes and glial reaction that occur during exposure of a Drosophila odorant receptor neuron population to continued exposure of a stimulating odorant. This work markedly advanced our understanding of cellular response to critical periods. This current Advance manuscript carries that work forward and examines the non-autonomous responses to constant odorant exposure. The authors discover that the changes to ORN populations are not accompanied by changes to either PN dendrite or PN axon volume, nor are they concurrent with changes in postsynaptic PN structures. These changes are, however, notable, accompanied by changes in Ca2+ and voltage responses in ORNs. Importantly, this set of responses is specific to the Or42a ORNs (that are highly sensitive to the odorant in question, ethyl butyrate) and not the Or43b ORNs (which respond to ethyl butyrate, but not as drastically). Finally, the authors include connectomics analyses showing that Or43b and Or42a ORNs differ in their synaptic input/output relationships.

      This is an excellent use of the Advance mechanism for the journal, as these are important follow-up findings for the parent story. The non-autonomous effects (or lack thereof) on PNs is an important part of the story, as is the functional response of Or42a ORNs and the differing response of similarly (but not identically) sensitive Or43b ORNs. The experiments are well-conceived, controlled, and conducted. Where the story falters a bit, though, is with the connectomics analysis. The authors show distinct differences between Or43b and Or42b ORN input-output relationships, and suggest that those differences may underlie the differences observed in their response to ethyl butyrate exposure during the critical period. This is certainly a possibility, but as it stands now, it is too disconnected to offer significant proof. There would have to be additional experiments to address this. Right now, the inclusion of the connectomics work feels like a distraction at best, and a complete non sequitur at worst. To be clear, the connectomics work is well done and I have no issues with its validity, but it is not helpful to the central thesis of the work. I would suggest the authors either remove it entirely or strongly rethink how it fits into the paper.

      We have tempered our stated interpretations of the connectivity analysis and include new experiments examining the impact of GABA signaling on pruning. We have therefore opted to retain the connectivity analysis as we feel that it has been better integrated into the overall narrative of the paper.

      Major Concerns:

      (1) The examination of PN axon terminals in the MB and LH is interesting, but it is only one possibility. Oftentimes, the volume of neurons remains constant with perturbation, while the synapse number is affected. Figure 1C and E would be greatly helped by examining synapse number (via Brp or Brp-Short) in the PN axons.

      We agree that the counting synapse number would provide greater resolution information about synapse function relative to axon volume and have added this analysis to what is now Figure 2.

      (2) The use of dlg1[4K] is a strong use of a new tool, but the result is surprising. The presynaptic ORN synapse number onto the PNs is notably changed, but that is not reflected in a postsynaptic PSD-95 change. That suggests a compensatory mechanism that the authors might explore. A good proportion of PN puncta should be postsynaptic to those ORNs, so why aren't they adjusted?

      We agree that this result suggests that a compensatory mechanism may be present. We have therefore added new text to point out this observation and potential explanation.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The interpretation of the voltage imaging results would benefit from clarification. If these signals are reduced because of upstream action potential changes rather than synaptic release, this should be explicitly discussed and illustrated with representative raw traces for both OSN types. The proposed link between inhibitory connectivity and selective vulnerability could be tested more directly, for example, by manipulating GABA receptor function in OSNs.

      We have now tested the link between inhibitory connectivity and susceptibility to glial pruning by testing the effects of GABA receptor knockdown in either Or42a or Or43b OSNs (fully described above).

      Adding an intermediate post-exposure time point for Or42a responses could help resolve whether suppression is immediate or develops over time.

      The suppression of Or42a odour-evoked responses is present immediately after the 2 day exposure period and responses remain suppressed until 25 days post-eclosion, indicating that the suppression is immediate and sustained. We therefore respectfully disagree that another physiological time point will help resolve whether the suppression is immediate or develops over time.

      In terms of presentation, the introduction could be tightened to reduce overlap with the 2024 paper, figures should have clear axis labels and consistent terminology for neuron types and glomeruli, and a schematic summarising key inhibitory connections for Or42a vs. Or43b would aid clarity

      We have now streamlined the introduction, improved clarity on axis labels and checked for consistency of terminology.

      Minor Concerns:

      (1) The dlg1[4K] is made with a V5 epitope but the authors have it labeled mCD8::GFP in Figure 1F. This is likely a typo and should be corrected.

      This typo has now been corrected.

      (2) Can the responses be separated in Figures 2A, C, and E? It is difficult to see the differences in oil and EB exposure. This would make it much more straightforward to tell the difference if both traces were clearly visible.

      Overlaying the averaged response traces for in Figure 2C, E and G (now Figures 1C, E and G) enables the reader to make direct visual comparisons between the responses of OSNs from flies in each condition to both mineral oil and ethylbutyrate. Separating the individual traces would make it much more difficult to make these comparisons.

    1. Author response:

      We would like to thank all the reviewers and the editors for their considerate evaluation of our study.

      We are pleased that overall the reviewers were positive about the bulk of our study establishing a role of tissue macrophage programming/specialisation in regulating the macrophage lipidome, in the peritoneum, including the exemplar sphingolipid class. The reviewers raise understandable issues about the specificity of the available inhibitory compounds, such as zileuton meaning that conclusive statements about the role of LTE4 are not possible.

      In a revised manuscript, we will address all points but predominantly focus on the second aspect of the study, ensuring that reviewers comments are addressed appropriately, detailing and weaknesses, or ambiguities, with our study. This will include, but will not be limited to:

      - Further commentary on the regulation of eosinophil numbers within the tissue;

      - Addressing the specificity of zileuton and the implications of this for interpretation of our results with respect to eosinophil biology;

      - More careful framing of the transcellular biosynthesis potential;

      - A detailed discussion of sex dependency with regard to eosinophil numbers in general and any potential effect on the reported Gata6-dependent phenomenon;

      We are grateful for the constructive comments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper describes an interesting phenotype of C. elegans lite-1 mutants. Previous work showed that lite-1 mutants lose a violet/blue light avoidance response. The authors show here that lite-1 mutants also show a defect in negative diacetyl chemotaxis. While wild-type worms avoid diacetyl at high concentrations, lite-1 mutants are instead *attracted* to it. The authors go on to perform Ca2+ imaging in sensory neurons and find that ADL and ASK neurons show altered Ca2+ responses to diacetyl in lite-1 mutants, suggesting LITE-1 is required for these responses. As unc-13 mutants with defective synaptic transmission show similar diacetyl Ca2+ responses as wild-type, this suggests these neurons respond cell autonomously to diacetyl. However, whether lite-1 also acts cell-autonomously is not discussed. Indeed, because unc-13 and lite-1 mutants show different ADL and ASK Ca2+ responses, it seems the diacetyl response regulated by LITE-1 is likely acting outside of those cells. An interesting result that is not commented on is the switching of the valence of the ASK Ca2+ response in lite-1 mutants. ASK neurons still respond to diacetyl, but instead of a strong increase in Ca2+, diacetyl appears to drive it strongly lower. This may be consistent with the switch in valence in the diacetyl chemotaxis assay. It also argues against the idea that LITE-1 is a low-affinity diacetyl receptor that drives avoidance or the Ca2+ responses in ASK, since it is still present in lite-1 mutants. The authors then use a strain that expresses LITE-1 in the body wall muscles and show this expression is sufficient to engender them with sensitivity to diacetyl, as measured through altered swimming and hypercontractility. The authors interpret this result as LITE-1 may act as a diacetyl receptor. The authors test whether a structurally similar molecule, 2,3-pentanedione, shows similar effects, and they find it does. Alpha-fold modeling and molecular docking analysis show where diacetyl might bind to the LITE-1 protein. They then test whether lite-1 mutants show chemotaxis defects to other molecules, as seen with diacetyl. Generally, they find that the observed diacetyl responses are unique, although lite-1 mutants do lose their avoidance response to 2,3-pentanedione. However, unlike the acquisition of diacetyl attraction in lite-1 mutants, 2,3 pentanedione avoidance is *lost*; it is not switched to attraction. Overall, I felt the description of the results and their implications could have been more in-depth. Further, the evidence that LITE-1 is a chemoreceptor itself, rather than acting in some way to shape chemoreceptor responses (via light or otherwise), remains unclear, as conceded by the authors.

      Strengths:

      Overall, the study follows up on an interesting and useful result. The experiments as presented are generally well-conceived and performed. The authors use a variety of behavioral and imaging approaches to test how LITE-1 mediates diacetyl avoidance.

      Weaknesses:

      The study is missing experiments needed to resolve whether LITE-1 is doing what they propose. The evidence that LITE-1 is a diacetyl receptor is lacking support since lite-1 mutants have their avoidance and calcium responses flipped, which would not be expected if it were acting solely as an avoidance receptor. Presumably, the authors are concluding that the attractive response that is left in the lite-1 mutant is mediated by ODR-10, but that experiment is not shown.

      We interpret the shift from avoidance to attraction in lite-1 mutants as consistent with the loss of an aversive sensory component in the presence of an underlying attractive response to diacetyl. We initially hypothesised that this residual attraction was mediated predominantly by ODR-10. To test this, we now generated and analysed lite-1; odr-10 double mutants. The double mutants retained an attractive response to diacetyl, indicating that ODR-10 alone does not account for the attraction observed in the absence of LITE-1 and that additional receptors or sensory pathways are likely to contribute. This finding is consistent with previous studies where loss of ODR-10 did not lead to a complete loss of diacetyl responsiveness.

      Similarly, the authors concede that "the use of lite-1 point mutants that affect specific LITE-1 function, such as light sensing, channel gating, or binding pocket, could further elucidate LITE-1 mechanisms." This reviewer agrees, and such experiments designed to localize diacetyl binding site(s) would be necessary to conclude definitively that LITE-1 is a diacetyl receptor. The body wall muscle assay used or some other heterologous experimental system could work for such a structure-function analysis. A concern is whether the extensive number of LITE-1 point mutants described in the literature affect cell surface expression vs. receptor function, which might complicate the interpretation of a result showing loss of diacetyl responses.

      We agree that structure-function analysis using LITE-1 point mutants could help identify regions or residues that contribute to the diacetyl response and is an important future direction for research, which we have included in the discussion.

      Reviewer #2 (Public review):

      Summary:

      Koh and colleagues investigate the broader sensory role of LITE-1, a gustatory receptor previously linked to UV light detection in C. elegans. Their study explores whether LITE-1 also mediates avoidance of specific chemical stimuli-namely, high concentrations of diacetyl and 2,3-pentanedione. They show that LITE-1 is required in the ADL and ASK neurons for calcium responses to diacetyl, and that its expression in body-wall muscles is sufficient to trigger hypercontraction upon odorant exposure. Molecular docking suggests both odorants may directly bind to LITE-1 with micromolar affinity. These findings suggest LITE-1 may act as a multimodal receptor for both light and chemical stimuli.

      Strengths:

      (1) Methodological Precision: The study is technically strong, with well-executed calcium imaging and quantitative behavioral assays that clearly show neural and muscular responses to chemical stimuli.

      (2) Novelty and Scope: The work presents a compelling case for LITE-1 functioning as a multimodal sensor, which is an intriguing expansion of its known role.

      (3) Potential Impact: If validated, the findings could significantly advance the understanding of sensory integration in C. elegans, and the tools developed may be broadly useful to the research community.

      (4) Relevance to the Field: The study adds to evidence that C. elegans uses non-canonical sensory pathways and may inspire further exploration of multimodal receptor functions in other systems.

      Weaknesses:

      (1) Lack of Rescue Experiments: The absence of rescue experiments makes it difficult to definitively link the observed phenotypes to loss of lite-1.

      We have now performed the rescue experiment expressing lite-1 in ADL, and showed that LITE-1 in ADL is sufficient for avoidance, although it is not a complete rescue to wild-type levels.

      (2) Single Loss-of-Function Approach: The reliance on a single genetic mutant limits interpretability. Additional strategies such as RNAi (e.g., neuron-specific knockdown) would provide stronger evidence.

      We observed the loss of avoidance in three independent lite-1 alleles. Combined with the new cell-specific rescue experiment, we think this provides sufficient support for the conclusion that the phenotype is due to loss of lite-1 function.

      (3) Unclear Neuronal Contribution: While calcium responses in ADL and ASK are reduced, it's unclear which neuron(s) are necessary for behavioral avoidance. Cell-specific rescue or knockdown experiments are needed.

      We have expressed lite-1 genomic DNA under the ADL-specific promoter srh-220, which restored the avoidance phenotype, although it is not a complete rescue of wild-type behaviour. Together with calcium imaging data, this suggests that proper avoidance likely requires input from both ADL and ASK neurons.

      (4) Unvalidated Docking Data: The molecular docking predictions lack experimental validation. Site-directed mutagenesis would be needed to support claims of direct interaction.

      We agree that the docking data does not in itself establish direct binding (we think the muscle expression and paralysis provides stronger evidence). Based on previously reported docking experiments, we wanted to check if diacetyl could occupy the same binding pocket. We have now also included docking data of the other odorants from the chemotaxis assays in the manuscript.

      (5) Limited Odorant Specificity Testing: Docking analysis does not include non-binding odorants, making it difficult to assess binding specificity.

      We agree and have now included docking data of the other odorants from the chemotaxis assays. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      (6) Incomplete Quantification: Some calcium imaging results (e.g., in AWA neurons of unc-13 mutants) lack statistical comparisons, which limits their interpretive value.

      We have generated the scatter plots of calcium imaging responses across the different sensory neurons, and the statistical significance was assessed using two-sided t-tests with FDR correction, which is now included in the manuscript.

      Reviewer #3 (Public review):

      In this work, Brown and colleagues report that the photosensor protein LITE-1 of the nematode C. elegans may also be a chemosensor that can be activated by high concentrations of the compound diacetyl. LITE-1 was described as a putative ion channel of the gustatory receptor family, which is mainly constituted by insect odorant receptors. These form tetrameric ion channels that can be activated by odorants. Specificity is achieved by forming heteromeric channels from three copies of the odorant receptor co-receptor (ORCO) and another subunit that resembles ORCO in the pore-forming C-terminus, but brings in a binding site for the respective odorant. LITE-1 has a very similar structure, according to Alphafold3 predictions, and also carries a binding pocket. In LITE-1, this was proposed to be occupied by a light-absorbing molecule that activates the channel when a photon is absorbed. Alternatively, compounds generated by absorption of high-energy photons may be formed in vivo and bound by the LITE-1 binding pocket. Koh et al. now demonstrate that another, non-light-activated compound, diacetyl, at high concentrations, can activate cells expressing LITE-1. Such (chemosensory) cells are also responsible for the avoidance of high concentrations of diacetyl. LITE-1 activation in excitable cells, i.e, muscles, causes strong body contraction and paralysis, and the authors show that this is also the case when diacetyl is presented. The authors further present molecular docking studies showing that diacetyl could occupy the binding pocket of LITE-1. Last, they show that another compound chemically resembling diacetyl, i.e., 2,3-pentanedione, can also induce avoidance in a LITE-1 dependent manner, though not as potently.

      The data are intriguing, and the demonstration of LITE-1 being a diacetyl chemosensor is interesting. Yet, there are a few questions arising that the authors should address.

      The authors identified mutants lacking diacetyl responses. In their chemotaxis assay (Figures 1A, B), they show that lite-1 mutants do not avoid high concentrations of diacetyl. However, the animals actually showed attraction, as the chemotaxis index was positive. If the lite-1 animals were insensitive, they should be indifferent, and the chemotaxis index should be close to zero. This means, other neurons contribute to the diacetyl response, and the result of these neurons being activated means/remains attraction? If so, the authors need to rule out any effects of these neurons on the effects they attribute to LITE-1 in the other assays.

      We have tested tax-4 mutants in the chemotaxis assay and found that, contrary to the predicted chemotaxis index of zero, these animals retained strong avoidance of high concentrations of diacetyl. This indicates that tax-4 mutants are not chemosensory null for this stimulus and that TAX-4 independent sensory pathways contribute to high diacetyl avoidance. We agree that these experiments cannot completely rule out indirect neuronal effects. We have therefore revised the text to acknowledge this limitation. Nevertheless, the rapid paralysis and contraction observed when LITE-1 is expressed specifically in body-wall muscle in a lite-1 mutant background support the idea that LITE-1 is sufficient to confer a diacetyl-evoked response in these cells.

      The effect of diacetyl on muscle cells (Figure 3C) is pretty rapid, i.e., already during 1 minute after application, the animals are almost maximally contracted. How fast is it really? Can the authors provide a time course with more time points during the first minute? This is a relevant question, as the compound would have to either pass the worm cuticle or enter through the gut and diffuse through the body to reach the muscle cells. Can one expect this to occur within (less than) a minute? In this context, the authors need to rule out that other mechanisms may be at play. E.g., diacetyl may be immediately sensed by ciliated chemosensory neurons that might release a signaling molecule that leads to activation of LITE-1 in muscles, or that sensitizes it somehow, responding to light used for filming animals. The authors should repeat this assay in a lite-1 mutant background.

      We repeated the paralysis assays under red-filtered illumination to minimise potential effects of light, with animals maintained in darkness from hatching to adulthood. We also included lite-1 mutants to assess whether neuronal LITE-1 contributed to the paralysis response. In addition, the assay was repeated with more frequent time points, revealing that paralysis and body contraction occurred within 10 s and neuronal LITE-1 does not contribute to the effect.

      Furthermore, the authors tested unc-13 mutants to rule out indirect effects on the neurons recorded. Likewise, they should eliminate neuropeptide signaling via unc-31 mutants (a recent paper cited by the authors showed involvement of neuropeptide signaling in LITE-1-mediated light avoidance behavior).

      We agreed and have acknowledged and discuss in the manuscript that contributions from gap junction-mediated communication, neuropeptide signalling and other chemosensory pathways cannot be excluded.

      Last, to demonstrate that effects are not indirect in response to chemosensory neurons, the authors should repeat the contraction or swimming assay in a tax-4 mutant, which largely lacks chemosensation. This also applies to the chemotaxis assay. Animals should exhibit a chemotaxis index to diacetyl of zero, then.

      We have tested tax-4 mutants, and like wild-type animals, retained strong avoidance of high concentrations of diacetyl, indicating that TAX-4-independent sensory pathways contribute to this response. This indicates that tax-4 mutants are not chemosensory null for this stimulus and that TAX-4-independent sensory pathways contribute to high diacetyl avoidance. Therefore, repeating the contraction or swimming assay in a tax-4 background would not completely exclude indirect input from other chemosensory neurons. In addition, rapid paralysis and contraction were observed when LITE-1 is expressed specifically in body-wall muscle in a lite-1 mutant background, and together with the calcium imaging and rescue data, they support a role for ADL and ASK in mediating high diacetyl avoidance. The tax-4 chemotaxis data is now included in the manuscript.

      Does diacetyl activate other neurons expressing LITE-1? A number of cells express LITE-1 at high levels, which the authors have not tested (they restricted their analyses to chemosensory neurons). This is important to address because it leaves the possibility that LITE-1 requires a specific partner only present in these chemosensory neurons to detect diacetyl. This partner would have to be present also in muscles, where diacetyl could activate ectopically expressed LITE-1. According to CeNGEN scRNAseq data, cells expressing LITE-1 can be identified. The ADL and ASH neurons actually come up only at the lowest threshold, so some of the other cells showing much higher levels of LITE-1 mRNAs, i.e., AVG, ALM, PLM, ASG, PHA, PHB, AVM, RIF, or some pharyngeal neurons, should be tested. ASG was among the cells the authors recorded from, but this neuron did not show a response.

      We have acknowledged and discuss in the manuscript that other non-sensory neurons may contribute to the avoidance behavioural, and which should be the future direction for investigation.

      The authors need to show that diacetyl responses of ADL and/or ASK can be rescued by expressing LITE-1 specifically in these neurons in a lite-1 mutant background.

      We have expressed lite-1 genomic DNA under the ADL-specific promoter srh-220, which restored the avoidance phenotype, although it is not a complete rescue of wild-type behaviour. Together with calcium imaging data, this suggests that proper avoidance likely requires input from both ADL and ASK neurons.

      Molecular docking studies are not described in detail. How was this done?

      Molecular docking was performed in two stages. First, diacetyl was docked to the tetrameric LITE-1 model using DynamicBind without a predefined binding pocket. The generated complexes were ranked using the DynamicBind confidence score, and the highest ranked poses were used to identify the candidate binding site. The top DynamicBind pose was then used to define the box region for redocking with Gnina. Gnina poses were ranked using the CNN score. A more detailed molecular docking procedure has now been updated in the Methods section.

      Diacetyl is a very small molecule. How well can docking algorithms assess this at all?

      We agree that the small size of diacetyl limits the precision of docking scores because it forms relatively few protein contacts. However, its small size and limited conformational flexibility also simplify pose sampling. To increase robustness, we used two conceptually different docking approaches. First, DynamicBind was used without a predefined binding pocket to identify candidate binding regions while allowing ligand-associated protein conformational adjustments. Second, the resulting pocket was subjected to focused redocking and CNN-based pose ranking with Gnina. The results are interpreted as a structural hypothesis for the probable binding site and relative affinity ranking, rather than as definitive proof of binding or an accurate quantitative affinity measurement.

      Did the authors preselect the binding pocket, or did the algorithm sample the entire molecular surface of the LITE-1 model and end up with the binding pocket?

      The binding pocket was not predefined. Diacetyl was first docked to the tetrameric LITE-1 model using DynamicBind without specifying pocket residues or grid coordinates. DynamicBind therefore performed global, pocket-agnostic docking. The highest-ranked poses identified a candidate pocket, which was then used for focused redocking with Gnina.

      The latter would be very convincing. The authors should provide control docking experiments with other molecules that caused avoidance in their hands (i.e. benzaldehyde, 2,4,5,trimethlythiazole, isoamyl alcohol, nonanone, octanone), but did not activate LITE-1. Also, they should try docking molecules related to diacetyl, and if there are some that do not dock under the same conditions, such molecules should be used in a behavioral experiment. Ideally, they should also not activate LITE-1. Examples could be, e.g., diacetyl monoxime or 2,4-pentanedione.

      We have now included docking data of the other odorants from the chemotaxis assays. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      Last, the authors should provide a PDB file with the docked diacetyl to allow readers to assess the binding for themselves. Since a large number of mutations of LITE-1 have been reported, it may be that amino acids shown to be essential for LITE-1 function are also required for diacetyl binding. If so, this could be backed up with an experiment.

      We agree that structure-function analysis using LITE-1 point mutants could help identify regions or residues that contribute to the diacetyl response and which we have highlighted in the discussion as an important future direction for research. Additionally, we have now provided the PDB file containing a representative DynamicBind derived docking pose of diacetyl within the LITE-1 binding pocket.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) lite-1 mutant animals, as described, fail to avoid' high concentrations of diacetyl, but isn't it more accurate to say that the valence of the response is changed from avoidance to attraction? Do the authors believe that attraction is mediated by ODR-10? Can you build an odr-10; lite-1 double mutant and determine if they lose this attraction to high concentrations of diacetyl and/or 2,3 pentanedione?

      Our initial interpretation was that, in the absence of LITE-1-mediated avoidance, attraction to high concentrations of diacetyl is driven by the low-concentration receptor ODR-10. Interestingly, however, the lite-1; odr-10 double mutants remained strongly attracted to high concentrations of diacetyl, suggesting that this phenotype is independent of ODR-10 and may instead be mediated by other, less specific odorant receptors. This is consistence with the odr-10 mutants not completely losing attraction to low concentration of diacetyl, suggesting the involvement of other potential/putative receptors (Sengupta et al., 1996; Taniguchi et al., 2015). The lite-1; odr-10 double mutant data is now added in the results section (lines: 82 to 90; Figure S1B).

      (2) For ADL, yes, it seems like lite-1 mutants have a reduced diacetyl response, but the ASK response seems... different. While it goes up (slowly) in wild-type, cellular Ca2+ levels in ASK (and maybe ADL) are *reduced* by diacetyl in lite-1 mutants. Can the authors comment on this, and the behavioral responses change in valence?

      ADL and ASK are involved in both attractive and aversive responses so it is possible that the attractive component of the diacetyl response suppresses their activity. In the absence of LITE-1 activation, the observed decrease in calcium responses may reflect the unopposed inhibitory input. The slower decay of the calcium signal in ASK neurons of unc-13 mutants further supports the presence of additional inhibitory signals influencing their activity. This explanation is now included in the results section (lines: 101 to 117).

      (3) ADL and ASK calcium traces in unc-13 mutants look generally similar and lack the effects seen in lite-1 mutants. Does that mean the lite-1 effect is in cells other than ADL or ASK? Can the authors spend more time discussing these differences?

      We acknowledge that LITE-1 is expressed in multiple cell types beyond chemosensory neurons, and that non-chemosensory neurons may also contribute to the observed phenotype. In the previous version, we had highlighted the interneuron AVG as a potential contributor, given its role in light-induced escape. In the revised discussion, we have now expanded this section to include additional possible contributors such as the LITE-1-expressing phasmid neuron PHA and the pharyngeal interneurons I2, both of which have been implicated in hydrogen peroxide sensing (lines: 177 to 181).

      (4) Chemotaxis responses to diacetyl and 2,3-pentanedione in lite-1 mutants are rather different. Diacetyl switches from repulsive (CI < 0) to *attractive* (CI > 0), which is not what would be expected for mutations that eliminate a receptor. In contrast, the 2,3-butanedione responses are more what would be predicted: diacetyl goes from inhibitor to no effect (CI ~0). Again, if the authors feel that this is because of the ODR-10 function, can they discuss whether 2,3-butanedione is predicted to bind ODR-10 like diacetyl?

      We think a switch to attraction is expected for the removal of a receptor for an aversive signal, in an attractive background signal. We did indeed think that this attraction was mediated by odr-10, but the double mutant results now show other receptors must be involved. This is also not wholly unexpected since previous work has identified other receptors of high-concentration diacetyl and the original odr-10 paper didn’t report a complete absence of diacetyl response, suggesting the presence of other receptors that mediate attraction to diacetyl (Sengupta et al., 1996; Taniguchi et al., 2014). The new data is now added in the results section (lines: 82 to 90; Figure S1B; Supplementary video 1 and 2).

      (5) Were the behavior experiments performed in the dark? I realize the calcium imaging experiments and some of the video behavior recordings are not possible in complete darkness, but maybe the authors made efforts to exclude visible light effects (e.g., infrared illumination, etc.) in some assays that might help determine whether light plays *no* role in the effects observed. Alternatively, the authors could try repeating their chemotaxis experiments in the dark or at least communicate in the methods that this was not viewed as a concern (and why). As the authors propose and discuss LITE-1 modulating diacetyl responses via light sensation as a possibility, it is incumbent upon them to communicate the steps they took to overcome this concern for themselves.

      We acknowledge this and have performed a chemotaxis experiment to compare assays performed under dark and ambient light conditions, and no significant differences were observed (results section: lines 75 to 80; Figure S1A, material and methods section: 241 to 246). Therefore, subsequent chemotaxis assays were carried out under ambient light while avoiding exposure to strong illumination.

      Paralysis assays were repeated under red-filtered illumination to minimise light effects, with animals maintained in darkness from hatching to adulthood. Additionally, the assay was expanded to include lite-1 mutants, ruling out contributions from neuronal LITE-1 to paralysis. The new data is now incorporated into the result section (diacetyl; lines: 135 to 137, figures 4 and S3; 2,3-pentanedione; lines: 152 to 154, figures 4 and S6), materials and methods section have been updated to reflect these changes (lines: 250 to 253 and 257 to 260).

      Reviewer #2 (Recommendations for the authors):

      Minor Issues:

      (1) Pmyo-3::LITE-1 worms shrink in the absence of odorants (Figures 3C, 4D); possible effects of ambient light should be discussed.

      We acknowledge the possibility that worms expressing LITE-1 in body-wall muscle experience minor contractions under ambient light, though it is not sufficient to cause paralysis.

      To minimise potential light-induced effects, the paralysis assays were repeated with red-filtered illumination to reduce light stimulation of LITE-1. Animals were maintained in darkness from hatching to adulthood, and in the updated assay, worm length remained relatively constant.

      The new data (Figures 3C, 4E, S4C and S6C) and materials and methods section has been updated accordingly (lines: 250 to 253 and 257 to 260).

      (2) The title is misleading, as ASH does not show altered activity in lite-1 mutants and should be removed from the claim.

      ASH has been removed from the title (line: 97).

      (3) Specific Kd values should be provided for the reported micromolar binding affinities.

      The values from DynamicBind and Gnina are provided in Figure S5.

      Recommendations:

      (1) LITE-1, a member of the gustatory receptor family, was previously shown to mediate UV light responses in C. elegans. In this study, Koh and colleagues demonstrate that LITE-1 is also required for the nematode's avoidance of high concentrations of diacetyl - an odorant that is attractive at low levels but aversive at higher concentrations. Using calcium imaging, the authors show that LITE-1 is necessary in the sensory neurons ADL and ASK for calcium transients in response to high concentrations of diacetyl. Additionally, they find that expressing LITE-1 in body-wall muscles causes hypercontraction upon diacetyl exposure. Similar LITE-1-dependent responses were observed for 2,3-pentanedione, another structurally related odorant. Molecular docking analyses suggest that both diacetyl and 2,3-pentanedione directly bind to LITE-1 with micromolar affinity.

      These findings are intriguing and have the potential to significantly advance our understanding of LITE-1 as a multimodal sensory receptor. However, several major issues need to be addressed to support the authors' conclusions:

      (1) Rescue experiments are missing. The authors should rescue at least one lite-1 mutant to confirm that the observed avoidance defects are specifically due to loss of lite-1.

      Because the avoidance defect was observed in three independent lite-1 alleles, we think background mutations are unlikely to be causal. We have now also performed a rescue experiment with lite-1 expressed in ADL neurons. In this strain, attraction is restored. The new data have been incorporated into the results section (lines: 119 to 121; Fig. 2D).

      (2) Alternative loss-of-function approach. To strengthen the findings, the authors should use a different method to disrupt lite-1 function-such as RNAi by feeding or cell-specific RNAi driven by the lite-1 promoter (see PMID: 17459615).

      We believe the multiple alleles (Fig. 1B) and new cell-specific rescue experiment (Fig. 2D) provide sufficient support for the conclusion that the phenotype is due to loss of lite-1 function.

      (3) Clarify the role of ADL and ASK neurons. While calcium imaging data show reduced activity in these neurons in lite-1 mutants, it remains unclear whether lite-1 is required in ADL, ASK, or both for avoidance behavior. Cell-specific rescue or RNAi experiments, along with additional calcium imaging, are needed to determine the contribution of each neuron.

      We agree that more in-depth work will be required to dissect the neuronal pathways involved in LITE-1-mediated diacetyl avoidance. We have avoided making specific comments on how exactly the observed imaging results relate to the behavioural phenotype. The new rescue experiment expressing lite-1 in ADL does at least show that LITE-1 in ADL is sufficient for avoidance, although it’s not a complete rescue to wild-type levels (lines: 119 to 121; Fig. 2D).

      (4) Validation of molecular docking results. While molecular docking suggests direct binding of odorants to LITE-1, experimental validation is needed. Mutations that reduce predicted binding affinity (engineered in transgenes or via CRISPR) should be tested for functional impact on avoidance behavior.

      We agree the docking does not in itself establish direct binding (we think the muscle expression and paralysis provides much stronger evidence). Based on previously reported docking experiments, we were simply curious whether diacetyl would be predicted to occupy the same binding pocket. We have now updated the discussion in the use of lite-1 mutants to test for impact on diacetyl avoidance (lines: 188 to 191).

      (5) Include analysis of non-binding odorants. Docking results should also be presented for odorants that did not elicit LITE-1-dependent avoidance, to help establish specificity.

      We have now included docking data of odorants from the chemotaxis assays, with the corresponding docking values shown in Supplementary Figure 5. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      (6) Figure 2C concerns. In neurons such as AWA, calcium transients in unc-13 mutants appear reduced compared to wild-type. A statistical comparison for all the neurons should be included to assess significance.

      Scatter plots of calcium imaging responses across the different sensory neurons were generated, and statistical significance was assessed using two-sided t-tests with FDR correction. Only ADL and ASK neurons showed significant differences between lite-1 mutants and wild-type animals. No significant differences were observed in any neuronal pairs between unc-13 mutants and wild-type, including AWA neurons, although the difference is close to significant (p = 0.07). The scatter plots were now included as Supplementary Figure 3.

      Minor points:

      (1) Figures 3C and 4D: Pmyo-3::LITE-1 worms appear to shrink even without diacetyl or 2,3-pentanedione. Could this be due to ambient light? The authors should discuss this possibility.

      We acknowledge the possibility that worms expressing LITE-1 in body-wall muscle experience minor contractions under ambient light, though it is not sufficient to cause paralysis. We repeated the paralysis assays with red-filtered illumination to reduce light stimulation of LITE-1. Animals were maintained in darkness from hatching to adulthood, to minimise potential light-induced effects, and in the updated assay, worm length remained relatively constant.

      The new data (Figures 3C, 4E, S4C and S6C) and materials and methods section has been updated accordingly (lines: 250 to 253 and 257 to 260).

      (2) Title revision needed: The title "Chemosensory neurons ADL, ASK, and ASH are involved in avoidance of diacetyl" is misleading, as calcium transients in ASH appear unaffected in lite-1 mutants. The title should reflect the actual data.

      ASH have been removed from the title (line: 97).

      (3) Binding affinity clarification: The authors report micromolar binding affinity for LITE-1 but should provide specific dissociation constants for clarity and completeness.

      The values from DynamicBind and Gnina are now provided in Supplementary Figure 5.

      Reviewer #3 (Recommendations for the authors):

      How did the authors measure body length if the animals were swimming in the diacetyl solution? Standard 6-well plates have an area of roughly 10 cm², meaning that if one adds 1 ml, the liquid level should be 1 mm. The animals would be able to move in 3D, so it is likely that animals swim up and down and do not move in one flat plane, i.e., head and tail would be out of focus, and only a projection image would be recorded that would lead to an underestimation of actual worm length.

      We acknowledge that this issue may led to an underestimation of worm length in the previous assay. However, every effort was made to exclude worms that moved out of focus. The paralysis assay has since been modified to include spreading a thin layer of solution across the worms, which keeps them mostly in focus, particularly those that are paralysed. Worms that were partially out of focus were excluded from the analysis.

      We have incorporated the new data into the results section, reflected in the updated Figures 3C, 4E, S4C, and S6C. Corresponding revisions have also been made in the materials and methods (lines: 250 to 253 and 257 to 260).

      Could diacetyl be a compound that results from UV absorption in cells? This may be worth discussing. What could be the precursor molecule?

      We are not aware of such a precursor, but we cannot rule it out. Even if UV absorption in cells leads to diacetyl production, it is likely that the resulting diacetyl levels are insufficient to activate LITE-1, as our data suggest that LITE-1 functions as a receptor for high concentrations of diacetyl. We have updated the discussion accordingly (lines: 169 to 173).

      In lines 57-61, the references to Edwards 2008 and Ward 2008 do not seem to fit the statements made in this sentence.

      The inclusion of Edwards et al., 2008 was an error, and it has now been removed. Ward et al., 2008 demonstrated that ASJ phototransduction requires cGMP and CNG channels, stating that “Our studies indicate that C. elegans photoreceptor cells also employ CNG channels and the second messenger cGMP for phototransduction.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper examines whether humans use protracted temporal integration in a noise-free, deferred-response contrast discrimination task, using a covert evidence-duration manipulation combined with EEG (SSVEP, CPP, Mu/Beta). The key finding is that evidence for protracted sampling is behaviorally and neurally supported, but even joint CPP + behaviour fitting cannot fully discriminate a standard integration (DDM) model from a novel "extremum-flagging" non-integration model. The paper is transparent about this outcome.

      Strengths:

      This is a well-conducted and well-written study that makes a genuine contribution to the perceptual decision-making literature by introducing a clean experimental design for probing temporal integration without participants adapting their strategy and demonstrating for the first time that a non-integration model (extremum-flagging) can replicate CPP waveform dynamics that have long been considered hallmarks of evidence accumulation. The transparent treatment of equivocal modelling outcomes is commendable.

      Weaknesses:

      My main concerns relate to statistical power, the under-specification of the and the extremum-flagging mechanism. Addressing these would greatly strengthen the paper.

      (1) The sample of 16 participants (15, after the exclusion of one participant) is described as "close to similar EEG studies" with no formal power analysis. Given that the paper's core claim rests on subtle quantitative differences between two model classes - differences that are, by the authors' own admission, not sufficient to declare a winner - even a modest increase in sample size might yield a more decisive outcome. At a minimum, the authors should report a sensitivity analysis or post-hoc power calculation to indicate what effect sizes the current N could reliably detect, particularly for the rmANOVA comparisons and the neural constraint fitting.

      We appreciate the reviewer’s concern regarding sample size and statistical sensitivity. To address statistical robustness throughout the paper, we have now reported effect sizes for our statistical tests (e.g. η2 for rmANOVA; including Tables S1 and S2), and we provide error-shading around the ERP waveforms to indicate the reliability of the key patterns our models are aimed at capturing (i.e. the dramatically higher and earlier CPP peak for high-contrast, and very little systematic differences across the four low-contrast durations - see revised Figure 3). We also conducted an indicative post-hoc power analysis using the G*Power software based on the behavioural data. Using the observed partial η2 = 0.44 for the test of duration effect on accuracy among only the low-contrast conditions, and the final sample of 15 participants, this amounts to a statistical power of 0.998.

      On the model comparison, while we agree that larger sample sizes are generally beneficial for population-level inferences, we respectfully maintain that our current sample size is sufficient to support the core claim that qualitative dynamics of neural signatures of decision formation, usually assumed to reflect temporal integration, can be successfully reproduced using non-integration models in the delayed-response task conditions we examine here. Any marginal changes in quantitative fit resulting from having a higher N contribute to grand averages are unlikely to substantively alter this conclusion of the model comparison. The statistical reliability of the data to which our models are fitted is also bolstered by the number of trials (about 256 per condition per participant). It is common in behavioural modelling studies for data to be collected from a much smaller sample (e.g., fewer than 10 subjects) but with a high trial yield - a relevant precedent for us being Stine et al., (2020), who provided a compelling demonstration of similar model fits for Extrema and Integration models using only 6 subjects. In sum, the key qualitative data patterns of accuracy improvements with duration and broader, lower and duration-invariant low-contrast CPPs are statistically robust and provide a strong basis to reveal the fundamental principle that the Extremum-flagging and Integration models are both able to produce these key qualitative dynamics.

      (2) The Extremum-flagging model is the paper's most novel contribution, yet its physiological basis is underspecified. The model posits that each decision-terminating bound-crossing triggers a stereotyped, half-sine-shaped centroparietal signal, but no neural circuit or computational mechanism is proposed for how the brain could detect the first bound-crossing event in a non-accumulating evidence stream or generate a temporally precise, fixed-amplitude signal in response. Possible connections to P3b theories of context updating and response facilitation are acknowledged, but these are vague functional descriptions rather than mechanistic accounts. I think the discussion should engage more directly with potential neural substrates that could generate this flagging signal, and whether these are consistent with the known generators of the CPP/P3b. Without this, the extremum-flagging model risks being viewed as a mathematical convenience rather than a biologically plausible alternative.

      We thank the reviewer for this constructive comment. While the focus of this paper was indeed on simple mathematical descriptions in the spirit of classical cognitive modelling, we agree that expanding on the potential neural substrates of the Extremum-flagging model strengthens its utility as an alternative framework. We have revised the Discussion to engage with potential biological mechanisms, particularly those previously proposed to underlie the P300/P3b, such as Nieuwenhuis’ (2005) proposal that it reflects a phasic arousal response mediated by the LC/NE system that serves to activate task-relevant areas following completion of a decision. We agree that aside from this, accounts of ERP component functions over the years have often been vague and non-mechanistic, but the idea that they reflect discrete neural activations marking an internal cognitive event in a stereotyped way persists, and remains a basic assumption of several new and influential ERP signal analysis toolboxes (e.g. Ehinger 2019; Weindel 2024). If the flagging signal’s fixed amplitude seems physiologically implausible, all-or-nothing neural activation events are not generally unheard of in neurophysiology, and, again, we are taking an approach favouring parsimony in the spirit of cognitive modelling, and we found that we did not need to assume any variation in the amplitude of the flagging signal in order to capture the key decision signal dynamics alongside behavioural accuracies in this particular case.

      We also discuss the study of Latimer et al. (2015), who demonstrated that discrete, step-function state transitions that on single trials may mark extrema detection events, can produce ramp-like signals when trial-averaged. While the biological plausibility of such step-function dynamics remains a subject of debate, it serves as another example of how continuous evidence integration is not the only way to reproduce the ramping neural signals traditionally observed in grand-average neural signals.

      (3) The Integration model at the preferred neural weighting estimates a high-to-low contrast drift rate ratio of 8.7, whereas the empirical Mu/Beta lateralization slopes suggest a ratio of approximately 3.5. The authors attribute this discrepancy to the nonlinear contrast response function of early visual cortex and the salience of the high-contrast evidence onset, but these explanations are speculative. These outcomes are arguably the most quantitatively damaging result for the integration model, so they deserve more than a brief discussion. I would recommend that the authors (a) estimate what range of contrast response nonlinearities would be required to close this gap, (b) test whether an alternative drift rate parameterization (e.g., scaling drift rates directly by SSVEP amplitude rather than contrast) reduces the discrepancy, or (c) be more explicit about treating this as a point against the Integration account.

      We agree that the quantitative discrepancy we demonstrated between the empirically observed buildup rate ratio in motor preparation signals (3.5) and the greater drift-rate ratio (8.7) required by the Integration model to fit the CPP waveforms is an important one that should be emphasised and discussed with greater depth and clarity. As we said, nonlinear contrast response functions and a boosting effect of the salient high-contrast step-change are two plausible ways that a drift rate might scale disproportionately more steeply with contrast, but in principle, assuming straightforward transmission of evidence accumulation to the motor level, Mu/Beta lateralization slopes should then reflect this steeper drift rate scaling, or at least approach it even when allowing for some temporal blurring. We have thus put more emphasis on the discrepancy by confirming that if we constrain the drift rates to be directly proportional to contrast, the Integration model is indeed significantly hampered in its ability to produce the much steeper CPP buildup for higher-contrast trials, much more so than the Extremum-flagging model (Figure 4 - Supplement 7). We have also applied a temporal blurring equivalent to the short-time Fourier Transform to the simulated motor preparation waveforms (convolving with a boxcar of the same duration as the Fourier window) in Figure 4N-P so that the real and simulated traces are on an equal footing in this respect. We have also revised the Discussion to elaborate on how this quantitative discrepancy represents a point against the Integration account, and possible ways it might be reconciled with an Integration account. One reason, for example, why the relative steepness of the Centroparietal ERP in the high-contrast condition so far exceeds that of Mu/Beta might be that additional processes are evoked by the very salient step-change, which may make a positive-polarity contribution to the centroparietal ERP waveform and hence cause overestimation of how early and steeply the underlying, high-contrast CPP decision signal rises. We looked into this by carefully examining time courses and topographies through the initial period of buildup, with no additional smoothing low-pass filter applied, now presented in Figure 3 - Supplementary Figure 1. While the smoothed waveforms that we show in the main paper and to which we fit models could be seen to have a brief inflection during the main buildup for the high-contrast condition, removing the smoothing shows that this arises not from random noise but from a distinct bimodal morphology, with a distinct early peak and lull during the buildup, which temporally coincides with a very strong bilateral occipital N2 (associated with a low-level evidence-onset detection or ‘target selection’ process - Loughnane et al 2016), in a way that suggests that the positive tail-end of the dipolar neural generators of the N2 may contribute to the initial part of the positive centro-parietal buildup. It is difficult to estimate the extent to which the neurally-constrained model estimate of high-contrast drift rate is inflated by this initial overlapping potential, because we can’t precisely know the ground truth of the N2 tail’s contribution, but this analysis provides a potential explanation that can be explored in future (e.g. through softening the strong-evidence onset with a ramp or use of auditory evidence). We thank the reviewer for raising this as we feel that this extra discussion positively adds to the theme of the paper to highlight methodological challenges with neurally-constrained modelling. In the process, we have updated the methods section to present in full detail the centroparietal electrode selection and waveform smoothing that was applied to provide the models with a relatively uninterrupted buildup signal to capture, which is important for readers to appraise the potential impact of this overlapping potential.

      (4) The sensitivity analysis over neural constraint weightings (w = 0.1 to 1000) is thoughtful, but the paper ultimately acknowledges that the preferred weighting is w=10, chosen because it achieves "a good fit to CPP dynamics without substantively sacrificing behavioral fit" - a qualitative criterion. No principled statistical framework is used to select the optimal weighting or to compare models at a given weighting. A Bayesian model comparison could provide a more formal framework for combining behavioral and neural fit components, and would allow a clearer statement about the relative posterior probability of each model.

      We agree with the reviewer that theoretically, the Bayesian framework provides a principled way to combine behavioural and neural evidence by weighting each source according to its statistical reliability. However, a Bayesian formulation typically quantifies reliability through across-trial variance, which applies quite differently for accuracy and EEG data. While the precision of EEG measurements can be estimated empirically (e.g., from noise characteristics), we currently lack a formal measure of uncertainty for the linking function itself, that is, the theoretical mapping between neural signatures and latent decision processes. This represents an unresolved methodological issue rather than a straightforward parameter estimation problem. Second, although hierarchical Bayesian approaches are well established for standard diffusion models, the mechanisms examined here for extremum flagging do not currently have tractable closed-form formulations suitable for Bayesian integration. Developing a dedicated hierarchical Bayesian framework for these non-standard mechanisms would require substantial methodological work and is beyond the scope of this research.

      Thus, rather than imposing a single assumed reliability relationship between neural and behavioural data, we chose to perform a systematic sweep across weighting values. We view this approach as a transparent sensitivity analysis that accommodates different scientific priors regarding the relative contribution of neural versus behavioural constraints. By presenting the full range of w (including in supplemental tables and figures), readers can directly evaluate how model behaviour changes when emphasis is shifted between behavioural data and neural data, transparently revealing how the behavioural and neural signal fits can trade against one another.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Hajimohammadi, Mohr, O'Connell and Kelly is intended to demonstrate that participants integrate evidence over time to make a decision, even in a noise-free, static decision context. This is validated by the observation that (1) participant accuracy improves with increased exposure to the stimulus; and (2) there is a correlation between participant accuracy and a neural index of evidence accumulation, as measured by centro-parietal positivity (CPP).

      Strengths:

      (1) Joint modelling of accuracy and CPP dynamics is a significant achievement, as behaviour alone often cannot distinguish between competing theories of decision-making. In the case of protracted sampling in particular, the absence of reaction times (RT) due to the delayed nature of the response makes this method highly appealing.

      (2) The experimental manipulations and the method used to extract the different neural indices are well chosen, enabling the mapping of putative cognitive processes such as evidence accumulation and motor preparation onto the recorded EEG with clarity.

      (3) The in-depth discussion of the results clearly articulates those reported by the authors and in previous works.

      Weaknesses:

      (1) One main issue to support the interpretation of the authors toward the need for protracted sampling is the timing of the evidence. By design, participants believe that the signal is present for 1.6 seconds (reinforced by the fact that easy trials were displayed for 1.6 seconds). However, the difference in stimuli is turned off either 1.4, 1.2, 0.8 or 0 seconds before the cue to respond. While this makes sense in the context of the authors' question, it also raises the possibility that participants will focus on the last samples before answering. Even if participants apply equal weighting, this still favours them delaying evidence accumulation until they are sufficiently certain that the evidence should be present (e.g. participants might start accumulating after the stimulus has disappeared in the 0.2 condition). I do not see an easy way to test these alternative explanations outside of running a study in which the evidence is always offset before the go cue.

      This is a reasonable question about the design - if participants were under the impression that they had a whole 1.6 sec of stimulation, couldn’t they afford to wait until later into the stimulus to start sampling? However, the task was designed to be so difficult that participants would be deterred from ignoring any initial evidence, and the fixed and explicitly instructed lead-in period as well as the interleaved easy trials, would have continually reinforced their ability to time their sampling onset quite precisely. Indeed, key aspects of the data confirm they did not appreciably delay sampling. First, accuracy in even the shortest (0.2 s) condition was reliably above chance (t(15) = 2.60, p = 0.0201) and improved steadily across evidence durations (Figure 1B). This places an upper bound on the accumulation onset: participants cannot have delayed accumulation until after the evidence disappeared and still achieve above-chance performance; if they only used the ‘last samples,’ at the end of the stimulus, they would have performed at chance level for all durations except 1.6 sec. Second, we fit a model that allowed for such a delayed sampling onset, captured in the parameter ‘sampT,’ which, across the range of neural weightings (Tables S3, S5-8), consistently landed within a few tens of msec of evidence onset (often slightly before rather than delayed), and improved the overall model fit very little relative to the addition of starting point variabilities or collapsing bound. The Methods section now addresses these aspects of task design.

      (2) Regarding the behavioural models, are these identifiable based on accuracy data alone? This should be addressed using a parameter recovery study, in which a set of parameters is used to generate data, and the same fitting routine used for the real data is used to estimate the parameters. This would enable us to determine what can be inferred from the model comparison presented. This is not a serious problem for the manuscript, as it specifically aims to go beyond behaviour. It is, however, worth noting that such a parameter recovery addition could be used to demonstrate the need for a joint modelling framework to answer the question of protracted sampling on delayed response times (RT).

      As the reviewer notes, we did have the specific aim of going beyond behaviour, and the need to do so is demonstrated in the inability to adjudicate between the alternative models based on behaviour alone. We took this as sufficient justification without a formal parameter recovery test to assess the degree to which behaviour-only models could accurately estimate parameter values. Still, we agree that it is valuable to address parameter identifiability in some way. Since a full parameter recovery covering the full possible parameter space for each of the many models would be too great in volume to add to this paper, we can instead address identifiability somewhat indirectly through parameter estimate consistency across the 10 fits we conducted with different instantiations of noise; we now provide the standard deviations alongside the mean parameter values for the D1, D2 and B parameters of each of the behaviour-only models in Table 1 - Table Supplement 1, which indicates that the parameter estimates were reliable across 10 different instantiations. 

      Minor comments:

      (1) I would advise authors to fix the D1 parameter and use it as a scaling parameter across all models. Currently, as I understand it, the models are scale-free, meaning the same fit is achieved by multiplying all parameters by two, for example. This makes the fit more complex (bounds on parameter values are required) and means that the models are less comparable in terms of their estimates. Perhaps I'm missing something, but I would have thought that fixing D1 (the common parameter across all models) would solve these issues.

      The models are not scale-free because they are constrained relative to a fixed sampling noise parameter value of s = 0.1; All tables in the main text have now been updated to make this more immediately clear. Aside from this being standard in diffusion modelling (Ratcliff & Smith, 2004), this enabled us to replicate the observation by Stine et al., (2020) that since the non-integration models depend on the magnitude of individual evidence samples rather than an integration of many, the drift rate values must be set much higher to achieve the same choice accuracy as the integration models (Table 1).

      (2) Why is the snapshot model so bad despite being a good model in Stine et al 2020? Can the authors speculate in the discussion?

      We thank the reviewer for querying this. We had originally thought that the poor performance of the snapshot model made sense because the continued presentation of zero contrast difference for short-evidence trials renders it a bad strategy. Because our main purpose was to briefly substantiate the principle that accuracies alone are an insufficient basis for model comparison and move on to the main goal of jointly modelling accuracies and CPP dynamics, we did not take the same level of care to ensure we attained the very best fit of the behaviour-only models, as we did for the neurally-constrained models. In the neurally-constrained modeling, we took care to check for every parameter whether the range of allowed values (Table S4) was narrow enough to avoid the optimisation algorithm getting lost in untenable parts of parameter space, yet wide enough to include the optimum point, and wherever we saw parameter values landing at or near the edge of the allowed range we expanded that range and re-ran the model fit. Applying these same checks to the behaviour-only fitting, we found that the SnapShot model needed a wider range on drift rate and when we applied this, the fit was much more competitive, in line with Stine et al., (2020), though it remained the worst-fitting model among all two-drift-rate behaviour-only models (see updated Table 1). We similarly conducted these checks across all behaviour-only models and re-ran them. The extrema detection model with last-sample default when no bound is hit also improved its fit, though again it did not fit better than the version with guess default. Thus, the point we were making with this section, that behaviour alone can be captured competitively by a range of integration and non-integration models, is bolstered by the updated model fits. Since the last-sample default was competitive in the behaviour-only fits, we also ran a version of the Extremum-flagging model jointly fit to accuracies and CPP dynamics with a last-sample rather than random guess default when a bound was not reached, and show in new Figure 4 - Figure Supplement 8 that the conclusions are the same. Again, thank you for prompting us to look back at those fits.

      (3) The meaning of the flag width is unclear. Figure 4 provides the reader with an intuitive understanding of the model that the authors have in mind. However, the tables in the appendices report values between 0.2 and 0.9. I understand that these values represent the width of the half-sine in seconds. This suggests that the actual estimated values for these flag events are much broader than those displayed in Figure 4. While this is probably fine for most models, it can be problematic for the extremum-flagging model, as it means that the rise to the peak takes between 0.1 and 0.45 seconds. While strictly speaking, this is still a 'flag' model, such a slow rise to the peak, given the usual expectation of evidence accumulation, would place this model closer to a smooth integration model than to a boundary-crossing flagging mechanism.

      We thank the reviewer for raising this about the flag width parameter. In so doing, they enabled us to catch that our schematic depiction of the model in Figure 4 was misleading, and have now revised it to make clear that the flag signal is a post-decision one triggered by the bound crossing, and we have updated explanations accordingly (in ‘Neurally-constrained models’ and Discussion). The reviewer is correct that the reported values in the supplemental materials (approximately 0.2–0.9 s) correspond to the width of the half-sine kernel used to model the post-decision flag event. However, the flag signal is stereotyped, evidence-independent, and is triggered once the decision threshold has already been crossed, so it does not share the key characteristics of evidence integration, regardless of how wide the model estimates it. In the extremum-flagging model, the boundary crossing remains a discrete event. The width parameter instead captures the temporal extent of the neural process that follows this commitment event. Such a post-decision neural process unfolding over several hundred milliseconds is in line with some classic theories of the centroparietal P300/P3b component, and we now expand our discussion point on this to address proposed neural substrates (e.g. Nieuwenhuis et al’s (2005) implication of a phasic noradrenaline system response).

      (4) In the modelling section, it is not clear overall (i.e. for G<sup>2</sup> and R<sup>2</sup>) how the participant dimension is taken into account. Are these individually fitted models, and if so, how are the secondary statistics generated from the individual estimates? Or were these fitted over all participants?

      All models were fitted to the grand-average neural and behavioural data across participants, rather than to individual participant data. We chose this approach as the CPP signal at the individual level is highly noisy, which can introduce substantial instability and noise into the model fitting procedure. We have revised the Modelling section to explicitly state that the reported G<sup>2</sup> and R<sup>2</sup> values are derived from models fitted to the grand-average data, and in the revised discussion acknowledged this as a limitation of the current modelling framework.

      (5) On page 7, in the last sentence of the first paragraph of the section titled 'Decision-Related Neural Signals', the authors state that 'this stable contrast-difference encoding suggests that a constant (i.e. non-adapting) drift rate is a reasonable simplifying model assumption'. However, I am not sure how this is true given that SSVEP quantifies encoding, yet the drift rate can vary due to non-sensory aspects (e.g. attention).

      The reviewer makes a good point - even if sensory encoding is stable, non-sensory factors like attention could cause dynamic changes in the effective drift rate independently of the sensory representation itself. However, our point in that section, which we have revised to put more clearly, was to test for one particular well-known time-varying effect that could impact drift rate, namely sensory adaptation, a well-established phenomenon behaviorally and at the level of sensory neuronal responses, where prolonged stimulation produces reductions over time in sensory neural activity. If strong adaptation were present in the sensory evidence representation indexed by the SSVEP, we would expect corresponding temporal changes in the signal. The absence of such changes lends support to the simplifying assumption (as in most accumulation models) that the drift rate is approximately stationary over time, even if we cannot be sure there isn’t a time-varying effect downstream.

      (6) The mu/beta lateralisation does indeed favor the integration model more, but in terms of boundary estimation and starting-point analyses, both models are pretty far apart. Providing an interpretation of this observation, e.g. regarding alternative linking functions for mu/beta, would add to the manuscript.

      In response to this comment, we revised the manuscript in the Discussion to say that in the current analyses, we implicitly assume an approximately linear mapping between Mu/Beta amplitude and decision units. However, the true relationship may instead reflect another monotonic transformation (e.g., involving power rather than amplitude, logarithmic scaling such as dB units, or a nonlinear saturating function). This uncertainty could affect the apparent correspondence between the neural signal and the model-derived estimates of boundary position or urgency dynamics. While our analyses support a close relationship between Mu/Beta lateralisation and the evolving decision process, the precise quantitative mapping remains uncertain. One possibility is that urgency itself evolves nonlinearly (e.g., decelerating over time), even if the measured neural trajectory appears approximately linear under the current transformation assumptions.

      Reviewer #3 (Public review):

      Summary:

      The authors aim to compare proposal models of perceptual decision making using a joint modeling approach, where they fit models to both behavioral outcomes as well as CPP. Most notably, they compare a standard evidence accumulation model with models that track the evidence without integrating it over time (extrema detection). The authors report that the joint CPP-behavioral data do not discriminate between two of their proposals.

      Strengths:

      This is an interesting finding that reinforces the idea that what we believe to see based on aggregation over trials may not be what happens on every single trial. The models are creative, and the simulations are convincing, relating the models to multiple neural markers of decision formation. These include the CPP but also mu/beta power spectra.

      Weaknesses:

      The paper makes some strong points, and the work seems generally well-executed. The weaknesses that I identified are twofold:

      (1) Embedding in the literature/exposition of the main argument.

      The focus in the introduction is on the noise-free nature of the stimulus and the prolonged presentation time. However, after reading the paper, I felt these were mostly experimental design choices that enable comparison of the different models using the CPP. Perhaps my misreading of the goals of the paper stems from two other observations:

      (a) The fact that the stimulus is noise-free does not entail that perception is noise-free. Thus, the argument that using a noise-free stimulus precludes the necessity of temporal integration seems not completely valid. Of course, one could argue that noise is limited in this case, but that makes a noise-free stimulus more of a design choice.

      (b) The focus on prolonged stimulus presentation, but at the same time the contrast with expanded judgement, did not make sense to me. Perhaps, as a non-native speaker, I am misreading the subtle difference between "protracted sampling" and "longer sampling", but again, the longer duration seems mostly a design choice.

      We thank the reviewer for this impression, which has helped us revise the introduction to more clearly motivate the paradigm as an interesting case for close examination. The primary driver of our choice of stimulus and task parameters was not to enable model comparison using the CPP; it was to examine a decision scenario that exists in everyday life but that has not been examined in terms of underlying decision mechanisms because it offers only sparse behavioural data - the scenario in which plainly visible objects (without noise or stochasticity, as in daylight conditions) need to be examined for a subtle feature difference to guide a later action. The reviewer echoes our point in the Intro, that despite the absence of physical noise in the stimulus, perceptual processing itself is not noise-free. Therefore, temporal integration is certainly not precluded, but its benefit is minimised and less obvious to the decision maker. Given the examples we raise where integration was found to not be employed to its optimal extent (e.g. bound setting foregoing accuracy improvements with duration), and the various theoretical accounts citing energy costs associated with integration and the fleeting nature of many natural environments where prolonged deliberation about a static stimulus is not the norm (e.g. Uchida et al 2006), it is quite hard to guess a priori whether humans will engage in protracted sampling and integration in this case, in practice, even if it is optimal under basic assumptions. As we make clear in our revised Intro, this theoretical interest in the uncertain case of long, noise-free stimuli where perfect, unbounded integration may be optimal but seems doubtful given extant empirical findings, is coupled with a methodological interest in the extent to which neural signatures of decision formation can ‘come to the rescue’ and provide grounds for reliable adjudication between competing mathematical models, when behavioural data fall short.

      More could be said about the optimality of the extrema detection methods. In particular, decades of work (centuries?) have shown that evidence integration is an optimal decision-making procedure: For example, the Sequential Probability Ratio Test is Bayes-optimal wrt mean RT (Wald, 1946); evidence accumulation together with collapsing threshold serves to maximize rewards in repeated choices (e.g., Bogacz et al., PsychRev, 2006; Boehm et al. APP, 2020). Given all this work, why would the brain have evolved to adopt a different mechanism? I realize that the paper is not about optimal decision making, but some discussion of this point seems warranted.

      We had a similar impression initially when reading Stine et al., (2020) where extrema-detection was pitted against integration - is extrema detection so suboptimal that it is too implausible to even consider? Ditterich (2006) argued that signal-to-noise ratio would have to be implausibly high for extrema-detection to produce the behaviour observed on typical decision tasks. However, the fact is, we do not know the effective signal-to-noise ratio, nor can we precisely quantify the costs associated with prolonged evidence accumulation, such as attentional or energetic costs (Drugowitsch et. al., 2012). Even if the extrema detection strategy appears implausibly suboptimal, it is an important principle to demonstrate how not only behavioural but also neural decision signal dynamics can be so nicely consistent with integration yet technically can be quantitatively captured with non-integration mechanisms.

      (2) Modeling choices.

      The authors introduce a parameter, sampT, that represents uncertainty in the sampling onset time. It was not clear to me whether this parameter represented an offset of all trials, or a distribution (probably the latter). I wonder how exactly this parameter was integrated into the models, and in particular, if and how it interacts with the starting-point parameters. My intuition is that on a single-trial, IF early sampling occurs, you can model that with either a negative sampT and z at 0, or with sampT at 0 but a shift in z. This would suggest trade-offs between these parameters, making them hard to estimate independently. Since the paper does not depend on the identification of parameter estimates, this may not be a huge problem, but nevertheless it is good to explore the consequences.

      We thank the reviewer for raising an important question regarding the relationship between sampT and starting-point variability (sz). Mechanistically, early accumulation onset can indeed generate effects that resemble starting-point variability: if accumulation begins during a period containing only zero-mean noise, then by the time informative evidence appears, the decision variable will already have diffused away randomly from zero. In this sense, negative sampT can induce variability in the state of the accumulator at evidence onset. However, the two mechanisms are not mathematically equivalent. The sz parameter assumes a uniform distribution over starting points, whereas the variability induced by early accumulation onset would instead reflect the distribution resulting from integrating zero-mean Gaussian noise over variable durations. Aside from this distinction between distribution shapes, the reviewer is correct that these parameters could partially trade off with one another when sampT takes negative values. In our model, however, sampT was allowed to take either positive or negative values. Positive values delay the onset of evidence integration relative to the evidence, thereby ignoring the first samples, very different from the effect of starting point variability. Nevertheless, to the extent that they can partially trade off each other to some degree, the consequent problem this might cause to accurately estimating both parameters is part of the reason we do not fit a model that includes both simultaneously.

      The way the Bounded Integration model (BIntg) is formulated seems very close to the EZ-diffusion model (Wagenmakers et al., PBR, 2007). This model states that the proportion of correct responses Pc = 1/(1+exp(-B*D/s^2), with B and D the bound and drift rate parameters, respectively. However, filling in the numbers for the high contrast condition from Table 2, and assuming that s=2 (because the model description states that dt=2, with s undefined), I get a Pc of 80% for the 1.6H condition. This seems substantially less than what Figure 2 suggests.

      As we had stated in the Methods section, the model used “a standard deviation of 0.1 arbitrary units” for the evidence. We now make it more explicitly clear that this corresponds to setting within-trial noise s = 0.1 as the scaling parameter (the first paragraph of ‘Model Fits to behaviour only’ and the first paragraph of ‘Integration models’ in Methods). Replacing (s = 0.1) in the suggested calculation yields a predicted accuracy close to 1 for the high-contrast condition, consistent with both the behavioural data and the model predictions shown in Figure 2. We have now clarified this parameter explicitly in the revised Methods and updated Table 1 and Table 2 to avoid confusion.

      On some occasions, it is unclear to me what modeling choices are being made:

      (a) It seems as if the models are fit on accuracy data alone (before introducing the neural data). This seems suboptimal given that the authors do report differences in RT.

      Because of the delayed-report feature of the task, RTs do not directly reflect decision termination time, which is the basis of the use of RT in cognitive modelling normally. Here, whether the decision process has concluded during the stimulus or not, indeterminate response-cue detection and motor execution processes intervene between the stimulus and RT, which would necessitate complicating the models with additional mechanisms, of which there are several possibilities as reflected in the response to the reviewer’s final comment about the RT effects below. This could potentially obscure the core mechanisms of decision formation during the stimulus itself, which was the focus of the study.

      (b) Are the models fit on all data combined, or on the data of individual participants? Fitting individual participant data is preferred, as combined or aggregated data may be distorted by individual differences.

      Because of the noise and variability of EEG data at the single-participant level, we model data averaged across participants, which we have ensured is clear in the revised paper. We provide individual accuracy trends in Figure 1, to verify that the accuracy improvements with increasing evidence duration seen on average are representative of the vast majority of individual subjects. We also added a comment on the limitation this incurs regarding individual difference analysis in the revised discussion.

      (c) The authors seem to suggest that the diffusion coefficient s is estimated (in the section "Integration models"). Most likely, however, this is set to a fixed value. Obviously, it matters for the model comparison using AIC whether this parameter was freely estimated or not.

      As noted in our response to an earlier Comment, the diffusion coefficient was fixed at s=0.1, and to make this explicit, we have entered it for all models in the revised Table 1 and Table 2.

      Not really a weakness, but I wondered about the effect of stimulus duration on RT. In particular, what hypothesis (or post hoc explanation) do the authors have for these RT effects? I could think of at least three hypotheses that are consistent with the behavioral data:

      (a) H1: The shorter the evidence duration, the more likely participants are to require a double-check before response execution, reflecting their uncertainty about their decision.

      (b) H2: There is a collapsing threshold that initiates at stimulus offset, leading to quicker responses on trials where there is more evidence.

      (c) H3: motor preparation is correlated with the evidence signal, which leads to faster responses on trials with more evidence.

      We thank the reviewer for these hypotheses. We agree that the RT effects admit multiple possible interpretations, and while we are cautious not to overinterpret them mechanistically in the paper given our focus on the decision process during the stimulus preceding these response-cue-triggered responses, we do take them to signify that the decision process has not always fully completed and been transformed to a finalised action plan by the time of response cue (start of Results section). To consider these interesting possibilities further:

      We agree that the longer RTs for shorter duration, more uncertain trials could reflect a “double-checking” process (H1), but it could alternatively reflect the fact that if a bound has already been reached during the stimulus, this commitment can be translated to fully-selected action plan that only needs triggering, whereas if a bound has not been reached by stimulus offset, which would occur more often for shorter evidence-duration trials, more of the motor action-selection process would yet need to be completed to initiate the action, causing the slight delay in RT. In other words, on longer-duration trials, the accumulated evidence is more likely to have already reached the bound before the response cue, allowing participants to both commit to a choice and prepare the associated motor response in advance. Therefore, RTs would be shorter as the remaining processes after the cue primarily involve cue detection and motor execution.

      This interpretation is broadly compatible with the reviewer’s H3 account, in the sense that motor preparation may track the evolving decision variable/evidence state. It is also possible that collapsing bounds are set on a post-stimulus, cue-evoked process (H2), which is not mutually exclusive with the above possibilities. It would be hard to determine whether such a process is primarily a response cue-detection decision process that is modulated by uncertainty state at stimulus offset, or a cue-triggered “double-check” process that perhaps operates on the iconic memory of the evidence, and our paradigm does not allow these alternatives to be cleanly dissociated.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you can see, the reviewers are positive about the work and highlight several strengths, while at the same time offering recommendations for improvement. Once these points are satisfactorily addressed, this may also lead to a revision of the eLife assessment below.

      Reviewer #1 (Recommendations for the authors):

      As outlined in my public review, my most important recommendations relate to sample size and possible neural bases of the extremum-flagging model.

      Minor Comments:

      (1) p. 4 rmANOVA statistic is listed as 12.35.31.

      This has been corrected, with thanks for spotting it.

      (2) The stable d-SSVEP amplitude during the evidence period is used to justify a constant (non-adapting) drift rate assumption. This is a reasonable inference, but the SSVEP reflects early sensory encoding rather than the decision variable per se. Neural adaptation or gain changes at later processing stages could still produce a non-constant effective drift rate even with a stable sensory representation. This inference should be qualified.

      This is true. We have clarified that these checks for one important potential source of a time-varying drift rate, namely adaptation at the level of early sensory representation, but admit that other effects may happen downstream.

      (3) The Methods describe a leaky accumulation extension; the Results note leak ≈ 0.0002 at w=10, which is effectively zero. This is a positive result (evidence against leaky integration) that should be stated more explicitly in the Results or Discussion rather than appearing only in supplementary tables.

      We thank the reviewer for highlighting this. We have pointed to this result now in the first paragraph of the Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) The panels in Figure 4 H, I, J are not discussed in the Neurally-constrained models Section, while I believe they are probably more informative than Figure 4E alone.

      The manuscript has been revised to explain Figure 4H,I, J fully under ‘Neurally-constrained models’.

      (2) In the method section, we don't know how many trials were rejected based on the chosen threshold.

      The Method section is updated with the rejected trials after preprocessing. It now reads, “After preprocessing, 12017 trials remained across all conditions, with an average rejection rate of 15% (± 12.9%) across participants.”

      (3) Figure 1: For b and c, data are mean {plus minus} s.e.m. after between-participant variance was factored out. -> reference or detailed method.

      We have now explained in the caption that this is done by subtracting the overall mean of each individual from their data and adding back the grand mean, retaining the between-condition differences - that is, we remove the component of variance that repeated-measures tests ignore.

      (4) Figure 2: Shouldn't the snapshot model only feature one sample, as in Stine et al. 2020? This figure and others would also benefit from a better resolution.

      The reviewer is correct regarding the schematic of the snapshot model presented in Stine et al., (2020). We used small, light orange dots to represent the evidence samples presumably being encoded and a larger orange dot to indicate the single randomly chosen sample used as evidence for the decision. In our revised figure, we have increased the visual distinction between the evidence dots and the selected dot, and pointed this out in the caption. We have also improved the figure resolutions.

      (5) Typo:

      - Semi-saturation Table S4.

      - pi missing in the text of the G^2 equation.

      Thank you for catching these.

      Reviewer #3 (Recommendations for the authors):

      Small, random points:

      (1) How was it ensured that participants indeed did not detect the change in contrast throughout the 1.6s interval? In previous work (Winkel et al., PBR, 2014) we did something similar in a random-dot motion task, but observed that participants always observed the change, if we did not slowly change the coherence of the stimulus (unfortunately, Winkel et al., 2014 is not explicit about the exact parameters of the change, but Figure 2 suggests that the change in coherence lasted 50ms, independent of stimulus strength).

      It is true that abrupt changes in stimulus strength are more salient and detectable than a ramped change. The experimenters tested the stimulus during the task design subjectively, to satisfy themselves that they could not tell when the contrast stepped back to baseline, but this was not verified systematically with psychometrics, nor can we be sure that a more sensitive observer couldn’t sometimes detect the change. However, based on the task design, stimulus properties, and both behavioural and neural data, we are confident that participants are very unlikely to have been sufficiently confident in detecting the step-back in contrast to ceasing their contrast-comparison decision process at that point:

      (1) Participants were naive to the underlying manipulation. They were informed that trials would naturally vary in difficulty. It was normal for them to perceive some trials as harder than others without suspecting a mid-trial structural change.

      (2) The contrast difference in the hard condition was very subtle, and though it is possible that the step-down in contrast could be detected with above-chance accuracy if instructed to do so, given there was no instruction on whether and when the step-down would happen, it is very unlikely they could be detected with sufficient confidence to be certain there is no remaining evidence in the stimulus. Furthermore, the rapid, flickering nature of the stimulus would have helped to mask the transition point, in comparison with a sudden change in a continuously-playing random-dot motion stimulus (as in Winkel et al., 2014).

      (3) Our behavioural post-cue RT data imply the participants did not cease decision formation at evidence offset. If they had, then they would have been afforded the most time to prepare their chosen action in advance of the response cue in the case of the earlier evidence offset (i.e. shorter durations), yet these were the conditions with the longest, not the shortest post-cue RTs.

      (4) The low-contrast CPP traces remained elevated for the full 1600 ms interval, unperturbed by the evidence offsets. If participants were explicitly detecting a sudden change in contrast, we would expect to see a transient evoked response locked to that change, marking that detection. Instead, the sustained elevation of the CPP is characteristic of a continuation of the decision process, uninterrupted, through the subtle offsets.

      (2) Could you include the regression coefficients of the statistical modeling of the behavioral data?

      The regression coefficient is now added to the second paragraph of the Results section.

      (3) I felt Figure 5B was a bit confusing: Are the dashed lines here the ipsilateral sides or the non-linear bounds? This was confusing because "data" only has a solid line in the legend.

      We agree that the subtle nonlinearity, which appears visually close to the linear model, may have caused this confusion. To improve clarity, we have added arrows to Figure 5B to explicitly indicate which legend refers to which panel, and distinguish the dashed nonlinear bounds from the other traces.

      (4) "amplitude variations [...] to be used as an independent evaluation of model fit": Could you refer to where these model predictions are presented? I think this is in Figure 4 - Sup 5?

      We thank the reviewer for raising this. The model-predicted waveforms showing amplitude variations across durations for all neural weightings are presented in Figure 4 - Supplementary Figures 2 and 5. We have clarified this in the revised manuscript.

      References

      Ditterich, J. (2006). Evidence for time‐variant decision making. European Journal of Neuroscience, 24(12), 3628–3641. https://doi.org/10.1111/j.1460-9568.2006.05221.x

      Drugowitsch, J., Moreno-Bote, R., Churchland, A. K., Shadlen, M. N., & Pouget, A. (2012). The cost of accumulating evidence in perceptual decision making. Journal of Neuroscience, 32(11), 3612–3628.

      Ehinger, B. V., & Dimigen, O. (2019). Unfold: An integrated toolbox for overlap correction, non-linear modeling, and regression-based EEG analysis. PeerJ, 7, e7838.

      Latimer, K. W., Yates, J. L., Meister, M. L. R., Huk, A. C., & Pillow, J. W. (2015). Single-trial spike trains in parietal cortex reveal discrete steps during decision-making. Science, 349(6244), 184–187. https://doi.org/10.1126/science.aaa4056

      Loughnane, G. M., Newman, D. P., Bellgrove, M. A., Lalor, E. C., Kelly, S. P., & O’Connell, R. G. (2016). Target selection signals influence perceptual decisions by modulating the onset and rate of evidence accumulation. Current Biology, 26(4), 496–502.

      Nieuwenhuis, S., Aston-Jones, G., & Cohen, J. D. (2005). Decision making, the P3, and the locus coeruleus–norepinephrine system. Psychological Bulletin, 131(4), 510.

      Stine, G. M., Zylberberg, A., Ditterich, J., & Shadlen, M. N. (2020). Differentiating between integration and non-integration strategies in perceptual decision making. Elife, 9, e55365.

      Uchida, N., Kepecs, A., & Mainen, Z. F. (2006). Seeing at a glance, smelling in a whiff: Rapid forms of perceptual decision making. Nature Reviews Neuroscience, 7(6), 485–491.

      Weindel, G., van Maanen, L., & Borst, J. P. (2024). Trial-by-trial detection of cognitive events in neural time-series. Imaging Neuroscience, 2, imag–2.

      Winkel, J., Keuken, M. C., Van Maanen, L., Wagenmakers, E.-J., & Forstmann, B. U. (2014). Early evidence affects later decisions: Why evidence accumulation is required to explain response time data. Psychonomic Bulletin & Review, 21(3), 777–784. https://doi.org/10.3758/s13423-013-0551-8

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers and editors for their time and valuable input into improving our manuscript. The reviewers recognized the value of this work while identifying places where further explanation and/or additional experiments could strengthen the manuscript. We greatly appreciate this feedback and have addressed reviewer comments through additional experiments and necessary textual edits.

      In response to the reviewer comments, our principal data-driven changes to revise the manuscript include: 1) testing how stabilizing HIF-1 in the ADF and NSM neurons affects healthspan; 2) measuring for interactions between HIF-1 stabilization in serotonergic neurons and the mt-UPR; and 3) measuring nlp-17 expression downstream of HIF-1 stabilization in serotonergic neurons. We also made textual changes including: 1) clarifying that our study focuses specifically on the vhl-1-mediated genetic hypoxic response; 2) summarizing which parts of the working model are experimentally validated vs. speculative; and 3) expanding our discussion of future directions needed to understand the epistasis of the many signals acting in this circuit. Our responses to each reviewer comment are below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study by Kitto et al., the authors set out to identify specific signaling components regulating the hypoxic response from the neurons to the periphery and which components are required for lifespan extension. Their previous work had shown that expression of a stabilized HIF-1 mutant in the nervous system extends lifespan through the serotonin receptor SER-7 and leads to the induction of fmo-2 in the intestine. In the current study, they mapped the precise neural circuits required for this response, as well as the signaling mediators. Their work reveals that neurotransmitters GABA and tyramine, and the neuropeptide NLP-17, act downstream of neuronal HIF-1 to convey a "hypoxic signal" to peripheral tissues. Through cell-type-specific expression studies, targeted knockouts, and comprehensive lifespan analysis, the authors provide robust evidence to support their conclusions. The insights gained from the study are both moving the field forward as they advance our understanding of neuro-peripheral hypoxic signaling, but they also lay the groundwork for potential therapeutic strategies aimed at the modulation of such signaling pathways.

      We appreciate the reviewer’s positive assessment of the topic and general interest in this work.

      Strengths:

      (1) This study provides new evidence further delineating signaling components required for hypoxic signaling-mediated longevity, from the nervous system to the periphery. Using a rigorous approach where they express stabilized HIF-1 mutant selectively in ADF, NSM, and HSN serotonergic neurons, followed by cell-type-specific tph-1 knockouts to pinpoint ADF-dependent serotonin signaling as essential for both lifespan extension and intestinal fmo-2 induction.

      This was followed by generating 11 transgenic lines that drive SER-7 expression under distinct neuron-specific promoters, to systematically tease out in which of 27 candidate neurons SER-7 functions to mediate hypoxia-induced longevity. This ultimately highlighted the RIS interneuron as the required signaling hub.

      (2) As the intestine lacks direct neuronal innervation, the authors employ neuron-specific RNAi (TU3311 strain) and dense core vesicle analyses to identify that the neuropeptide NLP-17 is required to transmit the hypoxic signal from RIS to induce fmo-2 in the intestine.

      (3) Overall, the paper is very well written. The experiments were carried out carefully and thoroughly, and the conclusions drawn are also well supported by the results they are showing.

      Weaknesses:

      Overall, I don't see many weaknesses. One point relates to their read-outs, which rely heavily on lifespan measurements and fmo-2 induction without evaluating other physiological processes that serotonin or NLP-17 might affect. For translational relevance, it would be valuable to assess or mention potential adverse effects, such as changes in reproduction, pharyngeal pumping, or proteostasis capacity (proteostasis capacity specifically in the tissue showing fmo-2 upregulation).

      We thank the reviewer for the positive review and fully agree and acknowledge that the primary readouts used in this work, fmo-2 induction and lifespan, may not reflect other important elements of health and physiology. To address this weakness, we have performed three measurements of healthspan (pumping, thrashing, and maximum velocity), in the ADF and NSM HIF-1 stabilized strains at young adulthood and at middle age. We focused on examining these HIF-1 stabilized strains rather than our nlp-17 knockout animals because stabilizing HIF-1 in the NSM or ADF neurons is sufficient to extend lifespan. nlp-17 is necessary, but it remains unclear whether nlp-17 signaling is sufficient to extend lifespan. Our new data show that both ADF- and NSM-specific HIF-1 stabilization had no effect on pumping rate in young (day 1 of adulthood) worms. In aged animals (day 12 of adulthood), however, NSM-, but not ADF-, specific HIF-1 stabilization rescued the pumping rate decline in hif-1 knockout compared to WT worms (new Fig. S1C). Similarly in the new thrashing data, NSM-, but not ADF-, specific HIF-1 stabilization rescued the thrashing rate decline in the hif-1 knockout young and aged worms (new Fig. S1D). Lastly, our new data show that ADF and NSM HIF-1 stabilization had no effect on average or maximum movement speed at days 1 and 5 of adulthood (new Fig. S1E-F). Together, these results indicate that genetic activation of the hypoxic response in NSM neurons but not in the ADF neurons could improve healthspan.

      While we did not examine reproductive capacity in these strains, we have expanded our discussion section to mention the importance of fully characterizing other elements of physiology, health, and behavior in future work.

      “Another limitation of this work is that it uses lifespan as the main readout for organismal health. While genetic manipulations that extend lifespan often improve stress resistance and healthspan [8,80], longevity manipulations can also have adverse effects on reproduction [81,82] and behavior [83,84]. In this study, we find that HIF-1 stabilization in the ADF neurons does not prevent the deleterious effects of hif-1 knockout on mobility, but that HIF-1 stabilization in the NSM neurons may attenuate age-related decline in pumping and thrashing (Fig. S1). However, future work should examine whether other modifications to this pathway, such as manipulations to RIM, RIS, or NLP-17 signaling, influence healthspan in addition to lifespan. It will be important for future studies to determine whether various components of this pathway affect both longevity and the response to different types of stressors like oxidative stress, proteotoxic stress, and infection, as HIF-1 activity also interacts with multiple stress responses [41,42,39,40,43].”

      While lifespan assays and fmo-2 expression do provide strong evidence, incorporating additional markers of stress resistance could strengthen the link between hypoxic signaling and organismal health as well.

      We also measured hsp-6 expression via qPCR in the serotonergic neuron-specific HIF-1 stabilized strains to determine whether these conditions that lead to upregulated fmo-2 may also affect the mt-UPR. Interestingly, we find that stabilizing HIF-1 in either the ADF or NSM serotonergic neurons decreases hsp-6 expression relative to WT worms (new Fig. S1G). This could suggest either that the mt-UPR response is impaired in these worms, or that HIF-1 stabilization decreases proteotoxic stress leading to a lower basal level of hsp-6. Although this method of measurement did not allow us to interrogate whether these changes occur in the specific tissues where fmo-2 is upregulated, we have expanded our discussion to emphasize that further investigation of cell- and tissue-specificity within the hypoxic response should be a focus of future work.

      Finally, we agree that it is important to examine whether activating the hypoxic response promotes stress resistance in addition to longevity. We did not focus on this element of the hypoxic response in this work because HIF activity is known to promote adaptive stress-responses to some stressors like infection [7,8] and oxidative stress [9,10], while simultaneously impairing the proteosasis stress response [11]. We have added this important information to our introduction section, and have expanded our discussion section to emphasize that a key future direction will be to test which specific components of this pathway also facilitate stress resistance.

      Introduction Section Modification:

      “However, the physiological changes induced by the hypoxic response are broad and involve adaptations such as increased vascularization, metabolic rewiring, and changes in cell survival pathways. In mammals, some of these same adaptations can be detrimental, as mutations in components of the hypoxic response have been linked to conditions like cancer and cardiovascular disease [12-15]. Additionally, HIF activity is essential to promote some forms of stress resistance to infection [9,10] and oxidative stressors [7,8] but can have detrimental effects on proteostasis [11].”

      Discussion Section Modification:

      “Another limitation of this work is that it uses lifespan as the main readout for organismal health. While genetic manipulations that extend lifespan often improve stress resistance and healthspan [1,2], longevity manipulations can also have adverse effects on reproduction [3,4] and behavior [5,6]. In this study, we find that HIF-1 stabilization in the ADF or NSM neurons has little effect on mobility in young and aged animals. However, future work should examine whether other modifications to this pathway, such as manipulations to RIM, RIS, or NLP-17 signaling, influence healthspan in addition to lifespan. It will be important for future studies to determine whether various components of this pathway affect both longevity and the response to different types of stressors like oxidative stress, proteotoxic stress, and infection, as HIF-1 activity also interacts with multiple stress responses [7,8,9-11].”

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to identify the specific neurons, neurotransmitters, and neuropeptides that mediate the longevity effects of the hypoxic response in C. elegans. By genetically dissecting the pathway downstream of HIF-1, they define a neural circuit involving ADF serotonergic neurons, the SER-7 receptor in the RIS interneuron, tyraminergic signaling from RIM, and neuropeptide NLP-17, ultimately linking neuronal hypoxic sensing to pro-longevity signaling in the intestine.

      Strengths:

      The study employs a diverse genetic toolkit, including neuron-specific transgenes, tissue-specific knockouts and rescues, RNAi knockdowns, allowing the authors to pinpoint causality, sufficiency, and necessity with high resolution. The comprehensive mapping of cell-nonautonomous signaling adds depth to our understanding of how HIF and serotonin signaling interface with aging pathways. The conclusions are supported by consistent survival assays and fmo-2 gene expression analyses.

      Weaknesses:

      A key limitation is the lack of clear evidence showing epistasis of so many identified molecular/neuronal components downstream of HIF-1 and serotonin. Thus, the mechanisms of how a diverse set of molecules/neurons coordinate and mediate neuronal HIF-1 effects on intestinal fmo-2 and longevity remain murky.

      We thank the reviewer for these important points. We agree that the epistatic relationships between ADF serotonin, RIM tyramine, RIS GABA, and neuropeptide NLP-17 signaling remain unclear within this pathway. Determining the epistatic relationships of each signal within this complex pathway will require: 1) generating genetic manipulations to each identified signaling component that may mimic vhl-1 knockout to promote longevity; 2) crossing these new strains into multiple genetic knockouts we identified as required for vhl-1 mediated longevity; and 3) measuring the lifespans of each double and triple mutant. We are very interested in testing the epistatic relationships between all of these molecules and neurons, and believe this extensive follow-up exploration will generate significant future results.

      To better address these limitations of the current work, we have added a summary table to Fig. 7 as well as a paragraph to our discussion section. This table and paragraph better explain which components of our working model have been tested for necessity, sufficiency, and epistasis, and which components of this model remain unclear (Fig. 7). See updated Fig. 7 with added table.

      Updated discussion section detailing this limitation:

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1 depletion, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant. This approach would also narrow down which signals are downstream of the genetic activation of the hypoxic response, and which are sufficient to extend lifespan upstream of the hypoxic response in a normoxic environment. One notable target for further exploration is the SER-7 expressing RIS neuron, which plays a role in sleep [16] and stress resistance [17], and can extend lifespan when optogenetically activated under normoxic conditions [18].”

      Some rescue strategies may inadvertently cause non-physiological expression.

      This is a great point. We bring attention to this limitation in the discussion section. Our cell-specific knockouts (ADF tph-1 KO, Fig. 1E) or ablations (RIS ablation, Fig. 2C) data showed that the neurons identified using rescue strains (i.e., tph-1 in the ADF and ser-7 in RIS) are likely not false positives. However, we did not generate a RIM-specific knockout or ablation strain to confirm our tdc-1 results and have added this limitation to the discussion of these results.

      Discussion of rescue strategy limitations and the need for a RIM-specific knockout:

      “The circuit-mapping approaches employed in this work are also impacted by limitations in cell-specific genetic modifications and in the use of RNAi knockdown. For example, cell-specific rescue constructs can sometimes lead to unintended rescues in other cell types due to cell-nonautonomous signaling. Because all serotonin-producing neurons also express the serotonin reuptake transporter mod-5, serotonin produced by one cell in our rescue strains could be taken up by other serotonin-producing neurons, leading to unintended signaling effects. This may also be true of the uv1 and RIM tyraminergic rescue strains, although little is known about tyramine reuptake in C. elegans. Two cells identified via tissue-specific rescue experiments (ADF and RIS, Fig. 1G and Fig. 2D) were also found to be necessary via cell-specific knockout (ADF tph-1 KO, Fig. 1E; RIS ablation, Fig. 2C), decreasing the likelihood of a false positive from the rescue strain technique. However, the role of the RIM neuron was identified via a tdc-1 rescue strain and was not validated using a RIM-specific knockout (Fig. 7B). Therefore, the contribution of RIM signaling to this circuit is less well-validated, and a RIM-specific tdc-1 knockout strain should be examined in future work.”

      Additionally, environmental hypoxia was not tested in parallel, so the claim on "hypoxia response" throughout the manuscript is not justified by genetic manipulation alone, and the translational relevance of the genetic manipulations remains somewhat uncertain.

      We thank the reviewer for identifying the need for additional specificity in our language. We agree that it is critical to make it clear that this paper only examines genetic mimics of hypoxia, rather than environmental hypoxia, and have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (e.g., “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We discuss this in Discussion and are excited to interrogate these differences in future work.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      Reviewer #3 (Public review):

      Summary:

      This study found that ADF serotonergic neurons have a significant role in extending lifespan mediated by HIF-1, as well as serotonin receptor SER-7 in the GABAergic RIS interneurons. The author focuses on the sufficiency and necessity of components from the central nervous system and how they contribute to aging upon hypoxia.

      Previous work from the lab has identified that the stabilization of HIF-1 in neurons is sufficient to extend lifespan through the serotonin receptor, SER-7, which subsequently activates fmo-2 in the intestine and leads to lifespan extension. Building on this, the author sought to determine which serotonergic neurons are involved and found that serotonin signaling in ADF neurons is required for lifespan extension mediated by HIF-1.

      The author next tested which subset of neurons requires Ser-7 expression to rescue hypoxic response. They found that ser-7 expression in multiple neurons is sufficient to induce fmo-2, with the top candidate being the RIS neuron. Ablation of the RIS neuron did not extend lifespan, suggesting that ser-7 expression in the RIS neuron is required for lifespan extension, positioning it as a key component in the longevity signaling pathway.

      The author also investigated neurotransmitters and found that GABA and tyramine are important components in this circuit. They showed that the tyramine receptor called tyra-3 is required for vhl-1-mediated longevity. Given that tyra-3 is expressed in oxygen- and carbon dioxide-sensing neurons, the author demonstrated that these sensing neurons work downstream of serotonin signaling. Lastly, the author screened neuropeptide/receptor binding pairs and identified NLP-17 as playing a role in hypoxia-mediated longevity.

      Originality and Significance:

      This research is significant in that it uncovers components that are sufficient and necessary for lifespan extension via the hypoxic response. It provides comprehensive data supporting longevity induced by HIF-1-mediated hypoxic response, in conjunction with fmo-2, a longevity gene, as demonstrated in previous work from the lab. Moreover, it provides a number of new transgenic worm tools for C. elegans and aging communities.

      We thank the reviewer for the positive assessment of the manuscript. We appreciate all suggestions for further improving the work and have made changes based on these suggestions (see details below).

      Data and Methodology:

      (1) The experiments were thoroughly conducted, especially the generations of strains using different neuron-type promoters and crossing into mutant strains to demonstrate sufficiency and necessity.

      (2) Some figure legends from the text do not match what the data show. (Figure 6E, F, G).

      We have made changes to the legends accordingly to make sure figures and legends are consistent.

      (3) The lifespan graph legends are confusing and could use some revamping for better clarification.

      We have updated the lifespan graph legends to provide the statistics in the same section as the labels, indicating which line is which condition (see updated Figure 1E). We hope this change better clarifies the lifespan graph legends.

      Conclusions:

      This study provides insights into how hypoxic response regulates aging in a cell non-autonomous manner, outlining a potential circuit involving neurons, neurotransmitters, and neuropeptides.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As suggested by the Reviewers 2 & 3, including environmental hypoxia will broaden the impact of the study. If not, the authors should consider clarifying their response as "genetic activation of the hypoxic response" (see Reviewer 2 below).

      We thank the editor and reviewers 2 and 3 for this clarification. We agree that it is critical to make it clear that this paper only examines genetic mimics of hypoxia, rather than environmental hypoxia, and have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (e.g., “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We have also expanded our discussion of this important future direction.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      Suggested minor changes by Reviewer 3 should also be made to improve the readability of the paper.

      We have made changes suggested by Reviewer 3 to improve the readability of the paper.

      Reviewer #2 (Recommendations for the authors):

      (1) Suggestions for additional experiments or analyses:

      (a) To clarify the hierarchical relationships among the components of the identified circuit, epistasis experiments between serotonin, tyramine, NLP-17, and oxygen-sensing neurons (e.g., double mutants or sequential rescues) would strengthen the proposed model and help determine whether these signals act in parallel or downstream of each other.

      We thank the reviewer for identifying this important caveat. As described in the public review response, we completely agree that understanding the epistasis of serotonin, tyramine, NLP-17, and oxygen-sensing neuron signaling within this pathway is important to fully test our working model. We hope to address these questions about epistasis and interactions between different signals in upcoming projects.

      To clarify that the order of many of these signals remains speculative in our working model, we have added a table to Fig. 7, that summarizes what is known and unknown about the epistasis of these pathway components. We have also expanded our discussion of this limitation. See updated Fig. 7 with added table.

      Updated discussion section detailing this limitation:

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1 depletion, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron we found to be required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant.”

      (b) Testing whether environmental hypoxia (e.g., 0.5-1% O₂ exposure) elicits similar neuronal requirements and fmo-2 induction as the genetic HIF-1 stabilization would validate that the described pathway is relevant to the actual hypoxic response and improve translational relevance.

      We thank the reviewer for identifying the need for clarification. As also discussed in the public review section, we agree that it is important to emphasize that this paper exclusively examines genetic mimetics of hypoxia, rather than actual exposure to a hypoxic environment. To clarify this point, we have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (ie “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We discuss this in discussion and are excited to interrogate these differences in future work.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      (c) Functional readouts beyond lifespan and fmo-2 induction (e.g., neuronal activity monitoring or optogenetic modulation of key neurons) could help clarify how information flows through the circuit.

      We thank the reviewer for this valuable suggestion. We are hoping to be able to implement these experimental techniques in future work. We agree that in combination with genetic epistasis analyses, direct measurements of neuronal signaling will greatly improve our understanding of the directionality and interactions between signals in this circuit. Our discussion section recommends employing these techniques in future studies.

      “Finally, while the use of RNAi knockdown and genetic knockouts establishes the necessity of many signals within the vhl-1-mediated longevity circuit, the exact directionality of these signals remains unclear. It is possible that increased, decreased, or pulsatile changes in signaling through these bioamines and neuropeptides are required for genetic activation of the hypoxic response to extend lifespan. Work on C. elegans reversal behavior has also revealed an antagonistic relationship between RIM and RIS activity facilitated by both chemical (neuropeptide and tyramine) and electrical (gap junction) signaling [16,21]. This known interaction should also be interrogated in the context of how these cells may communicate following genetic induction of the hypoxic response. Future work in this area could use tools to measure or modify neuronal activity, such as calcium imaging or optogenetics, to begin answering these questions.”

      (2) Recommendations for improving the writing and presentation:

      (a) The manuscript is well written overall, but clarity would be improved by explicitly stating in the abstract and introduction that the study is based on genetic activation of the hypoxic response, rather than environmental hypoxia.

      We appreciate this valuable suggestion and have updated the abstract and introduction to clarify that this work examines genetic activation of the hypoxic response.

      Updated sentences from the abstract:

      “Here, we interrogate the cell-nonautonomous signaling pathway downstream of genetic activation of the hypoxic response.” 

      “Together, these insights develop a circuit for how genetic induction of the hypoxic response cell-nonautonomously modulates ageing and suggests valuable targets for modulating ageing in mammals.”

      Updated sentences from the introduction:

      “In this study, we uncover key neural components of the longevity circuit initiated by genetic induction of the hypoxic response. Within this circuit, we identify individual cells, signals, and receptors necessary and/or sufficient to extend lifespan downstream of genetic activation of the hypoxic response. More specifically, we find serotonin signaling in the ADF serotonergic neurons is both necessary and sufficient to extend lifespan through genetic activation of the hypoxic response. This pathway signals through the serotonin receptor SER-7 in the RIS interneuron. We further demonstrate additional neurotransmitters (GABA and tyramine), and a neuropeptide (NLP-17) are critical for mediating these longevity effects. Finally, we identify that oxygen sensing neurons (URX, AQR, PQR and BAG) act downstream of neuronal HIF-1 in this circuit. Our insights into this longevity pathway provide a mechanistic understanding of how genetic activation of the hypoxic response delays aging and improves health.”

      (b) In the discussion, clearly delineating which parts of the proposed pathway are firmly established versus inferred would aid interpretation.

      We thank the reviewer for this idea, and have added the following text to the discussion:

      “Evidence for the necessity, sufficiency, and epistatic relationships between each signal are summarized in new Fig. 7B. In brief, all signaling molecules presented in this work are necessary for vhl-1 to extend lifespan. Rescuing ADF serotonin production, RIS ser-7 expression, and RIM tyramine production is sufficient for vhl-1 to extend lifespan. The sufficiency of oxygen sensing neurons BAG and UPA/PQR/AQR, and the neuronal signals of GABA, NLP-17, and TYRA-3 to restore vhl-1 mediated longevity remains unclear. All signals act downstream of vhl-1. ADF HIF-1 stabilization and SER-7 signaling act upstream of fmo-2 induction, and the oxygen sensing neurons (BAG, UPA/PQR/AQR) act upstream of or in parallel to ADF HIF-1 stabilization.”

      (c) Adding a summary table or schematic that visually distinguishes necessity vs. sufficiency for each component (e.g., ADF, RIS, RIM, NLP-17) would make the overall model more accessible.

      We thank the reviewer for this great suggestion and have added a table to new Fig. 7B that summarizes what is known about necessity for vhl-1, sufficiency for vhl-1, and sufficiency to extend lifespan independently of vhl-1 for each signal (table included in response to Reviewer 2, comment 1a).

      (3) Minor corrections and clarifications:

      (a) Define or replace "hypoxic response" with "genetically induced hypoxic response" where appropriate to avoid conflating genetic manipulations with actual environmental hypoxia.

      We appreciate this valuable suggestion and have replaced “the hypoxic response” and with “genetic activation of the hypoxic response” or “genetically induced hypoxic response” throughout the manuscript to clarify this point.

      (b) All genes and alleles should be italicized per worm nomenclature.

      We thank the reviewer for this comment and have reviewed the manuscript to italicize all gene names and alleles. In some locations, the protein is referred to instead of the gene using the conventional uppercase non-italicized format.

      Reviewer #3 (Recommendations for the authors):

      (1) Using vhl-1 RNAi as the sole approach to demonstrate hypoxic response appears somewhat limited, as vhl-1 is also involved in HIF-1 independent processes that can influence lifespan in C. elegans. Including additional downstream effectors of HIF-1, such as egl-9, or having HIF-1 nondegradable strain as validation could strengthen the findings.

      We thank the reviewer for this valuable comment and agree that an important next step is to test whether these signals are also required for other genetic (HIF-1 stabilized, egl-9) and environmental activators of the hypoxic response to extend lifespan. We have worked to clarify that this paper focuses primarily on vhl-1 mediated longevity throughout the text and have also included this limitation in our discussion section (for details, please see response to Reviewer 2 public review).

      (2) Investigating how the healthspan is affected by serotonergic neuron-specific hypoxic responses would be interesting and could enhance understanding of the physiological mechanisms underlying lifespan extension. 

      We appreciate this suggestion, and have performed three measurements of healthspan (pumping, thrashing, and maximum velocity), in the ADF and NSM HIF-1 stabilized strains at young adulthood and at middle age. We found that both ADF- and NSM-specific HIF-1 stabilization had no effect on pumping rate in young (day 1 of adulthood) worms. In aged animals (day 12 of adulthood), however, NSM-, but not ADF-, specific HIF-1 stabilization rescued the pumping rate decline in hif-1 knockout compared to WT worms (new Fig. S1C). Similarly, NSM-, but not ADF-, specific HIF-1 stabilization rescued the thrashing rate decline in the hif-1 knockout young and aged worms (new Fig. S1D). ADF and NSM HIF-1 stabilization also had no effect on average or maximum movement speed at days 1 and 5 of adulthood (new Fig. S1E-F). Together, these results indicate that genetic activation of the hypoxic response in NSM neurons but not in the ADF neurons could improve healthspan.

      (3) While the experiments were thoroughly performed, the connections between components such as NLP-17, GABA, and tyramine in regulating aging appear critical for establishing a "cell non-autonomous circuit." Additionally, how the potentially antagonistic roles of RIM and RIS neurons influence this axis could be interesting to further explore.

      We thank the reviewer for identifying this important caveat. As described in the public review response to Reviewer 2, we completely agree that understanding the epistasis of serotonin, tyramine, NLP-17, and oxygen-sensing neuron signaling within this pathway is important to fully test our working model. Our current data showed that intestinal fmo-2 is required for neuronal HIF-1 stabilization to extend lifespan, indicating information must be communicated between the nervous system and the intestine through serotonin, tyramine, NLP-17 and responsible neurons using a “cell non-autonomous circuit” [22]. However, we will continue to address questions about epistasis and interactions between different signals in upcoming projects to fully establish the circuit.

      With respect to RIM and RIS, we value this suggestion and agree that there could be interesting signaling occurring between RIM and RIS in this circuit, as is observed in initiation of reversal behaviors [16,21]. We have updated the discussion section to mention this interesting antagonistic relationship between RIM and RIS signaling in the context of reversal behaviors:

      “Finally, while the use of RNAi knockdown and genetic knockouts establishes the necessity of many signals within the vhl-1-mediated longevity circuit, the exact directionality of these signals remains unclear. It is possible that increased, decreased, or pulsatile changes in signaling through these bioamines and neuropeptides are required for genetic activation of the hypoxic response to extend lifespan. Work on C. elegans reversal behavior has also revealed an antagonistic relationship between RIM and RIS activity facilitated by both chemical (neuropeptide and tyramine) and electrical (gap junction) signaling [16,21]. This known interaction should also be interrogated in the context of how these cells may communicate following genetic induction of the hypoxic response. Future work in this area could use tools to measure or modify neuronal activity, such as calcium imaging or optogenetics, to begin answering these questions.”

      (4) Does serotonergic neuron-specific rescue impact the mitochondrial unfolded protein response (mtUPR), given that serotonin signaling has been shown to modulate mtUPR?

      We appreciate this question and suggestion. To determine whether activating the hypoxic response in serotonergic neurons modifies the mt-UPR, we measured hsp-6 expression via qPCR in the ADF and NSM-specific HIF-1 stabilized strains. Interestingly, we find that stabilizing HIF-1 in either the ADF or NSM serotonergic neurons decreases hsp-6 expression relative to WT worms (new Fig. S1G. This could suggest either that the mt-UPR response is impaired in these worms, or that HIF-1 stabilization decreases proteotoxic stress leading to a lower basal level of hsp-6. Although this method of measurement did not allow us to interrogate whether these changes occur in the specific tissues where fmo-2 is upregulated, we have expanded our discussion of these results to emphasize that further investigation of this response should be a focus of future work.

      (5) To remain consistent with the flow of Figure 1, the authors should include fmo-2 expression in NSM:HIF-1S in addition to the ADF:HIF-1S (Figure 1I).

      We appreciate this suggestion and have added NSM::HIF-1S data to Fig. 1I. We found that stabilizing HIF-1 in either the ADF or the NSM has a similar effect on fmo-2 induction.

      (6) It would greatly strengthen the RIS observation if the authors demonstrated that ablation of another neural subtype from their screen does not abolish lifespan extension by vhl-1 RNAi. This would be a good supplemental figure, but it is not necessary for the overall story.

      We thank the reviewer for this suggestion and agree that ablating a ser-7 expressing neuron that was not a hit from our screen would be an excellent additional control. We did find that ablating a non-ser-7-expressing neuron (the RIC, Fig. 3E) did not affect vhl-1 mediated longevity, suggesting that impairing the signaling of any interneuron is not sufficient to disrupt the phenotype. In addition, ablating a ser-7 expressing neuron other than RIS would provide much stronger support for this finding. We have suggested this approach to further validate this working model in the future directions section of our discussion.

      “Finally, additional genetic controls could better support the role of RIS-specific ser-7 expression in genetic activation of the hypoxic response. For example, a ser-7 expressing neuron that was not a hit in our screen could also be ablated and tested for necessity in vhl-1 mediated longevity. This experiment would test whether the ability of RIS ablation to attenuate vhl-1-mediated longevity is not a false positive driven by any disruption to ser-7 expression.”

      (7) For consistency with the rest of the manuscript, it would strengthen the hypothesis if modulating the expression of nlp-17 or its receptors impacted the intestinal activation of fmo-2 transcription.

      This is a great point. We attempted this experiment, but were unable to achieve consistent results (see Author response image 1). This result could be due to indirect effects of the overexpression of nlp-17 signaling modulating fmo-2 induction in a complicated circuit, variability in expression of its receptor, or other complexities within the circuit.

      Author response image 1.

      (8) It would strengthen the manuscript to determine whether serotonergic, GABA, and/or tyramine signaling activate the expression or secretion of this nlp-17 neuropeptide.

      We thank the reviewer for this great idea of experiment. To address this suggestion, we performed qPCR to measure nlp-17 mRNA in WT, hif-1 KO, ADF HIF-1 stabilized, and NSM HIF-1 stabilized strains. Compared to WT and hif-1 KO controls, we observed no change in nlp-17 expression when HIF-1 was stabilized in the ADF or NSM serotonergic neurons (new Fig. S5F). This could suggest either that nlp-17 signaling acts in parallel to serotonergic signaling following genetic activation of the hypoxic response. Alternatively, neuronal HIF-1 stabilization may modify nlp-17 splicing or translation without resulting in detectable differences in mRNA levels. Together, these data indicate NLP-17 signaling is required for longevity following genetic activation of the hypoxic response, although whether this peptide is synthesized or released in response to hypoxic response remains unclear.

      We agree that it is also important to connect nlp-17 expression and/or secretion to other components of this pathway. However, we believe the most effective experiment to confirm a connection between GABA and tyramine signaling and nlp-17 in the context of hypoxia would be to measure nlp-17 expression in strains that manipulate GABA and/or tyramine signaling in a manner that mimics vhl-1 knockdown and extends lifespan. Because we have not yet validated hypoxic-response mimetics for these specific signals, we hope to first generate these strains and then measure their effect on nlp-17 expression in future work. The importance of identifying manipulations to GABA and tyramine signaling that promote longevity has been added to our discussion section. 

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant. This approach would also narrow down which signals are downstream of the genetic activation of the hypoxic response, and which are sufficient to extend lifespan upstream of the hypoxic response in a normoxic environment. One notable target for further exploration is the SER-7 expressing RIS neuron, which plays a role in sleep [16] and stress resistance [17], and can extend lifespan when optogenetically activated under normoxic conditions [18].”

      References:

      (1) Calabrese, E. J., Dhawan, G., Kapoor, R., Iavicoli, I. & Calabrese, V. What is hormesis and its relevance to healthy aging and longevity? Biogerontology 16, 693-707 (2015). https://doi.org/10.1007/s10522-015-9601-0

      (2) Zhou, I. K., Pincus, Z. & Slack, J. F. Longevity and stress in Caenorhabditis elegans. Aging 3, 733-753 (2011). https://doi.org/10.18632/aging.100367

      (3) Yuan, R., Hascup, E., Hascup, K. & Bartke, A. Relationships among Development, Growth, Body Size, Reproduction, Aging, and Longevity - Trade-Offs and Pace-Of-Life. Biochemistry (Mosc) 88, 1692-1703 (2023). https://doi.org/10.1134/S0006297923110020

      (4) Mautz, B. S., Lind, M. I. & Maklakov, A. A. Dietary Restriction Improves Fitness of Aging Parents But Reduces Fitness of Their Offspring in Nematodes. J Gerontol A Biol Sci Med Sci 75, 843-848 (2020). https://doi.org/10.1093/gerona/glz276

      (5) Duric, V., Clayton, S., Leong, L. M. & Yuan, L.-L. Comorbidity Factors and Brain Mechanisms Linking Chronic Stress and Systemic Illness. Neural Plasticity 2016, 1-16 (2016). https://doi.org/https://doi.org/10.1155/2016/5460732

      (6) Mariotti, A. The Effects of Chronic Stress On Health: New Insights Into the Molecular Mechanisms of Brain–Body Communication. Future Science OA 1 (2015). https://doi.org/https://doi.org/10.4155/fso.15.21

      (7) Bellier, A., Chen, C.-S., Kao, C.-Y., Cinar, H. N. & Aroian, R. V. Hypoxia and the Hypoxic Response Pathway Protect against Pore-Forming Toxins in C. elegans. PLoS Pathog 5, e1000689 (2009).

      (8) Palazon, A., Goldrath, W. A., Nizet, V. & Johnson, S. R. HIF Transcription Factors, Inflammation, and Immunity. Immunity 41, 518-528 (2014). https://doi.org/https://doi.org/10.1016/j.immuni.2014.09.008

      (9) Vora, M. et al. The hypoxia response pathway promotes PEP carboxykinase and gluconeogenesis in C. elegans. Nature Communications 13 (2022). https://doi.org/https://doi.org/10.1038/s41467-022-33849-x

      (10) Nakazawa, S. M., Keith, B. & Simon, C. M. Oxygen availability and metabolic adaptations. Nature Reviews Cancer 16, 663-673 (2016). https://doi.org/https://doi.org/10.1038/nrc.2016.84

      (11) Fawcett, M. E., Hoyt, M. J., Johnson, K. J. & Miller, L. D. Hypoxia disrupts proteostasis in Caenorhabditis elegans. Aging Cell 14, 92-101 (2015). https://doi.org/https://doi.org/10.1111/acel.12301

      (12) Ohh, M., Taber, C. C., Ferens, F. G. & Tarade, D. Hypoxia-inducible factor underlies von Hippel-Lindau disease stigmata. Elife 11 (2022). https://doi.org/10.7554/eLife.80774

      (13) Wind, J. J. & Lonser, R. R. Management of von Hippel-Lindau disease-associated CNS lesions. Expert Rev Neurother 11, 1433-1441 (2011). https://doi.org/10.1586/ern.11.124

      (14) Lonser, R. R. et al. von Hippel-Lindau disease. Lancet 361, 2059-2067 (2003). https://doi.org/10.1016/S0140-6736(03)13643-4

      (15) Kaelin, W. G. Molecular basis of the VHL hereditary cancer syndrome. Nat Rev Cancer 2, 673-682 (2002).

      (16) Costa, S. W. et al. A GABAergic and peptidergic sleep neuron as a locomotion stop neuron with compartmentalized Ca2+ dynamics. Nature Communications 10 (2019). https://doi.org/10.1038/s41467-019-12098-5

      (17) Wu, Y., Masurat, F., Preis, J. & Bringmann, H. Sleep Counteracts Aging Phenotypes to Survive Starvation-Induced Developmental Arrest in C. elegans. Curr Biol 28, 3610-3624.e3618 (2018). https://doi.org/10.1016/j.cub.2018.10.009

      (18) Busack, I. & Bringmann, H. A sleep-active neuron can promote survival while sleep behavior is disturbed. PLOS Genetics 19, e1010665 (2023). https://doi.org/10.1371/journal.pgen.1010665

      (19) Powell-Coffman, J. A. & Coffman, C. R. Apoptosis: Lack of oxygen aids cell survival. Nature 465, 554-555 (2010). https://doi.org/10.1038/465554a

      (20) Kruempel, J. C. P. et al. Hypoxic response regulators RHY-1 and EGL-9/PHD promote longevity through a VHL-1-independent transcriptional response. Geroscience 42, 1621-1633 (2020). https://doi.org/10.1007/s11357-020-00194-0

      (21) Bach, M., Bergs, A., Mulcahy, B., Zhen, M. & Gottschalk, A. (2023).

      (22) Leiser, S. F. et al. Cell nonautonomous activation of flavin-containing monooxygenase promotes longevity and health span. Science (2015). https://doi.org/10.1126/science.aac9257

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether hepatic glucose metabolism is differentially regulated by the left and right sides of the LPGi and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi, which were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, and in changes in protein expression in the liver lobes. These data suggested lobe-specific modulation of HGP. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) Image clarity was improved in some cases, but not in others. For example, Figure 3I, showing c-Fos expression, is not convincing due to the image quality and lack of orientation.

      We sincerely apologize for the insufficient image clarity and anatomical orientation in the original Figure 3I. To resolve this issue, we have performed the following revisions in the revised Figure 3I:

      (1) Replaced the original panels with the high-resolution confocal images showing clear c-FOS immunofluorescence in the LPGi.

      (2) Included explicit anatomical orientation indicators (Bregma −6.75 mm) to clearly demarcate the boundaries of the LPGi.

      (3) Added ROI outlines surrounding the LPGi region.

      (2) The methods section states that 8-weeks-old male mice were used in the experiments without specifying the experiments (e.g., brain injection with AAVs or PRV organ inoculation). The authors should include these details.

      We thank the reviewer pointing out this oversight. We have updated the Methods section under "Animals" and specific procedure subsections to clearly state the exact age of animals.

      (1) For retrograde trans-synaptic PRV tracing, 8-week-old mice received intrahepatic viral injections and were sacrificed 5 days post-injection.

      (2) For chemogenetic and optogenetic manipulations, stereotaxic AAV injections were performed at 8 weeks of age. Mice were allowed 4 weeks for viral expression and recovery before undergoing metabolic tests or light stimulation at 12 weeks of age.

      (3) For chemical denervation (6-OHDA), 8-week-old mice were injected into targeted lobes and examined 7 days post-denervation.

      (4) For postnatal innervation mapping, neonatal mice at postnatal week 0 (P0), week 1 (P7), and week 2 (P14) were harvested for tissue clearing.

      (3) The authors should use the exact location of pre- and postganglionic neurons as they often refer to neurons in the sympathetic chain. Their findings should be compared with the existing literature on the location of preganglionic cells.

      We appreciate the reviewer for this feedback. We agree that our original description lacked precise anatomical localization regarding the pre- and postganglionic neurons, and it was inaccurate to state that descending fibers pass through the sympathetic chain (SyC).

      Based on our whole-mount tissue clearing data, we observed that the preganglionic neurons of the brain-liver sympathetic circuit are primarily located in the T6–T12 segments of the thoracic spinal cord. Accordingly, we have revised the text in Results 4 to specify these exact locations.

      Manuscript Revision (Results 4):

      "Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the thoracic spinal cord (T6–T12) send descending fibers via the splanchnic nerves to innervate postganglionic neurons in the CG-SMG (Figure 4A)."

      Furthermore, following your valuable suggestion to compare our findings with existing literature, we reviewed a recent study published in Nature Communications (Harima, Yukiko et al. Parallel labeled-line organization of sympathetic outflow for selective organ regulation in mice. Nat Commun. 2024;15(1):10478). In that study, researchers injected retrogradely transducible AAVs directly into the CG-SMG and traced the preganglionic neurons predominantly to the T8–T13 segments. Their results are largely consistent with our findings. Interestingly, the broader anatomical range observed in our trans-synaptic liver-to-brain mapping (T6–T12) compared to their CG-SMG-specific tracing (T8–T13) reveals a slight discrepancy. This observation suggests an intriguing anatomical hypothesis: a subset of sympathetic preganglionic nerves may bypass the CG-SMG relay entirely and project directly to the liver.

      (4) Figure legends should be revised and matched with the text.

      We apologize for the oversight. We have conducted a comprehensive audit of all figure and legends to ensure precise matching between the main text and the figures.

      Specifically, we have corrected a typographical error in the Figure 1 Legend where panel (C) was mistakenly labeled as a second panel (B), and we fixed a spelling error ("LPG" corrected to "LPGi"). Additionally, we corrected a miscitation in Results (Section 3) regarding Figure 3. In the original text, Figure 3C was incorrectly grouped with blood glucose data, whereas it actually displays the Western blot validation of sympathetic denervation.

      We have revised the corresponding sections in the manuscript as follows:

      Manuscript Revision (Figure 1 Legend):

      “(C) Quantification of PRV-labeled neurons in left and right LPGi across different hepatic lobes: left lateral, median, right posterior, right anterior, caudate, and porta hepatis (n = 3).

      (D) Sankey diagram showing projection patterns from left and right LPGi to individual hepatic lobes. (E and F) Representative slices of EGFP+ and mRFP+ neurons in left (top) and right (bottom) LPGi following PRV-EGFP (right anterior lobe) and PRV-mRFP (median lobe) injections. Proportions of EGFP+, mRFP+, and co-labeled neurons in left and right LPGi (F, n = 3). Scale bars, 100 μm.”

      Manuscript Revision (Results 3):

      “Despite the absence of directly sympathetic input to denervated lobes, systemic blood glucose levels were unchanged compared with controls (Figures 3A-3B, Figure S5A), indicating functional compensation through the remaining intact liver.”

      Reviewer #4 (Public review):

      Summary of General Strengths & Weaknesses:

      The studies here are highly informative for anatomical tracing and sympathetic nerve function in the liver in relation to glucose levels, but because they are conducted in a single species, it is challenging to translate them to humans or determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies provides mechanistically informative. Denervation studies lack proper controls, and sensory innervation in the liver is overlooked.

      We sincerely thank the reviewer for their time and evaluation. We respectfully note that these comments mirror those raised during the previous round of review. We would like to kindly direct the reviewer to the extensive revisions we implemented in our previous resubmission, which directly and comprehensively addressed these exact concerns. These revisions remain intact in the current version of the manuscript. Below, we briefly summarize how each point was previously addressed for your convenience.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      As addressed in our previous revision, we fully agree with this suggestion. We updated the title of the manuscript to explicitly include the species: "Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice." We also clarified the species used throughout the main text to ensure accuracy.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also labels a portion of sensory fibers that need to be ruled out in whole-mount imaging data.

      As detailed in our previous response, we acknowledge this important limitation. In our prior revision, we addressed this concern through both additional data analysis and text revisions:

      (1) We provided SyGlass 3D reconstruction data demonstrating that the TH-positive nerve fibers originate from the celiac-superior mesenteric ganglia (CG-SMG), a well-established sympathetic ganglion (Figure S5F).

      (2) In parallel, we collected dorsal root ganglia (DRG) from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While the T7-12 DRG segments are historically known to contain the sensory neurons that innervate the liver (Anat Rec A Discov Mol Cell Evol Biol. 2004; Auton Neurosci. 2024), we detected only a remarkably sparse number of PRV-positive neurons in these segments. This effectively functionally distinguishes this efferent pathway from primary sensory afferents (Supplementary figure B).

      (3) We explicitly added this methodological limitation to the Discussion section (paragraph 6) of the current manuscript, noting that more selective approaches, such as genetic targeting of sympathetic lineages, will be important for future validation."

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      As outlined in our previous response, we incorporated this crucial context into our revised manuscript. Specifically, we expanded the Discussion section (paragraph 3) to contrast our precise cell-type-specific chemogenetic and optogenetic approaches with historical studies that relied on coarse electrical stimulation. This addition highlights how our current methodology reveals the contralateral and lobe-specific architecture of brain-liver sympathetic control that was previously obscured.

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases in tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though that is clearly assumed. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We appreciate this insightful physiological perspective, which we addressed comprehensively in our previous revision. As we previously agreed, the central nervous system integrates a broad range of afferent signals, and compensatory sensory or parasympathetic mechanisms likely contribute to the observed LPGi activation following hepatic sympathetic denervation.

      To address this, we significantly expanded our Discussion section (paragraph 4) in the prior revision. We explicitly proposed a model wherein hepatic glucose production is regulated by an integrated afferent-central-efferent loop, acknowledging that our current study primarily resolves the efferent component. We clearly noted the lack of direct assessment of sensory or parasympathetic innervation as a limitation and highlighted this dynamic crosstalk as a critical avenue for future investigation.

      Comments on the revised version.

      Across all reviewer comments, the revised resubmission has adequately addressed all concerns.

      Recommendations for the authors:

      Reviewer #4 (Recommendations for the authors):

      No further recommendations aside from tempering the CGRP language, as marking all sensory fibers.

      We appreciate the reviewer for pointing out this important anatomical distinction. We entirely agree that CGRP specifically labels peptidergic sensory afferents and does not represent the entirety of the sensory nervous system.

      We have carefully reviewed the entire manuscript and tempered our language accordingly. Wherever CGRP is mentioned, we have clarified that it serves as a marker for peptidergic sensory fibers, rather than functioning as a pan-sensory marker.

      Manuscript Revision (Results 1):

      "Unlike the NTS, a well-established hepatic sensory center served here as a positive control, the LPGi contained few CGRP-positive cell bodies (Figure S1G), indicating a lack of peptidergic sensory projections."

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank both reviewers for their thoughtful and constructive evaluations of our manuscript. We are grateful that both reviewers found the study to provide a strong behavioral framework for defining sleep in Aedes aegypti and appreciated the breadth of the behavioral and genetic approaches used. We also appreciate the reviewers’ careful identification of several issues requiring clarification, particularly regarding the interpretation of post-blood-meal sleep, the support for the 10-min sleep threshold, possible nutritional confounds in the BSA experiments, the framing of the host-seeking model, and the description of statistical analyses. In the revised manuscript, we have addressed these concerns by clarifying our rationale, tempering several conclusions, revising the statistical reporting and methods, explicitly stating sample sizes, and expanding the Discussion to better acknowledge limitations and alternative interpretations. Where appropriate, we have also revised the text to distinguish more clearly between increased sleep and reduced locomotion, and to frame mechanistic conclusions more cautiously.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Conventionally, a coincidence of sleep increase and locomotion reduction would weaken the certainty of a sleep increase assessment. The authors implied this concurrence observed after blood meal is derived from internal "drowsy" neural state instead of physical "cripple", but they did not use their two high-resolution video tracking velocity or pDoze/Wake to clarify this.

      Thank you for addressing this point. We understand the need to validate locomotion when used as a readout of sleep. We note that analysis of waking activity is normalized to time spent awake, and therefore should be separate from the time spent inactive that is classified as sleep. Based on the reviewers’ suggestions we have reanalyzed some data and revised the relevant sections in include this analysis.: In brief we performed pDoze/pWake analyses on the two high-resolution tracking video from EthoVision XT system. A velocity threshold of 0.4 mm/s was used, with velocities above 0.4 mm/s defined as wake/activity and velocities below 0.4 mm/s defined as doze/sleep state. pWake and pDoze were defined as proportional time metrics of wake/active (velocity > 0.4 mm/s) and doze/sleep (velocity > 0.4 mm/s) within each LD cycle. The conclusion that sleep is increased following blood feeding is supported by these data. We also note (as described in response to Reviewer 2, that this paper represents a step towards describing sleep in mosquitoes. We hope that future application of approaches used in Drosophila, such as brain imaging and indirect calorimetry will further refine our understanding. Along these lines, we have also included a section in the Discussion about how additional measures, including systems like FlyVista might be applied in the future.

      (2) The major molecular component underlying blood meal effect on sleep/locomotion is less certain, because the BSA solution used for feeding contains ATP, which itself is able to enter haemolymph and potentially exerts sleep/locomotion effect. Additionally, the basal or control sleep recording is done after sucrose feeding. It is, however, unclear from the method if this is 10% too? And if the observed sleep level increase after a blood meal is a result of sugar level reduction in the blood (~0.1%).

      We thank the reviewer for raising this important issue. We think it is unlikely that the small amount of ATP used for feeding is driving the sleep phenotype. We have now included this point as a caveat within the discussion, and explained its inclusion.

      (2) Sucrose concentration in controls

      We apologize that this was not clearly stated. Yes, the control mosquitoes were maintained on 10% sucrose, and we have now clarified this explicitly in the Methods and figure legends where relevant.

      (3) Could the effect reflect reduced sugar intake rather than blood/protein?

      We think this is unlikely, however it cannot be ruled out based on the experiments we have run. We have added discussion of this point. However, we note in fruit flies, this has been studied extensively, and loss of sugar under certain contexts reduces sleep. The points above highlight the need for systematic analysis of the dietary components that contribute to sleep in mosquitoes. While we regret being unable to include them in this manuscript, we note that many of these experiments are challenging (with many controls) and have been ongoing for over a decade (with contributions from many labs) in Drosophila.

      Reviewer #2 (Public review):

      (1) The authors settle on a 10-minute immobility threshold, but their own data do not convincingly support this choice… A 15-minute threshold would be better supported by the data as presented.

      We appreciate this evaluation of the sleep threshold. We chose 10 minutes because the first significance in arousal threshold is at the time-point of 10-15 minutes. Therefore, we believe that sleep bouts longer than 10 minutes should be qualified as sleep. We are particularly interested in why arousal threshold continues to increas at 15 minutes. This is either incomplete sleep between minutes 10 and 15 or the presence of multiple sleep states. We have established a new system in the lab using Zantiks that we believe will allow for simultaneous recording of posture and arousal threshold. We now explicitly comment on this in the discussion, and the need for further analysis of the timeframe for which sleep is defined. Nevertheless, we believe we have honed in on a period of 10-15 minutes that serves as a good proxy for sleep regulation. We hope that this initial description of sleep in mosquitoes provides an initial step towards defining sleep, and that future studies that include techniques applied in Drosophila including brain imaging, indirect calorimetry and additional videography will define more nuanced changes in sleep. We have written in limitations and future opportunities to better define sleep throughout the manuscript.

      (2) The primary experimental paradigm measures sleep beginning at Day 4 post-blood feeding, immediately after oviposition... what is being measured as ‘sleep’ could reflect post-reproductive quiescence or recovery rather than diet-induced sleep per se. The BSA experiment partially addresses this, but since BSA also triggers vitellogenesis and egg production, the confound persists.

      We agree this is an important concern. Our intent in measuring sleep after oviposition was to isolate prolonged post-feeding effects from the well-established transient suppression of host-seeking that occurs during the first ~72 h after blood feeding. However, as the reviewer notes, this design does not by itself distinguish post-feeding sleep from other physiological processes associated with reproduction, including vitellogenesis, oviposition, or post-reproductive recovery. To address this issue, we included the experiment measuring sleep immediately after blood feeding, before oviposition. We agree, however, that this rationale should have been stated more clearly and that the limitation remains relevant, particularly because BSA can also support egg development. In the revised manuscript, we have therefore: In the current version we have clarified more explicitly that the immediate post-blood-meal recording was included to show that the sleep increase begins before oviposition; We have also tempered our interpretation of the Day 4–5 phenotype to avoid implying that it is purely diet-driven and fully independent of reproductive state; and expanded the Discussion to acknowledge that blood feeding, protein feeding, and reproductive physiology are closely linked in female mosquitoes and that our current experiments do not fully disentangle these processes. These changes frame the data more cautiously: blood/protein feeding is sufficient to induce a sleep-promoting state that begins immediately after feeding and persists into the post-oviposition period, but the relative contributions of nutrient sensing, egg development, and reproductive recovery remain to be determined.

      (3) The opportunistic vs. determined host-seeking hypothesis… requires actual measurement of host-seeking alongside sleep to be substantiated, or at least the caveats need to be discussed more explicitly.

      We agree with the reviewer. Our intention was to present this as a conceptual model motivated by the temporal dissociation between published host-seeking recovery and the prolonged sleep phenotype observed here, not as a demonstrated behavioral framework directly tested in this study. In the revised manuscript, we have substantially softened this section by clarifying that we did not directly measure host-seeking behavior in the current study; adding explicit caveats that the proposed framework remains speculative until sleep and hostseeking are measured simultaneously in the same animals across the same post-feeding time course. We appreciate this comment and agree that the distinction should be presented as a model for future testing rather than as a central conclusion established by the current data.

      (4) The methods describe ‘one-way ANOVA, followed by Mann-Whitney tests with Welch’s correction,’ which is an internally inconsistent combination…

      We thank the reviewer for catching this lack of clarity. We apologize for this inconsistency. We have fixed this error. In the revised manuscript, we have carefully rewritten the statistical analysis section to specify: which datasets were analyzed using parametric tests (e.g., ANOVA, with appropriate post hoc comparisons where assumptions were met), which datasets were analyzed using non-parametric tests (e.g., Mann-Whitney), and where Welch’s correction was applied, specifically for unequal-variance t-tests, not Mann-Whitney tests. We have also revised Methods, Figure legends and reporting throughout to ensure that the statistical test named in the text matches the reported test statistics. The changes include statistical methods rewritten for consistency and accuracy, and updated figure legends that include exact sample sizes. In addition, one summary spreadsheet of statistical analysis throughout this study is provided and will be submitted as a supplementary file.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear whether there are any systematic changes in preferences over the course of testing that could explain the observed changes in correlation with neural responses, such as changes due to learning (e.g., flavor nutrient conditioning, relief of neophobia), changes in deprivation state, or habituation to/proficiency with the BAT setup.

      For the revision, we have added analysis, including a new figure (Figure 3) between what are now Figures 2 & 4, testing the hypothesis that preference changes across testing days are non-random in direction (e.g., that they reflect attenuation of neophobia). This new analysis failed to reveal evidence supporting the hypotheses that: 1) preference for palatable tastes increases with experience (a result that would make sense given research on neophobia; 2) the preference for aversive tastes decrease with experience; or 3) absolute consumption of any particular taste changes in a reliable direction from session to session (lines 142-157 and new Figure 3).

      A secondary point is whether any changes in preference are attributed to internal individual versus external contextual factors. Both types of variation (i.e., across individuals and across time within an individual) are mentioned in the introduction, but it is not clear what the authors believe about the nature or neural representation of these sources of variation.

      While we assume that differences between rats are due to internal factors (given the controlled home-cage environment), we can’t be sure that some subtle, subthreshold (for us as observers) factor impacts taste preferences. Similarly, while changes across time within an individual is categorically within the individual, we cannot be sure whether some subtle facet of their experiences determines how preferences change (as opposed to it being purely internal). We have added prose to the Discussion session on this topic—including citation of Hilary Schiff’s recent work showing nurture-related preference changes as part of this new prose (lines 387-398).

      With respect to neural data analysis, no individual animal/day data are shown, making it difficult to assess the extent to which differences in correlation match individual differences in preferences and/or changes in preference with time within individuals.

      The revision now explicitly includes Figure panels (with analysis) showing the relationships between individual neural responses and consumption in the first and last BAT tests for a representative rat (lines 172-198; Figures 4A and 4D). As requested in the non-public comments, we have also added waveforms recorded for the representative neuron in an inset to Figure 4B.

      The correlation analysis is also lacking control for the fact that there is a certain degree of "chance" associated with behavioral and neural measures having matching ranks.

      Certainly chance cannot explain our results, which consist centrally of within-rat differences in match (that is, regardless of chance match levels, what we observed was specifically an enhancement of that match for the most recent behavioral assessment compared to an earlier assessment in the same rat)—a finding that is all the more surprising given that: 1) 2 weeks separate that behavior test and the electrophysiology session; and that 2) that gap between the ephys test and the (less well-matched) first behavioral test is only 1-3 days longer. Nonetheless, in appreciation of Reviewer 1’s concern, we have added an independent, convergent analysis to the revision, testing whether the observed pattern vanishes when we shuffle the preference ranks between tastes with neighboring ranks in the behavioral data (a more conservative test than complete shuffles among tastes). The results of this analysis, which are in the new Figure 5, provide further proof that our result is not based on chance—that they specifically reflect a match between neuronal activity and behavior (lines 242-251).

      Finally, …it is unclear to what extent changes in correlation may be attributed to overall changes in responsiveness of the neural population.

      We include several new analyses in the revision that test the hypothesis that the reduction in match between behavioral rankings and neural responses in the second electrophysiology sessions reflects spontaneous or taste-driven changes in neural excitability. These additional analyses reveal no clear between-session differences in baseline and/or taste-evoked responses, or in the percentages of neurons that are taste responsive and/or palatability-related (lines 292-309; Figure 7).

      Reviewer #2 (Public review):

      The manuscript could use additional corollary analyses to provide a more complete picture of the phenomenon. For instance, how many neurons (per animal and in total) have significant correlations with the final BAT patterns? And with the first BAT? Can a time course of such counts be provided? Can some decoding analyses be performed at a single session level to reconstruct a rat's behavioral preference pattern from its neural activity?

      These are all really good ideas. As noted in our response to Reviewer 1, we have implemented all but the last of the suggested analyses, which did not produce evidence suggesting that our results can be explained by changes in neuronal properties between the two recording sessions (lines 292-309; Figure 7). We have also made attempts to apply the decoding analysis; unfortunately, we don’t have large enough samples to obtain stable results such a subtle decoding task (reflecting the last BAT session’s preference pattern is significantly better than the first session’s pattern).

      The manuscript could benefit from additional polishing, both in the text as well as in the figures.

      An extensive holistic edit has been done, starting with suggestions made by Reviewer 2 in the non-public comments.

      Reviewer #3 (Public review):

      Without a behavioral measure collected after recording day 1 intraoral exposure, it is not possible to determine whether taste preference was altered by that experience…The authors' conclusion would be strengthened by adding an intervening brief access test between recording days 1 and 2.

      We very much appreciate Reviewer 3’s suggestion. Alas, the primary authors involved in data collection on this project have moved on, and we won’t be able to collect the additional dataset that would be required. Instead, we have softened the conclusion that we reached in the last section, and suggested the proposed experiment as a future direction (lines 366-374).

      The current experimental design exposes animals to 3 distinct sets of substances … [that] differ in identity … and concentration. Because palatability is known to be comparative depending on the other substances available and concentration-dependent, this introduces challenges to interpretation, [and] without more clarity, it is difficult to evaluate whether the interaction of different tastes within the sets of stimuli biases the main conclusions.”

      This is an interesting point. Analyzing each set of batteries separately and performing between-battery comparisons would require a larger number of experimental subjects then we have in our current sample size. That said, while we acknowledge that taste preference ranking is relative, we believe the ranking system used here deviates little, if any, from the 'true' ranking (and is therefore significantly relevant to gustatory activity). This is supported by our newly obtained result in response to Reviewer 1 & Reviewer 2 (see above), where an ancillary shuffle analysis (Figure 5C) showed that swapping adjacent preference orders eliminated the experimental effects across all batteries.

      Responses to sweet tastes are not reported in the electrophysiology data. This is seemingly the case because rats given set 1 received no sweet stimulus while rats given set 2 received to 2 distinct sweet tastes. Finally, rats given set 3 did not receive quinine, yet quinine is reported in electrophysiology data.

      We are unsure of the source of this confusion—in every case, the rat received the same tastes in the electrophysiology sessions that were delivered in the BAT preference tests—but in appreciation of Reviewer 2’s concern, we have modified the text and table to ensure: 1) that panels reflecting data from single example rats (panels that therefore necessarily include only a subset of possible tastes) are clearly marked as such; and 2) that the nature of which taste batteries were delivered is more explicit (lines 104-112; 172-178).

      The choice of reporting average lick cluster size is problematic because the authors use thirsty rats with 10-second-long trials. Thirsty rats are likely to lick in relatively long clusters, especially for neutral and palatable tastes. If the rat is mid-cluster when the trial ends, the final cluster would be cut off prematurely, resulting in shorter overall average lick cluster size, disproportionately affecting neutral and palatable tastes over aversive tastes.

      We have ourselves been deeply concerned with this issue, and in fact have recently published a paper that includes within it a direct test demonstrating that calculations of lick bout lengths from 10-sec BAT trials result in taste palatability estimates that are identical to (and less noisy than) those generated from more classically-used 15-min ad lib licking. We now cite this paper (Stone, Lin, et al., 2026) in the Methods section, along with text clarifying how we calculated lick clusters. We also conducted an additional analysis that estimates taste preference after removing these “prematurely ended bouts” without changing the observed pattern of results (lines 494-510).

      Of course, even if this last analysis had changed things, the result of clusters being cut short by the end of a trial would be an underestimation of the preference for the palatable tastes (which drive far more licking than aversive tastes and are therefore more likely to be mid-bout at the end of a trial). Such an underestimation would in turn be expected to reduce the observed neural-behavioral correlation. This fact highlights the robustness of our findings.

      Canonical palatability rankings may not apply to the concentrations selected in every stimulus set. This is particularly true for set 1, which included two concentrations of citric acid and quinine for the behavior. It is also not clear which concentrations are reported in Figures 3A2 and 3B2. Meanwhile, the concentrations of quinine and citric acid used for electrophysiology are quite low.

      In the revised Methods section, we explicitly motivate our reasoning (including citations) behind canonical rankings for each taste battery used (lines 513-522). Every taste used was of agreed-upon preference levels, and in the rare case that two concentrations of the same taste were used, both were known to have distinct palatabilities (e.g., 0.1M NaCl is preferred to 0.05M NaCl). This careful selection of tastes ensured that it was trivial to avoid misordering of canonical palatability rankings.

      And even if mistakes in canonical rankings had been made, the impact of these inaccuracies in these rankings would have been minimal. Our findings are primarily driven by high levels of inter-individual (between different rats) and intra-individual (day-to-day fluctuations within the same rat) preference differences. Given this variability, the fact that the brain-behavior correlation were consistently worse using these rankings almost certainly means that the neural activity matches preference behavior—our thesis. This conclusion is further supported by our shuffle analysis, which demonstrated that randomizing the order did not yield superior correlations between taste ranking and GC activity.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [<sup>3</sup>H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we have also explicitly stated this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we mention and describe alternative approaches for assessing parasite viability. This also includes the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We have revised the text and now explicitly mention the use of dual staining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al., 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non-commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer #2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We have clarified this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We have added this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post-antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We have revised the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      The authors observed differences between their i-lacZ assay and conventional PRR assays attributable to the accumulation of lacZ enzyme at higher levels of surviving parasites, followed by degradation. Have the authors tested how long lacZ remains stable in standard or overgrown parasite cultures?

      We thank the reviewer for this important question. We assessed the stability of β-galactosidase activity in parasite lysates stored under different conditions and observed that the enzymatic activity remained stable for up to 21 days when lysates were stored at either -20°C or 37°C (Hellingman et al., 2024).

      We have not systematically characterized β-galactosidase stability in standard or overgrown parasite cultures. However, in experiments involving overgrown cultures, we observed that the β-galactosidase-derived signal decreased rapidly in overgrown culture settings, suggesting that enzyme stability in overgrown cultures is lower than in standard cultures and parasite lysates.

      At what parasitemia were the counts reported in Figure 1D conducted at?

      We thank the reviewer for this clarification request. The measurements shown in Figure 1D were performed at approximately 3% parasitemia, assuming an erythrocyte infection rate of 10-fold within 48 hours as parasite cultures were initiated at 0.3% parasitemia and incubated for 48 hours under rapamycin before the measurements were performed.

      Figure 2: Is the increasing background in DMSO-treated parasites attributable to leakage of the di-cre system contributing to a background level of LacZ induction? To what extent would this impact results in the PRR assay format?

      We thank the reviewer for this important observation. We agree that low-level leakage of the loxP-DiCre system may contribute to the increased lacZ signal observed in DMSO-treated parasites under overgrowth conditions. However, this effect was only observed when parasites were allowed to proliferate extensively in the absence of effective drug pressure.

      In the context of the MULT-i<sup>2</sup> assay, these conditions correspond to compound concentrations that do not affect parasite survival or replication. Such concentrations are outside the range of interest for evaluating antimalarial activity, as they represent inactive treatment conditions. Therefore, although reporter leakage may contribute to background signal under extreme overgrowth conditions, we expect this effect to have a negligible impact on the interpretation of MULT-i<sup>2</sup> assay results.

      Please define the abbreviations used (e.g., NONMEM and DV).

      We thank the reviewer for pointing this out. We have revised the manuscript to define all abbreviations at their first occurrence in the text and have added the relevant terms to the abbreviation list.

      Line 297: cPRR assay - give citation.

      We thank the reviewer for pointing this out. We have added the appropriate citation for the cPRR assay at the indicated location in the revised manuscript.

      The 2 in MULT-i<sup>2</sup> is not always superscripted.

      We thank the reviewer for pointing this out. We have corrected the formatting throughout the manuscript to ensure that the “2” in MULT-i<sup>2</sup> is consistently presented as a superscript where appropriate.

      Reviewer #3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the MULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1) The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We have added more explanations to the Discussion including the strengths and limitations.

      Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study reports a novel function for syntaxin 11, a specialized SNARE protein critical for the immune system whose mutations cause familial hemophagocytic lymphohistiocytosis type 4. The data convincingly show that depletion of STX11 impairs store-operated calcium entry in Jurkat T cells and that this defect is recapitulated in primary cells from a patient suffering from the disease; the authors further show that the syntaxin interacts with the pore subunit of the ORAI1 channel and propose that it primes the channel by promoting the assembly of multimers before activation by its endogenous ligand, the ER Ca2+ sensing protein STIM1. This is a conceptually important claim that challenges the prevailing view that all structural transitions in ORAI1 are STIM-driven. The data are high-quality and broadly consistent with the interpretation, but alternative mechanisms for the defects are not considered; additional work should rule out vesicular trafficking, discuss other mechanisms, and address methodological issues.

      We thank the editor and reviewers for assessing our work. We have now included additional experiments in a new main Figure 2, which directly rule out any general or Orai1 plasma membrane trafficking defects in Syntaxin11-depleted cells. There are additional experiments and/or analysis in many other figures, throughout the paper. We have included new and missing methods, quantifications and calibrations, and provided response to each of the reviewer’s comments below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Patients with STX11 mutations develop familial hemophagocytic lymphohistiocytosis Type 4, a fatal immune disorder marked by defective T and NK cell cytotoxicity and cytokine storm. The conventional explanation attributes this to impaired cytotoxic granule release, but this has never fully accounted for the broader disease picture. This study proposes an alternative mechanism. The authors show that STX11 is required for store-operated calcium entry through ORAI1 channels, which are essential for both cytotoxic killing and NFAT-driven gene expression in T cells. In STX11-deficient cells, ORAI1 currents drop, NFAT nuclear translocation fails, IL-2 expression is suppressed, and degranulation is impaired. These defects are largely rescued by ionomycin or a constitutively active ORAI1 mutant, placing the primary lesion at calcium signaling rather than the fusion machinery. Mechanistically, STX11 binds the C-terminal tail of ORAI1 via its Habc domain and maintains ORAI1 in a state competent for productive assembly prior to STIM1-dependent gating, a step the authors call "priming."

      Strengths:

      The paper identifies a novel and disease-relevant role for STX11 in calcium channel regulation and raises the possibility of using channel agonists as a therapeutic strategy in the disease. The biochemical and functional data are of high quality and generally consistent with the interpretation. The proposal that a non-conventional syntaxin directly interacts with ion channels to prime its activation is novel and interesting.

      Weaknesses:

      For readers to appreciate the value of patient experiments derived from a single individual, the authors should quote prior studies showing that STX11 protein levels are abolished in all known human STX11 mutations. The priming model, while functionally well-supported, rests on indirect structural evidence, and the precise conformational transition involved remains to be defined. These are acknowledged limitations, but alternate mechanisms have not been explored and formally excluded. More direct evidence should be provided to exclude the possibility that STX11 could act as a conventional SNARE and sustain calcium fluxes by promoting the delivery of additional ORAI1 channels from vesicles.

      In the revised version, we have included references for all those prior STX11 human mutations that have been biochemically characterized till date. The reviewer has correctly pointed out that STX11 protein levels were almost abolished in almost all previously reported mutations. See line 168-173. Therefore, the prior STX11 patient mutations are essentially comparable to the frameshift mutation characterized in this study, in terms of STX11 protein depletion and, therefore, the mechanisms underlying the phenotypic defects reported here as well as earlier. We, therefore, believe that our data from even a single FHLH4 patient, with severely depleted STX11 levels, and additional knockdown studies across three different cell lines, are representative of majority of STX11 mutant FHLH4 patients that have been previously characterized.

      Regarding the Reviewers’ concern that absence of STX11 as a conventional SNARE could affect Orai1 channel delivery from intracellular vesicles. We would like to point out the following:

      (1) In Miao et al. 2013 (1), Figure 3C-D, we showed that expression of a dominant-negative mutant of NSF, a non-redundant protein in vesicle trafficking, impaired vesicle trafficking but did not affect SOCE. This experiment had essentially ruled out a role for vesicle trafficking in SOCE. In the same paper, we had also shown that Orai1 levels in the PM do not increase post-store depletion (Figure 3-figure supplement 2).

      (2) SNAP23/25 form a four helical bundle with R- and Q-SNAREs in orchestrating vesicle fusion. In this paper, we have ruled out a direct role for SNAP23, SNAP25 and SNAP29 in SOCE (Figure 2-figure supplement 3).

      (3) In v1 of this manuscript, we had shown that U2OS cells stably expressing Orai1-BBS-YFP have identical levels of Orai1 in the PM with and without STX11 depletion (Supplementary Figure 3B). This showed that the biosynthesis or delivery of Orai1 to the PM is not affected by STX11 depletion. The levels were also assessed in store-depleted U2OS cells but not included because in Miao et al. 2013 we had already established that levels of PM Orai1 remain essentially equal in resting versus store-depleted cells.

      In the revised version, we have included the data from store-depleted cells in U2OS and also done quantification of PM Orai1 in HEK293 and Jurkat T cells. In addition, we have added three independent membrane trafficking/ vesicle secretion assays performed in STX11-depleted cells (new Figure 2 and associated supplements). In all cases, we find no evidence of Orai1 in intracellular vesicles or a general defect in membrane trafficking/ secretion in STX11-depleted cells. Orai1 is constitutively and stably expressed in the PM in resting as well as store-depleted cells in three different cell lines.

      (4) Most importantly, in Figure 7I-J of this manuscript, we showed that calcium influx from a constitutively active mutant Orai1 (Orai H134S) is identical between STX11-depleted and scramble control cells. If wildtype Orai1 was indeed stuck in vesicles in STX11-depleted cells, then how would mutant H134S Orai1 be able to rescue the defect in SOCE? We have included the quantification of PM levels of Orai1 mutants w.r.t WT Orai1 in new Figure 8-figure supplement 3B and 3D.

      In summary, we have now done several new experiments to directly measure Orai1 levels in the PM and general vesicle trafficking assays in HEK293 and Jurkat T cells and have found no defects in these upon STX11 depletion.

      Regarding STX11 induced precise conformational transition, we are trying to setup collaborations with scientists who might be able to visualize this in situ. Please note that while purification of isolated pore subunits of ion channels followed by crystallization or expression in synthetic membranes for cryo-EM is currently considered a gold standard in the analysis of ion channel pore subunits, we have shown that ion channels are dynamic macromolecular complexes, in vivo (2), where synaptic proteins dynamically bind to induce conformational changes and affect their stoichiometry (2). Please also see (3) and (4). More advanced approaches, therefore, need to be developed to enable visualization of the dynamics of ion channel macromolecular complexes in their native environment in situ. In the absence of such approaches, the structural insights obtained from detergent-purified isolated subunits will remain incomplete.

      Reviewer #2 (Public review):

      Summary:

      Vig's lab delineates a critical role for STX11 in CRAC channel function, particularly in the context of the fatal immune disorder familial hemophagocytic lymphohistiocytosis type 4 (FHL4). They demonstrate that Syntaxin 11 directly binds and regulates Orai1, and that STX11 depletion abolishes CRAC currents and downstream signaling. Loss of STX11 reduces IL2 gene expression and impairs degranulation, both of which are rescued by the constitutively active Orai1 mutant H134S, whereas a gain‑of‑function mutant targeting the C‑terminus fails to restore these defects. The authors conclude that STX11 primes Orai1 for optimal local assembly that is independent of STIM1 yet required for CRAC channel gating.

      Strengths:

      This study is firmly grounded in disease biology and demonstrates that STX11 downregulation leads to profound functional defects. Using a comprehensive suite of methods and analyses, the authors interrogate the co-regulation of STX11 and Orai1 and present a near-complete view of STX11's modulatory role in CRAC channel function and downstream signaling pathways. The figures are clear, and the statistical analyses are rigorous and convincing.

      Weaknesses:

      The authors conclude that Syntaxin 11 directly binds Orai1. This conclusion is well supported by a multifaceted approach, including co-immunoprecipitation (co-IP), molecular dynamics simulations, co-localization/FRET assays, and targeted mutational analysis-all of which are thoroughly executed. While the interaction appears reasonably strong in co-IP experiments, the STX11-Orai1 interaction is comparatively weaker in pull-down assays, which the authors attribute to instability of the purified His-STX11 protein. A remaining gap is direct evidence of interaction in live cells; this is understandably challenging given that fluorescent tagging of STX11 is not feasible. Fully resolving this question lies beyond the scope of the present study and will require more advanced approaches to capture STX11 binding dynamics.

      We thank the reviewer for acknowledging that analysis of the dynamic binding of STX11 will require standardization of advanced techniques which are beyond the scope of the present study. We plan to continue developing methods that will allow us to visualize the binding and unbinding of STX11 to Orai1 in vivo.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Mechanistic issues:

      (1) More direct evidence should be provided to exclude the possibility that STX11 could act as a conventional SNARE and sustain calcium fluxes by promoting the delivery of additional functional channels, stored in secretory vesicles or recycling endosomes, to the membrane. A significant fraction of the ORAI1 channel is in vesicles, and the mobilization of this intracellular pool regulates the rates of calcium fluxes in HEK-293 cells (PMID 26116575) and effector T cells (PMID: 35217583). Mobilization of this pool could account for part or all of the functional effects reported here. The only evidence that STX11 depletion does not impact the plasma membrane availability of the channel relies on one single flow cytometry profile (Supplementary Figure 3). This experiment is performed in U2OS cells stably expressing a fusion protein containing an extracellular bungarotoxin site. This cellular system is only used here; the other data are obtained either in HEK-293 cells or in Jurkat T cells. The level of endogenous STX11 in U2OS cells is unknown, and the efficiency of protein depletion has not been assessed. The efficiency of depletion should be shown, and the total number of channels assessed by comparing the expression levels of permeabilized and non-permeabilized cells. A flow cytometry profile of cells treated with thapsigargin should be included to match the experimental conditions of functional recordings. It would be valuable to repeat this experiment in Jurkat T cells by expressing ectopically the channel tagged on the extracellular side. To further exclude the involvement of vesicular trafficking, the authors should also evaluate the contribution of VAMP8, as this R-SNARE has been proposed to interact with STX11 to regulate the exocytosis of specialized granules in cytotoxic T cells (PMID: 26124288).

      We have not observed significant levels of intracellular Orai1, as described in the PMID 26116575 paper. Multiple technical reasons could explain the artefactual appearance of intracellular Orai1. HEK293 is an embryonic kidney cell line, with most cells showing a distinct spindle shape and filopodia as shown on the ATCC website (https://www.atcc.org/products/crl-1573). The HEK cells shown throughout the Hodeify et al. 2015 paper (PMID 26116575) lack this typical morphology that majority of the HEK cells should show. Transfecting cells with high amounts of DNA and liposomes can severely affect the health and morphology of the cells leading to artifacts where PM proteins appear to be stuck intracellularly. The plasma membrane of the cell in movie 1 of PMID 26116575, for instance, also shows membrane blebs or membrane ruffles. The authors should have used a PM marker such as WGA (wheat germ agglutinin) to distinguish PM Orai1 from any intracellular Orai1 to establish whether what appear as intracellular Orai1 vesicles are not blebs of PM Orai1. Similarly, no endo/exocytotic vesicle marker is used to distinguish them from Orai1’s presence in apoptotic vesicles of unhealthy cells. In view of the overall abnormal morphology and any PM or intracellular vesicle marker, the claim that Orai1 resides in intracellular vesicles is unfounded.

      Similarly, the PMID: 35217583, quoted by reviewer #1, is about in vitro differentiated primary mouse T cells. The study lacks any detail of the generation of the HA-tagged mouse Orai1 plasmid, used in the study, or even a back reference to show how this plasmid was validated for normal expression in primary T cells earlier. Ectopic expression of CMV promoter-driven plasmids in mouse primary T cells is extremely challenging and often results in poor cell health and incomplete and selective expression in only 5-10% of cells. It is unclear if the same HA-tagged human Orai1 plasmid that was used in PMID 26116575 is being used in this study to express in primary mouse cells. In the absence of all this information, it is unclear whether there was an issue with the generation of a new HA-tagged mouse Orai1 construct, which made the protein get stuck intracellularly. Or potentially the expression of a human protein in mouse primary T cells is the problem. The functional verification of the construct by showing rescue of SOCE in Orai1-deficient primary mouse cells is an absolutely essential control but is missing. In the absence of these controls, one cannot disregard decades of robust data on the localization of Orai1 in the PM from multiple labs and papers that have established its PM localization conclusively (5) (6).

      Most importantly, in both the PMID 26116575 and PMID: 35217583, HA tagged-Orai1 is detected using a bivalent antibody, followed by secondary antibody. It is well known that cross-linking of cell surface receptors or proteins using bivalent antibodies has a major caveat that involves antibody-mediated clustering and capping, especially in lymphocytes, which typically induces rapid internalization of the entire antigen-antibody complex. The paper by Sekine-Aizawa et al., in 2004, showed that tagging of receptors/ channels with bungarotoxin binding site (BBS) followed by labelling with bungarotoxin (BTX) bypasses this confounding factor and, therefore, allows accurate estimation and localization of PM versus intracellular proteins. The artefactual bivalent antibody-induced endocytosis continues even when cells are incubated on ice as endocytosis is only slowed but not stopped on ice.

      Due to this challenge, we have generated BBS tagged Orai1-YFP. We use flow cytometry to show PM Orai1 because it is an unbiased and quantitative way of showing surface Orai1 expression (estimated by measuring the intensity of surface-bound BTX), simultaneously, in thousands of cells with no potential for visual bias in the selection and imaging of cells. BTX labelling is done on ice and cells are washed and fixed right after labelling with BTX-A647 to stop endocytosis. In the revised version, we have used both complementary approaches of flow cytometry and microscopy, showing representative images of cells, alongside quantifications. All new experiments were performed in HEK293 and Jurkat T cells to estimate surface versus total expression in a new main Figure 2. The post-store-depletion data have been added to the existing U2OS experiment in new Figure 2-figure supplement 2A-E.

      Our BBS-tagged Orai1 construct also has a YFP tag at the C-terminus. Since the flow cytometry experiment involves gating on YFP-positive cells, the total number of channels (biosynthesis) can be compared by looking at the YFP intensities in the scramble versus STX11-depleted groups. These data have now been included in the previous and new experiments. In none of the cases could we detect any difference in YFP or BTX-A647 intensities, pre- or post-store-depletion in scr or STX11-depleted cells which would indicate defects in biosynthesis or Orai1’s presence in vesicles in any group.

      Regarding a role for VAMP8, in Miao et al. eLIFE 2013 (1), Figure 3C-D, we showed that expression of a dominant negative mutant of NSF, a non-redundant protein in the vesicle trafficking pathway, impaired Transferrin receptor recycling within 20 hours but did not affect SOCE at all. This experiment had conclusively ruled out any role for vesicle trafficking in SOCE and therefore assessment of the role of each of the individual proteins involved in membrane trafficking becomes redundant. We have additionally ruled out a role for SNAP23/ SNAP25/ SNAP29 (old supplementary Figure 12, new Figure 2-figure supplement 3A-C). SNAP23/25 form a four helical bundle with most R-SNARE and Q-SNAREs to orchestrate vesicle fusion. R-SNAREs, typically, need SNAP23/25 to interact with Q-SNAREs. Since a role for these non-redundant proteins has also been ruled out by us, it is unlikely that VAMP8 plays a role in modulating the effects of STX11 in SOCE.

      (2) Both ORAI1 and STX11 are S-Acylated on cysteine residues, and this post-translational modification promotes their recruitment to the immune synapse (PMID: 24910990, 34913437). One possibility that should be discussed is that STX11 could enhance the recruitment of the ORAI1 channel into lipid domains rich in cholesterol, thereby favoring its activation. This type of priming would still require direct interaction between the two proteins but involve a different mechanism than the one discussed by the authors. This could be experimentally tested by expressing a STX11 mutant lacking the cysteine residues required for its S-Acylation. It would also be interesting to test whether depletion of STX11 impairs the recruitment of ORAI1 to the immune synapse forming between Jurkat T cells and antigen-presenting cells.

      In our experiments reported in this paper, we have used soluble anti-CD3 as well as plate-coated anti-CD3 in combination with soluble anti-CD28 to stimulate Jurkat and primary T cells. These antibodies are routinely used to stimulate T cells and this type of stimulus doesn’t depend on the formation of a classical immune synapse with an antigen-presenting cell (APC). Despite the absence of synapse, the T cells get fully activated and functional, as seen by NFAT translocation and secretion of cytokines, such as IL-2 in new Figure 4. Therefore, whether there is a defect in the recruitment of Orai1, or STX11, to T cell synapse formed with an APC is not within the scope of this study. Furthermore, accurate analysis of protein localization within the immune synapse requires a dedicated study employing sub-diffraction resolution microscopy approaches.

      Regarding PMID: 24910990, and the mechanism of recruitment of STX11 to the membranes. We believe this remains unknown. The frameshift mutant used in our study lacked all terminal cysteines, which have been earlier proposed to be crucial for membrane targeting, as well as a terminal part of the SNARE domain and yet it localized to the PM just as well as wild-type STX11 (See new Figure 5E). We have added this result in the text line 302-305 and removed the line stating that post-translational modifications of the terminal cysteines target STX11 to the PM as previously claimed in PMID: 24910990 and mentioned in v1 of this paper. In view of these new data, it currently remains unknown whether potential attachment to PIP2 in PM via various basic residues, spread throughout the sequence (7, 8), or binding to another protein targets STX11 to the PM (PMID: 26771955). There is no obvious poly-basic stretch in STX11 sequence, therefore, a systematic and focused deletion and mutagenesis study will be needed to individually assess the above possibilities which is outside the scope of the present study.

      Methodological issues

      (3) Since STX11 colocalize with Orai1 already in basal conditions, independently of STIM1, it could influence basal calcium levels. This cannot be appreciated from the data presented, because all the SOCE protocols start in calcium-free conditions, preventing baseline comparison between WT and STX11-deficient cells. A potential difference in basal calcium levels should be explored, and the impact of STX11 depletion on basal calcium fluxes should be documented by calcium shifts (2 mM → 0 mM → 2 mM) or manganese quenching approaches.

      The cells are typically loaded with Fura2 in 2mM calcium containing Ringer’s buffer. We switch the cells to 0mM right at the start of the SOCE protocol and start imaging within 5-10 seconds. Therefore, in our experience, the baselines of scramble versus STX11-depleted cells should show a difference even in the SOCE protocol if the basal calcium levels are affected because Fura2 is already present and bound to basal calcium present in the cytosol at the start of the protocol. Still, we have done the experiment suggested by the reviewer as shown in Author response image 1. The assay started with cells in 2 mM extracellular Ca<sup>2+</sup> followed by addition of 10 mM EGTA, which according to the following equation quenches the 2 mM extracellular Ca<sup>2+</sup> (https://somapp.ucdmc.ucdavis.edu/pharmacology/bers/maxchelator/CaEGTA-TS.htm).

      where, [Ca2+]<sub>Free</sub> is the free/unbound Ca2+, [Ca2+]<sub>Total</sub> is the total Ca2+, [EGTA]<sub>Total</sub> is the total EGTA concentration and Kd is the dissociation constant between Ca2+ and EGTA at 37°C and pH 7.4.

      To confirm complete sequestration of the extracellular Ca<sup>2+</sup>, we also repeated the assay with 20 mM EGTA but did not notice any difference between the 10 mM and 20 mM EGTA conditions. We do not see any differences in basal calcium, under any condition, between Scr and STX11 shRNA treated cells HEK or Jurkat T cells.

      Author response image 1.

      Representative Fura-2 traces of Scr (black) and STX11 (red) shRNA-treated HEK293 (A) and Jurkat (B) cells, where the cells were incubated with 2 mM Ca<sup>2+</sup> followed by addition of 10 mM EGTA to quench the 2 mM Ca<sup>2+</sup>. Since we do not have access to a perfusion system, we could not test Fura2 response after re-addition of 2mM calcium to the existing EGTA and Ca<sup>2+</sup> mixture. However, the transition from 0mM to 2mM is already shown in Figure 8.

      (4) The quantification of the calcium imaging data is problematic and requires clarification. In most figures, the data are shown normalized to the control condition. According to the method section (lines 721-724), 100% is the maximum value of the scramble shRNA group (amongst the three experiments). But what was measured here? The slope during calcium readmission? The peak amplitude after calcium readmission? Expressed as a ratio or as calcium values? The recordings are presented in micromolar calcium concentration. This implies calibration, but a calibration procedure is not mentioned. Please clarify. For the recordings of constitutive calcium entry in Figure 7, this normalization is not performed, and the data are expressed as ratio values. Here, it looks like the parameter quantified and compared is the absolute ratio value after calcium readmission. This is inappropriate. The trace in Figure 7G shows that the basal levels differ by more than two ratio units between control and STX11-depleted cells. Normalizing the data in Figure 7G to the basal ratio value would show no difference in the peak response amplitude between the two conditions. These data should be re-analyzed, and both the slope and the amplitude of the response should be presented, with statistics performed on independent recordings, not on individual cells pooled from different experiments. Cells from the same recording are not experimentally independent samples but rather replicates of the same experiment.

      We have updated the relevant method section with more details in lines 823-858. The peak amplitude after calcium readmission was measured and compared to the baseline as described in the updated methods. The Fura 2 calibration method has also been added to the revised version, we apologize for this omission earlier. Separate calibration was done for Fura 2 experiments in all figures except Figure 7 from version 1 (new Figure 8). The experiments in new figure 8 were done using a different objective (20X, water) and, therefore, although a separate calibration was done for these experiments it was not applied to the data. We apologise for this omission on our part and have now applied the respective calibration to these experiments.

      The trace in Figure 7G of version 1 should look different even in 0mM calcium in our opinion. The reason is the same as that explained in point 3 above; CAD-mediated constitutive activation of Orai1 is one of the strongest. The cells, even when they are being loaded with Fura2 in Ringer’s buffer with 2mM calcium, are constitutively recruiting calcium ions. This should result in a shift in Fura2 excitation due to higher levels of basal calcium. When the cells are switched to 0mM calcium and imaged within 4-5 seconds, the intracellular Fura 2 is still bound to all this extra calcium in the cytosol and therefore the baselines should show a significant difference. If the cells were imaged for several minutes in 0mM calcium, we might have seen the difference in basal ratios slowly reducing. However, we switched to 2mM calcium within 120sec. At this point any free Fura2 would be expected to bind incoming calcium again. For the same reason, the cytosol of Scr cells which would still have relatively higher levels of intracellular calcium concentration compared to STX11-depleted cells will show a smaller further increase in 2mM due to calcium-dependent inhibition of CRAC currents within the time frames we have measured. We have now added the Fura-2-calibrated response in the new Figure 8G-H. We have shown below the baseline-subtracted (normalized) response for Figure 8G. As you can see, there is still a significant difference in 2mM calcium between Scr and STX11-depleted cells but in this representation the important difference at 0mM is masked, we have therefore chosen to retain the original figure with Fura-calibrated values at 0 as well as 2mM calcium in Figure 8G-H.

      The cells shown in the old Figure 7G of version 1 were not from the same recording but from three different experiments. Because the cells were imaged with 20X objective in these experiments to allow selection of Orai1-CFP or mutant Orai1-CFP and CAD-YFP double-positive cells, the number of cells analyzed per experiment was less compared to other experiments. We have now shown Fura 2 calibrated values in the new Figure 8G-L. We have also re-done the statistical analysis on three independent experiments from each. As shown in Author response image 2, the difference is still statistically significant whether we show merged cells from all three experiments or single representative experiment out of three repeats. We believe merged cells have more information to offer and therefore have retained the same figures with the original analysis in the main Figure 8G-L.

      Author response image 2.

      Box plots representing quantification of individual repeats of constitutive calcium influx in Scr and STX11 shRNA-treated HEK293 cells expressing YFP-CAD and Orai1-CFP without (A) and with baseline subtraction (B). (C-D) Quantification of individual repeats of Scr and STX11 shRNA-treated HEK293 cells expressing Orai1-H134S (C) and Orai1-ANSGA (D) mutants.

      (5) Quantification of pull-down experiments. The binding data in Figures 4F, 5F, and 5J are presented largely qualitatively. Densitometric quantification with statistical comparisons across wild-type and mutant conditions would make these results more convincing, particularly given that the authors themselves acknowledge the interaction appears relatively weak in vitro.

      This has been done and included alongside the respective panels in the new Figure 6G and 6L (for old Figure 5F and 5J of version 1, where differences appeared relatively small in some experiments). The differences across lanes in both figures and their repeats were statistically significant.

      Figure 4F showed a clear and visually significant difference in binding across lanes and repeats and therefore no quantification is needed for these experiments in our opinion.

      Limitations of the study and mechanistic inferences.

      (7) Interpretation of the ORAI:ORAI FRET and crosslinking data. STX11 depletion increases basal ORAI:ORAI FRET (Figure 7A-C) and shifts crosslinked species toward higher molecular weights (Figure 7D-E). The authors interpret this as ORAI1 being trapped in an unprimed state, but higher FRET and higher-order species would conventionally suggest increased rather than decreased assembly. The paper needs a clearer mechanistic explanation of what "unprimed" looks like structurally. Is this aberrant crowding, non-productive oligomerization, or something else? The distinction between a change in intermolecular distance within existing oligomers versus an increase in oligomer density matters here and should be addressed.

      Higher ORAI: ORAI FRET and a shift in the size of crosslinked Orai1 oligomers, when analyzed together, suggests formation of ‘non-functional’ higher-order oligomers. Higher order does not necessarily translate to better function in the case of ion channels, it can also lead to non-selectivity or formation of ‘non-productive’ oligomers, as mentioned by the reviewer. It was shown by us earlier in Li et al. (2016) (2) that bigger oligomer size revealed by higher number of photobleaching steps of Orai1 did not translate to better function but led to non-selectivity.

      In new Figure 8, crosslinking with BS3, which has a spacer arm and working distance of ~11 Å, very likely reflects a change in the number of subunits within individual oligomers and not crosslinking of independent existing oligomers. This is because we show that neither total Orai1 expression nor Orai1 expression in the PM change in any group in new Figure 2. FRET works best within 1 to 10 nm distance, and therefore, in theory, can lead to energy transfer between neighbouring Orai1 oligomers in high Orai1-expressing cells. However, because there was no change in Orai1 abundance in the PM (new Figure 2) or distribution within PM (new Figure 7E,F,J,L) of any group, FRET changes also likely reflect intra-oligomer changes rather than inter-oligomer interactions. FRET changes can also arise from a change in the respective orientation of fluorophore pairs but when analyzed together with crosslinking studies, changes in pore assembly likely coincide with conformational shifts in Orai1 protomers. Furthermore, FRET has been used earlier to show shifts in conformation of other ion channels (9). Therefore, we believe that, when used together, these two approaches strongly suggest an intermediate conformational state along with a change in number of Orai1 subunits per channel since there was no evidence of overcrowding in the PM or obvious segregation of Orai1 in specific regions of PM in new Figures 2 and 7E,F,J,L.

      We could not assess whether the oligomers of Orai1 formed in the absence of STX11 possess an intact pore. The presence or absence of pore in STX11-depleted cells will require extraction of Orai1 oligomers from native membranes and performing systematic structural analysis using cryo-EM or related approaches which is outside the scope of this study.

      (8) The ANSGA versus H134S discrepancy. H134S ORAI1 rescues calcium influx in STX11-depleted cells (Figure 7I-J), but the ANSGA mutant does not (Figure 7K-L). The authors conclude from this that STX11 induces molecular shifts within ORAI1 transmembrane helices, and not in its C-terminal tail. This is an important mechanistic inference that needs more discussion. What does this imply about the conformational state of primed ORAI1? And why is straightening of the tails not sufficient for full opening without the correct TM helix arrangement? This distinction has implications for how the STX11-ORAI1 interaction should be modelled and should be engaged with more thoroughly.

      There is no discrepancy here, please also see our response to reviewer 2’s comment #4. The experiment implies that the conformational state of primed Orai1 involves shifts in the TM region of Orai1 and is different from unprimed state. The structural similarities between H134S and ANSGA Orai1 mutants have not been formally established. Unlike H134S, no structure exists for the ANSGA mutant. In the absence of this, it is impossible to comment on whether the two constitutively active mutants are structurally comparable or whether there are multiple ways to stabilize open states of CRAC channel pore, especially when using TM mutants of Orai1.

      The goal of this experiment was to determine what kinds of structural shits STX11 potentially induces in native Orai1. Using previously characterized constitutively active mutants and fusion proteins from the CRAC field, we have ruled out a potential role for STX11 in simply changing the orientation of Orai1 C-terminal tails. A discussion on the topic of why tail straightening of Orai1 is insufficient to open Orai1 is outside the scope of this paper. As pointed by reviewer 2, it is possible that C-term tails already exist pointing towards the cytosol in native, resting Orai1, although this has not been shown in any study using structure of full-length WT Orai1 and is purely speculative at this point. We prefer to not engage in speculative structural insights.

      Other points

      (9) Figures 1G and 1H. The patient-derived mutant STX11 band runs at approximately 37 kDa rather than the predicted 39.5 kDa. The authors suggest instability or reduced antibody reactivity, but premature translation termination is also a possibility that should be acknowledged.

      We have added this point in line 168.

      (10) Figure 2B. The traces and current voltage relationships should be rescaled to show the rectification and inactivation profile of the current in cells depleted of STX11.

      This has been done and modified in new Figure 3 (Figure 2 of version 1).

      While preparing source data files for all figures, we noticed an error in the value of the SE in the STX11-depleted group of old Figure 2C, which has now been corrected. The SE value in the older version was erroneously pasted from an adjacent data column.

      Similarly, in old Figure 3 (version 1), new Figure 4C, we noticed that some data points in the STX11 group were pasted twice in the same excel column. These cells were removed and additional cells were analyzed from the same experiment and added to this group. The overall result remains the same but the distribution of data points looks a bit different.

      (11) Figure 4B: This experiment should be repeated in cells treated with thapsigargin to deplete intracellular calcium stores, and the extent of colocalization quantified by measuring the Pearson's correlation coefficient.

      This has been done. Pearson’s correlation coefficient is included in new Figure 5D.

      (12) Figure 4F. Why is there no detectable band in the input lane of the left blot?

      Western blots show relative intensities of bands of proteins across lanes. A faint band in the input lane of old Figure 4F suggests that the IP/ co-IP/ pull down was robust. If we increase the exposure, the input band would become stronger but the pull-down band would become over-saturated and the difference in the intensities would not be linear. The faint non-specific bands in other lanes represent a fraction of soluble STX11 that tends to crash out of solution over time and gets spun down with the beads. See lines 535-541 explaining this.

      (13) Figure 5. Immunofluorescence data showing the membrane staining of the mutated syntaxin and channel should be included, as well as calcium recordings of cells expressing YFP-CAD with WT and mutated ORAI1.

      In version 2 Figure 6A, we have now also shown co-localization of mutant STX11 with Orai1-YFP in resting and store-depleted cells, in addition to WGA. Pearson’s correlation (not shown) did not show any significant difference in the localization of mutant synatxin 11 w.r.t Orai1. Calcium recordings of CAD-induced constitutive calcium influx from wild-type versus mutant Orai1 are now shown in new Figure 6O-P.

      (14) Figure 6B. A clear colocalization of CFP-O1 and STIM1-YFP is visible on the images, yet the authors conclude from morphometric analysis that the channel is not recruited into ER-PM clusters. Please show the difference in colocalization quantified by measuring the Pearson's correlation coefficient. Pictures should also be provided with the C-terminally tagged construct.

      The quantification of CFP-Orai1 localization inside Stim1-YFP puncta was already shown in old Figure 6E and F. The residence of Orai1 inside STIM1 puncta versus total Orai1 in the PM of STX11-depleted groups was clearly reduced. We have now also shown Pearson’s correlation coefficient for Stim Orai co-localization inside puncta in new Figure 7F.

      TIRF microscopy images of C-terminally tagged Orai1 were already included in Supplementary Figure 10. No defect in co-clustering of C-terminally tagged Orai1-YFP and N-terminally tagged CFP-Stim1 was seen and yet SOCE was inhibited. Therefore, we never concluded from Figure 6 that Orai1 and Stim1 fail to co-localize. We said, they fail to form ‘functional’ clusters. We have now moved the representative TIRF images from supplementary figure 10 to the new main Figure 7G. The Pearson’s correlation coefficient for Stim Orai1 co-localization inside puncta is shown in new Figure 7L.

      (15) Figure 6E and 6F show the same data.

      Figure 6E showed fraction of Orai1 inside Stim1 puncta divided by total Orai1, and 6F showed fraction of Orai1 outside puncta divided by total Orai1. The plots are different but we agree that the data are coming from same cells. We have removed old panel 6F and replaced it with Pearson’s correlation coefficient of Stim1:Orai1 colocalization in puncta in new Figure 7F.

      (16) Figure 7G-L. The difference in constitutive calcium fluxes should be confirmed by Manganese quench recordings. The surface expression of the Orai1 mutants should be shown.

      We have now shown the quantification of surface expression of Orai1 mutants for each respective mutant in the new Figure 8-figure supplement 3B and 3D. The Orai1 mutants we have used in this paper are well established in the literature, they showed clear surface localization and the differences in calcium influx between Scr and STX11 treated cells upon overexpression of Orai1 mutants in HEK are robust. Therefore, we do not see any compelling reason for repeating all of the experiments from Figure 7G to 7L to also show manganese quench recordings, as suggested by the reviewer. We have applied Fura 2 calibration done for these experiments to calculate intracellular calcium. These have been shown in the revised and new Figure 8G-L, where F340/380 ratios of representative calcium assays have also been replaced with the calibrated intracellular calcium concentration.

      (17) Supplementary Figure 12. The recordings show a very large variability between experiments. The different SNAREs that are depleted here could compensate for each other, accounting for this variability. It would be interesting to show the effect of the combined silencing of all the SNARES tested here. The efficiency of the protein depletion should also be documented.

      Genome-wide high- or medium-throughput screens are inherently noisy. None of the genome-wide high- or medium-throughput screens show evidence of protein depletion for each gene in any of the published screens to our knowledge. We chose to only characterize the candidates that reproducibly showed > 70% inhibition of SOCE, others were ignored as noise.

      Silencing of all SNAREs together will definitely lead to loss of morphology and early lethality as all membrane trafficking will be stopped. We never analyze cells that do not show normal morphology and have compromised viability for ablation of SOCE.

      (18) Lines 236-238. The authors note that STX11 harbors a stretch of C-terminal cysteines proposed to be essential for its membrane localization, but do not elaborate on the underlying mechanism. It would strengthen the discussion to explicitly acknowledge that this membrane anchoring is mediated by S-acylation of these cysteines PMID: 24910990 and to connect this to the known enrichment of Orai1 in lipid rafts and the immune synapse PMID 34913437. Both observations are relevant to understanding how STX11 and Orai1 are brought into proximity at the plasma membrane, and their omission leaves an explanatory gap in the proposed interaction model.

      Please see our response to point #2 above. We do not think C-terminal cysteines target STX11 to the PM. We have corrected this claim based on an earlier study, PMID: 24910990, in the revised version of this paper. Analysis of immune synapse and lipid rafts are outside the scope of this paper. The mechanism of PM targeting of STX11 is currently unestablished and will require a systematic and focused mutational analysis which is outside the scope and main focus of this paper.

      (19) Line 351. The statement that syntaxin depletion does not alter the structure or proximity of junctional ER to the plasma membrane is not supported by data. Neither electron microscopy nor TIRF imaging has been performed, which would be required to back up this claim.

      Because Stim1 itself can be used as a marker of ER-PM junctions, this statement was supported by data shown in Figure 6C, D, G, H of version 1 of this paper where the intensity and area of Stim1 clusters was assessed using TIRF microscopy and found to be indistinguishable between STX11 and scramble control cells. The imaging done in Figure 6G, H was TIRF imaging and this was already specified in the legend. We have now also done TIRF imaging of GFP-Mapper-expressing scr and STX11-depleted cells. Mapper is a genetically encoded fluorescent protein that was previously shown to mark ER-PM junctions (10). We found no significant difference in the area or intensity of GFP-Mapper puncta (new Figure 7O-Q), just like Stim1 puncta didn’t show any defect in STX11-depleted cells. Please see modified text from 416-423.

      (20) The molecular dynamics methods need more detail: force field, simulation length, water box dimensions, and convergence criteria should all be specified to allow replication. The supplementary RMSD plots (Supplementary Figure 5B) should also show individual replicate trajectories rather than averages only.

      We had already mentioned the force field (OPLS4) and simulation length (500ns) in the methods section. Also, the RMSD plots in Supplementary Figure 5B already showed individual replicates in version 1.

      We have now updated the methods with following additions:

      The OPLS4 force field was used for all 500 ns simulations in an orthorhombic water box with a buffer distance of 10 Å beyond the solute in each direction. Simulation stability was assessed based on the protein backbone RMSD over simulation time.

      Trajectory clustering was performed using the trajectory clustering tool in Schrödinger, which applies affinity propagation to the pairwise backbone RMSD-based similarity matrix. Within each affinity propagation run, convergence was defined as no change in the set of exemplar frames for 15 consecutive iterations, with a maximum of 400 iterations per run. If convergence was not reached, the damping factor was increased from 0.5 in increments of 0.01 until convergence.

      Reviewer #2 (Recommendations for the authors):

      Overall, this is a timely and impactful study supported by a broad set of methods and cell types. Before publication, the manuscript should address the following points.

      Major:

      (1) The authors note that STX11 contains cysteine residues that enable membrane association. What is the specific mechanism of membrane attachment? Could it occur via S-acylation (palmitoylation)? Both Orai1 and STIM1 are known to undergo S-acylation, which raises the possibility that this modification might also facilitate STX11 membrane anchoring and/or co-residence with Orai1. Is STX11 constitutively membrane-associated, or does it show preferential localization to specific membrane subdomains, particularly in proximity to Orai1?

      STX11 is constitutively membrane-associated and does not show any preferential localization to specific membrane subdomains in confocal images. Figure 4, panel A and B from version1 clearly showed this. In an earlier paper by Hellewell et al. 2014, PMID: 24910990, S-acylation of terminal cysteines of STX-11 was proposed to be crucial for membrane attachment of STX11 and its recruitment to the immune synapse. However, please see our response to reviewer 1’s comment #2 and a new Figure 5E for the localization of the frameshift FHLH4 mutant characterized in this paper. The frameshift mutant that we have characterized lacked all terminal cysteines as well as a short terminal part of the SNARE domain. Cloning and ectopic expression of this mutant still showed constitutive localization to PM and did not show preferential distribution to any specific regions. Therefore, we do not think that terminal cysteines of STX11 contribute to its membrane attachment, we have accordingly modified lines 291-293, 302-305, 540 in the revised version. Also see our response to your point#6 below.

      (2) Is there a possibility to monitor a dynamic change in STX11 co-localization from before to after store-depletion?

      We did not observe any change in the overall distribution of STX11 in cells expressing STX11 alone or co-expressing Orai1 with STX11, pre- or post-store-depletion (please see new figure 5B-C). In cells co-expressing ORAI1, STX11 and STIM1 (see new Figure 5M-N), we could not capture the dynamic segregation of STX11 into regions of PM devoid of STIM:ORAI puncta and therefore have only pre- or post-store-depletion images. Dynamic change in STX11 distribution would require live, multi-colour, high-resolution imaging of diffraction-limited ER-PM junctions and adjacent regions which is technically extremely challenging, and especially due to our inability to tag STX11 with a fluorescent tag without disrupting its localization. Also see our response to your point#6 below.

      (3) The authors use CAD to prove that the interaction with the R289A_E272A_E275A_E278A mutant is normal as for the wild-type. Does this also hold for STIM1 wild-type full-length?

      This is also true for full-length STIM1. The data have now been added to the new Figure 7-figure supplement1.

      (4) The authors state that STIM1 binds both the N- and C-termini of Orai1. While STIM1 binding to the Orai1 C-terminus is well established, the nature of its interaction with the N-terminus remains debated. Fragment-based assays suggest direct binding to the N-terminus; however, direct interaction with full-length Orai1 has not been conclusively demonstrated. This point should be phrased more cautiously to reflect the current uncertainty.

      We have re-phrased the sentence as follows in line 543: “The individual relevance of Orai1 N- versus C-terminus in the trapping versus gating of Orai1 remains unclear”

      (4) In the discussion, the authors report: "Though crucial for trapping and gating, Orai1 tails were missing from early structures of Drosophila Orai [28]". A previous NMR structure suggested that the C-terminal tails of two adjacent Orai1 subunits bend and pair with each other in an antiparallel fashion, and sit closely apposed to PM [37]. However, in recent structures of constitutively active H134 mutant Orai, the C-terminal tails were found to orient away from the membrane [28]. In STX11-depleted cells, switching the CFP-tag from the Orai1 N- to the C-terminus could rescue its clustering but not gating by Stim1. Furthermore, STX11 depletion inhibited the constitutively active ANSGA mutant of Orai1 [29], where the tails of Orai1 are proposed to be constitutively unlatched. These data essentially reinforce our conclusions that STX11 induced molecular shifts encompass Orai1 transmembranes.' However, the information provided here is not fully correct. The early Drosophila Orai structure lacks the full N-terminus but retains most of the C-terminus. It was the X‑ray, not cryo‑EM, structure that suggested an antiparallel arrangement of the Orai1 C-termini. Although the "open" X‑ray structure shows unlatching and straightening of TM4-C-termini, it remains uncertain whether these features reflect physiological gating or crystallization artifacts. It is also unclear whether the Orai1 ANSGA gain‑of‑function mutant adopts a similar unlatching; however, prior work indicates that ANSGA impairs proper coupling to the C‑terminal binding interface (in contrast to Orai1 H134S, which maintains effective coupling). This raises the key question: why do H134S and ANSGA respond differently to STX11 depletion? One possibility is that these mutants stabilize distinct conformations of the TM4-C-termini ("latched" vs "unlatched" states) that differentially dictate the requirement for STX11 in channel assembly or gating. We recommend refining the discussion

      We have changed the word ‘missing’ to ‘truncated’ in line 545 and 546.

      We agree that there is no evidence in literature that establishes similarity between H134S and ANSGA mutation-induced conformations of Orai1. It is, however, implied in most previous studies of mutant Orai1s that there is only one possible open state/conformation. We have added the suggested point and modified the discussion in line 551-556.

      (5) The authors highlight: 'A major problem with this interpretation is that even though amplification of CRAC currents was shown, none of the previous patch clamp studies established whether the higher currents resulted from a greater number of active channels or unchecked conductance per channel by performing single channel recordings.' It should be noted that CRAC channels have extremely low single‑channel conductance, making direct single‑channel recordings challenging. As a result, estimates of open probability and channel number typically rely on fluctuation (noise) analysis rather than direct measurements of single‑channel events (see https://doi.org/10.1085/jgp.200609588). We suggest acknowledging this limitation in the discussion to contextualize the interpretation of gating and channel density.

      We acknowledge how challenging it is to record the single-channel conductance from CRAC channels. We have added this fact to the discussion and the reference that the reviewer has suggested in line 570-572.

      (6) STX11 appears to shift Orai1 localization into puncta. Activated STIM1 is known to engage plasma membrane PIP2 to facilitate Orai1 coupling. How, if at all, is STX11 linked to PIP2 or PIP2-rich microdomains? Is there evidence for direct PIP2 binding by STX11, or for indirect recruitment via PIP2-binding partners? Any available data on STX11's lipid interactions or its enrichment within PIP2-enriched regions would help clarify this mechanism.

      We have not claimed that STX11 shifts Orai1 into puncta. We already showed in old supplementary figure 10 and Figure 6G-J of version1 (v1) of this paper that the C-terminally tagged Orai1-CFP can very well form puncta and co-localize with STIM1 in STX11-depleted cells. To avoid this confusion, we have moved the old Supplementary Figure 10 from v1 to the main figure in revised version, see new Figure 7 panel G. Despite the presence of ORAI1-CFP in puncta with YFP-Stim1, the SOCE was inhibited in STX11-depleted cells. Please also see new Pearson’s correlation coefficient for Stim1 and Orai1 colocalization in Figure 7 panel L. Therefore we concluded that, Orai1 forms ‘nonfunctional’ clusters with Stim1 in STX11 depleted cells, please see modified lines 408-412, clearly explaining this.

      Although syntaxin 1A has been shown to interact with cholesterol (11) as well as PIP2 (7, 8) using either a stretch of polybasic residues or basic residues spread throughout several domains. To our knowledge, these have not been proposed to recruit or segregate STX11 in membranes. STX11 doesn’t contain an obvious stretch of poly-basic residues in its sequence, either, to quickly mutate and address this question. Please also see our response to your point #1 and #2 above. Answering this question will require a systematic and dedicated mutagenesis study.

      Minor:

      (1) Please indicate in Figure 1 in the respective graphs in which cell type the Ca2+ imaging studies have been performed.

      Done.

      (2) Figure 4C, D: Why are the input bands so weak?

      Please see our response to reviewer #1’s similar comment 12 above.

      (3) Figure 5A: Please clarify what WGA is.

      WGA is wheat germ agglutinin which is used to mark PM in imaging experiments. It binds to N-acetyl-D-glucosamine and sialic acid residues found in mammalian cell membranes and glycoproteins. We have added the explanation to the new Figure 6A legend.

      The authors state that "the constitutively active ANSGA (261-265) mutant of Orai1 (Supplementary Figure 11G), which harbors 4 consecutive mutations in the Orai1 C-terminus ...". Please clearly state this is the nexus region connecting the C-terminus with TM4. The 5 aa stretch is not the C-terminus; it is just close to the C-terminus.

      We have modified this, as suggested, in line 470-471.

      References:

      (1) Miao Y, Miner C, Zhang L, Hanson PI, Dani A, Vig M. An essential and NSF independent role for alpha-SNAP in store-operated calcium entry. Elife. 2013;2:e00802.

      (2) Li P, Miao Y, Dani A, Vig M. alpha-SNAP regulates dynamic, on-site assembly and calcium selectivity of Orai1 channels. Mol Biol Cell. 2016;27(16):2542-53.

      (3) Chorev DS, Baker LA, Wu D, Beilsten-Edmands V, Rouse SL, Zeev-Ben-Mordehai T, et al. Protein assemblies ejected directly from native membranes yield complexes for mass spectrometry. Science. 2018;362(6416):829-34.

      (4) Dorwart MR, Wray R, Brautigam CA, Jiang Y, Blount P. S. aureus MscL is a pentamer in vivo but of variable stoichiometries in vitro: implications for detergent-solubilized membrane proteins. PLoS Biol. 2010;8(12):e1000555.

      (5) Vig M, Peinelt C, Beck A, Koomoa DL, Rabah D, Koblan-Huberson M, et al. CRACM1 is a plasma membrane protein essential for store-operated Ca2+ entry. Science. 2006;312(5777):1220-3.

      (6) Feske S, Gwack Y, Prakriya M, Srikanth S, Puppel SH, Tanasa B, et al. A mutation in Orai1 causes immune deficiency by abrogating CRAC channel function. Nature. 2006;441(7090):179-85.

      (7) Murray DH, Tamm LK. Clustering of syntaxin-1A in model membranes is modulated by phosphatidylinositol 4,5-bisphosphate and cholesterol. Biochemistry. 2009;48(21):4617-25.

      (8) van den Bogaart G, Meyenberg K, Risselada HJ, Amin H, Willig KI, Hubrich BE, et al. Membrane protein sequestering by ionic protein-lipid interactions. Nature. 2011;479(7374):552-5.

      (9) Miranda P, Contreras JE, Plested AJ, Sigworth FJ, Holmgren M, Giraldez T. State-dependent FRET reports calcium- and voltage-dependent gating-ring motions in BK channels. Proc Natl Acad Sci U S A. 2013;110(13):5217-22.

      (10) Chang CL, Chen YJ, Liou J. ER-plasma membrane junctions: Why and how do we study them? Biochim Biophys Acta Mol Cell Res. 2017;1864(9):1494-506.

      (11) Lang T, Bruns D, Wenzel D, Riedel D, Holroyd P, Thiele C, et al. SNAREs are concentrated in cholesterol-dependent clusters that define docking and fusion sites for exocytosis. EMBO J. 2001;20(9):2202-13.

    1. Author response:

      eLife Assessment

      This valuable study provides insights into the role of steroid signaling during tumorigenesis in the adult male drosophila accessory gland (functional equivalent of the prostate gland in mammals), hinting at a possible counterintuitive anti-tumoral role of sex hormones during prostate cancer in certain patients. While the Drosophila model provides an elegant way to study the hypothesis derived from the Cancer Atlas analysis, the analyses of public prostate cancer expression data are incomplete and critical knowledge on patients' treatment modalities and normalization across different datasets is missing. This work would be of interest to prostate cancer researchers as it suggests that the absence of androgen receptor signaling in humans could constitute a mechanism promoting tumor escape.

      We thank the reviewers for the time they have spent on the manuscript, the production of a public review and their useful recommendations. As a general goal for the corrected version, we will try to provide more insight on the data (especially the human data), and more controls, to strengthen our conclusions. We are aware of the lack of a definitive proof of the role of the apparent decrease in AR signaling on tumour progression, but hope that this manuscript will encourage medical scientists to test/challenge its counterintuitive results in large cohorts of tissues and mouse/human models.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this article, Vialat and his colleagues examine the early stages - which remain largely unknown - of the tumor escape process, particularly the basal extrusion of tumor cells following endocrine therapy for prostate cancer.

      They first used the "Prostate Cancer Atlas" database, which provides access to a vast amount of transcriptomic data, to perform high-throughput analyses. Interestingly, analyzing a series of Androgen Receptor (AR) target genes in castration-resistant prostate cancers, they concluded that the loss of the canonical AR signaling pathway may contribute to tumor resistance.

      Using a well-established model in Drosophila, they then replicated in vivo an endocrine therapy targeting the accessory gland by genetically inhibiting the expression of ecdysone, the only sex steroid present in Drosophila. These experiments induced basal extrusion similar to the mechanism observed in tumor escape in humans.

      These results suggest that the deprivation of sex steroids may play an important role in tumor progression.

      However, although the data from the "Prostate Cancer Atlas" constitutes a powerful tool that serves as the basis for this new concept, clinical validation using carefully selected human tumor samples would help strengthen the authors' conclusions.

      Strengths:

      (1) The Prostate Cancer Atlas is a comprehensive collection of clinical data derived from RNA sequencing and serves as a powerful tool for conducting high-throughput analyses in this paper.

      (2) The Drosophila model used in this article is well established and has already been the subject of publications by the team. In addition to being an in vivo model, Drosophila offers a threefold advantage for this study: i) the presence of an accessory gland, similar to the prostate, which allows for the simulation of tumor formation and, in particular, extrusion mechanisms; ii) its regulation by a single sex steroid, ecdysone; iii) the genetic ability to modulate or inactivate ecdysone expression, which allows for a parallel to be drawn with hormonal deprivation in humans.

      (3) This study presents interesting and original findings. The data are, for the most part, of high quality.

      Weaknesses:

      (1) The Prostate Cancer Atlas, which is an essential tool in this study, was described only briefly - if at all - in both the introduction and the "Materials and Methods" section. The selection criteria used to distinguish CRPC or NEPC from adenocarcinoma in the Atlas or as determined by the authors, as well as the analytical methods, were not specified. It is therefore difficult to be convinced by the results, particularly those presented in Figures 1 and 2.

      As the tool has been published in different articles, we chose to limit its description. However, we agree that explanation are necessary, that will be added in the new version. First, we initially used here just basic categories, in order to avoid any possible bias; so mCRPC includes rare DNPC and NECP patients. We will also put the data with true ARPC, with essentially the same results.

      For human data, we also expect to use transcriptomic data from an independent cohort to check whether the same loss of AR signaling occurs during progression. Furthermore, we consider to add the data showing that decrease in canonical AR signaling also (logically with the previous results) correlates with castration status or ADT exposure (these info are available on ProstateCancerAtlas). Interestingly, and this can be put in supplementary data, prostatecanceratlas detects changes in EMT genes or proliferation genes that are coherent with what is known about cancer progression, indicating that the apparent decrease in AR signaling should correspond to a real phenomenon.

      (2) Although the hypothesis put forward by the authors - that the deprivation of sex hormones contributes to tumor progression - is strongly supported by the Drosophila model and by the in silico analysis of transcriptomic data from the Atlas, this concept still needs to be clinically validated by analyzing a series of prostate cancer samples, either through transcriptomic analysis or by tracking gene expression in histological sections.

      There are many indirect evidences that loss of AR signaling induces tumor progression in mouse (as stated in the intro or the discussion of the manuscript). However, as suggested in the introduction of the letter, we believe that medical scientists are the most qualified to prove that sex steroid deprivation indeed induces tumor progression in human. We will add in any case data to at least reinforce this puzzling finding of a decrease in AR canonical signaling during progression.

      (3) With regard to the cells responsible for tumor escape, stem cells have been described as "candidates for the initiating resistant tumor growth" (lanes 50-55), but it is also essential to address the recent concept of "persistent cells". Indeed, these cells have been primarily associated with their tolerance to treatment (chemotherapy) and are referred to as "drug-tolerant cells". However, persistent cells could also correspond to cells that evade hormone therapy in the case of prostate cancer. This possibility should be discussed in the article.

      This is an interesting suggestion, which can be discussed: on the one hand, intrabasal cells may not have accumulated mutations to survive the loss of EcR signaling, as would do persistent cells. On the other hand, they strongly proliferate, and show no sign of senescence, behaving more like resistant cells. So, it does not look to us that we induced the appearance of persistent cells in the Drosophila accessory gland, except if these cells are quickly reactivating to give rise to intrabasal cells.

      Reviewer #2 (Public review):

      Summary:

      In this study, Vialat and collaborators study the role of steroid hormone signalling on the development of prostate cancer (patients) and of accessory gland tumours in Drosophila, a tissue functionally equivalent to the prostate. Mining publicly available prostate cancer expression data and using gene expression signatures, they uncover that androgen signalling is actually down-regulated in castration resistant prostate cancers (CRPC) compared to "primary" cancers, leading the authors to wonder whether down-regulation of canonical androgen signalling could represent an important event increasing tumour aggressiveness. They then take advantage of their recently published tumour model in the accessory gland of Drosophila adult males, in which cells are primed for tumorigenesis by the constitutive activation of the EGFR receptor, to test directly this hypothesis. They show that the genetic invalidation of ecdysone reception and signalling increases the aggressiveness of the "pre-cancerous" lesions, and that ecdysone-insensitive tumours present higher proliferation and initiate basal extrusion.

      Strengths:

      The authors bring original observations on the role of ecdysone signalling to prevent male accessory gland tumour development in Drosophila

      Weaknesses:

      (1) The link between the human data mining and Drosophila model is not straightforward.

      (2) Important information, in particular clinical information, is missing in the presentation of the cancer patients' data, making it complicated to grasp the solidity of the claims.

      (3) Data-mining insights should be validated by orthogonal approaches.

      (4) Ecdysone signalling activity should be monitored.

      While the two parts of the study both investigate the role of steroid signalling on tumour growth, the link remains slightly artificial. I think starting with Drosophila and then opening with some patient data would be better suited to the level of proof reached here, implying that the anti-tumoral role of steroids observed experimentally in the fly might be conserved based on data mining in patients, rather than trying to prove in the fly the hints gained from public data mining. Indeed, there are many important differences between the mammalian prostate and the fly accessory gland, as well as between sex hormone androgen signalling and developmental timing ecdysone signalling.

      This is an interesting suggestion. Actually, we first wrote the manuscript by starting with Drosophila data and then going to patients data, and previous reviewers said that this was not possible to directly go from Drosophila to human. So, we suppose that the real way to solve this will be by the validation or refutation of the data by other teams in different models.

      The prostate cancer data mining and re-evaluation brings some interesting observations that appear to challenge the androgen driver, contrary to the vast amount of literature. Indeed, the authors observe an apparent decrease in androgen signalling in the more advanced states of the disease, in particular CRPC. In order to better evaluate its clinical relevance, more background on the tumours analysed should be provided.

      This point is also important to Reviewer 1, and will be implemented.

      What treatments were received by the patients? Hormonotherapy? LH/RH analogues? +/- anti-androgens? Are these treatments still given when CRPC emerge and tissues were banked? Metastatic disease? Are these only primary tumours in situ? Are there metastases included in the analyses?

      Most CRPC come from metastatic sites. Most of the CRPC were treated by ADT. We chose to have an approach that included all the samples; but we will provide insight, whenever available, on these absolutely relevant questions.

      Frequently, castration resistance is associated with alternatively spliced variants of the AR (AR-V7) that become constitutive and could bind to new AR-sensitive enhancers, even in the absence of androgen. Is the splice variant status of patients known, or could it be inferred from the expression data? Would there be different responses according to AR-V7 status?

      As a first approach, from the cohorts that were used, it seems that in the PCA patients, there are around or less than 15% of patients harboring the AR-V7 driver. It could be of interest to test their behavior regarding the same set of genes, and it will be done if we can identify the patients.

      Regarding the signature used. Why not monitor PSMA, one of the major prostate cancer markers, which is regulated by AR?

      It will be done. PSMA behaves as the others, even though the drop between primary samples and ARPC samples is very limited and just statistically significant.

      Finally, to consolidate the surprising observation that AR signalling is repressed in CRPCs, the authors should back these in silico predictions with orthogonal approaches such as histochemistry on patients' TMA or tissues from mouse models, monitoring AR activity.

      As said previously, we believe that this specific work will be better done by medical scientists.

      Regarding the fly experiments, the observation that ecdysone signalling depletion cooperates with EGFR-lambda activation to generate big overgrowths that delaminate basally without passing through the muscular sheet is interesting. However, several important controls need to be provided in order to support the claims:

      Considering the comments regarding the fly experiments, we agree that, if experiences are taken individually, controls are lacking. However, we have to explain our strategy and why the results taken in their entirety have a significance. In our model of epithelial tumorigenesis, we have started to explore the EcR pathway after years of work on other pathways. At the first experiment (with the EcR RNAi line), we were struck by the intrabasal phenotype that did not occurred in our previous experiments, and especially for the 14 RNAi lines that we published in two independent articles on Ras/MAPK, Pi3K/Akt pathways and cholesterol metabolism. As justly said in the review, many unexpected effects can happen, so we decided to explore the role of five other genes of the same pathway to be sure of the reproducibility of the phenotype when we block the EcR pathway. The odds of having the same specific phenotype for 6 lines of the EcR pathway when there is always another phenotype for 14 lines targeting other pathways can be calculated: p = 0.00000494. So, the best control we offer, and it is largely significant, is the repetition of the experiments intended to downregulate the EcR pathway, that produce the same phenotypes independently of the target.

      Furthermore, all lines we used were previously tested, validated and most of the time published in other scientific works. This is essential to us, as one complexity of doing rare clones in an otherwise normal tissue is that decreasing an mRNA in less than 5% of the cells of course difficultly leads to a detectable drop in overall expression in the whole gland. This is also the reason why we always tried pairs of fly lines to block the receptor activity itself (RNAi EcR, RNAi Shd), the receptor's downstream targets (RNAi HR3, RNAi HR4), and the production of ecdysone (RNAi Sad, RNAi Phtm). As we validated RNAi Sad, we can try anyway to validate at least another RNAi of another category. Furthermore, we did use a RNAi White control: it behaves in the same way as the GFP control. We will put the results comparing the two lines in supplementary data.

      (a) The authors should use an ecdysone reporter (ERE-LacZ, ERE-GFP...) to monitor and show that Ecdysone signalling is indeed lower in the tumours after genetic manipulations, or that it is higher in EGFR-lambda small clones.

      This would be of interest to validate that the 6 lines are behaving in the same way (at least, they give similar phenotypes). However, in the adult accessory gland, ERE activity is largely lower than during development (DOI: 10.1016/j.jinsphys.2011.03.027), and to be able to decrease it, authors had to express notoriously strong dominant-negative EcR-DN. We can try the experiment but are really not persuaded that we will be able to see a drop of activity with only a decrease of expression of the gene. If we can think of another solution that could be more efficient, we will try it as the idea is of course interesting.

      (b) EcR is normally a repressor, which is turned into an activator in the presence of 20-hydroxyecdysone. The removal of EcR could lead to de-repression of genes and thus slightly activate the pathway. Monitoring ecdysone signalling activity is thus critical.

      Actually, there are different EcR isoforms. EcR-B1 is generally considered as the main activator of the pathway, as EcR-A is a repressor of the pathway. The EcR RNAi line which was used does not target a specific isoform.

      (c) The authors should also monitor the expression of Phantom, Shadow, Shade, and EcR in the different accessory glands (wild-type, EGFR-lambda, EGFR-lambda & EcR-RNAi). It is extremely surprising that systemic ecdysone has so little role since Phantom, Shadow, or Shade RNAi appear as potent as EcR-RNAi. This quantification has actually been performed for Sad in Figure S5, which is not even mentioned in the text. It should be done for Phtm.

      The levels of ecdysone are tenths of times lower in adult compared to the peaks during embryogenesis or metamorphosis. And one source of production is the epithelial cells of the accessory glands themselves. Considering that EcR is expressed in all the cells of the accessory gland (epithelial cells and muscle cells), it seems plausible that there is only a very little amount of ecdysone that can in fact be available for the other epithelial cells.

      A UAS-yellow-RNAi (or similarly irrelevant RNAi) rather than UAS-GFP should be used as a control for the EcR, Phtm, Sad, Shd, and Tub RNAi. Indeed, loading the RNAi machinery could have some unexpected effects not controlled by the UAS-GFP.

      The NLSGFP line we used here is the one we already published twice (and we compared it to RNAi lines), and this is the reason why we used this already validated control. However, we tested a RNAi White line, and it behaves in the same way.

      The authors should not use the term "sex steroid" when referring to ecdysone. It is a steroid hormone important for developmental timing and rate of growth, but is not a sex hormone, as sex is cell autonomously genetically determined in the fly.

      In human, sex hormones control sexual differentiation (up to adult characteristics) and sexual reproduction. In Drosophila, ecdysone controls sexual reproduction in both sexes and sexual differentiation at least in the female (DOI: 10.1007/s004270050186). From these results, we do not think that saying it is a sex steroid (not a sex hormone) is ill suited. We intend to precise what we put in this term in the introduction to avoid overinterpretation from our part.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript the applicants study two residues in the GHKL ATPase active site of Aq MutL and GyrB, and argue that the catalytic base function is shared between two conserved acidic residues that are 3 residues apart.

      In the manuscript, they generated mutant versions in MutL and GyrB (both ala and the appropriate Asn/Gln version) and performed ATPase analysis. They also generated high resolution crystal structures of the GyrB NTD with AMPPnP for WT and mutants of the two acidic residues. The data show that mutation in either of these residues does not fully kill activity (with the exception of the Alanine mutation of the first of the two, that interferes with ATP (or AMPPnP) binding). When the acidic residues are mutated to Asn/Gln, the catalytic water can still be positioned, and hence these mutants are more active than the Ala mutants. In both cases the double mutation is catalytic dead.

      The authors then perform phylogenetic analysis and ancestral gene reconstruction and based on this they argue that HSP90 forms a different class of GHKL ATPases, and lost rather than gained this separate status.

      Strengths:

      The biochemical analysis seems solid.

      Weaknesses:

      A major question that remains, is why the mutations have so much more detrimental effect in MutL (100-fold lower kcat/KM) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      The authors need to discuss this issue explicitly to make it clear that conservation of the mechanism is not complete and that other interpretations are possible.

      The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the the acidic residues and their mutants are not shown.

      This has been addressed.

      There are some issues with figure S2B and S5.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role of two conserved acidic residues rather than one. The authors have used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism.

      Through a detailed re-analysis of their previously published structure of the aqMutL NTD (ATPase domain) in complex with AMPPCP, they identified Glu29 and Glu32 as interacting with nucleophilic water for the catalysis. The authors carefully dissected the respective roles of these two acidic residues with a series of site-directed mutations. Mutations at Glu29 impaired ATPase activity without affecting protein secondary structure or ATP binding in the case of the E29Q mutant. Moreover, mutations at Glu32 did not affect secondary structure (except for E32G) but reduce ATPase activity. Activity was abolished when both residues (E29Q/E32Q) are mutated.

      The authors extended their study to another GHKL ATPase, aqGyrB. Their findings further supported the cooperative function of the corresponding acidic residues in aqGyrB (Glu48 and Asp51) during ATP hydrolysis. Mutation of these residues partially impaired ATP hydrolysis without affecting protein secondary structure. ATPase activity was completely lost in the double mutant E48Q/D51M. While the E48Q mutant retained the ability to bind ATP, the E48A mutant did not. High-resolution structures of the WT and E48A, E48Q, D51A and D51N mutants of the aqGyrB NTD demonstrated that nucleophilic water positioning depended on these residues. E48 played a dominant role in water positioning and is critical for stabilising ATP lid formation and associated conformational changes, whereas D51 contributed cooperatively to catalysis.

      The authors investigated the functional impact of mutating the corresponding residues in the human MutL homologs PMS2 and MLH1. Clinical variants consistently exhibited reduced or abolished ATPase activity, providing a potential molecular basis for Lynch syndrome, through impaired DNA mismatch repair.

      Lastly, through evolutionary analysis, the authors inferred that the second acidic residue was likely present in the common ancestor of MutL, GyrB, and MORC proteins, but was lost in the case of Hsp90.

      Strengths:

      (1) This study contains a detailed structural and biochemical analysis of a biologically important set of GHKL ATPases. The authors identify a second acidic residue that is conserved and contributes to catalysis in a large subset of GHKL ATPases. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed, which involves cooperative and partially overlapping roles for the catalytic residue pair. This revised mechanistic model is invaluable for the interpretation of clinical variants of GHKL ATPases such as PMS2 and MLH1.

      (2) The work described was performed to an excellent and rigorous technical standard. The structural and biochemical data are sound. The evidence supporting the claims is compelling.

      Weaknesses:

      (1) The identification in this study of a second acidic residue contributing to catalysis but not absolutely essential for catalysis is a useful finding. However, given that many structures of GHLK ATPases have been determined with different nucleotide analogs bound and that the essential role of the first acidic residue is well established, the importance and scope of the advances described here remain focused within the field of study of GHKL ATPases.

      (2) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of any effect of disease-linked mutations in GHKL ATPases would have strengthened this study.

      (3) The effect of other aqMutL NTD E32 mutants, particularly, the E32K mutant on ATP binding remains unclear, although experimental assessment of nucleotide binding would be challenging due to the high protein concentrations required for the equilibrium dialysis assay.

      We are grateful to the Editors and reviewers for their careful assessment of our revised manuscript and for identifying the remaining points that required clarification. We have addressed each of these comments in the present revision. We believe that these revisions have resolved the remaining concerns and have further improved the clarity and accuracy of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please discuss the large difference in the effect of the mutants on activity explicitly

      According to the reviewer’s suggestion, we have added the following discussion to the revised manuscript:

      “Although mutation of the two acidic residues impaired the ATPase activity in both aqMutL and aqGyrB, the magnitude of the effects differed substantially, with much greater reduction in the catalytic efficiency in aqMutL than in aqGyrB (Table 1). The molecular basis for this quantitative difference is currently unclear. One possible explanation is that subtle differences in the active-site architecture and surrounding residues alter the relative contribution of each acidic residue to catalysis, allowing aqGyrB to tolerate perturbation of either residue more effectively than aqMutL.” (p.6 line 257-262 in the revised manuscript)

      (2) Figure S2B is a completely different view from the other panels, please provide the correct one.

      We have revised Supplementary Fig. S2B so that the E48A structure is now shown from a viewpoint as similar as possible to those used in the other panels. We note, however, that the E48A structure cannot appear completely identical to the other panels because the E48A mutant does not bind the ATP analog and therefore does not undergo the nucleotide-binding-associated conformational changes observed in the other structures.

      (3) S5 : It is not clear to me what is meant by " The scale bar indicates the number of amino acid substitutions per site." : there is a 'Tree Scale 1" but no other numbers in my version.

      We thank the reviewer for pointing out that the scale bar in Supplementary Fig. S5 was insufficiently explained. The value “1” in the tree scale corresponds to a branch length of one amino acid substitution per site. To avoid ambiguity, we have revised the scale-bar label in Supplementary Fig. S5 to explicitly indicate “1 substitution/site” and have clarified its meaning in the figure legend:

      “Branch lengths are proportional to the evolutionary distances inferred by IQ-TREE. The scale bar represents an evolutionary distance of one amino acid substitution per site.” (p. 22 line 767-769 in the revised manuscript)

      Reviewer #2 (Recommendations for the authors):

      (1) P. 9, in the "Data Accessibility Statement", all three PDB codes (23UX, 23UY, and 23UZ) should be listed.

      We thank the reviewer for this comment. We carefully rechecked the Data Accessibility Statement and confirmed that all three PDB accession codes (23UX, 23UY, and 23UZ) are included in the statement.

      (2) Supplementary Figures S2 and S3. The authors have written "Asn33" and "Asn52", instead of "Glu32" and "Asp51" in both the figure and figure legend of Supplementary Figure S3. They have also written "TND" instead of "NTD" in the figure legend.

      “Asn33” and “Asn52” in Supplementary Fig. S3 are not typographical errors. Asn33 in aqMutL and Asn52 in aqGyrB are the residues that directly coordinate the Mg<sup>2+</sup> ion and are distinct from the acidic residues discussed in this paper. To avoid confusion, we have added the following sentence to the legend of Supplementary Fig. S3:

      “These Mg<sup>2+</sup>-coordinating asparagine residues are adjacent to, but distinct from, the second acidic residues Glu32 in aqMutL and Asp51 in aqGyrB examined in this study.” (p. 21 line 749-751 in the revised manuscript)

      We have also corrected the typographical error “TND” to “NTD” in the figure legend. (p. 21 line 749 in the revised manuscript)

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Bajohr and colleagues propose a transcription factor-driven approach to generating bonafide oligodendrocyte lineage cells (OLCs) from primary mouse astrocytes. Ectopic expression of Olig2, Sox10, or Nkx6.2 in isolated astrocytes produced a range of OLC-like cell states, with Sox10 emerging from lineage tracing and single cell RNA sequencing experiments as the most successful transcription factor in driving direct lineage reprogramming. The authors strengthened their claims with an unbiased, deep learning perturbation model to predict genetic drivers of the astrocyte cluster to OLC cluster transition observed in their scRNA seq dataset. Here, Sox10 surfaced in the top ten correlated genes, and the top transcription factor, mediating this fate shift. Altogether, this paper presents an interesting approach to generate OLCs, a cell type historically difficult to procure, from primary mouse astrocytes to study this lineage in development and disease and perhaps repopulate it in dysmyelinating conditions. While this certainly addresses a technical gap in the field, authors defined iOLCs as ones with lineage-specific gene expression and morphological characteristics, lacking any functional analysis to assess the reprogrammed cells' capacity to myelinate. This comment and other critiques are discussed below.

      While Sox10 and Mbp expression in iOLCs, as confirmed by IHC, is a promising result suggesting that ectopic Sox10 instructs transduced cells to develop into cells of myelinating potential, functional confirmation is essential. As mentioned in the discussion, the absence of a substrate for myelination may have also contributed to the low DLR efficiency. Co-culturing Sox10 iOLCs with primary neurons and examining the cells' potential to engage and enwrap axons would greatly strengthen the authors' claim that this could be an effective therapeutic approach to myelin regeneration in vivo, or even a technical approach to studying myelin dynamics in vitro.

      In Figure 1B, it appears that Mbp expression in tdTomato+ cells decreases in Sox10 transduced iOLs during the observed time period. Can the authors elaborate on this result, given that MBP expression is crucial for myelination and should, if anything, increase with time?

      The authors acknowledge that there is a conversion of tdTomato- zsGreen+ cells with an astrocyte-like morphology to OLC cells expressing Mbp following Sox10 induction (Supplementary figure 5C,D). While they note the diversity of the astrocyte lineage in the discussion, further analysis should be applied to this subset of cells to confirm the subset of astrocyte or progenitor-like cell type that gives rise to their cell endpoint of interest (Sox10-driven Mbp+ iOLs).

      Finally, ectopic expression of Olig2 and Sox10 in primary astrocytes resulted in very different OLC subtypes, as evidenced by OLC marker expression seen in IHC and the subclustering of these cell types in scRNA seq. Although this diversity in OLC type and generation efficiency follows with previous reports showing that these two transcription factors vary in effect, might the authors further discuss this discrepancy given that the two transcription factors regulate one another (as mentioned in the introduction) and should theoretically give rise to more similar cells? Perhaps due to the lower specificity of Olig2 in marking a pure OLC population relative to Sox10?

      We thank the editor and reviewers for their additional comments, which have significantly improved our manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors use Aldh1l1+ astrocyte with a GFAP promoter linked to TF expression, claiming that the converting cells are cortical astrocytes. However, during early mouse development, radial glia that can give rise to multiple cell types, including oligodendrocytes, also express Aldh1l1 and low levels of GFAP. Therefore, it is not proven whether the resulting iOLs came from mature astrocyte or a radial glia population. Especially since in Figure 3B,G it is shown that a majority of the D0 population expresses high amounts of Vimentin and Nestin, both markers associated with radial glia and immature astrocytes. It would be beneficial for the authors to confirm that either there are no contaminating radial glia or that the radial glia don't express the TF, especially since the iOL population is so small.

      Thank you for this comment. We agree that previous studies have demonstrated Aldh1l1 expression in radial glia cells [1]. However, our dissection protocol to obtain the postnatal astrocytes is cortex-specific and does not take portions of the VZ/SVZ, preventing radial glia contamination.

      Nevertheless, to confirm the astrocytic identity of our Aldh1l1+ starting cells we used AUCell enrichment scoring [2]. First, Aldh1l1+ cells were subclustered from our starting culture single cell dataset (Author response image 1A). We then defined two gene signature modules: an “astrocyte” module, comprised of canonical astrocyte markers (Aqp4, Gja1, Slc1a2, Glul, Aldoc, S100b, Nfia, Thbs1, Cst3, Clu), and a “radial glia” module, containing common radial glia and progenitor markers (Pax6, Fabp7, Sox2, Hes1, Prom1, Top2a, Mki67, Ube2c, Cdk6, Mcm2). When we scored all Aldh1l1+ cells (n=869) for enrichment of each signature, no cells were classified as radial glia (Author response image 1B). Instead, the astrocyte signature predominated, with 68.3% classified as astrocytes (Author response image 1B,C). The remaining cells (36.7%), were classified as transitional, reflecting substantial expression of genes from both modules (Author response image 1B,C). Therefore, although the Aldh1l1 astrocytes do express genes common to radial glia, there are no cells that express only progenitor markers. This is consistent with literature showing that many radial glia genes are commonly found in astrocytes [3], [4], [5].

      Taken together, our stringent dissection protocol and bioinformatic profiling of our starting Aldh1l1 cells suggests that the resulting Aldh1l1+iOLs are originating from astrocytes, rather than radial glia.

      In Figure 2D, Sox10 and Nkx6.2 have an n=4 while the Cre control has an n=3. Why is this the case? Was the 4th point excluded? Similar attention should be given to other panels in Figure 2 for consistency.

      We thank the reviewer for highlighting this. No outliers or datapoints were excluded from the analysis. Rather, the fourth culture for our control treated cells was not viable for analysis.

      Author response image 1.

      Astrocyte gene expression in Aldh1l1+ cells. (A) UMAP clustering of Aldh1l1+ cells in our starting cultures (0DPT, non-transduced). (B) UMAP clustering from (A) overlayed with cell classification based on AUCell gene signature expression scoring. (C) Feature plots showing AUCell enrichment scores (darker purple indicates higher enrichment) for the astrocyte signature (left), radial glia signature (middle), and the differential enrichment score (right, Astro_AUCell - RG_AUCell) (darker purple scores indicate higher astrocyte signature and negative scores (gray) indicate higher radial glia signature).

      Representative images in Figure 2 do not convincingly support the argument by the authors. It appears that some of the cells highlighted by the arrows are just background (e.g. PDGFRa and td Tomato in Figure 2E, or zsGreen in Figure 2F). Additionally, the authors should show a different representative image depicting astrocyte morphology in Figure 2G 7DPT.

      Thank you to the reviewer for this comment. We have replaced the images in Figure 2E,F to better represent our findings (updated manuscript Figure 2E,F). We have also adjusted the representative image in Figure 2G 7DPT to better visualize the astrocyte morphology (updated manuscript Figure 2G) as well as included as supplementary additional examples of pre-conversion astrocyte morphology to supplement our morphology analysis (Author response image 2).

      Author response image 2.

      Lineage tracing confirms true conversion of astrocytes to oligodendrocyte lineage cells. Representative images of astrocyte morphology observed prior to cell conversion (arrow indicates converting cells, scale bar =50um).

      References

      (1) L. C. Foo and J. D. Dougherty, “Aldh1L1 is expressed by postnatal neural stem cells in vivo,” Glia, vol. 61, no. 9, pp. 1533–1541, Sep. 2013, doi: 10.1002/glia.22539.

      (2) S. Aibar et al., “SCENIC: Single-cell regulatory network inference and clustering,” Nat Methods, vol. 14, no. 11, pp. 1083–1086, Nov. 2017, doi: 10.1038/nmeth.4463.

      (3) M. Götz and Y.-A. Barde, “Radial Glial Cells: Defined and MajorIntermediates between EmbryonicStem Cells and CNS Neurons,” Neuron, vol. 46, no. 3, pp. 369– 372, May 2005, doi: 10.1016/j.neuron.2005.04.012.

      (4) P. Malatesta, I. Appolloni, and F. Calzolari, “Radial glia and neural stem cells,” Cell and Tissue Research, vol. 331, no. 1, pp. 165–178, 2008, doi: 10.1007/s00441-0070481-8.

      (5) S. Clavreul, L. Dumas, and K. Loulier, “Astrocyte development in the cerebral cortex: Complexity of their origin, genesis, and maturation,” Front Neurosci, vol. 16, p. 916055, Sep. 2022, doi: 10.3389/fnins.2022.916055.

    1. Author response:

      eLife Assessment

      This study presents a useful database resource containing protein conformations generated through molecular dynamics simulations, with extensive quality evaluation and benchmarking. While the database is well-constructed and professionally organized, the evidence supporting its claimed representation of protein conformational landscapes is incomplete, as the short simulation times and starting structure bias prevent true Boltzmann sampling of the conformational space.

      We thank the editors for recognizing the usefulness of ProteinConformers and the value of its quality evaluation and benchmarking. We will revise the manuscript to clarify that ProteinConformers provides large-scale, energetically profiled descriptions of protein conformational landscapes, with broad coverage of locally stereochemically valid and energetic compatible structures from non-native to near-native regions, rather than a complete equilibrium sampling of all conformational states. These revisions will better define the scope of the resource while preserving its intended use for benchmarking and data-driven studies of protein conformational variability.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a new database that rigorously explores protein conformations.

      Strengths:

      It is extremely well done, using state-of-the-art tools by a group at the top of the field of structural modeling. The evaluation of qualities and the benchmarking of the structures are outstanding, and it is expected that the new database will have a significant impact on the field.

      We thank Reviewer #1 for the positive evaluation of our work and for recognizing the potential impact of the ProteinConformers resource.

      Weaknesses:

      The authors are using MD simulation to generate some of the structure, and therefore should have access to standard MD energies. I am surprised that no evaluation is provided based on these energies that can be extended to free energies.

      We thank the reviewer for this helpful suggestion. We reprocessed the original MD energy files and extracted three MD-derived energy terms, including total energy, system potential energy, and protein-only potential energy. These energy terms have been added to the ProteinConformers resource and the web portal. We will update the manuscript to describe these additional energetic annotations.

      Reviewer #2 (Public review):

      Summary:

      The authors developed a dataset of protein conformations by running molecular dynamics simulations starting from both native and decoy conformations for a large number of proteins. These conformations were put together as a dataset for querying and downloading, along with their energies under different force fields. The authors suggest that such conformations represent the proteins' conformational landscape, so that they will be useful for evaluating methods generating multiple conformations of proteins.

      Strengths:

      The dataset is online and working. It has good documentation for others to use.

      We appreciate Reviewer #2’s positive assessment of the online resource and documentation.

      Weaknesses:

      The biggest weakness is that the collected conformations very likely do not represent the true conformational landscape. To represent the conformational landscape, the structures need to be sampled based on the Boltzmann distribution. However, in this study, conformations are generated by running very short (125ps to 375ps) MD simulations starting from near-native conformations and decoys. Such short simulations will produce small fluctuations around the starting conformations, so the distribution of conformations is largely dominated by the distribution of the initial conformations, which by one means are Boltzmann distributed. A conformation might be physically plausible, but it might have very small weight in the Boltzmann distribution. On the other hand, conformations with large weights might not be in the dataset.

      We thank the reviewer for this important and constructive comment. We agree that the conformations in ProteinConformers should not be interpreted as an equilibrium ensemble sampled according to the Boltzmann distribution. Because the MD simulations used here are short, the resulting snapshots around each seed mainly reflect local relaxation and limited thermal fluctuation from that seed, rather than exhaustive equilibrium sampling. Therefore, the relative population of conformations in our dataset should not be interpreted as a Boltzmann weight, and some thermodynamically important states may be underrepresented or absent.

      Our goal in this work is different from conventional long-timescale MD studies that aim to estimate equilibrium populations from one or a few initial structures. ProteinConformers was designed as a large-scale, multi-seed, MD-refined conformer resource. The broad structural coverage comes primarily from initiating simulations from many diverse seed decoys for each protein, while the short all-atom MD protocol is used to relax structures under a molecular mechanics force field, remove structures that fail to converge, reduce steric clashes and unrealistic local geometries, and generate energetically annotated conformers. We will revise the manuscript to make this distinction clearer and to avoid implying that ProteinConformers provides a rigorous Boltzmann-sampled representation of the underlying thermodynamic landscape.

      We also agree that longer simulations and enhanced sampling methods, such as replica-exchange MD, metadynamics, or umbrella sampling, would be necessary to estimate equilibrium populations and improve sampling of rare but thermodynamically relevant states. We will add this point as a limitation and future direction in the revised manuscript. Thus, ProteinConformers should be viewed as a broad, energetically annotated, MDrefined conformer library for benchmarking, data-driven modeling, and descriptions of protein conformational landscapes, rather than as a complete equilibrium ensemble itself.

      Reviewer #3 (Public review):

      Summary:

      This manuscript describes a web-based tool that allows researchers to compare large numbers of representative ("plausible") conformations of proteins. It also includes energetic analysis from multiple widely used structure-prediction methods.

      Strengths:

      This tool will likely be useful for students who want to learn more about the ensemble properties of proteins. The resource is well organized and it represents a large amount of computing resources.

      We thank Reviewer #3 for the positive assessment of the ProteinConformers resource and for recognizing its potential value for community education.

      Weaknesses:

      It is not entirely clear how the database may be utilized by other groups to advance research. It could be helpful if the authors add a short section that provides example use cases that illustrate how this database can support new strategies for studying protein dynamics.

      We thank the reviewer for this constructive suggestion. We agree that the manuscript should more explicitly explain how other groups can use ProteinConformers to advance research. In the revised Discussion, we will add a short section describing concrete use cases of the database. In particular, we will emphasize that the benchmark analysis already presented in this manuscript provides a worked example of how ProteinConformers can be used by other groups. ProteinConformers-lite, together with the released evaluation metrics and codes, can serve as a standardized reference set for testing new multi-conformation or protein ensemble generation methods. Other groups can generate conformational ensembles for the same targets, compare their coverage of low-energy regions using the diversity metrics reported in Table S1, and evaluate the agreement of residue-pair geometric statistics using the plausibility metrics reported in Table S2.

      We will also describe additional use cases enabled by the full ProteinConformers resource. The dataset can be used as a training, validation, or pretraining resource for conformation generators, energy-aware ranking models, and model quality assessment methods, because each conformer is paired with structural similarity annotations and multiple energetic scores. The broad coverage from non-native to near-native conformations also enables systematic analysis of how local stereochemical validity, global structural similarity, and energetic evaluations co-vary across diverse conformational perturbations, and may provide useful structural proxies or starting points for modeling flexible or disordered-like protein states, where experimentally resolved structural data are often limited. In addition, the interactive portal allows users to filter protein-specific conformers by structural similarity, energetic annotations, and secondary-structure features, making it possible to construct customized subsets for downstream biomolecular modeling, hypothesis generation, and educational exploration.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work by Hall et al provides a novel and important new finding about communication between the anterior cingulate cortex (ACC) and the CA1 region of the dorsal hippocampus: there is a clear ability of ACC to predict CA1 activity, and that is modulated by learning/experience. Furthermore, they have some evidence that the modulation differs by whether the CA1 neurons were in the deep versus superficial sub-layer of CA1. The evidence is suggestive of new and exciting findings, but some gaps and weaknesses remain to be addressed before I believe all of the authors' claims can be supported. The figures also need to be slightly better organized, and the discussion is missing a major dimension in my opinion. Overall, this is a strong submission, but with some gaps to fill.

      Strengths:

      (1) This is a well-written manuscript - the introduction was especially clear, well-cited, and motivating.

      (2) The sub-layer specific communication between ACC and CA1 represents the discovery of a novel and functionally impactful piece of neurobiology.

      (3) Optogenetics was an important verification of ACC-CA1 communication, as was the analysis of neurons by waveform type.

      Weaknesses:

      (1) Figure 2: Why are the data separated into two groups from the outset? If all data are combined, is there a general drop in prediction gain from pre to post?

      Thank you for bringing this to our attention. In Figure 1F, all data is combined for GLM decoding. We found a significant pre-to-post decrease in prediction gain specifically using the –200 to 0 ms window to predict CA1 spiking during ripples. Figure 3 builds upon these findings to examine how prediction gain changes relate to task engagement.

      (2) 2b and 2c are important since they are complementary means to show the same thing, and it is important that they cross-validate each other, especially since the non-significant task active neuron difference in 2b appears to be nearly as strong as the significant difference to its left. A more holistic analysis can be done to compare these dimensions.

      We appreciate this feedback. In light of this comment, as well as similar concerns raised by other reviewers regarding the binary classification of neurons as task-active or task-inactive (modulation index > 0 or < 0), we adopted a more comprehensive analytical approach. Rather than relying on an arbitrary threshold, we first examined the continuous relationship between modulation index and prediction gain change across all neurons (Figure 2C). We then divided neurons into modulation index quartiles to assess whether prediction gain changes varied across different degrees of task modulation (Figure 2D). Follow-up analyses focused on the extreme quartile comparisons that contributed to the observed effects (Figure 2E). Overall, this new approach better captured the continuous nature of task-related modulation while avoiding the inclusion of a large population of neurons with modulation indices near zero that may not meaningfully differ in task engagement. However, we were unable to replicate modulation index changes for neurons split by prediction gain score percentile and removed those figures from the manuscript. We have adjusted the text accordingly to account for this new effect.

      (3) Sup vs deep neuron definition: Did the authors have any means to validate this anatomical separation using histology or otherwise? I don't believe they described anything like that, and instead use physiology to infer anatomical location. I understand anatomy-based methods may be practically impossible with tetrodes, but this limitation should at least be mentioned, and it should be explained that without something like silicon probes or histological validation, anatomy had to be inferred from physiology.

      We think this is an important limitation to address and thank you for bringing this to our attention. Given our technical restraints, we only validated radial position through physiological properties. This remains an outstanding limitation of this study. We have since added text into the main discussion bringing attention to this limitation. See below:

      “Lastly, our CA1 sublayer classification does come with its own caveats. Tetrode identification of CA1 sublayers is not a trivial matter. We implemented guidelines informed by past research (see methods) to help ensure we isolated CA1sup and CA1deep groups, only including neurons where we had the greatest confidence (Berndt et al., 2023; Mizuseki et al., 2011). In doing so, we excluded neurons which classification was uncertain. It is possible this excludes some meaningful populations of CA1 sublayers. Additionally, despite our approach, some cross-inclusion of sublayers may remain. Therefore, interpretations of CA1 sublayers difference should be considered with these limitations in mind. That said, the number of neurons and animals tested does provide overall confidence regarding our results. Future studies investigating this ACC to CA1 sublayer specific line of communication would benefit from the use of silicon probes or neuropixels that enable precise radial localization.”

      (4) Superficial vs deep differences in firing rate ratio based on PG: there are many fewer CAdeep neurons, but in 4c, the trends appear to be the same pre-training, top PG lower than others. It seems the lack of difference in CA1deep in 4c may be due to the much lower power/n. This should be discussed or addressed.

      We appreciate this feedback and since have added more recordings to address these lower Ns (CA1sup: Previous N = 71, Revisions N = 89; CA1deep: Previous N = 21, Revisions N = 61). Notably, we find that previous firing rate ratio (now called modulation index) is no longer significant with the inclusion of more neurons and is reported accordingly (Figure 2—figure supplement 1).

      (5) In Figure 5, the term "firing rate ratio" is used, and it sounds the same as in previous figures, but this is a different ratio (based on modulation by opto stim, not task).

      To improve clarity and avoid confusion with task-related modulation metric used across the paper, we renamed “Firing Rate Ratio" throughout the manuscript to "Modulation Index". We also relabeled the Figure 5D Y-axis as "Z-scored Firing Response" to more accurately reflect the plotted data and avoid confusion.

      (6) I would like to learn more about these v-type neurons. I understand we do not yet know about their molecular or morphologic correlate, but more analysis can be done with the current data.

      We thank you for this feedback. We have included further analysis into V-type properties. Namely, we performed autocorrelegrams, theta phase modulation, and burst index analyses. See Figure 6 and Figure 6—figure supplement 1.

      Additionally, we performed cross-correlogram analyses to examine whether V-type interneurons exhibited consistent temporal relationships with PV interneurons or other CA1 neurons. However, V-type and PV interneurons were sparse throughout our recordings, with most sessions containing two or fewer identified interneurons, which limited our ability to perform meaningful cross-correlogram analyses. Nevertheless, we examined the available recordings but found no consistent evidence of correlated firing between interneuron classes or between V-type interneurons and CA1 pyramidal neurons.

      (7) I would like more discussion of ACC-CA1 connectivity.

      We have since added greater discussion of ACC-CA1 connectivity into the discussion section. See below:

      “Finally, an important caveat to mention is that it remains an ongoing debate whether ACC directly projects to CA1 (Andrianova et al., 2023; Rajasethupathy et al., 2015; Shi et al., 2022). One lab reported clear monosynaptic ACC-to-CA1 connection (Rajasethupathy et al., 2015), while another lab replicated those same experiments and were unable to come to the same conclusions (Andrianova et al., 2023). Further studies report no direct connection (Shi et al., 2022). The contention in connectivity may arise from differences in targeting strategies, injection coordinates, or viruses used. Our findings reported an excitatory response in CA1 V-type interneurons in response to ACC stimulations, proposing another possibility for ACCàCA1 connectivity. Interestingly V-type interneurons responded with extremely low-latency as fast as 4.2 ms after stimulations, compatible with monosynaptic timing (Cho et al., 2013; Petreanu et al., 2007; Wang et al., 2009). If the ACC→V-Type connection was monosynaptic pathway, it could help explain discrepancies in the field, as the relative sparsity of V-type interneurons may reduce the likelihood of detecting ACC→CA1 connectivity. However, future anatomical studies are needed to conclusively determine connectivity.

      Alternatively, ACC→CA1 communication may be mediated by multiple intermediate structures (Behzadi et al., 1990; Oh et al., 2014; Shi et al., 2022; Souza et al., 2022). The ACC sends monosynaptic projections to the nucleus reuniens (RE) and median raphe (MnR), both of which project directly to CA1 (Oh et al., 2014; Shi et al., 2022). The RE has a known role in contextual discrimination learning and memory specificity (Ramanathan & Maren, 2019; Ramanathan et al., 2018; Ratigan et al., 2023; Silva et al., 2021; Xu & Südhof, 2013). Interestingly, RE→CA1 activity tuned to immobility (freezing) emerges only after shocks are presented, suggesting a learning-induced modification between regions, similar to that seen in our ACC–CA1 data (Ratigan et al., 2023). As for MnR, it receives dense inputs from the ACC (Behzadi et al., 1990; Souza et al., 2022), and its projections to the CA1 are predominantly glutamatergic (Jackson et al., 2009; Senft et al., 2021; Szonyi et al., 2016). Notably, these glutamatergic MnR inputs directly target CA1 interneurons, including CCK basket cells (Miettinen & Freund, 1992; Morales & Bloom, 1997; Senft et al., 2021), while avoiding PV interneurons and pyramidal neurons (Acsady et al., 1993; Freund et al., 1990; Halasy et al., 1992; Miettinen & Freund, 1992; Papp et al., 1999; Turi et al., 2019). This connectivity suggests that the ACC may indirectly modulate CA1 activity through the MnR, potentially suppressing PV interneuron and pyramidal neuron activity via local inhibitory circuits, thereby contributing to the regulation of hippocampal oscillations and memory consolidation (Huang et al., 2022; Wang et al., 2015). Ultimately, future experiments combining pathway-specific manipulations with simultaneous recordings will be necessary to distinguish direct from polysynaptic mechanisms.”

      (8) Some elements may be missing from the discussion, relating baseline functioning versus post-learning function.

      We thank the reviewer for this feedback and their recommendation for possible alternate explanations. We have added these discussions into the main text. See below:

      “Alternatively, ACC→CA1 communication may contribute to the homeostatic downscaling of memory-unrelated synapses during sleep. Evidence finds that slow-wave sleep is strongly linked to downscaling of non-learning related neuron activity (Gulati et al., 2017; Liu et al., 2010; Tononi & Cirelli, 2003; Tononi & Cirelli, 2006; Watson et al., 2016). Slow-wave sleep ripples in particular depotentiate memory-unrelated synapses (Gulati et al., 2017; Norimoto et al., 2018). In our study, we find that learning-related reduction in communication between ACC and CA1 were selective for task-inactive neurons. Therefore, ACC→CA1sup communication may not simply weaken following learning but rather becomes selectively disengaged from task-inactive neurons, enabling homeostatic downscaling while preserving behaviorally relevant synapses (Liu et al., 2010; Norimoto et al., 2018; Tononi & Cirelli, 2003; Tononi & Cirelli, 2006; Watson et al., 2016). Still, behavioral recruitment alone cannot account for the observed remodeling of ACC→CA1 communication, as CA1sup and CA1deep neurons did not display significant differences in task-related activity (Figure 2—figure supplemental 2). Instead, these learning-related changes of task-inactive neurons appear sublayer-specific.”

      Reviewer #2 (Public review):

      Summary:

      This study uncovers an inhibitory pathway from the anterior cingulate cortex (ACC) to pyramidal cells in the superficial sublayer of hippocampal area CA1 (CA1sup). As ACC neuron spiking tends to precede hippocampal ripples, this presents the intriguing possibility that ACC inputs are selectively inhibiting particular CA1sup neurons, which could play a role in the reactivation of task-related ensembles known to take place during hippocampal ripples. Indeed, through a generalized linear model (GLM) analysis, the authors demonstrate that the ACC activity within the 200ms immediately preceding the ripple is predictive of the ripple content.

      Strengths:

      The biggest strength of the work is the optogenetic manipulation experiments, which convincingly demonstrate that stimulation of ACC pyramidal neurons activates an interneuron population with symmetric spike waveforms, and inhibits parvalbumin interneurons and pyramidal cells in CA1sup but not CA1deep sublayer.

      An additional strength in the GLM analysis which consistently shows that ACC activity preceding the ripple is predictive of hippocampal activity during the ripple considerably more than in shuffled data for all cells and periods tested.

      Weaknesses:

      The major weakness of this work is that the link with learning and memory is not very well supported.

      The only evidence of rebalancing and reorganization appears to be a single statistical test (the test in Figure 1f, p=0.013) demonstrating a decrease of the GLM prediction gain from pre-task sleep to post-task sleep; the same test is repeated for subsets of the data in the rest of the figures. As the idea of rebalancing and reorganization is central to the paper as currently written, exploring it through another measure, independent of the GLM prediction gain, should be expected. The notion that this pathway is suppressed in sleep following learning can be supported by demonstrating a decrease in any of the following measures: ACC spike-triggered average CA1sup responses, cross-covariances (Wierzynski et al 2009) between ACC and CA1sup cells in post-task sleep, or ripple-triggered cross-correlations (Sirota et al. 2009).

      We thank the reviewer for this helpful feedback. We have added an additional analysis the reviewer pointed out to address this concern. Specifically, we performed an ACC spike‑triggered analysis. The ACC spike‑triggered average further supported the learning‑related decrease in ACC-to-CA1 activity (see Figure 1—figure supplement 1). We did not include a separate cross-covariance analysis because the ACC spike-triggered average captures essentially the same temporal relationship between ACC and CA1 activity.

      The differences between task-active and task-inactive neurons are not convincing. The separation between task-active and task-inactive neurons is to divide a distribution that is far from bimodal into what appears to be two arbitrary groups. Similarly, the authors divide cells relative to their prediction gain ("Top PG" and "Bottom PG" in Figure 2c), which fails to select for the population of significantly predicted cells (relative to the shuffle). Within CA1sup cells, after learning, there is a significant decrease in the prediction gain for "task-inactive" cells but not "task-active" cells, but it is important to keep in mind that the "task-active" group contains only 24 neurons, and there was no difference between the two groups of cells ("task-active" vs "task-inactive") when directly compared.

      We agree with this concern. To address this, we removed conclusion based on those arbitrary criteria instead opting for a more continuous approach. Specifically, we adopted a more comprehensive analytical approach. Rather than relying on an arbitrary threshold, we first examined the continuous relationship between modulation index and prediction gain change across all neurons (Figure 2C). We then divided neurons into modulation index quartiles to assess whether prediction gain changes varied across different degrees of task modulation (Figure 2D). Follow-up analyses focused on the extreme quartile comparisons that contributed to the observed effects (Figure 2E). Overall, this new approach better captured the continuous nature of task-related modulation while avoiding the inclusion of a large population of neurons with modulation indices near zero that may not meaningfully differ in task engagement.

      Finally, it is not clear whether the identity of the pathway-responsive CA1sup neurons is fixed or whether it may change with learning. A deeper analysis into the cell pair cross-correlations or the weights of the GLM analysis may reveal whether there is a reorganization of CA1sup responses (some cells that were inhibited are no longer inhibited, and vice versa) or a dampening (the same CA1sup cells are inhibited in both cases, but the inhibition is less-pronounced in post-task sleep). The possibility of a rigid circuit dampened immediately following fear conditioning, is not discussed by the authors.

      We appreciate this feedback. To address this concern without weight analysis, we examined the stability of prediction gain scores between pre- and post-training sleep. Preservation of neuronal prediction-gain rankings would suggest that learning weakens existing predictive communication while maintaining the relative contribution of individual neurons, consistent with a dampening response. In contrast, poor preservation of prediction-gain rankings would be more indicative of a reorganization of predictive relationships across the population. This led to interesting findings regarding sublayer differences: ACC→CA1sup communication is more dynamic and evolving following learning, whereas ACC→CA1deep communication remains comparatively stable (See Figure 2A&B; Figure 3 C–F).

      Reviewer #3 (Public review):

      Summary:

      In this study, Hall and colleagues investigate how the coupling of activity from ACC to CA1is altered by fear learning, showing that during sleep immediately before learning, there is evidence for increased coupling of ACC activity with neurons that will subsequently be inhibited during the learning process. They go on to show that this effect seems to be mediated most by a subpopulation of neurons in the superficial layer of CA1. This fits with previous reports suggesting that these superficial neurons are key for the flexible updating of memory. The authors then go on to show that artificial activation of ACC using optogenetics results in varied effects in CA1, including a subtle decrease in activity of superficial neurons that lasts longer than the stimulus itself. Finally, the authors present some preliminary data suggesting that different interneurons may be recruited by this optogenetic stimulation in different ways and at different times.

      Overall, this is an interesting paper, but much of the analysis is very preliminary, and much of the crucial data about the learning effects and alterations to cell firing are not presented clearly and fully. This is further confounded by a rather opaque description of the results and analysis in the text. Overall, there is something very interesting here, but there needs to be a substantial series of extra analyses to clearly say what this is. In many cases, more robust analysis may render the results underpowered, which could dramatically change the conclusions of the paper.

      Strengths:

      The authors performed difficult, dual-location recordings across a multi-day learning paradigm, which seems like it could be a really nice dataset. They delve into the circuit basis of an interesting finding regarding ACC to CA1 connectivity and how this changes before and after fear conditioning. They provide data to suggest this connectivity may be through specific and distinct subcircuits in CA1.

      Weaknesses:

      (1) There is essentially no information in the text or figures about what the actual learning was, how it was done, how individual animals performed, and how any of these metrics related to learning. Looking at the methods, the authors did a number of things never mentioned anywhere in the text or figures, including novel arena exposure, contextual reexposure in extinction after learning, etc. It seems that this is a very rich dataset that has not been presented at all. I would recommend at the very least:

      We appreciate the reviewers’ feedback and have worked to address these concerns. See below our response.

      (a) Plot all of the behavioural training data, and how each mouse relates to one another - did the mice learn? At this stage, we don't know!

      We have now plotted all contextual fear conditioning behavioral data for each mouse (Figure 2—figure supplement 1). All mice exhibited high level of freezing during the contextual fear test, suggesting successful learning of the context–shock association.

      (b) Explain in the text in detail exactly what was done and why, and what this tells us about the neuronal activity.

      We have now added text to describe in detail the behavioral results and how that may relate to our GLM analyses. We have also more clearly detailed our experimental objectives (what was done and why) utilizing contextual fear conditioning,

      “In this approach, we were able to examine ACC–CA1 communication prior, during, and after learning, enabling us to examine how this communication evolves across fear learning. Specifically, we emphasized investigation into communication changes between pre- and post-training sleep to understand whether functional connectivity undergoes learning-related reorganization.”

      “Lastly, we examined whether PG scores correlated with the freezing response in mice during recall. Across all mice, freezing was significantly higher during recall than pre-shock baseline during training (Figure 2—figure supplement 2A). Overall, we found no correlation between PG and freezing (Figure 2—figure supplement 2B–D). However, there was a trend for a positive correlation (p = .07) between ΔPG and freezing percentage. An important consideration is that the uniformly high levels of freezing in mice limited behavioral variability, potentially reducing our ability to detect relationships between ACC–CA1 communication decoding and behavioral differences.”

      (c) If there is variance in learning and or conditioning, does this relate to features in the analysis, such as the GLM result.

      We examined this question by first investigating whether prediction gain scores in pre-training, post-training, or overall ΔPG correlated with freezing percentage. We found no significant correlation between any of the variables (Figure 2—figure supplement 1). We speculate this may be a result of a relatively robust freezing response limiting the ability for our fine-grained GLM decoding analyses to detect those differences. We have added this consideration to the main text.

      (2) Along similar lines, a key metric for most of the paper is that neurons most coupled with ACC are more likely to be inhibited during training. However, there is nothing anywhere in the paper showing these data. How do neurons in general respond to contextual shocks? The methods describe this as the average firing rate during training, normalised to pre-sleep activity. This metric seems a bit coarse and may obscure really important task-relevant dynamics. Are the neurons active at specific times, are they tuned to relevant parts of the task, and do any of these features of the cell activity also relate to the coupling with ACC? Similarly, how did the authors mitigate the influence of electrical artefacts caused by the foot shock in their recordings? Again, there is a huge amount of data here that is not being described, and likely holds very valuable information about what is actually happening. The paper would really benefit from the inclusion of these data in an accessible form, such as heatmaps of spiking, how these patterns change over time, and around e.g., foot shock, etc. Also key is how these features are altered by the variability of learning across subjects.

      We thank reviewer for this feedback. As pointed out, electrical artifacts caused by the footshocks prevents our ability to examine, with temporal sensitivity, neurons’ responses to footshocks. Therefore, we are left to examine activity changes across longer timescales. We acknowledge that our current modulation index analysis is a bit coarse. One reason is that our preliminary analyses using more temporally sensitive approaches did not reveal robust CA1 activity associated with specific behaviors, such as freezing or transitions between mobility and immobility. Thus, we chose a more holistic approach looking at the full CFC session to include all components that CA1 may be encoding during the training session. For example, although the pre-shock baseline period does not contain any footshock stimuli, it serves a key part in the process as mice begin to encode their environment around them. Nevertheless, we have added an additional analysis to examine how modulation changes across the pre-shock versus post-shock window (Figure 5—figure supplement 1). Although this provides greater insight into how ACC and CA1 activity changes across different dimensions in the task, further investigation utilizing casual manipulations is necessary to elucidate which phase ACC→CA1 activity is most involved.

      (3) A number of the effects are presented by comparing a statistically significant effect to a non-statistically significant effect (e.g. in Figure 2b, Figure 2d, Figure 4 b,c, and others). This isn't really valid - the key test that the two groups are different is either with a direct test of the difference or an interaction term in an e.g., ANOVA test. In some places, I am not sure the same conclusions will be drawn from the data with these tests.

      We want to thank the reviewer for this critical feedback. We have since added the appropriate statistical measure including linear mixed-effects models and ANOVA tests and for our analysis to avoid our previous statistical errors.

      (4) To what extent is defining superficial and deep CA1 neurons solely by ripple waveform an accepted method? Of the two papers referenced for this approach, one is a 2-photon calcium imaging paper that does not do electrical recordings (as far as I am aware), and the second uses this as a descriptor after defining the positions of units on an array. It would be good to clarify how accepted this is, and also how robust this is. At the very least, some kind of metric or walkthrough in the supplement as to how this was done, and how well each cell was classified and with what confidence, or some metric of how distinct and separate the two populations were (or was it just a smudge).

      We appreciate this feedback. While the Berndt 2023 paper implemented 2-photon calcium imaging, they also used tetrode classifications for radial axes in that paper which help informed our approach. Ultimately, our tetrode classification remains an outstanding limitation which we have since added to the main text (See below).

      “Lastly, our CA1 sublayer classification does come with its own caveats. Tetrode identification of CA1 sublayers is not a trivial matter. We implemented guidelines informed by past research (see methods) to help ensure we isolated CA1sup and CA1deep groups, only including neurons where we had the greatest confidence (Berndt et al., 2023; Mizuseki et al., 2011). In doing so, we excluded neurons which classification was uncertain. It is possible this excludes some meaningful populations of CA1 sublayers. Additionally, despite our approach, some cross-inclusion of sublayers may remain. Therefore, interpretations of CA1 sublayers difference should be considered with these limitations in mind. That said, the number of neurons and animals tested does provide overall confidence regarding our results. Future studies investigating this ACC to CA1 sublayer specific line of communication would benefit from the use of silicon probes or neuropixels that enable precise radial localization.”

      (5) In the optogenetic experiment in Figure 5, the effect on the CA1 sup neurons seems to be driven by changes in a small subpopulation of this group, with no change in the others. Related to point 2, is there anything else in the data that can pull out what these cells are? More detailed analysis of the firing of these neurons might pull out something really interesting.

      We thank the reviewer for this feedback. Firstly, we want to clarify that optogenetic experiments were performed in a separate cohort of mice that did not undergo contextual fear conditioning. We have adjusted the text accordingly to make this distinction clearer. We have also added a per-animal separation of CA1 heatmap responses to ACC stimulations to demonstrate suppression is preserved across animals (Figure 5—figure supplement 3; Figure 6—figure supplement 2). Finally, our optogenetic experiments were primarily focused on understanding anatomical connectivity. Consequently, common waking behaviors (e.g., exploration and feeding) were not standardized across animals, and our analyses were therefore restricted to comparisons between slow-wave sleep and wakefulness more broadly.

      (6) Related to this - a number of comparisons simply pool neurons across mice and analyse them as if independent. This is done a lot in the past, but it would be better if an approach that included the interdependence of neurons recorded from the same mouse at the same time were used (such as a hierarchical model). While this is complex, a simpler approach would just be to plot the summary data also per mouse. For example, in Figure 5, how do the neurons inhibited by ACC activation spread across the different mice? Is the level of inhibition related to how well the mice learned the CS-US association?

      For dual-site analysis we have now added animal-level comparison for some key analyses (See Figure—figure supplement 1&2). As for optogenetic experiments, we have added per-animal heatmaps for ACC stimulation response (Figure 5—figure supplement 3; Figure 6—figure supplement 2).

      (7) Figure 6 is interesting, but very preliminary. None of the effects are quantified, and one of the cell types is not identified. I think some proper analysis needs to be done, again across mice, to be able to draw conclusions from these data.

      We thank the reviewer for this feedback. Reviewer 1 had a similar concern, and we have since added additional analyses to the revisions. Specifically, we performed autocorrelegrams, theta phase modulation analyses and a burst index analysis for V-Type interneurons. Importantly, however, these approaches still collapse neurons across mice. Given the sparse nature of V-type and PV interneurons, files often contain 2 or fewer interneurons making within-animal comparison difficult. That said, we have added supplemental figures displaying per-animal changes in response to ACC stimulations (Figure 6—figure supplement 1).

      (8) Finally, in general, I felt that the way the paper was written was very hard to follow, often relying on very processed levels of analysis that were hard to relate back to the raw traces and their biological meaning. In general taking more words to really simply and fully explain each analysis, and taking the words and figures to walk through how each analysis was done and what it tells us about the neuronal data/biology would be really beneficial, especially to someone who is not an extracellular electrophysiologist or immersed in the immediate field.

      We thank the reviewer for this feedback. Throughout the manuscript, we have revised the text to improve clarity in explaining our approaches and their results.

      In summary, while this manuscript explores an intriguing hypothesis about pre-learning circuit dynamics, it is currently held back by insufficient clarity in behavioural analysis, data presentation, and statistical quantification. Addressing these core issues would greatly improve interpretability and confidence in the findings.

      Additional comment:

      For the optogenetic experiments, we reprocessed and resorted the neuronal dataset to ensure accurate cell classification. Following this re-analysis, the principal findings remained unchanged. However, we found that sublayer-specific differences in response to ACC stimulation were restricted to the first second following ACC stimulation. Consequently, we removed the previous Figure 5E, which examined firing rate changes across successive 1-sec time bins, as the additional time windows did not provide further evidence of sublayer-specific effects.

      Acsady, L., Halasy, K., & Freund, T. F. (1993). Calretinin is present in non-pyramidal cells of the rat hippocampus--III. Their inputs from the median raphe and medial septal nuclei. Neuroscience, 52(4), 829-841. https://doi.org/10.1016/0306-4522(93)90532-k

      Andrianova, L., Yanakieva, S., Margetts-Smith, G., Kohli, S., Brady, E. S., Aggleton, J. P., & Craig, M. T. (2023). No evidence from complementary data sources of a direct glutamatergic projection from the mouse anterior cingulate area to the hippocampal formation. eLife, 12, e77364. https://doi.org/10.7554/eLife.77364

      Behzadi, G., Kalén, P., Parvopassu, F., & Wiklund, L. (1990). Afferents to the median raphe nucleus of the rat: Retrograde cholera toxin and wheat germ conjugated horseradish peroxidase tracing, and selective<span class="small">d</span>-[<sup>3</sup>H]aspartate labelling of possible excitatory amino acid inputs. Neuroscience, 37(1), 77-100. https://doi.org/10.1016/0306-4522(90)90194-9

      Berndt, M., Trusel, M., Roberts, T. F., Pfeiffer, B. E., & Volk, L. J. (2023). Bidirectional synaptic changes in deep and superficial hippocampal neurons following in vivo activity. Neuron, 111(19), 2984-2994.e2984. https://doi.org/10.1016/j.neuron.2023.08.014

      Cho, J. H., Deisseroth, K., & Bolshakov, V. Y. (2013). Synaptic encoding of fear extinction in mPFC-amygdala circuits. Neuron, 80(6), 1491-1507. https://doi.org/10.1016/j.neuron.2013.09.025

      Freund, T. F., Gulyas, A. I., Acsady, L., Gorcs, T., & Toth, K. (1990). Serotonergic control of the hippocampus via local inhibitory interneurons. Proc Natl Acad Sci U S A, 87(21), 8501-8505. https://doi.org/10.1073/pnas.87.21.8501

      Gulati, T., Guo, L., Ramanathan, D. S., Bodepudi, A., & Ganguly, K. (2017). Neural reactivations during sleep determine network credit assignment. Nat Neurosci, 20(9), 1277-1284. https://doi.org/10.1038/nn.4601

      Halasy, K., Miettinen, R., Szabat, E., & Freund, T. F. (1992). GABAergic Interneurons are the Major Postsynaptic Targets of Median Raphe Afferents in the Rat Dentate Gyrus. Eur J Neurosci, 4(2), 144-153. https://doi.org/10.1111/j.1460-9568.1992.tb00861.x

      Huang, W., Ikemoto, S., & Wang, D. V. (2022). Median Raphe Nonserotonergic Neurons Modulate Hippocampal Theta Oscillations. J Neurosci, 42(10), 1987-1998. https://doi.org/10.1523/JNEUROSCI.1536-21.2022

      Jackson, J., Bland, B. H., & Antle, M. C. (2009). Nonserotonergic projection neurons in the midbrain raphe nuclei contain the vesicular glutamate transporter VGLUT3. Synapse, 63(1), 31-41. https://doi.org/10.1002/syn.20581

      Liu, Z.-W., Faraguna, U., Cirelli, C., Tononi, G., & Gao, X.-B. (2010). Direct Evidence for Wake-Related Increases and Sleep-Related Decreases in Synaptic Strength in Rodent Cortex. The Journal of Neuroscience, 30(25), 8671. https://doi.org/10.1523/JNEUROSCI.1409-10.2010

      Miettinen, R., & Freund, T. F. (1992). Convergence and segregation of septal and median raphe inputs onto different subsets of hippocampal inhibitory interneurons. Brain Res, 594(2), 263-272. https://doi.org/10.1016/0006-8993(92)91133-y

      Mizuseki, K., Diba, K., Pastalkova, E., & Buzsáki, G. (2011). Hippocampal CA1 pyramidal cells form functionally distinct sublayers. Nature Neuroscience, 14(9), 1174-1181. https://doi.org/10.1038/nn.2894

      Morales, M., & Bloom, F. E. (1997). The 5-HT3 receptor is present in different subpopulations of GABAergic neurons in the rat telencephalon. J Neurosci, 17(9), 3157-3167. https://doi.org/10.1523/JNEUROSCI.17-09-03157.1997

      Norimoto, H., Makino, K., Gao, M., Shikano, Y., Okamoto, K., Ishikawa, T., Sasaki, T., Hioki, H., Fujisawa, S., & Ikegaya, Y. (2018). Hippocampal ripples down-regulate synapses. Science, 359(6383), 1524-1527. https://doi.org/10.1126/science.aao0702

      Oh, S. W., Harris, J. A., Ng, L., Winslow, B., Cain, N., Mihalas, S., Wang, Q., Lau, C., Kuan, L., Henry, A. M., Mortrud, M. T., Ouellette, B., Nguyen, T. N., Sorensen, S. A., Slaughterbeck, C. R., Wakeman, W., Li, Y., Feng, D., Ho, A., . . . Zeng, H. (2014). A mesoscale connectome of the mouse brain. Nature, 508(7495), 207-214. https://doi.org/10.1038/nature13186

      Papp, E. C., Hajos, N., Acsady, L., & Freund, T. F. (1999). Medial septal and median raphe innervation of vasoactive intestinal polypeptide-containing interneurons in the hippocampus. Neuroscience, 90(2), 369-382. https://doi.org/10.1016/s0306-4522(98)00455-2

      Petreanu, L., Huber, D., Sobczyk, A., & Svoboda, K. (2007). Channelrhodopsin-2–assisted circuit mapping of long-range callosal projections. Nature Neuroscience, 10(5), 663-668. https://doi.org/10.1038/nn1891

      Rajasethupathy, P., Sankaran, S., Marshel, J. H., Kim, C. K., Ferenczi, E., Lee, S. Y., Berndt, A., Ramakrishnan, C., Jaffe, A., Lo, M., Liston, C., & Deisseroth, K. (2015). Projections from neocortex mediate top-down control of memory retrieval. Nature, 526(7575), 653-659. https://doi.org/10.1038/nature15389

      Ramanathan, K. R., & Maren, S. (2019). Nucleus reuniens mediates the extinction of contextual fear conditioning. Behavioural brain research, 374, 112114. https://doi.org/https://doi.org/10.1016/j.bbr.2019.112114

      Ramanathan, K. R., Ressler, R. L., Jin, J., & Maren, S. (2018). Nucleus Reuniens Is Required for Encoding and Retrieving Precise, Hippocampal-Dependent Contextual Fear Memories in Rats. The Journal of Neuroscience, 38(46), 9925. https://doi.org/10.1523/JNEUROSCI.1429-18.2018

      Ratigan, H. C., Krishnan, S., Smith, S., & Sheffield, M. E. J. (2023). A thalamic-hippocampal CA1 signal for contextual fear memory suppression, extinction, and discrimination. Nat Commun, 14(1), 6758. https://doi.org/10.1038/s41467-023-42429-6

      Senft, R. A., Freret, M. E., Sturrock, N., & Dymecki, S. M. (2021). Neurochemically and Hodologically Distinct Ascending VGLUT3 versus Serotonin Subsystems Comprise the r2-Pet1 Median Raphe. J Neurosci, 41(12), 2581-2600. https://doi.org/10.1523/JNEUROSCI.1667-20.2021

      Shi, W., Xue, M., Wu, F., Fan, K., Chen, Q. Y., Xu, F., Li, X. H., Bi, G. Q., Lu, J. S., & Zhuo, M. (2022). Whole-brain mapping of efferent projections of the anterior cingulate cortex in adult male mice. Mol Pain, 18, 17448069221094529. https://doi.org/10.1177/17448069221094529

      Silva, B. A., Astori, S., Burns, A. M., Heiser, H., van den Heuvel, L., Santoni, G., Martinez-Reza, M. F., Sandi, C., & Gräff, J. (2021). A thalamo-amygdalar circuit underlying the extinction of remote fear memories. Nature Neuroscience, 24(7), 964-974. https://doi.org/10.1038/s41593-021-00856-y

      Souza, R., Bueno, D., Lima, L. B., Muchon, M. J., Gonçalves, L., Donato, J., Jr., Shammah-Lagnado, S. J., & Metzger, M. (2022). Top-down projections of the prefrontal cortex to the ventral tegmental area, laterodorsal tegmental nucleus, and median raphe nucleus. Brain Struct Funct, 227(7), 2465-2487. https://doi.org/10.1007/s00429-022-02538-2

      Szonyi, A., Mayer, M. I., Cserep, C., Takacs, V. T., Watanabe, M., Freund, T. F., & Nyiri, G. (2016). The ascending median raphe projections are mainly glutamatergic in the mouse forebrain. Brain Struct Funct, 221(2), 735-751. https://doi.org/10.1007/s00429-014-0935-1

      Tononi, G., & Cirelli, C. (2003). Sleep and synaptic homeostasis: a hypothesis. Brain Research Bulletin, 62(2), 143-150. https://doi.org/https://doi.org/10.1016/j.brainresbull.2003.09.004

      Tononi, G., & Cirelli, C. (2006). Sleep function and synaptic homeostasis. Sleep Med Rev, 10(1), 49-62. https://doi.org/10.1016/j.smrv.2005.05.002

      Turi, G. F., Li, W. K., Chavlis, S., Pandi, I., O'Hare, J., Priestley, J. B., Grosmark, A. D., Liao, Z., Ladow, M., Zhang, J. F., Zemelman, B. V., Poirazi, P., & Losonczy, A. (2019). Vasoactive Intestinal Polypeptide-Expressing Interneurons in the Hippocampus Support Goal-Oriented Spatial Learning. Neuron, 101(6), 1150-1165 e1158. https://doi.org/10.1016/j.neuron.2019.01.009

      Wang, D. V., Yau, H.-J., Broker, C. J., Tsou, J.-H., Bonci, A., & Ikemoto, S. (2015). Mesopontine median raphe regulates hippocampal ripple oscillation and memory consolidation. Nature Neuroscience, 18(5), 728-735. https://doi.org/10.1038/nn.3998

      Wang, J., Hasan, M. T., & Seung, H. S. (2009). Laser-evoked synaptic transmission in cultured hippocampal neurons expressing channelrhodopsin-2 delivered by adeno-associated virus. J Neurosci Methods, 183(2), 165-175. https://doi.org/10.1016/j.jneumeth.2009.06.024

      Watson, Brendon O., Levenstein, D., Greene, J. P., Gelinas, Jennifer N., & Buzsáki, G. (2016). Network Homeostasis and State Dynamics of Neocortical Sleep. Neuron, 90(4), 839-852. https://doi.org/https://doi.org/10.1016/j.neuron.2016.03.036

      Xu, W., & Südhof, T. C. (2013). A Neural Circuit for Memory Specificity and Generalization. Science, 339(6125), 1290-1295. https://doi.org/doi:10.1126/science.1229534

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggested fixes (these correlate with the numbered points in the Weaknesses section of the Public Review):

      (1) That seems like the larger finding to start with, before dissecting by Firing Activity Index. I'd first show the overall finding, then dissect it.

      We thank the reviewer for this recommendation. In our manuscript, figure 1F examines the overall pre-to-post prediction gain changes prior to any neuron separation. Further separation of neurons’ characteristics and classification occurs in subsequent figures.

      (2) The overall picture painted by Figure 2 suggests a correlation analysis should be carried out to search for a general property: prediction score gain vs firing activity index. Are those two variables considered significant by Pearson correlation? Did the authors try that and it didn't work, so they did these analyses? If there is no significant correlation, what does a more detailed look at those two variables on an x-y plot teach us? I suggest considering showing such a plot to readers, at least in a Supplement.

      We thank the reviewer for this feedback and have incorporated this approach in our revised manuscript. Specifically, we performed a correlation analysis between prediction gain change and modulation index (formerly called firing activity index). We uncovered a significant positive correlation between variables, suggesting that task engagement modifies ACC→CA1 communication (Figure 2C).

      (3) (No additional comments).

      (4) This weakness may be able to be addressed by doing a correlation of depth (LFP amplitude of sharp wave) versus the firing rate ratio. For example, the threshold used for deep may have been such that it reduced the number of detected deep neurons, but if a general relationship between depth and degree of FRR is found, it can remove issues from this difference in statistical power.

      We thank the reviewer for this comment. We were unable to perform a reliable link between correlation depth and modulation index. LFP amplitude can vary substantially between tetrodes due to differences in electrode impedance, placement, and recording conditions, making direct comparisons across animals difficult. While within-animal analyses could largely circumvent these issues, many recording sessions did not include tetrodes spanning the full superficial-to-deep CA1 axis, preventing a reliable assessment of this relationship.

      (5) I would give it a different name - "opto-modulation index" or "opto firing rate ratio" perhaps. This would make it clear that you are not measuring task-based modulation of firing.

      We have since modified our wording to improve clarity. Specifically, we renamed “Firing Rate Ratio" throughout the manuscript to "Modulation Index". We also relabeled the Figure 5D Y-axis as "Z-scored Firing Response" to more accurately reflect the plotted data and avoid confusion

      (6) Specifically: can the post-opto lag of v-type versus wide-waveform and PV-type neurons be analyzed? Are the V-type neurons increasing firing before the others decrease? What about cross correlograms between v-type and pyramidal neurons, either at baseline or post-stim?

      We appreciate this feedback. We have added a figure showing differences in response lags to the optostimulation. We demonstrate that V-Type interneurons clearly fire prior to PV and pyramidal cells (Figure 6—figure supplement 1E&F). Additionally, we performed cross-correlogram analyses to examine whether V-type interneurons exhibited consistent temporal relationships with PV interneurons or other CA1 neurons. However, V-type and PV interneurons were sparse throughout our recordings, with most sessions containing two or fewer identified interneurons, which limited our ability to perform meaningful cross-correlation analyses. Nevertheless, we examined the available recordings but found no consistent evidence of correlated firing between interneuron classes or between V-type and pyramidal neurons.

      (7) Can the authors discuss the candidate pathways for connectivity from ACC to CA1?

      We thank the reviewer for this feedback. We have since added discussions on ACC-to-CA1 connectivity and discussed possible relay brain regions between ACC and CA1.

      (8) There is mounting evidence about the role of sleep oscillatory events playing homeostatic roles, not only memory-based. The authors bring this up, but do not offer it as an explanation for their findings, but I believe they probably should. For example, Norimoto et al 2018 cited by the authors. Also, Gulati/Gunguly et al 2017 Nature Neuroscience suggests downscaling as a default NonREM activity. Gulati and also Roux/Buzsaki NatNeuro 2017 show that certain privileged or tagged neurons can be protected from this. This therefore reflects that default activity in nonREM may have a homeostatic role, but then learning may alter that default. I believe this should be discussed as a possible reason for the dissociation between ACC and CA1, the authors observe after CFC.

      In more detail, the authors state that CFC worsened ACC ability to predict ripple spike rate vectors. The authors suggest this may reflect "worsened" communication from ACC to HPC. It could also reflect a SHIFT (not worsening) in the information state of the hippocampus, where ripples reflect novel information and/or information coming from other brain regions. Essentially, the novel information may out-compete usual information flows. For example, ACC may be a default "feeder" into ripples (for example as part of default mode network) when there was no recent highly salient information, but under non-default conditions such as after CFC, ripple content may be fed from other sources (be they internal or external to the HPC). I believe this should be discussed.

      For example, were CA1 sup task inactive neurons basically DMN-active neurons? Figure 2 shows neurons with the highest pre-training ACC prediction were the ones that dropped the most in training - again suggesting these neurons may be tuned to internal or default dynamics rather than CFC (or other novel experiences).

      This shift from a default communication mode to a more experience-based one should probably be discussed as an alternative explanation, rather than simply "worsening" of communication.

      We thank the reviewer for this feedback and their recommendation for possible alternate explanations. We have incorporated many of the listed citations and ideas they discussed into our discussion section proposing homeostatic downscaling and a shift in the default mode network as possible explanations for our results.

      Minor Weaknesses:

      (1) Introduction Line 52: "during replays" should probably be "during replay events".

      Changed.

      (2) Introduction Line 76: "how communications" should be "how communication"

      Changed.

      (3) 200-0, 400-200, 600-400 time bins are a bit unclear in Figure 1f. Are they really negative times, rather than positive? Perhaps negative signs could be put into the legend of 1f, or the time windows can be shown on the left side of 1e. Or potentially 1f could use the same -0.6, -0.4, etc as 1e so readers understand they are linked (if I understand correctly).

      They are negative in the sense the occur before the ripple event. Figure 1e now displays the negative signs.

      (4) I don't believe the methods describe how many tetrodes are put into the ACC. It would seem this should be put in the ACC portion of the "Stereotaxic surgery" section.

      We now clearly explain the number of tetrodes (8) in the "Stereotaxic surgery" section.

      (5) Results line 125: "learning induced" should be "learning-induced".

      Changed.

      (6) In terms of display, deep and sup are swapped in the various figures in terms of which is shown first/second (at least for readers assuming left is first). I suggest putting deep first or sup first in all figures. To me, sup first seems more natural, but homogeneity seems best regardless. This will help readers easily track results.

      We adjusted the figures so that superficial is typically displayed first with some exceptions. For example, in Figure 3A, CA1deep is shown first to preserve the anatomical (dorsal-ventral) relationship.

      (7) Results line 178: "optogenetics stimulation" should be "optogenetic stimulation".

      Changed.

      (8) Results line 180: "upon stimulations" should be "upon stimulation".

      Changed.

      (9) Results line 180: "to different capacities" could be "to different degrees" or "in different manners".

      Changed.

      Reviewer #2 (Recommendations for the authors):

      The sleep scoring procedure is not described clearly. The text references delta waves and ripple oscillations, but the accompanying citation (Wang et al. 2015) does not use such a procedure. Since the post-task rest sessions are called "sleep sessions", there is some confusion about whether the data was restricted to slow wave sleep or not. If data from each sleep session were taken without restricting to actual sleep, that would be problematic because animals may be less likely to sleep immediately following fear conditioning, which could introduce some sleep/wake bias into the comparisons. In particular, the relationship between cortex and ripple activity has been reported to dramatically change between awake and sleep states (Tang & Jadhav, 2019). I am not including this point in the public review in case the data was in fact restricted for sleep, and it simply needs to be clarified in the text.

      Thank you for this feedback. The recordings were in fact restricted to sleep. We have added text to the manuscript to make this clearer. Moreover, we added further discussion on how sleep was calculated.

      There is a puzzling paragraph in the discussion, arguing that "Here, we add to this understanding with CA1sup neurons having a diminished role in fear memory formation". Sparse task-related activity in CA1sup does not imply that CA1sup is not involved in memory. Indeed, while the median of CA1sup neurons' firing rate ratio was below zero, there is a substantial proportion of neurons that are recruited, and these could be extremely important for memory. Most studies on reactivation and replay would only concentrate on cells sufficiently active in the task, and observe whether these patterns of activity are enhanced in post-task sleep.

      We thank the reviewer for this feedback and have removed the text claiming sublayer difference in fear conditioning.

      If the ACC is indeed inhibiting the CA1sup pyramidal cells through V-type interneurons, then one would expect the average GLM weights predicting the activity of those best-predicted CA1sup cells to be negative. If that is true, that could nicely tie the prediction effect to the optogenetic results, demonstrating that ACC's relationship to pyramidal cells is inhibitory in natural conditions as well.

      We thank the reviewer for this suggestion. Unfortunately, our primary GLM analysis did not properly save weight coefficients to perform such analyses. To address this as closely as possible, we performed a preliminary analysis using a modified version of our GLM to examine whether the coefficients predicting CA1sup pyramidal neuron activity exhibited a bias toward negative weights. While this modified analysis did not generate coefficients directly comparable to the prediction gain values reported in the manuscript, it allowed us to assess whether an overall difference in coefficient sign was evident between CA1 sublayers. We found no significant bias toward negative coefficients and no clear differences between CA1sup and CA1deep neurons. Although this result does not provide additional support for an inhibitory relationship under natural conditions, it does not necessarily contradict our optogenetic findings. GLM coefficients quantify statistical dependencies between neural activities and reflect not only direct interactions but also indirect network effects, shared inputs, and the model structure. Consequently, the sign of a GLM coefficient should not be interpreted as a direct measure of whether the underlying synaptic relationship is excitatory or inhibitory.

      Figure 1f is strangely missing comparisons for positive delays. If such windows were to be included and if the reactivation gain is lower for them, that could really drive home the point that communication takes place in the ACC->CA1 direction more than in the CA1->ACC direction.

      Our goal of this study was to examine how incoming information from the cortex may differentially drive CA1 sublayer activity. While we think examining the reverse direction offers a compelling future direction, it was beyond the scope of our present manuscript.

      There is some confusion about the N-s. There's a total of 190 CA1 neurons (Figure 2a legend). 21 of them are deep, and 77 are sup (Figure 3b legend), so presumably 92 would be neither. The legend of Figure 4b agrees with this: 24 task-active CA1sup and 53 task-inactive CA1sup cells, while in CA1deep, there were 14 task-active and 7 task-inactive cells, but in Figure 3c, there is a comparison of N=24 CA1deep cells and n=94 CA1sup cells (so 72 neither).

      For one dual-site animal, the CFC recording file was corrupted, while the pre- and post-training recordings remained intact. As a result, this animal was included only in the GLM analyses, leading to slight differences in sample size across analyses. We have clarified these sample sizes in the revised manuscript and highlighted this discrepancy in the Methods section.

      There appears to be a typo on line 213, the reference should be to Figure 6f.

      Changed.

      Reviewer #3 (Recommendations for the authors):

      To what extent do the authors think that this is learning dependent, as opposed to stress dependent? Not that I want an experiment here, but it is important to note that both of these regions are very much involved in stress responses. Do you think you would get the same result with a purely appetitive learning paradigm? Or is this specific to stress? It might be nice to add this to the discussion.

      We thank the reviewer for this raising this point of discussion. Future experiments utilizing appetitive behavioral tasks could help address these important questions. While we have avoided speculating too much, we agree it is valuable to acknowledge this possibility. Accordingly, we have added a brief discussion point to at least call attention to this possibility for the reader.

      “Another caveat to mention is that both the ACC and CA1 are involved in the stress response (Kim et al., 2015; Lamotte et al., 2021). Future experiments utilizing appetitive learning paradigms, rather than the aversive contextual fear conditioning used here, will help disentangle learning-related remodeling of ACC→CA1 communication from changes driven by stress.”

      There are a number of typos that confuse the message - for example, in the abstract, the authors say that ACC suppresses superficial CA1 interneurons. This seems most likely an error - I think the authors mean superficial CA1 neurons? Or PV interneurons? Similar errors exist throughout, as well as odd combinations of bold and italics across and within words, etc. Overall, especially in consideration of my final main point above regarding clarity of the text, it would be good to have a proper proofread to make sure the text is as clear as possible

      We thank the reviewer for identifying these issues. The specific error in the abstract has been corrected, and we have since carefully proofread the revised manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Okazaki et al. showed flickering stimuli to patients with unilateral spatial neglect (USN) and measured EEG responses. They compared this with another patient group (post-stroke, but no USN) and healthy controls. The author's rationale was to entrain intrinsic brain rhythms using the flicker of different frequencies (3-30 Hz). Effects found unique to the 9-Hz stimulation condition differentiate USN patients from the other groups, leading them to conclude that USN can be characterized by increased hemispheric alpha asymmetry, driven by a relatively increased response in the intact hemisphere.

      Strengths:

      This study is principled empirical work that benefits from access to special patient groups of considerable size (about 60 stroke patients in total, and 20 USN). The authors use state-of-the-art established methods to (1) deliver and (2) quantify the responses to the flicker stimulation in the EEG recordings. In addition, they use phase-coupling measures to investigate cross-frequency coupling (here: alphagamma) and a measure of directed connectivity between brain areas, transfer entropy. The results are supported by means of simulations using a coupled oscillators model.

      Weaknesses:

      In my eyes, the major conceptual weakness of the study is that the authors make the a priori assumption that the flicker stimulation entrains intrinsic brain rhythms, especially alpha (9 Hz). To date, there is no direct (and only equivocal indirect) evidence that alpha rhythms can be entrained with periodic visual stimulation. In the present study, the assumption of alpha entrainment permeates some analytical decisions - where it would be possible to separate stimulus-driven from intrinsic rhythms more strongly than is currently the case, potentially yielding deeper insights into the oscillopathy of USN - and, ultimately, the interpretation of the results. Another potential issue to consider here is the analysis of gamma rhythms in EEG data, absent a control of miniature eye movements, a known problem (YuvalGreenberg et al., 2008, https://doi.org/10.1016/j.neuron.2008.03.027) that may be exacerbated here, given that USN patients could show different auxiliary gaze behaviour.

      We thank Reviewer #1 for the careful and constructive evaluation of our study, and for recognizing the strengths of the patient cohort and our combined empirical and computational approach. We also appreciate the reviewer’s concern that our original wording could be read as assuming that flicker stimulation necessarily entrains intrinsic alpha rhythms. In the revised manuscript, we have clarified that our interpretation is based on frequency-specific stimulus-locked responses and model-based inference, rather than on an a priori assumption of entrainment. We have also revised the relevant parts of the Introduction and Discussion to distinguish more clearly between stimulus-locked responses and intrinsic oscillatory dynamics. In addition, we have expanded our discussion of the potential influence of miniature eye movements on gamma-band activity and PAC. These points are addressed in detail in our responses to the Recommendations for the Authors below.

      Reviewer #2 (Public review):

      This study investigates how altered neural oscillations may contribute to unilateral spatial neglect (USN) following right-hemisphere stroke. By combining steady-state visual evoked potentials (SSVEPs), phase-amplitude coupling (PAC), transfer entropy (TE), and computational modeling, the authors aim to show that USN arises from disrupted hemispheric synchronization dynamics rather than simply from lesion extent. The integration of empirical EEG data with a mechanistic model is a major strength and offers a valuable new perspective on how frequency-specific neural dynamics relate to clinical symptoms.

      The work has several notable strengths. The combination of experimental and modeling approaches is innovative and powerful, and the findings provide a coherent mechanistic framework linking abnormal neural entrainment to attentional deficits. The study also provides concrete evidence to support the potential for frequency specific neuromodulatory interventions, which could have translational relevance.

      At the same time, there are areas where the evidence could be clarified or contextualized further. The manuscript would benefit from more detailed characterization of lesions, since differences in lesion topography (white vs. gray matter, occipital vs. parietal areas) could greatly improve our understanding of the physiopathology causing unilateral spatial neglect and the altered neural oscillations reported. Methodological choices, such as focusing analyses on occipital electrodes rather than parietal sites, and the potential influence of volume conduction in transfer entropy analyses, also need clearer justification/elaboration. In addition, while the authors report several neural metrics, it is not always clear why SSVEP power was chosen as the primary correlate of clinical severity over other measures. More broadly, the manuscript would be strengthened by clearer definitions of dependent variables and reporting of software and toolboxes used.

      Overall, the study makes a significant contribution by demonstrating that USN can be conceptualized as a disorder of disrupted oscillatory dynamics. With some clarifications and expansions, the paper will provide readers with a clearer understanding of both the strengths and the limitations of the evidence, and it will stand as a valuable reference for future work on oscillatory mechanisms in stroke and attention.

      We thank Reviewer #2 for the positive assessment of our integrated empirical and computational approach, and for highlighting the potential contribution of our findings to understanding oscillatory mechanisms in USN.

      In response to the reviewer’s comments, we have added new supplementary figures showing lesion overlap maps and lesion-volume analyses (Supplementary Figure 1), additional analyses related to electrode selection and SSVEP topography (Supplementary Figure 2), and clinical-correlation analyses of hemispheric imbalance measures (Supplementary Figure 3). We have also revised the manuscript to better contextualize lesion topography and lesion extent, clarify the rationale for the occipital-electrode and clinical-correlation analyses, and provide additional methodological details. These issues are addressed in detail in our point-by-point responses to the Recommendations for the Authors below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I found your manuscript well-written and hence easy to follow. Allow me to provide some recommendations to tackle the "weaknesses":

      (1) I have mentioned that the entrainment assumption is not warranted based on the current evidence, but I also think that the analysis and interpretation is critically constrained by this. The way the analysis is carried out conflates stimulus-driven and intrinsic brain rhythms, especially in the alpha band. In other words, spectral representations of the EEG data will likely be dominated by natural alpha rhythms, whereas the stimulus-driven signals could be accentuated by a different analysis, time-locked to the stimulation. This is explained in greater detail in Keitel et al. (2019, https://doi.org/10.1523/JNEUROSCI.1633-18.2019). Looking into alpha and stimulus-driven responses separately, without the assumption of entrainment, may actually allow a more complete picture of the impact of USN because the stimulus driven SSVEPs are taken to indicate different cortical processes than alpha (see e.g., Duecker et al., 2021, https://doi.org/10.1523/JNEUROSCI.3134-20.2021, though for gamma). Importantly, this all does not exclude the possibility that alpha rhythms were indeed entrained here, but given the current situation, this should be an outcome of the study rather than an a-priori assumption. I suggest re-framing the manuscript this way.

      We agree that the current wording could be read as presuming alpha entrainment. We have revised the framing and terminology to reflect that entrainment is inferred from the 9-Hz–specific stimulation effect and the absence of a corresponding hemispheric bias at rest, with further support from the resonance mechanism demonstrated by our computational model. We also point readers to the Discussion “Alpha frequency-specific hemispheric bias and entrainment in USN patients”, where we explain why the present 9-Hz–specific findings are interpreted in terms of alpha-range resonance/phase alignment and how this is expressed in EEG. We revised the Introduction SSER description (p.4) to:

      “SSERs are elicited by rhythmic sensory stimulation and provide a measure of stimulus-locked neural responses. When the stimulation frequency is close to the system’s intrinsic resonance frequency, these responses may include an entrainment component, reflecting the alignment of endogenous oscillations to external input (Pikovsky, Rosenblum, and Kurths 2003; Okazaki et al. 2021)”

      We have also revised the Introduction hypothesis statement (p.5) to avoid implying an a priori entrainment assumption, replacing it with:

      “We hypothesized that frequency-specific stimulus-locked synchronization dynamics in response to rhythmic stimulation would be selectively disrupted in one hemisphere, resulting in an interhemispheric imbalance in USN.”

      Finally, to ensure consistent framing in the Discussion, “Alpha frequency-specific hemispheric bias and entrainment in USN patients”, we have revised the opening sentence (p.20) as follows:

      “We observed a hemispheric bias in stimulus-locked responses to flickering stimuli in USN patients in the alpha range, which corresponds to the natural frequency of the visual system (Rosanova et al. 2009; Okazaki et al. 2021).”

      (2) Transfer Entropy is used as a measure of directed connectivity and applied to narrow-band filtered EEG signals. Methodological issues have been pointed out with regard to that (Daube et al., 2022, https://doi.org/10.48550/arXiv.2201.02461). Has this been considered?

      We thank the reviewer for raising this important methodological concern. As Daube et al. (2022) noted, Transfer Entropy (TE) can be overestimated when applied to narrow-band signals with strong autocorrelation. However, in our analysis, TE was computed not from the narrow-band 9-Hz waveform itself but from its instantaneous amplitude (amplitude envelope), which fluctuates nonperiodically on a slower timescale, thereby reducing the risk of spurious causality driven by sinusoidal autocorrelation. We clarified this explicitly in the Methods, “Transfer entropy (TE)” (p.8) by adding:

      “After applying an 8.5–9.5 Hz FIR bandpass filter to the EEG responses to 9-Hz flickering stimuli, we extracted the instantaneous amplitude (amplitude envelope) from the analytic signal using the Hilbert transform. Because this amplitude envelope fluctuates nonperiodically at a slower timescale than the carrier 9-Hz oscillation, it provides a broadband measure of signal dynamics while minimizing the strong autocorrelation inherent in narrow-band periodic signals that can spuriously inflate TE estimates (Daube C et al., 2022).”

      We have also clarified the relationship between TE and volume conduction in the same section:

      “Importantly, because TE evaluates time-lagged prediction (from Y(t) to X(t+τ)), it is not designed to capture zero-lag common-source correlations (i.e., volume conduction) and therefore characterizes directed, nonzero-lag dependencies rather than an instantaneous coupling.”

      Finally, we have added an explicit interpretation emphasizing the direction-specific nature of the effect in the Discussion, “Biased information transfer in USN patients” (p.23):

      “This directional asymmetry argues against spurious overestimation of TE due to autocorrelation or volume conduction (Daube, Gross, and Ince 2022), because such pseudo-causal effects would be expected to manifest more symmetrically in both directions. In addition, by computing TE from the amplitude envelope of the 9-Hz activity, we reduced the influence of strong autocorrelation inherent in narrow-band oscillatory signals and thereby minimized the conditions that Daube et al. identified as leading to TE overestimation. Taken together, these points suggest that the observed TE asymmetry is unlikely to be explained solely by methodological artifacts and may reflect a genuine directional imbalance in interregional communication following right-hemisphere damage.”

      (3) If a closer control of miniature eye movements is not possible, I suggest removing any gamma analysis, or at least prominently mentioning the caveat that gamma activity may be contaminated by eye movement artifacts.

      We agree that the contribution of miniature eye movements cannot be completely excluded. However, we consider it unlikely that the hemispheric asymmetry in alpha– gamma PAC reported in this study mainly arises from eye-movement artifacts, for the following reasons. First, SP (saccadic spike potential)-related PAC would require saccade timing to be tightly phase-locked to the 9-Hz cycle. However, SPs are time-locked to saccade onset and typically cluster around 200–300 ms after stimulus onset (Yuval-Greenberg et al., 2008; Keren et al., 2010). Thus, it is unlikely that they would be consistently phase-locked to the 9-Hz cycle (≈111 ms), and even modest temporal jitter would markedly blur PAC. Second, because SPs are brief spike-like transients with broadband high-frequency components, periodic SP contamination would be expected to yield a broadband gamma profile, rather than the relatively narrow band (35–45 Hz) observed here. In addition, we directly compared gamma-band power (35–45 Hz) at O1 and O2 during 9-Hz stimulation using the same gamma range as in the PAC analysis and found no significant hemispheric differences in any group (Author response image 1), arguing against a systematic unilateral increase in gamma power driven by asymmetric saccade behavior. Taken together, these considerations make it difficult to attribute the observed alpha–gamma PAC asymmetry primarily to SPs arising from eye movements. Nonetheless, residual eye-movement artifacts cannot be completely ruled out, and we now explicitly state this limitation in the Discussion, “Limitations” (p. 25) by adding:

      “Fourth, the hemispheric bias in alpha–gamma PAC should be interpreted in light of potential contamination from miniature saccades (Yuval-Greenberg, Tomer, Keren, Nelken, & Deouell, 2008). However, several observations make it unlikely that such artifacts are the primary source of the effect. For saccadic spike potentials to account for the PAC under 9-Hz stimulation, they would need to occur in a highly periodic and phase-locked manner relative to the 9-Hz cycle. This scenario is unlikely given the stimulus-locked dynamics of miniature saccades. (i.e., post-stimulus inhibition followed by a rebound around 200–300 ms) (Yuval-Greenberg & Deouell, 2009). Moreover, SP-related contamination would be expected to yield a broadband gamma profile (~20–90 Hz) (Keren, Yuval-Greenberg, & Deouell, 2010; Yuval-Greenberg & Deouell, 2009), rather than the relatively narrow band (35–45 Hz) observed here. In addition, we found that gamma-band power in the 35–45 Hz range did not show any hemispheric difference between O1 and O2 during 9-Hz stimulation (data not shown).”

      Author response image 1.

      Hemispheric differences in gamma-band (35–45 Hz) power during 9-Hz flicker stimulation. Left (O1) − Right (O2) gamma power did not differ from zero in any group (non-USN: p = 0.45; USN: p = 0.34; healthy: p = 0.35). Error bars represent 2 standard errors of the mean.

      (4) Please provide more methodological detail on the resting state recordings. When, how, and under which circumstances were these recorded?

      Thank you for pointing this out. We clarified when the resting-state interval was taken in the Methods, “Steady-state visual evoked potential (SSVEP)” (p.7) by adding:

      “For the resting-state interval, power was estimated using the same procedure from the ‘off’ interval immediately preceding the 3-Hz flicker block (see Fig. 1)”

      (5) Provide power spectra of the EEG data for illustration - ideally for resting state and stimulation conditions. These should allow the reader to visually evaluate the effects of different stimulation frequencies, as well as differences between participant groups.

      Figure 2 already presents spectra normalized to the resting-state baseline. To make this explicit for readers, we clarified this point in the Figure 2 caption (p.11) by adding:

      “Spectra are normalized to the baseline from the resting-state interval.”

      Reviewer #2 (Recommendations for the authors):

      (1) The authors indicate L.608 "the precise extent and topography of brain lesions could not be fully homogenized across patients", but the manuscript would highly benefit from any additional detail that could be obtained from characterization of the lesion sites. At minimum, it would be important to indicate for USN and non-USN groups whether lesions predominantly affected white matter or gray matter, and whether occipital versus parietal cortices were involved (e.g., using MRI atlas templates). If this coarse information can be obtained, the authors could test whether any of these anatomical details can distinguish are different between USN and nonUSN patients. This information could help the reader evaluate whether the reported neural asymmetries might be driven by lesion topography rather than purely by oscillatory dynamics.

      We understand this comment as raising the important concern that lesion location and lesion volume may differ between the USN and non-USN groups and could contribute to the observed neural asymmetry. We agree that the presence of USN and the alteration of oscillatory dynamics should be interpreted in relation to which regions and networks are damaged, and to what extent. In response to the reviewer’s suggestion, we generated lesion overlap maps based on the available structural images and additionally quantified lesion volume for each patient (replaced Supplementary Figure 1). This analysis confirmed that lesion volume was significantly larger in the USN group than in the non-USN group. Thus, lesion volume is an important anatomical factor to consider when interpreting group differences in USN and neural responses.

      At the same time, the present results suggest that lesion volume and coarse lesion topography alone are not sufficient to explain the 9 Hz-specific imbalance in interhemispheric synchrony. Interhemispheric synchrony depends on distributed network functions involving multiple cortical and subcortical regions and their connecting pathways, and similar functional imbalances may arise from different patterns of anatomical damage. Moreover, although the newly added lesion maps confirmed more extensive lesions in the USN group, lesions in both groups predominantly involved the right MCA territory, and lesion extent and location varied across patients. Importantly, the hemispheric asymmetry in neural responses emerged selectively in the 9 Hz condition, whereas SSVEP responses at other frequencies were largely balanced between hemispheres. If lesion volume or coarse lesion topography alone were sufficient to explain the effect, one might expect a more uniform reduction across frequencies or a simpler pattern corresponding to lesion extent.

      Accordingly, we do not treat lesion topography and oscillatory dynamics as competing explanations. Instead, we regard them as hierarchically related: anatomical damage alters network components, and this in turn gives rise to a frequency-specific imbalance in synchronization capacity. To clarify this point, we replaced the previous Supplementary Figure 1 with a new figure showing lesion overlap maps and lesion volume information, and revised the Discussion section “Distinct neural responses in non-USN and USN patients” (p. 24) as follows.

      “To further characterize the anatomical background of these group differences, we generated lesion overlap maps and quantified lesion volume in the USN and non-USN groups (Supplementary Figure 1). Both groups predominantly showed lesions involving the right MCA territory, but lesion volume was significantly larger in the USN group than in the non-USN group (USN: 72,419 ± 73,778 mm<sup>3</sup>; non-USN: 11,739 ± 23,907 mm<sup>3</sup>; Mann–Whitney U = 361.0, p = 1.41 × 10<sup>-5</sup>). These anatomical differences indicate that lesion extent is an important factor associated with USN. However, they do not by themselves fully explain the frequency-specific neural effect observed here. Importantly, this frequency specificity coincides with the intrinsic alpha frequency of the visual system. This correspondence suggests that the present finding may not simply arise from lesion location or lesion volume alone, but may instead reflect a more complex mechanism involving selective functional disruption of oscillatory networks. From this perspective, lesion topography and oscillatory dynamics should be regarded not as competing explanations, but as different levels at which the same pathological condition can be understood. The key question is which network components are affected and how their dysfunction gives rise to the frequency-selective hemispheric imbalance in synchronization capacity at 9 Hz. Thus, even if USN and non-USN patients differ in lesion extent and aspects of lesion topography, this does not undermine the present interpretation, but rather highlights the need to examine how anatomical damage relates to frequency-specific network dysfunction.”

      We have also referred to our computational account and clarified its implication for the functional mechanism in the same subsection (p. 24):

      “Our computational model illustrates a plausible mechanism for such a process. When two coupled oscillators sharing the same intrinsic alpha frequency are connected via asymmetric interhemispheric couplings, the model selectively produces an imbalance in synchrony at the resonant frequency, whereas responses at non-resonant stimulation frequencies remain relatively balanced between hemispheres. This model result suggests that post-lesion asymmetry in interhemispheric coupling may bias alpha-band information processing (e.g. phase-dependent sampling/synchrony) between hemispheres, and may consequently manifest as systematic biases in perceptual and attentional allocation.”

      We have also revised the Limitations (p. 25) to acknowledge that the present lesion analyses characterize the distribution and extent of lesions across groups, but do not directly identify which anatomical network disruptions give rise to the 9 Hz-specific imbalance in interhemispheric synchrony.

      “Second, although we added lesion overlap maps and quantified lesion volume, these analyses primarily characterize where and how extensively lesions were distributed across the two patient groups. They do not directly identify which anatomical network disruptions give rise to the 9 Hz-specific imbalance in interhemispheric synchrony. Future studies with larger cohorts will be necessary to combine detailed lesion-symptom mapping and assessments of white-matter disconnection with neural synchrony analyses to determine which anatomical network disruptions lead to frequency-specific alterations in oscillatory dynamics after stroke.”

      (2) The rationale for focusing on O1 and O2 electrodes should be clarified. Given that neglect is classically associated with parietal dysfunction, one would expect analyses of parietal electrodes to be informative. Could the authors justify their choice, and possibly report whether similar effects were (or were not) observed at parietal sites?

      Our primary SSVEP analyses focused on O1 and O2 because flicker stimulation is designed to drive the visual system, and the fundamental SSVEP component is typically maximal over occipital electrodes in healthy participants (Norcia et al., 2015). We also note that hemispheric differences at central and frontal electrodes are already shown in Fig. 4 and described in the Results, demonstrating no significant left–right differences at these sites across any stimulation frequency conditions, including 9 Hz. In addition, we added a supplementary figure showing the scalp topography of the mean fundamental-frequency SSVEP power averaged across all stimulation conditions, which displays the expected occipital maximum and thus a typical SSVEP spatial profile (Supplementary Fig. 2, shown below). We added the rationale and the reference in the Methods, “Steady-state visual evoked potential (SSVEP)” (p. 7) section by adding:

      “SSVEP power at the fundamental (stimulated) frequency was then extracted for subsequent analyses. Because the fundamental SSVEP component is typically maximal over occipital electrodes in healthy participants (Norcia et al., 2015), we evaluated occipital electrodes (O1/O2) as primary sites. The scalp topography of the mean fundamental-frequency SSVEP power averaged across stimulation conditions is shown in Supplementary Fig. 2.”

      (3) For the mutual information and transfer entropy analyses, the potential influence of volume conduction should be acknowledged. Numerous studies mitigate this issue by applying source reconstruction or connectivity metrics that are insensitive to zerolag correlations. Even if the present study did not use such approaches, the authors should discuss the extent to which volume conduction might confound their results, and ideally provide some justification for why their findings remain valid.

      Regarding directed connectivity, our primary analysis uses Transfer Entropy (TE), which evaluates time-lagged prediction and is therefore, by definition, not directly sensitive to zero-lag common-source correlations. Consistently, our key finding is direction-specific (feedforward only), which is not readily explained by symmetric, zero-lag relationships typical of volume conduction. We stated this explicitly in the Methods, “Transfer entropy (TE)” (p.8) by adding:

      “Importantly, because TE evaluates time-lagged prediction (from Y(t) to X(t+τ)), it is not designed to capture zero-lag common-source correlations (i.e., volume conduction) and therefore characterizes directed, nonzero-lag dependencies rather than instantaneous coupling.”

      We have also clarified the direction-specific logic in the Discussion, “Biased information transfer in USN patients” (p.22):

      “This hemispheric difference appeared only in the Feedforward (visual-to-frontal) direction, and no such difference was observed in the Feedback (frontal-to-visual) direction. This directional asymmetry argues against spurious overestimation of TE due to autocorrelation or volume conduction (Daube C, Gross J and Ince RAA, 2022), because such pseudo-causal effects would be expected to manifest more symmetrically in both directions. In addition, by computing TE from the amplitude envelope of the 9-Hz activity, we reduced the influence of strong autocorrelation inherent in narrow-band oscillatory signals and thereby minimized the conditions that Daube et al. identified as leading to TE overestimation. Taken together, these points suggest that the observed TE asymmetry is unlikely to be explained solely by methodological artifacts and may reflect a genuine directional imbalance in interregional communication following right-hemisphere damage.”

      (4) In line 579, the authors write 'patients with USN may have larger lesions than those without USN.' It would be useful to clarify whether this claim is supported by the present data (i.e., lesion size comparisons between the two patient groups) or whether it reflects prior literature. If based on the current sample, please provide the corresponding statistical evidence.

      This issue has been addressed as part of our response to comment (1), with corresponding revisions made to the Discussion (p. 24), the Limitations (p. 25), and Supplementary Figure 1.

      (5) The dependent variable for SSVEP analyses is not clearly described, which makes it difficult to understand the results (e.g., L.214, L285-297). It seems that for all stimulation conditions, one value was extracted (and then compared between O1 and O2 electrodes), but is that the amplitude at the stimulated frequency (e.g., for a visual stimulation at f Hz, comparing the amplitude of the power spectrum at f Hz between O1 and O2)?

      Thank you for pointing this out. We clarified that the dependent variable for the SSVEP analyses was the power at the fundamental (stimulated) frequency f for each condition. We added this explicitly to the Methods, “Statistical analysis” (p.9):

      “The SSVEP power at the fundamental (stimulated) frequency f was analyzed for each condition as the dependent variable...”

      (6) Most analyses are described relatively precisely mathematically, but no toolbox or software is mentioned. Were they implemented using custom scripts? If toolboxes/software was used, the authors should indicate which ones and their versions, and add references. For instance, I might be wrong, but I guess the linear mixed model analyses were not performed using custom code. MATLAB is mentioned for the simulations, but no version is indicated. This is particularly important for reproducibility since the authors did not provide direct access to their scripts.

      The EEG preprocessing, PAC, and cluster-based permutation tests were conducted using the FieldTrip toolbox integrated with custom MATLAB scripts. Linear mixed-effects models were performed in SPSS. We added this information to the Methods, “Preprocessing” (p.7):

      “All EEG analyses were implemented using custom scripts in MATLAB (MathWorks, Natick, MA, USA) and the FieldTrip toolbox (Oostenveld R et al., 2011).”

      We have also specified the tools used for the linear mixed-effects models and the cluster-based permutation test in the Methods, “Statistical analysis” (p.9):

      “…using a linear mixed model in SPSS… a cluster-based permutation test (Maris E and Oostenveld R, 2007) using FieldTrip.”

      (7) The rationale for correlating BIT scores specifically with SSVEP power (rather than with the hemispheric imbalance measure, or with MI/TE results) is not clearly articulated. The results section highlights several neural metrics (SSVEP imbalance, PAC, TE), so it would strengthen the manuscript if the authors explained why SSVEP power was prioritized for correlation analyses. Is there a theoretical or empirical reason for this choice?

      While our main group-level results demonstrate a frequency-selective hemispheric imbalance, Fig. 5 suggests that the lesioned and intact hemispheres are not necessarily simple mirror images of each other and can show distinct response profiles. Therefore, to clarify which hemisphere’s response changes are associated with symptom severity, we assessed correlations with BIT using hemisphere-specific SSVEP power. At the same time, it is also informative to directly test the extent to which symptom severity covaries with hemispheric imbalance metrics (i.e., left–right differences in SSVEP, PAC, and TE). Accordingly, in the revised manuscript we additionally analyzed correlations between BIT and hemispheric imbalance measures of SSVEP, PAC, and TE, and report these results in the Supplementary Information (Supplementary Fig. 3). We added this clarification to the Results, “Correlation between SSVEP power and USN severity” (p.17):

      “In addition, correlations between BIT and the hemispheric imbalance of SSVEP power are provided in the Supplementary Information (Supplementary Fig. 3A). We also report, as additional exploratory analyses, correlations between BIT and hemispheric imbalance measures derived from PAC and TE (Supplementary Fig. 3B and 3C). None of these correlations reached statistical significance after correction, although the feedforward TE imbalance to the left frontal region showed a marginal uncorrected association with BIT score (r = -0.51, uncorrected p = 0.05).”

      We have also added the following statement to the Discussion, “Biased information transfer in USN patients” (p.23).

      “This interpretation is also consistent with the trend-level correlation shown in Supplementary Figure 3C. Although the correlation did not survive correction for multiple comparisons, patients with a stronger feedforward TE bias toward the left frontal region tended to show higher BIT scores. This observation raises the possibility that asymmetric information transfer from the right visual cortex to the intact left frontal cortex may contribute to compensatory network reorganization associated with milder neglect symptoms.”

      (8) The discussion could be enriched by considering whether the observed abnormal alpha-band entrainment relates to the well-known alpha-lateralization phenomenon in spatial attention. In healthy individuals, covert attentional orienting is typically accompanied by lateralized modulations of alpha power (increased ipsilateral, decreased contralateral). Could the asymmetric alpha entrainment reported here be interpreted as a pathological exaggeration of this mechanism, thereby linking the electrophysiological findings more directly to attentional orientation deficits in neglect?

      We thank the reviewer for raising this valuable point. To integrate our results with the alpha-lateralization literature, we added a new paragraph to the Discussion, “Local hemispheric bias of alpha-band entrainment and attentional dysfunction in USN patients” (p.22) subsection

      “Clinically, USN presents as neglect of the left visual field and is generally thought to reflect a relative dominance of rightward orienting. Because alpha power is often interpreted as reflecting functional inhibition (Worden et al. 2000; Kelly et al. 2006; Foxe and Snyder 2011), classic alpha-power lateralization associated with rightward orienting would typically predict increased alpha power over the right hemisphere and decreased alpha power over the left. However, in our data, right-hemisphere alpha power during stimulation was comparable to that of controls, whereas the intact left hemisphere exhibited stronger stimulus-locked synchronization (SSVEP) and enhanced alpha–gamma coupling. Thus, the present findings do not appear to reflect a simple amplification of tonic spatial alpha-power lateralization. Rather, they suggest a dynamic, temporally selective form of alpha-mediated inhibition that organizes the timing of local excitability. From this perspective, alpha-power lateralization and stimulus-locked entrainment may reflect related but distinct aspects of alpha-based inhibitory control, with the former regulating spatial gating and the latter regulating the timing of sensory processing (Jensen and Mazaheri 2010).

      Moreover, such a timing-based bias may arise even in the absence of explicit flicker. Visual input is continuously sampled under the influence of intrinsic alpha activity, and perceptual sensitivity fluctuates with the phase of ongoing oscillations (Romei et al. 2008; Iemi et al. 2017; VanRullen 2016). As a consequence, processing is relatively facilitated when incoming events coincide with high-excitability phases and relatively suppressed otherwise (Busch, Dubois, and VanRullen 2009; Mathewson et al. 2010). Thus, a hemispheric bias in alpha-band synchronization capacity could bias the timing with which sensory input is sampled across the two hemispheres, even without explicit rhythmic stimulation. This view is also consistent with the possibility that the bias is less apparent during eyes-closed rest with minimal visual input, yet becomes behaviorally expressed in everyday settings where continuous input is sampled in a phase-dependent manner (Landau and Fries 2012; VanRullen 2016). Overall, a hemispheric bias in synchronization capacity within the alpha range likely disrupts the temporally coordinated sampling of visual input required for balanced spatial attention, contributing to the attentional deficits observed in USN.”

      (9) All reported t-tests should include degrees of freedom (df). This is essential for transparency and allows readers to assess the robustness of the statistical results.

      We added the degrees of freedom for the t-tests reported in the Results, “Hemispheric imbalance of TE” (p.15):

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors tackle a long-standing question in developmental theory: given a gene-regulatory network that includes extracellular signaling, which topologies are even capable of transforming an initial spatial profile into a genuinely new pattern? Building on the classical reaction-diffusion framework in one dimension, but imposing biologically motivated constraints, they prove that every one-signal sub-network must be either Hierarchical (H), self-activating (L+), or selfinhibiting (L-). They further demonstrate that only three composite classes of full networks - pure H, a coupled L+ L- "Turing" pair, and an L- module fed by an intracellular positive loop ("noise-amplifying")-can create non-trivial spatial transformations. Analytical criteria and illustrative simulations are provided, together providing a closed taxonomy, which is supposed to be relevant for real systems.

      Strengths:

      (1) Useful classification framework. Reducing a vast number of possible gene circuits to three canonical patternforming motifs is a valuable organizing insight for both theorists and ---experimentalists.

      (2) Practical interpretability. Given a reaction network diagram, one can now decide (assuming the model applies to real systems) whether spatial patterning is even possible, saving experimental effort on in silico screens that could never succeed.

      Weaknesses:

      (1) After the resubmission, I still have concerns regarding the formal definition of "non-trivial transformations" (P1/P2) and its application to noisy or multi-dimensional systems. The criteria rely on counting "new" critical points (maxima/minima). In their response, the authors argue that the diffusion operator instantly smooths discontinuous white noise, allowing critical points to be properly defined. However, this very smoothing process passively generates a landscape of new, smooth local extrema from the initial noise. Consequently, trivial diffusive regularization could inadvertently fulfil the criteria for a "non-trivial" transformation, leaving the definition conceptually problematic.

      That is indeed the case: diffusion alone can generate new concentration maxima and minima; but these would be transient unless there are some self-activatory loops (as we detail over the article). If these concentration maxima and minima are transient in time, they do not count for pattern transformation. In the two version of the article we explicitly stated (in P1 in the introduction when defining pattern transformation) that we only consider resulting patterns that are stable in time. This implies that the concentration maxima and minima that may transiently arise from diffusion alone do not count for the definition of pattern transformation. In the current new version (the third) and after the reviewer’s suggestion, we insist (in red text) that the new concentration maxima and minima need to be stable in time (i.e. non-transient).

      Furthermore, when extending the framework to 2D/3D, the manuscript assumes that starting from a central "spike" will robustly preserve radial symmetry, yielding concentric rings or shells. This overlooks the fundamental nature of macroscopic mean-field models like reaction-diffusion equations. The realization of the final multidimensional pattern depends strictly on the stability of the solution against ubiquitous perturbations (including angular modes) rather than solely on the deterministic symmetry of the initial condition. It remains unclear how the current framework accounts for spontaneous symmetry breaking in cases where these angular modes become unstable, challenging the assumption that radial symmetry will strictly dictate the outcome. We note that the authors' use of noise as an initial condition does not resolve this fundamental issue. Reaction-diffusion equations inherently describe mean-field dynamics, meaning that microscopic fluctuations are continuously present in any real system, regardless of whether explicit stochastic terms are written into the equations. Ultimately, if a symmetric mean-field solution is structurally unstable to these inherent fluctuations, it simply cannot be realized in nature.

      If we understood right the reviewer is saying that there are always fluctuations everywhere and that, thus, the spike initial pattern should also have noise everywhere. In the discussion (and partially in the introduction) we have now added a discussion (in red) on how the resulting patterns possible from such spike-with-noise initial pattern can actually be understood, to a large extent, from those of the homogeneous-with-noise and spike-without noise. In that case, as the reviewer suggests, the resulting patterns do not necessarily have radial symmetry. We have kept the results on spike initial patterns without noise because they are very helpful to understand the pattern transformations possible from spike-with-noise initial patterns. Moreover, as we discuss now, the spatial fluctuations that occur everywhere can be really small compared to the concentration in the spike and, thus, the spike initial pattern without noise can be a good approximation in some cases, at least worth considering.

      (2) Theoretical limitations in the application of Linear Stability Analysis (LSA): I remain uncertain about the framework's reliance on LSA to categorize macroscopic transformations, especially those arising from large initial perturbations (spikes). In their rebuttal letter, the authors justify this by assuming the perturbation remains small over a short time interval. However, because the study aims to describe stationary, asymptotic states, applying a linear approximation that relies on transient t->0 conditions to predict long-term global stability is not fully resolved.

      We are not trying to predict long-term stability and we never intended to. We never claimed that to be the case. We are interested in stable in time spatial patterns but the long-term stability of resulting patterns is not something we intend to see from the LSA. The LSA just provides a necessary (but not sufficient) condition for non-trivial pattern transformations: any network unable to sustain the growth of small perturbations (i.e., any linearly stable network topology) cannot lead to a non-trivial pattern transformation, regardless of the nonlinear terms in the reaction term f(g). From the previous suggestions by the reviewer it could be the case that he/she thinks that since there is always noise everywhere and all the time we can never apply the LSA or that it cannot be applied along time. In fact, we do not try to applied over time (that would make no sense for our purposes). The LSA is only applicable and only informative when applied to the initial pattern (that is at time 0) to see gene network that cannot transform initial patterns into other patterns. As we discuss in the previous version and we further stress in the current one (in red in the LSA), large spikes do not invalidate this approach. Even if the spike would be large, it would only affect cells outside the spike through the diffusion of gene products from the spike. Thus, if a small enough time is considered, large spikes can be considered as small concentration perturbations outside the spike and, thus, in the the worse case scenario our whole approach may not be applicable inside the spike but outside of it (that is reasonable since spikes are by definition narrow).

      Here it is important to stress that many things that used the LSA in the original version of the article do not used it in the current version of the article. In fact, in this article we address two main questions: (i) which gene network topologies can produce non-trivial pattern transformations; and (ii) what can we say about the stationary patterns they produce. We acknowledge that a linear stability analysis alone is not enough to fully characterize the long-term behavior of the system, and this is why we only use LSA as an aid to answer question (i).

      This was not clear enough in the first version of the article but it was explicitly stated in the last version (in the Gene network classification and Linear stability analysis section). This allows us to discard many topologies, but further analysis is still required to assess whether linearly unstable network topologies can actually produce non-trivial pattern transformations.

      It is in this further study that question (ii) comes to play. Here we do not rely on LSA, but instead impose a series of requirements (R1-R5) on the reaction term f(g) that constrain the nonlinear dynamics in a biologically motivated way that prevents pathological behaviors (particularly, boundedness of solutions (R4) and monotonicity of the reaction (R5)). These requirements enable the qualitative analysis of the different unstable network topologies in order to say some things about the possible stationary patterns. This is complemented with numerical simulations and quantitative analysis of some prototypical examples in the supplementary information.

      (3) In the previous round of the review, I suggested that a biomolecular sink, such as A+B -> AB reaction, could break the approach. In their response letter, the authors defend their approach by arguing that such reactions can be accommodated by their abstract constraints (R1-R5) as long as the signs of the Jacobian elements remain invariant. However, the problem I see here is not the sign of the interactions, but the severe loss of spatial homogeneity.

      When a macroscopic initial perturbation (a "spike" of morphogen) is introduced into a domain with a strong bimolecular sink, it will inevitably cause massive local depletion of the consumed substrate near the source. Consequently, the background state of the system will rapidly evolve into a profile with macroscopic spatial gradients long before any spontaneous pattern-forming instability takes over. Mathematically, this dictates that the system no longer possesses a homogeneous steady state, and the Jacobian matrix becomes explicitly space-dependent, which should break the classical LSA approach.

      This criticism seems to be intimately related to the previous one. If we understand correctly, the argument is that in a very non-linear system, such as in the sink described, the spike will rapidly lead to a local change in the concentration of other gene products and that then the LSA is not applicable. This is true but what we care about is whether the initial pattern (that is the system at time zero) is actually stable or not. We care about it because as we explain, pattern transformation is only possible if the perturbation in the initial pattern (spike or noise) is unstable. This is simply a necessary condition (an initial pattern may be unstable to perturbation and still not produce nontrivially transformations). So whether the system will be suitable for a LSA some time after the initial pattern, as the reviewer suggests, is not something we need to know for our classification. Related to the other comments the reviewer may be concerned with whether other perturbations occurring everywhere (and all the time), that is noise, may actually affect the possible resulting patterns (that we described in the new section of the discussion).

      We want to thank the reviewer for this clarification, as we believe we had not fully understood their concern in the previous round of review. In our framework, the bimolecular sink the reviewer suggests corresponds to a three-gene-product network where A and B mutually inhibit each other and both activate AB.

      First it is important to consider that unless something else is specified an A+B→AB system with homogeneous initial pattern (with or without noise) is just not stable: the A and B gene products will decay to zero concentration and AB to a maximal concentration (that would be homogeneous over space if there is no noise). Besides the final stable state is totally stable since with no A or B left the concentration of AB cannot change. We explicitly state in the article (in the LSA section) that we apply the LSA to systems that, when unperturbed, are stable. For other systems we just wait for the system to stabilize and then ask whether pattern transformation is possible from that state if there is some perturbation, where we can now apply a LSA). So strictly speaking the system the reviewer is suggesting is outside the scope of the article and indeed unsuitable for LSA. Nevertheless, the system the reviewer proposes cannot lead to non-trivial pattern formation, as we detail below.

      The reviewer does not specify whether A, B or AB diffuse, we then consider all possibilities. We also assume that A is the gene product in the spike.

      Case in which no molecule diffuses. In this case a spike of A will simply lead to a valley of B (since A reacts with B to deplete B), and a spike of AB in the exact same location of the spike (i.e., no non-trivial pattern transformation occurs). If there is noise, each small fluctuation in the concentration of A or B would lead to a similar fluctuation in AB. Notice this case does not lead to non-trivial pattern transformations (the peaks in the initial pattern and resulting pattern are in the same places). This latter situation in fact we explain in the “Pattern formations from homogeneouswith-noise initial patterns in H networks section.”

      Case in which AB diffuses. Since nothing is promoting the production of A and B, their concentration will inevitably decay to zero and since AB diffuses its concentration on the long-term will inevitably become homogeneous (irrespectively of which initial pattern there may be).

      Case in which only A diffuses. In this case the spike of A will initially lead to a valley of B (since A consumes B) and a peak of AB (since this consumption leads to AB). However, since for each molecule of AB a molecule of both A and B are required and the concentration of B is homogeneous (since B does not diffuse), having more of A around the spike would not lead to more AB in the peak than elsewhere. The concentration of AB would thus become homogeneous over time. The same applies if B is the only molecule that diffuses.

      Case in which A and B diffuse and AB does not. In this case the spike of A will lead to a peak of AB. This would deplete B around the spike but since B can diffuse, new molecules of B would arrive and lead to a further growth in the peak of AB. As a result a stable peak of AB will form (even if A and B will ultimately decay to zero), just around the initial spike of A and both A and B will decay to zero (so no new peaks or valleys form and thus, no non-trivial pattern transformation).

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This is an extensively revised version of a previously submitted manuscript that, as detailed in their 20-page response to the first reviews, satisfactorily addresses the reviewers' comments. In particular, the revised manuscript makes it much clearer how this work fits into and advances the field. The added experiments strengthen the rigor of the manuscript as well. Overall, this paper is ready to go.

      We thank the reviewer for the positive evaluation and for recognizing the improvements made in our revised manuscript.

      Reviewer #2 (Public review):

      The revised manuscript has been substantially improved. The authors have addressed many of my previous concerns through the addition of new data, analyses, and discussion. The characterization of epithelial folding in the ascidian Ciona provides valuable insight into a comparatively less explored morphogenetic system, and the imaging and quantitative analyses are overall compelling. That said, a few important points remain to be addressed.

      We thank the reviewer for the positive assessment and for acknowledging the improvements made in our revised manuscript. We are also grateful for the reviewer’s continued constructive suggestions.

      One remaining issue concerns the mechanistic novelty of the actomyosin redistribution described in this study. The authors emphasize that the key novelty lies in the stepwise translocation of actomyosin from the lateral membrane to the apical domain during the initial stage (apical constriction), followed by redistribution from the apical domain back to the lateral domain during the accelerated stage (invagination). I agree that the dynamic redistribution itself is potentially interesting and may represent an underexplored aspect of epithelial morphogenesis. However, as I discussed in my previous review comments, from a mechanics perspective, the role of apical actomyosin in driving apical constriction and of lateral actomyosin in contributing to tissue folding/invagination have already been demonstrated in multiple systems, although to varying extents depending on the model. Therefore, while the current study convincingly documents a distinct spatiotemporal sequence of actomyosin localization in Ciona atrial siphon tube formation, it could be clarified further to what extent this work advances new mechanical principles underlying epithelial folding, as opposed to revealing a variation in the deployment of previously described force-generating modules.

      Importantly, I think the manuscript has the potential to provide deeper conceptual insight if the authors more explicitly consider the significance of the "redistribution" process itself. Redistribution does not only involve the appearance of actomyosin at a new membrane domain; it also necessarily involves its disappearance from the previous domain. The latter aspect has, in my view, been much less explored in the literature. For example: Is the removal of lateral actomyosin during the early phase important for efficient apical constriction? Conversely, is the reduction of apical actomyosin during the later accelerated phase important for proper invagination mechanics? These questions are particularly interesting because they address whether redistribution between domains serves an active mechanical regulatory role, rather than focusing on the role of force-generating actomyosin at a given location.

      I acknowledge that addressing these questions experimentally could be technically challenging. One potentially powerful way to address this would be through the revised computational model. For example, the authors could test whether tissue folding is altered when actomyosin is allowed to accumulate at a new domain without being concomitantly depleted from the original domain. Such analyses could help distinguish whether redistribution itself has functional mechanical importance, rather than merely reflecting sequential recruitment to different cellular regions. In my opinion, incorporating this aspect would substantially strengthen the conceptual and mechanistic novelty of the study.

      We thank the reviewer for raising this important point. We agree that the mechanical significance of actomyosin redistribution, beyond the individual roles of apical and lateral contractility, deserves further clarification. In our study, quantitative analysis of F-actin dynamics revealed a bidirectional reorganization of the actomyosin network during siphon morphogenesis: the increase of F-actin intensity in one domain was accompanied by its reduction in the other domains during both the initial and accelerated invagination stages. These observations suggest that actomyosin redistribution may represent an active mechanical regulatory process rather than merely sequential recruitment to different cellular domains. Following the reviewer’s suggestion, we used the computational model developed in this study to examine its functional significance.

      Specifically, we modified the temporal dynamics of actomyosin activity by maintaining apical contractility during the accelerated stage or by prematurely enhancing lateral contractility during the initial stage in simulations. We found that sustained apical tension during the accelerated stage primarily induced stronger central cell elongation (Figure 6—figure supplement 1A-C), whereas elevating lateral actomyosin activity during the initial stage (14–16 hpf) suppressed central cell elongation and reduced the inward movement of surrounding cells toward the central region (Fig. 6—figure supplement 1D-F), resulting in earlier bending deformation with a flatter invaginating morphology (Fig. 6—figure supplement 1F). Although both perturbations eventually converged to similar final shapes, likely because they reached similar final actomyosin distributions, their distinct morphogenetic trajectories demonstrate the mechanical importance of the temporal sequence of actomyosin redistribution. These results further clarify the significance of the previously identified apico-basal tension imbalance and lateral contraction by revealing how their sequential activation coordinates tissue deformation. The initial dominance of apical contractility, together with limited lateral contraction, promotes cell elongation and convergence of the active region, whereas the subsequent shift toward lateral contractility facilitates cell shortening and deep tissue invagination.

      Thus, our results highlight that the bidirectional redistribution of actomyosin is not merely a consequence of morphogenesis, but contributes to the dynamic regulation of epithelial folding. We have incorporated these new simulation results and analyses into the revised manuscript as Figure 6—figure supplement 1.

      My other concern relates to the new optogenetic data presented in Figure 4-figure supplement 2. In the "Dark" samples, active myosin does not appear to be clearly enriched along the membrane, but instead seems relatively diffuse within the cytoplasm. This appears distinct from the images shown in Figure 2, where active myosin exhibits clear membrane enrichment. Could the authors provide top-view images for the samples shown in Figure 4-figure supplement 2? This would help clarify whether active myosin is indeed enriched along the apical membrane at 16 hpf and along the lateral membrane at 17 hpf in the "Dark" condition.

      We thank the reviewer for this careful observation. We agree that the p-MLC signal in the "Dark" samples of the original Figure 4—figure supplement 2 appeared less membrane-enriched than in Figure 2. This was due to a technical adjustment: because the optogenetic system occupied the 568 nm channel, we have to switch the p-MLC signal from Alexa Fluor 568 anti-rabbit IgG to Alexa Fluor 647 anti-rabbit IgG, which yielded relatively weaker membrane signal under our imaging conditions. To address this, we have replaced the cross-sectional images with better-representative examples and added top-view images (Figure 4—figure supplement 2). These new panels clearly showed that active myosin was enriched at apical junctions (16 hpf) and lateral membranes (17 hpf) in the Dark condition, consistent with that in Figure 2. Since the Dark and Light groups were processed identically, the relative comparison remains valid.

      In addition, the tissue morphology in the "17 hpf Light 1 hr" panel of Figure 4-figure supplement 2 appears noticeably different from that shown in Figure 4. Specifically, the apical side of the tissue in Figure 4 appears substantially more relaxed than in Figure 4-figure supplement 2. Based on the authors' interpretation of the optogenetic experiments, apical active myosin is not strongly affected by the treatment described in Figure 4. If so, one would expect apical constriction to remain largely intact. However, the more relaxed apical domain shown in Figure 4 seems to suggest that apical constriction may in fact be perturbed by the optogenetic manipulation. This apparent discrepancy complicates the interpretation of the experiment and seems somewhat inconsistent with the authors' main conclusion from this figure.

      We thank the reviewer for this very careful and insightful observation. We agree that the tissue morphology in the "17 hpf Light 1 hr" panel of Figure 4—figure supplement 2 appears less relaxed than that in Figure 4B.

      We acknowledge that optogenetic inhibition of myosin activity might indeed have a partial effect on activity of apical myosin, which can lead to apical relaxation after the initial constriction, as shown in the representative embryo in Figure 4B. However, because Ciona embryos are not always perfectly synchronized at 16 hpf (the time point when light illumination was initiated), individual embryos exhibit slight variations in the extent of apical constriction that has already been achieved during the initial stage (13.5–16.0 hpf). For embryos that had completed a relatively stronger apical constriction by 16 hpf, the apical domain can maintain its constricted morphology even after light exposure (Figure 4—figure supplement 2). Embryos with relatively weaker apical constriction at the time of light onset are more prone to exhibit apical relaxation upon optogenetic manipulation, as illustrated in Figure 4B. This developmental heterogeneity is the primary reason for the morphological variability observed between individual embryos in the optogenetic groups. Importantly, despite this morphological variability, the quantitative comparison of apical p-MLC intensity between the Light and Dark groups in Figure 4—figure supplement 2B showed no statistically significant difference (t-test, ns), which is consistent with the fact that during normal development, apical myosin activity naturally declines after 16 hpf (Figure 2B). In contrast, lateral p-MLC intensity was significantly reduced in the Light group compared to the Dark control (Figure 4—figure supplement 2B). This reduction in lateral contractility is the key factor responsible for the blockade of invagination progression.

      We have revised the statements accordingly. Hopefully, these clarifications have adequately addressed the reviewer's concern.

      Reviewer #3 (Public review):

      Concerns raised in the initial submission were addressed in the revised manuscript.

      We thank the reviewer for the encouraging feedback and for acknowledging our revisions.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further revisions suggested.

      We thank the reviewer again for the positive assessment.

      Reviewer #2 (Recommendations for the authors):

      Here are several additional comments and suggestions in addition to the concerns described in the public review.

      Line 86 - 88: the authors state "However, this transition is from apical to basolateral (Sherrard et al., 2010), rather than a bidirectional redistribution between apical and lateral domains." This statement feels somewhat out of context because the concept of "bidirectional redistribution" has not yet been introduced. It may fit more naturally in the Discussion section, after the relevant observations and interpretations have been fully presented.

      We thank the reviewer for this suggestion. We have revised the sentence in the Introduction to avoid introducing the concept of "bidirectional redistribution" prematurely.

      Line 92 - 95: the authors raised the questions of "Whether a bidirectional redistribution of actomyosin between apical and lateral domains operates as a core mechanism for sequential invagination, and whether lateral contractility is essential for the accelerated phase, remain unclear." These questions also feel somewhat out of context, for the same reason mentioned above.

      We thank the reviewer for this suggestion. We have revised the text to avoid prematurely introducing the concept of "bidirectional redistribution" in the Introduction.

      Line 162 - 163: The authors state that "This redistribution pattern was consistent with that of F-actin in the corresponding phases." This conclusion should be revised, as the reported increase in apical F-actin and reduction in lateral F-actin during the initial stage do not appear to reach statistical significance, which is different from that of active myosin.

      We thank the reviewer for this careful observation. We agree that the F-actin changes during the initial stage did not reach statistical significance, unlike the active myosin data. We have revised the sentence to state that the myosin redistribution showed a similar trend to F-actin, while acknowledging the lack of statistical significance for F-actin.

      Line 349 - 350: "and the invagination speed (represented by slopes of curves in Figure 5A) gradually slows down at later stages." It seems that Figure 5A should be Figure 5B.

      Thank you for pointing out this typo. We have corrected it in the revised manuscript (now Figure 5C).

      Reviewer #3 (Recommendations for the authors):

      We appreciate the efforts made by the authors to address the questions and comments raised in the initial submission. The only remaining concerns are regarding grammar and spelling and a couple of minor errors.

      (1) Lines 346-347, "These trends are consistent with the experimental mutant data (Figure 5A, B)."

      This was confusing - was this meant to be Figure 3A, B?

      We sincerely apologize for this confusion. We have corrected it in the revised manuscript (now Figure 3B).

      (2) Lines 349-350, "... the invagination speed (represented by the slopes of curves in Figure 5A) gradually slows down at later stages..."

      Similar to above, was this meant to be Figure 5C?

      Thank you for pointing out this typo. We have corrected it in the revised manuscript (now Figure 5C).

      (3) There is a typo in Line 336 ("dimmish") and some minor grammatical concerns in the introduction.

      We have corrected the typo "dimmish" to "diminish". We have also carefully proofread the entire manuscript and fixed any remaining grammatical issues.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study presents a novel analysis of MRI resources for 16 avian species, spanning major (though not all) clades and ecological niches. This is a significant step towards large-scale datasets on internal parcellation and long-range connectivity, central to evolutionary studies for understanding the evolution of the bird brain.

      Strengths:

      The integration of high-resolution T2-weighted and diffusion-weighted MRI with histological validation (Nissl and Luxol Fast Blue staining) provides a strong, cross-validated framework for studying avian brain anatomy. Data on long-range connectivity are particularly useful for understanding how relationships between brain components evolved. The approach is also scalable, allowing for more detailed evolutionary analyses compared to what is currently possible.

      Weaknesses:

      The sampling supports evidence of modular evolution in the bird brain, but it is limited for broad evolutionary claims, as the effects of sizes and phylogenies can be hard to disentangle without enough species per clade.

      We appreciate the reviewer highlighting this limitation and agree that it warrants explicit consideration. Although the 16 species included in the current study encompass diverse avian lineages and ecological characteristics, our sampling was not designed to rigorously disentangle the effects of phylogeny and brain size within individual clades.

      To improve the taxonomic coverage of our dataset, we plan to expand the MRI dataset in the revised manuscript by including additional specimens that are currently available to us. Specifically, we will acquire and analyze MRI data from the slaty-backed gull (Larus schistisagus), black-tailed gull (Larus crassirostris), budgerigar (Melopsittacus undulatus), and crow (Corvus sp.). In addition, we plan to incorporate publicly available MRI data from the ostrich (Struthio camelus). These additions will broaden the phylogenetic and anatomical coverage of the comparative dataset.

      Nevertheless, we fully acknowledge the reviewer’s point that, even with these additional species, the number of species sampled within individual clades will remain insufficient to rigorously separate phylogenetic effects from size-related effects or to support broad generalizations across all avian lineages. We will therefore explicitly state this limitation in the revised manuscript and temper evolutionary claims that extend beyond what can be supported by the present taxonomic sampling.

      Accordingly, we will more clearly position the primary contribution of this study as the establishment of an MRI-based comparative resource and analytical framework applicable across diverse avian species, rather than as a comprehensive phylogenetic test of avian brain evolution. Within this scope, we will present the observed variation among brain compartments as evidence consistent with mosaic/modular diversification, while carefully limiting the broader evolutionary interpretation of these patterns.

      Tractography-based claims should be treated cautiously without sensitivity analyses. This is particularly important when comparing brains with different sizes and tissue properties.

      The reviewer raises an important point regarding the interpretation of tractography-based comparisons. We agree that particular caution is required when comparing datasets across species that differ in brain size, tissue properties, and imaging characteristics.

      In the revised manuscript, we will perform additional sensitivity analyses to evaluate the robustness of the major tractography-derived patterns. Specifically, we will examine the effects of varying the FA threshold used for tract reconstruction, rather than relying solely on the current threshold of 0.1, and assess whether the major reconstructed trajectory patterns and cross-species differences are robust to this parameter.

      We will also re-analyze the available diffusion data using Generalized Q-Sampling Imaging (GQI) as an alternative reconstruction approach and compare the resulting trajectory patterns with those obtained using the current DTI-based analysis. We recognize that the ability to resolve complex fiber configurations depends on the underlying diffusion acquisition parameters, including b-value and spatial and angular resolution. We will therefore interpret these comparisons within the limitations of the currently available datasets, without implying that either reconstruction approach provides a definitive representation of the underlying fiber architecture.

      In addition, we will provide a more detailed description of the relevant diffusion MRI acquisition parameters and spatial and angular resolution of the datasets so that potential technical differences among species can be more clearly evaluated. We will also discuss how such differences may affect cross-species tractography comparisons.

      Finally, we will revise the interpretation of the tractography results throughout the manuscript to more clearly distinguish diffusion MRI-derived reconstructed trajectories from anatomically demonstrated neuronal connections. We will avoid treating reconstructed streamlines as direct evidence of anatomical connectivity and will explicitly discuss the potential for false-positive and false-negative tract reconstruction, as well as other limitations inherent to diffusion MRI tractography.

      Existing literature is not acknowledged sufficiently. This makes some claims of novelty misleading, and prevents readers from understanding the current state of knowledge in this research area.

      We fully agree with this assessment. Although substantial comparative neuroanatomical work has established important principles of avian brain evolution, the current manuscript does not sufficiently cite or incorporate this literature into the discussion. As a result, some statements regarding the novelty of the present study are overstated.

      In the revised manuscript, we will substantially expand our discussion and citation of the relevant literature, including the extensive comparative studies of individual avian brain regions and brain subdivisions highlighted by the reviewers. We will revise the Introduction and Discussion accordingly to more accurately describe the current state of knowledge in comparative avian neuroanatomy and to clearly distinguish the contributions of previous studies from those of the present work.

      We will also carefully revise statements regarding the novelty of our study. Rather than implying that our study provides the first comparative framework for examining internal avian brain organization, we will more appropriately emphasize its contribution as an MRI-based comparative resource that enables standardized visualization and analysis of internal brain anatomy and connectivity across diverse avian species.

      Reviewer #2 (Public review):

      This manuscript presents a comparative MRI dataset from 16 avian species and uses MRI and tractography to examine variation in brain organization across birds. The authors argue that their analyses support mosaic brain evolution and provide a framework for comparative neuroanatomy. Although the dataset represents a useful resource, particularly given the inclusion of understudied species like penguins, toucans, and hornbills, I have substantial concerns regarding the novelty of the study, the anatomical interpretation of the results, and the validity of the tractography analyses. In its current form, I do not believe the manuscript provides sufficient new biological insight to support many of its conclusions.

      Major Concerns

      (1) The authors repeatedly state that comparative neuroanatomical studies in birds have largely been unable to examine internal brain organization or "internal parcellation". This claim is inaccurate and reflects limited engagement with a substantial body of literature. For decades, comparative studies have examined variation in the size of major avian brain subdivisions as well as specific sensory, motor, and associative nuclei. For example, the extensive work of Andrew Iwaniuk and colleagues has documented variation in numerous brain regions across birds and related these differences to ecology, behavior, and sensory specialization (e.g., Gutierrez-Ibanez et al., 2009; Iwaniuk et al., 2006, 2008, 2010; Corfield et al., 2015). Other authors have also made important contributions in this area (e.g., Boire and Baron, 1994; Burish et al., 2004; Moore and DeVoogd, 2011, 2017). Importantly, previous work has already examined variation in major subdivisions of the avian brain using relatively standardized datasets (e.g., Iwaniuk et al., 2004; Iwaniuk and Hurd, 2005), including datasets that contain more species and greater taxonomic diversity than the current study. In other words, these studies have already provided detailed analyses of internal brain organization across broad taxonomic samples.

      The manuscript should therefore be reframed as providing a new MRI-based resource rather than introducing the first comparative framework for studying internal avian brain organization. The current framing significantly overstates the novelty of the work.

      We appreciate the reviewer drawing attention to this issue. Although a substantial body of comparative neuroanatomical work has already examined internal brain organization in birds, the current manuscript does not sufficiently cite or incorporate this literature into the discussion. As a consequence, some statements regarding the novelty of our study are overstated.

      In the revised manuscript, we will expand our discussion and citation of previous work, including the studies highlighted by the reviewer on interspecific variation in major avian brain subdivisions, sensory and motor nuclei, and other anatomically defined brain regions. We will revise the Introduction and Discussion to more accurately represent the existing body of comparative avian neuroanatomy and to clarify how the present study complements and extends these established approaches.

      Most importantly, we will reframe the manuscript so that its primary contribution is presented as the establishment of an MRI-based comparative resource for visualizing and quantitatively analyzing internal brain anatomy across diverse avian species. We will remove or revise statements implying that the present study provides the first comparative framework for examining internal avian brain organization. Instead, we will emphasize the complementary advantages of the MRI-based approach, particularly its ability to provide non-destructive three-dimensional visualization of internal brain structures in intact specimens and to enable comparison of multiple brain compartments within a common analytical framework across diverse avian species.

      We believe that this revised framing will more accurately position the contribution of our study within the existing literature and clarify the specific value of the dataset.

      (2) A second significant concern is the lack of anatomical specificity in the tractography analyses. The authors repeatedly refer to regions such as "anterior cortex," "dorsal cortex," and "temporal cortex." These terms are not standard anatomical designations in avian neuroanatomy and provide little information about the actual structures being analyzed.

      For example, the "temporal cortex" could potentially include portions of the nidopallium (including the caudolateral nidopallium, NCL), mesopallium, and arcopallium. Similarly, the "anterior cortex" could correspond to the somatosensory or visual Wulst, the anterior nidopallium, or several other structures. The designation "dorsal cortex" is similarly difficult to interpret. Because these seed regions may encompass multiple functionally distinct systems, it is impossible to evaluate the biological significance of the reported connectivity patterns.

      I strongly encourage the authors to define their seed regions using accepted avian neuroanatomical terminology and to provide detailed anatomical maps. More informative analyses would focus on well-defined structures with known connectivity, such as the Wulst, arcopallium, entopallium, or NCL. As currently presented, the tractography results are too coarse to support meaningful biological conclusions.

      This is an important concern, and we agree that the anatomical nomenclature and definition of the regions used in the tractography analyses require substantial improvement.

      In the revised manuscript, we will carefully re-evaluate the anatomical description of each region with reference to established avian neuroanatomical terminology, anatomical atlases, and available histological information. Where the analyzed regions can be reliably assigned to established anatomical structures, we will replace broad mammalian-style positional terminology such as “anterior cortex,” “dorsal cortex,” and “temporal cortex” with more appropriate avian neuroanatomical terminology. For example, we will re-examine whether the region currently referred to as the “optic lobe” can be more precisely defined as the optic tectum, while the cerebellum can be retained as an anatomically well-defined region.

      At the same time, we recognize an important limitation of the present tractography analysis. The spatial resolution and analytical framework of the present comparative datasets do not necessarily permit reliable assignment of all analyzed regions or reconstructed trajectory patterns to fine pallial subdivisions such as the entopallium, arcopallium, or NCL across species. We therefore do not intend to assign such specific anatomical identities where they cannot be supported with sufficient confidence.

      Instead, for regions that cannot be unambiguously assigned to a single established anatomical subdivision, we will define their location and extent using reproducible anatomical landmarks and clearly indicate the level of anatomical resolution supported by the data. We will also provide revised anatomical maps showing the locations of the analyzed regions and their relationship to major avian brain subdivisions. This will allow readers to evaluate more clearly which anatomical structures may contribute to the reconstructed trajectory patterns.

      Accordingly, we will revise the biological interpretation of the tractography results to match this anatomical resolution. Rather than attributing the reconstructed patterns to specific fine-scale pallial structures or functional systems when these cannot be reliably distinguished, we will interpret them more conservatively as broad patterns of fibre organisation among anatomically defined brain regions. These revisions will improve the anatomical transparency of the analysis while avoiding anatomical or functional interpretations that exceed the resolution of the present datasets.

      (3) I am not an MRI specialist, but I have concerns regarding the interpretation of the tractography results. Bird brains are small, and diffusion MRI tractography is already known to be challenging even in substantially larger brains. The manuscript provides limited information regarding image resolution, diffusion sampling, and the expected accuracy of tract reconstruction in these specimens. More importantly, there is little validation of the tractography results. Diffusion tractography is prone to both false positives and false negatives, and reconstructed pathways cannot be assumed to represent true anatomical connections.

      The authors should provide evidence that their tractography pipeline can accurately recover known pathways. For example, they could compare reconstructed tracts with well-established anatomical pathways such as the anterior commissure or major visual pathways, which would substantially strengthen confidence in the results. Without such validation, it is difficult to determine whether the observed species differences reflect biological variation or methodological artifacts.

      We agree with the reviewer that the tractography results require cautious interpretation and that further evaluation of the robustness and anatomical plausibility of the tractography pipeline would strengthen the study.

      In the revised manuscript, we will provide a more detailed description of the diffusion MRI acquisition parameters, spatial and angular resolution, diffusion reconstruction, and tractography procedures so that the methodological limitations of the analyses can be more clearly assessed.

      Recent studies using DTI under imaging conditions comparable to those employed in the present study have demonstrated successful reconstruction of fiber trajectories in the mouse brain, which is smaller than the avian brains examined here (Janz et al., eLife, 2017). We therefore consider that the size of the avian brain itself does not preclude DTI-based fiber reconstruction and that such analyses are technically feasible at the brain sizes examined in the present study.

      Importantly, we will perform additional sensitivity analyses to evaluate the robustness of the reconstructed trajectory patterns. Specifically, we will examine the effects of varying the FA threshold used for tract reconstruction. We will also re-analyse the available diffusion data using Generalized Q-Sampling Imaging (GQI) as an alternative reconstruction approach and compare the resulting trajectory patterns with those obtained using the current DTI-based analysis. We recognize that the ability to resolve complex fiber configurations depends on the diffusion acquisition parameters, including b-value and spatial and angular resolution. We will therefore interpret the comparison between reconstruction approaches within the limitations of the currently available datasets, without implying that either approach provides a definitive reconstruction of the underlying fiber architecture.

      We will also evaluate the robustness of the k-means clustering used to summarize tractography patterns. Because the current choice of k = 10 was not based on an independently established biological criterion, we will examine alternative values of k and assess whether the major trajectory patterns and cross-species differences are robust to the choice of cluster number.

      In addition, we will examine anatomically well-characterized pathways, including major commissural and visual pathways, and assess whether the reconstructed trajectories are consistent with known avian neuroanatomy. We will use these comparisons as an assessment of anatomical plausibility rather than as definitive validation of tractography accuracy.

      Finally, we will revise the manuscript to more clearly distinguish diffusion MRI-derived reconstructed trajectories from anatomically demonstrated neuronal connections. We will avoid interpreting reconstructed streamlines as direct evidence of anatomical connectivity and will explicitly discuss the possibility of false-positive and false-negative tract reconstruction, as well as other limitations inherent to diffusion MRI tractography.

      (4) I also have some methodological concerns regarding the comparisons of anterior commissure (AC) size and cerebellar foliation. First, the authors measure the AC in a coronal section. I would recommend measuring the AC area in a midsagittal section instead. Furthermore, the authors use the cross-sectional area of the same coronal section as the scaling variable. This seems problematic because the area of any given section will depend on the angle of sectioning and other technical factors. If the objective is to compare the relative size of the AC, then total brain volume or telencephalon volume would be more appropriate scaling variables.

      With respect to cerebellar foliation, the authors developed their own metric. I would encourage them to use methods already established in the literature, such as the foliation index described by Iwaniuk et al. (2006). Their approach may yield similar results, but using the foliation index would facilitate direct comparisons with existing datasets and would allow incorporation of additional published data (e.g., Cunha et al., 2021, which includes foliation index measurements for 54 bird species). The authors should also be aware that the foliation index scales with body size. Consequently, the high degree of foliation observed in penguins may not necessarily indicate cerebellar expansion or increased demands for sensorimotor integration associated with their specialized locomotion. I therefore believe that the conclusions regarding variation in AC size and cerebellar foliation should be re-evaluated after more appropriate analyses are performed.

      These methodological suggestions are very helpful. We agree that both the anterior commissure analysis and the assessment of cerebellar foliation should be re-evaluated to enable more anatomically and quantitatively appropriate comparisons across species.

      Anterior commissure:

      We agree that the current analysis, in which AC area was measured from the coronal section showing the largest cross-sectional profile and normalized to whole-brain area in the same section, may be influenced by differences in brain geometry and section orientation among species. In the revised manuscript, we will re-evaluate the method used to quantify AC size, including measurements from sagittal or midsagittal views where anatomically appropriate. We will also examine normalization against volumetric measures, such as total brain or telencephalic volume, rather than relying solely on a single-section whole-brain area. The corresponding results and interpretations will be revised accordingly.

      Cerebellar foliation:

      We also agree that our current branch-counting approach should be considered in relation to established quantitative measures of avian cerebellar foliation. In the revised manuscript, we will evaluate whether the cerebellar foliation index described by Iwaniuk et al. (2006), or a comparable standardized measure applicable to our midsagittal MRI datasets, can be used to re-analyze cerebellar foliation. This will also allow us to place our observations more directly in the context of the larger comparative datasets reported previously, including that of Cunha et al. (2021).

      Importantly, we recognize that cerebellar foliation is strongly influenced by allometric and phylogenetic factors. We will therefore re-evaluate the interpretation of interspecific differences in cerebellar foliation in light of brain/body size relationships and the existing comparative literature. In particular, we will revise the interpretation of the pronounced cerebellar foliation observed in penguins and other species and avoid attributing these differences directly to locomotor or sensorimotor specialization unless supported by the revised analyses.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Naina Gour and colleagues provide a detailed observational study in which they demonstrate that MRGPRX4, a human G-protein coupled receptor (GPCR), is expressed exclusively in human melanomas and, when expressed in mouse melanocytes, drives the development of melanomas in mice. These findings provide evidence that MRGPRX4 has the properties of an oncogene, at least in certain cellular environments.

      Strengths:

      A strength of this work is the nice historical note in which overexpression of MAS1, a GPCR, led to classic studies of transformed fibroblasts in culture and tumors in nude mice. Cloning of MAS1 led to the identification of the MRGPR family of receptors, now known to be key players in neuroimmune and neurosensory phenomena. Here, the story comes full circle with a member of the MRGPR family being linked to a tumor, specifically melanoma. Perhaps the story is not entirely surprising given that the neural crest serves as a precursor for both nerves and melanocytes. But it is nice to see.

      Additional strengths include the vast array of tools and techniques employed, from public databases to engineered mice, to establish firmly that MRGPRX4 is expressed in melanomas, although not in every malignant cell.

      We thank the reviewer for the positive assessment of our work and for noting the arc back to the original MAS1 studies, which we agree makes for a satisfying full-circle story. We would add one nuance: while the shared neural crest origin of melanocytes and nociceptors offers a plausible explanation for why an MRGPR family member could be co-opted in melanoma, we were nonetheless surprised that this property is specific to MRGPRX4. When we tested the other MRGPRX family members, some of which are also known to be expressed in sensory neurons, using the same genetic strategy (MRGPRX1, MRGPRX2, MRGPRX3; Supplementary Fig. 2B), none induced melanoma, despite arising from the same family of receptors. This suggests the oncogenic property is not simply a generic consequence of neural crest ancestry, but reflects a biology specific to MRGPRX4 itself, which we find is one of the more intriguing open questions this work raises.

      Weaknesses:

      Given the power of the strengths of the data and story, the following comment is only sort of a weakness, as the topic is addressed while being saved for future studies. Specifically, what leads to the expression of MRGPRX4? The authors posit that it is an epigenetic phenomenon, look briefly at methylation, and rather than going down the proverbial rabbit hole of what comes first, have reasonably decided to punt.

      We agree with the reviewer that understanding the precise mechanisms through which MRGPRX4 gets activated during oncogenesis is important, and we had noted this as a limitation of current findings. As we mentioned in the discussion, we hypothesize that one or more environmental modulators - UV exposure, inflammatory cues, or an aging skin microenvironment could initiate the epigenetic reprogramming underlying this expression. This will require dedicated studies and is suited for future work.

      Another concern is that given what comes across as the initial observation of MRGPRX4 being expressed in melanoma, what do all of the additional studies add?

      The expression data in human melanoma are correlative and serve only as the entry point. The subsequent studies establish MRGPRX4 as a causal, druggable driver of melanoma: it is sufficient to induce 100% penetrant metastatic melanoma in vivo, required for melanoma proliferation and invasion, signals through ligand-independent basal PI3K-AKT/MAPK activity, and is pharmacologically targetable.

      For the non-cognoscenti, and to make the manuscript more accessible, the abbreviation NC/EMT, which is also inverted to EMT/NC, should be spelled out periodically as neural crest/epithelial-mesenchymal transition.

      This is corrected in the revised version.

      Please explain how this study came about. Was it a result of someone deciding to look at expression in the GTEx project and compare it to a tumor database?

      This was a serendipitous discovery. As MRGPRs can modulate immune cell function, we were driving expression of various human MRGPRX genes in immune lineages in mice. We noticed that mice overexpressing MRGPRX4 developed spontaneous melanoma-like growths on the ear and tail. Upon investigation, we found that while the Cre driver we had used is appropriate for immune cells, it is also expressed at low levels in melanocytes. We hypothesized that our initial "immune-knock-in" strategy likely resulted in low-level knock-in in mouse melanocytes, thereby allowing MRGPRX4 to directly drive melanocytic transformation. This led us to explore MRGPRX4 expression in human malignancies, where we found MRGPRX4 was highly expressed in melanoma. Next, to directly test our hypothesis that MRGPRX4 transforms melanocytes, we overexpressed MRGPRX4 using Tyrosinase-CreER, the standard Cre driver for melanocytes. This led to a fully penetrant spontaneous melanoma phenotype (Fig. 2).

      A comment could be made to explain that while NSG and normal mice were used, the former are immunocompromised, and drawing conclusions without specifying these differences is a weakness.

      This is a fair point, and we agree the distinction deserves clarification. The two mouse systems in this study serve different purposes and are not interchangeable. The transgenic TyrCreER+; MRGPRX4LSL+/- model, in which melanoma arises spontaneously from endogenous mouse melanocytes, was studied in immunocompetent mice, allowing us to fully characterize the tumour, including the tumour immune microenvironment (Fig. 2L–O), in an intact immune setting. In contrast, the human A2058 melanoma cell studies (proliferation and metastatic seeding; Fig. 4B–J) required an immunodeficient host, since human cells would otherwise be rejected by a competent murine immune system. For these experiments, we used NSG mice, the standard immunodeficient strain used in human xenograft studies. We have updated the text to reflect this.

      In Figure 1A, the p-value of -145 begs for a little explanation. I don't recall seeing such a p-value.

      The small p-value reflects two features of the comparison: (1) both GTEx normal skin and TCGA-SKCM contain large sample sizes (hundreds of samples per group), which gives the Mann-Whitney U test (also known as the Wilcoxon rank-sum test) statistical power, and (2) MRGPRX4 expression shows minimal overlap between the two groups; it's essentially undetectable in normal skin but broadly expressed across melanoma samples. Together, these yield such a p-value. This is expected for rank-based tests applied to large, cleanly separated datasets.

      Have you considered treating the murine melanomas with murine via PD-L1? I appreciate that this comment is somewhat superfluous given the inhibition of MRGPRX4 with compound 31-2, but given the human therapeutics combined with the fact that you have done 'everything else', I wonder what might happen.

      We agree with the reviewer that testing checkpoint blockade in the MRGPRX4-driven model is a relevant direction. Our finding that MRGPRX4-driven tumours develop a PD-L1hi, neutrophil-rich microenvironment (Fig. 2M–O) makes anti-PD-L1 treatment a natural next experiment, and we would predict it to be informative both on its own and in combination with an MRGPRX4 inhibitor, given that the two target distinct compartments- the tumor-immune microenvironment versus tumour-intrinsic proliferative/invasive signaling. We consider it an important direction for future work.

      Given the basal ligand-independent signaling, might engineering variants of MRGPRX4 that do not signal be of value?

      This is a great point. A variant that disrupts basal (ligand-independent) signaling specifically would be informative. Identifying and validating such a basal-activity-disrupting variant, including confirming normal receptor expression and trafficking, is an important direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This study presents a fundamental new finding - the identification of a sensory-neuron itch receptor, MRGPRX4, as an unexpected melanoma oncogene through a mechanism of lineage-inappropriate expression rather than mutation. The evidence supporting the core observation (tumor-specific upregulation, restriction to invasive transcriptional states, and sufficiency to drive fully penetrant metastatic melanoma in vivo) is compelling, drawing on convergent human genomic datasets and a well-controlled genetic mouse model. However, several of the mechanistic and translational claims - particularly regarding causal drivers of invasion, the immunosuppressive tumor microenvironment, and in vivo pharmacological efficacy - remain incomplete, relying on correlative evidence.

      Strengths

      The authors propose that MRGPRX4, normally restricted to a subset of peripheral sensory neurons, is aberrantly re-expressed in melanoma rather than through mutational mechanisms, and that this re-expression is sufficient to drive tumorigenesis through basal, ligand-independent GPCR signaling. This is a genuinely novel model for oncogenesis, and the manuscript deploys an impressive range of approaches - bulk and single-cell transcriptomics, spatial transcriptomics, proteomics, phosphoproteomics, and pharmacology - to support it.

      Strengths:

      The claim that MRGPRX4 is selectively upregulated in melanoma and confined to neural-crest-like/invasive transcriptional states is well supported, with consistent results across multiple independent human scRNA-seq datasets. The claim that ectopic MRGPRX4 is sufficient to drive melanoma is convincingly demonstrated by the fully penetrant, metastatic phenotype in the Tyr-CreER;MRGPRX4-LSL model, with appropriate specificity controls showing that MRGPRX1, MRGPRX2, and MRGPRX3 do not phenocopy this effect.

      The claim that MRGPRX4 signals through basal, ligand-independent activity is reasonably well supported by bile-acid quantification showing endogenous ligand concentrations well below the EC50 required for activation.

      We are thankful to the reviewer for the positive feedback.

      Weaknesses:

      The claim that MRGPRX4 remodels the tumor microenvironment toward an immunosuppressive state rests on flow cytometric frequency data (altered neutrophil/eosinophil ratios, increased PD-L1+ myeloid populations) but lacks any functional immune assay to demonstrate that this remodeling actually impairs anti-tumor immune responses.

      We acknowledge the reviewer’s suggestion that functional immune assays will demonstrate that MRGPRX4-driven tumor remodeling impairs checkpoint-driven anti-tumour immune response. We consider this an important direction for future work.

      The claim that the two MRGPRX4-enriched tumor subpopulations (ECM-rich and NC-like/invasive) underlie the observed invasive and metastatic phenotype is not directly tested; the authors appropriately acknowledge this as an open question, but it is worth noting explicitly that this leaves the mechanistic link between the identified cell states and the functional phenotype (proliferation, invasion, metastasis shown in Figure 4) unresolved.

      We agree with the reviewer. To test the direct role of these subpopulations in driving metastasis in our model, we need to ablate them specifically. This is possible with an intersectional genetics strategy, wherein we could use a state-specific Dre or Flp driver (e.g., a Prrx1-DreER or Prrx1-FlpO knock-in) crossed to a dual-recombinase-dependent effector allele (e.g., a Frt- or Rox-gated diphtheria toxin receptor), and then crossed to our TyrCreER-MRGPRX4-LSL model. This quadruple-transgenic strategy would allow ablation restricted specifically to MRGPRX4-driven tumour cells occupying the target mesenchymal state, for example. These genetic lines can be created but require generating or sourcing new dual-recombinase-dependent alleles and a substantially longer breeding and validation timeline (approximately 18-24 months). Together, these represent possible, if long-term, strategies for directly testing the causal contribution of these MRGPRX4-enriched subpopulations to melanoma invasion and metastasis.

      Finally, the comparison with BRAF- and NRAS-driven GEMMs (Figure 3K-L) establishes overlap in transcriptional cell states but does not report whether these canonical models themselves upregulate endogenous Mrgprx4. This omission leaves unclear whether MRGPRX4 acts as a convergent node downstream of canonical oncogenic signaling, or represents an independent, parallel route to a similar phenotypic endpoint - a distinction that matters considerably for how broadly the finding should be interpreted.

      MRGPRX4 is a primate-specific receptor with no mouse ortholog in melanocytes; thus, we cannot assess endogenous expression of MRGPRX4 in BRAF/NRAS-driven GEMMs. Nevertheless, the following observations argue against MRGPRX4 functioning solely downstream of canonical oncogenes. First, TyrCreER<sup>+</sup>; MRGPRX4LSL mice develop fully penetrant melanoma without engineered BRAF or NRAS activation or tumour-suppressor loss. Second, MRGPRX4 loss in BRAF V600E mutant A2058 cells reduces pERK1/2, pAKT, and pS6K, showing that MRGPRX4 sustains these signaling outputs even in the presence of activated BRAF. Together, these findings support a model in which MRGPRX4 could provide a distinct oncogenic input.

      Overall assessment:

      The manuscript's central, most novel claim - that lineage-inappropriate expression of a sensory GPCR is sufficient to drive melanoma - is compellingly supported. The secondary mechanistic and translational claims built around this finding are convincing and consistent with the broader literature but are currently supported by correlative rather than causal or functional evidence.

    1. Author response:

      We thank the editors and reviewers for their thoughtful and constructive assessment of our manuscript. We appreciate the reviewers’ insightful comments and suggestions, which will help strengthen the mechanistic rigor of our work. Below, we outline the key revisions we plan to undertake in the revised version.

      Response to Reviewer 1

      (1) Dynamic regulation of PA lactylation during infection.

      We plan to examine PA lactylation levels at multiple time points post-infection to assess whether PA lactylation changes dynamically during the viral replication cycle.

      (2) Epistasis experiments linking K605/K609 to lactate- or enzyme-dependent phenotypes.

      We acknowledge that multiple viral proteins undergo lactylation and that ATAT1/SIRT1 may mediate lactylation of multiple viral proteins. We will perform viral replication assays in the context of K605/K609 mutant viruses under lactate supplementation or ATAT1/SIRT1 manipulation conditions. These experiments will allow us to assess whether the effects of lactate or ATAT1/SIRT1 manipulation on viral replication are dependent, at least in part, on PA K605/K609.

      (3) Conservation across viral strains and cell lines.

      We will further examine key phenotypes in additional influenza A virus subtypes (e.g., H1N1 swine influenza and H9N2 avian influenza strains) and in additional cell lines to assess the extent to which the observed mechanism is conserved beyond the PR8 laboratory-adapted strain.

      (4) Direct biochemical evidence for ATAT1 and SIRT1 activity on PA.

      We plan to perform in vitro lactylation and de-lactylation assays using purified recombinant PA, ATAT1, and SIRT1 proteins to investigate whether ATAT1 and SIRT1 can directly modulate PA lactylation, respectively.

      (5) Disentangling lactylation from charge/structural effects.

      We acknowledge the reviewer’s point that K609R shows reduced lactylation without obvious changes in polymerase activity. We will include K-to-Q substitution mutants (e.g., K605Q/K609Q) in functional assays to further assess the functional consequences of these substitutions. Although K-to-Q substitutions do not strictly mimic lysine lactylation, these mutants may help distinguish effects related to lysine charge/chemical properties from those specifically attributable to lactylation. We will also temper our conclusions and acknowledge that charge and/or structural effects may contribute independently to the observed phenotypes.

      (6) Enzymatic activity dependence of ATAT1 and SIRT1.

      We will further investigate the enzymatic activity-dependent versus -independent contributions of ATAT1 and SIRT1 using catalytically inactive mutants, together with the epistasis experiments described above. We will also revise the text to clarify the interpretive limitations of these experiments.

      (7) Improved loss-of-function approaches.

      We will complement the existing siRNA experiments with CRISPR/Cas9 knockout cell lines for ATAT1 and SIRT1 and, where feasible, repeat key assays in ATAT1- and SIRT1-knockout cells.

      In the revised manuscript, we will also add a detailed methodological explanation in the figure legend and Methods section to clarify how the luciferase complementation system distinguishes asymmetric from symmetric polymerase dimers, show individual data points overlaid on bar graphs with error bars, and correct spelling errors throughout the manuscript.

      Response to Reviewer 2

      (1) Direct in vitro lactylation/de-lactylation assays.

      As noted above, we will perform in vitro modification assays with purified proteins to investigate whether ATAT1 and SIRT1 directly modulate PA lactylation and de-lactylation, respectively.

      (2) Functional assays with lysine-to-glutamine substitution mutants.

      Although we recognize that K-to-Q substitutions do not strictly mimic lysine lactylation, we will generate K-to-Q substitution mutants (e.g., K605Q/K609Q) and evaluate their effects on polymerase activity and viral replication to further assess the functional relevance of these sites.

      (3) Complementation experiments linking ATAT1/SIRT1 phenotypes to PA K605/K609.

      As noted above, we plan to perform viral replication assays in the context of K605/K609 mutant viruses under ATAT1/SIRT1 manipulation conditions to assess whether PA K605/K609 contributes to the effects associated with ATAT1/SIRT1 manipulation.

      (4) IP-MS comparison of PA WT versus mutant host interaction profiles.

      Our study focuses on the mechanism by which PA lactylation modulates viral polymerase activity and replication. A comprehensive host interactome analysis via IP-MS represents a broader systematic investigation beyond the scope of this focused work. We will discuss this as an important future research direction in the revised manuscript.

      (5) Exploration of host antiviral immunity downstream of PA lactylation.

      This work focuses on the direct effects of PA lactylation on viral polymerase activity and replication. As the PA mutants exhibit altered replication capacity, differences in IFN/ISG expression would be largely secondary and difficult to disentangle from the direct effects of altered viral replication. We therefore consider this question beyond the scope of the current study and will add it as a future research direction in the Discussion section.

      In the revised manuscript, we will also examine PA lactylation levels under increasing lactate concentrations to assess their relationship with the dose-dependent changes in viral titers. We will revise the text to clarify the interpretive limitations and, where feasible, perform endogenous co-immunoprecipitation experiments to further assess the interactions between PA and ATAT1/SIRT1 under physiological expression conditions.

      We believe these revisions will strengthen the mechanistic evidence and help address the core concerns raised by both reviewers.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors provide a detailed ultrastructural analysis of the larval pharyngeal sensory organs, including the dorsal pharyngeal sensilla, dorsal pharyngeal organ, ventral pharyngeal sensilla, and posterior pharyngeal sensilla. Using electron microscopy and 3D reconstruction, Richter et al., present a comprehensive mapping and classification of pharyngeal sensory structures, defining the morphological type of pharyngeal sensilla based on ultrastructure and generating a neuron-to-sensillum map. These findings significantly advance our understanding of internal larval sensory systems and establish a robust framework for future functional studies in coordination with external sensory systems.

      Strengths:

      The application of high-resolution electron microscopy and 3D imaging analysis successfully overcomes technical challenges associated with visualizing deep internal structures. This enables an unprecedented level of anatomical detail of the larval pharyngeal sensory system. Thus, the study complements and completes existing maps of larval sensory circuits, contributing a comprehensive neuroanatomical characterization of larval sensory input pathways. These insights will inform future studies on larval behavior, sensory processing, and may also have applied relevance for insect control strategies.

      Weaknesses:

      While the manuscript is concise, clearly written, and methodologically rigorous, it primarily addresses a specialized readership with expertise in insect neuroanatomy.

      We thank the reviewer for the positive assessment of our study and for the helpful suggestions. In response, we have clarified the visual presentation of the pharyngeal sense organs in Figure 1, expanded the discussion of adult pharyngeal sensory systems, briefly broadened the comparison to other insect species, checked and corrected the scale bars, and added further methodological detail where appropriate.

      Reviewer #2 (Public review):

      Summary:

      This manuscript documents the structure of the pharyngeal nervous system of the Drosophila larva. The authors wanted to achieve a detailed ultrastructural reconstruction of the gustatory sensory organs in the Drosophila pharynx. Using serial EM and the associated bioinformatics tools, they have achieved their goal. The paper is written clearly and illustrated beautifully with 3D models and annotated sections. The data will significantly enrich the field of Drosophila neurobiology.

      Strengths:

      Given the dataset, the findings presented are solid and will be an important work of reference for the future.

      Weaknesses:

      Previous work, including EM, on the pharyngeal sensory organ is not sufficiently referenced and used for comparison with the data presented in this study.

      We are grateful for the reviewer’s thoughtful comments and for the suggestion to strengthen the historical and comparative context of the work. We have revised the introduction to better acknowledge and discuss the relevant previous EM-based literature on adult and larval internal gustatory sensilla, clarified the organization of the shared pore structure in T1–T3, highlighted the DPO multidendritic neurons more explicitly, and added a comparison that emphasizes the added value of the complete serial EM dataset.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) For improved clarity, highlight the pharyngeal sense organs in Figure 1B. Consider using the color schemes to differentiate between peripheral and internal sensory organs.

      We thank the reviewer for this helpful suggestion. We have revised Figure 1 to more clearly separate the pharyngeal sense organs from the external sense organs in the head region. This revision improves visual clarity and accessibility for readers.

      (2) In reference to lines 80-84, expand the discussion to address how future studies could explore the conserved morphological and functional characterization of the adult pharyngeal sensory system.

      We appreciate this suggestion and have expanded the discussion accordingly. We now briefly address how future work could compare the larval and adult pharyngeal sensory systems to examine conserved morphological and functional features.

      (3) To broaden the manuscript's appeal and emphasize its relevance beyond Drosophila, briefly discuss similarities, differences, or conserved roles of pharyngeal sensory systems in other insect species.

      Thank you for this valuable recommendation. We have added a paragraph placing the Drosophila pharyngeal sensory system in a broader insect context, including similarities, differences and potential conservation across species.

      (4) Recheck the scale bars in all figures, including the supplemental material.

      We thank the reviewer for pointing this out. We carefully rechecked all scale bars across the main and supplemental figures and corrected the missing ones.

      (5) Consider including additional details on image processing or provide appropriate citations for further reading.

      We appreciate this suggestion. We have expanded the methods section to include additional information on technical details and provide the relevant reference for further reading.

      Reviewer #2 (Recommendations for the authors):

      (1) Line 57ff: The previous literature describes internal gustatory sensilla in considerable detail.

      (a) Adult: These sensilla form three complexes, the labral sensory organ, and the ventral and dorsal cibarial sensory organ (Nayak & Singh, 1983, 1985; Singh, 1997; Stocker & Schorderet, 1981; Kendroud et al., 2017). The work by Nayak and Sing includes TEM and presents detailed EM-based schematics. This should be referenced and discussed.

      (b) Larva: Gendre et al. 2004, describes the internal gustatory organs and relates them to their adult counterparts:

      - Dorsal pharyngeal sense organ (DPS) and dorsal pharyngeal organ DPO) are the forerunners of adult labral and ventral cibarial sensory organs

      - Posterior pharyngeal sensory organ (PPS) is the forerunner of the adult dorsal cibarial sensory organ

      - Ventral pharyngeal sensory organ (VPS), derived from the labial segment, undergoes apoptosis during metamorphosis

      This work, connecting larva and adult (and containing detailed diagrams comparing adult and larval pharyngeal sensilla) should be presented in the introduction.

      We thank the reviewer for this important comment. We have revised the introduction to better cite and discuss previous EM-based studies of internal gustatory sensilla in both adult and larval stages, and we now place our findings more explicitly in the context of this prior work.

      (2) Line 180: the relationship between the ending of T1-T3 in one shared pore, and the individually wrapped sensilla should be explained; maybe a simple diagram would help. I did not understand how it works. Normally, in a gustatory sensillum, you have one or more sensory neurons, surrounded by thecogen, trichogen, and tormogen cells. The trichogen generates the shaft with the pore at its tip. Now here, in T1-T3, you have three sets of thecogen/trichogen/tormogen. Do all three trichogen cells somehow participate in the shaft with the common pore? Or only a single one, and the other two generate no shaft? It is possible this cannot be resolved, but the authors should address the problem and suggest a possible scenario.

      We appreciate the reviewer’s concern and agree that this point required clarification. We have revised the relevant text to better explain the organization of T1-T3 and their shared pore and the organization of the support cells.

      (3) Line 205: the DPO multidendritic neurons with dendrites into the hemolymph should be shown; in Figure S4G, I could see only cell bodies. These MD neurons in the gustatory system are, I believe, a true novelty and should be emphasized more if the material allows (text figure!)

      Thank you for highlighting this point. We have revised the results and supplementary material to show these neurons more clearly and to emphasize their novelty and potential relevance to the pharyngeal sensory system.

      (4) A somewhat detailed comparison between the ultrastructure of the DPS as extracted from the serial EM stack of this study, and the conclusions of Nayak and Singh 1983 as depicted in their diagram Figure 7a would be productive. The idea being: what additional details can (only) a complete EM stack provide, compared to conventional EM.

      We appreciate this suggestion. Rather than directly comparing larval and adult structures in detail, we now emphasize what the complete serial EM dataset adds beyond conventional single-section EM, namely a more comprehensive and complete reconstruction of the sensory organs and associated cell types (multidendritic neurons, papilla sensilla, and chordotonal organs that were not described before, organization of support cells)

      (5) To round off the work and connect it to the previously published analysis of gustatory terminal arborizations and connectivity in the brain (Miroschnikow et al.,2018), it would be helpful to add an analysis of the distribution of axons from the different sensilla in the nerves. Miroschnikow analyzes the central terminations of the same sense for which the peripheral structure is described here, only that in their L1 connectome, the periphery was cut off. Do the findings of the current study match their predictions, as to the number of sensory neurons, etc? It should be possible to follow, even at the lower resolution of the dataset presented here, to follow axons of sensory neurons through the nerves to the neuropil entry, and thereby make the connection. I consider this to be of great importance for the field, for authors who want to use the data of this study, and the Miroschnikow et al analysis, for their own studies.

      We thank the reviewer for this thoughtful and constructive suggestion. We fully agree that linking the peripheral sensory anatomy described in this study to the central projections analyzed by Miroschnikow et al. would be highly valuable and of broad interest. However, a systematic analysis of axon distributions from the different sensilla through the nerves to their neuropil entry points is beyond the scope of the present work. Owing especially to the dataset’s resolution and inherent limitations, tracing the connections from sensory organs through the nerves to their projections in the brain is technically highly challenging and extremely time-consuming, since much of the process would need to be performed manually. We therefore do not include a detailed comparison with the predictions from Miroschnikow et al. in this manuscript. Nevertheless, we appreciate that such an analysis would be an important next step for the field and a useful resource for future studies.

      We are grateful for the reviewers’ thoughtful feedback, which has helped us improve the manuscript substantially. We hope that the revised version addresses the concerns raised and better conveys the significance of our work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study uses a tripartite transdiagnostic computational framework to distinguish depression-specific, anxiety-specific, and shared psychopathology dimensions, in their relationships to mood variability and mood reactivity to reward prediction errors across multiple large non-clinical cohorts and a clinical sample. The evidence is convincing overall because the study combines large samples, a well-characterized gambling task and in-depth computational and psychometric analyses, and it replicates the depression-specific association with blunted reward prediction error-sensitivity in a clinical sample. However, the anxiety-specific effects are less consistently supported across individual datasets, may be underpowered in the clinical cohort because of comorbidity, and some aspects of the factor-analytic, risk-attitude, and mediation analyses would benefit from clearer explanation. These findings advance a mechanistic account of how distinct symptom dimensions differentially shape reward-based mood updating and variability, providing a principled framework for future transdiagnostic modeling.

      We thank the editors and reviewers for this important assessment.

      Regarding inconsistent results for anxiety-related effects in healthy datasets. Although the anxiety-specific factor showed associations in the expected direction across healthy datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level. Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity in non-clinical participants (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). In addition, we conducted a mini meta-analysis, and results support that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE in non-clinical participants. We have clarified this point below:

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β<sub>RPE</sub> were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept to account for dataset-level variability. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001). These associations remained significant after controlling for gender, age, task earnings, and mood drift. In addition, we performed a mini meta-analysis on these correlation coefficients[49]. Results showed significant positive correlation for both mood variation and RPE-related mood sensitivity (mood variation: Z = 3.399, 95 % CI for correlation coefficient r [0.045, 0.166]; RPE-related mood sensitivity: Z = 2.618, 95 % CI for correlation coefficient r [0.021, 0.143]), supporting that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE in non-clinical participants.”

      We also admit that the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      Regarding be underpowered sample size in the clinical cohort. We agree that the clinical sample may have been underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). Although these covariate analyses support the robustness of the depression-related effect, they do not resolve whether the absence of the anxiety-related effect reflects limited power or true clinical discontinuity. We also revised the Discussion to explicitly acknowledge that the anxiety-related effect observed in the pooled non-clinical dataset was not replicated in the clinical sample. We now note two possible interpretations. First, this discontinuity may reflect limited statistical power in the clinical sample. Second, and more speculatively, it may reflect a disruption of mood homeostasis in affective disorders (Paulus, 2007). In non-clinical individuals, the counterbalancing associations of depression- and anxiety-related traits with mood variation may contribute to emotional equilibrium. In contrast, affective disorders may involve a loss of this regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals. We have revised the manuscript as follows:

      Pages 12-13:

      “To test whether abnormalities in RPE-driven mood fluctuations can serve as clinically relevant computational markers of depression- and anxiety-related symptom dimensions, we recruited patients with affective disorders (n = 116) to complete the same questionnaire battery and gambling task with momentary mood ratings (Figure 1). Demographic, psychological, and clinical characteristics are summarized in Table 1 and Table S8. We observed significant negative correlations between depression-specific scores and both mood variation (r = -0.239, p = 0.009) and RPE-related mood sensitivity (β_RPE; r = -0.216, p = 0.020). These associations remained significant after controlling for demographic and clinical covariates, task earnings, and mood drift (ps < 0.05). Bootstrap validation yielded consistent results. Mediation analyses further showed that reduced mood sensitivity to RPEs statistically mediated the association between depression-specific scores and lower mood fluctuations (a × b = -0.141, 95% CI = [-0.261, -0.038], p = 0.021; Figure 3). However, we did not observe significant correlation with anxiety (mood variation: r = -0.092, p = 0.327; β_RPE: r = -0.095, p = 0.311).”

      Pages 17-18:

      “Notably, the pattern of heightened RPE sensitivity observed in the pooled non-clinical dataset was not observed in the clinical sample. On the one hand, this discontinuity may reflect that the clinical sample was underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). On the other hand, it may reflect a disruption of mood homeostasis in clinical populations[41,58]. In non-clinical individuals, counterbalancing associations of depression- and anxiety-related traits with mood variation may help maintain emotional equilibrium. In contrast, affective disorders may involve a loss of such regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals.”

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      We have also revised the manuscript accordingly to make the factor-analytic, risk-attitude, and mediation analyses clear.

      Reviewer #1 (Public review):

      Summary:

      This is a very interesting paper. The research question is intriguing, allowing the authors to address commonly observed comorbidities between depression and anxiety and their dissociable and opposite relationship to mood fluctuations and sensitivity to reward prediction errors. The computational analyses are very in-depth, including many state-of-the-art checks and validations. Another strength is the inclusion of several large or very large samples, including a patient sample in addition to the general population sample.

      I have the following questions:

      (1) Factor analysis

      I found the hierarchical organization of the factors interesting. While this is a very common procedure in, for example, the field of intelligence (producing sub-scores and a general g factor), it is not yet very commonly used in the field of computational psychiatry (though it has been validated before for anxiety/depression, so it is used here with good reason). I was also impressed by the methodological depth. In particular, it was of note how thoroughly done it was (for example, repeating the EFA on the second half of the data set). I have one question though: is the sample size too small for the exploratory analyses, given the number of items? Given the stability across the half-split, I imagine it is not. Perhaps the authors could spell out how many items, what would be the recommended standard for a subject-to-item ratio, and comment on this. A very technical point, the authors should specify how they extracted the factor scores from the other data sets (is it using the Thurstone or Bartlett method)? From experience (though not doing a hierarchical factor analysis), Bartlett can be somewhat better compared to the default (Thurstone) - better as in the resulting factors more closely recapitulating the factor correlations in the original sample (and independence of responses of other participants in a sample for computing a person's factor score). Could you also comment on similarities or divergences in this hierarchical factor analysis approach from another one recently used transdiagnostically in Wise et al. (2026, Translational Psychiatry)?

      We thank the Reviewer for the positive evaluation of our hierarchical factor-analytic approach and for recognizing the methodological depth of our analyses. We are particularly grateful for the Reviewer’s constructive suggestions regarding the participant-to-item ratio, factor score extraction, and the relation between our approach and recent transdiagnostic hierarchical factor-analytic work.

      First, hierarchical organization. As the Reviewer noted, hierarchical and bifactor representations have a well-established tradition in intelligence research, where they are used to model the g factor alongside domain-specific abilities (e.g., Reise, 2012; Rodriguez et al., 2016). Crucially, the use of such hierarchical structures in the present study was motivated primarily by theory and evidence from anxiety and depression research, rather than by analogy to intelligence research alone. This tradition can be traced back to the tripartite model of anxiety and depression (Clark & Watson, 1991), which distinguished a broad shared component of general distress or negative affect from more specific anxiety- and depression-related components. Subsequent psychometric work has further supported bifactor and hierarchical representations of anxiety and depression symptoms, including models that separate a general internalizing/distress factor from symptom-specific dimensions (e.g., Simms et al., 2008). More recently, similar hierarchical symptom structures have also been adopted in computational psychiatry to relate shared and specific affective symptom dimensions to task-derived computational parameters (Gagne et al., 2020, 2022; see Wise et al., 2023 for a review). Thus, the bifactor structure used here provides a theoretically motivated way to capture both the variance shared by anxiety and depression and the symptom-specific variance relevant to our computational analyses. We have clarified this point in the revised manuscript as follows:

      Pages 3-4:

      “Recent work has used bifactor models of the tripartite model of depression and anxiety to clarify their distinct features and differential influences on decision-making[31,32]. The tripartite model of anxiety and depression proposes that these two symptom dimensions share a broad general distress or negative affect component while also including symptom-specific components: low positive affect/anhedonia is more specific to depression, whereas physiological hyperarousal is more specific to anxiety[30,33,34]. Bifactor analysis offers a way to model this structure statistically. In a bifactor model, symptoms load on a general factor reflecting their shared variance and on specific factors capturing residual variance in narrower symptom dimensions after accounting for the general factor. Although bifactor and hierarchical models have long been used in psychometrics, e.g., intelligence research[35,36], their application to anxiety and depression is grounded in the tripartite model and subsequent psychometric work distinguishing general internalizing/distress from symptom-specific dimensions. This framework has recently been extended to computational psychiatry, where shared and specific affective symptom dimensions have been linked to task-derived computational parameters. For example, Gagne et al. (2022) used bifactor analysis to show that depression was associated with weaker prior beliefs, whereas anxiety was associated with a stronger negative bias in belief updatin31.”

      Second, the participant-to-item ratio. We agree that the ratio of 450 participants to 128 items in the EFA split-half sample, approximately 3.5:1, is below some conventional sample-size recommendations for exploratory factor analysis, including the often-cited recommendation of five participants per item (Costello & Osborne, 2005). However, as the Reviewer noted, the split-half analysis showed a stable factor structure, and the independent CFA in the other split-half sample further supported the robustness of the solution. Importantly, participant-to-item ratios are only one criterion for evaluating factor recovery. De Winter, Dodou, and Wieringa (2009) demonstrated that reliable EFA solutions may be obtained even with relatively small samples when the data are well-conditioned, such as when factor loadings are high, the number of factors is small, and each factor is defined by multiple items. These conditions were largely met in our data. We have clarified this point in the revised manuscript as follows:

      Supplementary Page 3:

      “In addition, the ratio of 450 participants to 128 items in the EFA split-half sample, approximately 3.5:1, is below some conventional sample-size recommendations for EFA, including the often-cited recommendation of five participants per item[8]. However, the split-half EFA yielded a stable factor structure, and the independent CFA further supported the robustness of this solution. Moreover, reliable EFA solutions may be obtained even with relatively small samples when the data are well-conditioned, such as when factor loadings are high, the number of factors is small, and each factor is defined by multiple items9. These conditions were largely met in our data.”

      Next, factor score extraction. We followed prior work using bifactor modeling in computational psychiatry (Gagne et al., 2020) and extracted factor scores with the Anderson–Rubin method, implemented using psych::factor.scores with method = "Anderson". This approach yields standardized and mutually orthogonal factor scores, which is particularly appropriate for our subsequent correlation analyses because it produces orthogonal scores and therefore avoids multicollinearity among the general, depression-specific, and anxiety-specific factors. In addition, as suggested by the Reviewer, we extracted factor scores using the Bartlett method from an oblique bifactor model, which allowed the depression- and anxiety-specific factors to correlate. In the combined dataset (n = 1,026), anxiety- and depression-specific scores were significantly correlated when extracted using the Bartlett method (r = 0.638, p < 0.001), whereas, as expected, they were effectively uncorrelated when extracted using the Anderson–Rubin method (r < 0.001, p = 1.000). This comparison suggests that the Anderson–Rubin method is more appropriate for our analytic aim of estimating the unique associations of shared and symptom-specific components with task-derived parameters, because it separates the general distress/internalizing factor from the statistically separable residual anxiety- and depression-specific components. For this reason, we retained the Anderson–Rubin factor scores in the main analyses. We have clarified this point in the revised manuscript as follows:

      Supplementary Page 3:

      “For factor score extraction, we followed prior work using bifactor modeling in computational psychiatry[10] and extracted factor scores with the Anderson–Rubin method, implemented using psych::factor.scores with method = "Anderson". This approach yields standardized and mutually orthogonal factor scores, which is particularly appropriate for our subsequent correlation analyses because it avoids multicollinearity among the general, depression-specific, and anxiety-specific factors. As a robustness check, we also extracted factor scores using the Bartlett method from an oblique bifactor model, which allowed the depression- and anxiety-specific factors to correlate. In the combined dataset (n = 1,026), anxiety- and depression-specific scores were significantly correlated when extracted using the Bartlett method (r = 0.638, p < 0.001), whereas, as expected, they were effectively uncorrelated when extracted using the Anderson–Rubin method (r < 0.001, p = 1.000). This comparison suggests that the Anderson–Rubin method is more appropriate for our analytic aim of isolating the unique contributions of shared and symptom-specific variance, because it separates the general distress/internalizing factor from residual anxiety- and depression-specific components. Thus, we retained the Anderson–Rubin factor scores in the main analyses.”

      Finally, the similarities and differences between our hierarchical factor-analytic approach and the recent transdiagnostic hierarchical factor-analytic approach of Wise et al. (2026). Both approaches fit EFA models with different numbers of factors and use cross-level correlations to characterize hierarchical symptom structure. However, Wise et al. (2026) applied this framework to a broader symptom battery covering transdiagnostic and neurodevelopmental dimensions, identifying a hierarchy that included a general psychopathology factor and more specific dimensions such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal. In contrast, our study focused more narrowly on anxiety and depression dimensions, with the goal of deriving symptom factors that could be linked to task-derived computational parameters. Accordingly, whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-related variance remains unknown. Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026), to better determine whether the links among symptom dimensions, RPE-related mood sensitivity, and mood variability are disorder-specific or transdiagnostic. We have discussed this point in the revised manuscript as follows: Page 19:

      “Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026)[59], to better characterize whether links among symptom dimensions, RPE sensitivity, and mood variability are disorder-specific or transdiagnostic.”

      (2) Linking factors to task parameters

      As I understand it, the authors relate the orthogonalized depression/anxiety to task parameters (sensitivity to RPEs on mood and mood variations) using correlations. In order to have a better understanding of how this relates to other commonly used approaches, I would pose two questions:

      (i) What are the correlations when the full (non-orthogonalized) factor scores for depression and anxiety are used? Are the signs the same?

      (ii) What are the results when, instead of the independent correlations, the authors perform b_RPE ~ anxiety + depression (again using the non-orthogonalized factors)? I'm assuming all of these analyses should give the same results if the authors' hypothesis of opposing effects of anxiety and depression holds true.

      We thank the Reviewer for these helpful comments. Our original analyses used orthogonalized depression- and anxiety-specific factor scores because this approach is aligned with our analytic aim of separating shared and symptom-specific variance within the tripartite/bifactor framework of anxiety and depression (Clark & Watson, 1991), and has been used in prior work (Gagne et al., 2020, 2022; see Wise et al., 2023 for a review). Orthogonalization allows us to statistically separate the shared distress component from the symptom-specific components of anxiety and depression, which was central to our hypothesis regarding their opposing associations with task-derived parameters. As expected, the orthogonalized anxiety- and depression-specific factor scores were uncorrelated in the combined dataset (r < 0.001, p = 1.000; n = 1,026). By contrast, the full non-orthogonalized depression and anxiety scores retained substantial shared variance and were highly correlated (r = 0.638, p < 0.001; n = 1,026), making their separate associations less straightforward to interpret.

      Nevertheless, we agree that analyses using the full non-orthogonalized depression and anxiety scores provide an important comparison with more commonly used non-orthogonal symptom-score approaches. We therefore conducted the analyses suggested by the Reviewer. When the full depression and anxiety scores were entered separately into linear mixed-effects models predicting RPE-related mood sensitivity, with dataset included as a random intercept, the anxiety association was not significant (anxiety: b = 0.002, t = 1.130, p = 0.259; depression: b = -0.008, t = -3.732, p < 0.001). By contrast, when the full non-orthogonalized anxiety and depression scores were entered simultaneously in the same linear mixed-effects model, the original pattern was replicated: anxiety and depression showed opposing associations with RPE-related mood sensitivity (anxiety: b = 0.012, t = 4.619, p < 0.001; depression: b = -0.017, t = -5.851, p < 0.001). This pattern is consistent with a mutual suppression effect: shared variance between anxiety and depression may obscure their unique associations when examined separately, whereas the simultaneous regression model reveals their opposing symptom-specific associations.

      Together, these supplementary analyses support our original interpretation that RPE-related mood sensitivity is associated with the separable anxiety- and depression-specific components in opposite directions. We have revised the manuscript as follows:

      Supplementary Pages 3-4:

      “We further analyzed non-orthogonalized full depression and anxiety scores to assess the robustness of our results. When full depression and anxiety scores were entered in separate linear mixed-effects models predicting RPE-related mood sensitivity, with dataset included as a random intercept, the anxiety association was not significant (anxiety: b = 0.002, t = 1.130, p = 0.259; depression: b = -0.008, t = -3.732, p < 0.001). By contrast, when the full non-orthogonalized anxiety and depression scores were entered simultaneously in the same linear mixed-effects model, the original pattern was replicated: anxiety and depression showed opposing associations with RPE-related mood sensitivity (anxiety: b = 0.012, t = 4.619, p < 0.001; depression: b = -0.017, t = -5.851, p < 0.001). This pattern is consistent with a mutual suppression effect: shared variance between anxiety and depression may obscure their unique associations when examined separately, whereas simultaneous regression reveals their opposing symptom-specific associations. These results support our interpretation that RPE-related mood sensitivity is linked to the separable anxiety- and depression-specific components.”

      Minor comments:

      (1) The authors should write down when the data were collected for each study. This is because AI capabilities have massively increased since ~2020 in quite specific steps (with the public release of new AI models), meaning that AI is likely to have been able to do tasks and questionnaires without detection if data were collected recently.

      We thank the Reviewer for this important comment. We have now added the data collection periods for each dataset in Table 1. The laboratory and clinical dataset were collected in a controlled laboratory setting rather than through online testing. As shown in Table 1, all online experiments were conducted before November 2022, prior to the public release of ChatGPT and its broad entry into public awareness. Therefore, our data were unlikely to have been substantially affected by AI-assisted responding. We have clarified this point in the revised manuscript as follows:

      Page 20:

      “See Table 1 for demographic information and data collection periods. Because online data collection may raise concerns about AI-generated responses, we note that artificial intelligence tools, such as ChatGPT, became widely known to the public in November 2022, whereas all online experiments in the present study were conducted before November 2022 (see Table 1). Therefore, these data were unlikely to have been substantially affected by participants’ use of AI tools.”

      (2) The authors should include a statement in the methods section that checks for AI were done. If none yet, could you do any? Recent papers (Westwood, PNAS 2025; van der Stigchel PNAS, 2026) point to the risk since at least the release of o4-mini (used in the cited paper to create very human-like behaviour).

      We thank the Reviewer for this helpful comment. We have clarified this point in the revised manuscript as follows:

      Page 20:

      “See Table 1 for demographic information and data collection periods. Because online data collection may raise concerns about AI-generated responses, we note that artificial intelligence tools, such as ChatGPT, became widely known to the public in November 2022, whereas all online experiments in the present study were conducted before November 2022 (see Table 1). Therefore, these data were unlikely to have been substantially affected by participants’ use of AI tools.”

      (3) It would have been good to collect questionnaires of other, thought to be unrelated psychiatric traits, like compulsivity or schizophrenia symptoms, to check the specificity of the results, also under the assumption that higher scores on either of these skewed questionnaires can pick up individual differences in 'bad questionnaire completion'. The authors should comment on the absence of other questionnaires in the discussion in the limitations section.

      We thank the Reviewer for this helpful comment. We agree that the absence of broader psychiatric trait measures limits our ability to evaluate the specificity of the observed associations. Our symptom assessment focused specifically on anxiety and depression because the study was motivated by hypotheses about their potentially opposing links with RPE sensitivity and mood variability. Although previous research has shown intact mood sensitivity to RPEs in individuals with suicidal thoughts and behaviors (Wang et al., 2026), we cannot determine whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-style variance.

      Regarding the concern that higher scores on symptom questionnaires with skewed score distributions may partly capture individual differences in poor-quality questionnaire responding, we note that we implemented strict data-quality procedures for both questionnaire and task data. Four attention-check items were embedded throughout the questionnaire battery, requiring participants to select a prespecified response, for example, “Please select the second option for this item.” Similarly, four attention-check trials were embedded throughout the gambling task. For example, participants were asked to choose between a certain gain of 20 points and a gamble with possible outcomes of 35 and 55 points, for which the dominant response was to choose the gamble option. Participants who failed any of these attention checks were excluded. In addition, our behavioral and mood data reproduced key patterns reported in previous studies (Rutledge et al., 2014 & 2015) using momentary mood ratings during gambling tasks, including higher mood following gains than following losses and systematic mood drift over time (all ps < 0.001). These procedures and validation checks reduce the likelihood that the present findings were driven by poor questionnaire or task completion. We have clarified this point in the revised manuscript as follows:

      Page 19:

      “Second, our symptom assessment focused specifically on anxiety and depression. This choice was motivated by our primary hypotheses, but it limits our ability to evaluate the specificity of the observed associations. Recent work has shown that individuals with suicidal thoughts and behaviors exhibit reduced mood sensitivity to certain rewards (CR), but not to RPEs[49], suggesting that the current RPE-related effects are not driven by suicide-related processes. However, because we did not assess other psychiatric dimensions, such as compulsivity or schizophrenia-spectrum symptoms, we cannot determine whether the current findings are specific to anxiety- and depression-related symptom dimensions or instead reflect broader transdiagnostic psychopathology or nonspecific response-related variance. Future studies should include measures covering a wider range of psychiatric dimensions, such as internalizing, externalizing, inattentive/neurodevelopmental, mood/anxiety, and withdrawal dimensions identified by Wise et al. (2026)[59], to better characterize whether links among symptom dimensions, RPE sensitivity, and mood variability are disorder-specific or transdiagnostic.”

      Page 20:

      “Participants were excluded if 1) they failed any of the attentional checks (4 items); 2) they made the same choices for all items; 3) they responded with extreme inconsistency in two similar questionnaires (difference in z-scores out of ±2).”

      Page 21:

      “There were four items for attentional checks, which required the participants to make a specific choice and were embedded in the entire measurements, e.g., ‘please select the second option for this item’.”

      Page 22:

      “We also set 4 trials embedded in the entire task for attentional checks. For example, participants were asked to make a choice between a certain gain 20 and a gamble 35/55, where the correct response for this trial was the gamble choice.”

      Page 5:

      “Choice data (e.g., gambling rates) and mood data (e.g., initial mood, mean mood, and mood variation) showed patterns similar to those reported in previous studies measuring momentary mood during gambling tasks (Figure S2 & S3)[10,45]. We also replicated established effects on momentary mood: mood was higher following gains than following losses, and mood drifted over time (all ps < 0.001; Figure S4).”

      (4) The authors could include a more explicit sentence in the abstract stating that the anxiety result did not hold up in the clinical population.

      We thank the Reviewer for this helpful comment. We have clarified this point in the revised manuscript as follows:

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      Reviewer #2 (Public review):

      Summary:

      Despite their common co-occurrence, depression and anxiety are known to alter mood fluctuations in opposite ways. Here, the authors aimed at distinguishing depression-specific from anxiety-specific from psychopathology-general effects of reward processing on mood fluctuations, focusing on reward prediction errors (RPEs), which are known to be linked to mood fluctuations. This mechanistic study aims at uncovering the process through which these psychopathologies are associated with mood modulations. The authors were able to appropriately test their hypothesis and obtained results corroborating their conclusions.

      This work provides a convincing demonstration of the relevance of computational psychiatry (Huys et al, 2016) and the use of decision neuroscience to shed light on the interplay of anxiety, depression, and mood.

      Strengths:

      The authors used a tripartite model to distinguish depression vs anxiety, as well as a computational model distinguishing reward expectation (EV in the model) from outcome processing through RPE, which are two sequential cognitive processes.

      The manuscript adequately addresses the concerns one would have regarding risk-attitudes and regarding referring to trending statistical results.

      Weaknesses:

      The sample size of the clinical sample (N=116) may not be sufficient to detect anxiety-specific effects due to the high rate of comorbid anxious depression. It would be beneficial to include the number of MDD vs GAD vs anxious depression diagnoses in the clinical population, as this would likely shine light on the power limitations.

      We thank the Reviewer for this helpful comment. We agree that the clinical sample may have been underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (see Table S8 for diagnosis, illness duration, and medication status). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). Although these covariate analyses support the robustness of the depression-related effect, they do not resolve whether the absence of the anxiety-related effect reflects limited power or true clinical discontinuity.

      We also revised the Discussion to explicitly acknowledge that the anxiety-related effect observed in the pooled non-clinical dataset was not replicated in the clinical sample. We now note two possible interpretations. First, this discontinuity may reflect limited statistical power in the clinical sample. Second, and more speculatively, it may reflect a disruption of mood homeostasis in affective disorders (Paulus, 2007). In non-clinical individuals, the counterbalancing associations of depression- and anxiety-related traits with mood variation may contribute to emotional equilibrium. In contrast, affective disorders may involve a loss of this regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals. We have revised the manuscript as follows:

      Pages 12-13:

      “To test whether abnormalities in RPE-driven mood fluctuations can serve as clinically relevant computational markers of depression- and anxiety-related symptom dimensions, we recruited patients with affective disorders (n = 116) to complete the same questionnaire battery and gambling task with momentary mood ratings (Figure 1). Demographic, psychological, and clinical characteristics are summarized in Table 1 and Table S8. We observed significant negative correlations between depression-specific scores and both mood variation (r = -0.239, p = 0.009) and RPE-related mood sensitivity (β<sub>RPE</sub>; r = -0.216, p = 0.020). These associations remained significant after controlling for demographic and clinical covariates, task earnings, and mood drift (ps < 0.05). Bootstrap validation yielded consistent results. Mediation analyses further showed that reduced mood sensitivity to RPEs statistically mediated the association between depression-specific scores and lower mood fluctuations (a × b = -0.141, 95% CI = [-0.261, -0.038], p = 0.021; Figure 3). However, we did not observe significant correlation with anxiety (mood variation: r = -0.092, p = 0.327; β_RPE: r = -0.095, p = 0.311).”

      Pages 17-18:

      “Notably, the pattern of heightened RPE sensitivity observed in the pooled non-clinical dataset was not observed in the clinical sample. On the one hand, this discontinuity may reflect that the clinical sample was underpowered to detect anxiety-specific effects, especially given the high comorbidity between anxiety and depression in affective disorders (Table S8). Based on the effect size observed in the non-clinical datasets (r = 0.079), we estimated that a sample size of 1,226 would be required to detect this effect with 80% statistical power using a two-tailed test with α = .05. This estimate is substantially larger than the current clinical sample size (n = 116). On the other hand, it may reflect a disruption of mood homeostasis in clinical populations[41,58]. In non-clinical individuals, counterbalancing associations of depression- and anxiety-related traits with mood variation may help maintain emotional equilibrium. In contrast, affective disorders may involve a loss of such regulatory balance, reducing the ability to stabilize mood in the face of competing depression- and anxiety-related affective signals.”

      Abstract:

      “Results showed that depression was associated with dampened mood fluctuations due to mood hyposensitivity to RPE. Importantly, this pattern was also found in patients with affective disorders. In contrast, anxiety correlated with heightened mood fluctuations stemming from mood hypersensitivity to RPE in non-clinical participants.”

      Reviewer #3 (Public review):

      Summary:

      In this submission, Wang and colleagues jointly examine the association between depression and anxiety symptoms and individuals' affective reactivity to reward prediction errors in Ruttledge et al.'s gambling paradigm. Taking a bifactor approach to anxiety and depression in several non-clinical (and one clinical sample), the authors find that anxiety-specific symptoms relate to over-reactivity of mood to reward prediction errors (RPEs) as well as heightened mood variability, while depression-specific symptoms relate to blunted mood sensitivity to RPEs. These depression- but not anxiety-specific relationships replicated in patient samples.

      Strengths:

      I was impressed that the data-driven, transdiagnostic approach employed by the authors uncovered specific relationships between anxiety and depression-specific factors and RPE reactivity in a well characterized task and computational model, especially in a non-clinical sample. This sheds new light on how these affective processes may be perturbed-and importantly, in different ways-by anxiety and depression symptoms. Likewise, the replication of the depression-specific finding (RPE hypo-reactivity) in a clinical sample was nice to see.

      Weaknesses:

      (1) While the anxiety- and depression-specific factors had differential effects on mood variability (Figure 2A-D) and RPE reactivity (Figure 2E-G) in all samples, such that the correlations between the two factors and these mood parameters were significantly different, the anxiety factor was not consistently (significantly) associated with either mood-related parameter across samples. However, the authors resolve anxiety-specific predictive effects when they collapse across datasets. While it is intuitive that achieving a larger effective sample size would afford the power necessary to detect such individual differences, this struck me as a major caveat for this set of results.

      We thank the Reviewer for this important comment. Although the anxiety-specific factor showed associations in the expected direction across datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level.

      Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). In addition, we conducted a mini meta-analysis, and results support that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE.

      However, we have clarified in the revised manuscript that the anxiety-related effects were less robust than the depression-related effects and require further replication in larger samples.

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β<sub>RPE</sub> were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept to account for dataset-level variability. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001). These associations remained significant after controlling for gender, age, task earnings, and mood drift. In addition, we performed a mini meta-analysis on these correlation coefficients[49]. Results showed significant positive correlation for both mood variation and RPE-related mood sensitivity (mood variation: Z = 3.399, 95 % CI for correlation coefficient r [0.045, 0.166]; RPE-related mood sensitivity: Z = 2.618, 95 % CI for correlation coefficient r [0.021, 0.143]), supporting that anxiety is associated with intensified mood fluctuations and increased mood sensitivity to RPE.”

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      (2) The authors observe associations between the 'common factor' of depression and anxiety and risk-attitude tendencies, presumably the alpha (exponent) parameter in a prospect theory-type subjective value model. But where is this analysis explained? (i.e. how was this model formulated and how were risk attitude parameters estimated?) And what is the interpretation of this finding - is there precedent for looking at risk attitudes in this task? And why would these predictive effects only be observed in relation to the common, but not unique, factors of anxiety and depression?

      We apologize for the unclear statement. We have added a description of the computational modeling of choice behavior. Please see our revisions below:

      Page 13:

      “Choice parameters were estimated using an established approach–avoidance prospect theory model[10,45,49], which included loss aversion, domain-specific risk attitude parameters in the gain and loss domains, and value-independent Pavlovian approach and avoidance parameters (see Supplementary Note 7 for details of the computational choice models). In this model, risk attitude was quantified by the exponent parameter α in a prospect-theory-inspired subjective value function. Lower α values reflect greater risk aversion, whereas values closer to or above 1 reflect more linear or risk-seeking valuation.”

      Supplementary Pages 8-9:

      Note 7: Computational model of gambling choice

      To quantify how different events impacted participants’ momentary moods during the gambling In line with previous studies[14,15], our choice model space included expected value model (cM1), prospect theory model (cM2)[16], and approach-avoidance prospect theory model (cM3)[14]. For cM2 (Equations 6-9), there were 3 parameters, including risk aversion (α, range: [0.3, 1.3]), loss aversion (λ: [0.5, 5]), and inverse temperature (μ: [0, 10]).

      Where V<sub>gain</sub> and V<sub>loss</sub> are the objective gain and loss from a gamble, respectively. Please note thatV<sub>gain</sub> is 0 in loss trials and V<sub>loss</sub> is 0 in gain trials. V<sub>certain</sub> is the objective value for the certain option. U<sub>gamble</sub> and U<sub>certain</sub> denote the subjective utilities of the gamble and the certain option, respectively. Choice probability for gamble (P<sub>gamble</sub>) is determined by the softmax rule. Building on cM2, cM3 decomposes the decision process into risk-attitude-driven valuation (e.g., loss and risk aversion) and value-insensitive motivational components (Equations 6-8 & 10-12). That is, choice probability for P<sub>gamble</sub> in cM3 is jointly determined by the softmax rule and approach/avoidance parameters (β<sub>gain</sub>: [-1, 1], β<sub>loss</sub>: [-1, 1]). Approach/avoidance parameters are not applied in mixed trials. Please note that a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups. Risk attitude is indeed conceptualized in economics as the curvature of the utility function (i.e., the subjective value) of the objective outcomes, with concave curves associated with risk aversion, and convex curves associated with risk seeking[17,18]. By contrast, the approach or avoidance bias apply to all the value. A possible interpretation of the approach bias is that participant approach the option with the highest possible gain (the lottery) in the gain frame; the avoidance bias would then reflect a tendency to systematically avoid the highest potential losses (the lottery) in the loss frame.

      Model comparison using BIC revealed that the winning model for each dataset was the approach-avoidance prospect theory model (cM3; mean R<sup>2</sup> = 0.51 for the laboratory dataset, 0.49 for the online dataset1, 0.54 for the online dataset 2, and 0.40 for the clinical dataset; Table S9).

      Please also see our interpretation of this finding below:

      Page 18:

      “With respect to decision-making, prior literature using risky decision-making tasks without feedback has linked pathological anxiety to greater risk aversion[58]. In line with this, our results from a risky decision-making task with feedback suggest that the common factor, rather than anxiety-specific variance per se, is more consistently associated with risk aversion. This suggests that heightened gain-domain risk aversion may be a transdiagnostic feature of internalizing psychopathology, rather than being uniquely attributable to anxiety.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Thank you very much for giving me the opportunity to review this very interesting paper.

      Recommendations:

      (1) Add more specific ethics information than "study was approved by ethics committee of Beijing normal university".

      We thank the Reviewer for this important comment. We have added approval number. Please see our revision below:

      Page 20:

      “The study was approved by the Ethics Committee of Beijing Normal University (approve number: ICBIR_A_0016_028). Written or electronic informed consent was obtained from all participants before participation.”

      (2) Add information on how participants were recruited. I think the websites listed only hosted the experiment/questionnaires?

      We thank the Reviewer for pointing this out. We have revised the relevant text as follows:

      Page 20:

      “A total of 2634 participants via online platforms (questionnaires from https://www.wjx.cn and tasks from https://www.naodao.com) took part in five experiments, including a psychometric experiment, a laboratory experiment, two online replication experiments. Participants were recruited through participant pools and study advertisement. For online experiments, interested participants accessed the study through an online link and completed the questionnaires and task remotely. For the laboratory experiment, participants completed the study in a controlled laboratory setting.”

      (3) Typo in Figure 1A, grey panel - psychometric.

      We apologize for the typo. We have corrected typographical errors throughout the manuscript.

      (4) In the Discussion, there is a section on r-to-z transformations, and I was not quite sure what in the Results this links to.

      We thank the Reviewer for pointing out the unclear statement. Please see our revision below:

      Page 18:

      “First, although anxiety- and depression-related associations differed consistently, the anxiety-specific associations themselves were less robust across datasets.”

      **Reviewer #2 (Recommendations for the authors):&&

      The Results sections 2 (depression) and 3 (anxiety) could be improved by reducing the back and forth between factors throughout the results. It may be useful to split them into 3 sections: depression only, anxiety only, and depression vs anxiety.

      We thank the Reviewer for this helpful suggestion. As suggested, we have reorganized this part into three sections: depression, anxiety, and depression versus anxiety. Please see our revision below:

      Page 12:

      “Differential associations of depression and anxiety with mood fluctuations. To directly test whether depression- and anxiety-specific factors differed in their associations with mood dynamics, we compared the corresponding correlations. These comparisons showed that depression-specific associations were significantly more negative than anxiety-specific associations for both mood variation (laboratory dataset: Z = -1.84, p = 0.033; online dataset 1: Z = -5.36, p < 0.001; online dataset 2: Z = -3.42, p < 0.001) and β_RPE (laboratory dataset: Z = -1.77, p = 0.038; online dataset 1: Z = -3.67, p < 0.001; online dataset 2: Z = -4.00, p < 0.001; Figures 2A–C and 2E–G). These results support distinct associations of depression- and anxiety-specific factors with RPE-related mood dynamics.”

      In the discussion, the authors could have mentioned the brain areas most likely to be involved in these processes, both cognitive and psychopathological, as previous studies (such as Cecchi et al, 2022) have aimed at identifying regions involved in RPE processing while modulating mood in health. A short section on this would be useful to the neuropsychiatric community.

      We thank the Reviewer for this helpful suggestion. We agree that the Discussion would benefit from a more explicit consideration of the neural systems that may support RPE-related mood updating and their relevance to psychopathology. We have revised the Discussion accordingly, as shown below:

      Page 16:

      “Although the present study did not include neuroimaging, the observed computational dissociation may map onto partially distinct neural systems involved in reward learning, mood updating, and affective psychopathology. RPE processing has been consistently linked to striatal–midbrain dopaminergic reward-learning circuits[8,50]. The integration of these reward-learning signals into subjective mood and value-based decision-making may further involve the ventral medial prefrontal cortex and orbitofrontal cortex[44]. In addition, the anterior insula may be particularly relevant for integrating feedback-related signals with affective and interoceptive states[8,44], potentially linking RPE processing to anxiety- and depression-related mood dynamics. Consistent with this view, Cecchi et al. (2022)[51] used intracranial EEG to show that feedback-related neural activity tracks mood fluctuations and risky choice. Future neuroimaging studies should test whether depression-related reductions and anxiety-related increases in RPE-related mood sensitivity are associated with altered interactions among striatal, prefrontal, and insular circuits.”

      Reviewer #3 (Recommendations for the authors):

      (1) The authors need to present a clearer definition of the terms "bifactor analysis" and "tripartite model" in the Introduction. What does tripartite mean in this context? What are the assumptions of such bifactor analyses (e.g. as used in Gagne et al. and the present work) and how, in broad strokes, are they carried out? These are important constructs to clarify for readers outside the computational psychiatry niche.

      We thank the Reviewer for this helpful suggestion. We have revised the Introduction accordingly, as shown below:

      Page 3:

      “Recent work has used bifactor models of the tripartite model of depression and anxiety to clarify their distinct features and differential influences on decision-making[31,32]. The tripartite model of anxiety and depression proposes that these two symptom dimensions share a broad general distress or negative affect component while also including symptom-specific components: low positive affect/anhedonia is more specific to depression, whereas physiological hyperarousal is more specific to anxiety[30,33,34]. Bifactor analysis offers a way to model this structure statistically. In a bifactor model, symptoms load on a general factor reflecting their shared variance and on specific factors capturing residual variance in narrower symptom dimensions after accounting for the general factor.”

      (2) Previous examinations of depression and RPE reactivity in this task paradigm, as the authors note (e.g. Rutledge et al., 2017), observed that individuals diagnosed with depression showed an intact association between RPEs and mood. In other words, there was no previously observed relationship between depression and affective reactivity to RPEs in this task context. Here, the authors find that the "unique" depression factor identified by the authors (in a non-clinical sample) is associated with blunted RPE sensitivity - this is worth commenting on specifically.

      We thank the Reviewer for this helpful suggestion. We have discussed this point in the Discussion. Please also see it below:

      Pages 15-16:

      “Our computational model not only replicates the important role of RPEs in mood dynamics but also highlights the divergent mediating roles of RPE-related mood sensitivity in the associations of depression and anxiety with mood fluctuations. The opposite associations of depression and anxiety with mood sensitivity to RPEs complement previous findings of apparently intact RPE-related mood sensitivity in depression[12,25,37]. These findings further underscore the necessity of decomposing shared and specific components of depression and anxiety in studies of mood dynamics, which can enhance our understanding of their distinct associations with emotion processing and cognitive flexibility. This point is consistent with bifactor-based work showing that shared and specific symptom dimensions can have different computational correlates. For example, Gagne et al. (2020) showed that bifactor-derived symptom dimensions differentially relate to maladaptation to environmental volatility[32], complementing previous findings that trait anxiety is associated with inflexible adjustment to volatility[32].”

      (3) There is a note (line 247) about the interpretation of the correlations in Figure 2, which attempts to explain away the inconsistent relationships between the anxiety-specific factor and mood variability as well as RPE reactivity observed in Figure 2. I can't say I understand the authors' point here about "signs of positive correlations", so I would say the authors need to clarify their logic here. More to the point, the authors only resolve anxiety-specific predictive effects when they collapse across these datasets. As discussed above (see 'weaknesses'), this is a serious limitation in my view and needs to be discussed as such in the paper.

      We apologize for the unclear statement. Although the anxiety-specific factor showed associations in the expected direction across datasets, these associations were not significant in several individual datasets. Specifically, anxiety-specific scores were positively correlated with mood variation (laboratory dataset: r = 0.10, p = 0.531; online dataset 1: r = 0.08, p = 0.026; online dataset 2: r = 0.19, p = 0.004) and with RPE-related mood sensitivity (laboratory dataset: r = 0.04, p = 0.820; online dataset 1: r = 0.05, p = 0.216; online dataset 2: r = 0.19, p = 0.004; Figures 2A–C and 2E–G). This pattern may partly reflect limited statistical power at the single-dataset level.

      Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, we conducted pooled analyses to obtain a more stable estimate. Importantly, these analyses included dataset as a random intercept in mixed-effects models to account for between-dataset differences. Thus, the pooled analysis provides an integrated estimate across samples, conceptually similar to an individual-participant-data meta-analytic approach. The pooled results provided evidence for the expected anxiety-specific associations with greater mood variability and heightened RPE-related mood sensitivity (mood variation: t = 3.46, p < 0.001; RPE-related mood sensitivity: t = 2.60, p = 0.009). We have revised it to make it clear. Please see our revisions below:

      Pages 10-11:

      “Correlations between the anxiety-specific factor and mood variation were positive in direction across datasets, although they were not statistically significant in several datasets (the laboratory dataset: r = 0.10, p = 0.531; the online dataset 1: r = 0.08, p = 0.026; the online dataset 2: r = 0.19, p = 0.004). Similarly, correlations between the anxiety-specific factor and β_RPE were positive in direction but statistically inconsistent across datasets (the laboratory dataset: r = 0.04, p = 0.820; the online dataset 1: r = 0.05, p = 0.216; the online dataset 2: r = 0.19, p = 0.004; Figure 2A-C & 2E-G). Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect, we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis while accounting for dataset-level variability. Because these datasets used comparable task and questionnaire procedures and showed positive effect directions, and because reliable individual differences often require large samples to detect[48], we combined the laboratory dataset, online dataset 1, and online dataset 2 (total N = 1,026). This approach is analogous to an individual-participant-data meta-analytic analysis while accounting for dataset-level variability. We fitted linear mixed-effects models predicting mood variation and β<sub>RPE</sub> from the three bifactor scores, with dataset included as a random intercept. For mood variation, the anxiety-specific factor was positively associated with mood variation (t = 3.46, p < 0.001), whereas the depression-specific factor was negatively associated with mood variation (t = -6.13, p < 0.001). For RPE-related mood sensitivity, the anxiety-specific factor was positively associated with β<sub>RPE</sub> (t = 2.60, p = 0.009), whereas the depression-specific factor was negatively associated with β<sub>RPE</sub> (t = -5.30, p < 0.001).”

      Page 17:

      “Notably, the anxiety-related effects were less robust than the depression-related effects and were detectable only in the pooled dataset (n = 1,026); therefore, they require further replication in larger samples.”

      (4) I expected to see that the authors would also investigate relationships between anxiety/depression related factors and the decay (gamma) parameter in the 'Happiness equation', which is presumably estimated from the data here. While I don't have a strong intuition about directions of (or presence of) predictive relationships here, doesn't it stand to reason that different aspects of psychopathology examined here might map onto how long- (versus short-) lasting the effects of, say, RPEs are, upon mood?

      We thank the Reviewer for this important comment. In the healthy datasets, we fitted a linear mixed-effects model predicting the decay parameter (γ) from the three bifactor scores, with dataset included as a random intercept. None of the factors showed a significant association with γ (common: t = 0.708, p = 0.479; anxiety: t = 0.564, p = 0.573; depression: t = 1.146, p = 0.252). In the clinical dataset, we fitted a linear model predicting γ from the three bifactor scores and again found no significant associations (common: t = -0.036, p = 0.972; anxiety: t = -0.047, p = 0.963; depression: t = 0.052, p = 0.959). We have clarified this point in the revised manuscript as follows:

      Supplementary Page 5:

      “In the healthy datasets, we conducted a linear mixed-effect model against decay parameter (gamma) with all three factors, with dataset as a random factor. Results did not show significant effect (common: t = 0.708, p = 0.479; anxiety: t = 0.564, p = 0.573; depression: t = 1.146, p = 0.252). In the clinical dataset, we conducted a linear model against decay parameter (gamma) with all three factors and found no significant effect (common: t = -0.036, p = 0.972; anxiety: t = -0.047, p = 0.963; depression: t = 0.052, p = 0.959).”

      (5) The rationale for and interpretation of the mediation model, which presumably aims to explain the relationships between anxiety- and depression-specific factors, RPE reactivity, and mood variability was barely explained by the authors. At present, I'm not sure what the added value of this analysis is. The authors should either remove or explain/motivate the mediation more clearly.

      We thank the Reviewer for this helpful comment. The rationale for the mediation analysis is that mood variability in the task is not only a descriptive behavioral outcome, but may also arise from the degree to which momentary mood is updated by RPEs. Therefore, if depression is associated with reduced RPE-related mood sensitivity and anxiety with increased RPE-related mood sensitivity, these alterations should statistically account for their opposite associations with mood variability. The mediation model directly tested this possibility by examining whether RPE-related mood sensitivity accounted for the association between symptom-specific factors and mood variability. We have clarified this rationale in the revised manuscript as follows:

      Page 10:

      “Given the strong correlation between β<sub>RPE</sub> and mood variation (rs > 0.67, ps < 0.001), we further conducted a mediation analysis to examine whether individual differences in RPE-related mood sensitivity statistically accounted for the association between depression loading and mood variation. This analysis was motivated by the hypothesis that depression-related dampening of mood variability may arise, at least in part, from reduced mood sensitivity to RPEs.”

      (6) This submission would benefit from extensive English language copy editing. There are many passages in the paper (in fact, too many to list here) that suffer from either grammatical errors or clarity issues.

      We apologize for these mistakes. We have corrected the typographical errors throughout the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript is an excellent follow-up to your 2022 study, in which Sox17 expression was localized to the rete testis and shown to be required for proper formation of the Sertoli cell valve (transition region). By using Nr5a1-Cre to drive conditional deletion of Sox17 specifically in rete testis cells, you demonstrate that testis weights remain normal at 2 weeks of age but become significantly reduced by 8 weeks in Sox17-cKO males. At the later time point, the seminiferous epithelium is severely disrupted, with apparent arrest of spermiogenesis: the epididymal lumen is essentially devoid of sperm, and most tubules lack elongated spermatids.

      Strengths:

      The study clearly shows the role of Sox17 in Sertoli cells as being important to SV function. The SV (transition region) between the rete testis and seminiferous tubules remains an understudied domain of testicular biology. The present work, together with the authors' prior study, highlights intriguing mechanisms operating in this specialized niche.

      Weaknesses:

      At the same time, the available data do not yet fully explain either the developmental assembly of the Sertoli valve or the precise consequences of its functional disruption. These studies are nonetheless valuable precisely because they raise more questions than they answer; the conceptual implications are thought-provoking.

      Reviewer #2 (Public review):

      This manuscript investigates the role of SOX17 in the formation and function of the Sertoli valve (SV) at the interface between seminiferous tubules and the rete testis (RT). Building on previous work showing that rete testis-specific deletion of Sox17 disrupts SV formation, leading to defective spermiogenesis and male infertility, the authors explore how SOX17 overexpression in Sertoli cells regulates the SV of rodent testes.

      Using transgenic mouse models with ectopic Sox17 expression in Sertoli cells, the study demonstrates that SOX17 is not only required but can also modulate SV formation. Ectopic expression in Sertoli cells induces expansion of the SV structure and partially rescues SV defects and spermatogenesis in RT-specific Sox17 conditional knockout animals. The data support a model in which SOX17 acts through paracrine signaling to regulate SV formation, although the precise mechanisms remain to be clarified.

      Overall, this is a well-executed study with novel and significant findings. The ability to experimentally manipulate SV size is particularly compelling and provides a valuable framework to study fluid dynamics and epithelial interactions in the testis. This work will be of broad interest to the reproductive biology and developmental biology communities.

      Reviewer #3 (Public review):

      Summary:

      These studies are based on previously published work that showed that deletion of expression of the Sox17 gene in the testis essentially deleted the formation of the Sertoli valve in the Rete testis. The authors extended this work by constructing a vector that resulted in increased Sox17 expression by Sertoli cells and enhanced formation of the Sertoli valve in both wild type and Sox17 knockout mice. The work provides strong evidence supporting the requirement for Sox17 expression to allow formation of the Sertoli valve.

      Strengths:

      The general approach was to express Sox17 from a Tg mouse that expressed Sox17 from Sertoli cells. This Tg mouse was bred into both the WT and the Sox17 KO mouse. The Sertoli valve was enhanced in both the WT/Tg mouse and KO/Tg mouse, showing that ectopic Sox17 could compensate in the Sox17 Ko and act in a concentration-dependent manner in the WT mouse. The results are strong and support the conclusions from the authors. The results were as expected from the original paper describing the KO of Sox 17. These results strengthen these conclusions and provide ideas for additional conclusions. These studies were technically challenging, and the authors provided a very solid manuscript.

      Weaknesses:

      The authors refer several times to high or low expression, but it all appears to be based on immunohistochemistry, and there is no real quantification using PCR, for example. The process used for cell quantification lacks a rationale for why certain numbers were assigned.

      We sincerely thank the reviewers for their careful evaluation of our manuscript and for their constructive and encouraging comments. We are grateful for the recognition of the significance of the Sertoli valve as an understudied transition region between the rete testis and seminiferous tubules, as well as for the positive assessment of our genetic approach and the evidence that ectopic SOX17 expression can modulate SV formation. We have carefully considered all points raised in the assessment and have revised the manuscript accordingly. The major revisions include:

      (1) Clarification of the scope and limitations of the study (Reviewers #1 and #2):

      In response to the comments that the developmental assembly of the Sertoli valve and the precise consequences of its functional disruption remain incompletely understood, we clarified the scope and limitations of the present study at the end of 7th paragraph in the Discussion. Although our findings support a model in which SOX17 regulates SV formation through paracrine signaling, the downstream effectors and precise molecular mechanisms remain to be identified. We therefore revised the Discussion to avoid overinterpretation of the molecular mechanisms and to emphasize that comprehensive mechanistic analyses, including transcriptomic analyses using the Tg mouse model, represent an important direction for future research. We also added histological analyses of the earliest detectable lesions at 4 weeks of age and low-magnification images of adult Sox17 cKO testes (new Figure S1), revealing selective sloughing of round spermatids despite preserved Sertoli cell architecture and subsequent mosaic spermatogenic defects among individual seminiferous tubules. These observations provide additional insights into the altered luminal microenvironment and suggest that spermatogenic defects may progress in a tubule-by-tubule manner.

      (2) Clarification of quantitative analysis and methodology (Reviewer #3):

      In response to concerns regarding the basis and methodology of cell quantification, we revised the Methods to provide detailed information on tissue preparation, fixation, orientation of the rete testis–Sertoli valve region, and the criteria used for quantitative analysis of SV-associated Sertoli cells (new Figure S4). We clarified that Sertoli cells were counted within the SV region extending approximately 100 μm from the RT boundary, including Sertoli cells protruding into the RT lumen, based on previously established criteria (Aiyama et al., 2015).

      (3) Clarification of the limitations of expression-level assessment (Reviewer #3):

      In response to concerns regarding the quantitative assessment of SOX17 and other SV-associated molecules, we clarified the technical limitations of selectively isolating the very small SV region and obtaining sufficient material for quantitative molecular analyses such as qPCR at the end of 7th paragraph in the Discussion. We therefore clarified that expression of SV-associated molecules in the present study was primarily evaluated using histological and immunohistochemical approaches and added the relevant text to acknowledge these limitations.

      We also made additional revisions to clarify each mouse Tg line, phenotypic descriptions, standardize gene nomenclature, improve methodological descriptions, and refine the relevant Discussion where appropriate.

      We sincerely appreciate the reviewers’ thoughtful and constructive comments. Their feedback has helped us clarify the scope of our conclusions, strengthen the methodological descriptions, and improve the overall presentation of the study. All changes have been incorporated into the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (i) Although the current paper is not responsible for interpreting the 2022 findings, both datasets show reduced spermatid production accompanied by multinucleated giant germ-cell syncytia. This phenotype has been attributed to backflow of tubular fluid and consequent microenvironmental perturbation. While this is a reasonable hypothesis, it is not entirely consistent with earlier experimental observations. Complete ligation of the efferent ductules reliably produces giant cells, whereas estrogen-receptor knockout, which also causes massive luminal fluid accumulation, does not. In addition, ligation of the testicular artery itself can induce giant-cell formation. Although this may have already been answered in the papers, can you be sure that a direct or indirect effect on the vasculature can be excluded in the Sox17-cKO model?

      We thank the reviewer for this important comment. In our models, SOX17 expression was manipulated specifically in the Sertoli cell lineage, either by SF1-Cre-mediated Sox17 deletion or by ectopic SOX17 expression under the hAMH-promoter. SOX17-expressing vascular endothelial cells were not targeted in either model, making a direct effect of Sox17 manipulation on the testicular vasculature unlikely. Moreover, the partial rescue of the Sox17 cKO phenotype by hAMH-Sox17 supports the interpretation that the phenotype primarily results from altered SOX17 function in Sertoli cells and RT epithelia.

      However, indirect effects on the vascular or interstitial environment by aberrant luminal flow cannot be completely excluded, particularly with the substantial accumulation of sloughed round spermatids (giant cells) within the rete testis. Addressing the potential for an initial luminal flow defect, we newly added histological images of 4-week-old testes (Figure S1), where selective post-meiotic germ cell sloughing occurs despite preserved Sertoli cell process architecture, suggesting an altered adluminal microenvironment that impairs Sertoli–spermatid adhesion. Furthermore, low-magnification images of adult mature Sox17 cKO testes (Figure S1B) display a mosaic pattern of spermatogenic defects across individual tubules. While 3D reconstruction was not conducted, this structural pattern supports the view that spermatogenic failure progresses on a tubule-by-tubule basis, potentially linked to the structural integrity of individual Sertoli valves.

      (ii) A related and important unresolved issue is the total number of Sertoli cells per testis in cKO males. The number of Sertoli cells per tubule cross-section is reported to be equivalent to controls; however, the substantial reduction in testis weight implies a corresponding reduction in tubule length. Under these conditions, maintenance of a normal per-cross-section count would still be compatible with an overall decrease in total Sertoli-cell number. Although it is generally accepted that murine Sertoli cells exit the cell cycle around postnatal day 15, continued growth of the testis may still occur in the Sertoli valve region, where Sertoli cells retain proliferative capacity. Your discussion of possible heterogeneity in the embryonic origin of Sertoli cells near the rete testis is therefore particularly intriguing and commendable. Should this hypothesis be substantiated, it would raise the possibility that Sertoli cells derived from the valve region, especially those that migrate into the seminiferous tubules, are intrinsically less competent to support full spermatogenesis than those of classic gonadal-ridge origin.

      To help readers appreciate the overall severity and topographic distribution of the spermatogenic defect (particularly in tubule segments distant from the rete), inclusion of a low-magnification photomicrograph of a well-fixed (Bouin's) testicular cross-section would be very useful.

      We thank the reviewer for this important comment. We agree that the maintenance of Sertoli cell numbers per seminiferous tubule cross-section does not necessarily indicate preservation of the total Sertoli cell number per testis, particularly given the substantial reduction in testis size and potential reduction in overall seminiferous tubule length. Although total Sertoli cell numbers can theoretically be estimated using stereological approaches, such analyses are technically demanding and beyond the scope of the present study.

      We also appreciate the reviewer’s insightful suggestion regarding potential heterogeneity among Sertoli cell populations. Sertoli cells associated with the Sertoli valve region may have distinct developmental origins or functional properties compared with classical gonadal ridge-derived Sertoli cells, which could potentially influence their capacity to support complete spermatogenesis. Although this hypothesis was not directly tested in this study, we have expanded the Discussion to highlight the developmental and functional heterogeneity of Sertoli cell populations associated with the Sertoli valve as an important topic for future investigation.

      In addition, as requested, we have added a low-magnification image of well-preserved testicular cross-sections in Supplementary Figure S1B to better illustrate the overall severity and topographic distribution of spermatogenic defects throughout the testis.

      Specific Comments:

      (1) Figure 3A and associated fertility/histology data. The results state that epididymal spermatozoa were detected in only 2 of 7 cKO;Tg males at 8 weeks of age, yet Materials and Methods indicate that spermatogenesis was evaluated in only 5 males. a) Were the remaining two males also examined histologically? b) It would be interesting to determine if the severity of pathological changes was the same in regions more distant from the rete testis, or possibly different tubules. See: Nakata H, Wakayama T, Sonomura T, Honma S, Hatta T and Iseki S (2015). "Three-dimensional structure of seminiferous tubules in the adult mouse." J Anat 227(5): 686-694. c) In addition, mating trials were performed with four independent cKO;Tg males, two of which sired offspring. It is unclear whether the testes of these four mating males were included among the five (or seven) animals evaluated for histology, and whether the two fertile males correspond exactly to the two individuals that retained epididymal sperm. Please clarify these relationships explicitly so that readers can correctly interpret the link between histological findings and fertility.

      We thank the reviewer for this important comment. We apologize that the relationship among the groups of animals used for histological analysis, epididymal sperm detection, and fertility assessment was not sufficiently clear in the original manuscript. Because this study focused specifically on the anatomically minute RT–SV region, our sampling strategy had to prioritize the maximal utilization of this limited tissue. In this study, the RT–SV region, the remaining testicular tissue, and the epididymis were processed separately as three tissue blocks for each animal (Figure S4) and were independently evaluated for distinct analysis sets. Briefly, the proximal quarter containing the rete testis and Sertoli valve region was used for SV analysis, whereas the remaining three-quarters of the testis were used for evaluation of spermatogenesis, and the epididymis was analyzed separately for the presence of spermatozoa. Therefore, due to these technical requirements, tissue allocation, and independent analytical evaluation, the numbers of animals used for RT–SV analysis, testicular histology, epididymal sperm detection, and fertility testing were not identical.

      For quantitative histological analyses, we also used virgin males to minimize potential variation associated with mating experience and to allow comparison with age-matched littermate controls. Therefore, these animals were not used for fertility testing. Fertility assessment was performed using an independent cohort of cKO; Tg males that were subjected to long-term mating trials with wild-type females. Thus, fertility outcomes and histological findings were not designed to be directly matched at the individual level.

      In response to the reviewer’s suggestion, we have clarified the selection of experimental animals and the relationship among fertility assessment and histological analyses in the Materials and Methods and added a schematic illustration of the sampling strategy in Figure S4. We also corrected the citation for Nakata H et al., 2015 in the revised manuscript.

      (2) Page 8, line 301 (Sertoli-cell quantification). The description of the counting method-"counted in each ... (~100 μm from the edge of the RT; Fig. 4C)"-is ambiguous.

      (a) Does this mean that cells were counted beginning at the rete boundary and extending radially outward for approximately 100 μm, or is a circumferential sampling area intended? (b) Figure 4C shows a large standard deviation, indicating substantial variability with the current approach. An alternative strategy (for example, counting Sertoli cells within standardized areas or per tubule specifically within the valve region) might reduce variability and improve reproducibility. Regardless of the method ultimately chosen, a more precise, step-by-step description of the quantification protocol is required so that it can be reliably replicated by other laboratories.

      We thank the reviewer for pointing out that the description of the Sertoli cell quantification method was not sufficiently clear. The Sertoli cell quantification was performed using the same criteria as previously described (Aiyama et al., 2015; Uchida et al., 2022), in which SOX9-positive Sertoli cell nuclei within the SV-associated region were counted.

      In the revised manuscript, we have clarified that the SV region was operationally defined as comprising (i) the terminal 100 μm segment of the seminiferous tubule immediately adjacent to the rete testis (RT) and (ii) the protruded SV extending into the RT lumen. Based on the distribution of spermatogonial stem cells, the ~100 μm region extending from the RT boundary along the seminiferous tubule toward the ST side was defined as the SV region (Aiyama et al., 2015). Only sagittal sections showing a continuous RT–SV–ST axis and sectioning the SV approximately through its mid-sagittal plane were included for quantitative analysis.

      Furthermore, to improve reproducibility, we have added a more detailed description of the tissue preparation and quantification procedures in the Materials and Methods and provided a schematic illustration of the quantification strategy in the new Figure S4.

      Reviewer #2 (Recommendations for the authors):

      (1) Phenotypic differences between transgenic lines: the phenotypic differences between the tg26 and tg27 lines are intriguing and warrant further clarification. While tg27 mice exhibit infertility and defective spermatogenesis, tg26 animals remain fertile with SV expansion. Could the authors elaborate on the underlying causes of these differences? In particular, is infertility in tg27 mice due to excessive SOX17 expression impairing Sertoli cell function? A comparison of Sox17 expression levels between tg26 and tg27 lines would be informative. In addition, it would be useful to assess whether acetylated tubulin (Ac-Tub) expression is present in the Sertoli cells of the tg27 mouse testis.

      We thank the reviewer for this highly constructive and insightful comment. We clarified in the revised manuscript that the analysis of the Tg27 mouse was performed using the F0 founder male and added an explanation that only the Tg26 line could be established as its heterogenous SOX17 expression in Sertoli cells did not impair overall fertility. We agree that the phenotypic differences between the Tg26 line and the Tg27 mouse provide important clues regarding the dosage-dependent effects of SOX17 in Sertoli cells. Unfortunately, we were unable to establish a stable, multi-generational transgenic line from this Tg27 founder (F0) male. Consequently, we could not perform detailed molecular or immunohistochemical analyses on this line beyond the initial histological evaluation of the F0 generation presented in Figure 1. For this reason, we cannot provide a quantitative comparison of Sox17 expression levels or evaluate acetylated tubulin (Ac-Tub) expression in Tg27 Sertoli cells.

      To address the reviewer's concern without overstepping the available data, we removed direct quantitative comparisons of Sox17 expression levels between the two lines from the text. Instead, we added a clear description of their contrasting cellular expression patterns - specifically, the mosaic, heterogeneous SOX17 expression in Tg26 Sertoli cells versus the ectopic, uniform SOX17 expression in the infertile #27 F0 male - in the 'Animals' section of Materials and Methods. This mosaic pattern in Tg26 testes suggests the presence of Sertoli cells with low or undetectable SOX17 levels, which may be associated with sustaining overall fertility.

      (2) Mechanism of SOX17 action: although SOX17 is a transcription factor, the author's studies indicate it regulates SV formation via paracrine and/or autocrine signaling. The underlying mechanisms remain unclear. Which downstream factors mediate this effect? The observed upregulation of RSPO1 and WNT4 is suggestive, but more direct evidence would strengthen this conclusion. For example, does SV expansion in tg26 mice depend on the activation of RSPO1/WNT signaling? Additional molecular analyses, such as bulk RNA-seq comparing control and transgenic testes, could help identify pathways regulated by SOX17 and clarify its mode of action.

      We thank the reviewer for this important and insightful suggestion. At present, comprehensive analyses, including scRNA-seq of Sox17 cKO and littermate control testes, have not identified definitive downstream targets of SOX17 (Uchida et al., 2022). As the reviewer rightly points out, the Tg26 mouse model generated in this study represents a valuable tool for investigating SOX17-dependent molecular pathways. To this end, we are currently conducting transcriptomic analyses of Tg26 seminiferous tubules to identify genes altered in SOX17+ Sertoli cells. However, determining whether these candidate genes represent direct transcriptional targets of SOX17 and whether they function specifically in the rete testis-associated region during Sertoli valve formation will require extensive functional and expression studies. Therefore, we feel it would be premature to draw definitive conclusions regarding the underlying molecular mechanisms, including the precise involvement of the RSPO1/WNT signaling pathway, in the present manuscript. Accordingly, rather than overinterpreting the available data, we have revised the Discussion to clarify this limitation (at the end of 7th paragraph in the Discussion). Furthermore, incorporating initial insights from our ongoing Tg26 transcriptomic analyses, we have added a brief discussion, supported by relevant literature, on the possibility that SOX17 may regulate Sertoli valve formation by modulating cell adhesion and extracellular matrix (ECM) organization and altering the responsiveness of SOX17-positive Sertoli cells to morphogenetic signals originating from the rete testis (new 5th paragraph in Discussion).

      (3) Minor comment: Gene nomenclature should be standardized: e.g. line 245, Sox17 and hAMH should be italicized.

      We thank the reviewer for pointing this out. All gene names have been italicized throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      No suggestions except to quantify some of the changes in concentration of agents by PCR rather than eyeball levels with immunocytochemistry. Verify the cell quantification procedure used.

      We thank the reviewer for this comment. The Sertoli valve (SV) is an extremely small transitional structure, with only approximately 20 sites per mouse testis. As a result, selective isolation of the SV region to collect sufficient material for molecular analyses, such as quantitative PCR, remains technically challenging. We have therefore added this limitation to the Discussion.

      Regarding the cell quantification procedure, we have clarified the methodology in the revised Materials and Methods and added a schematic illustration in Figure S4. Specifically, the Sertoli cell number in the SV region was quantified by counting SOX9-positive Sertoli cell nuclei within a standardized SV-associated region in RT–SV–ST sagittal sections.

  2. Sep 2026
    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In this manuscript authors examined the effect of rif1 knockout on replication timing and transcription in early embryos of zebrafish. Contrary to the expectation, genome-wide replication timing domains did not significantly change upon Rif1 knockout, although the replication timing became less dynamic in the mutant, meaning the entire genomes are replicated toward the mid S. In contrast, transcriptional profiles change by rif1 mutation throughout the embryo stage. These effects were more predominantly observed after gastrulation at the early stages of zebrafish development.

      The results presented in this manuscript provide new information on the effects of rif1 mutation on early zebrafish development, although the underlying mechanism has not been explored. The information is useful for researchers in the field of early development, with specific focus on replication and transcription regulation.

      The genome wide analyses of replication timing has been conducted and analyzed properly. The transcriptional analyses are conducted by RNA-seq and SLAM-seq (determining the nascent mRNA), and the results convincingly show the overall transcriptional patterns at different developmental stages.

      This work shows that Rif1 regulates replication timing and transcription in zebrafish embryos, while the extents of the effects vary during the developmental process. Although the data convincingly illustrate the whole picture of Rif1 KO on replication and transcription during zebrafish development, the mechanistic insight is missing. Especially, how Rif1 may or may not coordinately regulate replication and transcription during the zebrafish development has not been addressed.

      We thank the reviewer for recognizing the value of combining genome-wide replication-timing, RNA-seq, and SLAM-seq analyses across zebrafish development. We agree that the original study did not establish a molecular mechanism linking Rif1-dependent transcriptional and replication-timing effects. To address whether these effects are locally coordinated, we added a gene-centred analysis comparing replication-timing values for genes with increased, decreased, or unchanged transcript abundance at Dome (Figure 5--figure supplement 2). Differentially expressed genes did not show a clear enrichment in early- or late-replicating regions, either at Dome or at pre-MBT. These results argue against replication timing state being the primary determinant of the Dome-stage transcriptional changes. We also expanded the Discussion to explain the limitations of the current study and the need for future measurements of origin use, fork progression, chromatin state, and cell-type-specific effects. The new discussion of Nakatani et al. (2025) further places our findings in the context of evidence that Rif1-dependent replication-timing changes can be uncoupled from transcriptional changes.

      Reviewer #2 (Public Review):

      This study by Masser et al. analyzes global replication timing and gene expression in rif-1 null zebrafish. This work is an extension of their previous report on the normal replication timing pattern during wild-type zebrafish development. The major valuable finding here is that Rif1 is not essential for viability in zebrafish, and - counter to expectation from studies in cultured cells and other species - late replication does not strongly depend on Rif1. Instead, the data suggest that Rif1 subtly sharpens replication timing pattern during normal development rather than function generally to delay replication timing. In the absence of Rif1, the normal pattern establishment is somewhat delayed. The authors also document some changes in expression during development with more genes being repressed by Rif1 than activated at some early stages.

      The study and analysis are generally rigorous, and the conclusions are supported by convincing data. The manuscript is well written, though there are aspects of the presentation that could be improved for a broader scientific audience. Given the strong link between replication timing and cell type/development, studying timing in a whole developing organism is important. The experimental approach is technically challenging, particularly the bioinformatic analysis. The scientific advance here is largely confined to documenting the timing of Rif1-affected transcription, the unanticipated effect of the rif1 deletion on replication timing and on sex determination, though the latter is not explored. The work is descriptive and feels like two relatively unconnected studies, transcription and replication plus a small bit of development, and the difference in timing of the transcription phenotypes and replication phenotypes suggests they may be very distinct Rif1 roles. There isn't a lot of new insight into the mechanism of how Rif1 affects either replication timing or gene expression. As such, the overall study is an useful set of findings and detailed data for future work, but it doesn't make a big step forward in understanding the role of Rif1 or the biological processes it affects.

      Weaknesses worth addressing include the following:

      (1) Loss of Rif1 did not affect viability, but it did strongly influence sex determination, resulting in a lower population of females. This effect is the strongest organismal phenotype, but the study provides no explanation for the loss of females from the data gathered here.

      (2) The approach to distinguish nascent zygotically expressed mRNAs from maternal mRNAs is a strength. Are the differentially expressed genes related at all to regions of the genome whose replication timing is most affected? Are any of them related to the sex determination or developmental phenotypes?

      We thank the reviewer for recognizing the rigor of the analyses and the value of studying replication timing in a developing vertebrate. We revised the manuscript extensively to make the experimental logic, zebrafish developmental context, replication-timing analyses, and figure legends more accessible to a broad audience. We also quantified the gastrulation phenotype, showing an approximately one-hour delay in completion of epiboly in maternal-zygotic rif1 mutants rather than a persistent developmental arrest.

      We agree that the mechanism underlying the sex-ratio phenotype remains unresolved. The transcriptomic experiments were performed in whole embryos at stages much earlier than zebrafish sex determination and therefore cannot resolve changes in primordial germ cells or supporting gonadal somatic cells. We have avoided making a mechanistic connection between the early embryonic transcriptional changes and the adult sex-ratio phenotype and identify this as an important area for future study. To address the relationship between transcription and replication timing, we added Figure 5--figure supplement 2. Genes with increased or decreased transcript abundance at Dome were not preferentially associated with early- or late-replicating regions. Together with the distinct developmental timing of the transcriptional and replication-timing phenotypes, this supports the interpretation that Rif1 has separable roles in the two processes rather than a single local mechanism that directly couples them.

      Reviewer #3 (Public Review):

      Using the zebrafish model system, this manuscript assessed the roles of Rif1 protein in replication timing control and transcription during early development, and successfully demonstrated the differential impact of Rif1 protein in replication timing control and transcription. Moreover, the comprehensive assessments of the impacts of mutating Rif1 on animal development (including animal survival and sexual development) were assessed. Although there are works that examined Rif1's implications in replication timing and transcription separately, this work is unique in assessing all these points at once.

      The strength of this manuscript is the genomic analyses of replication timing and transcription being combined in a single model system. Consequently, this manuscript clearly demonstrates the differential impact of Rif1 in these processes during zebrafish development.

      The weakness of this manuscript is, as the authors comment in the Discussion, analyses of replication timing and transcription were performed using bulk embryos. There is a possibility that tissue-specific changes could have been masked. Tissue-specific or single-cell analysis in the future will fill the gap in the knowledge.

      Some of the findings presented in this manuscript are consistent with previous findings using different models such as Drosophila and mice, whereas other findings do not necessarily agree. I hope further studies will reveal more clearly what is common in these systems, and what is different.

      Also, the suggestion that the Rif1 protein may be implicated in a function similar to Fanconi-Anemia genes/proteins is very intriguing.

      Overall, the data presented in this manuscript sufficiently justify the authors' claims. Moreover, this manuscript provides interesting insights into Rif1's function, as well as how development could be controlled.

      We thank the reviewer for highlighting the strength of analyzing replication timing, transcription, and developmental phenotypes in the same vertebrate model. We agree that bulk-embryo measurements may mask tissue- or cell-type-specific effects. We now emphasize this limitation and the need for future tissue-specific or single-cell studies, particularly in the cell populations relevant to sex determination. We also expanded the cross-species context by discussing the recent mouse-embryo study by Nakatani et al. (2025), which supports a conserved role for RIF1 in consolidation of the replication-timing program while also indicating that replication-timing and transcriptional effects can be uncoupled. We agree that defining which Rif1 functions are conserved across zebrafish, mouse, Drosophila, and other systems, including possible relationships to Fanconi-anaemia pathways, will be an important direction for future work.

      Reviewing Editor:

      While the paper was under revision, a relevant paper from the Torres-Padilla lab was published (Nakatani et al., Developmental Cell, 2025). It complements these studies and cites the previous version of this manuscript. I suggest adding a reference in the Discussion to support the conclusions.

      We thank the Reviewing Editor for bringing the recent study by Nakatani et al. to our attention. We have added a standalone paragraph near the end of the Discussion explaining how this work complements our findings, and we have added the complete reference to the bibliography. The new Discussion text reads:

      “A recent study in mouse embryos independently identified RIF1 as a regulator of the developmental consolidation of the RT program. RIF1 depletion produced a less-defined, developmentally immature RT program, while RIF1-dependent RT changes were not correlated with transcriptional changes (Nakatani et al., 2025). Together with our findings in zebrafish, these results support a conserved role for RIF1 in sharpening replication timing during vertebrate development and indicate that its effects on replication timing can be uncoupled from changes in gene expression.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The results presented in this manuscript provide new information on the effects of rif1 mutation on replication and transcription during early zebrafish development, although the underlying mechanism has not been explored. I suggest authors consider conducting the following experiments.

      (1) Does replication timing domains have any role in Rif1-mediated regulation of transcription? It is not clear from the data presented whether transcriptionally affected genes are in the early replicating domains or late replicating domains (that appear after the shield stage). This should be examined.

      We thank the reviewer for this helpful suggestion. To address whether transcriptional effects in rif1 mutants are associated with replication timing, we assigned each gene the nearest smoothed replication timing value and compared replication timing distributions for genes whose transcript levels increased at Dome, decreased at Dome, or were not significantly changed. This analysis is now shown in Figure 5—figure supplement 2. Genes with increased or decreased transcript abundance at Dome did not show a clear enrichment for either early- or late-replicating regions relative to genes with no significant transcript change. This was also true when replication timing was examined at pre-MBT, the stage preceding the major transcriptional changes detected at Dome. These results argue against replication timing state being the primary determinant of the Dome-stage transcriptional changes observed in rif1 mutant embryos. We have revised the Results to describe this analysis and added Figure 5—figure supplement 2.

      (2) It is of interest whether the Rif1-mediated regulation of transcription and replication are mediated by a common mechanism, e.g. through alteration of chromatin structures. Close look at the data in Figure 3D indicates that some genome segments convert replication timing or undergo significant changes of replication timing. It would be informative to know whether these segments (Rif1-regulated replication domains) are associated with the genes whose expression change upon rif1 knockout.

      We thank the reviewer for this insightful suggestion. We agree that an association between Rif1-dependent replication timing changes and Rif1-dependent transcriptional changes would be informative, and we considered this analysis. We attempted to identify Rif1-regulated replication timing domains using the same approach that we previously used to define developmentally regulated timing domains. However, the effect of Rif1 loss differed qualitatively from the developmental timing switches described in our prior work. Rather than producing a limited set of discrete timing-domain transitions, Rif1 loss caused a broad reduction in the dispersion of replication timing values across the genome, consistent with a general flattening of the timing profile. Under these conditions, an unbiased domain-calling approach preferentially identifies genomic regions with the most extreme early or late timing values in wild-type embryos, because these regions show the largest shift toward the mean in rif1 mutants. Thus, the resulting “Rif1-regulated replication domains” largely reflect the strongest wild-type timing domains rather than a discrete set of Rif1-specific regulatory intervals. For this reason, we do not think that assigning differentially expressed genes to such domains would provide a meaningful test of whether Rif1 regulates transcription and replication timing through a common local mechanism. Instead, we have now added a gene-centred analysis comparing replication timing values for genes with increased, decreased, or unchanged transcript abundance at Dome (Figure 5—figure supplement 2), which directly addresses whether transcriptionally affected genes are associated with early- or late-replicating regions.

      (3) Replication is analyzed only by timing analysis. Authors need to analyze frequency of origin firing and replication fork rate by DNA fiber analyses to see whether they are affected by rif1 knockout at various stages of development.

      We agree that measuring origin firing frequency and replication fork rate would provide valuable additional information about how Rif1 loss affects the replication program. However, performing DNA fibre analyses across multiple zebrafish developmental stages and genotypes would require substantial optimization and experimental expansion beyond the scope of the current revision. The current study was designed to measure genome-wide replication timing and transcript abundance across developmental stages, rather than single-molecule replication dynamics. We therefore have not added DNA fibre experiments. Instead, we have revised the Discussion to acknowledge this limitation and to clarify that replication timing reflects the combined effects of origin usage, fork progression, fork directionality, and fork stability. We added the following text to the Discussion:

      “A further limitation of this study is that we concentrated on replication timing without directly measuring other features of the replication program that contribute to this timing. These features include origin usage, replication fork spacing, fork directionality, fork progression, and fork stability. A more comprehensive understanding of how Rif1 loss affects these parameters will be important for defining the relationship between Rif1-dependent changes in replication timing and transcription.”

      Figure 4B, D and F: I did not see the blue lines which represent preMBT in the panels shown.

      We thank the reviewer for identifying this error. The pre-MBT data were not intended to be shown in Figures 4B, 4D, and 4F. We have corrected the figure legend by removing the reference to the blue pre-MBT line.

      Line 270: Figure 4G should be Figure 6G.

      We thank the reviewer for identifying this error. We have corrected the figure reference from Figure 4G to Figure 6G.

      No description of Figure 6E and 6F in the main text.

      We thank the reviewer for noting this omission. We have added text to the Results describing Figures 6E and 6F. The revised text explains that Dome Up-DEGs are normally upregulated from pre-MBT to Shield stages but show earlier upregulation in rif1 mutant embryos, whereas Dome Down-DEGs normally decrease between Dome and Shield stages but show earlier reduction in mutant embryos.

      Reviewer #2 (Recommendations For The Authors):

      (1) This study is an extension of the lab’s previous work which established the wild-type genome-wide replication timing pattern during zebrafish development. The experimental details and analysis are described in the methods, but the general strategy is sometimes treated very cursorily. A non-expert can only understand parts of it by going back to the Seifert study.

      We thank the reviewer for pointing this out. We agree that the replication-timing strategy should be understandable without requiring readers to consult our previous study. We have revised the manuscript to explain the general logic of the assay more clearly. Specifically, we now state that replication timing was inferred from copy-number differences between S-phase and G1-phase genomic DNA: genomic regions that replicate early in S phase are enriched in S-phase DNA relative to G1 DNA, whereas later-replicating regions are less enriched. We also clarified that pre-MBT, dome, and shield embryos were treated as S-phase samples because most cells are in S phase at these stages, whereas nuclei from bud and 24 hpf embryos were sorted by DNA content to isolate G1 and S-phase fractions. These additions make the experimental design and interpretation of the replication-timing profiles clearer in the main text and Methods.

      Figure 2 is meant to document developmental delay in early embryos, but the differences between the single wt and mutant examples in 2D are poorly described and labeled. Most readers will be unfamiliar with the specifics of zebrafish development. There is also no quantification of this developmental phenotype, and that quantification should be included along with better labeling and description of 2D.

      We thank the reviewer for pointing this out. We agree that the developmental delay shown in Figure 2D required clearer explanation and quantification for readers who are less familiar with zebrafish gastrulation. We have revised the Results to explain that epiboly is the process by which the blastoderm and yolk syncytial layer move toward the vegetal pole to envelop the yolk cell, and that zebrafish gastrulation stages are commonly described by the percentage of yolk coverage. We also added quantification of this phenotype. At 10 hpf, most wild-type embryos had completed epiboly, whereas most rif1 mutant embryos had not: 18 of 24 wild-type embryos, but only 2 of 24 mutant embryos, had reached 100% yolk coverage. By 11 hpf, all wild-type and mutant embryos had completed epiboly. These revisions clarify that rif1 mutant embryos show an approximately 1-hour delay in epiboly completion rather than a persistent arrest in gastrulation.

      (3) The presentation could be greatly improved with additional information about the experimental approach and display. As written, the text and figure legends assume readers are intimately familiar with replication timing experiments, zebrafish development, and differential gene expression analysis. Most of the figure legends are not sufficient to understand the figures themselves, and the necessary information is also not always in the results. An example is Figure 3 which is not well described (other than the PCA plots); the term “lag” which is the x-axis in 3C is not defined.

      We thank the reviewer for this helpful comment. We agree that several aspects of the replication-timing analysis required clearer explanation for readers who are less familiar with replication-timing experiments. We have revised the Results to explain the logic of the replication-timing assay more clearly and have added a more detailed description of the autocorrelation analysis in Figure 3C. Specifically, we now explain that autocorrelation measures how similar replication-timing values are across increasing genomic distances along the same chromosome, providing a quantitative readout of the peak-and-valley structure of the timing profile. We also clarified that increasing autocorrelation across hundreds of kilobases reflects the progressive establishment of broader replication-timing domains during development. In addition, we changed the x-axis label in Figure 3C from “lag” to “Genomic distance (Mb).” Together, these changes should make the experimental approach and display easier to understand without requiring readers to consult our previous replication-timing study.

      Figure 4 is generally poorly described and labelled (4B, D, and F graph legends indicate preMBT in the data, but there are no blue lines on the graphs), and Figures 6 and 7 are quite busy.

      We thank the reviewer for pointing this out. We agree that the Figure 4 legend incorrectly described the data shown in panels B, D, and F. The pre-MBT data were not intended to be plotted in these panels, and we have removed the corresponding reference from the figure legend. We recognize that Figures 6 and 7 contain several analyses, but we have retained the current organization because the panels in each figure address a connected set of questions. Figure 6 summarizes how Rif1 loss affects abundance of developmentally regulated transcripts, whereas Figure 7 extends this analysis by directly measuring nascent transcription using SLAM-seq.

      Reviewer #3 (Recommendations For The Authors):

      I do not think any additional experiments are required to justify the authors’ claims. Well done! However, for readers’ benefit, I propose the following changes or adding more explanations:

      (1) Page 2, line 86: I guess “single copy” means “single copy per haploid”. Better to clarify this point.

      We thank the reviewer for this helpful clarification. The reviewer is correct that “single copy” refers to a single copy per haploid genome. We have revised the text to state that the zebrafish genome has a single copy of the rif1 gene per haploid genome.

      (2) Related to the data presented in Figure 2C, do you have an explanation for why sex determination is affected in the heterozygotes, despite the change in Rif1 expression being subtle (Figure 1C)?

      We thank the reviewer for raising this point. We agree that the reduction in whole-embryo rif1 mRNA levels in heterozygotes appears modest relative to the sex-ratio phenotype. At present, we can only speculate about the basis for this difference. One possibility is that whole-embryo mRNA measurements do not accurately reflect Rif1 abundance in the specific cell populations that influence zebrafish sex determination, such as primordial germ cells or their supporting somatic cells. We have therefore avoided making a strong mechanistic conclusion from the heterozygous phenotype.

      (3) Related to the data presented in Figure 2D, did you observe a delay in heterozygotes?

      We thank the reviewer for this question. We have not quantitatively analyzed epiboly progression in heterozygous embryos. However, we did not observe an obvious developmental delay in heterozygotes during early development. The delay shown in Figure 2D was observed in maternal-zygotic rif1 homozygous mutants.

      (4) Figure 3D: it is not easy to distinguish WT and mutant lines, particularly for the Bud stage. Please consider changing the colour schemes or other aspects. For example, making colour lines thinner may help.

      We thank the reviewer for this helpful suggestion. We agree that the wild-type and mutant profiles in Figure 3D, particularly at the bud stage, were difficult to distinguish in the original version. We have revised Figure 3D by reducing the line width of the colored profiles, which improves the contrast between the wild-type and mutant traces.

      (5) Figure 3E: Could you avoid overlapping of WT and mutant plots?

      We thank the reviewer for this suggestion. We considered separating the wild-type and mutant density plots in Figure 3E, but we have retained the overlaid format because the purpose of this panel is to directly compare the distributions of replication timing values between genotypes at each developmental stage. Overlaying the plots makes the reduced dispersion of timing values in the rif1 mutants easier to visualize relative to the corresponding wild-type distribution.

      (6) Figure 4C and 4E: the point legends (WT and mutant) do not match the points used in the graph.

      We thank the reviewer for noting this potential source of confusion. In Figures 4C and 4E, point shape indicates genotype, with open squares representing wild-type samples and open circles representing rif1 mutant samples. Point color indicates developmental stage. We used separate visual encodings for genotype and stage to avoid a large legend containing every genotype-stage combination. To make this clearer, we have revised the figure legend to state explicitly that point shape denotes genotype and point color denotes developmental stage.

      (7) Figure 4D: Very difficult to recognise 24 hr mutant line. Please improve the way there are shown.

      We thank the reviewer for this helpful suggestion. We agree that the 24 hpf mutant profile in Figure 4D was difficult to distinguish in the original version. We have revised the figure by changing the appearance of the mutant lines to make them more visible while preserving the stage color scheme.

      (8) Related to data presented in Figure 4B. Is it possible to show a statistical evaluation of all (or a reasonably large number of samples from) DARs?

      We thank the reviewer for this suggestion. Figure 4A already provides a genome-wide analysis of the DAR set shown by example in Figure 4B. Specifically, Figure 4A plots the change in replication timing from shield to 24 hpf for all 2,498 putative enhancer-associated DARs in both wild-type and rif1 mutant embryos. The strong correlation between wild-type and mutant values indicates that DAR-associated timing changes are largely preserved in rif1 mutants. Because all DARs used for this analysis are included in the scatterplot, we did not add a separate statistical analysis of selected examples from Figure 4B.

      (9) Page 8, line 220: It is unclear what “all” means. Is it all the available replication timing values genome-wide? Please clarify.

      We thank the reviewer for noting this ambiguity. In this sentence, “all” refers to all genome-wide replication timing values calculated from the genomic windows used in our replication timing analysis. We have revised the text to make this clearer.

      (10) Figures 6C and 6D: Colour labels are too dark and it is almost impossible to read texts inside. Please reconsider the colour scheme.

      We thank the reviewer for pointing this out. We agree that the labels in Figures 6C and 6D were difficult to read because of insufficient contrast. We have changed the text colour inside the colored boxes to white to improve legibility.

      (11) Related to overall transcription studies: Is there any sign that Rif1 mutation affects the transcription of genes involved in sex determination?

      We thank the reviewer for raising this interesting question. We have not specifically analyzed whether genes involved in sex determination are differentially expressed in the early embryonic transcriptome data. Because zebrafish sex determination occurs substantially later than the embryonic stages analyzed here, and likely depends on specific cell populations such as primordial germ cells and supporting gonadal somatic cells, we do not think the current whole-embryo RNA-seq data can directly resolve this question. We therefore avoid drawing a mechanistic connection between the early transcriptional changes and the adult sex-ratio phenotype. Determining whether Rif1 mutation affects transcription in the cell populations that regulate zebrafish sex determination will be an important direction for future work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have addressed all the concerns and recommendations by the reviewers, in particular, the requested control experiments using alternative super-resolution microscopy approaches and analysis of the data using Voronoi tessellation in addition to DBSCAN, as requested by reviewer 3. We also provide additional data on tensin3 as suggested by reviewer 2. Finally, to provide a first insight on the role of mechanical forces in the distribution of integrin nanoclusters inside FAs as recommended by reviewer 1, we have performed experiments at different cell seeding times where it is known that FA maturation over time requires mechanical forces.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In recent years, it has become increasingly evident how beautifully intricate IAC are at the nanoscale. Studies like the one presented here that shed light on the precise inner organisation of IAC are thus quite important and relevant in order to obtain a better in-depth understanding of IAC functioning and the contribution of different integrin subtypes to cell adhesive and mechanotransductive processes.

      Interestingly, the authors found a distinct localisation of α5β1 and αvβ3 integrin nanoclusters within focal adhesion of human fibroblasts, with α5β1 integrin nanoclusters being at the periphery of IAC and αvβ3 integrin nanoclusters randomly distributed. Furthermore, a surprisingly high percentage of inactive integrins within IAC and relatively low spatial integrin colocalisation with adaptor proteins has been shown.

      Strengths:

      This is a very thoroughly performed STORM-based assessment of the nanodistribution of α5β1 and αvβ3 nanoclusters within IAC (and outside). The image quality is outstanding, and the authors have meticulously executed the experiments and the image analyses.

      We are grateful to the reviewer for acknowledging the strengths of our study.

      Weaknesses:

      The only weakness is maybe that the manuscript remains descriptive. However, the high quality of the "description" of the nano-organisation of IAC by this scrupulous study is really important to better understand the inner workings of IAC. It provides a very solid foundation to look deeper into the (patho)physiological implications of this organisation, see recommendations (which are rather suggestions in this case).

      We thank the reviewer for their feedback and have addressed their recommendations in our updated manuscript and accompanying reply (see recommendations to the authors). In summary, we have now performed experiments at different seeding times as FA maturation requires mechanical forces, and enquired whether forces might play a role in establishing the spatial distribution of the two different integrins within more mature IACs. The results are now shown as new Fig. 2 and discussed in pages 9 and 10. In addition, in order to get a first insight into the biological implications of our findings we performed dual-colour super-resolution experiments of tensin-3 and α<sub>5</sub>β<sub>1</sub> in FAs, as tensin-3 has been implicated in fibronectin fibrillogenesis. The results are now shown in Fig. S8 and we discuss their potential implications in pages 22 and 23 of the revised manuscript (see more details in the reply to the recommendation to the authors).

      Reviewer #2 (Public review):

      Summary:

      In this study, dual-color super-resolution microscopy analysis was performed to study the co-operation between integrins and focal adhesion proteins in human fibroblast cells. The study focused on two integrins which have been previously found to be mainly responsible for focal adhesions, namely α5β1 and αvβ3.

      Specifically, the study tried to shed light on the nanoclustering of integrins in focal adhesions.

      In the current study, more integrin nanoclusters were observed in focal adhesions compared to other cell-matrix adhesion structures. The study revealed that both α5β1 and αvβ3 form nanoclusters, and those appear segregated from each other. While αvβ3 nanoclusters organize randomly inside focal adhesions regardless of their activation state, α5β1 nanoclusters, and particularly the nanoclusters containing β1-integrin in active conformation, preferentially organized at the edges of focal adhesions. The nanoclusters formed by each integrin were similar in size.

      Cytoplasmic adapter proteins appeared less in nanocluster assemblies, suggesting that integrin nanoclusters are also forming without the studied cytoplasmic adapter proteins (talin, vinculin, paxillin). Active integrins were identified with the help of conformation-specific antibodies, and this enabled us to study the colocalization between integrins and their cytoplasmic adapter proteins. This analysis revealed that activated integrins are strongly engaged with adapter proteins.

      Strengths:

      The study stems from the thorough computational modelling of the nanoclusters, which enables quantification of the behavior of the clusters, including their mesoscale distribution.

      The study strengthens the view that α5β1 and αvβ3 have specific functions in focal adhesions, α5β1 nanoclusters localizing preferentially on focal adhesion edges. The study also revealed that nanoclusters localized at the edges of focal adhesion were enriched for talin and paxillin but not for vinculin.

      Analysis of adaptor protein nanoclusters (paxillin, talin, and vinculin) revealed that all adapter protein nanoclusters studied here close to active β1 nanoclusters are enriched on the focal adhesion edge region, whereas integrin adaptor nanoclusters far from active β1 appear to be more uniformly distributed.

      Importantly, the current study suggests that integrin subtype-specific nanoclusters are not only present at an early stage of adhesion formation, but integrin nanoclusters remain segregated from each other also in mature focal adhesions, maintaining their sizes and number of molecules.

      Interestingly, the study revealed that selected cytoplasmic adaptors (paxillin, talin, and vinculin), also form nanoclusters of similar size and number of single molecule localizations as the integrins, regardless of whether they locate inside or outside focal adhesions. The adapter nanoclusters are enriched in the focal adhesion "belt", colocalizing with the active α5β1 integrin nanoclusters.

      We are grateful to the reviewer for acknowledging the strengths of our study.

      Weaknesses:

      The current study is highly dependent on the antibodies. It is possible that antibodies containing two binding sites for antigen influence the nanoscale organization (and also activation) of the receptors. Control experiments to study the possible contribution of antibodies to the measured outcome should be performed to verify the main findings. One possible approach could be to use fluorescently tagged integrins available. Alternatively, integrins (or adapter proteins) could be tagged with a small ligand and detected using a monovalent binder.

      We understand the concern of the reviewer regarding the use of antibodies for imaging. Nevertheless, we would like to clarify that antibody labelling has always been performed after cell fixation, precluding potential cross-linking artefacts due to protein mobility and avoiding unwanted receptor activation.

      Nevertheless, and although it is highly unlikely to happen in fixed cells, there could be two potential sources of antibody (Ab) labelling artefacts. As the reviewer noted, a primary Ab containing two binding sites could bind to two adjacent proteins (within ~10 nm from each other), potentially underestimating the stoichiometry of the nanoclusters, i.e., number of receptors or proteins per nanocluster. However, in our manuscript we never attempted to provide an estimation of the nanocluster stoichiometry, as it is highly challenging (and prone to artefacts) to provide quantification of the number of proteins using super-resolution-based single-molecule localisation methods which rely on the stochastic blinking of individual fluorophores.

      A second source for potential artefacts comes from the use of the secondary Ab, which (albeit unlikely) could bind to two different primary Abs. To exclude this potential artefact, we performed super-resolution imaging using DNA-PAINT as a different imaging strategy. In this case, the DNA docking site is site-specifically coupled to one camelid single-domain Ab (sdAB), having a much smaller size as compared to a secondary Ab, reducing therefore linkage error and increasing the accessibility of primary Ab-labelled proteins. These new data are included now in Fig. S4. As can be observed, no differences in terms of nanocluster sizes and/or compositions were observed for any of the proteins investigated using DNA-PAINT as compared to our initial STORM data. These control experiments thus rule out any potential artefacts introduced by the secondary Ab (for more details, please see the reply to the recommendations for authors section).

      Only a limited number of integrin adapter proteins were investigated. Given the high number of identified adapter proteins, this is an understandable choice. However, it would be fascinating to understand if the nanoclusters of inactive integrins are dominantly bound with a certain adapter protein, such as tensin.

      We fully agree with the reviewer and have now performed dual-colour super-resolution STED microscopy of α<sub>5</sub>β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, instead of being an integrin inactivator, we found that tensin-3 is also highly enriched at the FA periphery where a large fraction of active β<sub>1</sub> integrins are located, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original data) or to tensin-3 (our new data shown in Fig. S8). We provide more details of our answer in the section of “recommendation to the authors”. Additional experiments, which in our opinion fall outside of the scope of this work, would be necessary to identify other potential integrin inactivator partners, but certainly a topic of future interest to our group.

      Reviewer #3 (Public review):

      Summary:

      In their study, the authors reveal using dual-color super-resolution STORM microscopy modality and immunolabeling in fixed adherent cells, that β1 and β3 integrins as well as adaptors (paxillin, talin and vinculin) are all organized in nanoclusters of similar size (50nm) and molecular density (20 copy number) inside FAs but also outside. Using activityspecific immunolabeling of β1 and β3 integrins, they revealed that active integrin subpopulations were both clustered but in distinct exclusive nano-aggregates in agreement with Spiess et al. (2018). Once more, the "active" integrin nanoclusters displayed similar properties in terms of size and molecular density, suggesting that molecular organization in nanoclusters is an intrinsic property of integrins in plasma membrane multimerizing independently of their location (inside or outside FAs), their level of activation, or their connection to the cytoskeleton. Then the authors followed up by analyzing at the mesoscale how these "universal" nanoclustered adhesive units are distributed spatially. Inspecting the surface density of nanoclusters revealed that the density of integrin nanoclusters in FAs was 5x larger, compared to integrin nanoclusters outside adhesions. Interestingly, whereas the density of total integrin nanoclusters was 2-4x larger than adaptor nanoclusters, the density of "active" integrin nanoclusters stoichiometrically matches that of talin and vinculin nanoclusters, and was slightly outnumbered by paxillin nanoclusters. These findings suggest that inside FAs, among the total number of integrin nanoclusters, the subset of "active" integrin nanoclusters could be engaged with "adaptor" nanoclusters on a 1:1 ratio. Using analysis of the nearest neighbor distance (NND) between distinct integrin clusters and each of the adaptors, the authors report that they found negligible spatial colocalization of integrins with these adaptor proteins and that spatial segregation is essentially determined by the density of nanoclusters within the FAs. As authors reported that α5β1 and αvβ3 do not intermix at the nanoscale, the authors finally highlighted how α5β1 and αvβ3 distinct nanoclusters are differently organized and segregated inside FAs. Adapting the NND analysis in order to inspect how far the nanoclusters are from the edges of FAs they are located in, authors revealed that α5β1 but not αvβ3 integrin nanoclusters are enriched on FA edges and that similar FA edge-enriched distribution for "active" α5β1 and adaptor protein nanoclusters was found for talin and paxillin but not vinculin. The latter results suggest that FA edges could constitute multiprotein hubs for enhanced colocalization and activation for α5β1 integrin nanoclusters and adaptors such as talin and paxillin. Unfortunately NND analysis could not confirm this enhanced colocalization hypothesis.

      General Assessment:

      While the study presents some valuable findings, it reads currently as a compilation of intriguing but preliminary observations derived primarily from a single methodology (dual-color STORM and DBSCAN clustering analysis). As the initial findings often lack confirmation through additional data analysis (such as the NND analysis the authors used), there's a critical necessity to bolster the methodological approach. This should involve replicating the main findings using alternative single-molecule super-resolution techniques (such as quantitative DNA-PAINT) or employing different clustering analytical tools (such as voronoi-tessellation). Furthermore, the manuscript feels incomplete, focusing solely on describing molecular organization without offering substantial insights into how these observations correlate with the regulation, activation, and functionality of integrins at the cellular level.

      We appreciate the comment of the reviewer and have taken their recommendation to heart in order to validate our methodology. In summary, we have now performed extensive DNA-PAINT to replicate most of our initial findings obtained by STORM, as requested by the reviewer. In addition, as a different super-resolution imaging strategy, we have also used STED microscopy to confirm the nanoclustering of integrins and some of the adaptors demonstrating now, by means of three different super-resolution techniques, that both integrins and their adaptors form nanoclusters of similar size and composition, regardless of whether they are inside or outside FAs. We have included these data as Figs. S3 and S4 and discussed the results in pages 8-9 of the main manuscript.

      Regarding the use of an alternative analysis for the data, we have now used the Voronoi tessellation algorithm to re-analyse our STORM data, as requested by the reviewer. The results of the analysis, which render similar sizes and number of localizations as obtained by DBSCAN, are now included in Fig. S5 and mentioned in page 8 of the main manuscript.

      The manuscript presents extensive datasets and utilizes methodologies in which the investigators demonstrate expertise. Nevertheless, there's uncertainty regarding the novelty and broad appeal of the findings. For instance, the observation of integrin nanoclustering has been previously reported in several publications (e.g., Changede et al., Dev Cell 2015; Spiess et al., JCB 2018; Fujiwara et al., JCB 2023). Similarly, the accumulation of specific proteins at the periphery of FAs has been documented elsewhere (e.g., Sun et al., NCB 2016; Stubb et al., NatComm 2019; Nunes-Vicente TCB 2023), as well as the differential dynamic organization of α5β1 and αvβ3 integrins inside FAs (e.g., Rossier et al., NCB 2012). Beyond the universal organization of adhesive proteins, there's a need to identify novel insights that significantly advance the field. One potential avenue could involve pinpointing the molecular determinant controlling the FA edge enrichment of active α5β1 integrins and talin nanoclusters. For instance, could there be an interplay between α5β1 and αvβ3 integrin nanoclusters visible on one's organisation when suppressing the other using deletion (KO) or depletion (SiRNA)? Also, could KANK, which also exhibits enrichment and regulates talin activity (e.g., Sun et al., NCB 2016), play a role in this process? Identifying the molecular players that regulate even partially the mesoscale organization of nanoclusters of proteins would really benefit the breadth of this manuscript.

      We could not agree more with the reviewer and in fact, we are currently investigating the mechanisms that control the enrichment of α<sub>5</sub>β<sub>1</sub> and adaptors at the edges of FAs. However, considering the amount of work needed to determine the spatiotemporal organization of other molecular players using super-resolution imaging constitutes a major tour de force.

      To get a first insight into the process of active α<sub>5</sub>β<sub>1</sub> enrichment at the FA edges, we hypothesised that mechanical forces exerted by the actomyosin machinery could influence the lateral distribution of both integrin subsets (α<sub>5</sub>β<sub>1</sub> and α<sub>v</sub>β<sub>3</sub>) inside FAs. Since FA maturation and strengthening over time requires mechanical forces, we performed experiments at different cell seeding times (90 min, 3 hours and 24 hours) and used STORM imaging to follow the evolution of integrin nanoclustering in time as well as their spatial distributions inside FAs. Interestingly, while nanoclustering of both integrin sub-sets inside FAs is not influenced by seeding times, their lateral distribution was markedly different, with α<sub>5</sub>β<sub>1</sub> nanocluster distribution being already established at earlier seeding times, while α<sub>v</sub>β<sub>3</sub> nanocluster distribution appeared as rather random at earlier seeding times and progressively organized reaching a well-defined lateral spacing at 24 hours of spreading time. These initial data strongly suggest that mechanical forces might play a role in the distinct lateral distribution of both subsets of integrin nanoclusters over time. We have now included these data as new Fig. 2 of the revised manuscript and discuss the results in the associated text (pages 9 and 10). We also discuss potential avenues for further research along the directions suggested by the reviewer.

      In addition, since it has been recently shown that tensin-3 interaction with talin drives the formation of fibronectin-associated fibrillar adhesions (Atherton et al, J Cell Biol 2022) which are enriched in β<sub>1</sub> integrins, we performed dual-colour super-resolution STED microscopy of β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, our initial data show co-enrichment of both tensin-3 and active β<sub>1</sub> nanoclusters at the FA periphery, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original manuscript) or to tensin. Our current working hypothesis is that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery serves to facilitate the translocation of α<sub>5</sub>β<sub>1</sub> integrins from FAs to fibrillar adhesions, most probably in a talin-tensin-dependent manner. We have now included these data as Fig. S8 and accompanying discussion in pages 22 and 23 of the revised manuscript.

      Echoing the previous concern, the manuscript described a novel and rather surprising finding related to molecular clustering of adhesion proteins. Indeed, the fact that nanoclusters exhibit uniform size and molecular density regardless of the protein type, location, or activation level is indeed surprising and raises many questions about the methodology used to assess molecular clustering. I feel that the description and characterization of integrin nanoclusters appear incomplete and need to be expanded by comparing different analytical strategies for protein clustering. Furthermore, a lack of the manuscript in its actual form concerns the quantification of integrin numbers inside the observed nanoclusters. I agree that the path from optical microscopy to protein stoichiometry quantification is hard and full of drawbacks. But the authors do not fully address these issues that are extremely important when discussing protein nanoclustering. This quantitative aspect should be discussed.

      We appreciate the comment of the reviewer as indeed, the existence of “universal” nanoclusters is intriguing. Recently, together with Prof. S. Mayor we have written a short review in Curr. Opin. Cell Biol 2024 proposing that nanoclustering constitutes a molecular-scale organisation principle that governs cellular information flow at the plasma membrane. Our proposal is supported by an extensive number of recent papers showing that most cell membrane receptors and downstream signalling components are organized as pre-assembled nanoclusters. We posit that these nanoclusters serve as modular units whose concatenation in a specific spatiotemporal sequence leads to distinct signalling outputs. Thus, the existence of universal nanoclusters of integrin receptors and adaptors is indeed intriguing but not surprising to us.

      In any case, the concern of the reviewer is well-taken, and as mentioned above, we have used a different algorithm to detect and quantify nanoclustering, obtaining similar values using either Voronoi tessellation or DBSCAN approaches. These data are now included as Fig. S5 in the manuscript.

      Regarding the quantification of integrin numbers inside the observed nanoclusters, we agree with the reviewer that determining protein stoichiometry using single-molecule localization microscopy or STED remains a major technical challenge and is highly prone to artefacts. For this reason, we refrain from making claims about absolute protein numbers per nanocluster. Our relative comparison of nanoclustering among the different proteins investigated is thus exclusively based on the number of single-molecule localisations contained in each nanocluster which is a fair approach since we always use the same reporter fluorophore and maintain similar excitation conditions throughout our experiments. We have now included a few lines on page 9 regarding quantification of the absolute protein numbers inside the nanoclusters and further discuss in the revised manuscript the limitations of single-molecule localisation methods towards the stoichiometry determination of the nanoclusters (see page 20 of the revised manuscript).

      First, it is crucial for the authors to carefully examine and discuss in their manuscript whether there are any potential biases or limitations in the experimental techniques (dual-color STORM) or data analysis methods employed (DBSCAN). Second, the authors did not in the current manuscript, but should provide control samples to demonstrate the sensitivity and dynamic range of their experimental strategy.

      As already mentioned, we have validated the STORM data using both DNA-PAINT and STED and, validated our data analysis obtained with DBSCAN using the Voronoi tessellation algorithm. See Figs. S3, S4 and S5. In terms of sensitivity and dynamic range of our methodology: our set-up has single-molecule detection sensitivity which is demonstrated by the fact that we observe and detect discrete blinking events, a property of single-molecule fluorescence emission and key ingredient to super-resolution single-molecule localisation microscopy. The dynamic range (if we understand correctly the question of the reviewer) is given by the number of frames used to accumulate single-molecule localisations. In our case, we stop acquisition after we deplete most of the single-molecule spots in the imaging view, which typically occurred after 70,000 frames acquisition, as correctly mentioned in the material & methods section.

      In STORM images displayed in Figure S1, the authors highlighted localization clusters detected by DBSCAN as a signature for integrin nanoclusters. But the authors do not discuss the localization spots that were not detected by DBSCAN. Could they be individual integrins? And if so, they should also be considered as useful information? This brings me to another related technical question about how DBSCAN handles the case where fluorescent molecules are blinking. This is important as multiple emissions by a single fluorophore could be detected as a nanocluster of several molecules where it would be an artefact due to the photophysics of the fluorophore. Could the authors comment on these points?

      As mentioned in the original manuscript, between 20-30% of the localizations were not assigned to nanoclusters (Fig. S1H, I) since we imposed a minimum of ten localizations within the radius defined by DBSCAN to be considered as a true nanocluster. This essentially means that regions with less than 10 localizations were not considered in our nanoclustering analysis. However, we cannot be certain as to whether these lower number of localizations correspond to individual integrins, stochastic blinking of the fluorophore or small aggregates containing only a couple of integrins, for the same reasons that we cannot provide quantification of the absolute number of proteins included in each nanocluster: stoichiometry determination by means of single-molecule super-resolution methods is highly prone to artefacts.

      Regarding the concern of how DBSCAN handles fluorophore blinking, the reviewer is completely right as the photophysics of the fluorophore can influence the analysis of the data and the identification of true nanoclusters. To decouple the photophysics of the fluorophore we first assess the number of blinking events within the DBSCAN radius, i.e., number of localizations corresponding to individual antibodies sparsely distributed on the glass surface. In our case, the median values for the two activator-reporter pairs corresponded to 5 localizations for Alexa 405-Alexa647-conjugated Abs and 3 localizations for Cy3-Alexa 647-conjugated Abs (see Fig. 1E). Yet, despite these median values, the number of localizations per individual Ab naturally shows a distribution. Thus, to avoid any overestimation in the degree of nanoclustering, we impose an additional constrain to our analysis and consider true nanoclusters only those ones containing at least 10 localizations. We have now significantly extended the explanation in the main text (see page 6) as well as materials & methods so that it becomes clearer to the reader.

      Also, using isolated and stochastically physisorbed fluorophores (Ab coupled with activator /reporter pairs used in this study) on glass helped define the signature in STORM of a single isolated molecule. To obtain the signature of clustered fluorophores, the authors could use anti-donkey antibodies to cross-link those STORM-specifically labeled Ab as a means to artificially obtain clustered fluorophores. Ultimately, to avoid the bias effect of the glass surfaces on the photophysics of fluorophores and be in the same imaging conditions as for the described nanoclusters, the authors should use model systems composed of multimers of GFP vs. single GFP, immunolabeled with a GFP-binding monoclonal antibody. This will permit evaluation of the cluster signature obtained with DBSCAN analysis of STORM data for single vs. multimers of known stoichiometry. This would constitute an undisputable molecular stoichiometry ruler.

      We appreciate the suggestions of the reviewer. Regarding the potential bias effect of the glass surface on the photophysics of the fluorophores we would like to clarify that the “calibration” for the number of blinking events per individual Ab on glass were performed on the same sample containing the cells that we image, so that we maintain exactly the same experimental and imaging conditions avoiding any potential artefacts. To our understanding this approach is more accurate than performing the calibration on glass substrates and then moving to samples containing the cells. This information is now contained in page 6 of the revised manuscript and in the materials and method section. Once the number of blinking events from individual Abs on glass within the DBSCAN radius are determined, one can then determine the number of localizations within the same DBSCAN radius on other parts of the sample. More localizations within the same DBSCAN radius basically means more molecules, and thus nanoclusters. This approach has been extensively used by other experts in the field as we properly acknowledge in our manuscript (Pageon et al, Mol. Cell. Biol 2016; Spiess et al. J. Cell Biol 2022).

      Using anti-donkey antibodies to cross-link those STORM-specifically labelled Ab in order to artificially obtain clustered fluorophores, as suggested by the reviewer, is indeed a sound approach to retrieve signatures of clustering. Nevertheless, we have preferred not to use this approach because those artificially induced clusters would have very little resemblance to the real nanoclusters and would only allow us to validate the performance of DBSCAN for cluster recognition. As mentioned above, DBSCAN is a well-established algorithm and used by many different experts in the field and thus can be trusted by the community. Instead, and following the recommendation of the reviewer, we now provide results using an alternative cluster analysis algorithm (Voronoi tessellation) reaching similar conclusions regarding the existence of integrin and adaptor nanoclustering inside FAs.

      Finally, the suggestion of using monomeric vs multimeric GFPs to determine the stoichiometry of the nanoclusters is highly appreciated. Indeed, we have used this approach in the past to identify nanoclustering of the chemokine receptor CXCR4 in living T cells (Mol. Cell 2018 and PNAS 2022). However, these experiments are best performed at sub-labelling conditions, which inherently underestimate the degree of nanoclustering. Combining GFPs with PALM to enable super-resolution is another approach but also subject to artefacts regarding the photo-conversion efficiency of GFPs as we reported earlier (Nature Methods 2017) and leading to underestimation of nanocluster stoichiometry.

      In summary, providing nanocluster stoichiometry from single-molecule localisation images remains a major technical challenge and is highly sensitive to methodological assumptions. We have therefore focused here on providing robust evidence for the existence of integrin and adaptor nanoclustering, using three different superresolution approaches and two independent analytical methods for cluster determination.

      Due to the surprising finding of the nanoclusters' "universality", it is imperative for the authors to validate the findings through complementary methodologies and analytical tools. This should involve replication of results using alternative super-resolution techniques (quantitative DNA-PAINT) and exploring different clustering algorithms (VoronoïTesselation) to ensure the robustness and reliability of the observations.

      As already mentioned, we have now performed extensive DNA-PAINT to replicate most of our initial findings obtained by STORM, as requested by the reviewer. In addition, as a different super-resolution imaging strategy, we have also used STED microscopy to confirm the nanoclustering of integrins and some of the adaptors demonstrating now, by means of three different super-resolution techniques, that both integrins and their adaptors form nanoclusters of similar size and composition, regardless of whether they are inside or outside FAs. We have included these data as Figs. S3 and S4 and discussed the results in pages 8-9 of the main manuscript.

      Regarding the use of an alternative analysis for the data, we have now used the Voronoi tessellation algorithm to re-analyse our STORM data, as requested by the reviewer. The results of the analysis, which render similar sizes and number of localizations as obtained by DBSCAN, are now included in Fig. S5 and mentioned in page 8 of the main manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This work already, as is, provides significant and novel information on IAC.

      The unexpectedly low spatial colocalisation of integrins with adaptor proteins might indeed be caused by the potentially quite long extension of talin upon force exposure and the ample zones of activity of IAC proteins, imaging the involved proteins in scale, as can be seen in Barnett and Goult (2022, doi: 10.3389/fncel.2022.1014629)? In super-resolution microscopy, this spatial separation might, in fact, become apparent. It would be interesting to see whether lowering the actomyosin contraction by different concentrations of blebbistatin lowers this separation. In general, it would also be interesting to understand whether lowering the forces can disrupt the nano-organisation and the strong separation of the two analysed integrin subtypes. It is true that nascent adhesion formation is force-independent, but maybe the forces play a role in establishing the particular integrin subtype nano-organisation within more mature IAC. I am also aware that a lot of work has already gone into the conclusion of this project.

      We thank the reviewer for these thoughtful comments and suggestions. Our most recent preliminary data (not yet included in this manuscript) indeed indicate that the physical separation of integrin nanoclusters (and adaptors) inside focal adhesions (FAs) is force-dependent. We are currently reproducing these experiments using lipid bilayers of varying viscosities and controlled ligand density to explore how ligand mobility (i.e., equivalent to force exerted from the extracellular side) controls the degree of IAC nanoclustering and their spatial segregation in FAs. This approach is more amenable to super-resolution microscopy as lipid bilayers are quite thin and optically transparent, yet the experiments are still challenging, time-consuming, and thus ongoing.

      To obtain a first hint as to whether forces might play a role in establishing the spatial distribution of the two different integrins within more mature IACs as the reviewer suggests, we have performed experiments at different seeding times (90 min, 3 hours and 24 hours). Our results show that even at earlier times (90 min), when a lower number of mature FAs are established, nanoclustering of integrins and main adaptors are similar to 24 hours. In contrast, and as suggested by the reviewer, the spatial distribution of the different subsets of integrin nanoclusters inside FAs is markedly different as a function of seeding time, with α<sub>5</sub>β<sub>1</sub> nanocluster distribution being already established at 90 min, while α<sub>v</sub>β<sub>3</sub> nanocluster distribution appears rather random at earlier seeding times and progressively organizes reaching a well-defined lateral spacing at 24 hours of spreading time. As FA strengthening over time requires mechanical forces, and α<sub>v</sub>β<sub>3</sub> is preferentially involved in FA strengthening (Roca-Cusachs et al PNAS 2009), these data strongly suggest that forces play a differential role in the lateral distribution of both integrin nanoclusters over time. We have now included these data as a new Fig. 2 in the revised manuscript and discuss the results in the associated text (pages 9 and 10). We also mention in the discussion additional experiments, as suggested by the reviewer, to further substantiate this hypothesis.

      Considering the high quality of the work and the new insight about the inner organisation of IAC, maybe the summary Figure 5 should be elaborated a bit, taking into account e.g. different lengths of extended talin proteins and also the various positions of vinculins on talin proteins (depending on opened cryptic binding sites), as well as the possibility that various actin filaments might be associated with single talins. What I mean is, the authors impressively demonstrate the complexity of IAC nano-organisation, which should be paid more tribute in the concluding figure. The quality of the figure should be adapted to the quality of the work.

      We have adapted Figure 5 (now Figure 6) as suggested by the reviewer.

      I would be curious to hear a bit more about the further speculations of the authors in the discussion, e.g., about why the integrin subunits are organised in this way. Why might the α<sub>5</sub>β<sub>1</sub> be preferentially located in the periphery? What is the potential physiological relevance of this organisation? Is this organisation different in other cell types (have the authors looked at other cells)? Is the organisation lost in pathophysiological situations, such as cancer?

      Although we do not know yet what drives the preferential location of α<sub>5</sub>β<sub>1</sub> nanoclusters to the FA periphery, it is known that Kank2 also exhibits enrichment at the FA periphery, regulates talin activity and it is involved in the formation of α<sub>5</sub>β<sub>1</sub>-enriched fibrillar adhesions (Sun et al, Nature Cell Biol 2016). Thus, it is highly probable that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery is a necessary step for their translocation from mature FAs to fibrillar adhesions to then assemble fibronectin into the fibrillar networks as found and needed in connective tissues. Consistent with this idea, we have observed similar α<sub>5</sub>β<sub>1</sub> distribution on other fibroblast cell lines (MEFS), which are the primary cells that produce fibrillar adhesions. Thus, α<sub>5</sub>β<sub>1</sub> nanocluster distribution inside FAs might be physiologically important for the process of fibronectin fibrillogenesis.

      Since it has been documented that tensin is important for fibronectin fibrillogenesis (Pankov et al J Cell Biol 2000) and more recently, it has been shown that tensin-3 interaction with talin drives the formation of fibronectin-associated fibrillar adhesions (Atherton et al, J Cell Biol 2022), we thought to investigate the spatial distribution of tensin-3 and its relationship with α<sub>5</sub>β<sub>1</sub> inside FAs by means of dual colour super-resolution STED microscopy. Interestingly, our initial data on HFF cells seeded for 24 hours show both enrichment of tensin-3 and α<sub>5</sub>β<sub>1</sub> nanoclusters at the edges of mature FAs, supporting our working hypothesis that α<sub>5</sub>β<sub>1</sub> enrichment at the FA periphery serves to translocate α<sub>5</sub>β<sub>1</sub> integrins from FAs to fibrillar adhesions, probably in a talin-tensin-dependent manner. While these initial data are quite exciting, many more experiments that include simultaneous super-resolution mapping of α<sub>5</sub>β<sub>1</sub>, talin and tensin in mature FAs are required to fully validate our hypothesis. Yet, because of their relevance we consider it appropriate to include these data as Fig. S8 and discussing their potential implications in pages 22 and 23 of the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform control experiments to confirm that the nanocluster size/composition is not affected by the antibodies used.

      As explained in the response to the public reviews, antibody labelling has always been performed after cell fixation, precluding potential cross-linking artefacts due to protein mobility and avoiding unwanted receptor activation. In addition, we have performed super-resolution imaging using DNA-PAINT as a different imaging strategy. In this case, the DNA docking site is site-specifically coupled to one camelid single-domain Ab (sdAB), having a much smaller size as compared to a secondary Ab, reducing therefore linkage error and increasing the accessibility of primary Ab-labelled proteins. As can be observed in new Fig S4, no differences in terms of nanocluster sizes and/or compositions were observed for any of the proteins investigated using DNA-PAINT as compared to our initial STORM data. These control experiments thus rule out any potential artefacts introduced by the secondary Ab. Finally, we would like to highlight that our results on the nanoclustering of integrins in terms of their size and number of localizations is consistent with previous results obtained by other groups around the world using similar labelling protocols as us (Spies et al, J. Cell Biol 2022), or relying on halo-tag strategies, as suggested by the reviewer (see Fujiwara et al, J. Cell Biol. 2023). The consistency of these results amongst different groups gives us further confidence that the nanocluster size/composition are not affected by the antibodies used.

      (2) Extend the study by inspecting a set of integrin adapter proteins for their association with inactive integrins, focusing on adapters associated with the maintenance of the inactive state. Possible candidates would be tensin and filamin, for example.

      We thank the reviewer for the suggestion and have now performed dual-colour super-resolution STED microscopy of α<sub>5</sub>β<sub>1</sub> and tensin-3 on HFF cells seeded for 24 hours. Interestingly, instead of being an integrin inactivator, we found that tensin-3 is also highly enriched at the FA periphery where a large fraction of active β<sub>1</sub> integrins are located, suggesting that at these particular regions, active β<sub>1</sub> could be either engaged to talin (as shown in our original data) or to tensin-3 (our new data shown in Fig. S8). These results might be surprising at first, since tensin competes with talin for the same binding site to the cytoplasmic β-tail of integrins, and thus believed to act as integrin inactivator, as the reviewer indicates. Nevertheless, recent data has shown that tensin is capable to activate integrins (in particular if β<sub>1</sub> is phosphorylated) by interacting with the actin cytoskeleton, providing mechanical coupling for integrin activation (Georgiadou & Ivaska, Trends Cell Biol. 2017). We have now included these new data as Fig. S8 in the revised manuscript. Additional experiments, which in our opinion fall outside of the scope of this work, would be necessary to identify other potential integrin inactivator partners, but certainly a topic of future interest to our group.

      (3) While the methods are described in sufficient detail, it is important to ask if the findings are based on sufficient data. Table S5 provides detailed information about the number of samples studied, and it appears that only small numbers of samples were investigated for certain protein pairs. This should be discussed, and perhaps more data should be obtained to strengthen the data.

      We have now performed additional experiments using DNA-PAINT as alternative super-resolution imaging technique (as also requested by reviewer 3) which adds additional data to the whole manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 M<sup>pro</sup> enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the M<sup>pro</sup> monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how M<sup>pro</sup> inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      The major concern is the absence of mutagenesis data to support the proposed inhibitory mechanisms, particularly regarding the role of the inhibitor binding site.

      We thank the reviewer for the recognition and comments. We agree that mutagenesis is critical for validating the proposed role of C300. We therefore generated the C300S and C300F mutants and characterized their oligomeric states and proteolytic activities. C300S was designed to remove the reactive thiol group while minimally affecting M<sup>pro</sup> structure and dimerization. C300F was introduced to mimic the steric perturbation associated with C300 modification and assess its impact on M<sup>pro</sup> dimerization. Native PAGE showed that WT and C300S M<sup>pro</sup> predominantly formed dimers, whereas C300F was mainly monomeric. Consistently, C300S retained approximately 70% of WT activity, whereas C300F retained only approximately 10%. Because C300F itself strongly disrupted dimerization, C300S was used as the principal mutant to evaluate the specific contribution of the C300 thiol to ebselen action. Native MS showed that ebselen could still bind to both monomeric and dimeric C300S M<sup>pro</sup> but did not markedly shift its monomer-dimer equilibrium toward the monomeric state. In parallel, ebselen reduced WT activity to approximately 53% of the untreated control, whereas C300S retained approximately 78% activity at the same 1:3 M<sup>pro</sup>-to-ebselen molar ratio. These results provide experimental support for the contribution of C300 to ebselen-induced dimer destabilization and functional inhibition, while the residual binding and inhibition observed for C300S suggest the involvement of additional C300-independent interactions. The corresponding revisions have been made to Methods (Lines 627–648), and Results (Lines 397–453) of the manuscript, together with the newly added figures (Figures S14–S16).

      Reviewer #2 (Public review):

      Summary:

      This is a mechanistic study that provides new insights into the inhibition of SARS-CoV-2 M<sup>pro</sup>.

      Strengths

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      We thank the reviewer for the positive comments and recognition of our study.

      Weaknesses:

      The primary weaknesses relate to linking the biophysical observations more directly to functional enzymatic outcomes and providing more quantitative rigor in some analyses. While the study is overall strong, addressing its weaknesses and limitations would elevate the impact and translational relevance of the current manuscript.

      We thank the reviewer for these comments, which have helped to iM<sup>pro</sup>ve the quality and impact of our manuscript.

      (1) Correlation with Functional Activity:

      The most significant gap is the lack of direct enzymatic activity assays under the exact conditions used for MS and HDX. While EC50 values are listed from literature, demonstrating how the observed dimer stabilization (by peptidomimetics) or dimer disruption (by ebselen) directly correlates with inhibition of proteolytic activity in the same experimental setup would solidify the functional relevance of the biophysical observations. For instance, does the fraction of monomer measured by native MS quantitatively predict the loss of activity? Also, the single inhibitor concentration used in each MS experiment needs to be specified in the main text and legends. A discussion on whether the inhibitor concentrations required to observe these dimerization effects (in native MS) or structural dynamics (in HDX-MS) align with EC50 values would be helpful for contextualizing the findings.

      We thank the reviewer for these important points. To link the biophysical observations more directly to function, we compared the oligomeric states and proteolytic activities of WT, C300S, and C300F M<sup>pro</sup>. C300F was predominantly monomeric and retained only approximately 10% of WT activity, whereas C300S remained predominantly dimeric and retained approximately 70% activity. We further evaluated ebselen inhibition using a matched 1:3 M<sup>pro</sup>-to-ebselen molar ratio. Ebselen reduced WT activity to approximately 53% of its untreated control but reduced C300S activity only to approximately 78%, demonstrating that removal of the C300 thiol significantly attenuated the functional effect of ebselen. These data support a relationship between C300-dependent dimer destabilization and reduced proteolytic activity. The Methods (Lines 627–648), and Results (Lines 397–453) have been revised accordingly, with Figures S14–S16 newly added, in the revised manuscript. We did not expect a linear relationship between the monomer fraction measured by native MS and enzymatic activity loss, because ebselen can modify multiple cysteine residues, and individual modification events may have distinct effects on M<sup>pro</sup> dimerization and catalytic function. The concentrations and molar ratios used in the native MS, HDX-MS, and activity assays have now been stated in the figure legends. The ebselen concentrations used for native MS and HDX-MS were optimized for biophysical characterization and comparison, and therefore, these concentrations might not be directly related to their IC<sub>50</sub> or EC<sub>50</sub> values. In these experiments, ebselen was applied at a 3-fold molar excess relative to M<sup>pro</sup>, consistent with the enzymatic assay. The observed dimer disruption and conformational changes were consistent with functional inhibition, supporting their mechanistic relevance.

      (2) For the two Cys residues found to be targeted by ebselen, what are their respective modification stoichiometry related to the ebselen concentration? Especially for the covalent binding site C300, which is proposed in this study to represent a novel allosteric inhibition mechanism of ebselen, more direct experimental evidence is needed to support this major hypothesis. Does mutation or modification of C300 affect the M<sup>pro</sup> dimerization/monomer equilibrium and alter the enzymatic activity? If ebselen acts as a covalent inhibitor linked to multiple Cys, why is its activity only in the μM range?

      We thank the reviewer for the insightful comments. Our LC-MS/MS data identified C44 and C300 as ebselen-modified residues, but they do not permit reliable site-resolved occupancy measurements because modified and unmodified peptides can differ in digestion efficiency and MS response. We have therefore clarified that these data provide qualitative site identification rather than absolute modification stoichiometry. To obtain direct functional evidence for C300, we generated C300S and C300F mutants. C300S preserved dimer formation and substantial activity, whereas C300F was mainly monomeric and showed severe activity loss. Importantly, although ebselen-bound C300S species were still detected by native MS, ebselen did not markedly redistribute C300S toward the monomeric state, and its inhibition was reduced from approximately 47% for WT to approximately 22% for C300S. These results indicate that C300 is an important contributor to ebselen-induced dimer disruption, while residual binding and inhibition indicate additional reactive sites. Corresponding revisions have been made to the (Lines 627–648), and Results (Lines 397–453) of the manuscript, together with the newly added figures (Figures S14–S16). The moderate micromolar potency of ebselen is consistent with its heterogeneous, multi-site covalent reactivity: modification occupancy and functional consequence are site-dependent, and not every adduct produces complete inhibition.

      (3) For the allosteric inhibitor pelitinib with low-μM activity, no significant differences in deuterium uptake of M<sup>pro</sup> were observed. In terms of the binding affinity, what is the difference between pelitinib and ebselen? Some explanations could be provided about the different HDX-MS results between the two non-peptidomimetic inhibitors with similar activities.

      We agree with the reviewer that the absence of significant HDX changes for pelitinib requires clarification. Different from ebselen that forms covalent bond with multiple cysteine residues of M<sup>pro</sup>, which could lead to sustained conformational changes that are more readily detected by HDX-MS, pelitinib non-covalently binds M<sup>pro</sup> and might not induce significant perturbations in backbone dynamics that are detectable at the peptide level by HDX-MS. These points have been integrated into the revised manuscript (Lines 333-337).

      (4) Native MS Quantification: 

      The analysis of monomer-dimer ratios from native MS spectra appears qualitative or semi-quantitative. A more rigorous and quantified analysis of the percentage of dimer/monomer species under each condition, with statistical replicates, would strengthen the equilibrium shift claims. For native MS analysis of each inhibitor, the representative spectrum can be shown in the main figure together with quantified dimer/monomer fractions from replicates to show significance by statistical tests.

      We thank the reviewer for the suggestion. We have performed a quantitative analysis of the monomer-dimer equilibrium based on triplicate native MS measurements for each condition. Representative spectra, quantified monomer/dimer ratios, and statistical analyses have been added to Figures 1 and S3. The quantitative results have also been described in the Results section (Lines 158–161, 165-168, 172-174, 177-179, 199-200).

      (5) Changes of HDX rates in certain regions seem very subtle. For example, as it states 'residues 296-304 in the C-terminal region of M<sup>pro</sup> were more flexible upon ebselen binding (Figure 4c)', the difference is barely observable. The percentage of HDX rate changes between two conditions (with p values) can be specified in the text for each fragment discussed, and any change below 5% or 10% is negligible.

      We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296–306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) The study lacks validation through inhibitor binding site mutagenesis assays, especially peptidomimetic inhibitor PF-07321332 and ebselen, which would strengthen the mechanistic conclusions.

      We appreciate this suggestion. For PF-07321332, the inhibitor forms a covalent interaction with the catalytic residue C145 and inhibits M<sup>pro</sup> activity through a distinct mechanism. Previous studies have shown that mutation of C145, such as C145A, completely abolishes M<sup>pro</sup> catalytic activity (Bhandari, D, et al. Communications Biology 2025, 8, 1061), making it difficult to directly evaluate the contribution of this residue to inhibitor-induced inhibition using enzymatic assays alone. This limitation and the relevant literature have now been discussed in the revised manuscript (Lines 272–277). Therefore, we focused on C300-dependent regulation of ebselen, which represents a distinct inhibitory mechanism involving modulation of M<sup>pro</sup> structural dynamics and dimer stability.

      To validate the role of C300 in ebselen-mediated regulation of M<sup>pro</sup>, we generated C300S and C300F mutants and performed additional biochemical and structural characterization. The enzymatic assay showed that the C300F mutation significantly affected M<sup>pro</sup> activity, and the inhibitory effect of ebselen on C300S M<sup>pro</sup> was markedly reduced compared with WT M<sup>pro</sup>. Furthermore, native MS analysis demonstrated that ebselen could still bind to C300S M<sup>pro</sup> but failed to induce a significant shift in the monomer-dimer equilibrium observed for WT M<sup>pro</sup>. These results indicate that C300 is not the only site involved in ebselen binding but is critical for mediating ebselen-induced structural perturbation and dimer destabilization. The manuscript has been revised accordingly for the (Lines 627–648), and Results (Lines 397–453), with new figures (Figures S14–S16) included, further supporting the functional contribution of C300 in ebselen-mediated M<sup>pro</sup> regulation.

      (2) MR6-31-2 is an ebselen derivative and exhibits a lower EC50 (1.78 μM) compared to ebselen (4.67 μM). It would be helpful to discuss why their activities differ, probably based on the assay conditions or binding behavior.

      We agree with the reviewer that the difference in antiviral activity between MR6-31-2 and ebselen requires further clarification. The lower EC<sub>50</sub> of MR6-31-2 may result from iM<sup>pro</sup>ved cellular properties, including compound stability, permeability, intracellular exposure, and potentially altered interactions with M<sup>pro</sup> and/or iM<sup>pro</sup>ved cellular properties. Although MR6-31-2 shares the ebselen scaffold, the modified chemical structure may affect its binding behavior and biological activity. However, EC<sub>50</sub> values obtained from cellular assays cannot directly reflect the biochemical inhibition potency against purified M<sup>pro</sup>. These points have been integrated into the revised Introduction (Lines 98–101).

      (3) In Figures 2, S1, S2, S4, S6, and S11, adding the drug name under each panel would make the data much clearer for readers.

      The corresponding drug names have been added to panels to iM<sup>pro</sup>ve figure clarity.

      Minor points:

      (1) Line 62-63 refers to the "long linker loop," while Figure 1a labels it as the "long loop linker." Please keep this consistent.

      The terminology has been unified as “long loop linker” throughout the manuscript.

      (2) Table 1 should be cited at line 80, and PDB code 7BAK should be included in Table 1.

      PDB code 7BAK has been included in Table 1, and Table 1 has been cited in the context, as suggested.

      (3) Figure 1a should include the corresponding PDB code in the figure legend.

      The corresponding PDB code has been added to the Figure 1a legend, as suggested.

      (4) It would be helpful to indicate in Figure 1a that the upper structure represents the dimer and the lower structure represents the monomer.

      The upper and lower structures in Figure 1a have been indicated as dimeric and monomeric M<sup>pro</sup>, respectively, as suggested.

      (5) In the Figure S1 legend, it should mention that some inhibitor structures (like ebselen and MR6-31-2) are not fully resolved. Also, the Se atom in ebselen should be shown in Figure S1f (PDB: 7BAK).

      The Figure S1 legend has been revised to indicate that some inhibitor structures, including ebselen and MR6-31-2, are partially unresolved, and the selenium atom of ebselen has also been shown in Figure S1f, as suggested.

      (6) Pelitinib is an allosteric, non-covalently binding inhibitor. However, in Figure S3, the native MS profile shows dimer species (13+ to 15+) compared with unbound M<sup>pro</sup> (14+ to 17+). Please clarify this difference.

      We thank the reviewer for raising this good point. Protein charge-state distributions can be influenced by solution-phase conformation, conformational flexibility, solvent properties, and electrospray droplet charging (Susa AC, et al. J Am Soc Mass Spectrom 2017, 28, 332-340). The observed shift in charge state distribution in native MS might suggest that the addition of pelitinib caused changes in the protein conformation, solvent property and electrospray droplet charging. The relevant literature and discussion have been added in the revised manuscript (Lines 200–204).

      (7) Line 172: "S1are" should be corrected to "S1 are."

      Corrected.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Comments on revised version:

      I thank the authors for their extensive revisions and detailed responses to the reviewers' comments. The manuscript has been substantially improved, and most of the major concerns raised in the initial review have been adequately addressed. In particular, the additional analyses of PKD2L1 channel activity, the incorporation of physiologically relevant pH conditions, the clarification of ASIC involvement, and the expanded Discussion have significantly strengthened the study.

      Major scientific concerns largely addressed:

      Quantification of PKD2L1 channel activity

      The authors appropriately addressed my previous concerns regarding the use of Po as the sole measure of channel activity. The inclusion of additional parameters such as apparent Po, open time, nmax, holding current, and membrane charge provides a more robust assessment of PKD2L1 activity and substantially strengthens the conclusions.

      Physiological relevance of pH modulation

      The inclusion of experiments at pH 6.5 and the additional analyses of holding current and resting membrane potential are valuable additions. These experiments considerably improve the physiological relevance of the study.

      ASIC contribution

      The additional pharmacological experiments using ASIC blockers are helpful and support the conclusion that the photolysis-evoked response in the apical process is predominantly mediated by PKD2L1 channels.

      Functional implications

      The expanded Discussion regarding Ca2+-dependent signaling, neurosecretion, and the potential physiological roles of CSFcNs considerably improves the manuscript.

      Remaining concerns:

      Continued overstatement regarding "exclusive" localization and function:

      Although the authors softened some statements in the revised manuscript, the term "exclusive" remains in several key locations, including the title.

      For example:

      "PKD2L1 channels segregated to the apical compartment are the exclusive dual-mode pH sensor..."

      The data clearly demonstrate strong enrichment of functional PKD2L1 channels in the apical process. However, the available evidence does not fully justify the term "exclusive," particularly because:

      - PKD2L1 immunoreactivity is still detectable outside the apical process.

      - ASIC-mediated responses are present in CSFcNs.

      - The authors themselves use more appropriate terminology such as "predominantly located" in the Discussion.

      Therefore, I recommend replacing "exclusive" with more conservative terminology such as:

      - predominant

      - predominantly localized

      - enriche

      - functionally segregated

      throughout the manuscript, including the title, Abstract, Introduction, Results, and Discussion.

      We agree with the reviewer that the world “exclusive” is misleading and should be replaced. Following the reviewer’s suggestions, we have deleted the word “exclusive from the title, which now reads: “PKD2L1 channels segregated to the apical compartment are the functional dual-mode pH sensors in cerebrospinal fluid-contacting neurons.”

      In addition, the word “exclusive” has been changed with more conservative terminology in other parts of the text: lines 80, 420, 466 and 551.

      Use of the term "tonic current"

      The manuscript continues to use the term "PKD2L1 tonic current."

      While the dibucaine-sensitive holding current is clearly present, the precise mechanism generating this current remains uncertain. Indeed, the authors themselves acknowledge in the Discussion that:

      - an alternative conducting state may exist, or

      - unresolved brief channel openings may account for the current.

      Therefore, the data support the existence of a sustained PKD2L1-associated current, but do not yet definitively establish a distinct tonic gating mode of the channel.

      I therefore recommend replacing:

      "tonic current" with a more neutral expression such as:

      - sustained current

      - PKD2L1-associated holding current

      - sustained PKD2L1-mediated current throughout the manuscript.

      Continued use of "off-current" and "off-response":

      The revised manuscript has improved considerably in this regard. However, the terms "off-current" and "off-response" still remain in portions of the text and figure legends.

      Because the manuscript itself demonstrates that the response reflects recovery from transient acidification rather than a separate OFF signaling mechanism, these terms remain potentially misleading.

      I recommend replacing them with terminology such as:

      - photolysis-evoked PKD2L1 current

      - recovery current

      - proton-removal-induced current

      throughout the manuscript, including figure legends.

      We apologize, as the word “tonic” and the terminology “off-current” should have completely disappeared after the first round of revisions. We have now replaced those all along the text. “Tonic” has been replaced by “sustained”.

      “Off-current” or “off-response” have been replaced by appropriate terms in lines: 339, 342, 544, 545, 546, 549, 552, 555, 556, 560, 576, 580, 585, 588, 807 and 929. We have nevertheless conserved the term “off-current” in line 552 as we are referring to terminology used by other authors.

      Minor editorial corrections

      Figure 1Bd Please change: "po" to "Po" for consistency with standard channel physiology nomenclature.

      Figure 1Ca Please add units (mV) to the voltage labels shown on the left side of the traces.

      Figure 3E Please change: "Norm po" to "Norm Po".

      Figure 4Fb Please replace: "sec" with "s" to conform with SI unit conventions.

      Done.

      The authors have addressed the majority of my previous concerns and the manuscript has been substantially improved. The remaining issues are primarily related to terminology and overinterpretation rather than experimental deficiencies.

      Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments-relates to their multimodal CSF-sensing function remains unclear.

      The authors confirm in the GATA3:GFP mice where these cells are labeled that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding on how these sensory cells can detect so finely the changes of pH in the CSF.

      Weaknesses:

      Not weaknesses, there are only minor requests to complete the beautiful study.

      (1) The apical extension's response to removal of acidification is nicely illustrated in Figure 4C,G. There's something puzzling there: while the response to Glutamate is immediate, the channel responses to H+ is extremely delayed by 100ms - 2s, and even sometimes came in bursts separated by few hundreds of ms. H+ diffuse even faster than glutamate. Why is that?

      I don't quite understand how the response is so delayed & how to explain the recurring bursts of channel opening in the figure panel ?

      The kinetic of the response to proton uncaging is analyzed in Figure 4E, where the charge of the current traces is plotted against time. What this analysis shows is that the response lasts a few hundred ms (τ 250 ms) and then the PKD2L1 activity increase subsides to baseline. The peak of the response is at 100 ms (Figure 4G), but the increase in activity happens as soon as the uncaging pulse ends (Figure 4D, G and H). This behavior has already been shown in expression systems, where the channel activity is blocked by protons and the blockage is released when the acid is withdrawn. In an intact cell as the CSFcNs studied here, the exact kinetics of the recovery response are probably more complex (and variable) than in expression systems. Indeed, it is known that the recovery of this current depends, for example, on pH and extracellular calcium. Also, PKD2L1 are inhibited by intracellular calcium (de Caen et al, eLife 2016) but are themselves permeable to Ca<sup>++</sup> ions. The interaction of these effects could give rise to the “bursts” that are observed in some cases. However, this is merely speculative at this point.

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      Following the reviewer’s suggestion, we have added a trace in Figure 4C (upper blue trace) showing the spontaneous activity of the cell, prior to uncaging, as it is already shown for another example in Figure 4D.

      - Could the authors use a fluorescent pH sensor to monitor pH in the extracellular space and in the cell ?

      This is an important point that was already addressed by the reviewing editors in the previous round of revisions. Indeed, we have attempted to perform pH calibrations in the setup using the pHsensitive dye pyranine (or HPTS: 8-Hydroxypyrene-1,3,6-trisulfonic acid). HPTS is a very useful tool for pH calibrations in the physiological range: its pKa value is close to 7.2 and it can be used as a ratiometric dye (its fluorescence is pH-independent at 405–410 nm and pH-dependent at 450 nm). Unfortunately, the calibration under the conditions of a real experiment is not possible because the photolysis in the slice occurs in a tiny volume (approximately 1 µm³ in a total bath volume of more than 1 ml). In these conditions, the 405 nm uncaging pulse bleaches the dye in the photolysis spot and any useful information is lost. In addition, our imaging system is not fast enough to follow the pH change. As discussed in the Materials and Methods section, subsection “Estimation of the pH drop induced by photolysis” (line 791), the fast protonation of bicarbonate indicates that the pH change induced by the photolysis recovers in the submillisecond range.

      - Could the authors investigate whether in the apical extension, PKD2L1 channels are mainly at the outer membrane in the apical extension OR whether many channels are located in inner membranes ?

      PKD2L1 channels are probably subject to a high rate of turnover, and they are certainly localized in the plasma membrane of the apical process as well as in the inner membranes. Although this is a very interesting point, we believe it is out of the scope of this work.

      (2) Suppl Fig 4 is very cool and should be moved to main figure. The coupling of Soma and AP is very tight, yet there is a clear difference in targeting of channels that respond to cues in the CSF. In the context of an intact spinal cord, we can wonder how and when the contribution from ASIC in the some would be relevant to physiology. Can the authors think of experiments with an intact central canal to test the sensitivity and condition of recruitment of pH sensing in the soma (ASIC) versus the apical extension (PKD2L1)?

      We have followed the suggestion of the reviewer and have made Supplementary Figure 4 a main figure.

      The fact that the normal interphase between the spinal cord parenchyma and the cc is lost is already acknowledged in the discussion, lines 486 to 489. As the reviewer suggests, PKD2L1 and ASIC channels seem both to be important in the response of CSFcN to pH changes. However, both channels are activated in very different physiological contexts, as is discussed in the section “The involvement of ASIC channels”. Keeping the central canal intact in order to be as close as possible to physiological conditions, as suggested by the reviewer, would be ideal. However, as CSFcNs are in the middle of the cord, it would require the use of optical techniques that allow to penetrate deep into the tissue (e.g., 2-photon microscopy) that unfortunately are not available in our labs.

      (3) The Reissner fiber is missing after slicing the spinal cord. From our observations in fish, the fiber being under tension triggers lots of activity in CSF-cNs (Bellegarda et al Elife 2023) that also relies on PKD2L1 (Bohm et al NC 2016; Sternberg et al NC 2019). Could the authors discuss the contribution of the Reissner fiber to the PKD2L1 mediated modulation of CSFcN excitability ? Could the authors conceive a way to slice along the anteroposterior axis (sagitally) the spinal cord to keep the Reissner fiber in the central canal when recording CSF-cN apical extension ?

      - The authors should show in Fig 4C,G the traces for 1-2 s before uncaging occurs so we can appreciate whether such events occur as well in baseline and discuss this further in revisions.

      As discussed in the previous point, the in vitro slice preparation has technical limitations that are mainly related to the alterations of the normal structure of the tissue. Although keeping the Reissner fiber intact in a sagittal slice seems possible, accessing the CSFcNs with electrophysiological methods would still be a challenge.

      We have now added a sentence in the Discussion, lines 561 to 564, where we discuss that CSFcN excitability is modulated by the Reissner fiber and that it remains to be explored whether in rodents the gating of PKD2L1 channels is modulated by the Reissner fibre, as has been shown in zebrafish.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study examines the context-dependent modulation of auditory cortical neurons in response to expected sensory input, either self-generated sounds or expected perturbations of self-generated sounds. Specifically, using songbirds, the authors ask whether social context (the presence of a female conspecific) affects 1) the response of auditory cortical neurons to the bird's own song when he is singing; and 2) the response of neurons to perturbations of auditory feedback that the bird has been trained to expect.

      Strengths:

      First, the authors report that across the population, the responses of the neurons does not differ when a male bird sings alone or if he sings to a female. A fraction of auditory cortical neurons, however, do show significant differences in the firing rate, precision, and/or degree of burst firing when males sing alone vs. when they sing to females. This finding is broadly consistent with the literature showing that sensory neurons (visual, auditory, somatosensory, etc.) can be rapidly reconfigured into different "information processing modes" depending on behavioral state (e.g., quiescence vs. vigilance).

      For the perturbation experiments, the authors trained birds to expect distorted auditory feedback during a particular syllable. They found that some neurons showed greater responses during perturbation when a female was present (compared to when males were alone) while other neurons had smaller responses during perturbation when a female was present. In addition, the response of a small number of auditory cortical neurons were not affected by behavioral state. These results contrast with their prior report that the responses of midbrain dopaminergic neurons that project to the basal ganglia are "uniformly reduced" in the presence of a female, raising a question of how an evaluation signal is transformed in the circuit from the primary sensory region to the midbrain.

      Weaknesses:

      While the experiments and analysis are solid, the finding that social context can alter responses of auditory cortical neurons in a multitude of ways (increase, decrease or no change) raises several questions that can be examined with additional analysis. For example, do context-dependent differences in auditory responses derive from context-dependent differences in the songs? Are context-dependent differences present in all classes of neurons and throughout the auditory system?

      The observed heterogeneity in the firing properties of auditory cortical neurons, both in response to self-generated sounds and during perturbations of auditory feedback, raises the question of which neurons are sensitive to social context (which likely can be addressed by the authors in a revision). The authors should provide additional details about the recordings:

      (a) What are the locations of the recording sites? Prior work has shown that there is an organized map of spectrotemporal features of sounds in the auditory cortex of songbirds; spectral tuning widths change along the medial-lateral axis and temporal tuning widths differ between the input and output layers of Field L. Were the recordings primarily in Field L2 (thalamo-recipient region), L1 or L3? Were some recordings lateral to Field L in secondary auditory regions? Were the neurons that showed context-dependent changes in firing properties localized or distributed throughout Field L (i.e., were the context-dependent differences in neural responses truly brain-wide)? At a minimum, the authors should include a schematic showing the different regions of Field L and a summary of the location of the recording sites. Images of the processed tissue with electrolytic lesions would also be helpful.

      We agree that the anatomical targeting and limits of localization should be made explicit. In the original manuscript, we referred broadly to recordings from "Field L" and described targeting coordinates in the Methods. In the revised manuscript, we have softened the anatomical claim from "Field L" to "auditory pallium" where appropriate, while explicitly stating that electrodes were aimed at Field L. We also added anatomical caveats and a new supplemental figure.

      The revised title and abstract now reflect this more conservative anatomical framing. For example, the abstract now states: "Here we recorded neural activity from the auditory pallium in zebra finches practicing singing alone and directing courtship songs to females." In the Introduction, we now explicitly state both the intended target and the limitation: "We targeted our recording electrodes to Field L, a primary auditory pallial area that projects into multiple higher auditory areas that, in turn, project to VTA."

      We then added the caveat: "Field L is composed of multiple subdivisions and surrounds the interfacial nucleus and because the implanted wire bundles spread in a small radius of up to ~0.5 mm, our recordings likely included large territories of the auditory pallium (Figure S1)."

      We also added mechanistic/anatomical context for why these recordings may reflect activity shaped by broader auditory forebrain circuitry: "Although Field L is classically described as a primary auditory thalamorecipient region, its activity may also be shaped by contextual signals related to the courtship context, potentially via recurrent interactions with higher-order auditory forebrain regions such as the caudal mesopallium (CM) and caudomedial nidopallium (NCM) (Bauer et al., 2008; Figure S1C)."

      Changes made in revision: We added Figure S1, which includes anatomical subdivisions, an example histological slice showing the area where the cannula was implanted, and auditory pathway connectivity. We also revised the wording throughout the manuscript from "Field L neurons" to more conservative phrasing such as "auditory pallium neurons" or "pallial auditory neurons" when appropriate. We did not claim layer-specific localization, because the revised manuscript explicitly states that we cannot make such claims.

      (b) Was the context-dependent modulation limited to a particular class of neurons (distinguished by spike waveform shape, spontaneous firing rate, or other feature)?

      We agree that identifying whether context-dependent modulation is associated with specific neuronal classes is important. In the revision, we added analyses examining relationships between spike width and various firing characteristics. We also looked for potential relationships between mean rate and DAF response, IMCC and DAF response, and found no clear trend. Overall, we did not observe any clear relationship between DAF response, DAF response modulation, and metrics like spike width or mean firing rate.

      The revised manuscript states: "Action potential width of individual neurons has previously been used to classify putative interneurons or putative principal cells in the zebra finch auditory pallium (Calabrese and Woolley, 2015)."

      We then describe the new analysis: "We tested if DAF-response scores, mean firing rates during singing, burst fraction, IMCC, and the change in all of these between undirected and directed singing was correlated with spike half width (spike half-width measured as peak-to-trough time; Figure S5)."

      The revised result is: "DAF response in either condition, the change in DAF response across conditions, and the change in firing rate, burst fraction, and IMCC were not significantly correlated with spike width (Figure S5A-B)."

      We also report that some general firing properties did correlate with spike width: "Consistent with previous literature, firing rate was significantly correlated with spike width (Pearson's correlation, p=0.003; Figure S5C). Interestingly, burst fraction (p=8.3x10-4) and IMCC (p=5.4x10-5) were also significantly correlated with spike width (Figure S5C)."

      Changes made in revision: We added Figure S5 and associated text analyzing whether context-dependent changes in DAF response, firing rate, burst fraction, and IMCC were correlated with spike half-width. These analyses did not support the conclusion that context-dependent DAF modulation was restricted to a waveform-defined neuronal class.

      (a) Prior work has shown that songs of zebra finches differ slightly when males sing alone compared to when they sing to females: songs are faster; pitch is less variable; and the number of introductory elements is greater when males sing to females. Do some of the observed social context-dependent differences in the responses of auditory neurons reflect differences in the songs in the two conditions? Did the authors of this study also find premotor activity in Field L, and if so, did it differ between the two social contexts? Might differences in Field L responses reflect motor/song differences?

      The revised manuscript now addresses the issues of context-dependent changes in song in several ways. First, we explain why motif-aligned comparisons are meaningful: "The acoustic structure of undirected and directed motifs is highly similar in adult finches, enabling singing-related neural activity to be precisely aligned and compared across contexts."

      Second, we added analysis and discussion of premotor-related activity. Changes made in revision: New results paragraph and new Fig S4. "A previous study recording from Field L in zebra finches reported neural activations prior to the onset of singing, consistent with premotor signaling (Keller and Hahnloser, 2009). We tested for context-dependent changes in premotor activity by examining peaks in neural activity aligned to motif onsets. Across the population neurons did not exhibit significant changes in the timing of motif onset-aligned activity (Figure S4)."

      (b) For the perturbation experiments, this raises a question of whether perturbation amplitude is different when a male is alone and when a female is present. It would be useful to know if (and how much) perturbation amplitude varied depending on the location inside the cage as well as whether the sound pressure level of the underlying song was higher (e.g., Lombard effect).

      We previously calibrated the perturbation amplitude in Roeser et al., 2023, in an identical recording setup. Two speakers deliver the feedback on either side of the bird's home cage. We acknowledge the possibility that the position and orientation of the bird can affect the way the sound hits either of the bird's eardrums and thus potentially affect a neural response. However, neural activations following the absence of distortion playbacks were a major feature of the dataset and were context-dependent in some cases.

      The Methods state: "DAF was implemented with a custom LabVIEW acquisition program that analyzed song syllables in real-time and delivered syllable-targeted feedback." and "DAF (50 ms broadband noise bandpass filtered at 1.5-8 kHz to match frequency range of zebra finch song) was played over speakers in the recording chamber on top of a specific target syllable randomly on 50% of motif renditions."

      The revised manuscript also makes clear that experiments occurred in the bird's home cage: "Experiments were carried out in the male's home cage, which was inside a sound isolation chamber."

      Importantly, the revised Results show that not all DAF-related responses were simple activations to additional sound. Some neurons were activated by the absence of distortion: "Unexpectedly, some neurons were not activated by the song distortion but rather by the lack of target syllable distortion." and "These activations following undistorted renditions could also depend on the courtship context."

      Changes made in revision: We clarified the DAF stimulus and recording setup in Methods and added Figure 3 showing neurons activated following undistorted renditions. These data argue that context-dependent responses are not simply explained by DAF sound amplitude, although we do not claim that position-dependent acoustic variation was fully eliminated.

      Finally, it would be helpful if the authors could include a model and/or more discussion of how the uniform attenuation in midbrain dopaminergic neurons may arise given the heterogeneous responses in Field L.

      The revised manuscript provides evidence for context-dependent retuning upstream of VTA, but does not offer a direct mechanistic explanation for the uniform attenuation seen in dopaminergic neurons. The revised Discussion states: "Because the main goal of this study was to test if courtship-associated reduction in DAF signaling, recently observed in VTA DA neurons (Roeser et al., 2023), resulted from a local process in VTA or reflected a retuning of auditory responsiveness, we explicitly tested for changes in DAF responsiveness between alone and female-directed singing."

      It then explicitly contrasts auditory pallium and VTA: "Surprisingly, we discovered that Field L neurons could retune at the transition from lone to courtship singing in diverse ways, consistent with a more widespread process in the brain that does not fully explain the uniform DAF-signal attenuation observed in VTA."

      Changes made in revision: We expanded the Discussion to explicitly state that auditory pallium retuning is heterogeneous and therefore does not fully explain the uniform attenuation observed in VTA. We do not present a formal circuit model, but we now more clearly frame the result as evidence for broader sensory retuning that is likely transformed downstream.

      Reviewer #2 (Public Review):

      Summary:

      In the manuscript, Jones and Goldberg study auditory cortex in male zebra finches. They explore song-related responses in two different contexts, when the male is either alone or in the presence of a female. They find a heterogeneity of responses, in line with auditory cortical neurons computing the social modulation of responses found in VTA.

      Weaknesses:

      Stability of responses has not been studied: some neurons seem to have responses that slowly drift in time, which could lead to observed differences between alone and with-female conditions. Also, possible motor confounds and sound-of-audience confounds should be addressed. The language is often imprecise.

      Stability and Reversal: It is a bit unfortunate that stability of effects seemingly has not been studied by reversing experimental conditions. The work would be much stronger if authors could show that audience-dependent tuning is robust in individual cells. Did they record from some neurons during reversal back to the alone condition?

      We agree that recording stability is essential. A reversal experiment was not feasible for this dataset, as it is difficult to confirm whether song motifs produced immediately following female presence represent undirected singing or are directed to an unseen but recently present female. Instead, the revised manuscript adds a strict unit-stability criterion based on waveform similarity across conditions.

      The revised Results state: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." The exact criterion is: "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."

      Changes made in revision: The strict waveform-stability inclusion criterion and new Figure S2 directly showcase unit stability across the time course of the experiments.

      Motor responses: Does DAF playback change song? If so, especially if it applies only in one of the two conditions (audience/no audience), then the observed response differences could be motor-related rather than auditory responses.

      We agree that motor confounds must be minimized. We previously found that DAF did not affect the acoustics of the subsequent syllable (Gadagkar et al., 2016). The revised manuscript clarifies that DAF and undistorted trials were randomly interleaved and analyzed by comparing matched renditions within conditions. Importantly, we only analyzed motif-aligned activity, ensuring that all syllables within the song motif are the same.

      Changes made in revision: We clarified the DAF analysis framework and added a more conservative permutation-based analysis comparing distorted and undistorted trials within each context, then comparing those DAF-response vectors across contexts. We do not claim that all possible motor consequences of DAF are eliminated, but the analysis directly tests neural responses to randomly interleaved distorted versus undistorted renditions.

      Similarly, motif-aligned spiking activity was time warped to the median duration of undirected or directed motifs. Could the shorter motifs during directed song lead to alignment differences that would account for the different error responses in alone/with-female conditions?

      We agree this is an important technical point. The time-warping we conducted, standard in the field, compensates for the tempo differences between directed and undirected song. Importantly, our main analysis of change in error response no longer uses a 100 ms response window, but rather includes all windows in the motif.

      Changes made in revision: We clarified that the revised DAF response analysis uses motif-aligned, time-warped spike trains. Importantly, the revised analysis moves away from relying on a single scalar response window and uses bin-wise permutation tests with family-wise error correction.

      Audience versus sound of audience: Is it truly the audience that causes the difference in error responses or is it the sounds the audience makes?

      We agree that the sensory cues defining "audience" cannot be fully separated in this experiment. The reviewer raises an important point that female zebra finches occasionally call at the male. We have excluded all song motifs from analyses that include an overlapping female call.

      The revised Methods now explicitly state that motifs overlapping with female calls were excluded: "Any motifs that had overlapping time with a female call in directed motifs was excluded from analysis."

      We also revised the Discussion to treat the mechanism by which auditory pallium receives information about the female as an open question: "An open question is how auditory pallium receives information about whether a female is present, and how this information influences neural activity."

      Changes made in revision: We excluded motifs overlapping with female calls and added discussion explicitly acknowledging that how female presence is represented in auditory pallium remains unresolved. We do not claim to distinguish visual, auditory, social, or motivational components of the female-present condition.

      Reviewer #3 (Public Review):

      Summary:

      In this study, Jones et al. examine how neural activity in a primary auditory area (field L) of singing male songbirds is modulated by the presence or absence of an audience (a female conspecific). Prior work has demonstrated that the presence of an audience attenuates the responses of dopaminergic neurons to distortions of auditory feedback (DAF). Here the authors report that even in a region that is primarily considered sensory, responses to DAF are also modulated by the audience, although in a heterogeneous manner. However, to be fully persuasive, additional analyses will be required to address how much of the apparent modulation by audience may be explained by other factors such as changes in recorded neurons or their properties over time.

      (1) A central concern relates to whether the main reported effects associated with differences in singing directed versus undirected song reflect only those changes in conditions, versus contributions from changes in unit isolation or response properties over time.

      We completely agree that unit stability is critically important in this study. To address this concern, we now quantify stability and apply strict inclusion criteria adopted from a study that assessed unit stability over days (Dickey et al., 2009). Additionally, we now include average waveform overlays for all example units across conditions as supplemental Figure S2.

      Changes made in revision: We added: "Importantly, because these neural recordings were performed over long time courses (~2-8 hours), a strict threshold for stability was imposed." and "A Pearson's correlation coefficient of at least 0.99 between the average neural waveform during undirected and directed singing was required for a unit to be considered stable (Dickey et al., 2009; Figure S2)."

      (2) A second concern has to do with the categorical definition of 'error neurons'. The authors define a subset of neurons as error responsive only if their responses to DAF exceed a specific threshold (2.5 standard deviations). The problem is that for some neurons categorically defined as being responsive to DAF in only one condition, there is almost certainly not a significant difference in the actual responses to DAF between conditions.

      We overhauled our analyses characterizing DAF responses. Rather than relying only on a 2.5 z-score threshold, we now use a more conservative permutation-based approach that directly tests DAF responsiveness and context-dependent changes in DAF responsiveness.

      The revised Results state: "Statistical tests defining auditory neurons as DAF-responsive or not in a binary fashion may not be suitable if the underlying population of DAF-related responses exist on a continuum from responsive to non-responsive."

      The updated result is: "This more conservative approach identified 48/147 neurons as DAF-responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."

      (3a) Some discussion of what is already known about the auditory tuning of Field L, and the extent to which responses associated with distortion of feedback may reflect the frequency tuning of Field L neurons versus something that might be construed as more specifically as detecting an error in perceived feedback.

      We agree that DAF-related changes in firing do not necessarily imply that neurons are explicitly detecting an "error" between predicted and actual feedback. Field L neurons can have spectrotemporal receptive fields and frequency tuning such that a broadband DAF stimulus could drive excitation or inhibition simply because the stimulus overlaps with excitatory or inhibitory regions of a neuron's receptive field. We therefore revised the manuscript to use more cautious language and to describe these responses as DAF-related or feedback-related signals rather than categorically as "error responses".

      Changes made in revision: The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We also added a sentence to the Discussion: "However, it is important to note that DAF-related changes in firing in auditory neurons do not necessarily imply that neurons compute sensory prediction errors. DAF-related responses could arise from ordinary auditory tuning to the broadband distortion stimulus."

      (3b) It would also be useful to discuss further previous work on differences in auditory tuning or responses between conditions when subjects are vocalizing, versus when vocalizations are played back, and to what extent efference copy signals might contribute to the processing of feedback distortions.

      We agree these are important points. Our experimental design did not include sufficient passive bird-own-song (BOS) playback trials to permit quantitative comparisons with vocalizing conditions, and we therefore cannot draw firm conclusions about the contribution of efference copy signals to the DAF responses described here. We did observe robust motif onset-associated neural activations, including some activity preceding motif onset, which were present across both social contexts (see new Figure S4). These observations are consistent with prior reports of premotor-related signals in Field L (Keller and Hahnloser, 2009), but whether such signals contribute differentially to DAF processing across contexts remains an open question that we now acknowledge in the Discussion.

      (3c) To what extent did the current study control for any vocalizations or other sounds produced by females during the directed singing, and could this have contributed to differences in Field L activity between conditions?

      Please see response R2.4 above, in which we describe the exclusion of all song motifs that overlapped in time with a female call. This exclusion criterion was applied throughout all analyses of directed singing.

      Figure 1D: In the directed condition there are no spikes at all following the first handful of motif renditions. Were the directed and undirected recordings interleaved here?

      Undirected and directed trials were not interleaved. The raster plots are presented in chronological order; however, for each behavioral condition, rows are sorted with the earliest renditions at the bottom and the most recent at the top. We have clarified this in the figure legend.

      A minor issue: the raw example trace with male alone does not seem to have a corresponding set of points in the raster plot. For panel E, I also cannot find rasters that correspond to the example recordings shown at top.

      In the original version, we randomly downsampled the condition with more trials to equalize trial counts across conditions in the example rasters, while performing all analyses on the full set of recorded trials. As a result, the example spike shown in the raw trace was drawn from one of the downsampled trials not displayed in the raster.

      Changes made in revision: For greater transparency, we now include all trials from both conditions for each example neuron in Figure 1.

      Figure 2A also shows a neuron that looks like it has non-stationarity; for the alone condition without altered feedback, the main peak has no spikes for the bottom half of the rasters.

      In the original version, example neurons were selected to illustrate the DAF-response scoring method, which in some cases highlighted neurons with less stable response profiles. In the revised manuscript, we have replaced this example with neurons that exhibit more robust and stable DAF-related responses, and we now provide a broader set of example neurons illustrating both increases and decreases in DAF responsiveness across conditions.

      Other figures show firing rate distributions that appear to be very non-Gaussian, with some motifs during which there is a lot of activity, and others in which there is little activity. Please consider applying non-parametric tests as appropriate.

      We agree. In general, some neurons exhibited non-uniform firing rate distributions across trials. All of our main analyses are now conducted using non-parametric permutation tests, which do not assume a Gaussian distribution of trial-by-trial firing rates.

      Approaches to addressing the non-stationarity issue could include more specifically indicating examples in which recordings from the alone condition and directed condition are interleaved and exhibit reversible changes in the pattern of responses.

      Unfortunately, nearly all of our undirected and directed recording periods were not interleaved, as the experimental design required a block of undirected singing followed by directed singing with female presence. We find it informative, however, that DAF-response modulation was observed in both directions, with some neurons losing DAF responsiveness during directed song and others gaining it, a pattern that is difficult to attribute to a simple unidirectional drift in recording quality. We now provide additional examples illustrating both directions of modulation in Figures 2 and 3.

      The methods and/or raster plots should include some further explanation of the time periods over which recordings were made in the alone versus directed conditions, and the extent to which they are interleaved or not.

      We have clarified this in the revised Methods. In brief, recording began when the home cage lights came on each day, with the male left to sing alone until at least 40 undirected song motifs were collected. A female was then introduced in approximately 10-minute intervals until at least 40 directed song motifs were collected. The total recording duration on a given day ranged from 0.56 to 10.27 hours, reflecting variability across birds in the time required to elicit sufficient singing in each context. We have added this information to both the Methods and relevant figure legends.

      It would be most helpful to assess the stability of waveforms and unit isolation across time.

      We now apply strict inclusion criteria based on waveform stability, as described in R3.1 above. SNR was quantified as Vpp/(2*sigma_noise), where Vpp was the peak-to-peak amplitude of each filtered spike waveform and sigma_noise was estimated from the median absolute deviation of the filtered voltage trace. This combines the peak-to-peak normalization used by Nordhausen et al. (1996) with the robust noise estimator described by Rey et al. (2015). Waveform overlays for all included example units are provided in Figure S2.

      It would be reassuring to see that significant differences between conditions are equally or more prevalent under the conditions of greatest unit isolation and recording stability.

      The average SNR of neurons ultimately included in the analysis was 9.47 +/- 3.57, with a minimum of 4.69. Neurons that exhibited significant DAF-response modulation did not have a significantly different SNR than neurons that did not exhibit significant modulation (Wilcoxon rank-sum test, p=0.38). The mean SNR for significantly modulated neurons was 8.70, compared to 9.5 for non-modulated neurons, indicating that the detection of context-dependent modulation was not systematically biased toward neurons with lower recording quality.

      One other way that the authors might be able to address the main concern would be to look at the stability of firing patterns within conditions.

      We agree that stability of firing patterns within conditions is an important consideration, and this concern directly motivated the adoption of the permutation-based analysis described above. In this framework, the observed DAF-response difference between conditions is compared to a null distribution generated by shuffling condition labels across trials. This approach inherently accounts for within-condition trial-by-trial variability and does not assume stationarity of firing rates.

      It would be helpful to have additional explanations of the criteria used for counting spikes, and assessing stability of recordings.

      Spike waveforms were visually inspected for consistency using our custom MATLAB GUI on a 12-second file basis. Interspike interval violations below 1 ms were explicitly checked as an indicator of multi-unit contamination. Detection thresholds were manually set, and each recording file included in the analysis was independently inspected. We have added a more explicit description of these procedures to the Methods section.

      For the specific examples shown in figures, it would be useful to indicate by small tick marks or otherwise which spikes were counted as single units.

      We appreciate this suggestion. In the revised figures, we have improved the clarity of the example raw voltage traces by annotating the detection threshold and, where multiple units were present on a channel, indicating the waveform amplitude range corresponding to the isolated single unit. We believe this provides sufficient transparency regarding spike identity without requiring tick marks on every individual spike, which would substantially reduce legibility of the example traces.

      What were the criteria for determining multi-unit versus single-unit activity?

      In the context of this manuscript, "multi-unit activity" refers to channels on which no single neuron could be reliably distinguished from others based on waveform shape and amplitude. Units ultimately included in the study were those for which a single, consistent waveform cluster could be identified and isolated in the custom GUI. In cases where a second distinguishable unit was present on the same channel, it was manually excluded from the sorted single-unit record. We have clarified this distinction in the Methods.

      Categorical scores: This definition results in cases where responses of 2.45 vs 2.55 are described as 'retuned', even if these responses are not significantly different. Retuning would be more persuasively demonstrated if the authors could provide a test of whether or not the responses for individual neurons differ significantly between conditions.

      We completely agree, and thank the reviewer for motivating us to develop a more rigorous statistical approach. Our revised analysis uses a non-parametric permutation test that explicitly tests for significantly different DAF responses between undirected and directed singing conditions, with correction for multiple comparisons. This replaces the previous threshold-based categorical classification and directly addresses the concern that neurons near the threshold boundary were being treated as categorically different.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Minor comments:

      (1) Please include a schematic of the brain, including the different subregions of Field L and the connections between auditory regions and the midbrain.

      Done. Figure S1 has been added, including a schematic of Field L subdivisions and auditory pathway connectivity.

      (2) The authors should include some additional information about the recordings, such as the proportion of Field L neurons that exhibited singing-related changes in firing rate. It would be helpful to include some examples of spontaneous activity when the bird is quiescent in Figs. 1-2, especially for cells that do not show firing locked to song.

      We appreciate this suggestion. Given the scope of the current revision and the primary focus on DAF-response modulation, we have elected not to add spontaneous activity examples to Figures 1-2 at this time. We agree this would be a valuable addition in future work and have noted it as a limitation in the Discussion.

      (3) Methods, p. 10: Surgery and awake-behaving electrophysiology: "The of the cannula" - this is the only mention of a cannula. Do the authors mean the ends of the probes?

      Cannula placement and wire bundle extension from the end of the cannula has been clarified in the Methods.

      (4) Bottom of p. 10: Fix reference for biorxiv paper: "ref andreas paper"

      Fixed.

      (5) Methods, p. 12: Redundant sentences regarding significant error response criteria.

      Fixed. The redundant sentences have been removed.

      Reviewer #2 (Recommendations For The Authors):

      (1) The abstract is too vaguely formulated. Authors should try to quantify the statements already in the abstract.

      We have reworded the abstract to align with the revision's more conservative claims regarding social context modulation of auditory feedback, and have added specific quantitative statements where possible.

      (2) Authors repeatedly refer to 'perceived song errors' without performing experiments or reporting on behavioral readouts of how birds perceive the jamming sounds. The wording should be changed to something more neutral, e.g. 'DAF responses'.

      We revised the manuscript throughout to use more neutral language centred on "DAF-related" or "feedback-related" responses rather than "perceived errors" or "mistakes". The title was changed from "Auditory cortical error signals retune during songbird courtship" to "Auditory cortical feedback signals are modulated during songbird courtship". We similarly revised the abstract and all relevant passages in the Results and Discussion.

      (3) Authors write that 33 neurons were DAF responsive in both conditions. How should we interpret this overlap relative to independence and identity assumptions?

      We agree that the original presentation made the interpretation of overlap across conditions unclear. The observed overlap is greater than expected under a strict independence assumption but smaller than expected if responsiveness were identical across conditions, consistent with partial but incomplete sharing of DAF responsiveness across social contexts. In the revised manuscript, however, we have moved away from this binary classification framework because DAF responsiveness appears to vary continuously across neurons. The permutation-based analysis now directly tests for changes in DAF responsiveness across contexts without requiring categorical assignment.

      (4) Only 10 neurons were not affected by courtship state or only 10 error responsive neurons were not affected? I suggest authors do a multivariate analysis or use a mixed effect model and summarize the result as a table.

      We agree that the categorical accounting of neurons across conditions was difficult to follow in the original manuscript. In the revised manuscript, we clarified the distinction between neurons responsive to DAF within a condition and neurons exhibiting significant modulation of DAF responsiveness across conditions. We now explicitly report: "This analysis identified 71/147 neurons as DAF responsive in at least one behavioral condition, whereas 76/147 were not responsive in either condition." and "This more conservative approach identified 48/147 neurons as DAF responsive in at least one condition, with 13 of those neurons exhibiting a significant modulation in their DAF response between undirected and female-directed singing."

      (5) It would help if authors could define 'z-scored difference'. Better known is d prime, is this the same?

      For each neuron, the z-scored DAF response was computed as the z-scored firing rate difference between distorted and undistorted trials. Importantly, our revised main analysis avoids any normalization such as z-scoring, and instead uses a permutation-based approach applied directly to spike counts.

      (6) Is the 'retuning' assessment a bit conservative? Neurons could also retune by showing error scores greater than 2.5 in both conditions but a shifted response time.

      We agree that neurons could retune by shifting the latency of DAF responses. Although potential latency shifts are beyond the scope of the current study, we did observe suggestive evidence of possible latency changes in some example neurons across conditions. We have noted this as an interesting direction for future analysis.

      (7) Could the stability of DAF response across trials be described? E.g. as the ratio between intra versus inter condition variability?

      We agree that stability of DAF responses across trials is an important concern. In addition to imposing strict waveform stability requirements, our permutation-based statistical test explicitly accounts for trial-by-trial variability by constructing null distributions from within-condition trial shuffles. We have also replaced the previously shown unstable example neuron with neurons that exhibit more consistent DAF-related responses across trials, and provide additional examples in Figures 2 and 3.

      Minor:

      (8) 'significant increase in burst fraction': specify effect size of t test in results section.

      We now specify in the main text: "A small but significant increase in burst fraction was observed (paired t-test, p=9.3x10-6, n=138 neurons, mean +/- SEM: 0.11 +/- 0.006 vs 0.15 +/- 0.007, Figure 1J)."

      (9) The IMCC parameter should be specified in the main text.

      The Gaussian smoothing parameter (20 ms) has now been specified in the main text.

      (10) Fig. 2: indicate the windows within which error scores are computed.

      This is no longer applicable, as the revised permutation-based analysis does not rely on scoring error responses within a fixed window.

      (11) In Fig. 2A, the neuron has an error score of -2.54 (significant), but the red and blue curves look almost the same.

      We agree that the previous error score quantification did not always capture firing rate differences in an intuitive way. This example neuron has been replaced in the revised manuscript, and the new analysis avoids scalar error scores in favor of the permutation-based approach.

      Reviewer #3 (Recommendations For The Authors):

      Minor points:

      (1) "(ref andreas paper)." Add reference here?

      Fixed.

      (2) Hessler and Doupe 1999 is a good reference for premotor signal re-tuning during courtship.

      We agree. The reference has been included in the revised manuscript.

      (3) Page 5: "discharge depended on courtship state, using" - should this be "depending"?

      The original wording was intentional: "we tested how discharge depended on courtship state." We have verified this reads correctly in context and made no change.

      (4) Page 9: "consistent with a brainwide process" - what is meant here?

      We have revised this wording. The revised manuscript replaces "brainwide process" with clearer language describing a distributed modulation of auditory responsiveness that is not confined to a single nucleus.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study provides mechanistic evidence that tea-adapted two-spotted spider mite overcomes green tea catechin defenses via the horizontally transferred dioxygenase TkDOG15, supporting a two-step adaptation model, combining enzyme refinement and inducible upregulation. The evidence is convincing because multi-omics signals converge with functional validation (RNAi knockdown and recombinant enzyme assays) and well-controlled behavioral/toxicity assays to link TkDOG15 activity and expression to survival and feeding on tea.

      We thank the editors and reviewers for this positive assessment of the importance of our study and the strength of the evidence. We would like to point out one factual correction. The assessment describes the tea-adapted mite as the "two-spotted spider mite" (TSSM, Tetranychus urticae), but the species adapted to tea in this study is the Kanzawa spider mite (KSM, Tetranychus kanzawai). We would suggest revising "tea-adapted two-spotted spider mite" to "tea-adapted spider mite" or "tea-adapted Kanzawa spider mite" accordingly.

      Reviewer #1 (Public review):

      Summary:

      This study investigates the molecular mechanisms allowing the KSM mite to infest tea plants, a host that is toxic to the closely related TSSM mite due to high concentrations of phenolic catechins. The authors utilize a comparative approach involving tea-adapted KSM, non-adapted KSM, and TSSM to assess behavioral avoidance and physiological tolerance to catechins. The main finding is that tea-adapted KSM possesses a specific detoxification mechanism mediated by an enzyme, TkDOG15, which was acquired via horizontal gene transfer. The study demonstrates that adaptation is a two-step process: (1) structural refinement of the TkDOG15 enzyme through amino acid substitutions that enhance enzymatic efficiency against catechins, and (2) significant transcriptional upregulation of this gene in response to tea feeding. This enzymatic adaptation allows the mites to cleave and detoxify tea catechins, enabling survival on a toxic host plant.

      Strengths:

      A multiomics approach (transcriptomics and proteomics) provided a compelling crossvalidation of its findings. Functional bioassays, such as RNAi and recombinant enzyme assays, demonstrated that the adapted mite has higher activity against catechins via TkDOG15. Other methodologies, like feeding assay using a parafilm-covered leaf disc, were effective in avoiding contact chemosensation.

      Weaknesses:

      Although TkDOG15 is assumed to "detoxify" catechins by ring cleavage, the study doesn't identify or characterize the breakdown metabolic products. If the metabolites are indeed non-toxic compared to the parent catechins, that would strengthen the detoxification hypothesis. Also, the transcriptomic and proteomic analyses identified other potential detoxification enzymes, such as CCEs, UGTs, and ABC (Supplementary Tables 3-1 & 3-2), which were also upregulated. The manuscript focuses almost exclusively on TkDOG15, potentially overlooking a multigenic adaptation mechanism, where these other enzymes might play synergistic roles, although it was mentioned in the discussion section.

      Reviewer #1 (Recommendations for the authors):

      There is no need for additional experiments, but I suggest revising the discussion section to mention the weaknesses pointed out above.

      We thank the reviewer for the positive assessment and helpful suggestions. We have revised the Discussion (L276-283) to address both points as limitations. First, we now note that we did not characterize the products of TkDOG15-mediated catechin cleavage, and that confirming their reduced toxicity relative to the parent catechins would further support its detoxification role. We note this as a direction for future work. Second, we note that DOG15 in KSM on tea was the only enzyme upregulated at both the mRNA and protein levels, whereas the CCEs, UGTs, ABC transporter, and other DOGs were enriched in only one dataset. We now state that tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15 and warranting functional validation.

      Minor corrections below:

      (1) Figure 1a: For better readability, I recommend adding "KSM" and "TSSM" to the two pictures, respectively.

      Done.

      (2) L165: tetur20g01790 refers to a TSSM gene, while TkDOG15 refers to a TSM protein. Revise it accordingly. (Same for L442 and L481).

      The reviewer is correct that tetur20g01790 is the TSSM gene ID. As the KSM genome is not yet available, we identified the TkDOG15 gene, the KSM ortholog of tetur20g01790, by de novo assembly of our RNA-seq reads. We have revised L169 and L455 accordingly.

      (3) L264: Supplemental Table 3-2.

      Done (L271).

      Reviewer #2 (Public review):

      Summary:

      The fascinating topic of the host range of arthropods, including insects, and the detoxification of host secondary metabolites has been elucidated through studies of the host specificity of two closely related species. The discovery that key genes were acquired from fungi through horizontal gene transfer (HGT) is particularly significant.

      Strengths:

      (1) The discovery that the TkDOG15 enzyme, acquired through HGT from fungi, plays a key role in the detoxification of green tea catechins in the Kanzawa mite, revealing a new mechanism of plant-herbivore interactions, is highly encouraging.

      (2) The verification of this finding through various experiments, including behavioral, toxicological, transcriptomic, and proteomic analyses, RNAi-based gene function analysis, and recombinant enzyme activity assays, is also highly commendable.

      (3) By proposing a two-step model in which amino acid substitutions and expression regulation of a specific enzyme gene (TkDOG15) enable host adaptive evolution, this study contributes significantly to our understanding of the evolutionary mechanisms of speciation and plant defense overcoming.

      Weaknesses:

      While transcriptome/proteome analyses reported changes in the expression of other detoxification-related enzymes, including CCEs, UGTs, ABC transporters, DOG1, DOG4, and DOG7, it is regrettable that the contribution of each enzyme, including its interaction with TkDOG15 and the functional analysis of each enzyme within the overall catechin detoxification system, was not investigated.

      We thank the reviewer for the encouraging assessment and this comment. We agree that the contributions of the other detoxification-related enzymes, including their interaction with DOG15, remain to be investigated. As this point overlaps with a comment from Reviewer 1, we have revised the Discussion (L276-283) to note that DOG15 was the only enzyme upregulated at both the mRNA and protein levels, whereas the CCEs, UGTs, ABC transporter, and other DOGs were enriched in only one dataset. We now state that tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15, and that the functional analysis of their individual and combined contributions warrants future work.

      Reviewer #2 (Recommendations for the authors):

      The manuscript titled "Adaptation of an Herbivorous Arthropod to Green Tea Plants by Overcoming Catechin Defenses" presents a well-designed, mechanistically insightful study that advances our understanding of herbivore adaptation to plant chemical defenses. The work is scientifically sound and of potential interest to a broad readership in chemical ecology and evolutionary biology.

      However, before the manuscript can be considered for acceptance, the authors must adequately address the comments outlined below regarding clarity, presentation, and interpretation across the manuscript.

      We thank the reviewer for the positive evaluation of our study. We have carefully addressed each of the specific comments below regarding clarity, presentation, and interpretation, and we believe these revisions have substantially improved the manuscript.

      Specific comments on each section:

      (1) Abstract

      (a) The authors are encouraged to add a concise concluding sentence summarizing the broader significance of the study and indicating potential future research directions or limitations, which would strengthen the impact of the abstract.

      We have added a concluding sentence to the Abstract summarizing the broader significance of the study and indicating future directions (L38-40).

      (b) The authors may consider adding representative quantitative results to the abstract, as this would enhance clarity and increase the impact and interpretability of the study for readers.

      We have added representative quantitative results to the Abstract. Specifically, we now state that the mRNA and protein levels of DOG15 in tea-adapted T. kanzawai are up to 31.6 and 12.1 times higher, respectively, than in T. urticae fed on tea plants (L30-32). For consistency, we now refer to the gene as "DOG15" throughout the Abstract (L29, L30, and L36).

      (2) Introduction

      (a) While the paragraph is informative, it reads more like a summary of the main results than a statement of study objectives. The authors are encouraged to reframe this section to explicitly define the study's aims and hypotheses.

      We have reframed the final paragraph of the Introduction to explicitly state the study's aims and hypotheses rather than to summarize the results (L73-81).

      (b) The authors should avoid excessive citation of multiple references for a single thematic statement when one key reference is sufficient. Where appropriate, inclusion of more recent literature is encouraged.

      We have reduced multiple citations for single statements to the most representative references: Cabrera et al. (2006) for the health benefits of catechins (L46) and Grbić et al. (2011) and Dermauw et al. (2013) for the DOG gene count (L66-67).

      (3) Materials and Methods

      (a) The Materials and Methods section is comprehensive and technically sound; however, its length and density reduce overall clarity. The authors are encouraged to streamline descriptions of standard or well-established protocols and rely on appropriate citations where possible.

      We agree that clarity can be improved by removing redundancy. The Materials and Methods are intentionally detailed to allow independent replication of our protocols, so we have retained this detail and instead removed the overlapping methodological descriptions from the figure captions, where the same information was repeated (see our response to comment 6a).

      (b) Greater consistency is needed in reporting biological and technical replicates across different experiments (e.g., performance assays, transcriptomics, proteomics, and enzymatic activity assays) to enhance reproducibility.

      We have standardized the reporting of replicates across all experiments to the format "x independent experimental runs (n = y per run)." Throughout the manuscript, "independent experimental runs" denotes biological replicates, with technical replicates specified separately where applicable (three technical replicates for qRT-PCR).

      (c) The authors should provide brief justification for key methodological parameters, such as catechin concentrations, exclusion criteria in behavioral assays, and thresholds used for defining DEGs and DEPs, to improve transparency and interoperability.

      We have added brief justifications for the three parameters. 1) The catechin concentration range was chosen to encompass the individual catechin levels measured in fresh tea leaves (L340-341). 2) In the behavioral assays, inactive mites were excluded because their movement was insufficient to determine chemo-orientation behavior, and escaped mites were excluded because they did not complete the assay (L365-367). 3) The thresholds for DEGs and DEPs follow criteria commonly applied in mite transcriptomic studies (Vidal-Quist et al., 2025, newly added to the references) (L414-416) and are consistent with our previous spider mite proteomic analysis (Arai et al., 2025) (L444-445).

      (4) Results

      (a) While significant differences in survival and fecundity are reported, briefly indicating the magnitude of these differences (e.g., percentage or fold change) would improve clarity and strengthen the presentation (Lines 91-96).

      We have added the magnitude of the differences (L94-97). The revised text now states that after 10 days, almost 90% of tea-adapted KSM survived, compared with about 5% of non-adapted KSM and 33% of TSSM, and that tea-adapted KSM laid up to about 2 eggs/surviving female daily, whereas the other two populations laid almost no eggs.

      (b) The final sentences include interpretative and concluding statements regarding catechins as key metabolites and mite adaptation. These statements would be more appropriate for the Discussion section rather than the Results (Lines 127-130). Follow the same for the rest of the Results section also.

      Following the reviewer's suggestion, we have removed the interpretive and concluding statements from the end of the Results section, so that it now reports only the observations (L129-130). The interpretation regarding the multiple modes of action of catechins and the insensitivity of tea-adapted KSM is already presented in the Discussion (L206-213 and Conclusions), so we did not duplicate it there. We also reviewed the remaining Results subsections and confirmed that they report the experimental observations and their direct conclusions without broader interpretation.

      (c) The comparison among catechin classes is clear; however, briefly listing the mean concentrations of each catechin (as shown in Figure 2a) in the text would improve readability without duplicating the figure (Lines 135-139).

      We have added the approximate mean concentration of each catechin to the text (L136-137).

      (d) Please clarify in the Results whether the same exposure concentration and duration were applied for all catechins and mite species, or explicitly direct readers to the Methods section (Lines 141-142).

      We have clarified in the Results section that all four catechins were tested at the same concentration series (0, 10, 10<sup>2</sup>, 10<sup>3</sup>, 10<sup>4</sup>, and 10<sup>5</sup> ppm) and the same exposure duration (24 h) for both mite populations (L143).

      (e) The phrase "lower sensitivity" should be explicitly linked to LC<sub>50</sub> estimates to ensure that the basis of comparison is immediately clear to readers (Lines 143-144).

      Following the reviewer's suggestion, we have linked the sensitivity comparison to the LC<sub>50</sub> values (L143-147). The comparison is now stated relative to TSSM based on the LC<sub>50</sub> estimates, and for ECg and EC we note that the LC<sub>50</sub> of tea-adapted KSM exceeded the highest concentration tested.

      (f) This section clearly identifies TkDOG15 as a key gene underlying tea adaptation in KSM; however, the authors are encouraged to briefly clarify the criteria used to define "highly enriched" mRNAs and proteins (e.g., fold-change and statistical thresholds) in the Results text or by explicitly directing readers to the Methods. This would improve transparency and facilitate interpretation of the multi-omics comparisons (Lines 147-173).

      We have added the criteria used to define the enriched mRNAs and proteins (log2 fold change ≥ 1 with adjusted p-value < 0.05 for mRNA and p-value < 0.05 for protein) and referred readers to the Materials and Methods (L159-160).

      (g) The enzymatic comparison between TkDOG15 and TuDOG15 is well presented; however, the authors are encouraged to briefly discuss whether the two amino acid substitutions (Q127A and T203A) were individually or jointly responsible for the increased catalytic efficiency, or to acknowledge this as a limitation and potential direction for future functional studies (Lines 176-190).

      We have added a brief discussion of whether the two substitutions (Q127A and T203A) act individually or jointly (L254-257). We note that T203A is adjacent to the active-site residue Y202 and may contribute more directly to catalytic efficiency, and we acknowledge that dissecting their individual contributions by site-directed mutagenesis is a direction for future work.

      (5) Discussion

      (a) The authors appropriately acknowledge that the molecular basis of chemosensory insensitivity and the contribution of additional detoxification enzymes remain unresolved. To further improve clarity, these statements could be explicitly framed as hypotheses or future research directions to clearly distinguish them from experimentally supported mechanisms (Lines 205-208; 266-270).

      We have reframed the statements on chemosensory insensitivity (L209-213) and the contribution of additional detoxification enzymes (L272-274) as hypotheses and future directions, distinguishing them from the experimentally supported mechanisms.

      (b) While DOG15 is convincingly identified as a key contributor to tea adaptation, a brief clarification of its relative importance compared with other upregulated detoxification enzymes would strengthen interpretative balance, even if the roles of these enzymes remain unresolved (Lines 259-265).

      DOG15 was the only enzyme upregulated at both the mRNA and protein levels (Figure 3d,e), and the only enzyme functionally validated in this study, by RNAi silencing (Figure 3f) and recombinant enzyme assays (Figure 4c). We have established that DOG15 contributes to tea adaptation, but because the other upregulated enzymes were not functionally tested, their relative contributions cannot be determined at this stage. As we note in the Discussion, tea adaptation in KSM may be multigenic, with these enzymes potentially acting synergistically with DOG15 (L277-283). We therefore did not add further text, to avoid duplication.

      (c) The discussion linking host plant adaptation to reproductive isolation and ecological speciation is interesting and well contextualized; however, these evolutionary implications should be slightly tempered or explicitly framed as potential long-term outcomes beyond the immediate scope of the present study (Lines 271-281).

      We have tempered the evolutionary implications (L292-294). The revised sentence now frames the link to reproductive isolation and ecological speciation as a potential outcome over longer evolutionary timescales rather than a direct finding of the present study.

      (6) Figure captions

      (a) The figure captions (Figures 1-4) are exceptionally detailed and, in several places, repeat methodological information already described in the Materials and Methods. The authors are encouraged to shorten the captions by retaining only information necessary to interpret the figures, while referring readers to the Methods for experimental details.

      We have shortened the figure captions (Figures 1-4) by removing methodological details that are described in the Materials and Methods, retaining only the information needed to interpret each figure. Where appropriate, readers are now referred to the Materials and Methods or to Supplemental Figure 1-2 for the full experimental procedures.

      (b) Several captions contain long, multi-sentence descriptions that may hinder readability. The authors may consider simplifying the wording, grouping related panels more concisely, and removing procedural details (e.g., extraction conditions, exposure durations, and instrument settings) to improve clarity and visual accessibility.

      As described in our response to comment 6a, we have simplified the figure captions by removing procedural details such as extraction conditions, exposure durations, and instrument settings, and by grouping related panels more concisely. These details are retained in the Materials and Methods.

      (c) In Figure 1, the panel labels (a-h) do not appear in a clear sequential order. For consistency with the other figures and to improve readability, the authors should ensure that panel lettering is arranged in a logical, sequential order throughout the manuscript.

      We appreciate the reviewer's attention to panel ordering. In the current layout, the panel lettering follows the order in which the panels are first cited in the text. Arranging the panels in a strict left-to-right, top-to-bottom sequence would require reducing the size of several panels, including the HPLC chromatogram in panel (e) and the survival and fecundity time courses in panels (c) and (d), which would compromise their readability. We have therefore retained the current arrangement, in which related panels are grouped together and the larger panels are kept at a legible size. We hope the reviewer finds this acceptable.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary: 

      The manuscript by Yang, Wang, and Cléry presents a lightweight pipeline for real-time identification of common marmosets in a laboratory setting. Models were trained and evaluated on data derived from a family of three closely related adults and a set of juvenile twins. Freely moving animals entered an enclosed space fixed to the housing cage door, which permitted the entry of individual animals for data acquisition. Utilizing YOLOv8-nano, identification was improved through the introduction of uniquely colored collar beads. Analyses of facial similarity showed close morphological relatedness amongst individuals and highlighted the need for highly discriminative classification. Overall, the authors offer a framework for identity tracking that prioritizes real-time inference. The authors demonstrate that combining facial detection with visual markers enables adequate identity assignment under controlled laboratory conditions with minimal cross-individual misclassification. 

      Strengths: 

      (1) The proposed pipeline offers a solution for real-time identity tracking in common marmosets. Its lightweight design enables deployment across a wide range of hardware configurations. Furthermore, if similar strategies are employed, this methodology is likely adaptable for other species with minimal modification. 

      (2) Evaluation of closely related individuals provides a necessary stress test for the discrimination of facial identity tracking. 

      Weaknesses: 

      (1) The pipeline's reliance on controlled animal isolation and small visual markers raises questions about the approach's generalizability to unconstrained multi-animal environments. The provided confusion matrices (Figures 6-8) indicate that the most common misclassifications are background-related, possibly suggesting that detection specificity is the primary source of error. All things considered, these findings raise concerns about performance in its use in socially dynamic and visually complex environments. 

      Thank you for the comment. The background column of the confusion matrix can be explained by several occasions: a) the model detects an object where there is no object, b) there is more than one prediction label for the same object, or c) an object appeared in the image but not manually labeled, however the program was able to detect that object. The value of the background column does not necessarily mean that the detection is incorrect, as the precision score for the detection labels are good. We have rephrased the relevant sections for clarification to include the sources of the increased value in background columns in confusion matrices, as follows:

      “The background class of the confusion matrix showed frequently predictions as marmoset faces and collar beads for the training (Figure 6A) and validation set (Figure 6B). However, it does not necessarily indicate incorrect predictions or misclassifications. Instead, these values were mostly explained by multiple detections of the same object class. For instance, additional marmoset faces were predicted when multiple animals were present within a single video frame. The long collar structure or motion blur of the marmosets could also cause multiple detections of beads that belong to the same collar. This also corresponded to the high precision and recall scores observed across prediction classes (Figure 5D), suggesting that the increased background false positives were mainly related to the object-count discrepancies, instead of poor detection performance.”

      Prediction misclassification is one source of the background false positive. The misclassification could not be avoided in automatic prediction algorithm, but we included the manual filtering and majority-voting during our real-time classification to reduce this effect. Multiple detection of the same class may also be considered as the background, since only one object may be labelled in the ground truth, such as multiple collar beads or automatic face extraction. In addition, blurry objects were not labeled manually during training but can be detected during prediction, which also resulted in background false positive. It was clarified in the main text as follows:

      “The normalized confusion matrices showed high accuracy and consistency of most marmoset faces and collars detection in training (Figure 8A) and validation (Figure 8B) tests, with some exceptions. Particularly, the background was frequently identified as the collar of Young2 marmoset. This elevated background score was likely contributed by the multi-color design of the Young2 marmoset collar, making it more difficult to distinguish compared to collars with a single bead color. In this occasion, if one bead is occluded, blurred, or outside the field of view, the other visible collar bead could affect the prediction and lead to an incorrect identification from the ground truth.”

      (2) The manuscript claims performance comparable to that of human experimenters but provides no explicit evidence to support these claims. While it is plausible that human experimenters may be less accurate in facial recognition tasks involving closely related marmosets, the authors don't provide evidence. Moreover, while that might be the case, the color-coded beads provide a salient identity cue for the model, which complicates the interpretation of this comparison grounded in facial recognition. 

      Thank you for pointing out this concern. The aim of the facial recognition tool is to collect data from marmosets without having experimenters to check the identity continuously. The program is not aimed at outperforming the experimenters’ role but avoid having constant human intervention that can disrupt a more ecological in cage data collection. It is also essential for having more flexibility to collect data in case a specific experimenter is not here and thus to not disrupt the project. Human experimenters have extensive experience closely working with marmosets, having the unique collar beads associated to each marmosets allows human experimenters to hardly make mistakes identifying marmosets and to do it quickly. We collected identification accuracy of human experimenters by presenting 10 clips of the five marmosets involved in the manuscript (2 clips per marmosets), with 2 random clips repeated twice. The results were plotted by each experimenter. The identification accuracy of the experimenters correlates with the time spent with the animals, as the animal health technicians (responsible for daily health check and husbandry) achieved 95.83% average accuracy in identifying the marmosets. We clarified those points in the main text as follows:

      “Its automated pipeline substantially reduces the time and work required for traditional manual identity labeling, while maintaining an expert-level human performance and reproducibility across experimenters (95.83% average accuracy for animal health technicians, responsible for daily health check and husbandry while lab experiments ranges between 25 to 80% of accuracy depending on the amount of time spent with each animal, Supplementary figure 1). The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.”

      We filtered the prediction of the collars, and the identification result solely based on the faces for the 2 young marmosets was correct. The prediction results were plotted on Video 7, Video 7—video supplement 1, and Video 7—video supplement 2 and added to the Results section as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Explanation for classification of marmoset faces and collar beads in the Discussion section:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”.

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      Reviewer #2 (Public review):

      Summary: 

      In this study, Yang et al. develop a real-time system for automatic face detection and identification of multiple unrestrained common marmosets in a home cage setting. 

      Strengths: 

      The study aims to address an unmet need in behavioral neuroscience: the ability to non-invasively identify animals is crucial to the automated and rigorous study of neural behaviors; this is especially true for common marmosets, which are rapidly becoming a model system of choice for the study of complex social cognition. By using a YOLOv8 backbone, the study achieve human level performance, both in terms of precision and recall of the trained models.

      Weaknesses: 

      The robustness of the system is not clear from the limited datasets presented. The use of color-coded beads undercuts the study's premise that the system achieves truly non-invasive tracking. Although the system achieves good performance in face detection, it does not perform as well for classification using faces alone (especially when the faces are similar, as in twin animals). Here, too, the color-coded beads play a key role in identity discrimination. The stated goals of the study and the actual results presented are therefore at odds.

      Thank you for the comment. First, we would like to clarify the role of the collar beads in our system. Compared to the faces, a unique identity marker, the collar beads were not used as the main identity classifier but rather as an external visual marker. The color-coded bead was not used solely for the purpose of marmoset video classification; it was also used as an additional source of identification for one marmoset. As the marmosets usually move very fast inside the cage, it is mainly used as a visual marker for experimenters to recognize them in a distance in a short time.

      The mislabelling is more frequent with the young twins not only due to their face similarity, but also due to the limited number of images being used for the model training, as discussed in the paragraph #4 of the Discussion section. Collar beads are small and less frequently detected by the camera, since it could be occluded by the marmoset fur. In addition, it was invisible to the camera if the marmoset turned sideways or was far from the camera. Therefore, higher weight was assigned to the beads due to their small size and less frequent detections compared to face labels, such that it was only an element to confirm the identity, instead of the main classifier.

      The inclusion of the collar beads doesn’t affect the prediction results of the marmoset faces. The model achieved a good precision/recall score for the identity labeling in the manuscript. In the revision, with the majority-vote strategy, we filtered the detection of all collar beads and showed that the model was able to correctly identify the marmosets solely by their faces. The Results section has been modified as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Explanation for classification of marmoset faces and collar beads in the Discussion section:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”.

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      Reviewer #3 (Public review):

      Summary: 

      In this manuscript, Yang et al introduce a new method for automatically identifying marmosets in their home cage using a supervised deep learning method that recognizes the face and colored beads on marmoset collars. The authors show a high precision rate of identifying marmosets to levels comparable to a human experimenter. The method overall seems robust at identifying marmosets at different life stages and different settings; however, given the current form, I'm struggling to see the generalizability and experimental utility of this method. 

      Strengths: 

      (1) The authors provide a near-perfect automatic identification of marmosets in their home cage. 

      (2) This method is robust across lightning, camera angles, etc., making it potentially useful for marmoset (and other NHP) identification outside the housing cage as well 

      Weaknesses: 

      (1) Despite the almost perfect precision, in its current form, I'm failing to see how this method can be useful to other labs. 

      Thank you for your comment. This Tools & Resources paper mainly described the development of the marmoset identification program and methods. Future work will focus on extending the program application on identification from different housing conditions, in combination with various behavioral tasks such as in-cage touchscreen system or manual tasks, and in the wild that precludes handling or isolation of marmosets for collecting behavioral data. The program solely requires a camera, a computing device, and marmosets, as there are no hardware restrictions. In addition, we are currently collaborating with other labs on the marmoset identification from videos taken from other setups. The program achieved effective face extraction from the marmoset in the video, without the need for additional program modifications.

      (2) This is a nice methods manuscript, but the authors do not present results to show how their method can be used outside of identifying marmosets inside their home cages in a small field of view. 

      Thank you for your feedback. The method developed was applied in combination with other touchscreen behavioral tasks, aiming to extract data without human intervention. This approach was discussed in paragraph #6 in the Discussion section. While this manuscript focuses on the methods of close-view face identification when marmosets perform behavioral tasks, the identification and automatic face extraction program could also be applied to marmoset videos taken from a larger view, including phone cameras. Even though the marmosets are still housed in their home cage, the example videos presented the program’s application in a larger field of view. We have added examples of the videos/photos from a different experimental setup to respond to this comment in the Discussion section as follows:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10).”

      (3) Reading the manuscript is strenuous, given its repetitive nature. Consolidating and shortening the results, as well as adding some definitions to the results section, would be helpful. 

      Thank you for pointing this out and your suggestions. We have rephrased the Results section for simplification to facilitate the understanding of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The weight of the color-coded beads was increased to improve identification accuracy. From a brief look at the code provided on GitHub, the weight assigned to the beads seems substantial. This calls into question the need to use facial recognition in the identification strategy. As the code currently stands, facial identity appears to serve primarily as a fallback when bead detection fails to register. To strengthen the methodological justification, the paper would benefit from the authors providing a rationale for choosing this weighting scheme and, if available, supplemental figures showing performance across a range of different weights to demonstrate why that specific value was assigned in the algorithm.

      We have added the model prediction results without the collar beads showing that the facial recognition algorithm works even without the collar beads and that those collar beads are not the main classifier. The Methods section has been modified as follows:

      “For each detected bounding box, the scripts returned a corresponding label of marmoset face and collar bead color. We assigned the detected collar beads as the corresponding marmoset identity with a higher weight, which improved the detection confidence across frames.”

      The Discussion section has been modified as follows:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”

      “With one separated model trained per family unit, our system can utilize distinct collar colors as an additional identifier when available, while facial features performed as the main biometric marker. Even though multiple marmosets with visually similar faces may present close to the camera, the additional collar information can improve confidence in identity prediction without replacing facial recognition as the primary mechanism of identification (Video 7).”

      The weight assigned to the beads in the GitHub page is the highest weight that we would suggest. The actual weight can be customized by the experimenters based on the actual experimental setup. For example, we used the weight of 2 in our real-time version of the marmoset face identification, while marmosets were presented with their corresponding tasks once identified. The GitHub page has been edited to clarify this point.

      (2) The overall utility of this approach, other than the real-time detection component, needs more clarification. It is currently unclear why this approach, and in which specific experimental or observational settings, is particularly advantageous compared to existing methods for assigning animal identity.

      In addition to the advantages mentioned in paragraph #1 of the Introduction and paragraphs #1-3 in the Discussion, we have added more details in the Discussion section:

      “While existing marmoset identification approaches usually utilize visible markers, Radio Frequency Identification (RFID), or observation, the manual works and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters to check the identity continuously, instead of outperforming the experimenters’ role.”

      (3) Although it appears that performance based on faces and color beads was evaluated separately, this was not clearly presented, leading to confusion about whether face detection performance also benefited from color beads on the animals.

      Prediction of the different labels in the same model is independent, so the prediction of color beads is not affecting the prediction results of marmoset faces. Correct identity classification could be achieved without depending on the color beads, as we have filtered out the color beads detection class. The Results section has been modified as follows:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary confirmation rather than the main classifier of the system (Video 7).”

      Reviewer #2 (Recommendations for the authors):

      (1) I found the paper quite confusing as written. The term "model" is overused and highly conflated: there are the YOLOv8 pre-trained models, the "face classification" model, and the "automatic facial and identity extraction" model. The flowchart in Figure 2A is equally confusing. The mapping from the flowchart to the results is not straightforward, and I needed several passes to grasp it. I would recommend that the authors simplify the terminology and the mapping of the methods to the results.

      Thank you for pointing this out! The YOLOv8 pre-trained models, the "face classification" model, and the "automatic facial and identity extraction" model were indeed separately trained object detection models. They all have different weights/parameters but share the same YOLOv8 architecture/backbone. We removed some of the “model” term in the manuscript and replaced them with “classifier/framework/pipeline” to avoid misunderstanding. This information has been clarified in the revised manuscript of the Methods section and Figure 2, which provides an overall clearer explanation of the workflow of the methodology of the program.

      (2) It is not clear how robust these results are, given the limited data sets analysed.

      We agree that only five marmosets were involved in this manuscript, this unfortunately limited the robustness of the prediction results. Indeed, the limited number of animals that can be used per study has been a main limitation in non-human primate research, as they are very valuable animal models. However, we included approximately 3400 images in the training dataset, which were collected across days. New videos and photos that were captured from different devices were also used in the testing to ensure that the program can be used on new marmosets, different housing cage, and from different recording devices as indicated here:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair system, and computer-based device.”

      (3) There are two paradoxes regarding the stated motives of the study:

      (a) If the objective was to truly use non-invasive methods for the identification of animals, then why use the color-coated beads?

      As mentioned previously, identity detection can be made without collar beads, still with correct prediction results as indicated here:

      “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Video 7).”

      The color-coated beads are used for easier and quick marmoset identification during daily care, health check, for weekend staff, training or handling.

      (b) If the objective was to achieve high identification performance, and the color-coated beads are sufficient for this purpose, then why bother with faces at all?

      Collar beads are small compared to the face, and less visible due to fur occlusion and motion blur. Moreover, it is possible that some marmosets do not have collar beads due to their young age or when involved in other procedures such as imaging scans. The collars need to be checked and changed regularly in growing marmosets and it is not always convenient (some marmosets do not support the collar, some can have sensitive skin that would lead to abrasion) thus the need to develop a facial recognition system. Furthermore, marmosets who are from other labs or in the wild might not wear a collar with colored beads, thus face is the main classifier in this model to be more generally applicable. It is highlighted here:

      “Facial features serve as the main and intrinsic biometric identifier for each marmoset, providing a unique source of individual recognition. Since collar-based confirmation could be affected by visibility limitations, we implemented the uniquely color-coded bead collars as an auxiliary cue to provide additional confirmation in identity prediction. For example, this issue can be caused by identical or similar bead colors between individuals (Video 5, 6) and beads that are occluded by fur (Video 3 - 6). In addition, collar beads may change over time or not be worn by all animals.”

      (4) I was puzzled by the face similarity results in Figure 9. It appears that the face similarity measures were stronger (higher cosine similarity, lower Euclidean distance) for the adult data set compared to the twin data set. If so, why was it more challenging for the system to handle the twin data set?

      Face similarity can only be compared within models (therefore within adults and within twins). As this is calculated from different models, the adult face similarity cannot be compared with twins’ face similarity. It has been clarified in the Methods section as follows:

      “Statistical tests were performed only within the face classifier of each marmoset family, as embedding spaces may vary in scaling, learned features, and baseline metrics making cross-model comparison of inter-individual face similarity unreliable (Bollegala, 2017).”

      And in the Results section as follows: “We performed the statistical tests only on the face classifier for the adult marmoset family, as the twin marmoset model only involved two individuals and thus not valid for within-model statistical analysis (Table 2, 3).”

      The twin dataset aims to represent a test for the program utility in new marmosets, especially for testing if the program can still distinguish between the marmosets with similar faces. Thus, the number of twin data collected is less than the adult dataset, as explained by Discussion paragraph #2 “While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition.” This explains the challenge the system faces when differentiating the twins, while increasing the training dataset is required to solve this issue.

      Reviewer #3 (Recommendations for the authors):

      Major issues:

      (1) My main issue is regarding the utility of this method in scientific experiments. This manuscript is a "methods paper" introducing a face recognition method to identify a single marmoset in their home cage in a very specific and confined field of view. This comprises a limitation on what experiments can be performed using this method. On the contrary, if (a) the authors can show that this method can be used for a bigger field of view, where the social structure/interactions can be studied for neuroethological, cognitive or social studies that will make this method significantly more robust; or (b) design an experiment that can be performed using the current method to show that this method in its current form is sufficient.

      (a) Our method worked in larger home cage (larger view) with videos taken inside the cage / outside the cage, with multiple marmosets moving around, while the camera and its fixation are also moving. A new figure (Figure 10) has been added to highlight this wide application:

      “The motivation for this real-time marmoset identity recognition program was to develop an easy-to-use, generalizable pipeline that could be applied across different marmosets and lab environments, such as using larger field of view or phone cameras (Figure 10). The pipeline was designed to have no specific hardware requirements and can be implemented for any standard recording device, including any commonly available cameras, primate chair system, and computer-based device.”.

      (b) We are currently using this method to collect in-cage touchscreen data with multiple marmosets without the need to isolate such animals to acquire the data, avoiding social separation. The collection of data in nonhuman primates is still a long process, so we wanted to share the facial recognition system first, aligned with our commitment towards open science, to benefit the broader community (we have already been contacted by two labs since the publication of this preprint) while we keep collecting data for the scientific project. We have added the touchscreen application as example in the Discussion section as follows:

      “Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.”

      (2) The authors claim a longitudinal identification of marmosets, yet I think the data to fully support this are deficient. This might be a result of unclarity of this experiment. How was this experiment done? Was the training done on the 7 months and then applied to the 11 months? Are there more continuous data that track the precision of the identification in time? For example, how does the twin identification evolve in time?

      This Tools & Resources paper mainly described the development of the marmoset identification program and methods. Ongoing work in the lab, the main research focus of which is the longitudinal assessment of cognitive functions, either during neurodevelopment or in preclinical ageing model, is benefiting from such algorithms to help identifying the animals to collect in cage behavioral data. As such, we have done some testing in one young cohort. The training of the young marmosets’ identification was done only on the 7-month data, and then we applied the identification program to the videos of the same marmosets when they were 11 months old and 16 months old (for the no-collar results) as indicated as follows in the Methods section:

      “Moreover, we evaluated the model performance and its generalization across developmental stages using new videos: 1) from the adult marmosets and 2) from the same young marmosets at 11 months, which were not involved during initial program training”.

      And in the Results section as follows: “We tested the prediction performance without collar and its longitudinal application, using the face-only prediction on the young marmosets at 11 months and 16 months (Video 7, Video 7—video supplement 1, Video 7—video supplement 2). The face classifier correctly identified the young twin marmosets solely based on their facial features, indicating that facial identity classification was performed independently of collar information and that the collar beads acted only as an auxiliary cue rather than the main classifier of the system (Video 7).”

      The identification program was shown to correctly identify the marmosets; however, we found that “While comparing between the adult and young marmoset datasets, we found that the adult marmosets’ face classifier, trained with a larger number of varied images, showed more reliability and efficiency in marmoset identity recognition.”

      The mislabeling was more frequent with the young twins not only due to their face similarity, but also due to the limited number of images being used for model training. The face images used in the identification model training were less compared to the adult model, which contributed to a less accurate prediction result. As the marmoset is still developing before adulthood, their face features will become more different as they age. By increasing the number of training images from different ages of the young marmosets, this could be solved as it is therefore possible to build efficient identification program for longitudinal study. Thus, instead of the current classifiers presented in this manuscript, we suggested that the method/tool could be beneficial for longitudinal studies, not restricting to the individual-based identification program mentioned in this manuscript, as “The tool’s advantages are particularly efficient for large datasets and longitudinal studies, where manual identity labeling becomes difficult, as variability and errors increase along with dataset size and experimenter number.”

      (3) How does this method compare to other methods that were used in the past?

      The advantages were mentioned in paragraph #1 of the Introduction and paragraph #1-3 in the Discussion. Current approaches for marmosets are usually visible markers (ear dye, collar, etc.), RFID, or observation, of which manual works and human interventions are required. These methods usually need continuous adjustment due to tighter collar, dye fading, etc. This can affect marmoset behaviors, especially during their behavioral task performance, as mentioned as follows:

      “While existing marmoset identification approaches usually utilize visible markers, Radio Frequency Identification (RFID), or observation, the manual works and human interventions involved can impact animal behaviors, especially during their behavioral task performance. The facial identification tool aims to collect data from marmosets without having experimenters to check the identity continuously, instead of outperforming the experimenters’ role.”

      (4) Did the authors think of adding a continuity or a space constraint? For example, video 6 shows misidentification of the twins; in this specific case, adding a probabilistic continuity or space constraint that will limit identity switches might be useful. This can also be using a retroactive correction - for example, video 3.

      We would like to thank the reviewer for this suggestion. We agree that these approaches will be valuable improvements for future offline analysis.

      The probabilistic continuity constraint can indeed help decrease identity switches. In our current application, we have implemented a temporal smoothing through majority voting across a 30-frame (1 second) window, of which the program outputs the most frequent prediction of identity. With this strategy, we could reduce the occasional frame misprediction and maintain the real-time performance. Our animals are free-moving and may appear in any location within the camera field of view and housing cage. Therefore, position is not strongly associated with the identity of individual.

      We agree that retroactive correction could improve the detection consistency for offline analysis by correcting past detection by future prediction results. However, the current pipeline is incorporated with behavioral tasks, meaning that the prediction results aim to be transmitted with minimal time delay. As additional frame analysis and extra computational power may be needed for retroactive correction, the increased latency can be limited to the utility of real-time system and task control.

      Minor issues:

      (5) In its current form, I think the manuscript can be significantly shortened and the results/figures can be consolidated (confusion matrices with validation figures for example).

      We have followed the reviewer’s suggestion and shortened the Methods and Results sections.

      (6) The term "unseen" that the authors use in their results is confusing. Are the authors referring to monkeys that are hidden from their view, or "unseen" before by the model? The video indicates the latter, but I think the term can be changed to something less confusing, like novel, new, etc.

      We have changed the term “unseen” by “new” to avoid confusion.

      (7) Can the authors add information about the relationship between the number of manually labeled images and the identification precision?

      The relationship between number of manually labeled images and identification precision has been described in the Discussion section as: “Moreover, the performance of the system is strongly dependent on the amount and variability of the training data, with identity classification improving as more marmoset images are involved in the model training.”

      This means that more manually labeled images (i.e. larger training datasets) could improve the identification precision. However, model performance will plateau regardless of training dataset size, referred in the Results: “Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs).”

      More manually labeled images could help improve the variability of model prediction, but too many of these images are also risky for overfitting. In this case, overfitted model might not be able to make valid predictions on new videos/images.

      (8) Can the authors expand a bit about the difference between YOLO Nano, small, and medium in the methods?

      Thank you for the suggestion. We have added a brief description of the pre-trained models in the methods: “These pre-trained models share the same object detection backbone but differ in number of parameters and computing power. Larger models, such as YOLOv8 medium, provide higher detection accuracy but require greater computational resources and longer inference time. In contrast, smaller models prioritize the computational efficiency.”

      (9) Clear and short definitions of what IoU, Recall, F1, and other terms represent should be added to the results section (not formulas, short sentences).

      The definitions and formula for the evaluation metrics were described in detail in the Methods section. To help readers while avoiding repetitions with the detailed methodology definitions, we have added a brief description of these terms at their first mention in the Results section: “Model performance included the precision (the proportion of correct positive predictions), recall (the proportion of corrected predicted ground-truth labels), and mAP@50–95 (the average detection accuracy across different IoU object localization thresholds; see Methods and Materials section for detailed definitions).”.

      (10) In the methods-"video collection" section, can the authors please include more information? Is this a motion-sensitive camera? Otherwise, what's the size of the data that is collected? This will help in reproducibility and system requirements. If this is not continuously collected data, discuss what can be done to make an online identification tool.

      We have included the camera as industrial color/RGB camera, which functions like any webcam and has no motion-sensitive functions. The size of data collected was described in detail in the Methods section that we slightly modified for clarity for: “Three adult marmosets from one family were recorded for 1 hour across each of the 5 recording days, with unrestricted voluntary access to the primate chair space. For the adult marmosets, the housing cage door was opened at the beginning of the recording session allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 1C). Two young marmosets were briefly isolated and recorded separately for testing and improving the automatic face extraction program. We recorded them at two developmental time points, 7 months old and 11 months old (an additional time point at 16 months old has been added for one marmoset to test the identification without collar). During video collection, a sliding panel and an in-cage box were positioned near the housing cage door to temporarily isolate individual marmosets from other family members. Individual isolation was kept brief (approximately 10 minutes) to prevent disturbance and potential stress due to family separation. “To capture sufficient variability in postures, individuals, and lighting conditions, clips were sampled throughout the adult marmoset videos (approximately 5 hours) (Figure 2B).”

      The size of the training dataset was also described in detail in the Methods section, referred as: “To minimize image computations and data storage, we created a dataset of 2498 annotated images from the three adult marmosets. All images were manually annotated to label marmoset faces, individual identities, and the collar bead colors (Figure 2C). The annotated images were used for training models of multi-marmoset face classification and the automatic identity extraction, which can automatically detect, localize, and identify marmoset faces (Figure 2A, Step 1 – 4). We created another dataset of two young marmosets at 7 months old (total images = 502) for testing the automatic facial and identity extraction (Figure 2A, Step 4 – 5). For both adult and young marmoset datasets, images were randomly divided into a training set and a validation set at a ratio of 8:2.”

      Our manuscript is not describing an online identification tool (i.e. the described program does not require connection to internet). Instead, once trained and the program is performing well with new marmoset videos, we could use the trained weights for real-time marmoset identification (no need to collect new training data) as we are doing it for our touchscreen data collection in cage. It is referred in the Discussion as follows: “Once trained, the system operates automatically to collect real-time identity and can work to present subject-specific behavioral or cognitive tasks based on the identity of the detected animal, with no work or presence needed on the user’s end. This tool has already been implemented in touchscreen-based marmoset cognitive tasks, including pairwise visual discrimination paradigm.”

      This program can be used for online applications if you are using a camera that is connected 24/7. For now, we are only using it while using our behavioral testing chair due to limitations issues (safety recordings, removing all electrical apparatus during night, overheating of camera if use continuously).

      (11) Would increasing the size of the beads help with their identification? In the images included in Figure 2C, it's very difficult to see these beads.

      Yes, collar beads can be occluded by fur, blurred by motions, or outside the field of camera view. We have now included increasing the collar beads size in the Discussion, referred as: “An alternate experimental solution is to improve collar visibility, including using distinct color code across individuals within a family, increasing the size of the beads, or increasing collar beads number to reduce occlusion.” However, the size of the beads needs to be appropriate to avoid being inconvenient and disruptive to the animals to ensure their welfare.

      (12) It's difficult to understand the setup of the camera in regard to the housing cage. Can the marmosets go into the primate's chair at any point (from the videos, it seems so, but the dashed line in 1C might indicate otherwise), or is the primate chair there only to mount the camera? Consider redrawing 1C in a more clear way.

      The figure 1C is a simplified drawing of the photo 1A, we have edited the drawing of Figure 1C to highlight the free access when the chair is mounted to the cage as stated here: “For the adult marmosets, the housing cage door was opened at the beginning of the recording session allowing them to enter and exit freely into the primate chair space for food rewards and observation (Figure 1C).”

      (13) Given the repetitiveness of the figures, an icon atop each figure specifying what's being tested will be helpful.

      We have added an icon for each figure 3-8.

      (14) I feel the supplemental figures for Figure 9 are more compelling than the main figure. Consider including some of the panels in the main figure.

      Thank you for the suggestion. We included heat maps of the adult family relationships in Figure 9, including the cosine similarities and Euclidean distances. The current legend of Figure 9 is changed as follows: “Figures 9. Across-model visualization of the face similarity between marmoset pairs. Four types of family relationships (mother-father, father-son, mother-son, and twin1-twin2) were compared, based on the training results of adult and young marmosets. The similarity was calculated using (A) cosine similarity and the (B) Euclidean distance. Heat map of the (C) cosine similarity and (D) Euclidean distance was plotted between the 3 relationship pairs in the adult family. The cosine similarity score ranged from 0 (very different) and 1 (exactly same) for marmoset faces. The Euclidean distance score ranged from 0 (exactly same) and 1 (very different) for marmoset faces.”

      (15) Related to this, I feel that the display of Figure 9 obscures the differences that the authors report.

      The display of Figure 9 has been improved following the above reviewer’s suggestions.

      (16) In Figure 9, given the large effect sizes but non-significant p-values, will adding more training points/epochs improve the differences.

      For the trained recognition programs, we have already implemented the “early stopping criteria” as shown in the Results section: “Each of the models was trained until reaching the early stopping criteria (i.e. no improvement within the last 100 training epochs).” This means that the program has already plateau with its performance.

      (17) Line 30: either "within" or "in".

      This has been corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This important study investigates how distinct honey bee viruses differentially alter flight performance through interactions with octopamine signaling pathways. The combination of behavioral flight assays, pharmacological perturbation, and transcriptomic analyses provides solid evidence that virus-specific effects on flight are associated with octopamine signaling. However, some of the stronger mechanistic conclusions regarding direct regulation of octopamine signaling remain incomplete without more specific validation of receptor-level effects and direct quantification of octopamine levels or signaling activity.

      We revised some of text in the manuscript, since we agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence—including gene expression analyses, OA supplementation experiments, and behavioral measurements—that collectively support a role for octopaminergic signaling in mediating the observed effects. The revised text better reflects the data included in this paper.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Kaku and Flenniken investigate the mechanistic pathways through which specific viral infections alter the flight capabilities of honey bees. Building on their previous discovery that DWV impairs flight while SBV unexpectedly enhances it, the authors hypothesized that these behavioral shifts are driven by interactions with the insect's octopamine (OA) signaling pathway, which is responsible for the "fight-or-flight" neurohormonal stress response and energy mobilization. To test this, the authors experimentally infected adult honey bees with DWV or SBV and pharmacologically manipulated the OA pathway using either octopamine supplementation or epinastine (EP), an OA-receptor antagonist. They then evaluated the bees' flight performance (distance, duration, and speed) on custom flight mills and profiled their gene expression using qPCR and RNA sequencing.

      Strengths:

      A major strength of this study is the high prevalence of preexisting background DWV and SBV infections in the honey bee cohorts, which meant there were no completely "virus-free" control groups. However, the authors successfully mitigated this limitation by rigorously quantifying viral RNA copies for every individual bee via qPCR and utilizing these viral abundances as continuous variables in powerful linear mixed-effect models.

      Weaknesses:

      The primary weakness lies in the methodology used for targeted pharmacological manipulations, as well as the lack of OA quantification across different treatments. Thus, their claims are not sufficiently supported by the current data.

      We thank Reviewer #1 for these comments.

      (1) The authors utilize Epinastine to block octopamine signaling, describing it as a highly specific OA receptor antagonist. However, pharmacological inhibitors often lack absolute specificity. Epinastine might bind to other octopamine receptor subtypes present in honey bee neural and flight muscle tissues, or it could potentially cross-react with tyramine and dopamine receptors. Without further genetic validation (e.g., RNA interference targeting specific receptors), it is difficult to definitively conclude that the altered flight performance is solely due to the blockade of the specific Oβ−2R pathway.

      We thank the reviewer for this thoughtful comment and agree that pharmacological approaches have inherent limitations with respect to receptor specificity. However, among the available octopamine receptor antagonists, epinastine is considered one of the most selective compounds for insect octopamine receptors. Roeder et al. (1998) reported that epinastine exhibits affinities for octopamine receptors that are at least four orders of magnitude greater than those for other insect biogenic amine receptors, including dopamine, tyramine, histamine, and serotonin receptors. We updated the text to include this information.

      Honey bees encode four β-adrenergic-like receptors (AmOARβ1- AmOARβ4) and one αadrenergic-like receptor (AmOARα1). Our transcriptomic analyses indicated that expression of AmOARβ2 was substantially higher than that of other octopamine receptor genes. Specifically, AmOARβ4 transcripts were not detected in our RNA-seq datasets, while AmOARβ1 and AmOARβ3 were expressed at very low levels in most samples (Supplementary Table S9; Figure S5). Although AmOARα1 transcripts were detected in some samples, expression levels were consistently lower than those of AmOARβ2. These observations support the interpretation that the physiological effects observed following epinastine treatment are primarily mediated through disruption of AmOARβ2 signaling. We updated the text to include this information.

      We agree that receptor-specific genetic approaches would provide valuable complementary evidence. RNAi-mediated knockdown of AmOARβ2 is an attractive future direction; however, RNAi efficacy in honey bees is variable and influenced by factors including transcript turnover rates. In addition, dsRNA treatments can induce sequence-independent antiviral effects that could confound interpretation in studies involving viral infection (Flenniken and Andino PONE 2013; Brutscher, Daughenbaugh, and Flenniken Sci Reports 2017). We have revised the manuscript to more explicitly acknowledge these limitations and to clarify the basis for our interpretation of the epinastine experiments.

      (2) As a natural neurotransmitter, insects have evolved highly efficient "cleanup" mechanisms. OA is rapidly cleared from the synaptic cleft via reuptake transporters and quickly inactivated by enzymes such as N-acetyltransferase (NAT) or Monoamine Oxidase (MAO). Consequently, an injection of OA produces only a transient "pulse" of activity. It is often a poor "tool" for inducing prolonged physiological effects compared to synthetic formamidines like Amitraz.

      We thank the reviewer for this important point regarding the pharmacokinetics of octopamine. We agree that octopamine is rapidly metabolised and cleared under physiological conditions and that exogenous administration is unlikely to precisely mimic endogenous signaling dynamics. Our goal was not to induce a prolonged pharmacological activation of octopamine signaling comparable to that produced by synthetic agonists such as amitraz, but rather to determine whether increasing octopaminergic signaling could mitigate the flight impairments associated with DWV infection. Octopamine was administered either by injection or through feeding (Lines 86-89), both of which resulted in significant improvements in flight performance in DWV-infected bees (Figure 2). The observation that two independent delivery methods produced similar outcomes supports the conclusion that enhanced octopaminergic signaling can partially rescue the DWV-associated flight phenotype. We have revised the manuscript to clarify this distinction and to acknowledge that exogenous octopamine administration likely produces transient elevations in signaling rather than sustained receptor activation.

      (3) The study relies heavily on transcriptomics and quantitative PCR to measure the mRNA expression of key synthesizing enzymes, namely tyrosine decarboxylase (tdc) and tyramine βhydroxylase (tβh), to infer the activation or suppression of the octopamine pathway. However, changes in enzyme synthesis at the RNA level are often insufficient to accurately reflect the true physiological levels of biogenic amines. To robustly prove the authors' hypothesis of a "feedback loop that regulates intracellular OA concentrations", direct quantification of actual octopamine and tyramine titers in the bees (e.g., using high-performance liquid chromatography or mass spectrometry) is necessary.

      We thank the reviewer for this comment and agree that octopamine and tyramine quantification would strengthen the mechanistic interpretation of our findings. Previous studies have successfully quantified OA in honey bees using HPLC-based approaches, including KayaZee et al. (2022, eLife), who measured OA in honey bee muscle tissue (both naturally occurring levels and levels post-treatment with 10 mM OA), and Cook et al. (2017, J. Exp. Bio) who quantified OA in pooled honey bee brain samples.

      Prior to submission, we inquired with our institutional mass spectrometry facility regarding the feasibility of measuring OA in individual honey bee samples. The expected concentrations of OA in our samples was below their limit of detection, so we did not pursue these analyses at that time. During the review process, we explored the possibility of analyzing a subset of samples at external facilities that may have the sensitivity required to quantify OA and tyramine in honey bee tissues. Since such analyses would require substantial resources, with estimated costs of approximately $5,000–10,000 for 12–15 samples that have been stored in the -80C since the study, rather than flash-frozen in liquid nitrogen as described by Zee et al. 2022. While we acknowledge that direct measurements of OA and tyramine would provide valuable complementary evidence, the current study relies on multiple independent lines of evidence— including gene expression analyses, OA supplementation experiments, and behavioral measurements—that collectively support a role for octopaminergic signaling in mediating the observed effects. We thank the reviewer for this valuable suggestion. While these analyses are beyond the scope of this study, we will consider using this approach in future studies.

      Reviewer #2 (Public review):

      Summary:

      This highly original and well-designed study provides insight into how honeybee picorna-like viruses, Deformed wing virus (DWV) and Sacbrood virus (SBV), affect flight performance, and reveals the role of the octopamine (OA) pathway in virus-honeybee interactions. The authors used a flight mill to quantify the flight performance of bees with different levels of DWV and SBV. Bees were treated with OA and/or epinastine (EP) - an OA receptor antagonist; the study also quantified virus loads and expression of two key genes involved in OA biosynthesis.

      The results showed that reduced flight performance associated with high DWV levels could be alleviated by OA administration. In contrast, increased levels of SBV had the opposite effect, leading to enhanced flight performance. This suggests distinct physiological responses to DWV and SBV infections. Administration of EP had led to a reduction of flight performance in SBVinfected bees, indicating the involvement of the OA pathway.

      The authors also quantified levels of mRNAs of enzymes involved in OA synthesis, tyrosine decarboxylase (TDC) and tyramine beta-hydroxylase (TbH), and concluded that DWV induced expression of TbH, while SBV upregulated expression of TDC. Furthermore, the study identified upregulated and downregulated genes in response to SBV, DWV and DWV in combination with OA.

      Strengths:

      The study reported opposing effects of infections of related viruses, SBV and DWV, on honeybee flight performance, and identified the central role of the octopamine (OA) signaling pathway in the effect of viruses on honeybee flights.

      These findings were achieved by using a combination of approaches, including experimental measurement of flight distance, virus infections, and introduction of OA and EP. Experimental work with honeybees is technically challenging and requires specialized expertise, which makes the results produced in this study more valuable.

      DWV and SBV are among the most important honeybee pathogens affecting honeybee health and threatening the pollination service. Therefore, an understanding of the mechanisms underlying DWV and SBV pathogenesis has the potential to develop novel approaches to mitigate the negative impact of these viruses.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      We thank Reviewer #2 for these comments

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions for the manuscript.

      (1) L. 45-46

      Please note that not only high virus levels have a negative impact on honeybees. Low levels of DWV, typical of covert infections, can have long-term deleterious effects on honeybee foraging and survival. Please include citation (e.g., Benaets et al, 2017, Proc Biol Sci (2017) 284 (1848): 20162149. https://doi.org/10.1098/rspb.2016.2149).

      We thank the reviewer for these comments and edited the text accordingly, and apologize for our inadvertent omission of Benaets et al 2017, which is cited in our previous publication.

      (2) L. 113

      Clarify what is meant by "high DWV levels"

      "i.e., 10^8 copies / 2 ug RNA" -> "i.e., above 10^8 copies / 2 ug RNA"

      We thank the reviewer for this comment and corrected this in the text.

      (3) L.115

      "..mock infected bees.." /Figure 2A.

      Did these bees have low levels of DWV, below 10^8 / 2 mg RNA? What was the level of DWV in these bees?

      Mock-infected bees had an average of 3x10<sup>5</sup> DWV copies and 2x10<sup>3</sup> SBV copies per 2 µg RNA (reported in Lines 107-109 in revised manuscript, a few lines before in original manuscript).

      (4) Figure 2 / Legends to Figure 2

      Note that in Figure 2 legends, the grey areas show 95% confidence intervals for regression lines.

      We thank the reviewer for this comment and added this in the text.

      (5) Figure 2 / Legends to Figure 2

      Consider including correlation coefficients (R) and p-values for each of the regression lines in Figures 2A-F. (These could be included in the Figure 2 legends).

      We thank the reviewer for this suggestion and agree that providing sufficient statistical information is important for data interpretation. Because the analyses presented in Figure 2 are based on linear mixed-effects models that incorporate both fixed and random effects, the statistical outputs are more complex than those associated with simple linear regressions. For figure clarity, we chose not to include all model statistics within the figure panels or legends and the key statistical results, including p-values and model fit metrics (R<sup>2</sup> values), are reported in the main text (Lines 121+). In addition, complete model outputs, including all relevant coefficients, correlation estimates, and associated statistics, are provided in Supplemental Data Sheet S4. To address the Reviewer’s comments, we revised the figure caption to improve clarity and include key p-values.

      We believe this approach better balances accessibility in the main figures with comprehensive reporting of the statistical analyses and thank the Reviewer for this useful suggestion.

      (7) L.357-377 - virus-specific responses

      A previous honeybee transcriptome analysis study, which showed different responses to DWV and SBV, could be cited (Ryabov E. 2016. PeerJ 4:e1591 https://doi.org/10.7717/peerj.1591).

      We thank the reviewer for this point and included this citation in line 338 of original manuscript (line 349 in revised, tracked-changes manuscript).

      (7) L. 412

      "bees were collected 24 hours prior to eclosion" -> e.g. "bees were collected at pupal stage 24 hours prior to eclosion"?

      Specify if dark-eyed pupae were collected to make sure eclosion in 24 hr.

      We thank the reviewer for making this point, and we revised the methods and results text to improve clarity.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This manuscript reports an important study in which the authors apply smFRET imaging to probe HIV-1 Env conformational dynamics in the presence of antibodies. Previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Through the cutting-edge application of smFRET imaging, the study provides convincing insights into the mechanisms of action of relevant antibodies.

      We appreciate this positive assessment and thank the reviewers for their time and constructive comments. We have made the following changes in the revised manuscript to address all points raised by reviewers.

      (1) Clarify the distinction between suppression efficiency and functional cost.

      (2) Add controls: smFRET experiments in the presence of monovalent 10E8.4 and iMab individually.

      (3) All of the smFRET population contour plots have been removed, as suggested.

      (4) Repeat neutralization experiments of tagged viruses (carrying nc-AA-incorporated, amber-suppressed Env), add and compare infectivity profiles between before and after click-chemistry labeling of tagged viruses.

      (5) Add a section (Complementary views from smFRET and structural studies) to the Discussion on how these approaches complement each other.

      (6) Further clarify three prefusion conformational states identified by smFRET, the relation with previously identified States 1, 2, 3, and asymmetry, the heterogeneity of Env presentations and virion morphology, and the focus of this study.

      Please find below our point-by-point responses to the public reviews and recommendations for the authors.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors have considered a panel of antibodies that target epitopes at the gp120/gp41 interface (8ANC195 and PGT151), the fusion peptide in the gp41 domain (VRC34), and the MPER region of gp41 (DH511.2_K3 and VRC42). They also investigate 10E8.4/iMab, which is an engineered bispecific antibody that targets the MPER and the CD4 receptor. On a technical note, they have applied a double amber codon-readthrough strategy to incorporate the non-natural TCO*A amino acid, which gets labeled through click chemistry. This approach should result in less disruption of the native Env structure as compared to the peptide insertion previously used for smFRET imaging of Env. Furthermore, previous implementations of smFRET imaging of HIV-1 Env, which focus on gp120 conformation, have yielded limited information on antibodies that target gp41. Altogether, through the cutting-edge application of smFRET imaging, the study provides novel insights into the mechanisms of action of interesting and clinically relevant antibodies.

      Thank you for the positive comments!

      In validating the functionality of the S401TAG/R542TAG Env, the authors performed infectivity assays and observed 20% infectivity as compared to wild-type (Figure S2A). However, the text equates this with "20% dual-amber suppression efficiency". This would benefit from some explanation. Why do the authors interpret infectivity as reporting on amber suppression efficiency, and not the functional cost of modifying Env, which is probably unavoidable? Or a combination of both? Is there data to suggest that 100% amber suppression would leave Env 100% functional? If so, this would be valuable to show. If not, the text should be clarified.

      We acknowledge this concern and have clarified the distinction between suppression efficiency and functional cost in this revised manuscript. The observed reduction in infectivity does not translate into functional loss; instead, it more reflects the efficiency of suppression (one of the critical limitations of applying genetic code expansion in mammalian cells). To support the preservation of Env functionality, we performed dose-response neutralization experiments of tag-free and 100% dual-ncAA-incorporated Env virions by two trimer-specific neutralizing antibodies, which exhibited similar dose-dependent neutralization sensitivity (Fig. 1D), providing stronger validation than infectivity assays. We also compared infectivity between labeled and unlabeled virions and observed no significant difference (Fig. S3B).

      We have previously discussed several limitations of amber suppression in mammalian cells when combined with smFRET viral systems (PMID: 38232732; PMID: 40716060) and, more recently, in our methodology chapters (PMID: 42349953; PMID: 42349954). In brief, orthogonal tRNA/aaRS pair–mediated amber suppression (reassigning/repurposing amber stop codons to non-canonical amino acids) of the introduced ambers in the target protein (Env in our case) must compete with the cellular translation system, particularly release factors that recognize amber codons and terminate translation. Readthrough of endogenous amber codons in virus-producing cells (in our case, HEK293T) can disrupt normal protein expression and virus production. Similarly, readthrough of pre-existing amber codons in HIV-1 ORFs other than the targeted ambers in Env can disrupt virus assembly, which we addressed by generating an amber-free provirus (PMID: 38232732). Introducing two amber codons into Env further reduces efficiency, as dual suppression requires two sequential successful suppression events within the same Env molecule.

      The authors state that the contour plots in Figure 2E reveal "dynamic sampling" of the observed FRET states. Strictly speaking, as presented, the contour plots (and FRET histograms) provide no information on dynamics per se. They indicate only the relative thermodynamic stabilities of the FRET states; transitions between states are a matter of interpretation. The TDPs, shown later in Figure 5A, nicely display the dynamics. More importantly, interpretation of the contour plots is challenging, as some seem to suggest an evolution toward lower FRET states. This is especially evident in Figures 2F and 3D, which suggest that the system evolves into a stable 0.1-FRET state (CO) after about 3 sec. Unless the authors want to conclude something from this, I would suggest that they consider removing the contour plots, since their interpretations are fully supported by the FRET histograms alone.

      We agree and have removed the contour plots, as they do not add meaningful information beyond what the histograms show.

      The data indicating that Env conformation is manipulated by 10E8.4/iMab is interesting. If I understand correctly, 10E8.4/iMab is an engineered antibody with one Fab targeting MPER and the second Fab targeting CD4. In the absence of CD4, could the difference between 10E8.4/iMab and the other MPER antibodies be due to 10E8.4/iMab being monovalent with respect to MPER binding?

      We appreciate this question. To address this, we have performed important controls: smFRET experiments in the presence of 10E8.4 and iMab individually in the absence of CD4. The results are shown in Fig. S9 in the revised manuscript, which indicates that 10E8.4 behaves similarly to other MPER-directed bNAbs we tested in this study, whereas iMab does not appear to affect the conformational populations of Env. The dual effect exerted by the bivalent 10E8.4/iMab is therefore very unexpected and thus interesting, as discussed in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Xu and co-workers unveil two distinct modes of neutralisation by gp41targeted broadly neutralizing antibodies on HIV-1 Env. So far, it was unclear as to how the mechanism of neutralisation occurred for this subset of neutralising antibodies (that can target the fusion peptide or the membrane proximal external region of the gp41 subunit). Thanks to single-molecule FRET, the authors show that the majority of broadly neutralizing antibodies stabilize the closed Env conformation (named State 1 since the original work by Munro and colleagues PMID: 25298114). Interestingly, the bivalent 10E8.4/iMab stabilized in turn a CD4-bound open state of Env. The two modes of neutralization described for these antibodies show previously unknown allosteric mechanisms that stabilize closed and open Env conformation, stressing the importance of Env conformational dynamics and its efficiency during the process of fusion.

      Strengths:

      The article is well-written, and the figures fully depict the data in a convincing way. The authors have used smFRET, which is now established in the field as a good tool to assess Env dynamics.

      We appreciate these positive comments!

      Weaknesses:

      (1) The limited controls on how click chemistry affects Env (as labelled Env HIV virions were not evaluated).

      We agree. Our previous validation focused on ncAA-incorporated Env HIV-1 virions, but not the fluorescently labeled virions. To address this, we have added infectivity results for labeled virions after click-chemistry labeling, compared with those before labeling. We did not observe any measurable difference in infectivity (Fig. S3B), indicating that the labeling procedure does not impair viral infectivity.

      We also attempted to perform dose-dependent neutralization after labeling. However, as anticipated in our provisional response, this remains technically challenging because the additional labeling and centrifugation steps substantially increase sample handling time, while the dual amber suppression system already limits virion production in cells. As a result, we were not able to obtain sufficiently robust datasets for this additional functional validation.

      Nevertheless, we have previously demonstrated real-time tracking of single click-labeled Env virions during internalization and intracellular trafficking in live cells (PMID: 38232732), providing independent evidence that click-chemistry-labeled Env retains functional competence.

      (2) Photobleaching of donor and acceptor molecules occurs right after 10sec exposure.

      We acknowledge this limitation and have included it in the revision.

      (3) Other limitations are well described in the corresponding section.

      We appreciate this comment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As a means of clarifying the mechanism of 10E8.4/iMab, the authors might consider performing separate smFRET experiments in the presence of the normal 10E8.4 antibody and the normal iMab antibody (a negative control). Alternatively, they could consider imaging in the presence of the DH511.2_K3 and VRC42 Fabs (as opposed to full-length Ig) to make a cleaner comparison, although this may be less informative given the high concentrations of antibodies used.

      We thank the reviewer for this excellent suggestion. To enable a direct comparison, we performed the most informative control by examining virus-associated Env in the presence of 10E8.4 alone and iMab alone. The corresponding smFRET results are presented in Fig. S9. We found that 10E8.4 behaves similarly to other MPER-directed antibodies, whereas iMab alone does not appear to have a notable effect on the conformational propensity of Env. For transparency and to facilitate future antibody design, we have also included the Fab region sequences of the antibodies in Table S2.

      Reviewer #2 (Recommendations for the authors):

      The article is well-written, the findings are of high interest for the community. The article should be shared once the points stated below are clarified and revised by the authors.

      We appreciate this comment and the points raised by the reviewer and have revised the manuscript accordingly.

      (1) In Figure 1C, the tomographic slices showing HIV-1 WT as compared to HIV-1 decorated with EnvBG505 S401ncAA R542nCAA are quite different morphologically. The micrographs chosen show a big particle with two capsids close to a smaller one without a capsid and at least in this plane bold (no Env incorporation) for the WT; whilst for the HIV-1 decorated with EnvBG505 S401ncAA R542nCAA no capsid is apparent in both particles, one (the right one) is very small and the right one does present a number of Envs but no apparent capsid is visible here. Please comment - perhaps it would be important to average the morphological traits of both and look at average diameter, average Env incorporation, morphology of the capsid, percentage of immature particles, capsid abnormalities (as the one shown in the upper micrograph).

      We thank the reviewer for this thoughtful comment.

      The tomograms in Fig. 1C were included to demonstrate the overall size and shape of the viral particles rather than to provide a quantitative structural comparison. HIV-1 viral particles are inherently heterogeneous, and the original slides were selected as representative examples. Following the reviewer's suggestion, we replaced the representative wild-type (Fig. 1C, top panel) and tagged virus (Fig. 1C, bottom panel) tomographic slides with those that better reflect the overall quality of each sample. To further address this concern, we refer the reviewer to the nanoparticle tracking analysis (NTA) shown in Fig. S3, which shows no significant difference in particle diameter between the wild-type and tagged viruses. In the revised manuscript, we now replace "morphology" with "shape" or "size," as these terms better reflect what our results can say.

      We agree that a quantitative analysis of capsid morphology, Env spike incorporation, and the proportion of immature particles would be informative. However, such analyses would require a substantially larger cryoET dataset, which is beyond the scope of the present study; nevertheless, it is certainly in our interest to pursue a cryoET-focused study of EnvCA interactions, with Env complexed with 10E8.4/iMab. Our primary objective is to study Env conformational dynamics by smFRET rather than viral morphogenesis or capsid maturation, whose relationship to Env dynamics remains largely unexplored. It is also worth noting that the optimal particle populations for smFRET and cryoET differ. smFRET measures the conformational dynamics of individual Env trimers and therefore selectively analyzes virions containing a single dually labeled Env trimer, whereas cryoET structural analyses typically benefit from particles with higher Env spike densities. Therefore, the particle populations favored for the two techniques are not identical.

      (2) In Figure 1D, there is a difference in neutralisation with PGT151 - how different are these two curves - how does the labelling affect neutralisation for bNAbs targeting gp41? Would it be possible to assess also the infectivity, fusion and neutralisation profiles of particles where the flurophores are included? This would be without diluting the Env for single particle analysis, but just to understand how harsh the organic reaction is and how it affects Env function (as all experiments and conclusions in the manuscript are based on labelled Env).

      Again, we sincerely appreciate these questions, which have helped us improve the manuscript. Neutralization assays for the tagged viruses were performed using the ncAA-incorporated, amber-suppressed viruses, whereas the engineered wild-type is amber-free. The differences between these two dose-response curves in the original Fig. 1D are small and within the experimental variation routinely observed under even identical conditions (same virus and same bNAb). We have repeated these experiments, and the new results are shown in the revised Fig. 1D. Although minor variations remain, the overall neutralization profiles and IC50 values are highly consistent.

      To assess whether the fluorophore labeling reaction affects Env functionality, as noted above, we have included infectivity results for labeled virions after click-chemistry labeling, compared with those before labeling. We did not observe any measurable difference in infectivity (Fig. S3B), indicating that the labeling procedure does not impair viral infectivity. We also attempted to perform dose-dependent neutralization after labeling. However, as anticipated in our provisional response, this remains technically challenging because the additional labeling and centrifugation steps substantially increase sample handling time, while the dual amber suppression system already limits virion production in cells. As a result, we were not able to obtain sufficiently robust datasets for this additional functional validation. Nevertheless, we have previously demonstrated real-time tracking of click-labelled Env virions during internalization and intracellular trafficking in live cells (PMID: 38232732), providing independent evidence that click-chemistry-labelled Env retains functional competence.

      We believe that the unchanged infectivity of labeled viruses relative to their unlabeled counterparts, together with our previously observed real-time trajectories of click-labeled virions in live cells, provides strong evidence that our labeling strategy does not measurably impair Env function.

      (3) In Figure 2E and 2G, the authors employ a three Gaussian fit approach to recover the three populations (pre-triggered - pre-fusion closed - CD4 bound open). Can you please relate these with State 1, 2 and 3 from the original article (PMID: 25298114). Comment on the possibility that more than three populations could be fitted and what this could mean - pre-triggered and partially open (one gp120 asymmetrically open) could occur? Could this labelling approach account for this asymmetry?

      Thanks for this suggestion. In this study, we compared our results obtained using the gp120-gp41 structural axis with those obtained using the referenced gp120 V1-V4 structural axis to confidently assign the FRET-identified states to the previously reported three primary populations. The referenced axis is comparable to those used in the original article (PMID: 25298114) and later confirmed using the amber-click strategy (PMID: 38232732). We observe the same structural changes from these two distinct structural angles, as probed under ligand-free conditions (Fig. 2E and 2G) and CD4-triggered open conditions (Fig. 2F and 2H).

      The pre-triggered state corresponds to State 1; the pre-fusion closed state corresponds to the symmetric State 2 (which the SOSIP-based soluble Env primarily adopts; PMID: 30971821); and the CD4-bound open state corresponds to the fully open State 3. The assignment of the FRET states observed from the gp120 V1-V4 structural axis to States 1, 2, and 3 was originally reported in two studies (PMIDs: 27795397 and 29561264). In the asymmetric trimer configuration, the State 2 FRET signal originates from the free protomer, while the other one or two protomers bind CD4 and adopt the open conformation (PMID: 29561264). The asymmetric intermediate (PMID: 29561264) was identified using a heterotrimer experimental design consisting of a mixture of wild-type and CD4-binding-incompetent D368R protomers, which was not used in the present study. Therefore, our labeling approach cannot unambiguously resolve this asymmetry.

      Regarding the possibility of more than three populations, evidence from current and previous studies (PMIDs: 25298114, 27795397, 29561264, 38232732, 30971821) strongly supports the presence of three primary states of virus-associated Env, with additional substates that can be resolved under specific triggering conditions (PMIDs: 30974085, 41326374, 39640534). The assignment of such substates requires well-controlled experimental designs (PMIDs: 30974085, 41326374, 39640534).

      We have related PT, PC, and CO to States 1, 2, and 3, and added comments on multiple states and asymmetry in the revised manuscript.

      (4) When comparing smFRET with CryoET or structure, one can see that in smFRET there are always many potential conformations for big sub-populations of Env. Indeed, there is a trend, and the addition of bNAbs (Figure 4) clearly has an impact on increasing and stabilizing a particular state as defined by the authors (e.g. PT at 45% upon addition of 8ANC195, but also 32% PC and 23% CO). I assume that when analysing single particle CryoET or single virus CryoET, one needs to discard after template matching different scenarios that do not necessarily contribute to the highest resolution and this information is not always discussed. It would be interesting to address this in the discussion as the effect on Env dynamics of adding ligands (including CD4 and 17b) is not inducing in all Envs a drastic conformational change - this could be derived from the Ka of the ligands, but also from the intrinsic Env heterogeneity in both dynamics and architecture - I think that addressing these matters in the discussion could be of interest for the community. In this regard, the transition density plots are very helpful.

      We completely agree and appreciate this insightful suggestion. We have expanded the Discussion to better address the complementary insights provided by smFRET and structural approaches. In single-particle cryoEM, we do not observe the full spectrum of Env conformations for technical reasons, not because particles are intentionally discarded to obtain only the highest-resolution structures. One reason is that open Env conformations are much more sensitive to radiation damage than closed Env. Likewise, ligand-free closed Env is more sensitive to radiation damage than a bNAb-stabilized closed Env. Thus, the outcome of an SPA cryo-EM study depends strongly on the biological question being addressed and the conformational state that is preferentially preserved under the experimental conditions. Although one could hypothetically collect much larger datasets to recover lower-abundance conformations, this would be both cost-prohibitive and unlikely to faithfully represent the relative conformational populations due to differential, conformation-dependent radiation damage.

      We agree that the smFRET data highlight an important aspect of Env dynamics. Ligands, including bNAbs, CD4, and 17b, generally shift the conformational equilibrium toward particular states rather than driving all Env trimers into a single conformation. This likely reflects both differences in ligand binding properties and the intrinsic conformational heterogeneity of Env. We therefore believe that structural studies and smFRET provide complementary information. Structural methods resolve the molecular architecture of individual conformational states at atomic (by cryoEM) and near-atomic (by cryo-ET) levels, whereas smFRET quantifies their relative populations and dynamic interconversion. As the reviewer pointed out, the transition density plots are particularly valuable in illustrating these dynamic changes.

      We have incorporated these points into the revised Discussion (Subtitle: complementary views from smFRET and structural studies in general).

      (5) One of the very interesting findings of the paper is that the effect of bNabs (at least the ones tested) have an impact on the Env dynamics and how this shift can alter entry - therefore the structural view is perhaps less important - In spite of this, we still employ a structural jargon to refer to "Env open conformation stabilisation" for instance - even if the data shows that upon ligand exposure dynamics are still important but shifted. Please comment.

      This is a great point. Our smFRET data show that bNAb binding generally shifts the conformational distribution toward and stabilizes particular Env states by lowering their free energy and increasing their occupancy. We have clarified that ligand-induced stabilization reflects a redistribution of the conformational ensemble while preserving the intrinsic dynamic nature of Env.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Fujita and colleagues investigated two selective peripheral nerve voltage-gated sodium channel inhibitors targeting either Nav1.7 or Nav1.8 on the excitability of human dorsal root ganglion neurons. The authors discovered that Nav1.8 inhibition is more effective at suppressing repetitive firing of DRG neurons, and this may explain the greater clinical efficacy observed for suzetrigine.

      Strengths:

      The study is interesting, and the findings are conceptually satisfying in that they may explain one aspect of Nav1.7 vs Nav1.8 targeting success.

      Weaknesses:

      (1) The use of postmortem human DRG neurons provides translational relevance, but the use of these cells is also a liability, given their high degree of variability. Of note are the 10 to 20-fold differences in baseline properties among cells, which dwarf the effects of the test compounds. The experiments may suffer from undersampling.

      We have added data from an additional 3 donors for the key results on increase in threshold and reduction of action potential upstroke, more than doubling the number of neurons for this data. We have also added a Supplementary Figure (Figure S2) that breaks out the effects on these parameters for each donor. This illustrates that there is a high degree of neuron-to-neuron variability in the effect of inhibiting Nav1.7 channels even within a single donor, even though we confined data to neurons that were verified to be capsaicin-sensitive. We also now note that there is a similar high degree of cell-to-cell variability in relative functional expression of Nav1.7 and Nav1.8 channels in capsaicin-sensitive mouse DRG neurons.

      (2) A potential confounder when using post-mortem human DRG neurons is heterogeneity of cell types. The methods clearly state that the cells selected for recording were of 'generally' small size, but specific criteria for what constitutes 'small' or other unstated selection criteria were not provided. A table of individual cell capacitance and input resistance values, along with information about individual donors (age, sex, ethnicity), is important to include. Additionally, some discussion of how DRG neuron heterogeneity impacts the findings. This relates to concern #1 about sample size determination and how cell heterogeneity factored into this calculation.

      We have added a figure (Figure S1) showing histograms and box plots of individual cell capacitance, input resistance, resting potentials, and maximum upstroke. We have also added a table with the information about donors (Table S1). As noted, we have also added a figure (Figure S2) that breaks out the effects on these parameters for each donor, illustrating that there is a high degree of neuron-to-neuron variability in the effects even within a single donor. We have added several sentences to the Discussion concerning the neuron-to-neuron variability in the effects of Nav1.7 inhibition, including the possibility that this may reflect heterogeneity of cell function.

      Reviewer #2 (Public review):

      Summary:

      The authors examine the functional role of Nav1.7 voltage-gated sodium channels in human sensory neuron electrogenesis using a Nav1.7 selective inhibitor and human dorsal root ganglion neurons obtained from organ donors. Patch-clamp electrophysiology is used at physiological temperature to measure the impact of Nav1.7 inhibition on sensory neurons' action potential firing. This is an important topic as Nav1.7 and Nav1.8 have been identified as therapeutic targets for the treatment of pain, but there has been mixed success with isoform-specific inhibitors in clinical trials. The data suggest that Nav1.7 and Nav1.8 have overlapping yet complementary functions in nociceptor neurons and that targeting both may be most effective for reducing nociception.

      Strengths:

      The data are of high quality. Action potential properties are measured at 37 degrees Celsius. Threshold is measured using brief pulses. The Nav1.7 inhibitor has been reported to be highly selective for Nav1.7 over Nav1.8 and moderately selective for Nav1.7 over Nav1.1 and Nav1.6. Data are collected using identical conditions and protocols to a previous study on the role of Nav1.8 in similar neurons.

      Weaknesses:

      The study relies on a single Nav1.7 inhibitor that has not been extensively characterized. One prior study indicates that the IC50 is around 140 nM, thus the 600 nM concentration used in this study could be predicted to reduce Nav1.7 currents by 80%. However, there is no voltage-clamp data in the current study to confirm this, and therefore, it is unclear if the batch of AM-2099 is as potent as reported in the paper that initially described its selectivity. The impact of Nav1.7 inhibition is compared to data from a previous study by this lab, and this is a minor concern. It would have been interesting to see if the combined inhibition of Nav1.7 and Nav1.8 completely blocked action potential generation in the human DRG neurons.

      We have done experiments to directly characterize the potency of the AM-2099 sample we used on both cloned human Nav1.7 channels and on native currents in the DRG neurons. Using a stable cell line expressing human Nav1.7 channels, we determined dose-response curves at both 22°C and 37°C, using an automated patch clamp instrument. These results are shown in a new Figure 1. Interestingly, we found that the IC<sub>50</sub> is substantially higher at 37°C than at room temperature. We also did experiments quantifying the effect of 100 nM and 600 nM AM-2099 on native sodium currents in the human DRG neurons, which align well with the results on the cloned Nav1.7 channels in suggesting that at 37°C, 600 nM AM-2099 inhibits Nav1.7 channels by about 85%.

      We have also added a new figure (Figure S3) showing the effects of a different Nav1.7 inhibitor, PF04856264. The effects of this inhibitor were qualitatively identical but quantitatively smaller than those of AM-2099. When we realized this, we did voltage clamp experiments on cloned Nav1.7 channels and discovered that the potency of PF-04856264 at 37°C was weaker than expected from the published IC<sub>50</sub>, which was determined at room temperature.

      We are currently doing experiments testing combined inhibition of Nav1.7 and Nav1.8 channels whenever we can obtain human neurons. Because it is of interest to examine effects of partial as well as full inhibition of each channel type, there are multiple permutations of inhibitors combined and alone that are of interest to characterize, and these studies are still in progress. We think the results in the present manuscript stand on their own and together with previous data on effects of Nav1.8 inhibitors alone provide a foundation for on-going and future studies on combinations of inhibitors by ourselves and others.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Fujita/Jo/Stewart/Osorno et al. investigate the contribution of Nav1.7 in regulating the excitability and firing properties of human dorsal root ganglion (hDRG) neurons in vitro. The authors characterize the effects of a previously reported Nav1.7-selective blocker AM-2099 in cultured hDRG neurons from postmortem organ donors. The authors observed modest changes in many of the properties expected by inhibiting Nav channels, including decreased action potential upstroke rate and amplitude, while increasing the voltage and current thresholds for spike generation. However, AM-2099 did not change the maximum number of APs in response to suprathreshold stimulation, leading the authors to conclude that Nav1.7 inhibition alone has limited efficacy in reducing the firing properties of hDRG neurons and that Nav1.7 blockers may have limited efficacy as analgesics. This is surprising, given that patients with loss-of-function mutations in Nav1.7 suffer from congenital insensitivity to pain. While it may indeed be true that pharmacological inhibition of Nav1.7 is unlikely to produce analgesia, the present study was limited to a single concentration of AM-2099. The manuscript would be significantly strengthened by a more careful and thorough pharmacological characterization of this compound, which has not been widely used or validated in native human DRG neurons.

      Strengths:

      Experiments are well-designed and executed, and the results presented are convincing. The focus on voltage-gated sodium channels in native human DRG neurons is highly relevant to recent efforts to develop safer analgesic options for chronic pain in people.

      Weaknesses:

      Only a single concentration of AM-2099 was used for all experiments. This compound was reported to be selective for cloned human Nav1.7 channels in heterologous systems, but has not been validated in other studies after the original publication in 2016. Since the original study reported a substantial statedependent block of recombinant Nav1.7 channels, more detailed pharmacological characterization of AM-2099 is needed in human DRG neurons to fully support these claims. This study would be significantly strengthened by the inclusion of dose-response curves to assess how much of the sodium current is inhibited at this concentration, confirming selectivity in hDRG, and whether maximal inhibition of Nav1.7 still has limited efficacy in reducing the firing of native human sensory neurons.

      We have added results from experiments to directly quantify the potency of AM-2099 on both cloned human Nav1.7 channels (new Figure 1) and on native currents in the DRG neurons (new Figure 2). These show that 600 nM AM-2099 produces about 85% inhibition of Nav1.7 channels at 37°C. We have added a paragraph to the Discussion explaining that we chose this concentration of AM-2099 to produce reasonably complete inhibition of Nav1.7 channels while minimizing potential inhibition of a component of non-Nav1.7 TTX-sensitive current.

      With regard to the broader point about reconciling the variable and sometimes relatively modest effects of Nav1.7 inhibition with the complete loss of pain sensation in humans with loss-of-function mutations, we have modified the Introduction and Discussion to eliminate any implication that the results in the manuscript suggest that pharmacological inhibition of Nav1.7 is unlikely to produce analgesia. Our experiments are only on action potential firing in the cell body, and it is perfectly possible that inhibiting Nav1.7 channels in the axon could disrupt generation or propagation of action potentials, either in the main axon or in the fine axon terminals in the spinal cord. We have modified the Discussion to explicitly point this out, which would reconcile the loss of pain sensation in humans with loss-offunction mutations with the incomplete effects of Nav1.7 inhibitors on excitability of cell bodies.

      Recommendations for the authors:

      Reviewing Editor comments:

      In addition to the points noted in the eLife assessment summary above, the study has several important strengths, including use of human primary neurons and recordings performed under physiologically relevant conditions (at 37 {degree sign}C using brief current injections). However, reviewers also identified several key weaknesses that must be addressed to support the central conclusions. In particular, multiple reviewers raised concerns regarding the lack of voltage-clamp data evaluating the efficacy and specificity of AM-2099 inhibition of Nav1.7 currents. A single dose of 600 nM was used based on the report of Marx (2016) in recombinant systems (Marx, 2016). Since no other studies other than the single Amgen report exist on this compound, it is important to validate its effects directly in the human DRGs used here. Additional concerns include the lack of dose-response analysis, as well as the large variability in baseline properties, which complicates the interpretation of the results. To assist the revision of the study, we outline below the key issues that should be addressed.

      Recommendations for authors:

      (1) Add voltage-clamp experiments to directly measure Nav1.7 current inhibition by AM-2099 in hDRG neurons. Given the limited previous characterization of this compound, it is important to confirm that the concentration used here (600 nM) effectively blocks Nav1.7 currents in the native system used here.

      (2) Related to point 1 above, perform a dose-response of AM-2099 on hDRGs on Nav1.7 currents in human DRGs. Since this study, at least in part, is framed as a comparative analysis of Nav1.7 vs Nav1.8 channel subtypes in DRGs, it seems important to establish pharmacological equivalence to ensure that the comparisons are made at functionally comparable levels of channel block.

      We have added results from experiments to directly characterize the potency of AM-2099 on both cloned human Nav1.7 channels and on native currents in the DRG neurons. Using a stable cell line expressing human Nav1.7 channels, we determined dose-response curves at both 22°C and 37°C, using an automated patch clamp instrument. These results are shown in a new Figure 1. Interestingly, we found that the IC<sup>50</sup> is substantially higher at 37°C than at room temperature. We also did experiments quantifying the effect of 100 nM and 600 nM AM-2099 on native sodium currents in the human DRG neurons, which align well with the results on the cloned Nav1.7 channels in suggesting that at 37°C, 600 nM AM-2099 inhibits Nav1.7 channels by about 85%.

      (3) Reviewer 1 notes that there seem to be 10-20-fold differences in baseline firing properties, which would exceed the effects of the test compound. This raises concerns about undersampling. Additional analysis or experiments would strengthen the conclusions.

      We have added data from an additional 3 donors for the key results on increase in threshold and reduction of action potential upstroke, more than doubling the number of neurons for this data. We have also added a Supplementary Figure that breaks out the effects on these parameters for each donor. This illustrates that there is a high degree of cell-to-cell variability in the effects even within a single donor, even though we confined data to neurons that were verified to be capsaicin-sensitive. Reviewer 1 made the excellent suggestion that because of the neuron-to-neuron variability in baseline properties, the effects of compounds could be better illustrated by displaying changes from baseline. Following this suggestion, we have added Tukey-style box plots displaying the data in this way. Together with the donor-to-donor breakout of data in the new Figure S2, these plots make it clear that the neuron-to-neuron variability reveals genuine differences in the channel make-up of each neuron and not experimental error.

      (4) Reviewer 2 notes an interesting experiment: does a combined block of Nav1.7 with the AM compound and Nav1.8 block action potential generation? If Nav1.7 controls threshold and Nav1.8 controls firing, then the combined inhibition should be highly effective in blocking nociceptive output, which could have therapeutic relevance.

      We are currently doing experiments testing combined inhibition of Nav1.7 and Nav1.8 channels whenever we can obtain human neurons. Because it is of interest to examine effects of partial as well as full inhibition of each channel type, there are multiple permutations of inhibitors combined and alone that are of interest to characterize, and these studies are still in progress. We think the results in the present manuscript stand on their own and together with previous data on effects of Nav1.8 inhibitors alone provide a foundation for ongoing and future studies on combinations of inhibitors by ourselves and others.

      Reviewer #1 (Recommendations for the authors):

      Concerns in addition to those in the Public Review:

      Major:

      (1) As per point 1 of the weaknesses in the Public Review, I'm concerned that the experiments suffer from undersampling. This requires a discussion of how the sample size was determined.

      We have added experiments from an additional 3 donors to the key results in Figures 3-5, more than doubling the number of neurons for these measurements.

      (3) The effect of compounds could be better displayed as a change from baseline in Figure 1C-E. Also, are the AP traces and phase plots shown in Figures 1AB and 2AB averages or representative?

      Thanks for this excellent suggestion. We have added box-plots that show changes from baseline for the various parameters. We have also clarified that the action potential traces and phase plots are from application of AM-2099 in a single representative neuron.

      Minor:

      (1) Provide source of VX-548 and report the purity of both compounds.

      We have provided the information for VX-548 and added the information on the purity of both compounds

      (2) Clinical failures of Nav1.7 blockers may not be solely due to pharmacodynamic limitations as implied by this study. Pharmacokinetic differences and toxicity (e.g., effects on the autonomic nervous system) may also have contributed.

      Thanks for raising this important point. We have added this point to the Introduction.

      Reviewer #2 (Recommendations for the authors):

      It is an interesting study, and the conclusions are reasonable. However, it would have been good to see validation of the potency of AM-2099 on native DRG sodium currents and/or recombinant human Nav1.7 channels expressed in a heterologous expression system.

      We have added results from experiments to directly quantify the potency of AM-2099 on both cloned human Nav1.7 channels (new Figure 1) and on native currents in the DRG neurons (new Figure 2).

      Minor comments:

      (1) Page 3, middle paragraph - there is a "(" missing before Renganathan.

      Thanks, corrected.

      (2) Page 4: Is anything known about AM-2099 in terms of state-dependence? It seems like Marx 2016 is the only previously published study using it, so additional information on the inhibitor would be helpful.

      We have not characterized the state-dependence of AM-2099, but we characterized its potency in voltage clamp using holding voltages similar to the average resting potentials of the cells in current clamp conditions.

      (3) Page 6 discusses that there might be differences between human and rodent DRG neurons in terms of Nav1.7 and Nav1.8. It would be nice if this were directly tested with these same Nav1.7 and Nav1.8 inhibitors.

      We have recently done such a study on mouse DRG neurons which has just been published (J Physiol. 604:6104-6127, doi: 10.1113/JP290574).

      (4) Figure 2A, right panel: I could not figure out the difference between the red and green traces. Perhaps this could be explained in the figure legend?

      Thank you for pointing out that this was confusing. These two traces showed two different subthreshold responses, one of which was slightly regenerative without generating a full-blown spike. We have simplified the figure by now showing only a single subthreshold response.

      Reviewer #3 (Recommendations for the authors):

      (1) The conclusion that Nav1.7 inhibition has limited efficacy for inhibiting the firing of human DRG neurons is not fully supported by the data. This may be true, but it cannot be concluded without a more thorough pharmacological characterization of this compound. Dose-response curves and experimental confirmation of Nav1.7 selectivity (maybe just total Nav current, TTX-sensitive and TTX-resistant components) are needed.

      We agree and have now added two new figures with this data.

      (2) How was the 600 nM concentration chosen? Given that AM-2099 was reported to exhibit state dependent block, how much of the Na current is inhibited by this concentration at the initial voltage used in current clamp experiments (~-80 mV)?

      We have added a paragraph to the Discussion recognizing the limitation that 600 nM AM-2099 produces ~85% rather than complete inhibition of Nav1.7 current and explaining that we chose this concentration of AM-2099 to produce reasonably complete inhibition of Nav1.7 channels while minimizing potential inhibition of a component of non-Nav1.7 TTX-sensitive current.

      (3) It appears that the effects of AM-2099 on the refractory period are bimodally distributed, where neurons that recovered more slowly at baseline were preferentially affected by AM-2099 (Figure 4). Do these reflect different neuronal populations (e.g. smaller or larger diameter DRG) or different resting voltages in these experiments?

      We agree that there seem to be two groups based on initial refractory period. Examining the parameters for the cells, there is no clear correlation between the effects of AM-2099 on the refractory period with resting potential or cell diameter. At this time, it is not obvious what determines the differences in refractory period. We speculate that neuron-to-neuron differences in the potassium conductances that generate the after hyperpolarization may be different in these cells but it will take further work to explore this.

      (4) How much of the sodium current is mediated by Nav1.7 in hDRG neurons? How does inhibition of both Nav1.7 and Nav1.8 affect hDRG excitability?

      The new Figure 2 shows data quantifying the AM-2099-sensitive current in the DRG neurons. With regard to combined Nav1.7 and Nav1.8 inhibition, we are currently doing experiments examining inhibition of excitability by combined Nav1.7 and Nav1.8 inhibition, which we agree is the logical next step in exploring how the two components of current control excitability. These are still in progress. Because designing and interpreting these experiments is facilitated by the current experiments with Nav1.7 inhibition alone, we believe that reporting the current results now will serve the scientific community better than waiting to obtain and interpret a body of data on dual inhibition in a sufficient number of donors, which we obtain only sporadically.

      (5) Donor information and soma diameters should be included. Capsaicin sensitivity testing was mentioned in the methods, but I was unable to find any inclusion of these data in the results. These may be useful to potentially infer effects in different cell types.

      We have added a figure (Figure S1) showing histograms and box plots of individual cell capacitance, input resistance, resting potentials, and maximum upstroke. We have also added a table with the information about donors (Table S1). We have now clarified that data were confined to cells verified to be capsaicin-sensitive and that ~95% of all cells tested were capsaicin-sensitive.

      (6) Please check statistical tests and reporting. Several graphs do not appear to have paired responses (e.g. Figure 1E, Figure 4B). As a result, two-tailed Wilcoxon tests would not be appropriate. Also, check reported p-values (e.g. p=.0002), which are identical for multiple panels in the Results section.

      We have clarified that the symbols of action potential width in control without a corresponding value after AM-2099 represent neurons in which the action potential in AM-2099 had a peak < 0 mV. These cells were not included in the data set of paired parameters used for the Wilcoxon test. We have also checked and verified all statistical tests.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Freas and Wystrach present a computational and experimental study of ant navigation. The main innovation of the computational model is the insertion of an oscillatory element between the steering signal and the motor control that results in a trajectory whose heading oscillates around a goal direction. Additionally, the model imposes periodic cessations of forward movement and inversely couples rotational speed to forward velocity. As a result the model periodically makes larger reorientations reminiscent of those seen in behaving ants.

      The behavioral data consists of two experimental sets: experienced Melophorus bagoti foragers, recorded in 2010 and inexperienced M. bagoti foragers, recorded in 2023-2024 at the same site. The behavioral data is qualitatively compared to the model in Figures 3 through 6. In figures 3-5, all ant sets are grouped together while in Figure 6 they are separated. In Figure 6, the authors should do a careful job of making sure the reader is aware that comparisons are being made between behavioral data sets captured more than a decade apart and of justifying the validity of a quantitative comparison between these sets.

      We now make explicit in the methods and figure that the two datasets were recorded at different times: experienced foragers from 2010 (Deeti et al., 2023) and inexperienced foragers from 2023. Their comparison is used to test a qualitative difference predicted by the model.

      The manuscript also describes Myrmecia ants and makes comparisons between modeled Myrmecia ants and supplemental videos of these ants (Videos 3,4). These videos are not described in the methods. While the captions describe these as ants "homing in an unfamiliar environment," the videos show tethered ants walking on a ball. Without more information and absent any analysis, it is difficult for me to understand how these videos support granular points in the text about coupling between rotation and forward velocities.

      We have added a description of Videos 3 and 4 to their respective captions and now state explicitly that these clips were recorded from ants on a tethered trackball apparatus and that they provided as qualitative examples of behaviour discussed in the text.

      Strengths:

      The manuscript's main thesis, that an oscillatory element interspersed between the control signal and the motor unit can reproduce aspects of ant navigation, appears supportable.

      Weaknesses:

      Qualitative agreement between aspects of a model and aspects of a behavioral measurement do not prove the correctness of a model. In the section (802), "An ancestral design? Striking parallels with crawling Drosophila larvae," the authors argue that behavioral data in larvae support their model, despite the larva's lack of a (known) central complex. C. elegans navigation can also be segmented into longer runs and shorter exploratory behaviors (Chen 2025), comparable to the runs and scans described here. C elegans definitively does not have a central complex. In general, multiple internal mechanisms are capable of producing the same macroscopic behavioral outcome. This fact limits the ability of behavioral data to confirm the details of a particular model; it does not imply that observation of similar behaviors in multiple species shows that a particular model is correct or generalizable.

      Here the ability of the behavioral data to confirm or constrain the model is further limited by the qualitative nature of the comparisons. Some of the comparisons are trivial (e.g. Figure 5E-F: any first order process will produce a Poisson distribution, and in the model a Poisson process was explicitly coded in with parameters chosen (1070) to match the behavioral data). Finally, the number of adjustable parameters (13) is comparable to the number of comparisons made; it is unclear that the model could not be adjusted to fit any set of behavioral measurements.

      Our model is a minimal neuro-mechanical model. It is not a mathematical model where each parameter can be optimised to a final output.

      From our 13 parameters, 10 were either taken directly from independent prior studies. 5 concern the oscillator, and have been arbitrarily chosen and simply need to produce regular oscillation (as explained in supplemental material). 4 are the necessary motor gain and noise, which scale neural activations values into movement units, note that this conversion is backed up by previous evidence in drosophila and present in previous ant models). 1 parameter specifies the width of the bump of activity in the CX, and is roughly matched to neural imaging data in flies. None of these parameters have been introduced or adjusted to back up our claim.

      Only 3 parameters have been added to produce scannings. Two of them were tuned to match local scanning-specific data (the probabilistic trigger to stop (p_stop) and the threshold for triggering a saccade (θ_CPG) enabling us to tune fixation duration). This level of parametrisation enables the model to reproduce realistic scans, but does not influence the qualitative predictions of this article. For instance, we agree that the probabilistic trigger producing the Poisson distribution of scan duration (Figure 5’s E) is used to parametrise the model to scanning data, which does not constitute an emerging prediction of the model. We do not include it as evidence (see ~490). Finally, the CX_output_gain, is a new parameter we invoke to implement our hypothesis that CX steering modulates the oscillator. From these three added parameters emerge the large array of behavioural signatures and predictions. These are emerging consequences of the model's architecture rather than curve-fits.

      While the introduction is improved, there is still room to eliminate confusion as to what aspects of the model reflect hypothesized rather than measured neural circuits. For instance, if there is data showing LAL oscillations in insects, the authors should cite it and call it out clearly.

      Alternatively they should say that the oscillator is hypothesized based on measured bistability. They should also clarify whether they are discussing neural oscillations or motor oscillations and whether these oscillations are measured, modeled, or hypothesized.

      As one example: Lines 283-284 "This oscillator [referring to the model's intrinsic oscillator described in the previous paragraph], which is widespread in insects (Cheng, 2024; Kanzaki, 2005; Kanzaki and Mishima, 1996), resides in the lateral accessory lobes (LAL)" reads as though it is known that a neural oscillator occupies the LAL. Cheng 2024 is a brief review of behavioral oscillation. Kanzaki et al. 2005 describes numerical modeling and simulation with a physical robot. Kanzaki and Mishima, 1996 demonstrates bistability (flip-flopping) in moth descending neurons. None of these show neural oscillations and none of them describe the LAL. The authors should review the paper and be scrupulously careful that the claims made in the text are supported in the cited references. These difficulties were pointed out in a previous round of review; hopefully they can be fully corrected this time.

      Kevin S. Chen, Jonathan W. Pillow*, Andrew M. Leifer*, "State-switching navigation strategies in C. elegans are beneficial for chemotaxis," arXiv:2508.00191 31 July 2025.

      We have softened the text to be more explicit about what is modelled versus what is neurally shown (labelled each as behavioural, electrophysiological, or modelled)

      Reviewer #2 (Public review):

      The paper by Freas and Wystrach is an interesting computational study, exploring the detailed mechanisms of how simple neural circuits could explain complex behavioral patterns observed in navigating ants. The authors compare detailed, high speed video recordings of Australian desert ants (Melophorus bagoti) with predictions made by their new computational model and find convincing similarities between the model and the behavioral data, at a level of detail not previously studied. Particularly interesting are emerging properties of the model, yielding behavioral motifs it was not designed to reproduce, but which occur in natural ant behavior.

      A strength of the study is that the model is based on previous models, without making major novel assumptions. It combines existing models of the insect central complex with a model of the lateral accessory lobe and adds a stochastic inhibition of forward velocity to the interaction of central complex and lateral accessory lobes. In essence, the central complex provides corrective steering signals when the goal direction and the current heading of the insect are not aligned, while the lateral accessory lobes provide an intrinsic oscillator underlying the behavioral oscillations shown by walking ants at all times. These background oscillations are modulated by the steering signals from the central complex. Depending on which phase of the intrinsic oscillations coincides with the corrective signals, and how fast the ant is moving forward during this time, a complex set of behaviors emerges.

      Most prominently, scanning behaviors, which are regularly carried out by the ants, are recapitulated in great detail by the model. Additionally, other behaviors, such as full loops, emerge naturally from the model. While computational models are not to be seen as definite evidence for any biological reality, they can provide strong support for particular neural implementations. The current study is an excellent example in that it provides evidence for a serial arrangement of central complex circuits upstream of the lateral accessory lobe circuits, modulated by speed regulating input. While the latter is hypothetical, it yields a clear hypothesis that can be validated by connectomics studies and functional work in the future.

      The computational model is explained in detail and information about all model parameters is provided in an accessible way. The approach is thus transparent and reproducible, leaving it to the readers to assess the assumptions made in the model and how the studied complex behaviors emerge. This also provides the possibility to combine this new model with existing models to expand the scope and to more comprehensively capture the behavioral repertoire of ants, and insects in general.

      Importantly, the study shows that even complex behavioral motifs do not require dedicated neural modules, but can rather emerge from the interplay of already known circuits - highlighting the efficiency of insect brains and possibly providing the path towards embodied hardware solutions of such circuits in autonomous agents.

      We thank Reviewer 2 for this assessment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The paper would benefit if the authors would take a more traditional/formal approach to the presentation and interpretation of results. They should avoid drawing conclusions or presenting interpretations in the figure captions (e.g. caption to figure 6) and avoid unnecessary modifiers (e.g. just use "supports" instead of "strongly supports"). This might help correct the tendency of the manuscript to overstate or over-interpret the correspondence between the model and the data.

      Figure captions are now revised to remove interpretive discussion. We also removed unnecessary modifiers throughout the manuscript.

      We also added clarifying text to the “An ancestral design?” section to make clear that we are not implying homologous neural implementations across taxa.

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed my comments fully and I only have a few minor, mostly editorial points:

      line 142: it appears that the references should refer to goal encoding, but both references are head direction papers (one review, one research paper). The only paper showing goal encoding in the CX is Mussels-Pires et al 2024. This should be fixed to ensure the citations are not misleading.

      Changed citations.

      line 165: maybe remove 'intrinsic' to not suggest that this reflects what it known from Biology? It is clear that in the model it is an intrinsic oscillator, but it should not leave the impression that this is an established fact for the LAL

      Removed when not discussing the model.

      line 168: remove either 'a diversity' or 'key qualitative'

      Changed to reproduce multiple key qualitative…. (~Line 170)

      section: 'Neural substrate of insect navigation':

      - '... compares to output steering commands.' grammar is misleading as 'output' might be read as an adjective rather than a verb, replace with 'generate'?

      Changed.

      - as above, the only paper showing goal directions is Mussels-Pires et al 2024. Any other paper either assumes this in models (such as Stone et al and the Honkanen review) or deal with head direction encoding. Pfeiffer and Homberg, 2014 is a general review. Please ensure that references are used more accurately.

      We revised the text to clarify the specific evidence provided by each cited reference.

      - The goal direction in the CX can be updated by various pathways....' After this, behavioral and modeling papers are cited, which is misleading. None of these papers deal with the CX or the neural representation of goals. Same with the rest of the sentence, referring to MB output and PI, which is only shown in modeling, not data.

      We now explicitly distinguish between behavioural evidence and modelling evidence.

      line 210: use CX, not central complex

      Changed.

      Figure 2: I suppose all data shown are modeling data? This should be more explicit in the figure caption (a bit unclear what 'using the neural circuit model.' in the caption heading means. Maybe rephrase to: 'Schematic of neural circuit model and its outputs across navigation relevant brain regions.' (or something like that, putting model first, not brain regions)

      Changed.

      Figure 7: Axis labels in the graphs are still much too small to be read on a printed version (ensure at least 5pt font size in the actual figure on the printed page)

      Enlarged axis labels.

      Line 654: mirror, not mirrors

      Changed.

      line 816: insert 'fly' before larva, as otherwise one might assume this refers to ant larva

      Added (~line 822).

      line 819: The CX does (as we currently know) not exist in fly larvae. At least not as a brain structure, or a set of homologous neurons. There might be equivalent circuits for action selection, but they have not yet been convincingly described. I suggest to rephrase to: 'Although no present as neuropils in fly larvae, the CX and LAL....'

      Changed as suggested (~Line 830).

    1. Author response:

      We thank the reviewers and editor for their thoughtful and constructive comments. Our goal was to connect theoretical work on habituation with empirical findings on intracellular habituation in Stentor coeruleus. We developed the phase-portrait analysis to provide a more formal way to examine habituation dynamics and to address the concern that apparent potentiation might simply reflect incomplete recovery. We are glad that the reviewers found this approach useful, and we hope to build on it in future work through mechanistic modeling.

      We will submit a revised version with the following changes:

      (1) We will include a supplementary figure that provides visual intuition for the habituation curves and phase portraits.

      (2) We agree that the anomalous behavior of the 2 min ISI / 1 hr ITI condition is noteworthy. This behavior arises from a subtle difference in the fitted shape of the trial 2 habituation curve: its Hill coefficient is less than 1, so the curve has no inflection point and its initial slope has nonzero magnitude. As a result, the corresponding phase portrait begins away from the x-axis, unlike the other conditions, whose Hill coefficients are greater than 1 and whose phase portraits are U-shaped. We will discuss this explicitly in the revision.

      (3) We will reconsider the layout to make the background and results easier to follow. In particular, we will consolidate the repeated material while preserving the context needed to interpret the theoretical consequences of the empirical findings.

      (4) We will revise the conclusions and discussion to leave claims about the decay of potentiation more open-ended.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Weaknesses:

      (1) The derivation of the main error term misses some important steps, which complicates peer review at this stage. In particular, factorisation of the covariance into noise and the inverse of the observation covariance matrix needs a more thorough justification. The cited sources do not contain the derivation for a noise term with full covariance, which is essential for deriving this error term.

      The derivation of the main error term misses some important steps, which complicates peer review at this stage

      We thank the reviewers for this careful observation. This concern is associated with the error term. Thus, we first clarified the noise assumption explicitly. We assume that ξ(t) is i.i.d. over time with zero mean and an arbitrary (not necessarily diagonal) positive definite covariance matrix .

      In particular, factorisation of the covariance into noise and the inverse of the observation covariance matrix needs a more thorough justification. The cited sources do not contain the derivation for a noise term with full covariance, which is essential for deriving this error term.

      The cited sources do contain the derivation for a noise term with full covariance. We have updated the citation that directly supports Eq. (S.2): Proposition 11.1 of Hamilton (1994, TimeSeries Analysis), which establishes the asymptotic distribution The proof, given in Appendix 11.A of Hamilton (1994), proceeds via a CLT for martingale difference sequences. See Theoretical Details of the Supplementary Materials.

      (2) The practical recommendation at the end of the paper also requires clearer guidance on how the design perturbations are constructed, and how many times and for how long the system is stimulated in each iteration of the experiment.

      Thank you for this helpful suggestion. We agree that the practical implementation of the experimental design should be explained more clearly. We have addressed this concern in two ways. First, we have revised the manuscript to explicitly describe the parameter design procedure. Second, we have revised the manuscript to clearly provide a reference to the detailed experimental condition table in the supplementary material. See Results - Main Manuscript.

      (3) Finally, there is no analysis of model mis-specification. In particular, the true dynamics are unlikely to be linear; the noise is unlikely to be either Gaussian or uncorrelated across time; and the B matrix is unlikely to be known perfectly. We’re not suggesting that the authors consider a more complex model, but it’s important to know how sensitive their method is to model mismatch. If nothing can be done analytically, then simulations would at least provide some kind of guide.

      We thank the reviewer for raising this important point regarding model mis-specification. We agree that it is important to run simulations to assess the impact of these model mismatches, therefore we conducted preliminary simulations to assess the sensitivity when two primary assumptions are violated: linear state dynamics and a perfectly known input matrix B. We added these preliminary results in the Supplementary Material, and revised the main manuscript to include a mention of these results. In summary, the simulations showed:

      The model estimation error increases with the strength of the nonlinearity; however, perturbation can increase the information and lead to accurate estimation of the hidden mode even under nonlinearity.

      The matrix A can be estimated roughly even if the assumed input matrix B differs from the true matrix in some cases.

      The matrices A and B can be estimated jointly without bias, provided the stimulation pattern excites the full state space.

      In such joint estimation, preferentially exciting the hidden modes directly leads to more accurate estimation of A than stimulating other modes. See Background and Results - Main Manuscript, Experimental Conditions and Results - Supplementary Material.

      Recommendations for the authors:

      (1) Please tell us what tACS, tDCS, and TMS are, and how much control the experimenter has over them. That’s important, because they are going to be used as control signals, so we need to know how accurately u(t) can be specified, and what its range is.

      Please tell us what tACS, tDCS, and TMS are, and how much control the experimenter has over them.

      We appreciate the reviewer’s helpful comment. We agree that it is important to describe these stimulation methods with appropriate references. We have added a new section titled “Neural Stimulation as Control Inputs” to the background, which connects our theoretical framework to practical experimental settings.

      We need to know how accurately u(t) can be specified, and what its range is.

      The accuracy and range of control inputs vary substantially depending on the specific stimulation technique and experimental setup, and a thorough discussion would require a dedicated review beyond the scope of this manuscript. Instead, we have added a sentence acknowledging the gap between practical experimental implementations and the theoretical formulation, and cited relevant references for readers interested in further details.

      Background - Main Manuscript

      “Neural Stimulation as Control Inputs

      This section describes how commonly used neural stimulation techniques can be related to input signals in control theory. Their adjustable parameters vary depending on how the stimulation inputs are modulated.

      Three non-invasive electrical stimulation methods illustrate how stimulation paradigms map onto basic control inputs. Transcranial magnetic stimulation (TMS) induces brief and transient perturbations via electromagnetic pulses [19], which are naturally represented as a sequence of impulse-like inputs, where the timing and intensity of each pulse are the primary controllable parameters. Transcranial direct current stimulation (tDCS) primarily modulates neural activity through approximately constant inputs [26], which can be viewed as a step-like signal whose main controllable parameter is the amplitude of the applied current. Transcranial alternating current stimulation (tACS) delivers oscillatory inputs [7], corresponding to sinusoidal signals characterized by amplitude, frequency, and phase. In control theory, impulse, step, and sinusoidal inputs are the basic components used to characterize system responses and dynamics [22, 21].

      The control input framework extends beyond non-invasive techniques to invasive and optogenetic stimulation. Invasive electrical stimulation, including intracranial microstimulation and deep brain stimulation (DBS), enables direct delivery of electrical inputs to neural tissue [17], providing flexible control over amplitude and timing through pulse trains or temporally structured waveforms. Optogenetic stimulation allows genetically targeted activation or inhibition of specific neurons using light [5], providing fine-grained control over multiple input dimensions, including amplitude (light intensity), temporal pattern, and cell-type specificity. In particular, recent developments enable stimulation at the level of individual neurons with high temporal precision [27, 16], allowing flexible construction of spatiotemporal input patterns.

      These stimulation examples demonstrate that the theoretical framework developed in this paper connects to practical experimental settings. While a substantial gap remains between idealized control inputs in theory and experimentally realizable stimulation, the core principles established in the following sections provide a foundation that naturally extends to these practical stimulation paradigms.”

      (2) Is the solid curve in the last panel of Figure 4b the prediction? If not, are the data points consistent with the prediction in Equation 8? This should be clear.

      Thank you for this question. The solid curve shows the empirical eigenvalues of the state covariance matrix, not the eigenvalues of matrix A. As shown in Equation 9, the estimation error is proportional to the inverse of the sum of the eigenvalues of the state covariance matrix. We have clarified these points by revising the main text and the figure captions. See Results - Main Manuscript.

      (3) Page 10, "Nodes 7 and 8 have only outgoing edges". It looks like node 7 has an incoming edge from node 8 (Fig. 6b). Or are we misinterpreting something?

      Thank you for pointing out this inconsistency. You are correct. Node 7 did have an incoming edge from Node 8 and contradicted the statement in the text. Moreover, we now think that two hub nodes (Node 7 and 8) are not necessary for demonstrating our primary theory. To resolve these issues, we have revised the simulation so that the network now contains a single hub node with only outgoing edges. Please refer to Fig. 6.

      (4) Figure 6f, g: why isn't the estimation error proportional to tr[Sig_x^{-1}]?

      We thank the reviewer for this observation. In the original manuscript, the estimation error in Figure 6f, g was plotted on a logarithmic scale, which obscured the proportional relationship with . The underlying values are indeed proportional, consistent with our theoretical prediction.

      In the revised manuscript, we have substantially reorganized Figure 6 to make the theoretical reasoning more transparent by using 1/µ instead of following Eq. 9. The sum across the column in panel (g) is proportional to the estimation error shown in panel (h).

      (5) Page 11, "The simulation was conducted with a single node receiving an impulse-shaped perturbation input." What’s an "impulse-shaped perturbation input"? A delta function? Please make this clear.

      Thank you for this clarifying question. We clarified the explanation as follows.

      Results - Main Manuscript

      “Each node was individually perturbed by an impulse input. Here, an impulse input is defined as a Kronecker delta at t = 0 with fixed amplitude α = 10, with no external input applied at any subsequent time step.”

      (6) Page 11, "Fig. 6d shows that the system possesses modes with small absolute eigenvalues." According to Figure 6d, all the absolute eigenvalues are between 9.9 and 9.98. So this statement appears not to be correct. Are we missing something?

      Thank you for pointing this out. You are correct. The previous statement was inconsistent with the figure. In the revised simulation, we have redesigned the network so that it clearly contains modes with distinct damping characteristics: heavily damped modes with absolute eigenvalues below 0.8 and lightly damped modes with absolute eigenvalues close to 1.0 (see revised Fig. 6c). This makes the relationship between mode damping and perturbation effectiveness much more transparent.

      Results - Main Manuscript

      “Figs. 6c shows the damping rates |λA| of the eigenvalues of A: the mode formed by Nodes 1 and 2 (15Hz) is heavily damped, while those formed by Nodes 3–6 are moderately damped. Node 7 serves as a hub with only outgoing edges.”

      (7) Page 11, "Crucially, the eigenvectors corresponding to these rapidly decaying modes (e.g., evec 7, evec 8) have their largest components concentrated at Nodes 7 and 8." If "nodes" are the same as "eigenvalue index", then the components on node 8 are zero (Figure 6e, bottom). In any case, it should be clear what you mean.

      Thank you for this comment. We agree that the previous description regarding eigenvectors was unclear and inconsistent with the figure. We now think that explaining with eigenvectors is not necessary for demonstrating our primary theory. We have therefore replaced the plots of eigenvectors with the reciprocals of the eigenvalues of Σ<sub>X</sub>, which are directly and rigorously explained by Equation (9). These reciprocals clearly show that the modes associated with Nodes 1 and 2 are heavily damped and applying perturbations to these nodes contribute to the estimation error and the perturbations to hub node (Node 7) broadly increases all eigenvalues of Σ<sub>X</sub>, thereby reducing the estimation error across all modes. This change ensures that all simulation results are grounded in the theory presented in the paper. Please refer to the revised Fig. 6d–g for details.

      (8) It’s not clear to us what’s plotted in Figure 6e. The real part of the eigenvectors? Which would explain why some of the eigenvectors are the same (e.g., 1 and 2). But that does not seem like a good idea, since the eigenvectors can be rotated by an arbitrary complex phase. Also, eigenvectors 7 and 8 seem totally opposite, and there’s no weight at all on index 6 and index 8. Could there be a mistake in the figure? In addition, Figure 6e is explained and interpreted in two different paragraphs, discussing the same observation. It would be easier to understand if they were moved to the same paragraph. Also, in the second paragraph where plot 6e is referenced (page 11, line 19), what does ’concentrated’ mean?

      Thank you for this comment. As described in our response to Comment (7), we have removed the eigenvector-based analysis from the revised simulation to prevent confusing readers. The revised results focus on quantities that are directly explained by Equation (9). Please refer to the revised Fig. 6 for the updated results.

      (9) Page 11, "This simulation demonstrates that, given a tentative connectivity matrix, an effective perturbation input (such as TMS or tDCS) can be designed by targeting the node with the highest weighted out-degree." "Demonstrates" seems strong. There is only one simulation, and that wasn’t totally convincing: a perturbation applied to node 7 did almost the same as a perturbation applied to nodes 1-6, even though it had a much higher outdegree than those nodes. It would be very helpful if you provided theoretical reasoning for why perturbing nodes with high-weight connections (Figure 6) minimises the prediction error. In particular, which properties of a hub node make it effective as a stimulation target? Why is it just its total output weights and not also its number of edges/centrality? How is the high output weight of a node related to its alignment with other eigenvectors, and is this always the case or just in this example?

      Demonstrates seems strong.

      Thank you for this important comment. We agree that "demonstrates" was too strong and have replaced it with "illustrates."

      It would be very helpful if you provided theoretical reasoning for why perturbing nodes with high-weight connections (Figure 6) minimises the prediction error.

      We agree with this comment. We have revised the simulation and accompanying text to connect the results directly to the main theory based on . As responded to Comments (7) and (8), we have removed the eigenvector-based analysis and instead plotted the reciprocals of the eigenvalues of Σ<sub>X</sub>, which are directly related to the estimation error via Equation (9). This change allows us to provide a clear theoretical explanation for why perturbing certain nodes minimizes the prediction error.

      Is this always the case or just in this example?

      We have also explicitly stated the limitations: whether a hub node or a specific subnetwork node is more effective depends on factors such as the outgoing edge weights from the hub and the individual damping rates of each mode. The revised text emphasizes that this simulation presents one example of perturbation location design, and the optimal strategy must be evaluated case by case using the theoretical framework of Equation (9).

      Results - Main Manuscript

      “As this simulation represents only one example of location design, its limitations and the corresponding countermeasures should be stated. In actual experiments, the most effective perturbation location depends on factors such as hub-node connectivity and modal damping rates, and B itself may not always be known a priori. In such cases, approaches such as iterative optimization of (described in a later section) and joint estimation of A and B (Supplementary Material B.3) provide systematic alternatives. Nevertheless, the results presented here provide an intuitive guideline: stimulation directed at hub nodes or at nodes driving heavily damped modes effectively excites the full set of dynamical modes and minimizes estimation error.”

      (10) What are the physical units for the time scales and stimulation amplitudes? For instance, on page 13, there is an argument that "In practical experiments, such long windows are unrealistic because neural states change rapidly over time." However, it is unclear whether T=100 a.u. or T=1000 a.u. etc. is realistic. By relating it to the eigenspectrum of A, which is supposed to be physiologically realistic, one can estimate the length of the stimulation window and support the above statement. Similarly, the impulse amplitude on page 13 is alpha=10<sup>20</sup> (a.u.). Also, Table 1 in the Supplementary has an extremely wide range of stimulation amplitudes. Is 10<sup>20</sup> a.u. a feasible amplitude in practice? And finally, please tell us which nodes the input was applied to.

      Thank you for your incisive comments. We have addressed each comment as follows.

      What are the physical units for the time scales and stimulation amplitudes?

      They don’t have physical units. This study is a theoretical investigation that focuses on the relative differences between passive observation and perturbation-based approaches, rather than providing precise predictions for specific experimental settings. The time scales and stimulation amplitudes are therefore expressed in arbitrary units.

      For instance, on page 13, there is an argument that "In practical experiments, such long windows are unrealistic because neural states change rapidly over time." However, it is unclear whether T=100 a.u. or T=1000 a.u. etc. is realistic.

      We agree that the original expression “unrealistic” was not appropriate given the arbitrary units. We have revised the text to clarify that the time scales are in arbitrary units and that the main point is about the relative difference in required data length between passive observation and perturbation-based approaches, rather than making an absolute claim about feasibility.

      Results - Main Manuscript

      “To obtain estimates under the passive condition that are comparable to those derived under perturbation, it is necessary to experimentally observe extensive time-series data. Figure 7f illustrates the LDA projection and classification accuracy for different time-series lengths. The leftmost LDA plot (T = 20) corresponds to the passive condition shown in Fig. 7c, indicating that the estimation performance in the passive condition becomes comparable to that in the perturbation condition only when the time window reaches approximately T = 200, a 10-fold increase compared to T = 20. While the absolute duration depends on the interpretation of the time unit, such time windows may not be prohibitive in some experimental settings. Nevertheless, our results consistently show that passive observation requires substantially longer recordings to achieve comparable performance, highlighting the efficiency of the perturbation-based approach when the available data length is limited.”

      Is 10<sup>20</sup> a.u. a feasible amplitude in practice?

      Although we have already stated that the stimulation amplitudes are in arbitrary units, we agree that the original value of 10<sup>20</sup> was excessively large and could be misleading. We have revised the impulse amplitude from 10<sup>20</sup> to 10<sup>2</sup>, as 10<sup>20</sup> is physically unrealistic—it would imply a stimulus intensity many orders of magnitude beyond any conceivable experimental setting. The revised value of 10<sup>2</sup> is more plausible: for reference, TMS stimulation voltages exceed typical EEG amplitudes by roughly 4–6 orders of magnitude. The classification accuracy decreased slightly; however, our main conclusion regarding the efficiency of the perturbation-based approach under limited data remains unchanged. See Table B.1.

      Finally, please tell us which nodes the input was applied to.

      The stimulus location was optimized to minimize the . We have clarified this procedure in the main text and added a visual indication of the selected stimulation site (red circles) in Fig.7.

      Results - Main Manuscript

      “The neural signals were simulated under five different task conditions and two stimulation conditions: passive observation and external perturbation. Perturbation was applied as impulse-type inputs, such as TMS. The stimulus location was determined for each task condition by applying an impulse to each node and selecting the one that minimized . The resulting time-series data are shown in Fig. 7b. Using this data, we estimated the underlying dynamical system via a control-based identification approach presented in Eq. 5, which corresponds to an estimation of functional connectivity.

      (11) Page 13: "The controlled transition test was run with T = 1." Previously, T referred to the number of time steps. Is that the case here? If so, that seems hard to justify. If not, please tell us what T is (and, ideally, use a different symbol).

      Is that the case here?

      No. In this context, T does not denote the number of time steps.

      If not, please tell us what T is (and, ideally, use a different symbol).

      Thank you for your helpful comment. We have standardized the notation throughout the manuscript. In this paper, T consistently denotes the number of time steps (i.e., data length). In the sections “Neural State Classification” and “Neural State Transitions,” we had mistakenly used T to refer to time length. To resolve this inconsistency, we have added a separate column labeled “Data Length (T)” to Table B.1 for clarification and replaced the previous usage of T with “data length” where appropriate.

      Results - Main Manuscript

      “To obtain estimates under the passive condition that are comparable to those derived under perturbation, it is necessary to experimentally observe extensive time-series data. Figure 7f illustrates the LDA projection and classification accuracy for different time-series lengths. The leftmost LDA plot (T = 20) corresponds to the passive condition shown in Fig. 7c, indicating that the estimation performance in the passive condition becomes comparable to that in the perturbation condition only when the time window reaches approximately T = 200, a 10-fold increase compared to T = 20. While the absolute duration depends on the interpretation of the time unit, such time windows may not be prohibitive in some experimental settings. Nevertheless, our results consistently show that passive observation requires substantially longer recordings to achieve comparable performance, highlighting the efficiency of the perturbation-based approach when the available data length is limited.”

      Results - Main Manuscript

      “The controlled transition test was run with T = 50. The control objective was to set nodes3 and 4 to 25 while keeping all other nodes at 0 without any movement. See Table B.1.”

      (12) Page 14: "where the matrix A (Fig. 10a) is designed to have 16 modes." What do you mean by has "16 nodes"?

      Thank you for pointing this out. By “16 modes,” we refer to 16 dynamical eigenmodes (i.e., 16 eigenvalue pairs). In the real-valued state-space representation used in the simulations, each complex conjugate pair corresponds to a 2-dimensional real block, resulting in a 32-dimensional system (32 nodes). We revised the wording to clearly distinguish between the number of dynamical modes and the dimensionality (number of nodes) of the state vector, to avoid confusion.

      Results - Main Manuscript

      “Iterative refinement of both the perturbation design and the estimation process progressively improves the accuracy of A. The time-series data is collected from 32 points, where the matrix A (Fig. 10a) is designed to have 16 oscillatory modes (i.e., 16 complex-conjugate eigenvalue pairs, yielding 32 eigenvalues in total).”

      (13) In Figure 10b, the y-axis should start at zero; otherwise, it’s a bit misleading how much the active perturbation helps. This will make it clear that the estimation error drops by about 33%. It would be worth commenting on whether this is typical; after all, potential users of this method would want to know how much improvement they’re likely to see.

      It would be worth commenting on whether this is typical; after all, potential users of this method would want to know how much improvement they’re likely to see.

      Thank you for this insightful comment. We agree with the reviewer that quantifying the expected improvement would be valuable for experimental practice. However, this simulation is a theoretical demonstration. Its primary purpose was to show that an iterative active perturbation approach can progressively converge to an optimal perturbation design even without prior knowledge of the true system, rather than to quantify a universally expected improvement rate.

      The y-axis should start at zero; otherwise, it’s a bit misleading how much the active perturbation helps.

      We believe that the y-axis should start at the accuracy with optimal perturbation because the primary purpose of this simulation was to demonstrate that an iterative active perturbation approach can progressively converge. If we had started the y-axis at zero, the message would be visually obscured.

      Potential users of this method would want to know how much improvement they’re likely to see.

      We acknowledge that this is one example of the application of our method, and the magnitude of improvement depends on various factors such as network structure, noise level, and stimulation design. We should not mislead the readers. We have therefore maintained the y-axis starting point and added a clarifying statement in the revised manuscript to indicate that the simulation is case-specific rather than universal.

      Results - Main Manuscript

      “These results should be interpreted as a case-specific illustration rather than a universal gain, as the magnitude of improvement depends on factors such as network structure, recording duration, noise level, and stimulation design. This simulation demonstrates that our perturbation design framework enables the step-by-step refinement of system identification even without prior knowledge of the system.”

      (14) Please provide a derivation for the factorised covariance in Equation S.2. This is the equation that underpins the main result of the paper - the error in dynamical system estimation. Currently, it appears to be taken from [Hamilton, J. D. Time Series Analysis], yet we were unable to find this result in the book. Most of the derivations in Chapters 8.1 and 8.2 assume diagonal noise covariance, and even isotropic noise (cov = sigma*2 * I), which simplifies the particular case of the derivations. Could the authors provide a reference or the derivation for the case with full covariance?

      As described in our response to the Weakness above, the relevant result is Proposition 11.1 of Hamilton (1994), not Chapters 8.1–8.2. Proposition 11.1 states the asymptotic distribution of the vectorised OLS estimator in a VAR model with i.i.d. innovations whose covariance matrix Ω is an arbitrary positive definite matrix. The factorisation in our Eq. (S.2) corresponds directly to the revised manuscript we explicitly cite “Hamilton (1994), Proposition 11.1” at Eq. (S.1) and in Hamilton’s Proposition 11.1, with and . In state the i.i.d. assumption on ξ(t) in the preceding paragraph, so that the connection to this result is unambiguous. See Theoretical Details Supplementary.

      (15) Strongly perturbing/Exciting fast-decaying modes to give them more ’runway’ and increase observed variability makes intuitive sense for a normal system. But what would happen in the case of non-normal dynamics, where stimulating one dimension only transiently amplifies it, but then excites other dimensions? Non-normality breaks the alignment between PC components and dynamic modes (see Kumar, Ankit, Loren M. Frank, and Kristofer E. Bouchard. "Identifying feedforward and feedback controllable subspaces of neural population dynamics." arXiv preprint arXiv:2408.05875 (2024)), so eigenvectors of A and Σ<sub>X</sub> won’t align for non-normal dynamics. This alignment appears to be a hidden assumption of this paper. Should the normal dynamics then be stated as an assumption/limitation of the framework? Is the example in Figure 6a highly non-normal? Does considering out-degree provide an empirical approach, an alternative to ’enlargement’, to dealing with non-normality?

      We thank the reviewer for this insightful comment regarding non-normal dynamics.

      Should normal dynamics be stated as an assumption/limitation of the framework?

      No. Our framework does not assume normal dynamics. For example, Figure 5 and the revised Figure 6 illustrate non-normal cases. In Figure 5, the row and column norms of A differ substantially (row norms ≈ [1.96, 4.05, 1.33, 4.70]; column norms ≈ [6.14, 1.94, 1.39, 0.87]). In Figure 6, Node 7 is a hub with only outgoing edges, so row 7 of A is zero while column 7 has nonzero entries. In addition, the Nodes 5–6 and Nodes 3–4 are strictly one-way, leaving the corresponding off-diagonal block upper-triangular. Both features break the symmetry required for normality.

      We additionally computed a commutator-based non-normality index which equals 0 for any normal matrix and approaches for the canonical maximally non-normal example (the 2×2 nilpotent Jordan block). The non-normality index is 1.34 for Figure 5 and 0.41 for Figure 6. These values confirm that both are clearly non-normal.

      Is the example in Figure 6a highly non-normal?

      Yes. As stated above, the network structure of Figure 6a clearly indicates non-normality.

      Does considering out-degree provide an empirical approach, an alternative to ‘enlargement’, to dealing with non-normality?

      No. In the previous manuscript, the results of out-degree and eigenvector structure were provided as supplementary intuition. However, the central contribution of our framework lies in maximizing the minimum eigenvalue µ of the observed state covariance (Equation 9), and for systems where non-normality is strong and subnetwork structure is less modular, the iterative optimization of provides a principled, assumption-free method. We clearly mentioned this point in the revised manuscript as follows. See Results in the main manuscript.

      (16) In the iterative experiment at the very end of the paper (Figure 10), what was the strategy for designing ‘u<sub>design</sub>´? How many stimulations were applied in an iteration? Are you stimulating along eigenvectors? Do you sample from components randomly, or perturb each of them individually, with the amplitude proportional to reciprocals?

      Thank you for this important question. We have clarified the design rule and stimulation protocol in the revised manuscript, and address each sub-question below.

      How many stimulations were applied in an iteration?

      One stimulation session was applied in an iteration.

      Are you stimulating along eigenvectors?

      The optimal stimulation is designed as the target node of the perturbation is determined through numerical optimization that minimizes

      Do you sample from components randomly, or perturb each of them individually, with the amplitude proportional to reciprocals?

      The optimal stimulation is designed as a composite-frequency sinusoidal input encompassing all modes of the estimated Â, and the target node of the perturbation is determined through numerical optimization that minimizes . See Results - Main Manuscript.

      (17) While it is clear that the proposed active method performs better than passive observation, some results lack a comparison with stimulating random directions/nodes with a comparable control energy (Figure 7e & Figure 10b).

      Thank you for your helpful comment. Although the optimally designed stimulation outperforms random stimulation, the previous stimulation settings were not configured to explicitly demonstrate this difference. Therefore, we modified the network structure, recording length, and stimulation intensity so that both the main messages and the difference from random stimulation can be shown simultaneously. Accordingly, we have added a comparison with random stimulation in both Figure 7e and Figure 10b.

      In Figure 7e, we included a random stimulation condition where the target node is selected randomly, and the results show that the optimized stimulation outperforms random stimulation.

      These additions strengthen the evidence for the effectiveness of our proposed method compared to non-optimized approaches.

      (18) Is it reasonable to assume full observability of the system? It would be interesting to consider biases arising from the partial observability of the system, in the spirit of Figure 9, which looked at partial controllability.

      We thank the reviewer for this suggestion. We agree that partial observability is an important consideration, however it is out of scope for the current work. Thus, we have added a future direction in the Discussion addressing partial observability. We note that Takens’ embedding theorem and Hankel DMD enable recovery of a system’s eigenvalues from partial observations, and since our framework relies on the eigenvalue structure of A, the proposed perturbation design remains applicable under partial observability.

      Discussion - Main Manuscript

      “Two directions warrant further investigation: extending the framework to partial observability, and validating it through stimulation experiments. In experimental neuroscience, recordings are often limited to a subset of neural populations, resulting in partial observability. A growing body of work has leveraged delay-embedding techniques, represented by Takens’ embedding theorem [25], to reconstruct hidden dynamics from partial observations [2, 3, 23, 11]. Applying such techniques enables the estimation of the full connectivity matrix, thereby extending our framework to settings with partial observability. The second direction concerns experimental validation. Validating a theoretical framework through experimental design is an essential in bridging the gap between theory and practice.”

      Recommendations for improving the writing and presentation.

      (1) The word ’state’ is overloaded in Figure 7. When talking about neural state classification, the ’state’ refers to a regime guided by a distinct dynamics A (should it be A<sub>i</sub>? Figure 7A bottom). However, each dynamical system also has a ’state’. A different word should be used in Figure 7A and the corresponding text.

      We agree that the terminology is potentially confusing. To avoid the confusion, we replaced the term “state” with “task condition” and “neural signal” in the main text, and revised Fig. 7.

      Results - Main Manuscript

      “We designed a neural network with clearly distinct task conditions and considered a simulation setting in which these conditions are classified using signals of a fixed duration. These distinct task conditions consist of five types, each defined by a unique linear dynamical system characterized by differing eigenvalue spectra and connectivity topologies of matrix A (Fig. 7a). These task conditions are intended to mimic different cognitive or behavioral contexts. For example, in a typical motor task experiment, such conditions could correspond to motor execution or imagery involving the left or right hand, or resting state [1, 24]. The neural signal was simulated under five different task conditions and two stimulation conditions: passive observation and external perturbation. Perturbation was applied as impulse-type inputs, such as TMS. The stimulus location was determined for each task condition by applying an impulse to each node and selecting the one that minimized . The resulting time-series data are shown in Fig. 7b. Using this data, we estimated the underlying dynamical system via a control-based identification approach presented in Eq. 5, which corresponds to an estimation of functional connectivity.”

      (2) It would be helpful to provide dimensions of matrices around Equation S.2, since vectorization makes dimensions hard to track.

      We thank the reviewer for this helpful suggestion. We agree that explicitly stating the matrix dimensions improves readability, particularly around the Kronecker product where vectorization can obscure the size of the resulting covariance matrix. We have revised the text as follows (the equation S.2 is 3 now). See Theoretical Details of the Supplementary Materials.

      (3) The background sections of the paper would benefit from referring to similar active perturbation methods: Wagenmaker, Andrew, et al. "Active learning of neural population dynamics using two-photon holographic optogenetics." Advances in Neural Information Processing Systems 37 (2024): 31659-31687. Minai, Yuki, et al. "MiSO: Optimizing brain stimulation to create neural activity states." Advances in Neural Information Processing Systems 37 (2024): 24126-24149.

      We thank the reviewer for these helpful suggestions. We have incorporated Wagenmaker et al. (2024) and Minai et al. (2024) into the Introduction. Their works focus on developing algorithmic approaches to active stimulation design for specific experimental platforms, while our work aims to establish a general theoretical framework for why and which perturbation inputs are effective has yet to be established. We cited these works and have clarified this distinction in the revised manuscript as follows.

      Introduction - Main Manuscript

      “In this paper, we propose a framework for designing the optimal perturbation input through control theory in neuroscience. We interpret neural dynamics as a control system [8, 6, 14, 12, 20, 24], and treat external perturbations as control inputs to design properties of neural stimulation (Fig. 1d). If the optimal perturbation input can be systematically designed, it becomes possible to steer the neural system toward states that are maximally informative (Fig. 1e), thereby enhancing the accuracy of the inferred connectivity (Fig. 1f). While recent studies have begun to develop algorithmic approaches to active stimulation design for specific experimental platforms [18, 28], a general theoretical framework for why and which perturbation inputs are effective has yet to be established. We first describe how to formulate neural dynamics as a control system and how to estimate the model parameters from observed data. Building upon this formulation, we derive a theoretical basis that enables us to design the optimal perturbation inputs for the neural system identification. We demonstrate the validity and utility of this theoretical basis by exploring its implications for optimizing parameters of common neurostimulation techniques and by applying it to practical examples, including neural state classification [4, 9, 1, 24] and control of neural states [8, 13, 14, 12]. In these demonstrations, we define concrete problems and apply the theory to validate its practical utility.”

      Minor corrections to the text and figures.

      (1) In Figure 3, it would be helpful to point out that the plots are in the subspace spanned by the first three principal components.

      Thank you for pointing out. We revised the caption of the Fig.3 as follows. See Results - Main Manuscript.

      (2) Figure 4 caption: "eigenvectors" –> "eigenvectors of Σ<sub>X</sub>", just to make it crystal clear (since A also has eigenvectors).

      Thank you for your suggestion. We have revised the caption of Figure 4. See Results - Main Manuscript.

      (3) Page 5, Equation (8): Capital xi should be introduced in the main text of the paper as the covariance matrix for the noise term xi. Now it can only be understood after reading the Supplementary. Or maybe call the covariance matrix Σ<sub>ξ</sub>rather than Σ<sub>ξ</sub>? That will probably make it clearer.

      Thank you for this suggestion. We have made both changes. First, we introduced the noise covariance matrix explicitly in the main text immediately after Eq. (1), defining. Second, we replaced the notation Σ<sub>ξ</sub> with Σ<sub>ξ</sub>(lowercase subscript matching the noise variable throughout the main text and supplementary.

      (4) Page 7, Equation (14): The description of the equation states "The covariance matrices of x(t) for impulse inputs can be written as:... ". We assume this is supposed to be the "state vector," not "covariance matrices".

      Thank you for catching this. We have corrected the wording to “state vector” instead of “covariance matrices". See Results - Main Manuscript

      (5) Page 10, after introducing Figures 6a-b, potentially a sentence is missing (’...’ in the first line of the last paragraph).

      Thank you for pointing this out. We have removed the placeholder along with updating the stimulation settings for Figure 6.

      (6) Page 10, Figure 6 (e): A clearer labelling would be helpful, e.g., a title for the legend (e.g. node index) and a more informative title (e.g. Eigenvector alignment with nodes), a caption (what is the take-home message?), and y-axis labels (what is the ’value’?).

      Thank you for this helpful suggestion. We have revised the Figure 6 taking your suggestion into account. The new figure includes a clearer title, axis labels, and an informative caption that highlights the key take-home message. See Results - Main Manuscript

      (7) Page 11, paragraph 1: The sentence "...(as discussed in Section)" is missing a section reference.

      Thank you for pointing this out. We have replaced the section reference placeholder with the equation reference to Equation (9), which is the relevant theoretical result.

      (8) Page 13, caption of Figure 8c: compputed -> computed.

      We have corrected this typo. We also reviewed the manuscript for any similar typographical errors and corrected them.

      (9) Page 14 Figure 9: The y-axis label for "Controlled State Process" plots is missing.

      Thank you for catching this. The y-axis was hidden by other elements in the figure. We have revised the figure layout to ensure that the y-axis label is visible.

      (10) Page 16: The sentence "As described in Section, a preliminary ... " is missing a section reference

      Thank you for pointing this out. We have corrected the missing reference. The sentence now reads “as shown in Fig. 10” rather than the incomplete “as described in Section.”

      (11) Page 16, Fig. 10c: It looks like the errors have inconsistent color ranges. A shared colorbar would help.

      Thank you for your suggestion. We have revised Figure 10c to use a shared colorbar across all subplots.

      (12) Page 17, Paragraph preceding eq. 16: "the eigenvectors of X" --> "the eigenvectors of Sigma_X".

      Thank you for catching this. We have revised the text. See Methods - Main Manuscript.

      (13) S.1: This is a GLS, not an OLS estimator, if this assumes full noise covariance.

      As stated in our response to the Weakness above, we assume the noise term ξ(t) is i.i.d. with covariance matrix Σ<sub>ξ</sub>, and the estimator we analyze is the OLS estimator. See Theoretical Details - Supplementary Material.

      (14) S.2: Unclear that⊗is the Kronecker product (not outer), as it is not defined.

      We thank the reviewer for pointing out this ambiguity. In the revised Supplementary Material, we have explicitly explained the Kronecker product with a reference to Hamilton [10, Appendix A.4, p. 732]. See Theoretical Details - Supplementary Material.

      (15) S.51: There is an accidental comma between alpha and A after "xdiff(t) ="

      Thank you for pointing this out. We have removed the accidental comma. See Theoretical Details - Supplementary Material.

      (16) Section B.2: It would be useful to have the simulation details for that section (as is given for the other section in B.1).

      We thank the reviewer for this helpful suggestion. We have added a dedicated “Simulation details” paragraph to Appendix B.5 (the section containing Fig. B.4) so that the setup is now described with the same level of specificity as the other appendix sections. See Experimental Conditions and Results - Supplementary Material.

      References

      (1) Irma N Angulo-Sherman, Marisol Rodríguez-Ugarte, Nadia Sciacca, Eduardo Iáñez, and José M Azorín. Effect of tDCS stimulation of motor cortex and cerebellum on EEG classification of motor imagery and sensorimotor band power. J. Neuroeng. Rehabil., 14(1):31, April 2017.

      (2) Hassan Arbabi and I Mezić. Computation of transient koopman spectrum using hankeldynamic mode decompoisition. APS, page G1.009, November 2017.

      (3) Steven L Brunton, Bingni W Brunton, Joshua L Proctor, Eurika Kaiser, and J Nathan Kutz. Chaos as an intermittently forced linear system. Nat. Commun., 8(1):19, May 2017.

      (4) Adenauer G Casali, Olivia Gosseries, Mario Rosanova, Mélanie Boly, Simone Sarasso, Karina R Casali, Silvia Casarotto, Marie-Aurélie Bruno, Steven Laureys, Giulio Tononi, and Marcello Massimini. A theoretically based index of consciousness independent of sensory processing and behavior. Sci. Transl. Med., 5(198), August 2013.

      (5) Karl Deisseroth. Optogenetics. Nat. Methods, 8(1):26–29, January 2011.

      (6) Shikuang Deng, Jingwei Li, B T Thomas Yeo, and Shi Gu. Control theory illustrates the energy efficiency in the dynamic reconfiguration of functional connectivity. Commun. Biol., 5(1):295, April 2022.

      (7) Shrey Grover, Renata Fayzullina, Breanna M Bullard, Victoria Levina, and Robert M G Reinhart. A meta-analysis suggests that tACS improves cognition in healthy, aging, and psychiatric populations. Sci. Transl. Med., 15(697):eabo2044, May 2023.

      (8) Shi Gu, Fabio Pasqualetti, Matthew Cieslak, Qawi K Telesford, Alfred B Yu, Ari E Kahn, John D Medaglia, Jean M Vettel, Michael B Miller, Scott T Grafton, and Danielle S Bassett. Controllability of structural brain networks. Nat. Commun., 6:8414, October 2015.

      (9) Mark Hallett, Riccardo Di Iorio, Paolo Maria Rossini, Jung E Park, Robert Chen, Pablo Celnik, Antonio P Strafella, Hideyuki Matsumoto, and Yoshikazu Ugawa. Contribution of transcranial magnetic stimulation to assessment of brain connectivity and networks. Clin. Neurophysiol., 128(11):2125–2139, November 2017.

      (10) James Douglas Hamilton. Time Series Analysis. Princeton University Press, Princeton, 1994.

      (11) Ann Huang, Mitchell Ostrow, Satpreet H Singh, Leo Kozachkov, Ila Fiete, and Kanaka Rajan. InputDSA: Demixing then comparing recurrent and externally driven dynamics. arXiv [q-bio.NC], November 2025.

      (12) Shunsuke Kamiya, Genji Kawakita, Shuntaro Sasai, Jun Kitazono, and Masafumi Oizumi. Optimal control costs of brain state transitions in linear stochastic systems. J. Neurosci., 43(2):270–281, January 2023.

      (13) Teresa M Karrer, Jason Z Kim, Jennifer Stiso, Ari E Kahn, Fabio Pasqualetti, Ute Habel, and Danielle S Bassett. A practical guide to methodological considerations in the controllability of structural brain networks. J. Neural Eng., 17(2):026031, April 2020.

      (14) Genji Kawakita, Shunsuke Kamiya, Shuntaro Sasai, Jun Kitazono, and Masafumi Oizumi. Quantifying brain state transition cost via schrödinger bridge. Netw. Neurosci., 6(1):118– 134, February 2022.

      (15) Hassan K Khalil. Nonlinear systems. Prentice-Hall, Upper Saddle River, NJ, 2002.

      (16) Paul K LaFosse, Zhishang Zhou, Jonathan F O’Rawe, Nina G Friedman, Victoria M Scott, Yanting Deng, and Mark H Histed. Single-cell optogenetics reveals attenuationby-suppression in visual cortical neurons. bioRxivorg, page 2023.09.13.557650, May 2024.

      (17) Andres M Lozano, Nir Lipsman, Hagai Bergman, Peter Brown, Stephan Chabardes, Jin Woo Chang, Keith Matthews, Cameron C McIntyre, Thomas E Schlaepfer, Michael Schulder, Yasin Temel, Jens Volkmann, and Joachim K Krauss. Deep brain stimulation: current challenges and future directions. Nat. Rev. Neurol., 15(3):148–160, March 2019.

      (18) Yuki Minai, Matthew Smith, Joana Soldado-Magraner, and Byron Yu. MiSO: Optimizing brain stimulation to create neural activity states. In A Globerson, L Mackey, D Belgrave, A Fan, U Paquet, J Tomczak, and C Zhang, editors, Advances in Neural Information Processing Systems 37, volume 37, pages 24126–24149, San Diego, California, USA, 2024. Neural Information Processing Systems Foundation, Inc. (NeurIPS).

      (19) Davide Momi, Zheng Wang, and John D Griffiths. TMS-evoked responses are driven by recurrent large-scale network dynamics. Elife, 12(e83232), April 2023.

      (20) Ali Moradi Amani, Amirhessam Tahmassebi, Andreas Stadlbauer, Uwe Meyer-Baese, Vincent Noblet, Frederic Blanc, Hagen Malberg, and Anke Meyer-Baese. Controllability of functional and structural brain networks. Complexity, 2024(1), January 2024.

      (21) Norman S Nise. Control Systems Engineering. John Wiley & Sons, 8 edition, 2020.

      (22) Katsuhiko Ogata. Modern Control Engineering. Prentice Hall, 2010.

      (23) Mitchell Ostrow, Adam Eisen, and Ila Fiete. Delay embedding theory of neural sequence models. arXiv [cs.LG], June 2024.

      (24) Yumi Shikauchi, Mitsuaki Takemi, Leo Tomasevic, Jun Kitazono, Hartwig R Siebner, and Masafumi Oizumi. Quantifying state-dependent control properties of brain dynamics from perturbation responses. J. Neurosci., page e0364252025, December 2025.

      (25) Floris Takens. Detecting strange attractors in turbulence. In David Rand and Lai-SangYoung, editors, Dynamical Systems and Turbulence, Warwick 1980, volume 898 of Lecture Notes in Mathematics, pages 366–381. Springer, Berlin, Heidelberg, 1981.

      (26) Liam C Tapsell, Matheus D Pinto, Ann-Maree Vallence, Casey Whife, Maria Luciana Perez Armendariz, Shaswat Senger, Jack Andringa-Bate, Dana Hince, and Myles C Murphy. What are the optimal transcranial direct current stimulation parameters and design elements to modulate corticospinal excitability? a systematic review and longitudinal meta-analysis. Neurol. Res. Pract., 7(1):86, November 2025.

      (27) Lei Tong, Shanshan Han, Yao Xue, Minggang Chen, Fuyi Chen, Wei Ke, Yousheng Shu, Ning Ding, Joerg Bewersdorf, Z Jimmy Zhou, Peng Yuan, and Jaime Grutzendler. Single cell in vivo optogenetic stimulation by two-photon excitation fluorescence transfer. iScience, 26(10):107857, October 2023.

      (28) Andrew Wagenmaker, Lu Mi, Marton Rozsa, Matthew S Bull, Karel Svoboda, Kayvon Daie, Matthew D Golub, and Kevin Jamieson. Active learning of neural population dynamics using two-photon holographic optogenetics. Adv. Neural Inf. Process. Syst., 37:31659–31687, 2024.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Summary of Revisions Performed:

      We have clarified the qPCR methodology in the methods section and stated the housekeeping gene GAPDH to address potential misunderstandings.

      We have assessed hearing in the generated HA-tagged mouse lines and included an adequately powered ABR measurements analysis in the revised manuscript as a supplemental figure.

      We have included powered DPOAE experiments in both ATP8B1 and TMEM30B KO mice to strengthen the findings of the ABRs.

      We have clarified the presentation of the z-stack in Figure 1F.

      We have elaborated on the analysis for Figure 7B to strengthen comprehension by readers.

      We have revised the statement to read: “No IHC stereocilia-enriched P4-ATPases were detected under the conditions examined.”

      While we appreciate the suggestion to examine TMEM30B localization on the ATP8B1 KO background, this is not feasible within a reasonable timeframe; we have clarified this limitation in the manuscript.

      We have incorporated relevant prior work (e.g., George and Ricci, 2026) demonstrating minimal Annexin V labeling prior to P6 and lack of PS externalization in TMC1/2 double knockout models.

      We have clarified that hearing thresholds for TMEM30B-HA and ATP8B1-HA lines were addressed in this study, while additional HA-tagged flippase lines (ATP8A1, ATP8A2, ATP11A) are part of ongoing work to be reported separately.

      We have softened statements regarding HA-tag insertion and clarified that, to our knowledge, localization and function are not disrupted, while acknowledging this as a potential limitation.

      We have revised the Methods section to clarify differences in fluorescence measurements across experiments.

      Public Reviews:

      Reviewer #1 (Public review):

      Figure1D.

      The authors should clarify how the qPCR data were normalized and specify the reference (housekeeping) genes used. This information is necessary to evaluate the robustness and comparability of the gene expression data.

      We thank the reviewer for this comment. qPCR data were normalized to GAPDH as the reference (housekeeping) gene. We have clarified this in the Methods section to ensure transparency and reproducibility.

      (2) Figure 1F.

      The lack of F-actin staining at the hair cell base raises the possibility that the permeabilization conditions may have limited antibody access to certain membrane regions. This is especially important given that the authors used a gentle permeabilization agent such as saponin to preserve membrane integrity. Because the authors conclude that ATP8B1 and TMEM30B are localized "almost exclusively to OHC bundles and the apical membrane, with minimal staining in the remaining plasma membrane," (line 128). Including co-labeling with a plasma membrane marker or more comprehensive F-actin visualization of lateral and basal regions would help ensure that the restricted localization is biological rather than technical. In the absence of such controls, the localization claim may be somewhat overstated and should be tempered accordingly.

      We thank the reviewer for this important point. The image shown represents a single z-slice from a larger stack, and the hair cell body lies outside the plane of this section. To clarify this, we revised the accompanying text.

      (3) Figure 7B.

      Although quantification of ATP8B1-HA intensity at the bundle appears similar between WT and Cib2 KO samples, the representative image suggests that some bundles lack detectable labeling. To better capture phenotype variability, it would be helpful to include an additional quantification showing the fraction or number of bundles with detectable ATP8B1-HA signal in Cib2 KO mice.

      We thank the reviewer for this suggestion. We have clarified the quantification of the fraction of hair cell bundles with detectable ATP8B1-HA and TMEM30B-HA signal per field of view. Although the representative images may give the impression that some hair bundles lack staining, this is due to changes in ATP8B1-HA and TMEM30B-HA distribution within the cell body. In all cases, detectable ATP8B1-HA and TMEM30B-HA signal remained present in the hair bundles.

      (4) Lines 346-349

      The manuscript suggests that IHCs lack stereocilia-enriched P4-ATPases. However, this conclusion is not directly supported by the presented data. The authors should either provide supporting localization or expression data for other P4-ATPases or soften the statement to indicate that no stereocilia-enriched P4-ATPases were detected under the conditions examined.

      We agree with the reviewer and have revised this statement to read: “No IHC stereocilia-enriched P4-ATPases were detected under the conditions examined.”

      Recommendations:

      (5) The authors convincingly demonstrate that TMEM30B loss results in ATP8B1 mislocalization. While not essential to the central conclusions, examining TMEM30B localization in ATP8B1 KO hair cells would clarify whether this interdependence is reciprocal, as described for other P4-ATPase-CDC50 complexes.

      While we agree that this experiment would provide valuable information, performing it would require generation of a compound mouse line carrying both the TMEM30B-HA allele and the ATP8B1 knockout allele. This work is beyond the scope of the current revision and cannot be completed within a reasonable timeframe.

      (6) Lines 359-374. The discussion of Annexin V labeling is careful and balanced. This paragraph would benefit from referencing other studies that showed minimal Annexin V labeling in healthy P6 organ of Corti, reinforcing that robust PS externalization in the present study is pathological rather than developmental.

      We thank the reviewer for this suggestion and have incorporated relevant prior work, including George and Ricci (2026), which demonstrates minimal Annexin V labeling prior to P6 and further supports our interpretation.

      (7) Lines 392-399.

      The proposed feedback model linking MET activity and ATP8B1-TMEM30B localization is compelling. The discussion could be strengthened by noting that in TMC1/2 double knockout hair cells, PS externalization is not observed, consistent with the idea that flippase activity becomes critical specifically when scrambling occurs. The mislocalization observed in Cib2 KO hair cells further supports the coupling between TMC-mediated scrambling and flippase-mediated membrane restoration.

      We agree and have revised the text to include that TMC1/2 double knockout hair cells do not exhibit phosphatidylserine externalization, supporting the idea that flippase activity becomes critical in the context of scrambling.

      Reviewer #2 (Public review):

      Weaknesses:

      (1) Are the HA tags causing any functional issues? Function and localization of tagged proteins can sometimes be compromised. It would be good to know, for each knock-in model (TMEM30B, ATP8B1, ATP8A1, ATP8A2, and ATP11A), whether the HA-tagged protein is causing any issues with the mice and particularly with hearing (ABRs). Are these mice normal? Can they hear? These data are missing.

      We thank the reviewer for raising this important point. In this study, we focus on TMEM30B-HA and ATP8B1-HA mouse lines, while additional HA-tagged flippase lines (ATP8A1, ATP8A2, ATP11A) are part of ongoing work to be reported separately.

      Both TMEM30B-HA and ATP8B1-HA mice are viable and exhibit normal breeding and ageing. We have included adequately powered ABR measurements of both TMEM30B-HA and ATP8B1-HA which indicate wild-type–like hearing thresholds.

      (2) Following on the point above, is it possible that ATP8B1-HA is well localized, but localization for the other three flippases (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA) is compromised by the tag? Is this potential mislocalization causing any functional phenotypes? (ABRs of point 1). I find it surprising that there are flippases only in outer hair cells and only formed by ATP8B1. A possible explanation is that the tag is interfering with trafficking. If so, there should be a phenotype (ABRs), although this might be masked by redundancy among these flippases or caused by systemic issues (admittedly difficult to sort out). Given that this manuscript will likely become foundational, and that there is evidence that at least two of the other flippases are involved in hearing loss, it would be good to provide more information about the mice and HA-tagged proteins in the other knock-ins (ATP8A1-HA, ATP8A2-HA, and ATP11A-HA). Depending on the data available for the knock-ins, the authors may want to discuss these scenarios and soften the statement indicating that inner-hair cells may lack flippase activity altogether.

      We appreciate this concern. To our knowledge, the HA tag does not appear to disrupt localization or function of the tagged proteins. However, we agree that this cannot be fully excluded. We have therefore softened our conclusions about IHC flippases and clarified that additional flippases (ATP8A1, ATP8A2, ATP11A) are under investigation and will be described in a separate study.

      (3) Expression of ATP8B1 at P0 (Figure 1D), when there should not be protein in outer hair cells yet seems high. Does this mean that other cells in the cochlea also express ATP8B1? Is this a concern?

      We thank the reviewer for this observation. We interpret the elevated ATP8B1 transcript levels at P0 as reflecting transcription that precedes detectable protein accumulation in OHC stereocilia. While expression in other cochlear cell types cannot be excluded, we did not detect ATP8B1-HA immunolabeling outside hair cells in the knock-in model.

      (4) Fluorescence scales in Figure 6 B and D and Figure 7 B and D are very different. So are the values for WT. One would expect that the WT would be similar in all cases (at least within the same compartments), given that the methods section indicates that "All images were collected using identical acquisition parameters, including zoom and laser power, across genotypes". If WT shows such variability, how can we compare?

      We appreciate the need for clarification. Identical acquisition parameters were maintained within each experiment used for direct comparison (e.g., within a given panel). However, different panels (e.g., Figures 6B vs. 6D) were acquired on different days using different imaging settings.

      We have revised the Methods section to explicitly state this and clarify that comparisons are intended only within panels, not across experiments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 42: When discussing TMC similarity to TMEM16 scramblases, it may be helpful to mention that some TMEM16 family members (TMEM16A and B) function as ion channels, highlighting the dual ion/lipid functionality within the superfamily. The similarity to TMEM63/OSCA ion channels and lipid scramblases could also be noted. The fact that TMC, TMEM16, and TMEM63/OSCA belong to the same superfamily would provide a broader context. I also suggest referencing the work that initially suggested this relationship: (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0192851)

      We have included this citation and expanded on this discussion in the revised version.

      (2) Line 45: Consider including the recent Cryo-EM structure of CeTMC2 (PNAS, 2023), which provides structural insight into TMC-lipid interactions.

      We have included this citation in the revised version.

      (3) Line 62: For precision, consider removing "calcium-activated," as caspase-activated scramblases also disrupt membrane asymmetry.

      For precision, we have removed “calcium-activated” in the revised version.

      (4) Line 71: Clarify whether this refers to "fusion of membranes" or "cell fusion."

      We have clarified this statement to mean cell-cell fusion.

      (5) Line 85: Consider citing studies showing constitutive PS externalization in TMC1 mutant mouse models linked to deafness.

      We have added a citation to show that constitutive PS externalization is linked to deafness (Ballesteros and Swartz, 2022, and Beurg et al. 2025).

      (6) Line 93: TMEM30C is not discussed. A brief comment on its expression or relevance in hair cells would provide completeness.

      We have added a brief statement regarding TMEM30C and cited prior work describing its expression pattern (Osada et al. 2007).

      (7) Figure 1A: Use distinct colors for the P4-ATPase and CDC50 subunit rather than a rainbow scheme to improve clarity.

      We have retained the original color scheme in this panel.

      (8) Figures 3C-D and 5C-D: Increase legend symbol size for clarity. Update Y-axis labels to "Number of OHCs/100 μm" and "Number of IHCs/100 μm." Correct "um" to "μm."

      We changed the legend to improve the presentation of these panels to be more legible and changed the measurement to μm.

      (9) Figures 3F, 5F, 5H: Add scale bars.

      We have added scale bars to these figures.

      (10) Figure 7: The confocal images (A, C) show the bundle on top and cell body below, but the quantification (B, D) is in the opposite order. Reorganizing the panels for consistent orientation would improve clarity.

      We have reorganized the panels to improve clarity.

      (11) ABR measurements: Please specify the sex of the mice tested or clarify whether both sexes were included.

      We have included both male and female mice in this study as there were no differences in hearing function. We have added this clarification to the methods section under hearing tests.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 1A, the panels show CDC50. I would either change to TMEM30B or mention in the caption that TMEM30B is also known as CDC50 as labeled in the figure.

      We have changed CDC50 to TMEM30B.

      (2) Figures 1F and 1G are missing scale bars.

      We have added scale bars to these figures.

      (3) Figures 2 C, D, and 5 C, D - difficult to tell what's what in the legend. Perhaps make symbols larger in front of WT P17, KO P17, etc.?

      We changed the legend to improve the presentation of these panels.

      (4) Text under "TMEM30B is required for hearing and OHC maintenance". There is a difference in phenotype between the TMEM30B (Figure 5C) and ATP8B1 (Figure 3C) knockouts that is not discussed, as apical and middle cells seem to be okay. Should this be discussed?

      We appreciate this observation. We have elected not to expand the discussion of these regional differences because apical and middle hair cells also undergo degeneration at later ages (after P30), suggesting that the observed differences primarily reflect the timing of degeneration rather than distinct underlying mechanisms.

      (5) In the discussion text, under "Why do ATP8B1/TMEM30B-deficient OHCs die?", "Tmc1/2 or Cib2" should probably be "TMC1/2 or CIB2" or "Tmc1/2 or Cib2"

      We have changed this to read TMC1/2 or CIB2.

      (6) The methods section states "..., whereas non-significant comparisons are not shown." However, non-significant p values are shown in Figures 7B and D (bottom panels).

      We have removed the nonsignificant comparisons from Fig 7B and D to be consistent with the rest of the paper.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a study utilizing several types of analyses (computational modeling, neuronal cultures, rodent epilepsy model, and human intracranial multi-scale recordings) to address a highly relevant conceptual question: Are fast ripples (FRs) distinct pathological entities or largely emergent products of stochastic spike clustering? The results can potentially reshape current approaches to incorporating fast ripples into the epilepsy surgery evaluation.

      Strengths:

      The conceptualization of fast ripples as potentially arising by chance is highly novel and builds effectively on questions raised in prior studies that have never been satisfactorily resolved.

      The integration across biological scales and models is a major strength. The state dependency analysis provides additional, strong support. The methodology and statistical approaches used are thoughtfully presented and rigorously applied.

      In particular, this paper provides a strong response to the findings from Gliske et al, Nat Commun 2018. This study utilized long-term data analysis to uncover low rates of FRs detected from most recording sites, suggesting spurious detections, although FRs were concentrated within seizure onset areas.

      We fully agree with this comparison. Although we had already cited this paper, we now further emphasize this observation in the Discussion:

      “Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence.”

      Weaknesses:

      The authors clearly aimed to use a statistical rather than a mechanism-based approach in this work. However, the paper's framing of true fast ripples as oscillatory events with stochastic fast ripples considered as confounders does not take prior investigations into biological mechanisms, particularly prior studies that point to an important role for stochastic fast ripples in some contexts. Incorporating recognition of these mechanisms would strengthen the manuscript and provide a more complete and nuanced characterization.

      Some examples from the literature:

      Eissa et al, eNeuro 2016, a paper that closely parallels this manuscript but took a mechanistic rather than statistical approach, showed that fast ripples can arise from population paroxysmal depolarizations - a key feature of epileptiform discharges - as temporally clustered, jittered population firing, with FRs appearing in LFP or EEG due to summated postsynaptic potentials (which are slower than action potentials and can generate signals in the high gamma range).

      Foffani et al., 2007, Neuron, and Ibarz et al., 2010, J Neurosci, argue that FRs are pseudo-oscillations created by jittered neuronal populations in the setting of altered spike timing.

      Smith et al., 2020, Sci Rep, contrasts FR characteristics in different regimes, i.e., intact inhibition early in a seizure vs. implied collapse of inhibition after recruitment. Schlingloff et al., 2025, J Neurosci, reported analogous findings in an animal model.

      We agree with the reviewer that even stochastic events may be of biological importance and an increase in stochastic events will occur when there is an increase in synchronisation and excitability, two properties of pathological cortex. We also don’t disagree that FRs can occur as distinct entities, although our work indicates that most are due to chance.

      To address this point, we have clarified our claims in the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      In addition, we expand on these points at various junctures in the Discussion. In particular, we reiterate our assertion that FRs may still be a useful biomarker, but that their interpretation should be moderated to reflect the fact that they often occur by chance:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      The computational model and subtraction approach provide a strong case for the random emergence of clustered activity in the high gamma band, given its assumptions. However, any such modeling effort needs to account for inhibitory activity, including impaired inhibitory function that is expected in epileptic brain regions, which has a strong modulating effect on excitatory firing and is thought to play a significant role in FR generation.

      We appreciate the reviewer’s concerns, but we believe that the impact of inhibitory interneuron activity on excitatory firing rates and synchronisation is incorporated indirectly into our simulations, while keeping our model as parsimonious as possible by not directly incorporating interneuron activity into our simulations. We have addressed this point in the Methods section:

      “Varying synchrony allowed us to test the impact, on the network, of inhibitory cells, which have been shown to favour synchrony (Bocchio et al., 2024; Cobb et al., 1995).”

      The shuffling procedure aims to preserve the power spectrum but randomizes high frequency phase (>200 Hz). However, this procedure removes biologically meaningful spike timing correlations, as well as structured cross-frequency coupling. The subtraction method thus likely underestimates the incidence of structured "distinct" FRs, while perhaps overestimating "chance" FRs due to biologically infeasible activity, making the statement that most FRs are due to chance correlation too strong.

      We appreciate this concern, which is especially important given that our results depend crucially on the validity of our shuffling procedure (as described in the Discussion). To address this issue, we have implemented an additional shuffling algorithm that preserves cross-frequency coupling (see last section of the Results, especially Supplementary Fig. 10f). This method showed no qualitative difference, compared with other alternative methods presented in Supplementary Fig. 10. These new results are described in the Methods section:

      “Last, we also implemented a method based on wavelet-IAAFT with preservation of cross-frequency coupling, since fast ripples are typically locked to low-frequency phase (Sheybani et al., 2019). The code detects the highest phase-amplitude coupling (PAC) in the original signal between [300-6000 Hz] for amplitude and several low-frequency bands ranging from 2-20 Hz, bandwidth of 3 Hz. PAC is computed using the modulation index (Tort et al., 2008). Then, in the shuffled signal under construction and during convergence testing of PSD (see above), the PAC between high-frequency part of the signal (300-6000 Hz) and the identified low frequency for phase is normalized to that of the highest PAC identified earlier.”

      The kainate findings underscore this point: the increase in the number of FR detections could be, as the authors state, an increase in chance clustering due to increased network excitability generally. However, the likelihood of a parallel increase in pathological FRs cannot be ruled out, given likely pro-epileptic alterations in spike timing and circuit function.

      We appreciate the reviewer’s point but wish to re-emphasise our interpretation of these findings – that the observed increase in the incidence of FRs occurs as a result of increased network excitability/synchrony, secondary to the pathological mechanisms of epilepsy. We have updated the Discussion accordingly:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-duration FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      To further emphasise this important point, we have also updated the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      Reviewer #2 (Public review):

      Summary:

      This paper asks an important question that has not been discussed much in the extensive literature on the High Frequency Oscillations (HFOs) that have been extensively studied in patients with epilepsy and experimental models of epilepsy. The question is whether the Fast Ripples (FRs), the HFOs in the 250-500 Hz frequency band, represent a pathological phenomenon or represent a physiological phenomenon that occurs in the healthy brain but happens to be more frequent in epileptic tissue. It is an important question that has not been systematically addressed until now. The authors conclude, from very extensive simulations, from extensive experimental animal studies (the systemic kianate model of epilepsy in rats), and from a modest amount of human data, that FRs occur in healthy brains as a result of the chance occurrence of bursts of action potentials, and that in epileptic tissue, their frequency of occurrence is approximately 30% higher than what is expected by chance. They conclude that FRs are not a separate phenomenon of epileptic tissue. This finding is reinforced by the recent findings of FRs in experimental models of Alzheimer's disease.

      Strengths:

      This is a valuable study because it asks an important and original question and because it evaluates it from several angles (simulation, tissue culture, experimental animals, and human patients). The simulations and the analyses of real data are performed very carefully and with original and solidly documented approaches, using extensive simulations and extensive data sets in the cultured cell data and in the in vivo experiments. The paper is clearly written and well-illustrated.

      Weaknesses:

      I found only one serious weakness in this study, but it is one that is of importance. Although the original work on FRs was done in an experimental model of epilepsy, the field really became prominent when ripples and fast ripples were found first in microelectrode recordings of epileptic patients and then in the intracerebral EEG of such patients. Numerous studies have been performed since then, with a valuable meta-analysis including 700 patients (Wang Z, Guo J, van 't Klooster M, Hoogteijling S, Jacobs J, Zijlmans M. Prognostic Value of Complete Resection of the High-Frequency Oscillation Area in Intracranial EEG: A Systematic Review and Meta-Analysis. Neurology. 2024 May 14;102(9). Although the consensus at this point is that FRs are not the ideal and totally specific marker of epileptic tissue that many thought it could be, FRs are nevertheless much more frequent in epileptic tissue than in non-epileptic tissue and are a solid biomarker.

      We agree with the reviewer, and do not intend to challenge the role of FRs as a marker of the seizure-onset zone, and potentially the epileptogenic zone. Instead, the aim of this study was to address the question of whether FRs are generated by intrinsic pathological mechanisms, or whether they arise due to the chance co-occurrence of action potentials that follow different dynamics in epileptogenic parenchyma. We have updated the Discussion accordingly:

      “Importantly, we do not question the potential of FRs to delineate the seizure-onset zone. Instead, our results suggest that the observed increase in FRs within the epileptogenic zone is an emergent phenomenon – arising due to changes in secondary network properties such as excitability and synchronization, not as a direct result of some pathology that is specific to epilepsy. In addition, we show that long-durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      To further emphasise this important point, we have also updated the Abstract:

      “This work does not rule out FRs as potential indicators of epileptogenic tissue, but it does challenge prevailing assumptions about their generation and specificity. Their higher prevalence in epileptogenic tissue is likely primarily due to increased excitation and/or neural synchronization, rather than peculiar abnormalities in network behavior.”

      It is also well established that they are much more frequent in NREM sleep than in wakefulness, as reported in the original paper of Staba et al (Staba RJ, Wilson CL, Bragin A, Jhung D, Fried I, Engel J Jr. High-frequency oscillations recorded in human medial temporal lobe during sleep. Ann Neurol. 2004 Jul;56(1):108-15., not mentioned in this paper) and in the study of Bagshaw et al (2009). In this last paper, using SEEG in various brain regions, the average rate of FRs in NREM sleep is about 6 times that in wakefulness. In the paper by Staba, with microelectrodes in mesial temporal structures, it is about twice. As a separate issue, the paper of Fraucher et al (Frauscher B, von Ellenrieder N, Zelmann R, Rogers C, Nguyen DK, Kahane P, Dubeau F, Gotman J. High-Frequency Oscillations in the Normal Human Brain. Ann Neurol. 2018 Sep;84(3):374-385), which is not quoted, found that, in an extensive sample, non-epileptic human tissue sampled with SEEG generated extremely rare FRs (an average rate of 0.04/min/channel, i.e. 1 every 25 min).

      The results above are mentioned because they do not fit with the data provided in the present study: FRs are much more frequent in NREM sleep than in wakefulness in human epileptic patients, and they are much more frequent (not 30% more, but many hundreds of percent more) in epileptic tissue than in non-epileptic human tissue. The fundamental phenomenon of interest is, I believe, the FRs in epileptic patients. The animal experiments, tissue studies, and simulations are models to study the human phenomenon. With respect to the modulation by sleep and the differentiation between epileptic and non-epileptic tissue, it seems that the systems studied in this paper are not good models of the human condition. The human results presented in the study only reflect wakefulness recordings, which is not the condition in which most HFO studies have been done and in which most HFOs occur. The authors refer to the study of long-term fluctuations in HFO rates by Gliske et al. (2018) to say that one has to be careful with the results regarding sleep, for example, Bagshaw et al (2009), but the clear predominance in of HFOs in NREM sleep has been observed by many studies. The cautions regarding fluctuations over extended periods also apply to the awake human data analyzed in this study. The study's conclusions regarding the generation of FRs are therefore questionably applicable to the human condition. I do not dispute their validity for the models and situations in which they were studied.

      We looked at this in more detail. Our simulations were intended to test how the incidence of FRs can vary with different parameters of network activity (neuronal count, firing rate, synchronization). Indeed, since their incidence is known to vary across regions and within regions and across states, we wanted to test how FRs are controlled by different factors. As such, we do not wish to draw firm conclusions about the observed sleep-wake changes in FR incidence in rodents, and how it relates to humans – evidence shows that pathological FRs in rodents do not display state-specific preferential occurrence (Ewell et al., 2019). We have added new text to the Abstract and Discussion to emphasize this.

      Abstract:

      “Our simulations showed that chance aggregation can generate fast-ripples and that their incidence changes depending on brain state, an observation that we confirmed in our rodent data.”

      We acknowledge that previous publications have reported higher rates during sleep, although with shorter recordings than in our rodent recordings (Staba, 2004: one night; Bagshaw, 2009: 10 min; Frauscher, 2018: 20 min – only sleep recordings). We have rewritten the part of the Discussion on the effect of the sleep-wake cycle on FRs incidence:

      “In our rodent data, we were initially surprised to find a higher rate of FRs during wakefulness, which contrasts with previous reports in humans (Bagshaw et al., 2009; Staba et al., 2004). However, previous studies only indicate that physiological vs pathological FRs are more easily distinguished during NREM sleep (von Ellenrieder et al., 2016) and that their incidence varies during sleep (Von Ellenrieder et al., 2017), but in hours-long recordings, no differences in incidence have been reported in the mesial temporal lobe (Dümpelmann et al., 2015). Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence. Last, but not least, another report did not find a state-dependent expression of FRs in the kainate rat model of temporal lobe epilepsy (Ewell et al., 2019), thus indicating that the variability of FRs across sleep and wake is still an open question, at least in rodents. Hence, the main conclusion on the effect of sleep-wake transitions is that these transitions impact the likelihood of stochastic events, more than dictating the direction (increases vs decreases) of change. It also highlights that the specificity of FRs to epileptogenic parenchyma could vary across the sleep-wake cycle, which would be crucial in epileptology (Dimakopoulos et al., 2024; Roehri et al., 2018; Sheybani et al., 2019, 2018; Zijlmans et al., 2012, 2009). Hence, FRs reflect and are highly susceptible to changes in network excitability.”

      Reviewer #3 (Public review):

      Summary:

      An outstanding question in the field of high-frequency oscillations (HFOs) in the context of epilepsy is how these oscillations emerge, considering that they occur at such high frequencies, i.e., 250Hz, well above the firing ability of single neurons. One hypothesis that has been suggested in the past is that neurons that fire in an out-of-phase fashion, or rather at random intervals, may contribute to a spectrum of HFOs ranging from 250-500Hz that are observed in epilepsy. However, how possible it is that random action potentials could aggregate to the extent that they could give rise to HFOs in the so-called fast ripple (FRs) frequency range (>200 according to the authors) remains unclear. To test this hypothesis, they used computational modeling to randomly insert action potentials in a signal, and they found that this approach is sufficient to generate FRs. Some of the predictors of whether FRs could occur were neuronal count, firing rate, and synchronization. Besides computational modeling, they used different model systems to test whether that would be possible to be observed in neuronal cultures, in epileptic rats (intrahippocampal kainic acid model), and human data. Neuronal cultures treated with picrotoxin did not show evidence that FRs could be generated beyond chance aggregation of action potentials. They then asked whether synchronization and firing rate could play a role in the emergence of FRs. They found that changes in neural firing and synchronization, such as those occurring during differences phase of the sleep-wake cycle, could affect the number of FRs occurring by chance aggregation, with more FRs seen during periods of wakefulness, a result that they replicated in human data.

      The authors largely achieve their proposed aims of demonstrating that random neuronal firing can, in principle, generate FRs. Results from this study could influence current thinking around mechanisms generating FRs in epilepsy. The use of different computational approaches and model systems could offer new analytical methodologies for the study of FRs in the context of brain disease.

      Strengths:

      (1) The authors used a multi-level approach combining computational modeling with experimental datasets, including neuronal cultures, a rat model of temporal lobe epilepsy, and human data.

      (2) Identification of key parameters such as neuronal count, firing rate, synchronization, and brain state in observed incidence of FRs generated through random aggregation of neural firing.

      (3) Cross-species validation increases the likelihood of generalizability of the findings.

      Weaknesses:

      (1) Some of the simulated FRs appear short in duration and may not meet standard detection and definition criteria, potentially influencing validity.

      We thank the reviewer for raising this important concern. To address this issue, we quantified and compared the duration of FRs in original and shuffled rodent data. Consistent with the reviewer’s suspicions, we found that FRs in shuffled signals are shorter than FRs in original signals. This is important because it shows that: (i) a longer duration should be considered a core feature of genuine FRs; and (ii) depending on the basal duration of FRs, the shuffling procedure will lead to different ratios of genuine to stochastic FRs. We have updated the Results accordingly:

      “These findings demonstrate the challenge of identifying distinct FRs within a composite population of distinct and stochastic events. One parameter that could help disentangle these events is their duration. Indeed, one might expect stochastic events to be more likely to be short-lived, since the probability of consecutive APs continuing to co-occur across neurons decreases over time. Hence, we next compared the distribution of FR durations between original and shuffled rodent data and found that FRs in shuffled data are shorter than those in original data (Supplementary Fig. 9). This makes duration a key feature that could help identify distinctly generated FRs.”

      And Discussion accordingly:

      “In addition, we show that long durations FRs are more likely to be distinct oscillations than stochastic events; and so FR duration is a key parameter that should be considered in future studies.”

      (2) The neuronal culture approach does not directly test random insertion of action potentials, limiting interpretation.

      Neither the neuronal culture approach, the rat data or the human data directly test random insertion of action potentials. The insertion of random action potentials is only performed in the simulated data to test if FRs can arise from the chance insertion of action potentials. Once this was confirmed in the simulations, we then used the shuffling procedure in biological data to test if FRs are more frequent than expected by chance.

      (3) Sleep is treated as a homogeneous state in the rat dataset, without accounting for stage-specific differences in synchronization, which may affect the results and interpretation.

      We agree with the reviewer, but our primary aim was to answer the question of whether FRs can arise by chance. Although it was interesting to see that, in our longitudinal rodent data, the incidence of FRs varies across the sleep-wake cycle, any sleep-stage-specific changes are beyond the scope of this work.

      (4) The analyses conducted in human data lack direct comparison with sleep data.

      We agree that it would have been useful to investigate variations in the incidence of FRs across the sleep-wake cycle in human microelectrode recordings. Unfortunately, however, such sleep recordings were not available. Hence, while we cannot compare variations in FR incidence across brain states between humans and animal models, our conclusions that FRs arise mostly by the chance co-occurrence of action potentials still holds.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please indicate where corrections for multiple comparisons were used.

      P-values corrected for multiple comparisons are indicated by the accompanying phrase: “adjusted p-value”. We had previously omitted to mention this once in the Results, which we have now corrected.

      (2) Delta amplitude is likely sufficient for detecting sleep-wake transitions, but the beta/delta ratio is better supported in the literature. Do the results change if beta activity is incorporated?

      We have now computed the beta (15-40 Hz) to delta (0.5-4 Hz) ratio and find that this is closely correlated with delta across time. We have updated the Results accordingly:

      “Importantly, these findings were robust to the specific method used to detect FRs (Supplementary Fig. 4c) (Padmasola et al., 2024; Sheybani et al., 2019, 2018). Also our use of delta power to identify periods of presumed wakefulness and sleep was highly (negatively) correlated with an alternative method of using the beta-to-delta ratio across time (another marker of increased vigilance; (Fraigne et al., 2023), see Supplementary Fig. 4d).”

      Methods:

      “We further verified that delta power across time displayed similar fluctuations to beta (15-40 Hz) to delta power ratio, another marker of vigilance (Fraigne et al., 2023).”

      And we updated Supplementary Fig. 4d

      “(d) Beta to delta power ratio across time is superimposed over delta power across time. There is a strong (inverse) correlation between the two time-series (inset), which is confirmed by the correlation coefficient across animals (right).”

      (3) Figure 2's axis labeling with the 3D plots is hard to read.

      We have enlarged the font size.

      (4) The scaling of the histogram in Figure 3 is unclear.

      This was on omission. The scale has now been added to Figure 3.

      (5) There is a risk of overfitting in the regression model. Was cross-validation used?

      We have now repeated this analysis with cross-validation, without any qualitative impact on the results (e.g. the model still performs well above chance). We have updated the Methods:

      “To further confirm the performance of GBT, we used a cross-validation procedure where the GBT is trained on 80% of data and then tested on the 20% remaining. The procedure is repeated 1000 times and the r<sup>2</sup> is saved at each round. We repeated the analysis with randomization of the outputs across 1000 rounds and saved this null distribution r<sup>2</sup>. We then compared the performance against original data.”

      Legend of Fig. 3:

      “(e) Performance of the GBT classifier using cross-validation (training: 80% of data; test: 20% remaining) using original (orange) and shuffled (blue) data. The difference is significant (paired t-test, p<0.0001).”

      And Results:

      “Furthermore, using a cross-validation approach with 80% of the data as training set and the remaining 20% as the test set, we obtained a significantly higher explained variance than when outputs were shuffled across the 125,000 solution points (paired t-test, p<0.0001, Fig. 3e), […]”

      Reviewer #2 (Recommendations for the authors):

      Maybe I missed it, but I did not find the length of human data analyzed or how the sections were selected.

      Apologies for this omission. The methods have been updated accordingly:

      “Microwire signals were selected based on high signal-to-noise ratio, as reflected by the detection of ≥ 1 single unit. Duration of recordings was of (median, interquartile range) 10 min and 17 s [3-13 min] and number of electrodes per patient was 4.5 [2.75-8].”

      The authors use the term "virtual simulation", which I find odd. I think the simulation is very real in the sense that it simulates reality, and I do not understand how a simulation can be virtual.

      We have updated the manuscript accordingly.

      Reviewer #3 (Recommendations for the authors):

      Major Comments:

      (1) In Figure 1, the authors suggest that random insertion of action potentials in a signal is sufficient to yield FRs. However, the observed FRs shown in panel 1b (also in supplemental Figure 5) seem pretty short in duration and may not meet the mentioned criteria in methods that require at least 4 cycles and ".whose amplitude is 3 times that of the surrounding baseline..". Moreover, in panel 1b, it seems that the FR shows a candle-like appearance, which has often been associated with filtering of sharp transients. How did the authors validate that the detected FRs were "real" FRs?

      Given the very large amount of data, it was not possible to visually verify all FRs. However, FRs were detected with published methods (Roehri et al., 2016; Roehri et al., 2017; and Sheybani et al., 2018 for confirmation of 24-hour variability in rodents) that have subsequently been used in several publications.

      Regarding the candle-like appearance of the spectrogram, the Delphos algorithm precisely looks for isolated “islands” of increased power (see Roehri et al., 2018, Ann Neurol), thus excluding any candle-like appearance. Similarly, the detector in Sheybani et al. (2018) J Neurosci first detects candidate FRs but then excludes those that are associated with a peak in lower frequencies, thus also limiting the risk of detecting candle-like events.

      Regarding duration, we have compared the duration of FRs in original and shuffled rodent data and found that FRs in original signals are indeed longer. This makes duration a key feature to identify distinct FRs. We have updated the Results accordingly:

      “These findings demonstrate the challenge of identifying distinct FRs within a composite population of distinct and stochastic events. One parameter that could help disentangle these events is their duration. Indeed, one might expect stochastic events to be more likely to be short-lived, since the probability of consecutive APs continuing to co-occur across neurons decreases over time. Hence, we next compared the distribution of FR durations between original and shuffled rodent data and found that FRs in shuffled data are shorter than those in original data (Supplementary Fig. 9). This makes duration a key feature that could help identify distinctly generated FRs.”

      (2) In the context of neuronal cultures, it is unclear how it could be deducted that the result relates to chance incidence of action potentials considering that no random action potentials were inserted, but only random shuffling of the high frequency component of the signal was attempted "Hence, neural networks with limited complexity (Kim et al., 2020; Saglam-Metiner et al., 2024; Sanchez-Vives and McCormick, 2000; Timofeev and Chauvette) fail to generate FRs beyond that expected from the chance coincidence of APs, even after increasing network excitability."

      FRs arise from series of action potentials occurring at a delay corresponding to their oscillatory frequency (250-500 Hz). Simulations demonstrated that FRs can occur by chance. When the EEG is shuffled, the only FRs that remain are those occurring by chance, because those occurring as individual entities have been broken up. Hence, if the original EEG displays more FRs than the shuffled EEG, then it means that these additional FRs were generated as individual entities. We have improved the Results section to clarify this:

      “We hypothesized that if FRs arise purely from chance firing, then temporally shuffling these recordings while conserving their spectral properties (Supplementary Fig. 3) would disrupt any oscillatory structure, leaving only FRs that occur due to chance.] Any additional FRs in the original data, compared to the number of FRs in the shuffled EEG, should thus be assumed to be individual entities.”

      (3) In the rat dataset, sleep was treated rather homogenously, without accounting for the sleep stage that is characterized by different synchronization and firing. An analysis of different sleep stages would be valuable.

      Although we agree that it would be scientifically interesting, we believe that our claim – that the ratio of genuine to stochastic FRs changes across the sleep-wake cycle – would hold. Unfortunately, lack of EMG prevents us from performing reliable sleep scoring. However, we do now include an alternative method for differentiating sleep from wake using the beta-to-delta ratio, which was highly correlated with delta activity, supporting our previous approach. Please refer to Supplementary Fig. 4d for further information.

      (4) The authors found that chance aggregation was highest during periods of wakefulness. Analyses of human data also confirmed that FRs could occur by chance aggregation during wakefulness. However, a comparison with sleep data would further strengthen this finding.

      We fully agree, but unfortunately, we do not have sleep data using microwires. Although our central claim – that FRs can occur by chance clustering of action potentials – would hold, we agree that it would have been scientifically interesting to add sleep data.

      (5) The statistics section would benefit from addressing how normality was determined and power analysis, as well as the inclusion of the exact sample size for all experiments.

      With large sample sizes, ANOVA and linear mixed models are robust to non-normality. Given the large sample sizes of our data, we thus used ANOVA and linear mixed model. For tests with small sample sizes where normality was violated, we used non-parametric tests, indicated by their name, e.g., Wilcoxon test for Supplementary Fig. 3b.

      (6) Greater discussion on the implications of this study for proposed in-phase or out-of-phase FR generation mechanisms is suggested.

      We have added further discussion on this. In the aim to keep the Discussion short and impactful, we could not elaborate too much. We have synthetized other parts of the Discussion to keep it within the right length. Here is the additional part:

      “It has been argued that the very high frequency that can be obtained during FRs are due to out-of-phase firing of excitatory neurons (Foffani et al., 2007; Ibarz et al., 2010), which is also consistent with our concept of stochastic firing. The conceptual difference is the degree to which there is any underlying organization of this firing. We argue that in the majority of cases there is no organization, although a substantial minority cannot be explained on a stochastic basis.”

      (7) More explanation around why wakefulness may drive chance aggregation and the clinical relevance of it, as often presurgical epilepsy recordings are being evaluated during sleep.

      We have profoundly rewritten the Discussion regarding the effect of the sleep-wake cycle on FRs incidence:

      “In our rodent data, we were initially surprised to find a higher rate of FRs during wakefulness, which contrasts with previous reports in humans (Bagshaw et al., 2009; Staba et al., 2004). However, previous studies only indicate that physiological vs pathological FRs are more easily distinguished during NREM sleep (von Ellenrieder et al., 2016) and that their incidence varies during sleep (Von Ellenrieder et al., 2017), but in hours-long recordings, no differences in incidence have been reported in the mesial temporal lobe (Dümpelmann et al., 2015). Furthermore, the variability of FRs across time (Gliske et al., 2018) indicates that longer nocturnal recordings in humans are necessary. It also suggests that changes in excitability across time could explain this change in FR incidence. Last, but not least, another report did not find a state-dependent expression of FRs in the kainate rat model of temporal lobe epilepsy (Ewell et al., 2019), thus indicating that the variability of FRs across sleep and wake is still an open question, at least in rodents. Hence, the main conclusion on the effect of sleep-wake transitions is that these transitions impact the likelihood of stochastic events, more than dictating the direction (increases vs decreases) of change. It also highlights that the specificity of FRs to epileptogenic parenchyma could vary across the sleep-wake cycle, which would be crucial in epileptology (Dimakopoulos et al., 2024; Roehri et al., 2018; Sheybani et al., 2019, 2018; Zijlmans et al., 2012, 2009).”

      Minor Comments:

      (1) Abstract, please include the frequency range of fast ripples explored in this study.

      The abstract has been updated accordingly.

      (2) Abstract, consider including the exact epilepsy model system in rats instead of "a rodent model of hippocampal epilepsy".

      The abstract has been updated accordingly.

      (3) Line 87, while Ylinen uses the term "high frequency oscillations" to refer to ripples up to 200Hz, which are different from the ones discussed here, better to rephrase or use another reference.

      The reference has been changed for Bragin et al. (1999), Epilepsia

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Ghosh and colleagues investigates the transcriptional changes within the oligodendrocyte lineage that contribute to age-related declines in oligodendrocyte differentiation and myelination. Combining bulk RNA-Seq on acutely purified oligodendrocyte lineage cells with bioinformatic approaches, the authors identify groups of genes that show different patterns of dynamic regulation during differentiation (which they term "switch" genes, or "switches"). A subset of these switch genes is differentially regulated with age. The authors identify two transcription factors, Bcl11a and Foxm1, that are downregulated during differentiation, have predicted binding site enrichment at other switch genes, and are downregulated in aged OPCs. Functionally testing Bcl11a, the authors show that Bcl11a knockdown inhibits the differentiation of young OPCs in culture, whereas overexpression promotes the differentiation of aged OPCs. Viral expression of Bcl11a in Sox10-expressing cells accelerates the formation of Plp1+ oligodendrocytes in aged rodents following lysolecithin induced demyelination.

      Strengths:

      The work is clearly presented and addresses an important biological problem. The bioinformatic approaches used in the manuscript are powerful, and the identification of Bcl11a as a modulator of oligodendrocyte differentiation is a novel finding. The combined in vitro and in vivo approaches to assess the function of Bcl11a in oligodendrocyte differentiation are a substantial strength of the work.

      We sincerely thank the reviewer for their positive assessment and for recognising the significance of our study, as well as the bioinformatics approach and tool developed as part of this work.

      Weaknesses:

      Although the PCA plots show distinct and reproducible global gene expression differences between the different isolated cell populations, the authors do not present a figure showing expression levels of typical stage-specific markers (e.g., Pdgfra, Pcdh15, C1ql1 for OPCs, Bcas1, Enpp6, Gpr17 for preOLs, Mobp, Mog, etc. for OLs) or confirm the absence of markers of other lineages (astrocytes, neurons, microglia, etc.). This makes it difficult to evaluate the success of their cell isolation strategy at different ages without reanalyzing the raw data.

      Thank you for this suggestion. We have presented markers expression in a new figure (Supplementary Figure 1) and included a description in the new Supplementary text.

      We observed elevated expression of Hes1 in OPCs as compared to both PreOL and OL, consistent with its role as a Notch effector that maintains the OPC progenitor state and inhibits oligodendrocyte maturation (PMID: 19104146, PMID: 21167918).

      Compared with PreOLs, adult OPCs isolated from 2–3-month-old rats did not show higher RNA expression of canonical OPC markers: Pdgfra, Pcdh15, and C1ql1. However, as expected, OPCs expressed higher levels of these markers than mature OLs.

      One possible explanation is the intrinsic heterogeneity of adult OPC populations. Adult OPCs exist in multiple transcriptional states, including quiescent-like and differentiation-primed states. During early differentiation, OPC markers such as Pdgfra are not immediately extinguished, and PreOLs may transiently retain these transcripts. The PreOL population captured in our study represents intermediate states transitioning from OPC to OL, potentially still carrying residual OPC-associated RNAs from activated OPCs. Therefore, comparing PreOLs with the total heterogeneous OPC pool, which includes quiescent-like OPCs, may give the appearance of higher canonical OPC marker expression in PreOLs.

      Among the PreOL-specific markers, Gpr17 clearly distinguished the PreOL state in our data, showing higher expression compared with both OPCs and OLs. Bcas1 and Enpp6 showed higher expression in PreOLs compared with OPCs. However, when PreOLs were compared with OLs, Bcas1 appeared to be lower in PreOLs, whereas Enpp6 expression remained largely unchanged.

      The OL markers Mobp and Mog showed significantly higher expression in OLs compared with OPCs, whereas their expression was not altered between OPCs and PreOLs. However, the canonical OL maturity marker Mbp showed a progressive and significant increase during differentiation, with expression levels clearly following the expected pattern OL > PreOL > OPC.

      We did not find any difference of astrocytes marker Gfap in those cell types comparison, suggesting similar level of unavoidable contamination which will not affect determination of differential gene expression. Regarding this please also see reviewer #2 major point 1.

      We now included this in the supplementary text:

      “Please see Supplementary Figure 1. We observed elevated expression of Hes1 in OPCs compared with both PreOLs and OLs, consistent with its role as a Notch effector that maintains the OPC progenitor state and inhibits oligodendrocyte maturation (Brosnan et al, 2009; Ogata et al., 2011).

      Compared with PreOLs, adult OPCs isolated from 2–3-month-old rats did not show higher RNA expression of canonical OPC markers: Pdgfra, Pcdh15, and C1ql1. However, as expected, OPCs expressed higher levels of these markers than mature OLs. One possible explanation is the intrinsic heterogeneity of adult OPC populations. Adult OPCs exist in multiple transcriptional states, including quiescent-like and differentiation-primed states. During early differentiation, OPC markers such as Pdgfra may not be immediately extinguished, and PreOLs may transiently retain these transcripts. The PreOL population captured in our study represents intermediate states transitioning from OPCs to OLs, potentially still carrying residual OPC-associated RNAs from activated OPCs. Therefore, comparison of PreOLs with the total heterogeneous OPC pool, which includes quiescent-like OPCs, may give the appearance of higher canonical OPC marker expression in PreOLs.

      Among the PreOL-specific markers, Gpr17 clearly distinguished the PreOL state in our data, showing higher expression compared with both OPCs and OLs. Bcas1 and Enpp6 showed higher expression in PreOLs compared with OPCs. However, when PreOLs were compared with OLs, Bcas1 appeared lower in PreOLs, whereas Enpp6 expression remained largely unchanged.

      The OL markers Mobp and Mog showed significantly higher expression in OLs compared with OPCs, whereas their expression was not altered between OPCs and PreOLs. In contrast, the canonical OL maturity marker Mbp showed a progressive and significant increase during differentiation, with expression levels clearly following the expected pattern: OL > PreOL > OPC.

      We did not detect any difference in the astrocyte marker Gfap across these cell-type comparisons, suggesting a similar level of unavoidable astrocytic contamination across groups. Therefore, such contamination is unlikely to confound the interpretation of differential gene expression among OPCs, PreOLs and OLs.”

      In the main text we have added the following text:

      “The expression patterns of cell-type-specific markers were consistent with their being distinct OPC, Pre-OL, and OL populations (Supplementary Figure 1, see Supplementary text for detailed description).”

      Please note that a detailed discussion of marker expression in the main text will disrupt the flow of the manuscript in manner we feel would detract from its clarity. We have therefore provided this discussion in the Supplementary Text.

      In addition, other publicly available datasets (e.g., the Barres lab bulk RNA-Seq datasets from PMID 25186741 or the Castelo-Branco lab single cell datasets from PMID 27284195) do not show downregulation of Bcl11a during OL differentiation as is described here - this apparent discrepancy is not discussed.

      Thank you for raising this point. We have now included new data as a Supplementary Figure 4. We performed RT-qPCR (reverse transcription followed by qPCR) to quantify Bcl11a expression and found that it was significantly lower in OLs than in OPCs, and significantly lower in aged OPCs than in young OPCs. These data were presented together with stage-specific markers.

      Regarding the comparison with PMID: 25186741: we extracted Bcl11a FPKM values from their dataset (GSE52564) and plotted, as shown in Author response image 1. We found that Bcl11a expression is downregulated during differentiation. However, the dataset contains only two replicates, and the SEM between the two OL replicates is very high, which may have contributed to the apparent lack of clarity. With such high SEM and only two replicates, the statistical power is poor, making robust statistical inference difficult.

      Author response image 1.

      Plotting of FPKM values of Bcl11a (obtained from GSE52564). mean+SEM shown along with individual data points. OPC: Oligodendrocytes progenitor cells, NFO: Newly formed oligodendrocytes, MO: myelinating oligodendrocytes. Dotted red line: linear regression line.

      Regarding comparison with PMID 27284195: we contacted the Castelo-Branco laboratory, and they kindly provided us with the analysis shown below in Author response table 1. This analysis showed that Bcl11a expression is lower in myelinating oligodendrocytes (MOLs) compared with OPCs. The apparent discrepancy observed in the web interface is likely because MOLs are displayed separately by subtype in the online resource. In single-cell datasets, particularly earlier pre-10x datasets with relatively lower cell numbers and sparser transcript detection, visualisations such as violin plots or t-SNE plots can be difficult to interpret when expression is distributed across multiple subclusters. Therefore, directly examining the differential expression statistics, including fold-change and significance values, provides a clearer and more quantitative assessment of the expression change.

      Author response table 1.

      Bcl11a expression difference in MOLs vs OPCs (dataset: GSE75330)

      FC: fold change, p_val_adj: adjusted p-value.

      Therefore, our bulk RNA-seq and RT-qPCR analyses presented in this paper are consistent with the Barres laboratory bulk RNA-seq dataset (PMID: 25186741) and the Castelo-Branco laboratory scRNAseq dataset (PMID: 27284195).

      Reviewer #2 (Public review):

      Aging poses a significant challenge to the regenerative capacity of oligodendrocyte precursor cells (OPCs) to differentiate and myelinate neuronal axons. Myelin abnormalities accumulate with age, and it is likely that the ability of OPCs to differentiate into myelinating oligodendrocytes becomes progressively impaired during aging, leading to inefficient turnover of damaged myelin and oligodendrocytes, as well as reduced adaptive myelination. Understanding the molecular mechanisms underlying the compromised capacity of aged OPCs is therefore critical for addressing age-related white matter decline.

      This study aims to decipher the intrinsic molecular changes that occur in aged OPCs. By profiling differentially expressed transcription factors (TFs) between young and aged OPCs, and by employing a novel bioinformatic tool to identify key TFs that undergo dynamic changes across distinct stages of OPC differentiation, the authors identify Bcl11a as a potential regulator. Bcl11a is highly expressed in young OPCs but markedly reduced in aged cells. Functional experiments further demonstrate that while Bcl11a does not affect OPC proliferation, it significantly promotes the differentiation of aged OPCs. Importantly, this effect is also observed in vivo following demyelinating injury in aged mice.

      While the study provides compelling evidence that BCL11A represents a limiting factor for OPC differentiation during ageing, the downstream targets and molecular mechanisms through which BCL11A exerts its effects are not directly addressed. As such, the work should be interpreted primarily as identifying a key regulatory node rather than a fully defined molecular pathway.

      Overall, this study offers valuable insights into the age-related loss of regenerative capacity in the central nervous system and introduces a computational framework that may be broadly useful for investigating dynamic gene regulation in other biological contexts.

      We are grateful to the reviewer for their supportive comments and for highlighting the broader relevance of our computational framework beyond our specific subfield.

      Major Points:

      (1) MACS mouse anti-A2B5 microbeads are not OPC-specific and may also label astrocyte precursor cells or immature astrocytes. How do the authors justify this caveat? Could some of the claimed "OPCspecific" switch genes in fact be enriched in astrocyte lineage cells?

      We thank the reviewer for raising this important point. While anti-A2B5 is a well-established and widely used antibody for isolating OPCs, we nonetheless agree that no technique can isolate a specific cell type with 100% purity, and this also applies to OPC-specific isolation using a validated anti-A2B5 antibody.

      To check whether astrocyte contamination could be an issue in determining differential expression, and specifically whether the OPC population was affected by astrocyte contamination, we checked the relative expression and statistical significance of the astrocyte marker Gfap. We refer to our new Supplementary Figure 1 and Supplementary text. This suggests that no difference exists in Gfap levels when comparing OPC, PreOL and OL populations. Therefore, we contend that it is unlikely that the differential expression observed in any cell population is actually due to astrocyte contamination, or that the OPC population is selectively contaminated by astrocytes.

      We now included the following in the Supplementary text:

      “We did not detect any difference in the astrocyte marker Gfap across these cell-type comparisons, suggesting a similar level of unavoidable astrocytic contamination across groups. Therefore, such contamination is unlikely to confound the interpretation of differential gene expression among OPCs, PreOLs and OLs.”

      (2) Overall, Figures 1 and 2 are not very informative in terms of biological insight. The authors should provide more detail in the main figures regarding the enriched gene sets associated with each of the Type 1-4 switch categories. For example, summarizing the top Gene Ontology terms for each switch type would greatly enhance interpretability.

      We agree that GO analysis can add further interpretability. We have now prepared a new Supplementary Figure 3A to summarise the significant top GO-term enrichment for switch Types 1– 4, for which gSWITCH-identified patterns are presented in Figure 1C. We also prepared a Supplementary Figure 3B to summarise the top significant GO-term enrichment for the 135 Type 3 switch genes affected in ageing, presented in Figure 2C. Please note that only 8 Type 4 genes overlapped with differentially expressed genes in ageing. Due to this small number, we could not identify any significant GO-term enrichment, and therefore this was not plotted.

      (3) A similar issue applies to Figure 3. The authors should explicitly specify the transcription factors in the main figure, particularly the 27 TFs identified through theENCODE/ReMap2 analysis.

      Thank you for raising this point. We have now prepared a new Supplementary Table 3, where we list 27 TFs and highlight, with light grey shading, the 5 TFs that overlapped with Type 3 switches.

      (4) Have the authors validated Bcl11a expression across different CNS cell types and between young and aged conditions using independent methods such as qPCR, immunofluorescence, or western blotting?

      Thank you for this suggestion. We performed qPCR and presented this data in a new Supplementary Figure 4. We found that Bcl11a expression is lower in OLs than in OPCs (Supplementary Figure 4A). We also observed reduced Bcl11a expression in aged OPCs compared with young OPCs (Supplementary Figure 4B). (see also response to Reviewer 1’s recommendations).

      (5) Regarding OPC aging, an open question is whether the reduced differentiation capacity of aged OPCs is an intrinsic property of the cells themselves or whether it results from prolonged exposure to an aging environment that induces non-cell-autonomous epigenetic or genetic changes, thereby rendering OPCs less efficient at differentiating. It would be helpful if the authors could expand on this point in the Discussion, with reference to relevant previous studies and experimental evidence.

      We thank the reviewer for suggesting this important aspect be discussed. We have now included the following paragraph in the discussion section:

      “The extent to which the reduced differentiation capacity of aged OPCs is intrinsically encoded within the cells themselves or induced by prolonged exposure to an aged tissue environment is an interesting question. Based on our previous work, we favour the view that loss of OPC function is primarily determined extrinsically since various manipulations of the aged environment such as heterochronic parabiosis (Ruckh et al., 2012), fasting and calorie restriction mimetics (Neumann et al., 2019), and niche biomechanics (Segel et al., 2019) can all alter the cell-intrinsic state, reverting aged cells to a ‘youthful state’. Significantly, when aged OPCs are transplanted into the neonatal CNS they proliferate and differentiate as if they were neonatal OPCs (Segel et al., 2019). The reversion of aged OPCs to a functional state by changes in their external environment necessarily operates through changes in cell intrinsic function, suggesting that the same intrinsic mechanisms could be targeted directly to restore declining OPC function—for example through epigenetic regulation of differentiation inhibitors (Shen et al., 2008) or overexpression of transcriptional regulators such as c-Myc (Neumann et al., 2021, Dimas et al., 2025).”

      (6) Do the authors observe a change in the number or density of OPCs between young and aged mice?

      Thank you for asking this important question. In 2002 we reported that there was no difference in the OPCS density between young adult and old adult rats, at least in the deep cerebellar white matter (Sim et al. 2002 - PMID: 11923409). We also refer the reviewer to Figure S1 of another previous study, published in Cell Stem Cell in 2019 (PMID: 31585093). We did not find any difference in OPC number between young and aged brains. Quantification was performed using FACS, where freshly isolated cells were stained with A2B5 (OPC marker), CD11b (microglia marker), and MOG (oligodendrocyte marker). Thus, we do not find any evidence for an age-related decline in OPC densities.

      (7) The in vivo characterization of Bcl11a overexpression using the AAV-based approach appears incomplete. Do aged mice overexpressing Bcl11a in Sox10⁺ cells exhibit reduced age-related myelin degeneration under baseline conditions? In the LPC model, do the authors observe differences in lesion size and/or remyelination efficiency?

      Again, we thank the reviewer for raising these interesting points. To assess whether Bcl11a overexpression in Sox10+ myelinating oligodendrocytes exhibit less age-related myelin degeneration would, we suspect, require long-term experiments. For this to be the case would require a role for Bcl11a in myelin maintenance – and interesting question but one we feel (and hope the reviewer agrees) is beyond the scope of the current study. We do not see any difference in lesion size (and would not expect the expression of elevated levels of Bcl11a to protect against the membrane-solubilising effects of LPC) but do see changes in remyelination efficiency as shown in Figure 6.

      (8) Are the authors presenting gSWITCH for the first time in this manuscript? Given that the gSWITCH framework is novel and central to the study, its conceptual contribution could be emphasized more strongly. A brief comparison with existing trajectory- or pattern-based methods-ideally in the main text around Figure 1-would help readers better appreciate its novelty.

      We thank the reviewer for this important suggestion. Yes, gSWITCH is presented for the first time in this manuscript as a new computational framework and web application. We agree that its conceptual contribution should be made clearer in the main text itself, although we explained its concept in detail in ‘Materials and Methods’ and in the supplementary Figure 2 (which was Supplementary Figure 1 in first version of this manuscript).

      We now included the following paragraph in the manuscript:

      “Existing computational tools such as Monocle (Trapnell et al., 2014), tradeSeq (Van den Berge et al., 2020) and maSigPro (Nueda et al., 2014) are highly valuable for identifying genes with dynamic expression changes across pseudotime or time-course data. gSWITCH addresses a different question. It does not aim to infer trajectories. It works with a user-defined ordered series of biological states or time points and asks a more specific question — does this gene show a statistically supported "switchlike" change in expression as cells move through these states, and if so, what shape does that change take? It combines GLM-based statistical testing with criteria that capture where a gene reaches its highest or lowest expression and whether its expression changes steadily in one direction across the ordered series. To our knowledge, no existing tool combines significance testing with this type of explicit, shape-based classification into discrete, interpretable switch categories. gSWITCH sorts genes into four biologically meaningful patterns, rather than producing only a ranked list of significant genes based on pairwise comparisons between multiple conditions or states. gSWITCH also flags which of these switch genes are transcription factors, making it easier to prioritise candidates for follow-up experiments.

      This biologist-friendly tool is freely available as a web application requiring no programming, works with experimental designs containing three or more stages or time points with at least two replicates per stage (no upper limit on either), and can be applied to bulk RNA-seq or to single-cell RNA-seq data aggregated as pseudobulk.”

      (9) The evolutionary analysis also appears somewhat disconnected from the rest of the study. Could the authors leverage available public datasets to test whether a similar Bcl11a expression trajectory is observed in human oligodendrocyte lineage cells?

      We thank reviewer for mentioning this. We would like to clarify that the evolutionary analysis was included to examine whether Bcl11a sequences across vertebrates, including humans, show evidence of selective constraint, meaning that the sequence has been preserved during evolution because changes in it are likely to be disadvantageous. This analysis was therefore intended to provide broader evolutionary support for the functional importance of Bcl11a, rather than to stand as a separate or disconnected component of the study.

      For this analysis, we included Bcl11a DNA and protein sequences from 23 vertebrate species, including humans. We refer the reviewer to the Methods section of this paper, under “dN/dS analysis”, for further details. To provide further clarity regarding the different species used in this study, we have now prepared a new Supplementary Table 4, listing the 23 species together with their DNA and protein sequence accession numbers for Bcl11a.

      We also added this sentence in the main text:

      “We included twenty-three vertebrate species, including humans (Supplementary Table 4).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Given how central the isolated cells are to the subsequent analysis, the manuscript would be strengthened by a figure showing expression of stage and lineage-specific markers.

      Ideally, the authors would provide some sort of orthogonal experimental approach to confirm downregulation of Bcl11a during oligodendrocyte differentiation and loss with age (e.g., IF or RNAScope in conjunction with stage-specific markers in tissue, or western blot in culture).

      Thank you again. We have performed these. Please see the Reviewer #1 comment (above).

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1A: It should be 'anti-O4' instead of 'anti-04'.

      This is now corrected. Thank you.

      (2) Figure 1C: The authors should specify what the connecting lines indicate (e.g., gene sets or gene modules).

      Each coloured line represents one gene and connects its log2 fold-change values across the three oligodendrocyte lineage states: OPC, PreOL and OL. The connecting lines are used to visualize gene-wise patterns of expression change across these cell states. For example, in Type 1, each line shows a pattern in which gene expression increases progressively from OPC to PreOL to OL, with the highest expression change observed in OLs: OL > PreOL > OPC.

      We now included the following line in the figure legend:

      “Each coloured line represents one gene and connects its log<sub>2</sub> fold-change values across the three oligodendrocyte lineage states: OPC, PreOL and OL. The connecting lines are used to visualise gene-wise patterns of expression change across these cell states.”

      (3) Figure 2C: The authors should specify "DF genes" in the figure legend.

      Thank you for pointing this out. This was a typo: it was written as DF, but it should be DE (differentially expressed) genes. We have now corrected this in the figure and spelled out the abbreviation in the legend. Also, DE gene list is accessible through GEO accession: GSE303317. This also mentioned in the figure legend as:

      “DE: Differentially expressed. DE gene list is accessible through GEO accession: GSE303317.”

      (4) Figure 4C & Figure 5B: the title for the y-axis of the bar graph is confusing. The authors should specify what "#" indicates. Does it represent the counts? What are the thresholding criteria to judge whether an Olig2 cell is MBP-positive or not? It's unclear what the unit is here for the 0-100 scale.

      We apologise for the confusion. We used ‘#’, which is a common notation in mathematical and quantitative contexts, to denote counts, so you are correct. We now mentioned in the legend: “The symbol “#” indicates cell count.”

      We counted the number of MBP+OLIG2+ cells, divided this by the total number of OLIG2+ cells, and expressed the value as a percentage. For greater clarity, instead of writing #MBP+/#OLIG2+, we have now written #MBP+OLIG2+/#OLIG2+.

      Regarding the 0–100 scale, the unit of the Y-axis is percentage, as stated in both figure legends.

      The criterion for classifying an OLIG2+ cell as MBP+ was morphological: an OLIG2+ nucleus, shown in white, had to be surrounded by MBP+ staining, shown in red. Cells meeting this criterion were counted as MBP+OLIG2+ cells. Manual counting was performed blinded to sample identity.

      (5) Figure 6B: To discriminate from IF staining, the authors should use italic'Plp1' to indicate the RNA in situ results.

      Thank you for pointing this out; we have now corrected it.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewer’s for their thoughtful comments that have significantly strengthened the paper. Below, we have outlined our responses to both the public reviews and recommendations.

      In addition to the alterations to the manuscript based on the reviews, during our review of the data analysis we uncovered some small errors that we have now corrected. In looking back over the image registration, we identified three animals whose olfactory bulbs did not register properly and one with poor cell counting in the telencephalon. To account for these issues, we imputed the missing data using an iterative soft-threshold singular value decomposition (described on lines 779-783 of the updated manuscript). This update had little impact on the results. We also identified a small error in how we determined ‘unique’ and ‘overlapping’ edges in the network analysis (Figure 8). In the previous analysis we had incorrectly noted that all ‘unique’ edges did not have an overlapping confidence interval with the two other networks (i.e., the networks for evading freezers, freezers, and non-reactive). Instead, the ‘unique’ edges in the prior version of the manuscript did not have an overlap with at least one other network. We have now updated the analysis so the reader can distinguish between edges that are truly ‘unique’ versus those with ‘1 overlapping confidence interval’ or ‘2 overlapping confidence intervals’ with other networks. As before, this update and change to the analysis does not materially affect the results or conclusions.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      The neural analysis part is very comprehensive. Figure 5 and Figure 6 are independent but complement each other very well. They together support that the cerebellar system is the key brain component for a freezing response. Their extreme focus on high-level analyses, however, came at the expense of biological intuitions. I suggest adding some figure panels and result/discussion paragraphs to help with that aspect.

      Thank you for the suggestion. We have made extensive edits to the manuscript to include additional discussion and biological intuition. Specifically:

      We added a supplemental figure (Figure S6-2) that has scatterplots showing how cfos levels vary with the different behavioral contrasts. Although the PLS analysis is multivariate, this univariate analysis should help give readers a better intuition of how the behavior relates to brain function.

      We have also rewritten the results sections for both the PLS analysis (lines 303-361) and network analysis (lines 396-437) to incorporate more of a discussion about the biological context of different regions identified. Thank you for this suggestion, we feel that this significantly strengthens the biological interpretation of the data for the reader.

      Reviewer #2 (Public review):

      (1) My first concern relates to the claim in the abstract that "We found that fear memory behavior fell into four distinct groups: non-reactive, evaders, evading freezers, and freezers".

      In my opinion, the "freezing" aspect is well supported as being both triggered by the CAS and for memory effect upon re-exposure to the tank, but I am less convinced about the "evasive" behaviour. In Figure 2, it appears that "evasiveness" is generally not increased in both the Exposure or Memory phases for many groups, and in Figure 5, it appears that "evasiveness" is expressed by nearly 50% of the fish in the pre-exposure condition before CAS addition and in all phases in the vehicle condition. Therefore, it appears that most of the expression of this behaviour is independent of any memorybased effect.

      We thank the reviewer for this suggestion and we agree that this line in the abstract was unintentionally misleading. We have now altered this line in the abstract (lines 34-36) to read:

      “We also found that that behavior fell into four distinct groups: non-reactive, evaders, evading freezers, and freezers with the evading freezer and freezer groups most clearly associated with memory formation.”

      On the larger point of the inclusion of evasion as part of the fear response, we believe this is warranted for the following reasons: (1) evasive behavior has long been acknowledged as a highly variable aspect of how fish respond to alarm substance where some fish exhibit evasion and others do not. This observation goes back to the original work from Karl von Frisch in minnows (von Frisch, 1938), and others in zebrafish (e.g., Suboski et al, 1990). One goal of our paper (and the work from the lab in general) is to try dissecting out this individual variation that can get lost when only considering population averages. (2) The unsupervised clustering also suggests that there are two distinct types of freezing clusters (Figure 4B) where some fish freeze intermittently with normal swimming and others freeze intermittently with evasive behavior. This suggests that evasion is increased in response to CAS, but only in a subset of fish. (3) The brain networks from the evading freezer and freezer groups are distinct (Figure 8A) despite having equally high levels of freezing behavior (Figures 4B and C). This means the difference we’re able to distinguish behaviorally is also manifesting in the brain, suggesting that it is not anomalous. Thus, while we agree that freezing is definitely the strongest and clearest behavioral response to CAS, we believe the analysis of this large dataset supports the interpretation that, in a subset of fish, increased evasive behavior in response to CAS is also a part of the response.

      (2) My second concern relates to the claim in the abstract that "background strain and sex influenced how fish respond to CAS, with males more likely to increase evasive behaviors than females and the TU strain more likely to be non-reactive."

      My understanding, based on the introduction and on the methods, is that it is likely important that the CAS be prepared from conspecifics of the same strain and sex, and for this reason, they prepared different CAS specific for each strain and each sex. Therefore, the "CAS" that is applied is necessarily different for each condition, and I am concerned about if the differences observed could relate more to variation in the quality, purity, concentration, etc. of the specific CAS samples for different groups, rather than their reactivity to the substance or their ability to form memories based on such experiences.

      The CAS was prepared by mixing extracts from all four strains and both sexes (so 8 fish per batch). Thus, all the fish were exposed to the same CAS mix derived from the same donors. This is described in the methods (lines 626-629). However, to ensure that this is clear to readers, we’ve now included a line indicating this in the results section (lines 123-124).

      (3) My third concern relates to the interpretation of the cFos data.

      As I mentioned above, I feel as though the behavioural analysis is perhaps more complex than is warranted via the inclusion of evasiveness, and I wonder if the conclusions from the experiments would be simpler if analyzed only from the perspective of freezing.

      We agree that the freezing response is driving the majority of the neural cfos response that we are seeing (e.g., Figure 6A-C). However, we feel that the network analysis (Figure 8) justifies the distinction between freezers and evading freezers. This is because the brain networks for these two groups (freezers and evading freezers) are quite distinct, even though these groups both have the same levels of freezing behavior (Figure 4). This stark difference in patterns of neural activity suggests the brain of a freezer and an evading freezer are engaging with the world in two distinct ways that is worth noting. We’ve updated the abstract to make this point clearer (abstract: lines 39-48) and discuss the biological interpretations of patterns of brain activity unique to evasion or evading freezers in more depth (lines 303-361; lines 396-437).

      Reviewer #3 (Public review):

      (1) The three-day contextual fear paradigm, as implemented - one CAS pairing on day 2 followed by a single recall test on day 3 - inevitably conflates acquisition and long-term memory, making it impossible to know whether strains like TU truly recall the association poorly or simply learn it more slowly. For example, given that TU fish extinguish fear faster than AB or TL strains in extended protocols, they may simply require additional or repeated CAS pairings to achieve the same asymptotic performance. To disentangle learning kinetics from recall strength, the assay could be revised to include multiple acquisition trials (e.g., conditioning on two or more consecutive days) with an immediate post-conditioning probe to assess acquisition independent of consolidation, and continuous measurement of freezing and evasive behaviors across each trial to fit learning curves for each strain. Such refinements - even if on a subset of the strains - would reveal whether "non-reactive" phenotypes reflect genuine recall deficits or merely delayed acquisition.

      We thank the reviewer for this thoughtful comment. We agree that it is difficult to disentangle acquisition from consolidation. Indeed, the TU fish do appear to have lower levels of freezing in response to the CAS (Figure 2A), supporting the idea that reduced performance at memory day could be due to some sort of deficit at acquisition. However, pursuing a detailed examination of strain dependent differences in fear memory acquisition versus consolidation is beyond the scope of the current paper where we primarily focus on individual differences in behavior. Nonetheless, we have included this important point in the discussion (lines 470-471).

      (2) My second major question is with respect to Figure 3 panel B. This is a complex figure, and I can understand the gist of what the authors are attempting to show, but it is difficult to understand as it is. Can this be represented in a way that is clearer and explained a bit more easily?

      We agree that this figure is one of the more complex in the paper. However, we’ve struggled to come up with a better way to present it. We have improved the presentation based on other reviewer comments by making the vehicle and CAS groups more easily distinguishable by using open versus closed circles. We’ve also included additional interpretations of the data in the results, which we hope will help guide readers through this figure better (lines 208-223).

      (3) The brain mapping is by far one of the most interesting aspects of this study, and the methods that the group used are interesting. The brain mapping, however, relies on generating "contrasting" groups (Figure 6A), and I was not clear as to how these two groups were formed. Could the authors elaborate a bit?

      These contrasting groups (contrast 1, contrast 2) arise analytically from the partial least squares (PLS) analysis; they are not defined by the experimenter. In brief, PLS is a multivariate technique that identifies latent variables that capture axes of maximal covariation between two datasets: behavior and brain activity. As an analogy to a more widely known technique, principal components analysis (PCA) uncovers axes of maximal variance within a single dataset. PLS, in contrast, simultaneously analyzes the covariation in two datasets. The contrast groups in Figure 6A represent the behavioral weights of the latent variables that capture the most covariance, which illustrates how the four behaviors load onto these top two contrasts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) The c-Fos analysis in Figure 5 is very comprehensive and convincing, but lacks intuitive presentations. In my understanding, the increase in c-Fos expression in red areas means increased freezing behavior for Contrast 1 for the PLS analysis? Do you have representative c-Fos expression images between different groups of fish?

      We decided not to include a representative cfos image because the data is derived from a large number of fish (N=87) and thus it can easily be cherry-picked to choose images that match the narrative. Instead, to more accurately capture the breadth of the data while providing a more intuitive presentation, we have included an additional supplemental figure that includes scatterplots of scaled cfos data against behavioral scores for each of the two contrasts (S6-2). We believe this more fully and accurately captures the relationship between behavior and brain activity. We included six different example brain regions and scatterplots for cfos activity against behavioral scores for contrasts 1 and 2, demonstrating a range of relationships. However, we should note that PLS is a multivariate technique, and so this univariate analysis does not fully capture the subtleties of the PLS analysis. Nonetheless, we think this will help give a more intuitive interpretation of the data to readers. We have also referenced this additional data in the manuscript (lines 307-309). We thank the reviewer for this excellent suggestion that improves the ability of readers to understand the paper.

      (2) Also related to Figure 5, the result section only describes the PLS statistics and does not try to describe the biological interpretation. Do the authors think the c-Fos expression directly represents lowlevel behavior, such as swimming, or a high-level behavioral state or learning? Maybe different areas mediate different aspects?

      For example, the medullary locomotor areas, which are usually highly correlated with swimming in terms of neural activity, seem to have higher c-Fos expression in freezing fish. I'm not saying this shouldn't be the case. c-Fos expression in this area was not elevated in larval fish during OMR in Shainer et al., 2023, indicating that it doesn't linearly reflect neural activity. But discussing a bit of intuition on the connection between c-Fos expression and biological process, rather than just saying "the cerebellum could regulate emotional states", would help us guide through this highly complex analysis.

      We have now added more interpretation of the data in both the PLS and network analysis sections (lines 303-361 and lines 396-437). Again, thank you for this excellent suggestion. This helps make the biological interpretation of the data clearer.

      Minor points:

      (1) Figure 2B titles: please write "memory" on the right side.

      We considered writing ‘memory’ on the right-hand side, but we thought this may add confusion because it would not apply to both graphs in the row. The left-hand graphs are the responses during ‘exposure’ and the right-hand graphs are the responses during the ‘memory’ phase. This is indicated by the titles above the left and right-hand sets of graphs.

      (2) Figure 2C: needs legend lines.

      We have now moved the legend lines from the top of the graphs to below the graph to make them more visible to readers.

      (3) Line 187: "aggregated" data.

      This has now been changed to ‘aggregated’ (now line 194).

      (4) Line 371: I'm not sure what "Beyond" means.

      We have now significantly changed this part of the paper and we no longer use the word ‘beyond’ here.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding point (1) in the Public Review:

      I would encourage the authors to consider whether this study might be better focused exclusively on the freezing behaviour, which does appear to be reliably expressed during CAS exposure and in the memory phases, and would significantly simplify the subsequent analyses of neural activity, and perhaps may lead to a more coherent conclusion.

      As noted in our response to the public review, we appreciate this suggestion, but we have decided to keep the inclusion of the evasive behavior. This is because (1) evasive behavior has long been acknowledged as a highly variable aspect of how fish respond to alarm substance where some fish exhibit evasion and others do not. This observation goes back to the original work from Karl von Frisch in minnows (von Frisch, 1938), and others in zebrafish (e.g., Suboski et al, 1990). One goal of our paper (and the work from the lab in general) is to try dissecting out this individual variation that can get lost when only considering population averages. (2) The unsupervised clustering also suggests that there are two distinct types of freezing clusters (Figure 4B) where some fish freeze intermittently with normal swimming and others freeze intermittently with evasive behavior. This suggests that evasion is increased in response to CAS, but only in a subset of fish. (3) The brain networks from the evading freezer and freezer groups are very distinct (Figure 8A) despite having equally high levels of freezing behavior (Figures 4B and C). This means the difference we’re able to distinguish behaviorally is also manifesting in the brain, suggesting that it is not anomalous. Thus, while we agree that freezing is definitely the strongest and clearest behavioral response to CAS, we believe the analysis of this large dataset supports the interpretation that, in a subset of fish, increased evasive behavior in response to CAS is also a part of the response.

      A more minor concern related to the analyses in Figure 2: in the figure legend, it is stated that "*-P < 0.05 compared to vehicle treated fish via t-tests". How are the authors dealing with the multiple comparisons problem? Would something like an ANOVA not be more appropriate?

      Thank you for bringing this point up. We did not initially correct for multiple comparisons because we considered each of these experiments across sex and strain separate since we did not compare across strains. However, the way we’ve grouped the data together in figure 2 makes it appear as if they are one large experiment. To alleviate any concern about multiple testing, we have now corrected for multiple comparisons using the false discover rate (FDR) correction. The statistics in the figure and captions have now been updated.

      (2) Regarding point (2) in the Public Review:

      If the authors agree with my concern regarding potential variability in the CAS samples, I would suggest either testing for differences among strains using the same batch of CAS, or including and explaining this caveat in the text.

      As noted in our response to the public review, the CAS was the same for all the fish. Each batch was derived from 8 donor fish, one fish from each strain and sex (described in lines 123-124 of the results and lines 626-629 of the methods).

      (3) Regarding point (3) in the Public Review:

      I feel like the standard in the field for such conclusions would be after

      (a) Direct analyses of the activity states in these areas. I was surprised not to see a direct analysis of the cFos stainings in the cerebellum relative to freezing behaviour, for example, ideally in a different animal cohort.

      The PLS analysis does relate activity in the cerebellum (and other brain regions) to specific behaviors via the the behavioral contrasts (Figure 6A). We believe this approach (instead of dividing fish into ‘high and low freezers’) is a more powerful way to leverage the data from all the animals tested (87 fish). However, we appreciate that the interpretation of the PLS analysis is not as intuitive as seeing scatterplots or bar charts comparing neural activity. For this reason (and in response to a comment from reviewer 1), we have included as a supplementary figure (Figure S6-2) scatterplots showing how standardized c-fos activity varies with the behavioral scores from the contrasts identified from the PLS analysis. Given that contrast 1 weights heavily in the positive direction on freezing, these figures can essentially be read as looking at cfos activity as a function of freezing levels. What can clearly be seen is that for regions of the cerebelleum (E.g., the LCa and CC) there is a clear positive relationship between cfos activity and the behavior scores for contrast 1.

      (b) Some kind of manipulation of the brain area resulting in the relevant behavioural modification.

      We completely agree with the reviewer. However, at the moment, we do not have the tools to do this in adult zebrafish. It is something we’re actively working on.

      Of course, I appreciate that such experiments might not be possible or feasible, and in which case I would suggest adjusting the claims accordingly and highlighting the caveats to their interpretations.

      We have incorporated the caveat that we have not directly altered neural activity into the discussion (lines 542-543) and adjusted how we discuss our findings in the abstract (lines 39-41) to more accurately represent the type of evidence we provide. Hopefully we’ll be able to do so in the near future!

      MINOR CONCERNS:

      (1) In Figure 3, how is the end of a behavioural epoch defined? I am surprised to see that you consider transitions between the same behavioural state. How does erratic swimming -> erratic swimming differ from a longer single epoch of erratic swimming? In general, I find this analysis confusing, and I am not sure if it adds significantly to the message of the paper.

      Thank you for this question as it prompted us to realize we were missing this in our methods section. We have now updated the methods to include how we calculated the behavioral transitions (lines 644-650). In short, we used a 750 ms behavioral epoch time that corresponds to the size of the sliding window we used for the random forest model.

      We have also updated the description of this analysis in the results to indicate the main finding from it (lines 207-223). In brief, the main finding is that exposure to CAS results in longer bouts of evasive behavior without increasing its frequency. Whereas CAS induced freezing arises from both longer bouts and likelihood of occuring. While we agree that this is a relatively minor finding in the paper, one of our goals is to provide as comprehensive analysis of fear behavior as possible to help guide future researchers interested in using fish for understanding different aspects of fear-related behaviors.

      (2) In the PLS analyses, two measures of evasion are used: evasion time, and evasion as a percent of active behavior. I don't understand the justification for both of these being used rather than one. Again, my overall recommendation is to reduce the focus on the analysis of evasion behaviour, but if you do not choose to do this, I think the rationale of how both measures are used and why needs explanation.

      We chose to incorporate two different measures of evasion throughout the study because the high levels of freezing in some animals results in little opportunity to express other behaviors (like evasion). Thus, to better capture what fish may be doing in the absence of freezing (i.e., when they are active) we also calculate the amount of active time spent performing evasive behaviors (instead of normal swimming). We have now included an explanation for this earlier in the results section when we first use this metric (lines 149-152).

      (3) In the methods, I don't understand this: "Animals that were assigned the wrong sex were removed from data analysis, as well as its paired fish (< 2%)".

      We determine the sex of fish when we set them up for dual housing. However, we occasionally make errors in sex determination. To ensure we properly sexed the fish, at the end of experiments, we euthanize the fish and check for the presence of eggs. If we incorrectly assigned the sex to a fish, they are removed from the experiment alongside the other fish they were dual housed with. This is because we want to ensure all fish are housed in the same way (i.e., a male fish with a female fish).

      (4) How was this determined differently from the first time, resulting in exclusion?

      After experiments, fish were euthanized and we checked for the presence of eggs (line 599-601). We’ve now added a line in this other part of the methods referring back to where we describe this (lines 676678).

      Reviewer #3 (Recommendations for the authors):

      Here are some minor concerns and errors found in the manuscript:

      (1) For Figure 2B and Figure 3B, can the group make the lines solid and dotted? The circle or triangle designation is difficult to see, and since the crux of the figure depends on comparing Veh and CAS, it would be easier to see if the lines were altered.

      Thank you for this suggestion. Instead of making the lines solid and dotted, we decided to make both the CAS and vehicle group circles and then have open and closed circles. We believe this solves the issue of being able to distinguish these groups and makes the data more readable.

      (2) Figure 2C: It appears that the line colors in the legend are missing.

      We have moved the line colors below the graphs to make them more obvious.

      (3) Figure 8A: Same thing here - could the text be enlarged? It's really difficult to make out each node, and when I zoom the text becomes pixelated. This is an important figure and one that will likely be referenced, and making it clear would be helpful.

      This one is difficult. We have made the network images as large as would fit on a page. We have now uploaded vectorized versions of the images so that they do not become pixelated when zooming in. As part of our supplemental materials we also include a cystoscope file that can be explored in greater depth as well.

      (4) The paper is really well written: I found a few typos, though:

      (a) Line 529: "Institutional Cara and Use Committee" should be "Institutional Animal Care and Use Committee" (Change cara to care and add animal).

      (b) Line 274: "hybdridization" should read hybridization.

      Thank you for catching these typos. They have now been fixed.

    1. Author response:

      We thank the editorial team and all three reviewers for their time and attention to detail in reviewing our manuscript. We are particularly grateful for comments recognising the “direct practical value” and “comprehensive analysis” in our work (Reviewer #2) as well as its “thorough and methodical” approaches (Reviewer #1).

      We also appreciate the reviewers’ constructive comments to improve our manuscript, which we plan to address in a revised version. Specifically, we plan to more explicitly acknowledge some of the limitations of our study, including clearer highlighting of the experiments performed only in females (Reviewer #1); the inability of our experimental design to control for chloride concentration (Reviewers #1 and #2); the limitations of transgene induction using the AGES system in adult flies (Reviewer #2); and the rationale for using different concentrations of auxin across different experiments (Reviewer #3). We also note that Reviewers #1 and #2 have provided additional recommendations beyond the public reviews (largely relating to helpful ways to clarify our text and more explicitly acknowledge limitations), which we also plan to address.

      In addition, we plan to perform additional experiments to address specific points raised by reviewers in both their public reviews and additional recommendations. In the first instance, we plan to follow Reviewer #3’s suggestion to characterise induction in the brain using 10mM auxin, as well as Reviewer #2 and #3’s suggestions to explore feeding behaviour of flies fed auxin within our experimental setup.

      We look forward to submitting an improved manuscript guided by the reviews, with the aim of strengthening this “well-needed validation” study that also “highlights critical caveats that will support future studies” (Reviewer #2).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (R1C1) The only minor weakness that I found is the assumption of independence of bacterial species, which is expressed as the well-stirred approximation. One could imagine that bacterial species might cooperate, leading to non-uniform distributions that are real. How to distinguish such situations?

      I believe that this method can be extended to determine if this is the case or not before the application. For example, if the bacteria species are independent of each other and one can use the binomial distributions, then the Fano factor would be proportional to the overall relative fraction of bacterial species. Maybe a simple test can be added to test it before the application of REPOP. However, I believe that this is a minor issue.

      This is an interesting point raised by the reviewer.

      First, we need to clarify an important point: we do not make a well-stirred assumption. Samples can be drawn and plated from any region of space however small and that region’s population can be quantified using our method. The stirring only occurs after we collect a sample in order to dilute the contents and pour the solution homogeneously over the plate.

      As such, learning multiple independent species is possible and not impacted by the dilution (“well-stirred” assumption). In the new first paragraph of the methods section, we made it clear that this assumption concerns the dilution process. REPOP is designed to recover the true underlying heterogeneity in species abundance (even from limited data) by leveraging a Bayesian framework that remains valid regardless of whether species are independent or correlated.

      If the method is applied to multiple species as currently implemented, REPOP can recover the marginal distribution of each species, provided that the species are either selectively cultured or produce sufficiently distinguishable colonies on the same plate. To demonstrate this, we have added a new Results subsection with a synthetic two-species example in which the species abundances are correlated across samples.

      However, in order to learn the joint distribution and capture correlations between species within samples, the method would need to be extended. At present, in Eq. 5 we sum the likelihood over all values of n, using a data-driven cutoff (twice the largest naïvely estimated count times the dilution factor). Extending this to multiple species adding up to (n<sub>1</sub>,n<sub>2</sub>), while retain the generality of the method, would require quadratically scaling memory with this cutoff in the population number. For this reason while we comment on this in the new paragraph in the conclusion, it is not implemented as part of REPOP.

      Reviewer #2 (Public review):

      (R2C1) A more thorough discussion of when and by how much estimated microbial population abundance distributions differ from the ground truth would be helpful in determining the best practices for applying this method. Not only would this allow researchers to understand the sampling effort necessary to achieve the results presented here, but it would also contextualize the experimental results presented in the paper. Particularly, there is a disconnect between the discussion of the large sample sizes necessary to achieve accurate multimodal distribution estimates and the small sample sizes used in both experiments.

      That is a great suggestion from the reviewer. To address it, we expanded Appendix B. We know report (1) the relative error in the estimated means (as already done for Fig. 4 formally 3), and (2) the Kullback-Leibler (KL) divergence between the reconstructed and ground-truth distributions. These metrics will are show as a function of the size of the dataset, for the examples in Fig 3. enabling a direct assessment of how the sampling effort affects the precision of the inference.

      That said, we now highlight in the Conclusion that, by explicitly modeling the dilution process within a Bayesian framework, REPOP extracts the maximum information available from each individual sample at a given sample size. This strategy therefore enables more accurate inference with fewer measurements, which is particularly important in applications such as plate counting, where data acquisition is labour-intensive.

      Reviewer #3 (Public review):

      (R3C1) While the study is promising, there are a few areas where the paper could be strengthened to increase its impact and usability. First, the extent to which dilution and plating introduce noise is not fully explored. Could this noise significantly affect experimental conclusions? And under what conditions does it matter most? Does it depend on experimental design or specific parameter values? Clarifying this would help readers appreciate when and why REPOP should be used.

      We agree with the reviewer that this is an important point, and we expanded Appendix B to include a quantitative analysis using simulated data (Fig. 3, formely 2), reporting both relative error and KL divergence as a function of dataset size. This complements our response to R2C1 clarifying when REPOP offers the greatest benefit.

      In addition, we will expand the discussion on how modeling dilution noise becomes essential when learning population dynamics. In particular, we emphasize? the role of Model 3, especially relevant when working with multiple plates and approaching the asymptotic regime; an aspect that was alluded to in Fig. 3 but not fully explored.

      (R3C2) Second, more practical details about the tool itself would be very helpful. Simply stating that it is available on GitHub may not be enough. Readers will want to know what programming language it uses, what the input data should look like, and ideally, see a step-by-step diagram of the workflow. Packaging the tool as an easy-to-use resource, perhaps even submitting it to CRAN or including example scripts, would go a long way, especially since microbiologists tend to favor user-friendly, recipe-like solutions.

      In the new paragraphs of the introduction, we made clear that REPOP is written in Python (PyTorch), installable via pip, and designed for ease of use. We are also expanding the tutorials to include clearer guidance on data formatting and common workflows. The new workflow figure (Fig 2) better illustrates the full process.

      (R3C3) Third, it would be great to see the method tested on existing datasets, such as those from Nic Vega and Jeff Gore (2017), which explore how colonization frequency impacts abundance fluctuation distributions. Even if the general conclusions remain unchanged, showing that REPOP can better match observed patterns would strengthen the paper’s real-world relevance.

      We thank the reviewer for this interesting suggestion. We agree that applying REPOP to additional existing datasets would make REPOP’s relevance clearer. However, the Vega and Gore datasets lack the information required. REPOP requires the plate count measurement process to be specified, including the dilution factors used for each measurement. Furthermore, we can leverage on additional information about the experimental procedure when the colony cutoffs and dilution schedules used are reported. Without the dilution factors, the likelihood connecting the observed colony counts to the underlying population size is not possible. We hope this clarification will help make future datasets made available publicly more useful for purposes of uncertainty propagation.

      (R3C4) Lastly, it would be helpful for the authors to briefly discuss the limitations of their method, as no approach is without its constraints. Acknowledging these would provide a more balanced and transparent perspective.

      We agree with the reviewer. We have added two new paragraphs to the conclusion highlighting important current constraints and future development directions of the framework. In particular, we now discuss that, in its present implementation, REPOP focuses on the population distribution that maximizes the posterior, rather than returning posterior uncertainty over the reconstructed distributions themselves. We also note the computational demands of the method, making GPU acceleration highly beneficial and more complex multi-population inference computationally challenging. This discussion synthesizes points raised throughout our response to R1C1 and the reviewers and provides a more balanced perspective on the current scope of the method.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.

      To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to some neural changes across sessions. However, several aspects of our behavioral design and results argue against extinction or repeated nonreinforced retrieval as the primary drivers of the observed effects. Importantly, discrimination ratios remained stable or increased across time rather than progressively diminishing as would be expected under extinction (this new analysis will be added to the resubmission). Nevertheless, we will address this important point in the Discussion and explicitly acknowledge that repeated retrieval may contribute to some component of the observed representational changes.

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      We thank the reviewer for raising this important point, which was also noted by the other reviewers. To address this issue, we will implement Reviewer 3’s suggested Generalized Linear Model (GLM) analysis using inferred spiking activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. Because freezing behavior varies across trials whereas stimulus identity is fixed, this approach will allow us to dissociate their respective contributions to neuronal activity. If, after accounting for freezing behavior, responsive neurons continue to exhibit graded coding consistent with inferred threat value, this would strengthen the interpretation that the identified ensembles reflect generalization gradients related to aversive value rather than freezing behavior alone. Otherwise, we will adjust the conclusions according to the interpretation that freezing itself drives the generalization gradients.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      We thank the reviewer for this suggestion. We will include measures of registration quality in the resubmission.

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      We will correct correspondence between text and Figure.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      Thank you for this suggestion. We will clarify the labelling of the Figure 2a and call the graphs “activity-plots”.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      We will redo the Statistics of Figure 3 to take into account the following variables: group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30).

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term “inference” may overstate the cognitive processes engaged by the current task. Accordingly, we will revise the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

      We will correct the language and use “threat value”.

      Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      We thank the reviewer for these important points, which we address in further detail below.

      Ongoing learning during test sessions: The reviewer correctly notes that unreinforced test presentations may constitute extinction-learning trials and that some neural changes across days could therefore reflect ongoing learning rather than spontaneous ensemble reorganization. However, new analyses indicate that extinction is unlikely to be the primary driver of our findings. Discrimination ratios do not decay over time; instead, they either sharpen or remain stable across sessions (new analyses to be included in the resubmission). These results argue against robust extinction as the primary source of the neural changes observed across sessions. This interpretation is also consistent with the strength of our conditioning protocol, which used 10 CS+ shock pairings and 10 CS− no-shock pairings specifically to minimize extinction across repeated testing sessions. Nevertheless, we acknowledge that the current design cannot fully dissociate time-dependent consolidation from retrieval-induced plasticity, and we will explicitly discuss this limitation in the revised Discussion.

      Stability reflecting behavioral consistency: We agree this alternative cannot be fully excluded. However, the cluster stability analyses assess identity at the level of response profile across all four frequencies, not response magnitude alone. Tone-selective clusters, which also show consistent behavioral correlates (firing rate correlates with threat-value, Fig. S8), do not show equivalent profile stability, suggesting that the stability of graded clusters is not simply a consequence of behavioral consistency. This point will be added to the Discussion in the resubmission.

      Language of "inference" and "integration": The reviewer is correct that responses to novel tones are consistent with graded stimulus generalization. We will substantially revise the manuscript to replace "inference" and "integration" with more precise language describing graded frequency generalization gradients.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      We thank the reviewer for these specific and constructive criticisms. We will revise the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.

      (a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      The hypothesis will be rewritten as: "We hypothesized that generalization to tones acoustically similar to the CS+ and CS− depends on the emergence of stable ensembles encoding threat value, and that population-level response similarity across stimuli would correlate with the degree of behavioral fear generalization, consistent with prior work in auditory cortex [1]."

      (b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      The summary statement will be rewritten: "Our results show that stable cortical sub-ensembles preserve the emotional content of the fear memory over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity in response to tones associated with threat correlates with the degree of behavioral threat generalization."

      (c) 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS⁺15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      We will adopt the reviewer's suggested rewording: "In CS+15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization."

      We will systematically review the entire manuscript to ensure consistency with this revised framing.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we will follow Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using inferred spiking activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis will allow us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we will revise and appropriately adjust our conclusions.

      In addition, we will revise the section heading and surrounding text to remove the terminology of “learned and inferred valence.” Instead, the findings will be described more conservatively as: “PL population responses reflect behavioral generalization to auditory stimuli following discriminative fear conditioning.”

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      The reviewer is correct that controls do not form threat associations; however, these animals still could respond differentially to distinct frequencies, something that is not reflected in the data. We will correct the section indicating that distinct neutral frequencies do not produce graded responses: "graded responses require associative learning" will be retained but reframed simply as: "graded frequency-dependent population responses were absent in animals that did not receive fear conditioning." The concluding statement of the paragraph will be rewritten as: "PL population responses reflect behavioral generalization to acoustically similar stimuli following discriminative conditioning," in line with the reviewer's suggestion.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      We respectfully disagree with the reviewer’s interpretation. While repeated testing cannot be entirely excluded as a contributing factor, several lines of evidence suggest that it cannot fully account for our observations.

      Regarding extinction: discrimination ratios between CS+ and all other frequencies either remained stable or increased over time (new analysis included in resubmission), indicating that animals continued to discriminate threat value across the testing period rather than showing the progressive suppression expected under extinction — the opposite of what we observe.

      Regarding the recruitment of new neurons: repeated non-reinforced tone exposure would be expected to produce stimulus-specific adaptation — characterized by reduced, less discriminative neural responsiveness and flatter tuning profiles [2]— not the progressive sharpening we observe. The same would be expected if these neurons represent or are associated with new extinction learning.

      Finally, sharpening of generalization gradients during repeated within-subjects testing has been reported previously [3], suggesting that successive exposures may promote more precise discrimination in some cases. Consistent with this, discrimination learning has also been shown to narrow or sharpen fear generalization gradients rather than broaden them [4], supporting the idea that discriminative conditioning enhances stimulus specificity during testing. Although we cannot exclude the possibility that more extended training could eventually broaden the generalization gradient, under the training parameters and temporal window used in our study, the data support a progressive sharpening of the gradient over time. In the revised Discussion, we will present systems consolidation as the primary interpretive framework and further elaborate on why repeated testing is unlikely to account for the full pattern of behavioral and neural findings reported here.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS⁺. In CS⁺15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      The GLM analysis described in our response to reviewers 1 and 3 will directly address the contribution of freezing. We will report these results in the resubmission and revise the interpretive language in the manuscript accordingly.

      However, regarding the analysis of population vector similarity, we need to clarify a point of confusion. The reviewer states “Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained”. The similarity vectors were calculated by correlating activity across all tone presentations within each testing day, not only the first two presentations. In Fig. 4, “Early” and “Late” refer to the order of a tone within a trial, which we will clarify more explicitly in the resubmission. Notably, repeated-measures analyses did not reveal any effect of the time variable (Fig. 4e,f), indicating that similarity across tone presentations remained high for tones associated with high threat value. Importantly, our data showed no evidence that responses to 11 kHz or 15 kHz in the CS15 group, or to 3 kHz in the CS3 group, exhibited extinction-like patterns at either the behavioral or neural level. Therefore, the persistence of high population similarity across time provides additional evidence against extinction as the primary explanation for our findings.

      We will remove the word "determines" from the manuscript, as our data cannot conclusively establish a causal relationship.

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      We will make this correction.

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      We thank the reviewer for raising this point. The phrase "integrated learned valence across tones" refers specifically to a subpopulation of neurons that respond to all four frequencies in a graded manner, with response magnitude scaling according to threat value. This is distinct from tone-selective neurons, which respond preferentially to a single frequency. The neurons responding to all tones in a graded manner are present only in conditioned animals and not in no-shock controls, demonstrating that their graded response profile is shaped by associative learning.

      We agree, however, that the phrase "integrated learned valence" is unnecessarily opaque and we will replace it with more precise language: these neurons will be described as showing graded frequency-dependent responses whose magnitude scales with threat value. We believe this subpopulation represents a genuinely novel finding that complements the behavioral generalization literature by identifying a specific neural substrate for the generalization gradient within PL.

      (8) Another example of what has been a common theme in this review:

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      As stated in our previous responses, in the resubmission we will determine the contribution of freezing. If we find that freezing predicts graded neural responses, we will adjust the language of the manuscript.

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      We thank the reviewer for highlighting the lack of clarity in this passage and agree that the original phrasing was insufficiently precise. What we intended to convey is that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. Therefore, we suggest that neurons recruited into the active population after conditioning — likely frequency-selective neurons — contribute to the graded population responses through changes in their firing-rate activity, which is modulated by threat value (Fig. S8). We will rewrite this passage in the resubmission to make this interpretation explicit rather than leaving it to the reader to infer.

      Regarding the reviewer's suggestion that the characteristics of newly recruited neurons more likely reflect learning from non-reinforced exposures during repeated test sessions, we respectfully maintain that this interpretation is difficult to reconcile with two aspects of our data. First, graded-response neurons are absent in no-shock controls that are exposed to nonreinforced repeated testing. Second, as detailed in our responses to previous points, the progressive sharpening of population responses over time is inconsistent with what would be expected from repeated non-reinforced exposure, which would more plausibly produce broader or flatter tuning profiles.

      We agree that the phrase "shaped by associative processes" was ambiguous and will replace it with explicit language clarifying that we refer to fear conditioning as the associative process driving the emergence of graded responses, rather than any learning occurring during the test sessions themselves.

      (10) The following points all relate to the Discussion and reiterate many of the points above. 

      (a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      We will modify the language as stated in the prior points.

      (b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      We will incorporate new analysis (GLM) to better address this point and conclusions.

      (c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      We thank the reviewer for this comment and agree that both phrases require clarification or revision.

      By "graded representational axis" we intended to convey that PL population activity varies systematically as a function of stimulus similarity to the conditioned tone — that is, population responses are not categorical but scale continuously with spectral proximity to the CS+. We agree this was not clearly stated and will revise the manuscript accordingly.

      Regarding recurrent connectivity, we agree with the reviewer that nothing in our data directly measures or manipulates connectivity between neurons. This statement was intended as a speculative interpretive hypothesis in the Discussion, motivated by the established literature linking strong recurrent connectivity in prefrontal circuits to stable population-level representations [5]. However, we acknowledge that invoking it in this context, without direct evidence, risks overstating our conclusions. We will revise this sentence to make its speculative nature explicit and ground it more carefully in the cited literature rather than presenting it as an inference from our own data.

      In summary, we will ensure our conclusions will be restricted to population-level coding of learned threat value and its generalization across auditory frequencies. We will revise the relevant passages in the Discussion to ensure that speculative interpretations regarding emotional memory content are either removed or clearly flagged as speculative hypotheses.

      (d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

      We thank the reviewer for this point. Our interpretation that these neurons reflect pre-existing PL sensory-driven properties was based on the observation that tone-selective responses were present in control animals that never received conditioning, consistent with prior reports of sensory responsiveness in PL cortex ([6, 7]. Because these responses emerge from the first time we expose mice to the intermediate frequencies, they cannot be explained by repeated exposure. Moreover, we did not observe progressive refinement, emergence of discrimination-like changes, or suppression of responding to non-reinforced tones in control mice. This difference between conditioned and control animals indicates that repeated tone exposure alone is not sufficient to produce the observed dynamics — associative learning is necessary. We therefore maintain that the tone-selective responses of these neurons reflect pre-existing sensory-driven properties of PL cortex that are present independently of conditioning history.

      In summary, we thank the reviewer for suggesting clarifications to our interpretation, for raising the possibility that freezing behavior may contribute to graded neural responses, and for raising the question of whether repeated tone exposure may contribute to the properties of neurons recruited after conditioning. In the revised manuscript, we will include additional analyses to better dissociate the contributions of freezing behavior and tone identity, clarify passages that were insufficiently precise, and include a paragraph in the Discussion addressing potential alternative explanations alongside our own interpretation of the data.

      Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded-response ensembles may partially reflect freezing-related activity rather than mnemonic or salience-related representations of the conditioned stimuli themselves. In the revision, we will acknowledge that prior work has identified PL neurons that encode freezing independently of stimulus identity or associative content. Furthermore, we will implement the reviewer’s suggested generalized linear model (GLM) approach using inferred spiking activity derived from the Ca2+ signals. Specifically, we will include both stimulus identity and freezing behavior as predictors. Because freezing varies across trials whereas stimulus presentation is fixed, this analysis will allow us to dissociate the relative contributions of stimulus-related versus freezing-related activity to the graded neuronal responses. We thank the reviewer for this excellent suggestion.

      If graded stimulus coding remains significant after accounting for freezing behavior, this would strengthen the interpretation that these ensembles encode learned salience or associative properties of the stimuli rather than behavioral output alone. Conversely, if freezing explains a substantial proportion of the variance, we will revise our interpretation accordingly.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

      We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such holographic two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.

      To provide an indirect link between ensemble organization and behavior within the constraints of the current dataset, we will examine inter-individual variability in the revised manuscript. Specifically, we will test whether the proportion of neurons participating in stable graded-response ensembles versus dynamic stimulus-specific ensembles predicts individual differences in freezing behavior and fear generalization across retrieval sessions. If animals with a higher proportion of stable graded-response neurons show stronger discrimination and less generalization to non-conditioned tones, this would strengthen the association between ensemble organization and behavioral outcome, while remaining correlational in interpretation.

      We will modify the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the associative nature of our conclusions.

      References

      (1) Aschauer, D.F., et al., Learning-induced biases in the ongoing dynamics of sensory representations predict stimulus generalization. Cell Rep, 2022. 38(6): p. 110340.

      (2) Kato, H.K., S.N. Gillet, and J.S. Isaacson, Flexible Sensory Representations in Auditory Cortex Driven by Behavioral Relevance. Neuron, 2015. 88(5): p. 1027–1039.

      (3) Vervliet, B., et al., Generalization gradients in human predictive learning: Effects of discrimination training and within-subjects testing. Learning and Motivation, 2011. 42(3): p. 210–220.

      (4) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.

      (5) Mante, V., et al., Context-dependent computation by recurrent dynamics in prefrontal cortex. Nature, 2013. 503(7474): p. 78–84.

      (6) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.

      (7) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript provides a comprehensive and mechanistic analysis of how tsetse flies feed on blood across a wide range of host skin types. The authors combine detailed anatomical characterization of the feeding apparatus with quantitative measurements of mechanical properties, probing forces, and blood uptake, complemented by experiments using artificial skin. They show that tsetse flies do not rely on extreme forces or uniquely specialized structures, but instead on subtle and highly efficient structural and mechanical adaptations (such as the toothed labellum and coordinated proboscis movements) to achieve effective blood pool feeding. The study successfully moves beyond descriptive anatomy to a quantitative, functional analysis that explains how feeding is accomplished across diverse substrates.

      Strengths:

      A major strength of the work is the impressive integration of multiple complementary approaches. Advanced imaging tools provide a convincing three-dimensional view of the proboscis, labellum, and associated structures, while direct force measurements and blood intake quantification place these observations on a solid quantitative footing. The use of artificial skin with different mechanical properties is particularly powerful, as it allows structure-function relationships to be tested under controlled and reproducible conditions. Together, these datasets provide strong and coherent support for the authors' central conclusions. The quantitative treatment of feeding mechanics represents a significant advance over largely descriptive prior work by others (e.g., Gibson W et al 2017) and establishes a valuable mechanistic insight for studying blood feeding in insect vectors more broadly.

      Weaknesses:

      The study focuses almost entirely on uninfected flies and does not address how infection might alter feeding mechanics or performance. Previous work has shown that trypanosome infection can affect salivary gland function and feeding time (Van Den Abbeele et al 2010), and even cause damage to mouthparts, all of which can influence feeding behavior and efficiency. While this does not detract from the technical quality or the core findings of the study, a more explicit discussion of these biological variables would help place the results in a broader transmissionrelevant context and clarify how generalizable the conclusions are to natural infection settings.

      We thank the reviewer for this important comment. While our study focused on uninfected flies, we agree that parasite infection may influence feeding performance and should therefore be considered when assessing the broader relevance of our findings. Previous studies have shown that trypanosome infections can alter salivary gland physiology and saliva composition (Van Den Abbeele et al., 2010; Matetovici et al., 2016). In addition, transcriptomic analyses suggest that infection with T. congolense may affect the molecular and physiological state of the proboscis (Awuoche et al., 2017). However, there is currently no direct evidence that these changes translate into fundamental alterations of the mechanical properties or function of the mouthparts themselves, which were the primary focus of our study. We have now expanded our manuscript to discuss this (lines 508-520):

      "It is also important to note that our experiments were conducted using uninfected flies. Previous studies have shown that trypanosome infection can alter feeding behaviour, leading to increased probing activity and prolonged feeding times (Jenni et al., 1980; Van den Abbeele et al., 2010). These effects have primarily been attributed to infection-induced changes in saliva composition and the resulting interactions with host blood (Van den Abbeele et al., 2010). However, a different study found no significant effects of infection with either salivary gland-resident T. brucei or with proboscis-colonizing species such as T. congolense and T. vivax on Glossina feeding behaviour (Moloo, 1983). Furthermore, although infection-associated transcriptional changes in the salivary glands and proboscis have been reported (Awuoche et al., 2017; Matetovici et al., 2016), there is currently no direct evidence that trypanosome infection alters the mechanical properties or function of the mouthparts themselves."

      Overall, this is an outstanding and carefully executed study that will have a significant impact on the fields of vector biology and parasite transmission.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an impressively detailed, multidisciplinary analysis of the mechanics of blood feeding in Glossina spp. Combining SEM, CLSM, µCT, FIB-SEM, macro-videography, and quantitative force measurements, the authors characterize the structures and biomechanics of attachment, proboscis deployment, tissue penetration, and blood uptake. They also examine interactions with diverse host-type substrates, from human skin equivalents to cow, deer, and lizard skin, and integrate these with force measurements to quantify penetration and retraction dynamics.

      The work's key conclusion is that the tsetse fly does not rely on any single exceptional morphological innovation, but rather uses a suite of subtle structural features and retractive forces to feed efficiently across diverse hosts. This result is novel, insightful, and evolutionarily compelling. Overall, this is a strong manuscript that combines methodological sophistication with biological relevance. It should be of high interest to researchers studying vector biology, biomechanics, parasite transmission, and vector-host interactions.

      Strengths:

      (1) The combination of SEM, CLSM, µCT, and FIB-SEM provides an unusually comprehensive anatomical characterization of the tsetse feeding apparatus.

      (2) The direct measurement of proboscis penetration and retraction forces across diverse substrates is highly original and fills a major knowledge gap in vector-host interaction mechanics.

      (3) The study bridges morphology, mechanics, behavior, and host tissue properties, which strengthens the overall conclusions.

      (4) Imaging of trypanosomes within the hypopharynx and surrounding tissue during feeding provides new information about parasite delivery mechanisms.

      Main Comments:

      (1) The authors conclude that feeding versatility arises from the sum of subtle adaptations. This interpretation is reasonable, but it would help to sharpen which findings most robustly support this statement. For example, the relative similarity of proboscis forces across skin types is compelling evidence that the proboscis is broadly tuned rather than specialized. The observation that tsetse targets softer interscale regions on lizard skin suggests behavioural selectivity, not morphological specialisation. It would strengthen the discussion to highlight which data most directly refute the hypothesis of a unique specialization.

      We thank the reviewer for this comment. To address this point more explicitly and to sharpen the interpretation of our findings, we have expanded the final conclusion in the Discussion (lines 528544):

      "Ultimately, the objective of this study was to investigate how tsetse flies can feed on a seemingly random selection of animals with highly diverse skin structures. In our detailed anatomical studies and force measurements, we did not identify a single dominant trait that explains the fly's feeding versatility.

      Instead, our results indicate that this capability emerges from the combined effect of multiple, more subtle traits. In particular, the proboscis generates broadly similar penetration forces across a wide range of skin types, suggesting a generalised mechanical mechanism rather than hostspecific optimisation. The intricate architecture of the labellum and the strong retractile forces during probing likely contribute to efficient penetration and blood pool formation across heterogeneous substrates. Behaviourally, tsetse flies further increase feeding success by flexibly targeting mechanically favourable sites, such as the softer interscale regions on lizard skin, rather than relying on specialised morphological adaptations.

      This composite strategy likely reflects evolutionary fine-tuning that enables the broad host range of tsetse flies. By allowing efficient blood feeding across diverse vertebrate hosts, this versatility may also have facilitated the ecological success and transmission opportunities of African trypanosomes.”

      (2) A central finding is that retraction forces exceed penetration forces across substrates, implying that backward pulling is a key component of wound creation. However, the biological interpretation could be deepened. Specifically, do the authors believe retraction serves primarily to enlarge the pool-feeding site? How does this compare mechanically to mosquito fascicle oscillation or other blood-feeding arthropods (especially other flies such as those in the tabanidae family)? Could retraction forces contribute to anchoring or resisting host grooming behaviors?

      The stronger retraction forces observed during probing indeed suggest that backward pulling is not a passive withdrawal, but likely an active component of tissue disruption. As discussed in the manuscript (lines 475–482), we interpret these repeated pullback movements, together with the outward-facing prestomal teeth of the everted labellum, primarily as a mechanism to enlarge the feeding lesion and improve access to blood, consistent with the blood pool feeding strategy of tsetse flies. To make this more clear, we have added a half sentence to line 482 "..., thereby creating a larger blood pool for feeding."

      We also already compare this mechanism to mosquito feeding mechanics in the discussion (starting from line 487). In mosquitoes, high-frequency fascicle oscillations are thought to reduce insertion resistance and facilitate minimally invasive capillary feeding. Although we also observed oscillatory movements during tsetse feeding (Video 4), the underlying mechanical strategy appears fundamentally different. In contrast to the mosquito’s system optimized for delicate penetration, the tsetse proboscis appears adapted for forceful tissue disruption during pool feeding. Notably, the oscillations observed in tsetse flies seem to occur during active blood uptake rather than initial tissue penetration. Consequently, the functional role of these oscillations in tsetse flies remains unclear. We have now addressed this more specifically in the discussion (lines 490-495):

      "Oscillatory movements were also observed during tsetse probing (Video 4). Notably, these oscillations appeared predominantly during active blood uptake rather than during the initial penetration phase, suggesting that they are associated with ingestion rather than insertion. Whether they facilitate blood flow, prevent occlusion of the feeding canal, or simply reflect pump activity remains unknown."

      When looking at other species, stable flies (Stomoxys) may represent a particularly relevant comparison because they employ a similar penetration mechanism and are pool feeders with prominent prestomal teeth (Krenn and Aspöck. Function and evolution of the mouthparts of blood-feeding Arthropoda. Arthropod structure and development, 2012). In contrast, tabanids employ a different mouthpart architecture with rasping/cutting structures but without comparable prestomal teeth. Whereas mosquito mouthparts have been described as functioning like a syringe, and we compare the tsetse proboscis to a saw, tabanid mouthparts have been likened to scissors (Krenn and Aspöck. Function and evolution of the mouthparts of blood-feeding Arthropoda. Arthropod structure and development, 2012). Although tabanids are also known to inflict substantial tissue damage, it remains unclear whether their feeding movements produce retraction-dominated force patterns comparable to those we observed in tsetse flies.

      Lastly, we agree that the elevated resistance generated during retraction may contribute to withstanding host defensive behaviour such as shake-off responses. Structurally, the orientation of the prestomal teeth and the architecture of the everted labellum could provide temporary anchoring during feeding, as we have already briefly discussed in the manuscript (lines 458– 461). However, while stronger anchoring may increase feeding stability, it could also increase the risk of injury to the fly if detected by the host. Compared to other pool-feeding flies such as stable flies, tsetse flies have been reported to respond more readily to host defensive behaviour (Schofield and Torr. A comparison of the feeding behaviour of tsetse and stable flies. Medical and Veterinary Entomology, 2002). We therefore currently consider anchoring to be a possible secondary function but lack direct experimental evidence to assess its practical importance.

      (3) The study analyzes a diverse set of substrates, which is a strength. However, some caveats deserve explicit discussion. Human skin equivalents and dermal equivalents lack the full mechanical complexity of real skin (e.g., innervation, perfusion, tension). Frozen or ethanol-stored samples, particularly reptile skin, may also exhibit altered mechanical properties compared to live tissues. These limitations do not undermine the findings but should be explicitly acknowledged as they influence the interpretation of absolute force magnitudes.

      The reviewer raises a valid point regarding the interpretation of absolute force magnitudes across the measured substrates. We have therefore added a clarifying statement to the discussion (lines 469-476):

      "When interpreting absolute force magnitudes, it is important to bear in mind that our samples do not fully recapitulate physiological conditions. Skin explants and skin equivalents may behave differently to skin under active perfusion and native tissue tension, as may our fixed and frozen animal skin samples. Nevertheless, comparative force measurements revealed consistent biomechanical signatures across substrates, suggesting that the observed force patterns reflect fundamental aspects of the feeding mechanism that are likely relevant in vivo.”

      (4) The SEM and FIB-SEM images showing trypanosomes in the hypopharynx and surrounding tissue during penetration are visually striking and suggest rapid dispersal. It would be helpful to connect these observations more clearly to the kinetics of parasite deposition and whether mechanical tissue laceration is likely to increase inoculation efficiency. Without conducting additional experiments, the authors could discuss whether these findings support or modify existing models of salivary-gland-derived parasite release.

      We have now expanded the Discussion to clarify that our observations of trypanosomes in the hypopharynx are consistent with the established model of salivary-gland-derived parasite release during probing and feeding, in which infective metacyclic trypanosomes are delivered with saliva into the host tissue. Furthermore, the presence of trypanosomes beyond the immediate feeding canal supports rapid parasite dispersal following inoculation, as described in previous work (Reuter et al., 2023). In this context, the tissue laceration generated by the tsetse proboscis may facilitate local parasite distribution by creating a larger, mechanically disrupted feeding lesion. However, our data provide high-resolution structural snapshots and were not designed to quantify deposition kinetics or inoculation efficiency. We therefore refrain from concluding that mechanical laceration increases transmission efficiency and instead view this as a plausible consequence that should be tested directly in future work. Specifically, we have added this paragraph to the discussion (521-527):

      "Overall, our observations of trypanosomes within the fly's hypopharynx, labial gutter, and host tissue are consistent with the established model of salivary-gland-derived parasite release during probing and feeding. Their presence beyond the immediate feeding canal is consistent with rapid local dispersal following inoculation, as described previously (Reuter et al., 2023). This process may be facilitated by the extensive tissue disruption caused by the tsetse mouthparts, although this hypothesis will require direct experimental testing."

      (5) The authors demonstrate that tsetse attachment abilities fall within the range of generalist insects and are far lower than those of obligate ectoparasites. However, the manuscript could discuss how attachment forces relate to the tsetse's ecological context, e.g., whether their attachment is generally brief, whether host shaking strongly selects for grip strength, etc. Is there evidence that other Glossina species or tabanids with different host preferences show variation in attachment performance? This would broaden the relevance of the findings.

      Tsetse flies are obligate blood feeders, but host contact is typically brief and frequently interrupted by host defensive behaviour. As a result, selection may favour rapid and efficient feeding rather than exceptionally strong attachment. This interpretation is supported by Schofield and Torr (A comparison of the feeding behaviour of tsetse and stable flies. Medical and Veterinary Entomology, 2002), showing that tsetse flies experience more feeding interruptions than the stable fly Stomoxys calcitrans, despite completing successful blood meals in less time. These differences are consistent with life-history theory (Anderson and Roitberg. Modelling trade-offs between mortality and fitness associated with persistent blood feeding by mosquitoes. Ecology Letters, 1999), which predicts that long-lived species with low reproductive rates, such as tsetse flies, should be less willing to risk injury by persisting on a host than shorter-lived, more fecund species. Against this background, our finding that tsetse attachment forces fall within the range reported for generalist insects, appears biologically plausible. Their attachment performance needs to be functionally sufficient for brief feeding events rather than maximized for prolonged host retention.

      We are not aware of comparative biomechanical data on attachment performance across different Glossina species or tabanids. We agree that such comparative studies would be valuable to test whether differences in host preference and feeding ecology correlate with variation in attachment capacity.

      (6) In video 4, could the authors clarify whether the observed maxillary vibrations are hypothesized to reduce penetration resistance or serve another function?

      The vibrations of the maxilla specifically appear during active blood uptake rather than during initial tissue penetration, suggesting they are linked to the ingestion phase. Whether they serve a mechanical function, such as facilitating blood flow or preventing canal occlusion, or represent a passive consequence of pump activity, remains unclear. We consider this an open and interesting question that warrants dedicated investigation.

      We have therefore clarified that the functional significance of these oscillations remains unresolved to date (lines 490-495). This reads: “Oscillatory movements were also observed during tsetse probing (Video 4). Notably, these oscillations appeared predominantly during active blood uptake rather than during the initial penetration phase, suggesting that they are associated with ingestion rather than insertion. Whether they facilitate blood flow, prevent occlusion of the feeding canal, or simply reflect pump activity remains unknown.”

      Reviewer #3 (Public review):

      Summary:

      Human and animal trypanosomiasis are fatal illnesses caused by African trypanosomes transmitted by tsetse flies during a bloodmeal. Thus, tsetse fly feeding is the key physical step in disease transmission to mammals. Tsetse fly feeding is not a new story, but it is revisited here through the application of sophisticated imaging techniques and novel biomechanical methods of analysis. The authors aim to provide a high-resolution picture of the structures and forces involved in feeding to provide mechanistic insights into the process of feeding, from attachment, penetration, drinking and retraction of the feeding parts.

      Largely, the authors have achieved their aims. They (i) examine the structures and forces involved in attachment; (ii) they provide detailed multi image analysis of the proboscis providing insights into its probing ability and physical mechanism of penetration; (iii) they conduct a controlled analysis of the physical forces involved in penetration and report that they are in the low nM range, not especially strong but much higher that the mosquito bite and finally they provide a first analysis of blood uptake during feeding.

      Strengths:

      The study images the tsetse fly feeding structures in unprecedented detail, with resolution to the uM scale, in 3-D, and during feeding. The resulting images are dramatic and insightful (and beautiful and frightening!), so researchers interested in trypanosomes, tsetse flies, or blood feeding by flies in general will want to see.

      They conclude that flies attach strongly to smooth surfaces because of interactions possible via the array of acanthae of the pulvillus pad at the ends of the tarsi. The estimated attachment forces are similar in male & female flies, in the low mM range (they look impressively strong in video 1). They provide a very striking analysis of the proboscis and labellum and associated tooth structures (Figures 4 & 5). I recall many years ago observing that tsetse flies are messy feeders, and these structures, especially the rasping teeth structures on the reverse folded labial tips, explain why! This seems more like a chainsaw than a jigsaw in action, but the authors are probably correct that these structures and the probing/retraction mechanism explain many features of tsetse fly feeding and their ability to feed on a wide range of hosts with very different skin types.

      We agree that “jigsaw” may be too specific and not fully appropriate in this context. We have therefore replaced it in the manuscript with the more general term “saw.”

      The impressive aspect of this paper is the range of imaging techniques (CLSM, SEM, uCT, FIB SEM), the quality of the images, which attests to the obvious care taken with sample preparation. The biomechanical analysis, especially the penetration analysis, is impressive. Finally, the paper is clearly written and presented; it was a very easy read and, overall, a very engaging study.

      Weaknesses:

      I suppose it could be said that the paper is a descriptive study; it doesn't really test a hypothesis, but that is not a prerequisite for sharing it. Perhaps the least convincing parts are the imaging of the flexible versus rigid parts of the structures, which is based on the amount of resilin (flexible) and chitin-protein (stiff), based on their autofluorescence. It seems odd that the joints would be less blue (stiffer) in Figure 1i, or what the blue structures correspond to in Figure 6B-D.

      Our analysis is based on established CLSM approaches that use exoskeleton autofluorescence as a proxy for relative differences in cuticular composition and material properties (Michels & Gorb, 2012; Michels et al., 2016). In the tarsus, the observed differences in inferred stiffness are relatively subtle, with most regions exhibiting broadly comparable material properties. This becomes particularly evident when compared with the proboscis, where the contrasts in cuticular composition are much more pronounced (Figure 6). We also note that locally stiffer regions at joints are not unexpected, as stiffness gradients in arthropod joints can provide mechanical support and help constrain the direction of movement. Importantly, our images show a flexible, ring-like blue region directly at the articulation, surrounded by slightly stiffer material. We therefore interpret this pattern as a combination of a flexible hinge region and adjacent supporting structures that together enable controlled joint motion.

      The blue structures in Figure 6B–D correspond to flexible regions of the furca (f). Because this spring-like cuticular element undergoes substantial configuration changes during labellar eversion, the presence of highly flexible regions is consistent with its proposed mechanical function.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further experiments or analyses are suggested. However, the Discussion would benefit from briefly acknowledging how trypanosome infection can alter feeding behavior and mouthpart function, based on prior work, to place the mechanical findings in a more biologically relevant transmission context.

      We thank the reviewer for this suggestion. Previous studies have indeed shown that trypanosome infection can alter tsetse feeding behavior, primarily through changes in saliva composition. Van den Abbeele et al. (2010) demonstrated that infection with T. brucei significantly impairs the anti-haemostatic activity of tsetse saliva, resulting in prolonged prefeeding probing and therefore extended feeding times. These findings are consistent with earlier observations by Jenni (1980), who reported increased probing frequency in infected flies.

      Jenni (1980) also proposed that these behavioral changes might be linked to altered mechanoreceptor function. However, Van den Abbeele et al. (2010) argued against this interpretation for T. brucei, noting that this parasite does not colonize the mouthparts where these mechanoreceptors are located. Taken together, the available evidence suggests that the observed changes in feeding behavior are mediated primarily through altered interactions with host blood rather than through direct effects on the mouthparts themselves.

      It should be noted that this conclusion is specific to T. brucei. Other tsetse-transmitted trypanosome species, such as T. congolense, do colonize the proboscis. However, a comparative study examining flies infected with T. brucei, T. congolense, or T. vivax found no significant effects of infection on feeding behaviour relative to uninfected controls (Moloo, 1983). To our knowledge, there is currently also no direct evidence that any trypanosome species alters the physical properties or mechanical function of the mouthparts, or causes damage that would directly affect feeding performance. We have added this paragraph to the Discussion (lines 508520):

      "It is also important to note that our experiments were conducted using uninfected flies. Previous studies have shown that trypanosome infection can alter feeding behaviour, including increased probing activity and prolonged feeding times (Jenni et al., 1980; Van den Abbeele et al., 2010). These effects have primarily been attributed to infection-induced changes in saliva composition and the resulting interactions with host blood (Van den Abbeele et al., 2010). However, a different study found no significant effects of infection with either salivary gland-resident T. brucei or with proboscis-colonizing species such as T. congolense and T. vivax on Glossina feeding behaviour (Moloo, 1983). Furthermore, although infection-associated transcriptional changes in the salivary glands and proboscis have been reported (Awuoche et al., 2017; Matetovici et al., 2016), there is currently no direct evidence that trypanosome infection alters the mechanical properties or function of the mouthparts themselves."

      Reviewer #2 (Recommendations for the authors):

      Several figures (particularly SEM-based ones) contain very dense labeling. Consider providing simplified overviews or annotated "orientation guides" in figure supplements to improve navigability for readers unfamiliar with proboscis anatomy.

      We thank the reviewer for this helpful suggestion. While we agree that orientation aids can be valuable, we have decided not to include additional simplified overview figures, as we consider that introducing separate schematic summaries could potentially complicate rather than improve navigation of the structural detail. We therefore rely on consistent labelling within the existing figures and detailed captions to guide interpretation.

      The manuscript uses appropriate non-parametric tests, but could benefit from reporting effect sizes and indicating sample sizes on all plots.

      Sample sizes are reported in the figure legends, Methods section, and Supplementary material for all experiments. We agree that reporting effect sizes can be informative and will consider this in future studies. However, because the primary objective of the statistical analyses in the present work was to support comparisons between experimental conditions rather than to estimate effect magnitudes, and because the figures are already information-dense, we therefore decided not to further modify the graphical presentation in this revision.

      Reviewer #3 (Recommendations for the authors):

      (1) P5 L111. Perhaps indicate these knobs on the image Figure 1S). I assume these are the structures visible under the pointer labelled spa? Maybe highlight some of the worn areas in Figure 1G.

      The knob-like structures in Figure S1 are highlighted in green and we have now revised the figure description from:

      “…showing fine crests on the underside and surface modifications (green) on the upper side.”

      to:

      “…showing fine crests on the underside and knob-like surface modifications (green) on the upper side.”

      Regarding Figure 1G, the purpose of the panel is to illustrate the contrast between deformed spatulae (Figure 1G) and intact spatulae (Figure 1H). We therefore chose to retain the original presentation, as we feel that additional markings would not substantially improve interpretation and could obscure structural details. We hope that the direct comparison between the two panels provides sufficient visual guidance.

      (2) P9. The frictional force (and P38/39) has the units of N (kg.m/Sexp2). The safety factor is this force divided by the weight of the fly? So are there units (Kg/sexp2) or are these not shown? Perhaps this is a convention.

      The safety factor is defined as the ratio of the total frictional force to the fly’s weight force (m·g), where m is body mass and g is gravitational acceleration. Since both quantities are express in Newtons (kg·m·s<sup>-2</sup>), the safety factor is dimensionless.

      We agree that the terminology in the original manuscript may have been ambiguous, as “body weight” is sometimes used colloquially to refer to body mass. To avoid confusion, we have revised the text to explicitly refer to weight force and now define the safety factor as the total friction force divided by weight force (mg, where m is body mass and g is gravitational acceleration). We have clarified this in the main text, the Figure 2 legend, and the description of Supplementary Material 1.

      (3) P10 Figure 2G & H. It is not very clear...are these the data, the average of all readings across all surfaces in E and F? If so, why is this value useful...how does it add to what is already shown?

      The figures 2G and 2H summarize the friction forces (G) and safety factors (H) across all tested substrates, based on the values from the male (B, E) and female (C, F) datasets. The purpose of these panels is to provide an overall comparison between sexes independent of substrate type. While this information can also be inferred from the substrate-specific plots, the sex-separated presentation does not make the absence of an overall sex difference immediately obvious. Figures 2G and 2H therefore serve as concise summary plots highlighting this result.

      (4) P12. For the nonspecialist, it might be useful to draw a cartoon showing the organisation of the labium, labrum and the hypopharynx...this is visible in Figure 4i but not in the dissected proboscis and labellum ....only the labium as the labrum doesn't extend this far?

      To clarify the anatomical arrangement in the dissected specimen, we have added the following statement to the Figure 4 legend (lines 224–226):

      “In an intact fly, the labrum would be positioned within the empty groove of the labium visible in J; however, it is absent in this dissected preparation.”

      (5) P17 legend to Figure 5. Include the Lm abbreviation in the legend, and maybe a close-up of the rsp teeth?

      We have added “lm, labellum” to the Figure 5 legend (line 250), as this abbreviation was previously missing. Panel J is a close-up of the rasping teeth.

      (6) F3S and Video 3. Are the images in B and C taken from the FIB SEM video images? It is not clear. A small legend descriptor for video 3 would be helpful.

      The images in Supplementary Figure 3B and C are reconstructed from the same FIB-SEM dataset shown in Video 3, but they are displayed in a different orientation. This is indicated schematically in Supplementary Figure 3A, which illustrates the viewing plane used for the reconstruction.

      We already included the following legend for Video 3 (lines 1146–1149):

      "Video 3: FIB-SEM of the tsetse labellum. Sequential cross sections reveal internal ultrastructure progressing from near the tip of the labellum downward. Data were acquired on a Crossbeam 540 (Zeiss) with the EsB detector in continuous milling mode."

      To improve clarity, we have now added a sentence to the video legend linking the figures to the video: (lines 1149-1150)

      “Reconstructed images from this dataset are shown in Figure 5A and Supplementary Figure 3B and C.”

      In addition, we have now explicitly cross-referenced Video 3 in the legends of Figures 5 and S3 to make the connection clearer for the reader.

      (7) Figure 7. These are amazing images, especially G-I.

      Thank you for this positive feedback, we appreciate it.

      (8) P24. It is really good to see that there is a difference in force penetration for full skin v dermal...this deserves a comment.

      We agree and have revised the text accordingly. We replaced:

      "Human skin substrates required the lowest penetration forces, with 0.97 mN for full-thickness skin equivalents, 0.67 mN for dermal equivalents, and 0.85 mN for native skin explants (Figure 8C, D)."

      With this (lines 363-367):

      "Human skin substrates showed the lowest penetration forces, with dermal equivalents requiring less force (0.67 mN) than full-thickness skin equivalents (0.97 mN), reflecting the additional mechanical resistance of the epidermal layer absent in dermal-only constructs. Native skin explants fell intermediate at 0.85 mN (Figure 8C, D)."

      (9) P26 Figure S5. Panel c, there seems to be a big scatter in the drinking time. Was there an outlier?

      Indeed, the observed scatter is due to a single fly with an unusually long drinking time of 184.44 seconds, which is approximately six times the median duration. We have verified the underlying data and found no indication of a measurement error; the value therefore remains included in the analysis. The data for the plots in Supplementary Figure 5 are also available in Supplementary Material 3.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental yet poorly understood biological process.

      Comments on revised version.

      The manuscript overall is significantly improved and authors addressed majority of my concerns. The addition of the computational vertex model (Figure 7) as well as Atg1 RNAi (Figure 4) to inhibit cell fusion provide more mechanistic insight to their study. However, the analysis of Atg1 RNAi wound assay falls short as it does directly measure changes in syncytium frequency nor size to confirm that cell fusion is reduced. The authors should quantify the number of nuclei per syncytium over the 2hr wound healing period as performed for WT in Figure 1C. It would have been ideal if they could have also performed the Act-GFP spreading assay in WT and Atg1 RNAi strains to determine if Act-GFP movement is dependent on cell fusion as purposed. At the least, further quantification of Atg1 RNAi phenotype is warranted to support their conclusions.

      In response to the reviewer's comment, we have repeated the analysis of Fig 1C and generated a new panel, Fig. 4C, which is directly comparable to the control and shows that syncytial size is dramatically reduced in the Atg1 knockdown area. Unfortunately, we cannot perform the second analysis of actin-GFP spreading in the Atg1 knockdown cells because we need Gal4 for labeling individual cells and for knocking down Atg1, and we can't do both at the same time.

      Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laser-induced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFP-positive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to syncytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Comments on revised version:

      The authors have extended their original manuscript by adding two key parts. First, they show a role of Atg1 in mediating cell fusion (Figure 4). Second, they provide additional evidence for a contribution of radial border fusions to wound closure through its effect on tissue fluidity and through computational modelling (Figure 7).

      This new version of the manuscript is greatly improved and provides significant new insights into the role of syncytia in aiding wound repair. There are just a few minor, yet important, additions needed to back up Figure 4 which should not require new experiments.

      Minor but important points:

      The authors show a role of Atg1 in mediating syncytia formation in Figure 4. However, since the Pnr>+ side of the wound closes slower than the non-Pnr side (control side), a few additions to this figure would be important and should not require additional experiments.

      (1) The authors should show, similar to the data shown in Figure 4D of the wound radius over time for control versus Pnr>Atg1RNAi, also the same type of data for control versus Pnr>+.

      The data the reviewer requests is available in our bioRxiv manuscript, in Fig. 6B (Hua, Krystofiak, Pumford, Page-McCaw, and Hutson, https://doi.org/10.64898/2026.05.31.728998). These experiments were all done at the same time. As you can see, the difference in closure rate is quite subtle in control wounds.

      (2) Since Pnr>+ also slows down wound healing, albeit to a lesser extent than Pnr>Atg1, the authors should also show an extra graph that provides evidence that Pnr>Atg1RNAi reduces syncytia formation more than Pnr>+ does. E.g. Two graphs could be added that show individual cell size at 4 or 5h post wounding for control versus Pnr>Atg1RNAi as well as for control versus Pnr>+ and also another graph with the same data but comparing cell size between Pnr>+ and Pnr>Atg1RNAi. Otherwise, if the expected minimum cell size for a syncytium is easy to estimate, a graph could be added that shows the percentage of cells that are above this threshold (e.g. above 100 square micron) for control versus Pnr>Atg1RNAi and control versus Pnr>+ and Pnr>+ versus Pnr>Atg1RNAi.

      In response to this comment and the comment from reviewer 1, we have now added new Fig. 4C, which addresses the reviewer's question about the comparative frequency of fusion in pnr>Atg1RNAi and pnr>+. These graphs show that Atg1 knockdown significantly reduces the size of syncytia.

      Reviewer #3 (Public Review):

      In this revised manuscript, White et al. aimed to understand the wound-induced syncytia formation behavior in wound repair of Drosophila melanogaster pupal notum. For this purpose, the authors characterized two different types of adherens junctions' outcomes during syncytia formation around the wound region - border breakdown versus apical shrinking which appear to happen in different time points and for different time durations. The authors characterized cell-cell fusion events using cytoplasmic, junctional and nuclear markers. They determined that about half of the cells within 70 um radii from the wound undergo cell-cell fusion. They studied wound induction on the border between control epithelia and pnr domain suggesting that Atg1 is required for post-wound syncytia formation and wound closure. They showed that during wound closure syncytia gradually invade the wound leading edge mostly by radial fusion events. The data suggests that intercalation of cells from the leading edge slows down the wound closure process. They propose that cell fluidity of syncytial cells plays a role in wound closure speed. Finally, the authors showed that actin is concentrated to the front edge of syncytia located in the wound leading edge. The authors described some aspects of syncytia formation during wound closure using different approaches. Some clarifications are needed as described below.

      Major suggestions:

      (1) Introduction, page 4. The examples of developmental syncytia formation of invertebrates and vertebrates are confusing. The authors may want to make the examples clear and add additional examples. Currently, readers may assume that C. elegans cell fusions occur only in the hypodermis - other structures can be mentioned like the vulva, pharyngeal muscles, glia, tail. In addition, the authors may want to add injury-induced fusions like the C. elegans' PLM and PVD neurons (Ghosh-Roy et al., 2010; Newman et al., 2015; Oren-Suissa et al., 2017).

      We appreciate the suggestions and have included the additional examples of C. elegans vulva and PLM and PVD neurons. We are limiting ourselves to those because we don't want to focus too heavily on C. elegans examples, as that's not the direction this paper is heading.

      (2) In cases where it is not clear whether fusion has occurred or whether mononucleated cells were ejected from the leading edge, membrane markers can be used. Page 6. Lines 96-99. The authors may want to use a membrane marker like RFP-PH driven by the epithelial cell promoter.

      At this point in the manuscript, we are introducing syncytia and are not concerned yet with their origin. 

      (3) Pages 8-10. The authors may want to clearly explain that apical junctions shrinking is a post fusion event. That the apical shrinking is caused by the expansion of fusion pores and the migration of apical junctions towards the basolateral domain. This is something that was clearly shown during physiological epidermal cell-cell fusion in C. elegans by Mohler et al., 1998 and 2002. A cartoon showing the process of cell-cell fusion, pore expansion and apical junction dynamics would make the manuscript much clearer.

      Apical shrinking cannot be caused by the "migration of apical junctions towards the basolateral domain" because that is not what we observed -- rather, we observed labeled adherens junctions remaining at the apical surface while the area they enclose becomes smaller (shrinks). Further, despite close reading of the Mohler papers, it is not clear how similar the apical shrinking events of this manuscript are to the fusion events described there. Finally, we do not want to include a schematic describing this process because that would suggest certainty that we do not have. Unlike in C. elegans development, wound-induced cell fusion is stochastic, not stereotyped; with cells that display apical shrinking, the fusion partner of a labeled cell is difficult to identify because it is often not a neighboring cell. These factors make it difficult to describe this process in detail, but we have sufficient data to conclude that these are indeed cell fusion events.

      (4) Page 9. Line 170. "...as these cells represent fusion initiation events (fusion pore) but were unable to productively stabilize and expand the site of fusion and so returned to the diploid state." The authors may want to make clear that this is an assumption that needs to be tested. Live imaging using a membrane marker may resolve whether a reversible fusion pore was generated.

      Thank you for the suggestion; we have updated this text to make it clear that this is an interpretation.

      (5) Page 11. It is not clear whether Atg1 is directly required for cell fusion, or that autophagy is required for efficient cell fusion or both Atg1 and autophagy participate in the fusion process.

      Our data show that Atg1 is required for cell fusion. The work that inspired this experiment, Kakanj et al 2022, concluded from their more comprehensive studies that the process of autophagy was required. We have clarified the text.

      (6) Page 12. Line 235. "Indeed, we observed that several hours after wounding, the entire leading edge was occupied by syncytia." This observation is based only on the adherens junction marker. Can they test basal cell membrane marker? Is it possible that the mononucleate cell in the leading edge is under the two syncytia?

      Unfortunately, there are not good basal markers -- the recently reported basal spot markers also label adherens junctions. Nonetheless, we are confident that the mononuclear cell is not under the syncytia because we image Z-stacks and thus can detect cell overlap.

      Recommendations for the authors:

      Reviewer #3 (Recommendations For The Authors):

      Minor suggestions:

      (1) Figure 1. The authors may want to add an image immediately after laser ablation of the actual wound and the area around the wound. Add an arrow to mark the wound.

      With this wounding modality, the extent of the wound is unclear for ~30 min. As we reported in O'Connor et al, PLoS One, 2021, there is a gradient of damage emanating out from the center of the wound, and cells with greater amounts of damage die while those with less damage repair and survive. Immediately after laser ablation, very little visible damage is evident by 120ctn-RFP and Histone-GFP (the markers in Fig. 1) until the cells die and the surrounding cells respond.

      (2) Page 6. Line 86. "A mitotic tissue utilizes cell-cell fusions during wound repair." replace "during wound repair" with "after wound induction" since in this section the authors do not show that this process is part of wound repair.

      Thank you for the suggestion - we reworded this heading to remove "wound repair".

      (3) Page 6. Line 92. The authors may want to be consistent with the terms used in the text and in the figure - His2GFP in the text versus Histone GFP in the figures.

      Thank you for the suggestion, we have revised for consistency.

      (4) Figure 1 - supplement figure 1D. The "v" of Div panel moved below D.

      Thank you, we have corrected it.

      (5) Figure 1 - supplement figure 1G. add "i" to second Gii to make it Giii.

      Thank you, we have corrected it.

      (6) Page 24. Figure 1H legend. 3 or 4 wounds?

      Thank you for catching this error - 4 wounds.

      (7) Page 7. Line 124. "GFP mixing always preceded border breakdowns (n=11)" instead of "always" use "in all observed cases".

      We have made this change.

      (8) Figure 2. Switch the writing "Apical Shrinking: Nuclear Transfer" since apical shrinking represented in panel 2A and Nuclear Transfer in panel 2B. If this description applies only to panel 2B, make it clear.

      We consider this heading to apply to panels A and B together (as they show the same sample, just different channels).

      (9) Figure 2C. Is ActinGFP a cytoplasmic GFP driven by actin promoter or Actin-bound GFP? Cytoplasmic GFP versus membrane-cortex GFP?

      It is a transgene expressing an actin-GFP fusion protein, as noted in the key reagents table and discussed in Fig. 8. We corrected the manuscript to ensure it is always referred to now as Actin-GFP in the text, figures, and legends.

      (10) Video 3 - Impressive movie!

      Thank you!

      (11) Page 9. Line 155. "In both these cells, as the cell lost its basal volume, cytoplasm moved laterally to join the neighboring syncytia." It seems that the apical shrinking cells' cytoplasm joined the neighboring syncytia even before.

      Because both indicated cells (yellow and white arrows) and the neighboring syncytium are all labeled with GFP, it is not possible to determine precisely when the cells' cytoplasm joined the syncytium.

      (12) Page 9. Line 158. "...but fusions associated with apical shrinking occurred later and were more numerous." Did the fusion occur later or the apical shrinking itself as was mentioned before and shown in Figure 2F?

      We have changed the wording, as for many apical shrinking events we cannot tell exactly when the fusions were initiated.

      (13) Page 25. Figure 3A legend. What is the meaning of morphological fusion? Border breakdown and apical shrinking? The authors may want to define it.

      We have defined it now in the legend.

      (14) Page 26. Figure 3B-C legend. "Panel C shows that apical shrinking fusion and border-breakdown fusion occur at similar distances from the wound." It seems that fusion by apical shrinking mostly occurs within 60-70 um from wound center and fusion by breakdown occurs equally at all distances up to 80 um.

      We don't disagree with your comment, but we feel the dataset is too small to make such a statement. The data is presented so the interested reader can make their own conclusion.

      (15) Page 9. Line 165. "...but infrequently (n=3) with GFP mixing and no subsequent cell fusion..." Does this mean that there were GFP mixing without border breakdown or apical shrinking?

      Yes, that is correct. We assume that in this case a fusion pore opened and then closed again. We have added a phrase to clarify.

      (16) Page 9. Line 175. "...the spatial distribution of fusing cells that shrank vs. lost borders was similar (compare Figures 1G and 2E)." Even though the visual comparison suggests similar spatial distribution, the carefully quantified distribution in figure 3C suggests more fusion by shrinkage at 60-70 um from wound center of the 5 tested wounds.

      As we noted to comment 14, we feel the data set is too small to make such a statement. The data is presented so the interested reader can make their own conclusion.

      (17) Figure 3. The shown pies sum the results from 5 wounds. It would be interesting to add a graph comparing the percentage of fused and persisted cells per wound to see the variability, if exists.

      Unfortunately, the number of fused/persisting cells in each wound is greatly affected by the heat-shock conditions that generate the labeled clones; even the ratio of these fates would be heavily influenced by noise because the numbers are small in each animal. Further, the frequency of fusion is determined by the wound size as shown in Fig. 1. Because of these variables, such data could be easily misinterpreted.

      (18) Figure 3 - figure supplement 1D. Even though it was mentioned that the duration of some border breakdown is finished within minutes it is worth comparing it with shrinking duration on one graph.

      Unlike apical shrinking, it is difficult to identify exactly when border breakdown concludes, so this data is difficult to compare. We have provided several examples of border breakdown in the manuscript that give an overview of the process.

      (19) Video 1 is not mentioned in the main text.

      Thank you for catching that omission. We now refer to it in the first paragraph of the results.

      (20) Figure 4B. The difference between the treated group and the control group is unclear. Add arrows.

      We have added some arrows to Fig. 4B.

      (21) Figure 4C. For consistency use percentage for both border breakdown and shrinking cells.

      In response to the reviewer's comment, we now provide the consistent metric of number of lost borders and number of shrinking cells.

      (22) Page 11. Did the authors try other wound types (e.g. mechanical/chemical wounds)? May other wound causes besides laser ablation result in different response? This may help to answer whether there is a causation between syncytia formation and speed wound closure.

      There are reports of puncture and pinch wounds inducing cell fusion. Perhaps the reviewer is suggesting that we might be able to identify a wounding method that does not induce cell fusion and then compare the rate of wound closure. However, another type of wound would probably inflict different amounts of cell damage and so would be hard to compare. Overall, we think the half-and-half system of comparing responses on the two sides of the wound is the best, most controlled comparison.

      (23) Figure 4F-G. It was mentioned that there is less syncytia formation in Atg KD cells, however the difference in cell area between control, WT and Atg KD is not obvious. The authors may want to mark the dots that represent syncytia to distinguish them from mononucleated cells.

      The point we are trying to make (now Fig. 4G-H) is that cell area is related to distance moved, regardless of how cell area is determined. We do not have the ability to count nuclei in the control sides (nuclei are labeled only on the pnr side), and further, we have reported separately (White et al, 2024) that there is a limited amount of endocycling in these cells, which should also increase area.

      (24) Figure 5G. y axis. The authors may want to change "small cells" to "mononucleate cells".

      We changed it to "unfused cells" which is the term we used in the legend. In these wounds we were unable to visualize nuclei.

      (25) Page 12. Line 243. (Figure 5D,G) instead (Figure 5D).

      We changed it to read (Figure 5D,G).

      (26) Page 12-13. Lines 241-246. The description of "mononuclear cells removed" and "syncytia outcompete unfused cells" may be clearer if explained here as mononuclear cells joining the syncytium by cell-cell fusion.

      Here we are describing a different phenomenon - not that fusion is removing all the smaller cells but rather that the syncytia are faster/better/more effective at wound closure than the smaller cells. This is illustrated in Fig. 5Cii-Ciii.

      (27) Page 13. Line 261. "Thus, there were about five-fold more tangential borders lost to fusion than radial" Is this conclusion also true when analyzing each wound individually?

      This is a reproducible finding, that there is more fusion across tangential borders than across radial borders. The ratio of tangential-border loss: radial-border loss for each wound is as follows:

      wound 1, 63:12

      wound 2, 39: 11

      wound 3, 44:8

      wound 4, 50:8

      (28) Page 31. Figure 6 - figure supplement 1 legend, Line 652. Make "B)" bold.

      Done.

      (29) Page 14. Lines 270-275. If there is an advantage to radial fusion versus cell intercalation for wound closure speed, how do the authors explain that the percentage of radial fusion is lower than the percentage of intercalation? (Figure 6D) How does the wound affect the molecular level (fusogen expression?) of the surrounding cells? 

      We expect that radial fusion specifically reduces the need for intercalation at the leading edge, as shown in Fig. 6C. Both would speed closure, however, as any increase in cell area will allow more efficient redistribution of resources such as actin and will also reduce the total number of junctions needing to be remodeled as the wound closes. Since we don't know the fusogen, we can't say how the wound affects its distribution.

      (30) Page 14. It is not clear where the experimental data ends and the model starts. For example, in line 276, it would be clearer to describe the "tissue fluidity as measured" or is it more precise to write instead "as estimated/calculated". The fusion between observations and model is confusing and maybe this should be unfused.

      This text, referring to the analysis in Fig. 7A, B, is not a computational model but rather a quantitative analysis of tissue fluidity as measured by a pre-existing metric, the shape index. This is experimental data. The computational model begins in the next paragraph, accompanying Fig. 7C, D. We edited the language slightly in this paragraph to clarify.

      (31) Figure 8 versus Figure 2C. Actin-bound GFP versus cytoplasmic GFP? Both mentioned as Actin GFP. Make it clear.

      They are indeed the same thing, actin protein fused to GFP, as described in the text and legend, and they are labeled identically.

      (32) Figure 8Biii, Div. Nice presentation of signal distribution between the cells.

      Thank you!

      (33) Page 30. Figure 8G legend. Lines 623-628. Is the shown mean profile plot based on specific images shown in Fi and Fv or just the cells represented there? Since Fi is a single z slice and Fv is maximum intensity projection which are not comparable.

      In response to the reviewer's question, we reanalyzed the image. Fig. 8G compares Z-projections.

      (34) Page 15. Line 303-304. "Tangential border fusions allow resources from distant cells to be mobilized to the wound edge." Does not this leading-edge actin localization happen in radial fusions close to the region of the wound?

      Fusions along radial borders, as shown in the top panel of Fig. 6A, would not offer the opportunity to move actin from distant cells to cells nearer to the wound.

      (35) Page 15. Did the authors test any predictions from the simulations of the model experimentally?

      This isn’t so much a predictive model as an exploratory model that addresses one question: is it plausible that the presence of syncytia can speed closure by reducing the need for intercalations, even if the syncytia have no other special properties. The only prediction would be that inhibiting fusion would slow wound closure.

      (36) Page 17. It would be interesting to discuss the following questions: (A) Is autophagy required for fusion. (B) Is Atg1 required for epithelial cell fusion? (C) Is autophagy required for wound repair? Are any of the combinations correct (A&B, A&C, B&C, A&B&C)

      The role of autophagy in wound-induced cell fusion was thoroughly explored in the 2022 EMBO J paper from Maria Leptin's lab, "Autophagy-mediated plasma membrane removal promotes the formation of epithelial syncytia" by Kakanj et al. We merely knockdown a gene they discovered to be important for wound-induced epithelial fusion, Atg1, as one means of investigating how syncytia contribute to wound closure. Our results don't add to their findings, and the role of autophagy is not what we want to focus on in our Discussion.

      (37) Page 19. Line 381-384. "If N represents the number of cells that fused, our results suggests that syncytia can apply up to N times more actin to the leading edge; considering that we observed syncytia with dozens of nuclei, this could represent a significant enhancement of actin at the leading edge. Increased actin might explain the ability of syncytia to outcompete diploid cells at the leading edge." To enhance this suggestion, the authors may want to compare actin signal in the leading edge of different size syncytia.

      We thought a lot about this experiment because reviewer 2 asked for it in the previous round of review, but as we said then, we can imagine too many caveats to the interpretation to make it worthwhile.

      (38) Page 22. Line 447-448. "...fusion would act the fastest after wounding because there is no need for DNA replication." There may be a potential need for protein (fusogen) synthesis.

      The timing of fusion, which we report here begins within 10 minutes after wounding, suggests that if there is a fusogen, it is already present in the cells before wounding.

      (39) Page 34. Line 716. Add "C" to "29{degree sign}".

      Done.

      (40) Page 38. "Wound closure analysis" part. Can the wound closure be visualized using brightfield?

      The scar also impedes imaging through bright-field microscopy.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment:

      This valuable study addresses the effects of selection on aggression on fitness and life-history trade-offs in Drosophila melanogaster. However, the evidence presented is incomplete and does not support the claims proposed in the study of increased survival of highly aggressive males at the expense of reproductive success and shorter mating duration. The main limitation of the study is the choice to use males from only one aggressive Drosophila line in combination with Canton-S females, that do not allow disambiguation between nonaggression-related factors, such as hybrid vigor and aggression-related factors influencing mating and lifespan.

      We would like to clarify the points raised in the eLife assessment.

      The report states that we relied on a single line of hyper-aggressive males tested with Canton-S females, and implies that Bully and Cs have not co-evolved. This is a misunderstanding: Bully flies were derived from Cs population. Thus, Bully and Cs have co-evolved. In addition to the Bully A line presented in the main figures of the manuscript, we replicated several of our findings with a second independent selected line, Bully B. Results from courtship assays involving both Bully A and Bully B couples males and females were presented in Figure Supp1. We apologies for not having made this more explicit in the original manuscript, which we will correct. These experiments should alleviate the concerns from the reviewers; they demonstrate that our conclusions are supported by two independent hyper-aggressive lines, and these include assays with selected male and female flies.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study asks how selection for male aggressiveness affects life-history and reproductive fitness traits in Drosophila melanogaster males.

      Strengths:

      Multiple comprehensive assays are used to address the question.

      We thank the reviewer for recognizing these strengths.

      Weaknesses:

      (1) The flies used for comparisons are inadequate. Behavioral assays compare Bully males mated to non-coevolved Cs females with Cs males mated to coevolved Cs females.

      We thank the reviewer for this comment, which made us realize that we had not sufficiently highlighted some of our experiments. The Bully lines used in our work were derived from Canton-S flies and thus did co-evolve with Cs. As originally described by Penn et al. (2010), highly aggressive “Bully” lines were generated through selective breeding from Canton-S males that consistently won aggressive encounters. After 34–37 generations, stable Bully lines were established. Thus, 1) Bully and Cs flies have co-evolved and 2) the selection applied was male-specific. Independent selection replicates produced distinct lines, including Bully A and Bully B. Previous studies only characterized Bully A (Penn et al., 2010; Chowdhury et al., 2017), but our work includes both Bully A and Bully B (Fig. S1).

      The rationale for pairing Bully or Cs males with Cs females (with which both male types co-evolved) follows the approach used by Dierick et al. (2006), who investigated how the male-specific selection for aggression affected courtship and mating behaviors by testing them with standard Canton-S females. This design allows to isolate the effects of male genotype and behavior on courtship and mating outcomes, avoiding confounding effects from female behavioral changes.

      We initially compared selected Bully pairs (Bully males × Bully females) (Fig. S1) with Cs pairs and observed similarly shortened mating durations in both Bully × Bully and Bully × Cs matings (Fig. S1, Fig. 1F and G). Thus, the reduction in mating duration arises specifically from Bully males. We therefore chose to use Cs females as a standard background to assess the consequences of male-specific selection for aggression on reproductive behaviors.

      (2) Lifespan analysis is done on male progeny of Cs females mated to either genetically more distant Bully or co-evolved Cs males; the longer lifespan and performance on the former is interpreted as a trade-off with aggressiveness, rather than a simple explanation of hybrid vigor.

      We appreciate this comment, which again stems from a poor explanation from our part about the origin of the Bully line in the original manuscript. The Bully flies were derived from the same original population as the Cs line. Hybrid vigor typically arises when crossing individuals from distinct populations, which is not the case here as both Bully and CS come from the same population.

      To further support our conclusions, we conducted additional experiments using progeny from within-line crosses (Bully males × Bully females) and results revealed the same phenotype: the progeny of these flies also exhibited significantly longer lifespans than Cs males x Cs females progeny. This finding argues against hybrid vigor as the main explanation for the observed phenotype, since both the Bully and Cs crosses result in inbreeding, yet give longer lifespan in Bully. We will include these additional longevity data (currently not included in the manuscript) to strengthen our results and reinforce our interpretation.

      (3) Differences in CHCs between Bully and Cs males and Cs females mated to those males are not shown to cause differences in measured behavioral outcomes.

      We thank the reviewer for raising this important point regarding causality. One way to establish a causal link between differences in CHCs observed in Bully and Cs flies and the corresponding behavioral outcomes would be to experimentally manipulate CHC profiles. For instance, one could perfume oenocyte-less males with the compounds found in higher abundance in Bully flies, then perform behavioral assays to assess causality. We agree that such experiments would be highly informative in determining the functional roles of specific CHCs elevated in Bully males. However, this approach is technically challenging, as the perfuming technique must be optimized to transfer precise amounts of each compound. For example, this method can be used to gradually perfume flies to assess dose–response behavioral effects, whereas matching exactly the natural concentrations found in individuals, especially given inter-individual variability, remains difficult.

      We considered conducting such experiments during our study but did not pursue them for these technical reasons. Nevertheless, we can include a statement in the Discussion acknowledging this as an important future direction to test the causal relationship between CHC variation and behavior.

      Reviewer #2 (Public review):

      Summary:

      The authors compare "Bully" lines, selected for male aggression, to Canton-S controls and find that Bully males have lower mating success, shorter mating durations, and remate sooner. Chemical analyses show Bully males have distinct cuticular hydrocarbons (CHC) signatures and transfer markedly less cVA to females, offering a plausible mechanistic link to weaker mate-guarding.

      Paradoxically, Bully males live longer and remain fertile at older ages when CS males no longer mate, indicating a shift in the reproduction-survival trade-off in aggression-selected populations.

      Importantly, the work sheds light on proximate mechanisms, demonstrating that shifts in CHCs and pheromone transfer co-occur with changes in fitness traits, thus offering new entry points for understanding life-history evolution.

      We thank the reviewer for this positive summary of our work.

      Strengths:

      The manuscript's strengths lie in its comprehensive and integrative approach framed within an evolutionary context. By combining behavioral assays, chemical profiling, and lifespan measurements, the authors reveal a coherent pattern linking aggression selection to life-history trade-offs. The direct quantification of cVA in female reproductive tracts after mating provides a particularly compelling mechanistic correlate, strengthening the link between behavior and chemical signaling. Findings on altered 5-T and 5-P levels further highlight how chemical communication shapes mating and mate-guarding strategies. Analytical approaches are largely rigorous, and the results provide valuable insights into the pleiotropic effects of selection on socially relevant traits. The study will be of interest to Drosophila biologists working on sexual selection, behavioral evolution, and aging.

      We thank the reviewer for recognizing the integrative design and mechanistic contributions of our study.

      Weaknesses:

      The weaknesses are primarily conceptual rather than procedural. The generality of the findings is uncertain, as selection appears to be represented by only one (and a second closely related) Bully line, limiting conclusions about selection responses versus line-specific drift or founder effects. The causal link between aggression selection and increased longevity is not established: the data show a correlated shift but do not identify mechanisms underlying lifespan extension. In several places, the manuscript uses causal language (e.g., that selection 'influences' longevity or mating strategy) where association would be more accurate; this should be toned down to avoid overstatement. Ecological relevance is also not addressed, since laboratory conditions may bias the balance between costs and benefits of aggression compared with variable natural environments. Addressing these points would strengthen both the impact and clarity of the study.

      (1) Generality of findings and potential line effects

      We agree that our results presented in the main figures of the manuscript relied mainly on one Bully line (Bully A). To address potential line-specific effects, we replicated key courtship experiments with another independent line, Bully B, selected in parallel from the same Canton-S stock but through distinct selection replicates. The results obtained from Bully B closely matched those from Bully A, suggesting that the observed phenotypes are consistent consequences of aggression selection rather than random drift or founder effects.

      (2) Causality versus correlation

      We concur that some sentences in the manuscript could overstate causal interpretations. We will revise the text to clearly distinguish correlation from causation and to avoid implying direct causal relationships where data only support association.

      (3) Ecological relevance

      We appreciate this point. Our experiments were performed under controlled laboratory conditions, which may not fully capture the ecological contexts shaping the costs and benefits of aggression. We will acknowledge this limitation and expand the Discussion to consider how environmental variability could modulate the fitness trade-offs associated with aggression in natural populations.

      We thank both reviewers for their constructive feedback, which will help us strengthen the rigor and clarity of the manuscript. We believe that the additional results and revisions will satisfactorily address their concerns.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The major weaknesses raised by the reviewers, namely the flies used in the study (CsxBully compared to CsxCs) where the effect on lifespan could be explained by hybrid vigor and the use of only one Bully and one Cs line that does not allow to link unambiguously the observed effect to the selection for aggression, should be addressed by a different experimental design and additional lines to exclude the effect of non-aggression related factors.

      We thank the Reviewing Editor for these comments.

      (i) Experimental design and hybrid vigor:

      Hybrid vigor typically arises from crosses between genetically divergent populations. In our study, Bully lines were derived from Canton-S background and are not therefore not genetically distant from controls. To directly address this concern, we included new data from Bully × Bully pairs (Figure 1), using independently selected Bully lines. These experiments reproduce the key aggression and courtship phenotypes observed in Cs × Bully assays, indicating that the effects are not attributable to hybrid vigor.

      (ii) Use of additional selected lines:

      We now include data from two independently selected lines (Bully A and Bully B), both derived from Cs, which show consistent behavioral phenotypes. This supports the conclusion that the observed effects are associated with selection for aggression rather than line-specific artifacts. We note that generating such lines is time- and labor-intensive, and only a few laboratories have established aggression-selected lines in Drosophila melanogaster (e.g., Penn et al., 2010; Dierick et al., 2006; Edwards et al., 2006). Accordingly, we have revised the manuscript to explicitly acknowledge this limitation and to frame our conclusions in terms of association rather than causation.

      Reviewer #1 (Recommendations for the authors):

      I can't see any way to interpret the data using CsxCs vs CsxBully comparisons.

      We thank the reviewer for this important point. This concern appears to arise from the assumption that Bully and Cs represent genetically distinct or non-coevolved populations. However, Bully lines were directly derived from Cs and therefore share a common genetic background. We have clarified this point in the Introduction (lines 104-106) and Results (lines 132-137).

      Importantly, we now include additional data showing that key phenotypes, including reduced mating duration, are also observed in Bully × Bully pairings and across independently selected Bully lines (new Figure 1). These results demonstrate that the observed effects are driven by the male genotype and do not depend on the female background.

      Because selection for aggression was applied specifically to males, we used Cs females as a standardized background to isolate male-specific effects while minimizing variability arising from female genotype or behavior. This rationale is now explicitly stated in the Results (lines 163-166) and at the beginning of the Discussion (lines 326-330). This experimental design allows interpretation of male-specific effects, and the observed differences cannot be attributed to cross design artifacts or hybrid vigor.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) Several passages currently imply causality, whereas the data support correlations between selection and trait differences rather than direct causation. This overstatement also appears in section subheadings within the Results, such as "Hyper-aggressive males display reduced mate-guarding efficiency, without compromising female fertility." Please consider toning down the wording by replacing active causal verbs with more neutral phrasing. Additionally, it would be important to include a clear, explicit sentence in the Discussion acknowledging this caveat, as the existing phrase "is associated with changes in reproductive traits" does not fully convey this nuance.

      We thank the reviewer for this important comment. We have revised the manuscript throughout, including Results subheadings, to replace causal language with association-based phrasing. We also rephrase the first sentence of the Discussion to clarify that our conclusions are correlational (see line 320).

      (2) Figure 1 would benefit from a simple schematic of the behavioral paradigm and the arena, since the authors' arena design minimizes manual handling; a cartoon would help readers quickly grasp the assay flow and the conditions under which interactions occur.

      We thank the reviewer for this helpful suggestion. We have added a schematic to Figure 1 illustrating the behavioral paradigm and arena design. Additional details are provided in the Materials and Methods (Trannoy et al., 2015). This improves clarity and accessibility of the experimental design.

      (3) In Figures 2A-B and A'-B', the higher post-mating UWE in Bully males is intriguing, but these panels do not actually measure the refractory period. It would be helpful to include 'latency' in the first UWE after mating in these swapped-female conditions. This could also be repeated with pheromone-standardized (cVA/CHC-equalized) decapitated females to disentangle effects of female pheromone load from male sensory perception. In addition, a baseline courtship control (naive males with decapitated virgins) is necessary to test whether Bully males simply have a lower threshold for initiating courtship.

      We thank the reviewer for this suggestion. The referenced panels are now shown in Figure 3. We quantified post-mating courtship latency; however, latencies were very short across conditions, and no differences were observed between genotypes. We therefore revised the text to interpret these results in terms of post-mating courtship motivation rather than refractory period. The baseline courtship control with decapitated virgins is provided in Fig 2G. These changes clarify the interpretation of post-mating behavior and address the reviewer’s concerns.

      (4) Related to my above point, the results in Figure 2B-B′ raise the possibility that Bully males have reduced perception or neural sensitivity to anti-aphrodisiac pheromones deposited by CS males, which could account for their elevated post-mating courtship; the authors might consider experiments that directly test male sensory responsiveness to these cues or mention this possibility in the Discussion.

      We thank the reviewer for this point. We performed additional assays to test males’ sensory responsiveness using binary choice assays and measured the time spent performing UWE towards decapitated females versus males. These results were added in Figure 3-Figure Supp 1, and indicate that both Cs and Bully males displayed courtship preferentially towards females, providing a control for sensory perception.

      (5) In multiple figure panels, virgin and mated females are depicted with the same symbols, which makes interpretation confusing.

      Thank you for pointing this out. We have updated the figure panels to use distinct symbols for virgin and mated females to improve clarity.

      (6) For Figure 4, it would be helpful to provide standalone KM curves for Bully versus CS males, including a separate panel for isolated (never-mated) males, and present mating counts in a separate panel while reporting survival models that incorporate mating frequency (or use it as a time-dependent covariate). Although Figures 4C-D report median survivals, full KM plots and an isolated-male curve are important since mating itself elevates mortality and can otherwise confound intrinsic lifespan differences.

      Thank you for this important point to improve clarity of this figure. We have reorganized this figure (now Figure 5) to now, present survival curves first (isolated and group-housed males), followed by lifetime mating counts. This reorganization separates survival from mating activity and addresses the potential confounding effect of mating on lifespan. For the lifetime mating data, we used bar plots rather than curves to better visualize individual mating events.

      (7) The following sentences overstate the results and imply causality; consider toning down: "These findings suggest that 5-P and 5-T might contribute to promoting remating in females that have previously mated with Bully males (Figure 2F' and G'). Given that Bully males also showed higher levels of both 5-P and 5-T compared to naïve Cs males (Figure 3B), it is likely that the elevated levels of these compounds observed in females result from their transfer during mating."

      We thank the reviewer for this important point. We have rephrased these sentences to remove causal language and instead describe associations between CHC profiles and behavioral outcomes. In particular, statements implying that 5-P and 5-T promote remating or are directly transferred during mating have been revised to reflect correlational evidence only (see lines 253-256).

      (8) It is not entirely clear how aggression was quantified in each generation, what proportion of males were selected to breed, and whether the findings generalize beyond a single Bully line (Figure Supplement 1 shows data from a closely related Bully line). Without independent replicate lines or sham-selected controls, it remains difficult to rule out drift or line-specific artifacts, and this limitation should be explicitly acknowledged.

      We thank the reviewer for this important point. We have clarified the aggression selection procedure by adding methodological details from Penn et al., including how aggression was quantified and how breeders were selected (lines 104-106 and 131-137). Briefly, independent selection replicates were initiated from the same Canton-S population, generating three lines (Bully A, B, and C), which were maintained separately.

      To address generality, we now include data from multiple lines. In particular, a new Figure 1 presents aggression and courtship phenotypes across Bully A, B, and C, and key behavioral results are consistent across independent lines.

      We acknowledge that additional independent lines would further strengthen generality; this limitation is now explicitly stated in the Discussion (lines 325-326).

      These additions clarify the selection procedure and support that the observed phenotypes are associated with aggression selection rather than line-specific artifacts.

      Minor Comments:

      (1) Exact sample sizes for every experiment should be included in the main figure legends.

      Thank you. We have added the number of replicates in each figure legends.

      (2) In Figure 1-Supplement 1, the orientation for depicting mating success is reversed compared to Figure 1, which is a bit jarring; it would be clearer to keep the orientation consistent with the main figure.

      Thank you. We have incorporated the results initially presented in Figure 1-Sup 1 into a new Figure 1 with additional results, and have taken into account reviewers’ comment.

      (3) For multivariate analyses, I suggest including important details such as group sample sizes, p-value, and the percent variance, etc., in the figure legend rather than keeping this only in Supplementary Table S1.

      We have inserted these details directly into the figure legends for clarity.

      (4) Why was the food cup used for arenas where decapitated virgins were used in mating assays?

      Thank you for pointing this. We now have clarified the experimental procedure in the M&M of the revised manuscript (lines 465-469).

      (5) For cartoons in Figure 2, the current yellow background makes it very difficult to distinguish flies drawn in yellow or green. Please adjust to a higher-contrast background or add darker outlines so that the cartoons are clearly legible.

      Thank you. We have increased the contrast of the female bodies to ensure the cartoons are clearly distinguishable (now figure 3).

      (6) Addition of line numbers in the manuscript would be helpful during the review process.

      Line numbers have been added throughout the manuscript.

      (7) I noticed a few typos in the manuscript. For example, in the Introduction, "seminal fuids" should be corrected to "seminal fluid." In the Discussion, the phrase "CHCs profiles compared those" requires a "to" before "those." Please carefully review the manuscript for similar errors.

      Thank you for pointing this out. We carefully reviewed the manuscript for typos and corrected all identified errors.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the editors and reviewers for your careful evaluation of our manuscript and for the constructive recommendations that have helped us improve the rigor, clarity, and balance of the study. We are pleased that the reviewers recognized the potential value of linking red light exposure to SIRT4 downregulation, fatty acid metabolism, H3K9 acetylation, and attenuation of ageing-related phenotypes. We have revised the manuscript extensively in response to the reviewers’ comments.

      In particular, we have clarified the wavelength specificity of the red-light response, reanalyzed and more cautiously interpreted the omics data, improved the presentation and quantification of semi-quantitative experiments, revised statistical reporting, corrected gene/pathway annotations, toned down mechanistic claims where direct evidence was insufficient, and expanded the Discussion to integrate recent evidence on red-light-induced fatty acid oxidation and AMPK/ACC signaling. We also added a dedicated limitations paragraph addressing the use of female mice, the absence of a complete in vivo wavelength-control and source-blocked sham cohort, and the need for future direct metabolic flux and isolated mitochondria studies.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      This is a challenging hypothesis that would require some additional experimental controls. The pathway dissection, while extensive, is sometimes approached in unconvincing ways, and the results are not always evident to judge or interpret. Technically, the western blots and transcriptomic analyses require notable improvements.

      We would like to thank the reviewer for the careful and patient examination of the issues identified in our manuscript. The poor quality of some of the Western blot bands in Figure 4 may have been caused by inappropriate electrophoresis conditions during the Western blot experiments. In the revised manuscript, we will optimize the electrophoresis conditions to obtain higher-quality protein bands and update the quantitative data. Regarding the quantification format, we believe that heatmaps provide a more intuitive representation of trends in protein expression across different treatment groups. This approach more accurately reflects the results of our biological replicates than simply analyzing the significance of differences in the grayscale values of protein bands. For the analysis of transcriptomic data, we will conduct a more detailed analysis of signal pathway enrichment and the identified differentially expressed genes to ensure that predicted genes are excluded from our current results and redundant data presentation is removed.

      Regarding additional experimental controls, such as incorporating experimental data under blue light treatment conditions as a control for red light. While exploring the optimal red light irradiation dose at the cellular level, we simultaneously conducted experiments on the effects of blue light irradiation at the same dose on keratinocyte activity. The results indicated that as the blue light irradiation dose increased (0–160 J/cm<sup>2</sup>), the keratinocyte activity exhibited a dose-dependent decline. This indicates that blue light is phototoxic to keratinocytes. The relevant experimental results have already been published in our previous study (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). Taken together with the data from our study, this demonstrates that the anti-ageing effects of red light reported in the current manuscript are indeed driven by red light.

      Reviewer #2 (Public review):

      Weaknesses:

      The paper does not evolve to use the mechanistic discoveries of the manuscript to help our community to identify the mechanism of photobiomodulation, which is not known so far.

      I would like to draw attention to a recently published paper by Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), which shows that red light (660 nm) stimulates mitochondrial fatty acid oxidation in keratinocytes via AMPK‑dependent phosphorylation of ACC, without altering expression of electron transport chain complexes. I believe this paper is highly complementary to the current study.

      Herrera et al. demonstrate that red light increases basal, ATP-linked, and maximal oxygen consumption rates in keratinocytes specifically through enhanced fatty acid oxidation (inhibited by etomoxir). This independently validates the central finding of the current manuscript, i.e., red light boosts lipid metabolism, strengthening the robustness of this concept.

      While the current manuscript focuses on the SIRT4-MCD axis, Herrera et al. identify AMPK phosphorylation and ACC inhibition as key effectors. The authors can integrate and expand their discussion, since SIRT4 downregulation may converge on AMPK activation, or they may represent parallel, reinforcing mechanisms. This would enrich the mechanistic model and open new hypotheses.

      The mechanism of photobiomodulation: Herrera et al. explicitly challenge the prevailing paradigm that red light acts solely via cytochrome c oxidase (by showing long-lasting effects, unchanged OXPHOS protein levels, and no difference in permeabilised cells). The current finding (red light acts through SIRT4 downregulation, i.e., not direct enzymatic activation) aligns perfectly with Herrera´s critique.

      Long-term metabolic effects-Herrera et al. show that a single red light exposure elevates oxygen consumption for up to 2 days. The current study focuses on changes at 12-24 h. Their data extend the time window and suggest that the metabolic reprogramming you describe may persist longer than currently discussed, which is clinically relevant.

      Discussing Herrera et al.'s results would not only acknowledge independent, corroborating evidence but would also allow the authors to position their SIRT4-centric mechanism within a broader, emerging understanding of red-light photobiomodulation.

      We would like to thank the reviewer for providing us with constructive suggestions for discussion. Our results showed that under red light conditions, both glycolipid and lipid metabolism were activated in keratinocytes, and cellular metabolic flux increased. The activation of lipid metabolism directly led to an increase in metabolism-associated H3K9ac and drove the upregulation of anti-ageing-related genes; we believe this is key to the anti-ageing effects of red light. Mechanistic analysis combining proteomics and acetylation proteomics revealed that red light significantly downregulated SIRT4 expression and increased the acetylation of MCD, a protein regulated by SIRT4 that governs cellular fatty acid oxidation rates. Through validation using cell-level knockdown and inhibitors, we confirmed that SIRT4 inhibition exerts anti-ageing effects in vitro and that inhibiting MCD function under red light conditions suppresses H3K9ac. These results establish the role of the SIRT4-MCD signalling axis in mediating the anti-ageing effects of red light.

      The study by Herrera et al. included a substantial body of validation data confirming the role of red light in promoting fatty acid oxidation, providing robust empirical support for our research. Furthermore, Herrera et al. revealed that red light-induced fatty acid oxidation depends on AMPK and ACC phosphorylation. This mechanism of red-light photobiomodulation may refute the notion that its bio-regulatory effects rely solely on the action of mitochondrial cytochrome c oxidase. Furthermore, together with our study revealing that red light exerts anti-ageing photobiomodulatory effects via the SIRT4-MCD signalling axis, these findings independently confirm that red light regulates cellular fatty acid oxidation, thereby demonstrating the pivotal role of activated fatty acid oxidation in the bio-regulatory effects of red light. In the revised manuscript, we will include a discussion on the potential link between the red light-driven downregulation of SIRT4 and the phosphorylation of AMPK/ACC. This will be of positive value in elucidating how SIRT4 exerts its anti-ageing effects by regulating lipid metabolism, as well as in explaining the possible mechanisms by which red light downregulates SIRT4.

      Recommendations for the authors:

      Summary of Major Revisions

      Changes made in the revised manuscript:

      (1) Added a clearer explanation of why the 625-635 nm red-light regimen was considered the active intervention and how the available blue-light data from our previous work support wavelength-dependent effects on keratinocytes.

      (2) Revised the language describing inflammatory regulation. We now avoid presenting red light as producing a uniform anti-inflammatory effect and instead describe selective remodeling of ageing-associated inflammatory and SASP signatures.

      (3) Improved figure presentation and quantification for immunofluorescence, metabolite, and western blot assays; clarified image-analysis regions, replicate numbers, and normalization procedures.

      (4) Reanalyzed transcriptomic, proteomic, and acetyl-proteomic datasets with appropriate multiple-testing correction and corrected erroneous pathway/gene annotations in metabolic gene panels.

      (5) Replaced overly strong causal wording with more conservative language, especially regarding PI3K/Akt/mTOR, cytochrome c oxidase, SIRT4 localization, PPARα immunofluorescence, and direct fatty acid oxidation flux.

      (6) Expanded the Discussion to incorporate Herrera et al. (FEBS Letters 2025, doi:10.1002/1873-3468.70195), highlighting convergence between the SIRT4-MCD model and AMPK/ACC-dependent fatty acid oxidation.

      (7) Corrected typographical, nomenclature, and figure-legend inconsistencies throughout the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) Wavelength specificity and need for a non-red-light control

      As a reader, one is left wondering whether the effects are due to red light specifically. An important control would have been to irradiate mice and cells with another light wavelength, such as blue light.

      We agree that wavelength specificity is a critical issue for interpreting photobiomodulation studies. In the revised manuscript, we have clarified that the anti-ageing and metabolic effects described here apply specifically to our 625-635 nm red-light regimen, rather than to visible light in general. We have also added a discussion of our previously published blue-light experiments, in which keratinocyte viability decreased in a dose-dependent manner across the same 0-160 J/cm<sup>2</sup> dose range (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). These data indicate that blue light and red light produce distinct biological outcomes in keratinocytes. Because high-dose blue light was cytotoxic under comparable cellular conditions and because the present study was designed to investigate the long-term effects of red light in aged mice, we did not perform prolonged in vivo blue-light irradiation as an ageing intervention.

      Changes made in the revised manuscript:

      Clarified in the revised Introduction and Discussion that the conclusions are specific to 625-635 nm red light under the irradiation parameters used in this study. (Lines 92 to 94, Lines 1146-1149)

      Added text summarizing the published blue-light comparison data from our previous study (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1), including the dose-dependent decline in keratinocyte activity after blue-light irradiation. (Lines 90 to 92, Lines 1146-1149)

      Added data on the wavelength range of the red light used in this study. (Lines 529 to 531, Fig S1a)

      (2) Complexity of inflammatory effects

      The manuscript repeatedly emphasizes anti-inflammatory effects, yet some cytokines such as IL-18, Ccl2, TNF-α, Ccl2, and IL-8 appear increased. This suggests that the effects may be more complex than presented and may require additional readouts or stronger statistical power.

      We unanimously agree that the inflammatory response to red light should not be described as a simple, uniform suppression of all cytokines. We have demonstrated that changes in the levels of the senescence-associated secretory phenotype (SASP) at the cellular level and in skin tissue following red light treatment not only indicate that red light-induced metabolic activation can reduce the age-related inflammatory baseline, but also reveal red light-driven short-term reparative effects or stress-related cytokine responses. We consider this to be consistent with the findings, and the downregulation of NF-κB-related signalling observed in skin tissue following periodic red light irradiation of aged mice further supports the conclusion that red light alleviates the age-related inflammatory baseline. In fact, in our previous study, we did observe that red light treatment promoted increased levels of the cytokine Ccl2, which plays an important positive role in rapid wound healing (Communications Biology 2024, doi: 10.1038/s42003-024-06973-1). In the revised manuscript, we have reworded the relevant Results and Discussion sections to indicate that red light remodels ageing-associated inflammatory signalling rather than globally reducing every inflammatory mediator. We have also toned down statements suggesting that red light ‘reverses’ or ‘suppresses’ inflammation where the underlying data support a more selective effect.

      Changes made in the revised manuscript:

      Replaced broad terms such as “anti-inflammatory effects” with more precise wording such as “remodeling of ageing-associated inflammatory signaling” where appropriate. (Lines 541 to 543, Lines 573 to 574, Lines 893 to 895)

      Expanded the Discussion to explain that red-light-induced metabolic activation may simultaneously reduce senescence-associated inflammatory tone while allowing transient reparative or stress-related cytokine responses. (Lines 1135 to 1142)

      (3) Figure clarity, semi-quantitative methods, western blot quality, and inconsistent band patterns

      Many differences are difficult to see or require orthogonal validation. Some tissue-specific signals and western blots are difficult to judge. Several western blots are of poor quality, and multiple markers show inconsistent band profiles across experiments, including SIRT4 in Figure 5.

      We thank the reviewer for highlighting these technical and presentation issues. We have reviewed the semi-quantitative data and revised the presentation of the figures to improve their interpretability. For Western blot experiments, we optimised the electrophoresis and transfer conditions, replaced low-quality representative images where possible, and updated the semi-quantitative results. Furthermore, regarding the lack of clarity in the SIRT4 protein band, we have conducted repeat experiments and updated the main text to include a clearer image of the band. The issue with the annotation of the protein location was in fact due to an oversight during the data analysis process; we have carried out a detailed review and provided the uncropped full-length Western blot images for all experiments in the Supplementary Materials for the reviewers’ scrutiny. Finally, we would also like to point out that factors such as sample origin, protein extraction, electrophoresis conditions, antibody exposure time, and potential non-specific detection may all contribute to differences in band patterns. At present, the core conclusions regarding SIRT4 are supported by multiple lines of evidence, including mRNA analysis, immunofluorescence, Western blotting of bands at the expected sizes, and SIRT4 knockdown experiments, rather than being based solely on any single semi-quantitative Western blot result. We therefore believe that the conclusions drawn from the data presented in the revised manuscript are equally convincing.

      Changes made in the revised manuscript:

      Replaced or improved low-quality Western blot panels and updated quantitative analyses in revised Figures 3-5 and associated supplementary material. (Fig 3m, Fig 4, Fig 5f)

      Clarified the normalization approach for H3K9ac/H3 and target/loading-control comparisons, and the use of Actin or H3 as appropriate loading controls. (Lines 283 to 293)

      Bands with nonspecific profiles were excluded from quantitative conclusions and the manuscript conclusions no longer depend on those ambiguous signals.

      Revised the Results (Repeat the experiment to update the low-quality Bands) to avoid overstating changes that are not clearly visible or not supported by statistical analysis. (Fig 4n and r)

      (4) Choice of pharmacological agents and need for genetic strategies

      The choice of drugs in Figure 4 is puzzling. More specific and widely used inhibitors could be used to block PI3K/Akt or mTOR, and natural agonists such as insulin or EGF could be used. Genetic strategies should complement these observations.

      We agree that pharmacological perturbation experiments should be interpreted with caution. In this study, our criteria for selecting inhibitors were based on transcriptomic and proteomic analyses; we sought to determine how the most direct inhibition of red light-activated signalling pathways would affect H3K9ac levels. In the revised manuscript, we have clarified the rationale for the compounds used and have reduced the causal weight assigned to these inhibitor/agonist experiments. These data are now presented as supportive evidence that red light is associated with metabolism-related signalling changes, rather than as definitive proof that PI3K/Akt/mTOR is the primary upstream mechanism. We have also emphasised the genetic SIRT4 knockdown experiments as a more direct mechanistic test for the SIRT4-centred part of the model. We acknowledge that additional experiments using more selective inhibitors, physiological agonists such as insulin or EGF, and genetic perturbation of PI3K/Akt/mTOR components would be valuable for future studies.

      Changes made in the revised manuscript:

      Revised the text describing pharmacological experiments to distinguish supportive pathway modulation from direct causal evidence. (Lines 787 to 789)

      Added a limitation and future direction noting that genetic perturbation of PI3K/Akt/mTOR and physiological pathway activation with insulin or EGF would strengthen the model. (Lines 1190 to 1194)

      (5) Serum NADH measurement

      In Figure 1t, the authors measure serum NADH. NADH is poorly detectable in serum or plasma, and changes may reflect blood-cell lysis during collection rather than circulating NADH.

      We appreciate this technical concern. We have revised the manuscript so that serum NADH is no longer used as a central mechanistic readout. We now treat this measurement only as an exploratory indicator of systemic redox-related changes and explicitly acknowledge that serum or plasma NADH is vulnerable to artifacts from blood-cell disruption during sampling. The mechanistic interpretation has been shifted toward cellular and tissue measurements, including intracellular NADH/NADPH/GSH, ATP, acetyl-CoA, fatty acid uptake, and H3K9ac, which are more directly relevant to keratinocyte metabolic remodeling.

      Changes made in the revised manuscript:

      Removed serum NADH from the main causal argument linking red light to metabolic flux and H3K9ac.

      Placed greater emphasis on cell-based metabolite assays, tissue acetyl-CoA, and H3K9ac measurements as the main metabolic-epigenetic evidence. (Lines 582 to 587, Fig 1s)

      (6) Direct assessment of glycolysis and fatty acid oxidation

      The authors propose that red light increases glycolysis and fatty acid oxidation, but this could be assessed directly rather than through surrogate measures.

      We agree. Our current data include multiple metabolic readouts, including glucose and fatty acid uptake, ATP, NADH/NADPH/GSH, triglycerides, fatty acids, pyruvate, lactate, acetyl-CoA, and MCD-dependent changes; however, these assays are not equivalent to direct flux measurements such as Seahorse extracellular flux analysis, isotope tracing, or etomoxir-sensitive respiration. We have therefore revised the wording throughout the manuscript to distinguish metabolic remodeling and fatty-acid-oxidation-related signatures from direct measurements of fatty acid oxidation flux. We also incorporated the recent independent work by Herrera et al., which directly measured oxygen consumption and demonstrated red-light-induced fatty acid oxidation in keratinocytes. This external evidence supports the biological plausibility of our SIRT4-MCD model while making clear which aspects are directly measured in our study and which are inferred.

      Changes made in the revised manuscript:

      Replaced overstrong language such as “red light increases fatty acid oxidation” with “red light promotes PPAR-α-related fatty acid metabolism pathway” where direct flux data were not measured in our experiments. (Lines 882 to 883)

      Expanded the Discussion to integrate direct FAO evidence from Herrera et al. and to place the SIRT4-MCD axis within a broader red-light metabolic framework. (Lines 1169 to 1189)

      (7) Incorrect annotation of metabolic genes in Figure 3e

      Acss2, Aldh3b1 and Aldh3a1 are not glycolytic enzymes, Aldh3a3 does not appear to exist, and several enzymes classified as FAO are fatty acid synthesis enzymes. This questions the interpretation of the data.

      We thank the reviewer for identifying these annotation errors. We have rechecked the gene names and pathway assignments in the transcriptomic analysis and corrected the metabolic gene panels. We have confirmed that Acss2 is an acetyl-CoA synthase involved in the metabolism of acetate to acetyl-CoA. The Aldh family genes, meanwhile, are associated with aldehyde metabolism and detoxification. We have made the corresponding adjustments in the manuscript. The incorrectly listed Aldh3a3 entry has been removed. Furthermore, we have categorised genes involved in fatty acid metabolism as ‘fatty acid metabolism-related genes’, rather than grouping them all under the FAO category. These revisions have significantly improved the accuracy of the metabolic interpretation.

      Changes made in the revised manuscript:

      Reannotated Figure 3e and the corresponding Results text to correct glycolysis, TCA cycle, pentose phosphate pathway, and fatty acid metabolism categories. (Fig 3d)

      Removed the erroneous Acss2 and Aldh family genes. (Fig 3d)

      Revised the metabolic model to avoid using incorrectly grouped genes as evidence for direct fatty acid oxidation. (Lines 707 to 709)

      (8) Transcriptomic analysis and implausible volcano-plot p-values

      The transcriptomic analysis raises concerns. For example, the volcano plot in Figure 3d appears incorrect, with -log<sub>10</sub>(P-value) around 300 despite n=3 biological replicates.

      We thank the reviewer for pointing out this important issue. We have reopened the transcriptomic data and found that the extremely high -log<sub>10</sub>(P-value) in the original volcano plot were caused by the automatic replacement of very small P-values—generated during the differential expression analysis—with zero in the tabular data. To avoid misleading visualisations, we have regenerated the volcano plot using Q-values in place of the original P-values. Differentially expressed genes were defined as those with a Q-value < 0.05 and |log<sub>2</sub> fold change| > 1. Furthermore, for visualisation purposes only, the upper limit for q-values was set to 1 × 10<sup>-50</sup> for values below 1 × 10<sup>-50</sup>. This adjustment does not affect the statistical classification of differentially expressed genes but prevents over-interpretation of extremely small values. The revised volcano plots and legends have been updated accordingly. To avoid any potential misinterpretation arising from these updates, the updated volcano plots are presented in the supplementary materials.

      Changes made in the revised manuscript:

      Reanalyzed transcriptomic data using appropriate multiple-testing correction and revised the volcano plot. (Fig S3a)

      Corrected the y-axis transformation and removed implausible -log<sub>10</sub>(P-value) presentation. (Fig S3a)

      Updated Methods to specify the statistical workflow for transcriptomic differential expression and pathway enrichment. (Supplementary materials Lines 46 to 52)

      Moved the analysis of metabolic pathways based on transcriptomic data to the supplementary material, thereby reducing the reliance of the conclusions on transcriptomic data (Fig S3b).

      (9) Need to tone down mechanistic claims regarding PI3K/Akt/mTOR, cytochrome c oxidase, SIRT4, and PPARα

      The mechanisms proposed must be toned down. PI3K/Akt/mTOR should not be called glycolytic pathways, the link to red light or cytochrome c oxidase is vague, SIRT4 reduction requires mitochondrial counterstaining, and PPARα appears cytosolic after SIRT4 knockdown.

      We agree and have substantially revised the mechanistic language. PI3K/Akt/mTOR is no longer referred to as a ‘glycolytic pathway’; instead, it is described as a metabolism-related signalling axis that may influence glucose uptake, growth and nutrient-responsive metabolism. We have also toned down statements attributing red-light effects directly to cytochrome c oxidase, as our study primarily examines downstream metabolic and epigenetic remodelling rather than direct photoreceptor activation. With regard to SIRT4, we have revised the text to avoid interpreting changes in SIRT4 immunofluorescence alone as evidence of altered mitochondrial abundance or mitochondrial localisation. The conclusion is now based on a combination of SIRT4 mRNA levels, western blot bands of the expected size, immunofluorescence trends, and SIRT4 knockdown phenotypes. With regard to PPARα, we have re-examined the PPARα antibody used for the cellular immunofluorescence experiments. In the original Figure 5p, we mistakenly used a PPARα antibody (PPARα, Abclonal, A25296) that is only suitable for Western blot (WB) experiments; we believe this was the cause of the mislocalisation of the fluorescent signal; Consequently, we conducted new experiments using a PPARα antibody (PPARα, Abclonal, A22887) specifically designed for cellular immunofluorescence. The relevant experimental data have been corrected in the manuscript.

      Changes made in the revised manuscript:

      Replaced “PI3K/Akt/mTOR glycolytic pathway” with “The PI3K-AKT signalling pathway is involved in the regulation of glucose metabolism” throughout the revised manuscript. (Lines 701 to 705, Lines 809 to 810, Lines 813, Lines 1193)

      Reduced mechanistic certainty around cytochrome c oxidase and framed it as a possible upstream photoreceptor rather than an experimentally proven mechanism in this study. (Lines 827 to 830, Lines 842 to 848)

      Repeat the PPARα immunofluorescence staining experiment. (Fig 5p)

      (10) Need for isolated mitochondria experiments and red/blue light comparison of mitochondrial respiration

      If the effect of red light relies on mitochondrial cytochromes, additional proof would be needed, potentially using isolated mitochondria and comparing how red and blue light influence respiration capacity.

      We agree that isolated mitochondria experiments would be an important way to test direct mitochondrial photoreception. Because the present study was designed around cellular and in vivo metabolic-epigenetic remodeling, we did not perform isolated mitochondria irradiation experiments. To address this concern, we have toned down statements implying direct cytochrome activation and revised the Discussion to distinguish between direct mitochondrial photoreceptor models and downstream metabolic reprogramming. We also added a future direction proposing isolated mitochondria or permeabilized-cell experiments comparing red and blue light effects on respiration, ATP-linked OCR, maximal respiration, and FAO-dependent respiration. The revised manuscript now emphasizes that our data support a downstream SIRT4-MCD-H3K9ac mechanism after red-light exposure, while the proximal photophysical event remains to be fully defined.

      Changes made in the revised manuscript:

      Added discussion of the need for isolated mitochondria, permeabilized-cell, and wavelength-comparison respiration experiments. (Lines 1194 to 1197)

      Reviewer #2 (Recommendations for the authors):

      (1) Statistical reporting, post-hoc tests, normality/equal-variance tests, exact p-values, and FDR control

      The manuscript states that one-way ANOVA followed by Tukey or Dunnett tests was used, but it does not consistently specify the post-hoc correction for each figure. Normality and equal-variance tests are not reported, p-values are shown only as asterisks, and FDR control is not mentioned for transcriptomics and proteomics.

      We agree that the statistical reporting needed to be more complete. We have revised the Statistics and reproducibility section and the figure legends to specify the statistical test used for each experiment, the post-hoc correction applied after ANOVA, the number of independent biological replicates, and the definition of error bars. Where multiple comparisons were performed, we now state whether Tukey’s or Dunnett’s correction was used. Regarding P-value presentation, we have retained the use of asterisks in the figures as visual indicators of statistical significance to maintain figure readability. For transcriptomic, proteomic, and acetyl-proteomic analyses, we have revised the Methods section to state that multiple-testing correction was performed using the Benjamini–Hochberg false-discovery-rate procedure. Adjusted P values or Q values were used for differential-expression and pathway-enrichment analyses. These revisions clarify the statistical workflow and strengthen the reproducibility of the study.

      Changes made in the revised manuscript:

      Revised the Statistics and reproducibility section to define statistical tests, post-hoc corrections, assumption checks, and multiple-testing correction. (Lines 509 to 522)

      Updated relevant figure legends to include n values, statistical tests, post-hoc corrections, and definitions of significance symbols. (Lines 515 to 516)

      Added FDR control details for RNA-seq, proteomics, acetyl-proteomics, and pathway-enrichment analyses. (Lines 450 to 455)

      (2) Figure clarity and quantitative analysis of fluorescence, JC-1, metabolite, and western blot data

      Several figures lack clarity or appropriate quantification. Figure 1i-j H3K9ac quantification should be based on whole-image or multiple fields; Figure 2e JC-1 should include red/green ratio quantification; Figure 2k-p metabolite data should include absolute concentrations; Figure 3j needs appropriate loading controls.

      We appreciate these specific suggestions and have revised the figure presentation accordingly. For H3K9ac immunofluorescence in skin sections, we have clarified the anatomical region quantified and performed a more objective quantification using multiple fields/regions per section rather than relying on a visually selected dashed area. The dashed regions in the representative images were made clearer and the quantification criteria were added to the Methods and legend. For JC-1 staining, the bar chart on the right-hand side of the mitochondrial membrane potential fluorescence image in Figure 2e shows the quantitative data for the red/green fluorescence ratio obtained from independent experiments; compared with providing only a representative image, these data offer a more easily interpretable quantitative measure of mitochondrial membrane potential. To avoid any potential misunderstanding, we have corrected the vertical axis. For metabolite assays, we clarified normalization to cell number or protein content and revised the data presentation to include absolute or normalized concentrations where available, rather than relying solely on fold changes with variable y-axis scaling.

      Changes made in the revised manuscript:

      Revised Figure 1i-j quantification using multiple fields/regions per mouse section and improved dashed-region visibility. (Lines 431 to 440, Fig 1i and j)

      Corrected the vertical axis of the quantitative data for the JC-1 red/green fluorescence ratio. (Fig 2e and Fig S2a)

      Updated metabolite panels and/or source data to include absolute or protein-normalized values where available, and standardized y-axis interpretation. (Fig 2k-p and Fig 4s and v, Given the diversity of intracellular fatty acid and triglyceride species, absolute quantification based solely on absorbance measurements would be technically challenging and may not accurately reflect the content of each molecular component. Therefore, we presented the changes in fatty acid and glycerol levels as percentage-normalized relative absorbance values, which allowed consistent comparison among the experimental groups.)

      (3) Experimental design limitations: sex of mice and sham control

      Only female C57BL/6 mice were used, although aging and metabolic responses can be sex-dependent. The thermal-control argument lacks a true sham control in which mice are placed in the same apparatus with the light blocked at the source.

      We agree with these points. We have added a section on limitations stating that all aged mice used in this study were female, and that sex-dependent responses to red light, SIRT4 regulation, metabolism and skin ageing should be investigated in future studies using both male and female cohorts. Furthermore, regarding the design of the non-irradiated control group: although the control mice underwent the same depilation and routine procedures, they did not receive red light irradiation. However, we also acknowledge that establishing a sham-irradiated control group with light shielding would allow for stricter control of factors such as restraint, contact with equipment and procedural stress. However, given that this experiment involved a continuous cyclic photoperiodic treatment lasting two years, we were unable to supplement the study with a control experiment involving only red light shielding. Nevertheless, based on the fact that we observed only minimal changes in the mice’s skin temperature following red light irradiation, we believe that the primary factor driving the alleviation of the skin ageing phenotype in the mice remains red light-induced.

      Changes made in the revised manuscript:

      Added a limitation noting that the study used female C57BL/6 mice only and that sex as a biological variable should be addressed in future studies. (Lines 1198 to 1201)

      (4) Textual errors, nomenclature inconsistencies, and ChIP-qPCR normalization

      Several textual errors and inconsistencies should be corrected, including Pparg1a/Ppargc1a, Ricotr/Rictor, Pi3k/PI3K, Sirt4/SIRT4 protein nomenclature, and the use of RPL30 normalization in ChIP-qPCR without showing that RPL30 is unchanged.

      We thank the reviewer for their careful reading. We have corrected the typographical errors and standardised gene and protein nomenclature throughout the manuscript and figure legends. Specifically, Ppargc1α has been corrected to Ppargc1a, Ricotr to Rictor, and the capitalisation of PI3K has been standardised. We now use Sirt4 for the mouse gene and SIRT4 for the protein, applying the same convention to other genes and proteins. For ChIP-qPCR, we have revised the Methods and Results sections to describe normalisation against input and IgG controls more clearly, and to specify the role of the RPL30 locus as an internal control. We have also included data in the Supplementary Materials showing relative enrichment of H3K9ac in the RPL30 promoter region in PAM212 cells before and after red light irradiation; the results indicate that H3K9ac enrichment at the RPL30 locus remained stable across treatment groups after normalization to input DNA and correction against IgG background. This result indicates that the use of RPL30 as an internal control in ChIP-qPCR experiments is feasible.

      Changes made in the revised manuscript:

      Corrected Ppargc1a, Rictor, PI3K, Sirt4/SIRT4, and related nomenclature throughout the manuscript.

      Supplement the experimental results on the effect of red-light irradiation on the level of H3K9ac enrichment at the RPL30 locus in keratinocytes. (Lines 339-352, Lines 539 to 541, Fig S1d)

      (5) Additional Revision Addressing the Public Review and Herrera et al.

      The reviewer suggested integrating the recent study by Herrera et al. showing that 660 nm red light stimulates mitochondrial fatty acid oxidation in keratinocytes through AMPK-dependent phosphorylation of ACC, without changing electron transport chain complex expression. The reviewer also noted that these findings may complement the SIRT4-MCD axis and challenge a cytochrome-c-oxidase-only model of photobiomodulation.

      We are grateful for this constructive suggestion. We have expanded the Discussion to incorporate Herrera et al. and to place our SIRT4-MCD-centered mechanism within the broader emerging model of red-light-driven metabolic remodeling. Herrera et al. provide direct oxygen-consumption evidence that red light enhances fatty acid oxidation in keratinocytes and that this effect involves AMPK/ACC signaling. This is highly complementary to our data, in which red light decreases SIRT4, increases acetylation of MCD, promotes fatty-acid-metabolism-related signatures, elevates acetyl-CoA, and increases H3K9ac. In the revised Discussion, we propose two nonexclusive models: red light-induced SIRT4 downregulation may converge with AMPK/ACC-dependent relief of fatty acid oxidation, or the two pathways may represent parallel reinforcing mechanisms that together enhance lipid metabolic flux.

      Changes made in the revised manuscript:

      Added a paragraph discussing Herrera et al. in the revised Discussion. (Lines A1169 to 1189, Lines 1206 and 1208)

      Revised the conceptual model of red-light photobiomodulation to emphasize downstream metabolic reprogramming rather than direct cytochrome c oxidase activation alone. (Lines 827 to 830, Lines 843 to 844)

      Added future directions to test whether red-light-induced SIRT4 downregulation causally affects AMPK/ACC phosphorylation and FAO-dependent respiration. (Lines 1169 to 1189)

      We again thank the editors and reviewers for their thoughtful and constructive comments. The revised manuscript now provides a more rigorous and balanced presentation of the evidence, distinguishes direct measurements from inferred metabolic flux, corrects pathway annotations, improves figure quantification and statistical transparency, and places the SIRT4-MCD-H3K9ac mechanism within a broader framework of red-light-induced fatty acid metabolic remodeling. We believe these revisions substantially strengthen the manuscript and clarify both the significance and the limitations of our findings.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      All experiments in the manuscript use optogenetic activation of DANs, thus it is not clear what kind of memories are formed. Several stimuli can be used as punishment, such as electric shock, salt, bitter, and light - it is not clear what kind of memory the authors investigate here. The findings could be discussed in the context of what DANs respond to.

      This is indeed a caveat of our study and we discuss this issue now in lines 557-566. We refrained from testing necessity to specific US on purpose as we knew that another research group was focussing on this question in parallel (Weber et al., 2023, also published in eLife) and therefore rather focussed on complementary experiments. That study also includes a rather deep discussion about the inputs to individual DANs. We briefly refer to this discussion (lines 442-444) but decided to not go into detail to avoid too much overlap.

      Furthermore, studies in adults and larvae showed that most DANs can code for both valences - etc., aversive DANs can be activated by punishment, and inhibited by reward. Thus, safety learning might be a result of a decrease in activity in DANs during odor presentation. The authors also do not discuss possible feedback loops from MBONs to DANs across compartments. Could such connections allow for safety learning in larvae?

      We thank the reviewer for raising these points and included a brief discussion of both scenarios in lines 469-474.

      The authors show that artificial activation with different light intensities can form different memories and that increasing the light intensity sometimes leads to no memories. Also, using different optogenetic tools reveals different results. This again raises the question of how applicable the results will be for learning with real stimuli. Is there a natural stimulus that only induces safety learning, but no punishment learning?

      We do not know of such a stimulus. Based on our data, a US that only activates a single DAN should only make safety memory – however, the available data of which US activates which DAN is very limited in larvae and currently no such US is known. We discuss this point briefly in lines 557-563 and 572-574.

      The authors provide a detailed behavioral analysis of locomotion behavior; however, the detailed analysis seems unnecessary for that dataset. Modulation of speed and bending rate has been described before with simpler methods (specifically for MBONs). The revealed locomotion phenotypes probably affect larval locomotion during memory recall with light activation, thus the authors should show that larvae are potentially able to move during light-on memory tests.

      We expanded our locomotion analysis of the innate and learned odor preference experiments (new Fig. 6) and show that in these experiments, even with TH-DANs being activated, larvae indeed can move relatively normal and the existing locomotion phenotypes are not correlated to their olfactory choices.

      We do not agree that the locomotion analysis is unnecessary. Modulations of speed and bending have been described for MBONs but to our knowledge not for DANs. It is not trivial at all that DANs and MBONs cause the same behavioral modulations (see, for example, this adult study: Mohammad et al., 2024 Plos Biol). There is extremely limited knowledge about the motoric effects of dopaminergic neurons in larvae - we therefore find it important to describe our results in detail. We added some further rationale of why we think it is crucial to explore the functions of DANs for learning and movements together (lines 81-86).

      Reviewer #2 (Public review):

      Weaknesses:

      (1) The authors have done a great job at structuring the figures. But some main figures would benefit from including the controls instead of placing them in supplementary.

      We had decided to put the controls into the supplement in some cases to prevent the main figures to be overcrowded. We revised this decision upon the reviewer’s comment for Fig. 8 (previously Fig. 7) but decided to keep other figures unchanged as we feel that the current design best fits the purpose of each figure. We provide a figure-for-figure rationale in our response to the recommendations for the authors.

      (2) The paper would benefit from a deeper discussion regarding molecular mechanisms underlying their results. It would be interesting to see what the authors think about different Dopamine receptors and how they relate to the findings of this paper.

      We thank the reviewer for the suggestion. Although we agree that such a discussion would be interesting, we hesitate to expand on this topic, as the discussion is already quite long and our study does not contribute any new data to clarify the molecular dopaminergic mechanism.

      (3) Throughout the paper, the authors have been clear and comprehensive, but in some cases, further explanation of their choices were missing. For example, the choice to compare bending and tail velocity over other parameters within the same clusters is unclear.

      We understand that this choice was not clearly explained and expanded on our rationale in lines 244-251.

      Reviewer #3 (Public review):

      Weaknesses:

      The larvae exhibit directed locomotory action to express punishment or safety memory. If the larvae did not move, we would not be able to assess memory function. Hence, functional activation of DANs could result in one action, which seems like two different functions of memory expression and locomotion. It can also be argued that activation of DANs represents a teaching signal to the KCs, and then eventually, downstream of the MBONs, it results in locomotion modulation. Hence, the seeming functional diversity could be a function of different downstream neuronal pathways and not molecular context-dependent diversity inside dopaminergic neurons. The authors should address this possibility or point out the fallacy in the above argument.

      We thank the reviewer for raising this issue. To the first point, we expanded our locomotion analysis of the innate and learned odour preference experiments (new Fig. 6) and show that in these experiments, even with TH-DANs being activated, larvae indeed can move relatively normal. In addition, the existing locomotion phenotypes in these experiments were not correlated with the animals’ olfactory choice. This makes it unlikely that the changed locomotion directly determines our observation during the olfactory experiments.

      We do agree that it is possible that both the preference after learning and the changed locomotion could come through the same dopaminergic mechanism via diverse downstream pathways. We cover this hypothesis in Fig. 10G and address this question briefly in lines 580-582.

      The finding that activation of TH-GAL4 conveys aversive valence and R58E02-GAL4 conveys appetitive valence seems redundant (Figure 6). I understand they say this in the context of locomotion. However, they may not have mentioned similar findings in adults. In adults, artificial activation of DANs covered by the same GAL4 lines acts as aversive and appetitive teaching signals for memory formation. These references should be cited appropriately in the results and discussion if not currently included.

      We thank the reviewer for this comment and tried to include the relevant adult literature (see e.g. lines 351-356 and 583-604). In particular, we added a quite detailed discussion about a paper published after our initial submission that performed similar experiments for the adult PAM-DANs Lozada-Perdomo et al., 2025, iScience).

      We do not agree, however, that the experiments in Fig. 7 (previously Fig. 6) are redundant. Recent studies in adults found no correlation between the rewarding/punishing effects and the innate valence a given dopaminergic neuron induces (Rohrsen et al., 2021, bioRxiv; Mohammad et al., 2024, PLOS Biol; Lozada-Perdomo et al., 2025, iScience). To our knowledge, no such studies have been carried out in larvae so far. Therefore, we think that it is not only important to test it but that the respective results compared to the results in adults are of relevance for the readership.

      The evidence for the role of dopamine (Figure 7) can be bolstered by using other available RNAi lines against TH. A valium20 vector-based shRNA line is recommended. The current evidence is based mainly on non-specific pharmacological intervention with 3IY.

      We agree to this caveat and made it transparent now in lines 387-389 and 401-403. We nevertheless chose, for the time being, to not include further experiments to address this point in the current study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Activation of specific or multiple DANs seems to increase naïve odor preference (f1 or TH). Is this due to locomotion defects in TH - how does the odor preference develop over time? Can this increased odor preference explain the safety learning - where they also approach the odor stimulus?

      The reviewer is right that in presence of light, we see increased odor preference both innate and after unpaired training – theoretically, that could be the same effect. However, this would not explain why we see the same increased preferences after unpaired learning with all driver strains but increased innate preferences only for some of them. Moreover, when activating TH, we also see increased odor preference in absence of light after unpaired training but not innately. We therefore think that these are independent effects.

      We also include a new Fig. 6 providing additional information, including the development of preference over time, and addressing the question whether the modulations in locomotion can explain differences in odor preference.

      Locomotion behavior was assessed in 30s light-on periods - was the behavior different from the memory test or naïve preference test which had light on for 3 (or 2.5) minutes - which light intensity was used for the data shown in Figure 5.2? How did the larvae move in the high light concentration/ or with ATR - did they not show memory due to impaired locomotion?

      Light intensity for all odour preference and learning experiments with ChR2-XXL was 100 µW/cm<sup>2</sup> (except Fig. 2 – S2C), i.e. equivalent to the experiments in Fig. 5 – S1 and Fig. 9 – S2 (weak light). We replaced Fig. 5 – S2 with a new expanded Fig. 6 analysing the locomotion of our experiments shown in Fig. 2F, 3F and Fig. 2 – S1F. We show in this figure that the larvae can move relatively normal and that the locomotion effects are weaker than in our experiments with 30s light periods. Unfortunately, we do not have videos available for all experiments and therefore cannot make a similar analysis for Fig. 2 – S2C and D when we used strong light or ATR feeding. The experimenters did not notice impaired locomotion during the experiment and the animals did show normal odour preferences similar to those shown in Fig. 3 – but the preferences were the same after paired and unpaired training, resulting in zero Memory Scores. Therefore, we do not think that the locomotion prevented the memory expression.

      In several experiments, even genetic controls seem to show learning with blue light activation. Thus, the light itself seems to activate DANs. The authors should discuss these effects and explain what this could mean for the findings. The light stimulus might not just activate the specific DAN that expresses the optogenetics, but also additionally other DANs which respond to light.

      We thank the reviewer to point this out and point out this caveat in lines 159-165.

      The authors speculate about the function of potential MB circuits - the DAN-MBON circuit is not well described so far and might be required for the US in test memory recall. A straightforward experiment to investigate the involvement of this circuit in punishment or safety memory recall would be to block dopamine receptors in the MBON.

      We very much agree to this suggestion, but believe these experiments are beyond the scope of the current study. We therefore decided to not perform these experiments for the current paper.

      Reviewer #2 (Recommendations for the authors):

      (1) As self-explanatory as the figures are, it would be interesting to also see controls in some of them. For example, in Figure 5, the effect size graph (Figure 5C) clarifies to an extent the difference between control genotypes and the experimental genotypes. It would be nice to see the results of genetic controls in Figures 5A and 5B instead of in Figure 4 - supplement 2.

      We originally decided to put the controls into the supplement to prevent the main figures to be overcrowded. We revised this decision upon the reviewer’s comment for Fig. 8 (originally 7). For Fig. 4 and 5 specifically, we decided to keep the current layout because each serves a different purpose: Fig. 4D-L, Fig. 5 – S1 and S2 present the actual data with all genotypes that were made in parallel and therefore can be compared directly. Fig. 4 – S2 aims to visualize the effect of the light by comparing all controls across all experiments, normalized to the same starting value. Fig. 5A and B aim to compare the shape and effect size of activating DANs on top of the effect of the light - therefore, we subtracted the controls in each experiment from the experimental group. We think that adding the controls’ behaviour to Fig. 5A and B would undermine the aim of this figure.

      (2) It is a bit unclear why bending and tail velocities were the parameters chosen to compare between groups while in most cases they were of lower relative importance according to Figure 4 - supplement 1. Elaborating on this would strengthen the differences in behavior and also the claims of this study.

      We thank the reviewer for the suggestion and tried to make our choice clearer. Please see our answer to the respective part of the public review.

      (3) In adults, it has been shown that the same DAN can encode opposing valence depending on whether it was activated before or after odor presentation. Discussing the importance of temporal order of stimulus processing would bolster the results regarding paired and unpaired training in Figure 3.

      We thank the reviewer for this very good suggestion – also in larvae, this temporal function has been described. We discuss these observations in relation to our results in lines in 481-494.

      Reviewer #3 (Recommendations for the authors):

      Toshima et al., as stated in the public reviews, have done an admirable job with this manuscript. Below are specific suggestions that could improve the manuscript. It is, of course, up to the authors to decide which ones to attend to.

      Treat controls consistently. In Figure 2 and others, parental controls are not pooled, but in Figure 3, for odor preference, controls are pooled.

      We agree that the same things should be treated in the same way throughout a study and normally adhere to this principle. We nevertheless made an exception for Fig. 3 only because its goal is to provide a post-hoc analysis across several replications of experiments, some of which included genetic controls, others not (from Fig. 2, Fig. 2-S2 and S3). Due to relatively small effect sizes and high variability in odor preferences, to answer the question of paired and unpaired learning, we need higher sample sizes than each individual experiment provided. We therefore decided to pool all “equivalent” data across all these experiments. We do agree that this is a suboptimal approach but hope the reviewer can understand the rationale behind it. We explained our rationale clearer now (lines 180-183).

      I prefer to see all data points in a graph. It is more transparent than the box plots. Also, could you note why the data median is preferable to show over the mean?

      Although we in principle agree to the notion that presenting all data points is more transparent, we opted against it as it makes some graphs harder to read in particular with high sample sizes – in some of our figures, we have hundreds of data points per group. We explain our choice, including for using the median, in the method section (lines 835-839).

      Please undertake another round of language editing to handle spelling errors, etc. Use consistent British/American English.

      We thank the reviewer for their suggestion and tried our best to fix any spelling and grammar errors.

      I urge the authors to move beyond the false dichotomy of 'p' value statistics to using the statistical framework of estimation statistics for data analysis. I understand switching from familiar statistical analysis in such a late manuscript stage is very difficult. However, the authors can consider the estimation statistics framework in subsequent studies. https://www.estimationstats.com is a good starting point for biologists to get to know a framework that has been extensively worked on and is arguably a more 'honest' way of analyzing data. Disclaimer: I am not associated with the above website.

      We agree that the p-value has problems and are aware of the estimation statistics framework. We had considered applying it here, but we decided against switching to a completely different statistical framework for a research project that was ongoing since several years. However, we are sincerely considering it for our current research projects.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) A major question that remains is why the mutations have so much more detrimental effect in MutL (100-fold lower k<sub>cat</sub>/K<sub>M</sub>) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      We agree that the quantitative effects of the mutations differ between MutL and GyrB. However, we do not think that this difference argues against conservation of the catalytic mechanism. The trends of the mutational effects are highly consistent between the two enzymes. In both proteins, replacement of the conserved catalytic glutamate with Ala (E29A in aqMutL and E48A in aqGyrB) abolished ATPase activity and ATP binding, whereas replacement with the isosteric amide residue (E29Q and E48Q), which preserves hydrogen-bonding capability but lacks proton-accepting capacity, retained ATP binding and measurable ATPase activity. Likewise, substitutions of the second acidic residue (E32Q in aqMutL and D51N in aqGyrB) also retained substantial ATPase activity despite the loss of proton-accepting capability. Most importantly, simultaneous substitution of both acidic residues (E29Q/E32Q in aqMutL and E48Q/D51N in aqGyrB) completely abolished ATPase activity in both enzymes. Therefore, although the magnitude of the activity reduction caused by the individual mutations differs between MutL and GyrB, the qualitative pattern is essentially identical. We therefore think that the proposed catalytic mechanism is conserved, while the quantitative differences likely reflect differences in the local catalytic environment rather than differences in the underlying mechanism.

      (2) The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the acidic residues and their mutants is not shown.

      We have added Supplementary Fig. S2, which shows the 2F<sub>o</sub>–F<sub>c</sub> electron density maps around residues Glu48/Asp51 (or their substituted residues) in the wildtype and all mutant structures. The following sentences have been added in the revised manuscript:

      “To show that the introduced substitutions were unambiguously supported by the crystallographic data, electron density maps around residues 48 and 51 are shown in Supplementary Fig. S2.” (p. 4 line 198-200 in the revised manuscript)

      Reviewer #2 (Public review):

      (1) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of the effect (or likely effect) of disease-linked mutations in GHKL ATPases would have strengthened this study.

      We agree that extending the analysis to additional disease-associated variants in other human GHKL ATPases would further strengthen our understanding of the conserved catalytic mechanism and its clinical relevance. However, we believe that such a comprehensive analysis is beyond the scope of the present study, which focuses on establishing the fundamental catalytic mechanism shared between MutL and GyrB. We consider systematic functional and structural analyses of disease-associated variants across the GHKL ATPase family to be an important direction for future research. We have now added a statement to the Results and Discussion section to acknowledge this limitation and highlight this future perspective:

      “Although we focused here on pathogenic variants in the MutL homologs MLH1 and PMS2, extending similar structural and biochemical analyses to disease-associated variants in other human GHKL ATPases will be important for evaluating the generality and clinical relevance of the conserved catalytic mechanism proposed in this study.” (p. 6 line 303-306 in the revised manuscript)

      (2) In MLH1, the E37K mutation completely abolishes ATPase activity, but the corresponding mutations in aqMutL, aqGyrB, and PMS2 do not. It remains unclear why E37K in MLH1 leads to complete loss of activity, as the authors propose that water molecule positioning via the first acidic residue, as well as ATP lid stabilisation and associated conformational changes, should still be possible.

      We agree that the complete loss of ATPase activity caused by the MLH1 E37K variant cannot be explained solely by loss of the catalytic carboxylate. However, we note that the corresponding aqMutL E32K variant analyzed in this study also exhibited essentially no detectable ATPase activity, indicating that this phenotype is not unique to MLH1. It can be thought that the severe defect of the lysine variants arises not merely from loss of the acidic side chain but from charge reversal. We have clarified this point in the Results and Discussion sections:

      “In contrast, the E37K mutation in the MLH1 NTD completely abolished the ATPase activity under our assay conditions (Fig. 5B and Table 1) unlike the corresponding glutamine substitutions, which retained substantial residual ATPase activity in aqMutL, aqGyrB, and PMS2 NTDs. A similar complete loss of ATPase activity was also observed for the E32K mutant form of the aqMutL NTD. These observations suggest that the severe defect caused by the lysine substitution cannot be attributed simply to loss of the catalytic carboxylate. Instead, introduction of a positively charged side chain (charge reversal) is likely to perturb the local electrostatic environment. Structural characterization of the MLH1 E37K and aqMutL E32K mutant forms will be required to clarify the molecular basis of this severe functional defect.” (p. 6 line 281-289 in the revised manuscript)

      (3) The authors do not examine ATP binding in the E32 mutants of aqMutL NTD and the D51 mutants of aqGyrB, or AMPPNP binding of the NLH1 and PMS2 mutants. Hence, the relative contributions of the acidic residues to ATP binding and hydrolysis remain partially unclear.

      We performed additional ATP-binding experiments using the aqMutL NTD E32A and aqGyrB NTD D51A mutant forms. Both mutant forms exhibited ATP-binding activities comparable to those of the corresponding wildtype forms. These results support our conclusion that the second acidic residue primarily contributes to ATP hydrolysis rather than ATP binding, whereas the first acidic residue plays dual roles in ATP binding and catalysis.

      Although we agree that nucleotide-binding analyses of the MLH1 and PMS2 variants would be informative, these experiments were not feasible because the equilibrium dialysis assay requires high protein concentrations, which we were unable to obtain for the recombinant human MLH1 and PMS2 N-terminal domains.

      We have incorporated these new data into the Results and Discussion section:

      “In contrast to the E29A mutant form of the aqMutL NTD, the E32A mutant form exhibited ATP binding ability comparable to that of the wildtype form (Supplementary Fig. S1A), indicating that Glu32 does not contribute to ATP binding.” (p. 3 line 143-145 in the revised manuscript)

      “The D51A mutant form of the aqGyrB NTD retained ATP binding ability comparable to that of the wildtype form, indicating that Asp51 is not required for nucleotide binding (Supplementary Fig. S1B).” (p. 4 line 183-185 in the revised manuscript)

      (4) The ATPase assays for PMS2 and MLH1 (Figure 7 and Table 1) were performed with purification/solubility tags still present. Hence, it cannot be ruled out that these tags influence the measured activities.

      We thank the reviewer for raising this important point. We agree that the possible influence of the purification/solubility tags on the absolute ATPase activities of the PMS2 and MLH1 NTDs cannot be completely excluded. However, the wild-type and mutant forms for each homolog were analyzed using identical constructs under the same experimental conditions. Therefore, the affinity/solubility tags are unlikely to affect the relative comparisons of the mutational effects. Furthermore, because the affinity tags are located at the N terminus and are distant from the ATPase active site, they are unlikely to directly perturb the catalytic center.

      (5) The authors suggest that the two-acidic-residue mechanism proposed in this study could be shared among several GHKL ATPase families, yet they also state that the hydrogen-bonding network was not observed in MutL and MORC family proteins. This raises doubt about how conserved the mechanism is, e.g., in MutL and MORC proteins.

      We thank the reviewer for this insightful comment. Our proposed mechanism is based on the cooperative catalytic roles of the two conserved acidic residues, namely the involvement of the first acidic residue in ATP binding and nucleophilic water positioning and the role of the second acidic residue in proton abstraction. In contrast, the Glu48–Gln340 hydrogen-bonding interaction described in aqGyrB was proposed only as a structural feature that may modulate the contribution of the first acidic residue to ATP binding. It is not an essential component of the catalytic mechanism proposed in this study. Therefore, the absence of this particular hydrogen-bonding network in the currently available structures of MutL and MORC proteins does not argue against conservation of the catalytic mechanism itself.

      Recommendations for the authors:

      Reviewing Editor Comments:

      One of the structures (Crystal Structure of the E48A variant) has relatively poor statistics in the PDB validation report. Please improve this structure.

      We performed additional refinement of the E48A crystal structure. This resulted in a clear improvement in the overall model quality, with the Ramachandran favored residues increasing from 93.4% to 95.4%, the percentage of side-chain outliers decreasing from 6.1% to 1.4%. The refined structural model has been used throughout the revised manuscript, and the updated refinement statistics are provided in Table 2.

      Reviewer #1 (Recommendations for the authors):

      Please show conventional density maps (e.g., sigmaA weighted 2fo-fc maps).

      This comment is closely related to Comment (2) in the Public Review by the Reviewer #1. In response, we have added Supplementary Fig. S2, which presents conventional σA-weighted 2F<sub>o</sub>–F<sub>c</sub> electron density maps around the catalytic acidic residues in the wild-type and mutant aqGyrB structures.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding the analysis of clinical variants, it would be informative to note that the second allele is lost before tumor growth in Lynch syndrome.

      “Therefore, these variants might contribute to the development of Lynch syndrome by weakening the ATPase-driven regulatory functions of MutL.” (p. 6 line 280-281 in the original manuscript) has been changed to:

      “In individuals carrying these germline variants, subsequent loss or inactivation of the remaining wildtype allele would leave only the ATPase-defective MutL protein, thereby compromising mismatch repair and promoting tumorigenesis.” (p. 6 line 296-298 in the revised manuscript)

      (2) P. 4, in the paragraph "Conserved roles of two acidic residues of aqGyrB in ATP hydrolysis", the E48Q mutant retains approximately one third of the WT activity, not one quarter as stated in the text (Table 1). Additionally, later in the article, the D51 mutant is reported to retain approximately one-sixth (~17%) of the WT activity, rather than ~25% as written.

      We thank the reviewer for carefully identifying these inconsistencies. The text has been corrected to accurately reflect the data presented in Table 1: “…one third of the wildtype activity” (p. 4 line 180) and “…retaining ~16%...” (p. 4 line 186 in the revised manuscript)

      (3) P. 6, lines 275-276, this sentence should be rephrased for clarity, as the authors note at the end of page 5 that not all members of the GHKL ATPase family possess this second acidic residue.

      “…this second acidic residue plays a conserved and functionally significant role in ATP hydrolysis across the GHKL ATPase family.” in the original manuscript has been changed to:

      “…this second acidic residue plays a conserved and functionally significant role in ATP hydrolysis among some members of the GHKL ATPase family.” (p. 6 line 292 in the revised manuscript)

      (4) P. 9, in the "Data Accessibility Statement", the PDB code 23UY is missing. This entry corresponds to the crystal structure of the D51A mutant of aqGyrB NTD and should be included.

      The Data Accessibility Statement has been revised to include the code 23UY. (p. 9 line 454 in the revised manuscript)

      (5) P. 14, the table should be labelled "Table 2. Data collection and refinement statistics for the aqGyrB NTDs", rather than "Supplementary Table 2", to ensure consistency with how it is cited in the main text.

      The table title has been corrected from "Supplementary Table 2" to "Table 2”. (p. 4 line 198 in the revised manuscript)

      (6) It is difficult to determine from the figures whether the magnesium ion is positioned equivalently in aqMutL and aqGyrB. Did the authors observe any differences in ion positioning?

      To facilitate direct comparison of the catalytic Mg<sup>2+</sup> ion between the aqMutL and aqGyrB NTDs, we have added Supplementary Fig. S3, which shows a structural superimposition of the ATPase active sites of the two proteins:

      “Structural superposition of the aqGyrB NTD and aqMutL NTD revealed that the catalytic Mg<sup>2+</sup> ion occupies essentially the same position in the two ATPase active sites (Supplementary Fig. S3), indicating that the metal-binding geometry is highly conserved, where the Mg<sup>2+</sup> ion is coordinated by the side chain of the conserved Asn, AMPPNP, and surrounding water molecules. Neither Glu48 of aqMutL nor Asp51 of aqGyrB directly coordinated the Mg<sup>2+</sup> ion.” (p. 5 line 201-205 in the revised manuscript)

      (7) The authors should discuss the interaction between the aqGyrB NTD, Mg<sup>2+</sup>, and ATP during the binding step. In the case of the E48A mutant, where ATP binding is lost, does E48 directly establish contacts with Mg<sup>2+</sup>, or is another residue involved (with conformational changes preventing this interaction)?

      Our structural analyses indicate that Glu48 does not directly coordinate the catalytic Mg<sup>2+</sup> ion. Instead, as shown in Supplementary Fig. S3, the Mg<sup>2+</sup> ion is coordinated by the side chain of Asn52, AMPPNP, and surrounding water molecules. We have clarified this point in the Results and Discussion sections:

      “Structural superposition of the aqGyrB NTD and aqMutL NTD revealed that the catalytic Mg<sup>2+</sup> ion occupies essentially the same position in the two ATPase active sites (Supplementary Fig. S3), indicating that the metal-binding geometry is highly conserved, where the Mg<sup>2+</sup> ion is coordinated by the side chain of Asn52, AMPPNP, and surrounding water molecules. Neither Glu48 nor Asp51 directly coordinated the Mg<sup>2+</sup> ion.” (p. 5 line 201-205 in the revised manuscript)

      (8) Figures 1 and 3: Use ribbon representation and no shadows, at least for the inset panels, to enhance clarity and interpretability.

      We have revised Figures 1 and 3 by displaying the protein structures in ribbon representation and removing shadows from the inset panels.

      (9) Combine Figures 1 and 2, and combine Figures 3 and 4.

      Following the reviewer's recommendation, we have combined the original Figures 1 and 2 into a single figure and the original Figures 3 and 4 into another single figure.

      (10) Figure 5: Zoom in further and remove shadows. The current panels are not very effective in highlighting how ATP is bound by the different protein variants.

      Figure 5 has been revised by increasing the magnification of the ATP-binding sites and removing shadows from the structural renderings.

      (11) Figure 8. Add a scale bar to show evolutionary distance.

      We thank the reviewer for this helpful suggestion. To provide information on evolutionary distances while preserving the clarity of the main figure, we have added a new Supplementary Fig. S5 showing the same phylogenetic tree with branch lengths proportional to the inferred evolutionary distances and an evolutionary distance scale bar. Figure 6 has been retained in its simplified form with equal branch lengths to facilitate visualization of the ancestral-state reconstruction, and we have clarified this distinction in the Materials and Methods section:

      “For visualization purposes, branch lengths were not scaled and were displayed with equal lengths in Fig. 6. The corresponding phylogeny with branch lengths proportional to the inferred evolutionary distances is provided in Supplementary Fig. S5.” (p. 9 line 429-432 in the revised manuscript)

    1. Author response:

      We would like to thank the editor and reviewers for their thoughtful and constructive feedback. We appreciate the time and care devoted to reviewing our manuscript, as well as the recognition of rigorous experimental design, the technical challenges involved in conducting an fMRI study with two sensory modalities and two tasks in both deaf and hearing participants, and the value of the findings for understanding crossmodal plasticity and cortical organisation in deafness. We are encouraged by the overall assessment of the study, and appreciate the suggestions for strengthening the manuscript. Below, we provide a summary of how we plan to address the reviewers’ comments in our formal revision of the manuscript:

      (1) Additional analyses

      (a) We will incorporate behavioural performance measures into the relevant analyses to disentangle potential behavioural contributions to the observed effects.

      (b) We will calculate the noise ceiling value for each of the RSA analyses. 

      (c) We will conduct a correlation analysis between RDMs of auditory and control regions, to investigate the similarity between these computations and whether this is influenced by sensory experience.

      (d) Regarding the suggestion to conduct whole-brain searchlight analyses, we respectfully do not believe that this approach would address the primary research question of the study, namely whether and how representations within auditory cortex differ between deaf and hearing individuals. Our central hypotheses specifically concern representational content within predefined auditory cortical regions, making the ROI-based approach the most appropriate and sensitive method for testing these questions.

      Furthermore, the searchlight approach would require adequately powered group comparisons at the whole-brain level. Given the challenges associated with recruiting deaf native signers participants and the resulting sample size, we do not believe the study is sufficiently powered to draw reliable conclusions from this analysis. We will further clarify this rationale in the revised manuscript.

      (2) Presentation of the results

      We will revise the presentation of the findings to better guide the reader through the analyses and facilitate interpretation of the figures. In particular, we will ensure that significant effects, interactions, and their relationship to the corresponding figures are described more explicitly throughout the manuscript.

      (3) Revision of the discussion

      Following the reviewers’ feedback, we will revise the Discussion to more clearly distinguish between results that directly support a conclusion and hypotheses that remain speculative.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper leverages 7T fMRI data from the Natural Scenes Dataset to investigate whether retinotopic coding, the position-selective organization of visual response structures, spontaneous resting-state interactions between the Default Network (DN) and the Dorsal Attention Network (dATN). Using individualized network parcellations and population receptive field (pRF) modeling, the authors show that DN voxels can be split into two subpopulations based on their response to visual stimulation: those with position-specific positive BOLD responses (+pRFs) and those with position-specific negative BOLD responses (-pRFs). Critically, these subpopulations relate differently to the dATN during rest: -pRFs are anticorrelated with the dATN, +pRFs are positively correlated, and non-retinotopic DN voxels show no coupling. The anticorrelation (and positive correlation) is enhanced when DN and dATN voxels share visual field preferences. An eventtriggered analysis suggests that retinotopic coding shapes both "top-down" (DNinitiated) and "bottom-up" (dATN-initiated) spontaneous activity transients, supporting the claim that the retinotopic scaffold is intrinsic to the DN. These findings challenge the prevailing view of global DN-dATN antagonism and suggest retinotopic coding as an organizing principle for cross-network communication.

      Strengths:

      The central finding that what looks like network-level independence between DN and dATN decomposes into structured, bivalent interactions organized by voxellevel visual field preferences is a compelling demonstration that macro-scale network descriptions can hide meaningful substructure. The logic of the analysis is clean: pRF properties are estimated from retinotopic mapping data and then used to predict resting-state coupling in completely independent scanning sessions. This cross-session, cross-modality design rules out many circularity concerns.

      The use of individualized multi-session hierarchical Bayesian parcellation (Kong et al.) to define DN and dATN boundaries within each subject is the right methodological choice for this question. Network boundaries in posterior cortex, where DN and dATN interdigitate most closely, vary considerably across individuals, and group-average approaches would introduce exactly the kind of misassignment that would most confound the result.

      The matched-vs-random pRF analysis is well-controlled. The authors demonstrate that cortical distance between matched and randomly-matched dATN pRFs does not differ, effectively ruling out spatial proximity on the cortical surface as a confound. tSNR controls further show that signal quality differences do not drive the effect.

      The event-triggered analysis (Figure 3) is creative and adds genuine value. Showing that retinotopically-specific coupling persists during DN-initiated activity transients, not only dATN-initiated ones, is the key piece of evidence for the claim that the code is intrinsic to the DN rather than passively inherited through bottom-up visual drive.

      The result is observed consistently across all individual participants, which provides strong evidence for the robustness of the qualitative pattern despite the small sample size inherent to densely-sampled designs.

      Weaknesses

      (1) The nature of negative pRFs requires more scrutiny

      The entire interpretive framework depends on treating negative pRFs in the DN as genuine position-selective neural responses (suppression). However, negative BOLD signals are well known to arise from non-neural sources, specifically, vascular stealing (where activation in nearby tissue diverts blood from adjacent voxels) and macrovascular draining vein effects that produce spatially displaced signal inversions. These concerns are amplified at 7T, where T2*-weighted GEEPI carries substantial macrovascular weighting. The DN and dATN interdigitate extensively in the posterior cortex, often within millimeters. A negative pRF in a DN voxel adjacent to a positive dATN voxel could, in principle, reflect the hemodynamic shadow of its neighbor rather than an independent neural response.

      The spatial dispersion control (matched vs. random pRFs have similar cortical distribution) is valuable but addresses long-range confounds, not local hemodynamic crosstalk. The reliability of sign and center position across runs is reassuring but does not exclude a vascular origin, as vascular architecture is itself stable across sessions. I would encourage the authors to test whether the matched-vs-random effect survives exclusion of voxels near large pial vessels (identifiable from T2* contrast or the venograms available in the NSD). These analyses would not be dispositive, but they would meaningfully strengthen the neural interpretation.

      The reviewer raises an important concern about the interpretation of negative pRFs in the DN, namely that spatially specific negative BOLD responses could, in principle, reflect local vascular effects rather than genuine position-selective suppression. The reviewer suggests excluding voxels near large vessels to address this issue.

      Based on the reviewer’s suggestion, we repeated the pRF matching analysis excluding any voxels within 3mm of a major vein, as identified using the time-of-flight (TOF) MR venography included in the NSD. This analysis therefore tests whether the retinotopically specific DN–dATN coupling persists after removing voxels most likely to be affected by vascular signal.

      Excluding these voxels did not impact our results: we found preferential coupling according to response valence and center position, with stronger correlation between matched +DN and +dATN voxels (t(6) = 6.054, p < 0.001), and a more pronounced negative correlation between matched -DN and -dATN voxels (t(6) = -5.0448, p < 0.01). We have added these results to the supplemental figures (Fig. S7), and also added to the text (Pg. 8). Together with the run-wise reliability of pRF sign and position, and the persistence of the matched-versus-random effect after vessel exclusion, this analysis supports the interpretation that negative DN pRFs reflect structured, spatially specific responses rather than a vascular artifact.

      “Finally, to rule out any possible influences from vascular stealing (i.e. the shunting of blood into active tissue from nearby regions), we repeated the matching analysis after excluding any voxels within a 3mm radius of a major vessel (Fig. S7; see Methods). Both matching effects remained after excluding vascularly susceptible voxels (+DN x dATN: t(6) = 6.054; p < 0.001; -DN x dATN: t(6) = -5.0448; p < 0.01).”

      (2) Amount of retinotopic mapping data and choice of pRF pipeline

      The NSD includes 6 runs of retinotopic mapping (~5 minutes each; 3 baraperture, 3 wedge/ring). The authors use only the 3 bar-aperture runs (~15 minutes total per subject) and fit their own pRFs using AFNI's 3dNLfim procedure, rather than using the pRF estimates provided as part of the NSD release (which were fitted using the analyzePRF toolbox with all 6 runs).

      Fifteen minutes of bar data is quite limited for reliable voxel-wise pRF estimation, especially in regions far from the early visual cortex, where signal-to-noise is inherently lower. Standard recommendations for robust pRF mapping in higherorder regions generally suggest substantially more data. The variance-explained threshold is close to the noise floor by design, meaning that a non-trivial number of the "retinotopic" DN voxels may be poorly estimated. Given that the core analyses depend on both the sign and the center position of these pRFs, the limited data is a significant concern.

      The authors do not explain why they chose to re-fit pRFs rather than use the NSD-provided estimates. If the motivation was methodological (e.g., the NSD pRF pipeline does not readily yield signed amplitude, or the bar-only fits were judged more appropriate for detecting negative responses), this should be made explicit. If the NSD-provided pRFs can reproduce the key findings, this would substantially increase confidence in the results. If they cannot, that divergence itself would be important to understand. I would ask the authors to address this choice and, if feasible, to report whether the core results replicate using the NSDprovided pRF estimates and/or whether using all 6 runs of retinotopy data changes the findings.

      The reviewer raises two related concerns: first, that the amount of retinotopic mapping data available in the NSD may be limited for estimating voxel-wise pRFs in higher-order cortical regions; and second, that we re-fit the pRF model using AFNI rather than relying on the pRF estimates provided with the NSD release. We appreciate the opportunity to clarify both points. We agree with the reviewer that more travelling bar data would be preferable and would likely yield more robust model fits, particularly in higher-order regions with lower SNR. This is a limitation of our paper that we now acknowledge in the discussion section. However, we do not think that more data would fundamentally change the pattern of our results for the following reasons.

      First, we implemented a novel data-driven approach to derive a threshold for thresholding significant pRF fits (a noise floor). Importantly, our noise floor estimation yields a conservative threshold (R<sup>2</sup> > 0.14), which is greater than both our previous work characterizing cortical pRFs (Steel et al. 2024: R<sup>2</sup> > 0.08) and other work exploring visual responses in the default network (Klink et al. 2021: R<sup>2</sup> > 0.05; no threshold: Szinte and Knapen 2020; Knapen 2021).

      Second, the key pRF features used in our analyses – response sign and centre position – were reliable across retinotopic mapping runs. This reliability is important because our central matching analysis depends on voxel-wise estimates of both response valence and visual-field position.

      Third, the matching analysis asks whether pRF parameters estimated from the retinotopic mapping task predict functional coupling measured during independent resting-state scans. Noisy or unstable pRF estimates should weaken this relationship, because they would degrade the accuracy of voxel-wise matching. Thus, parameter instability would be expected to obscure retinotopically specific coupling rather than systematically produce the observed matched-versus-random effects.

      To the reviewer’s question about our decision to re-fit the pRF model using AFNI, the reviewer is correct that this was motivated by the requirements of our analysis: we re-fit the pRF estimates using AFNI because it allows for both positive and negative signed amplitudes. The pRF model fits provided with the NSD do not allow bivalent amplitude estimates. We have made this decision clearer in the text, reproduced below (Pg. 4-5; Pg. 14-15).

      “We chose to re-fit the data using a simple Gaussian approach as implemented in AFNI to allow for both positive and negative signed amplitudes.”

      “The limited amount of pRF mapping task data included in the NSD posed a challenge for establishing reliable visual response estimates. Here, we addressed this issue by developing a novel thresholding method to establish robust voxel-wise model fits. Among voxels that passed this empirical threshold, we observed a significant correlation in voxel-wise estimates of centre position and visual response amplitude. In addition, our pRF matching results were based on the relationship between the voxel-wise estimates of centre position and response amplitude with resting-state fMRI – a completely independent measure. Crucially, noisy estimates of pRF parameters would obscure this relationship and make our results less likely. Therefore, despite the relatively limited pRF mapping data available, unstable pRF estimates are unlikely to drive our results.”

      (3) pRF model adequacy for the Default Network

      The isotropic Gaussian pRF model was developed for and validated in early and mid-level visual cortex, where it captures the dominant spatial selectivity of neuronal populations. In DN voxels where the model explains comparatively little variance, it is less clear that the model is capturing the right quantity.

      Specifically, the negative pRFs could conceivably be described by a model with a dominant suppressive surround (e.g., a difference-of-Gaussians model), in which what appears as a "negative pRF" in the standard model is actually the surround component of a center-surround mechanism whose center is poorly resolved. This distinction matters: a genuine inverted code (negative center response) implies a qualitatively different computation than inherited surround suppression from nearby visual cortex.

      The authors should consider discussing why the standard model is sufficient for the questions asked, or ideally, testing whether the sign distinction survives under alternative pRF model specifications.

      We appreciate the reviewer’s comment about the limitations of a single gaussian pRF model. We chose the single gaussian model as a direct extension of prior work from our lab and others (Steel et al., 2024, Klink et al., 2022, Szinte and Knapen, 2021). We agree that a negative response in this model could, in principle, reflect a more complex spatial profile, such as a dominant suppressive surround. However, adjudicating among alternative pRF models would require more retinotopic mapping data than are available in the NSD, particularly for higher-order cortex. Thus, we feel that it is outside the scope of the current work. We now address this limitation in our discussion (Pg. 15).

      “Relatedly, here we used a single gaussian model, consistent with prior work on negative visual responses in memory systems (31, 33, 34). However, other models of visual receptive fields might offer further insight into the DN’s visual responsiveness, such as double gaussian models of surround suppression (65) or compressive summation (66). Future studies might directly compare different visual models to further refine the computations underpinning visual responses in the DN.”

      (4) Interpreting resting-state transients as top-down vs. bottom-up The event-triggered analysis labels high-amplitude DN pRF activations as "topdown events" and dATN activations as "bottom-up events." This is a reasonable inference given experience-sampling studies showing that rest involves alternation between internal and external attention, but it remains an inference. Without concurrent experience sampling, eye-tracking, or physiological monitoring, we cannot establish that a spontaneous DN transient reflects memory retrieval or internally-directed thought rather than a global arousal fluctuation. Similarly, dATN transients during rest could reflect covert shifts of spatial attention to remembered or imagined locations rather than bottom-up processing per se. I would ask the authors to soften this framing or to discuss what additional data would be needed to validate the top-down/bottom-up attribution.

      The reviewer raises an important concern about the strong interpretation of elevated BOLD activity detected in the DN and dATN as top-down and bottom-up events. We agree that the limitations of fMRI in our current data prevent these strong claims about the origin of these signals. We have therefore softened this framing throughout the manuscript, and we now refer to these events as DN-driven and dATN-driven. We think that this more directly describes the analysis: events were defined by transient high-amplitude activity in DN or dATN pRFs, respectively.  

      (5) The "retinotopic code" vs. "visual field bias" distinction The paper uses the language of a "retinotopic code" throughout and correctly distinguishes this from a "retinotopic map," noting that DN voxels do not form a continuous topographic representation on the cortical surface. This distinction deserves greater emphasis. In vision science, retinotopic maps carry computational significance through their topographic continuity and relationship to cortical wiring. A distributed collection of voxels with coarse visual field preferences but no cortical topography is a fundamentally different organizational feature. Recent reviews have drawn an explicit distinction between retinotopic maps and visual field biases (Groen, Dekker, Knapen & Silson, TiCS 2022), and the present findings may be more accurately characterized as the latter. Perhaps the authors think that the distinction is merely a signal-to-noise distinction, in which case I would invite them to clearly speak to this interpretation. In any case, this is not a criticism of the findings themselves, but clarity on this point would prevent conflation of two different organizational principles and would help position the work for both the vision and network neuroscience communities.

      The reviewer raises a valuable point about the distinction between a retinotopic code, a retinotopic map, and a visual field bias, and we are happy to add discussion of this topic to our manuscript.

      Our results show that the DN does not exhibit a continuous retinotopic map in the sense used in early visual cortex. Rather, our results suggest a distributed voxel-level code for visual-field position: individual DN voxels show reliable spatial preferences, and these preferences predict retinotopically specific functional coupling with dATN voxels. This voxel-level organization is analogous to other distributed spatial codes, such as head-direction coding in retrosplenial cortex, where spatial variables are represented by population activity without requiring a topographic map on the cortical surface. This differs from a coarse visual-field bias, including preferential responses to the contralateral visual field, although we do also observe such biases. We have added text unpacking this important distinction to the Discussion (Pg. 15-16):

      “Prior work has emphasized the visual response bias in regions where voxel-wise retinotopic responses lack a map-like organization(35); overall, the DN does exhibit this kind of bias. However, our results show that the voxel-scale activity underpinning this bias reflects the latent connectivity of those voxels. Thus, we adopt the term “retinotopic coding”, because this voxel-scale coding scheme exists without a map-like organization on the cortical surface. For example, rodent and bat head direction cells are not laid out in a literal ring, but the population code of these neurons forms a ring manifold(68, 69).”

      Reviewer #2 (Public review):

      Summary:

      Using a public dataset of retinotopic mapping and resting-state data, the authors find that the default mode network has voxels that respond (positively or negatively) to visual stimulation at specific retinotopic positions, and that restingstate activity in these voxels is correlated with activity in more traditional sensory voxels with the same visual-location preference. The retinotopic specificity is bidirectional, such that high activity in default mode voxels drives activity only in voxels with matching receptive fields in sensory cortex, and vice versa. These findings are at odds with traditional views of the default mode network as having abstract (non-retinotopic) representations and competing (rather than cooperating) with external sensory representations.

      Strengths:

      This study continues an intriguing line of research about how default mode regions interact with the sensory cortex. Demonstrating that there are structured interactions between these regions at rest, and that these interactions are in fact organized according to retinotopic location (as opposed to traditional views of representational format in the default mode network), provides a new framework for thinking about large-scale internal and external brain networks. The authors make use of a well-powered public dataset that allows for precise estimates of pRFs and individual-specific resting-state networks, and develop a number of interesting analyses that characterize the relationships between DN and dATN voxels. The findings are exciting and could have a major impact on future studies in cognitive neuroimaging.

      The authors mention that these findings could shed light on internal/external interactions such as "anticipatory saccades or memory-guided attention," which is true, though I would argue that constructing DN representations of external stimuli is in fact even more fundamental than these specific cases (e.g., see Barnett and Bellana, 2025, "Situation models and the default mode network"). The "highways" identified in this study could play a vital role in real-world perceptual processes that are constantly translating external input into internal mental models.

      Weaknesses:

      (1) The criterion used for defining voxels as retinotopic seems very liberal. The authors show that only 5% of voxels have R^2>0.14 in a null analysis, and therefore define voxels with R^2>0.14 as retinotopic. Although all the networks in 1C show voxel distributions that differ from the null, the number of false positives above R^2>0.14 seems problematic, especially for the DN positive pRFs (red distribution) and to a lesser extent the DN negative pRFs (blue distribution). From visual inspection of the plot, the false discovery rate (fraction of voxels labeled as retinotopic that are false positives) looks like it would be greater than 50% for the DN-positive pRFs. The authors do show that the positive pRF voxels have abovechance consistency across runs, again providing evidence that there are true positive voxels in this set, but perhaps a stricter criterion (such as having consistent negative fits across runs) would provide more targeted identification of the DN voxels with true retinotopic sensitivity.

      We thank the reviewer for giving us the opportunity to discuss this important decision. We agree with the reviewer that a stricter R<sup>2</sup> criterion could result in more targeted pRF identification. Motivated by the reviewer’s suggestion, we repeated the cross-region pRF matching analysis across multiple R2 thresholds.

      The retinotopic matching effects were not dependent on the original threshold. In fact, we found that the pRF matching effects are enhanced as the R<sup>2</sup> value increases (Fig. S5). This pattern suggests that any false-positive voxels admitted near the original threshold would dilute, rather than drive, the observed matched-versus-random effects. We have added text to the results highlighting this finding (Pg. 7):

      “In contrast, DN voxels that responded positively to visual stimulation (DN positive pRFs, +pRFs) had a positive correlation with the dATN (mean correlation = 0.284±0.152, t(6) = 4.96, p = 0.0025), while DN voxels with systematic negative responses to visual stimulation (DN negative pRFs, -pRFs) were anti-correlated with the dATN (mean correlation = -0.21±0.149, t(6) = -3.75, p = 0.0094). This relationship was further strengthened by adopting more conservative R<sup>2</sup> thresholds up to 0.30 despite the overall number of included voxels decreasing, suggesting that this effect is not driven by false-positive voxels at the edge of our threshold criteria (Fig. S5).”

      (2) The claim that "opponency at rest between the DN and dATN appears to be driven by the subset of DN voxels with negative retinotopic tuning" is not well supported. The fraction of DN voxels with negative pRFs is small: 9.42% of DN voxels have pRFs, and 58.77% are negative, so about 6% of DN voxels have negative pRFs. The fact that any DN voxels have negative pRFs is notable, but the authors do not provide evidence that these 6% are driving the overall behavior of the DN. They do show (e.g., in Figure 2B) that negative and positive pRFs have opposing influences, but the overall correlation with dATN does not look similar to the negative pRF connectivity. I'm also unsure whether "opponency" is a reasonable description for two networks that are "independent (i.e., not correlated)" in this analysis.

      The reviewer raises an important point about whether negative DN pRFs should be described as driving the overall DN–dATN relationship. We agree that this language was too strong. Negative pRFs constitute a small subset of DN voxels, and our analyses show that this subset has a distinct pattern of functional coupling with the dATN, not that it explains the global relationship between the DN and dATN as a whole.

      We have therefore revised the manuscript to avoid implying that negative DN pRFs drive overall DN–dATN opponency. Instead, we now frame these voxels as an important retinotopically tuned subpopulation nested within broader network dynamics. Specifically, our results show that visually responsive DN voxels are not homogeneous: positive and negative DN pRFs show opposing patterns of coupling with dATN pRFs, and these interactions are strengthened when voxels share visual-field preferences. This suggests that a small but structured subset of DN voxels may provide a route for retinotopically specific communication between internally and externally oriented networks, without implying that this subset determines the mean activity pattern of the entire DN:

      “Spontaneous DN and dATN activity during rest is uncorrelated at the network level. However, voxel-scale functional coupling across networks is shaped by the latent visual field preferences of individual voxels in each network, as measured during independent retinotopic mapping.” Abstract (Pg. 2)

      “This result shows that voxel-level visual response profiles shape DN-dATN coupling during spontaneous resting-state activity. Specifically, the DN and dATN activation is independent during rest. However, at the voxel-level, specific sub-groups of DN voxels have distinct coupling patterns with the dATN that depends on the valence of voxels’ visual responses. DN and dATN voxels with positive visual responses show a positive relationship during rest, and a notable subset of DN voxels with negative visual responses display the canonical opponency with dATN voxels. This suggests that retinotopic coding may be a mechanism that enables visual information to be exchanged between these large-scale brain systems. Specifically, opponency at rest between the DN and dATN appears to be driven by the subset of DN voxels with negative retinotopic tuning.” Results (Pg. 7)

      These findings offer a multi-scale account of neural communication, in which interactions among sub-populations of voxels with shared tuning preferences are nested within macro-scale network dynamics. Nesting multiple neural codes might enable ongoing computations within a larger brain system (e.g., attending to internal mental states within the DN during memory recall), while simultaneously allowing for the sharing of fine-grained representations across brain systems (34). Discussion (Pg. 14)

      (3) The event-triggered analysis is effective at testing the bidirectional relationship between DN and dATN, with high activity in either network triggering a response in the other network. However, it would be helpful to show more validation that these "events" are meaningful windows of time to study. First, is 13 TRs a typical length of time that activity is elevated during one of these events? Second, the top-down and bottom-up terminology is perhaps too loaded and not well-justified; if the negative pRFs in the DN reflect a meaningful coding system, then couldn't low (rather than high) activity indicate a top-down event?

      We thank the reviewer for these helpful suggestions. To the best of our knowledge, there is not currently a widely agreed-upon time window for performing event-based fMRI analyses. We chose a 13 TR time window to balance between sufficiently capturing BOLD signal related to the chosen event while also minimizing influence from other signal fluctuations, based on the procedure adopted in Gordon et al. (Nature, 2023) and Mitra et al. (J. Neuro Phys, 2014), which considered temporal relationships among brain regions over comparable timescales. In our analysis, this window considered 6 TRs (9.6s) on either side of the detected event, which we felt comfortably captures the peak BOLD signal that would result from an impulse at the event time, and responses that may reflect upstream activity leading into it.

      The reviewer has raised an additional comment about the terms “top-down” and “bottom-up.” These concerns were shared by Reviewer 1. Based on these comments, we have adopted the terms “DN-driven” and “dATN-driven”, which we think aligns more closely with our analysis approach.

      (4) The framing of this paper relative to the authors' past work, such as Steel et al. 2024 ("A retinotopic code structures the interaction between perception and memory systems"), could be improved. The existence of negative pRFs in the DN and a functional relationship between these pRFs and the sensory pRFs have already been described in prior work. My understanding of the primary novelty here is that this paper examines resting-state data, showing that there are widespread spontaneous interactions between broad internal and external networks, but this distinction is not made explicit in the Introduction.

      We appreciate the opportunity to clarify the novel aspects of our paper. The reviewer correctly identifies the extent of prior work, which identified -pRFs in regions of the canonical default network (Szinte and Knapen, 2021; Klink et al., 2022) and characterized the local interactions between adjacent perceptual and mnemonic regions (Steel et al., 2024). Our current work builds upon these findings in two key ways.

      First, we explore the effect across individually-defined whole brain networks. While the DN and dATN are often adjacent, these networks are spatially discontinuous and are comprised of distinct sub-regions (e.g. in prefrontal cortex). Whether retinotopic patterning of activity would persist in distributed networks could not have been extrapolated from our prior work. We think that finding will be of broad interest to the community studying perception and memory systems, because it offers a mechanistic account of how information is read in/out of memory.

      Second, here we considered whether spontaneous activity across networks would be structured by a retinotopic code. Our previous work characterized activity during tasks that depended on visual information: either scene perception or mental imagery. While the prior work was an important first step, it left open the possibility that retinotopic coding may only be relevant in visual tasks. By demonstrating that the retinotopic coding structures voxel-specific coactivation during rest, which entails no overt visual demands, we provide evidence that retinotopic features are a general, mode-agnostic code between regions.

      (5) The definition of the default mode (DN) in this study aligns with past research, but the definition of the dorsal attention network (dATN) seems at odds with standard terminology. For example, the authors cite Fox et al. 2006, which depicts the dATN as including regions such as IPS, FEF, SMA, and MT+. Here, however, the "dATN" seems to be primarily lateral and ventral visual cortex (e.g., Figure S5). The exact location of these sensory pRFs is not critical to the authors' claims, but this labeling seems incorrect, and the motivation for defining/selecting the sensory network in this way is not described.

      We thank the reviewer for this insightful comment and their careful consideration of our network definition.

      Our method for network identification, and the topography of the resulting networks, are broadly consistent with more recent conceptualizations of the DN and dATN (e.g. Du et al. 2024, Gordon et al. 2017, Braga and Buckner 2017). Relatedly, because we defined brain networks based on the unique connectivity patterns of each individual participant, we expect them to differ from previous group-level network descriptions. The increased resolution of the 7T data in the NSD may also result in greater departure from prior definitions compared to previous work done at 3T.

      Further study into dATN differences between group-level 3T, individualized 3T, and individualized 7T networks could be a valuable future direction, but this is outside the scope of this work.

      Reviewer #3 (Public review):

      Summary:

      This paper addresses an important question (the relationship between DN and dATN, and the role of retinotopic coding) and uses a set of novel analyses.

      Strengths:

      Important question, novel analytical approaches (pRF-informed functional connectivity analysis).

      Weaknesses:

      Some of the key claims are not fully supported by the data presented. There is also a concern about over-interpretation of the results. Key issues:

      (1)  The authors claim that retinotopic coding scaffolds the interaction between DMN and dATN. However, retinotopically tuned voxels account for a mere 9% of DMN voxels. So this appears to be a major overstatement. For instance, the statement that "these findings would position retinotopy as a unifying framework for brain-wide information processing" is not justified given the presented data.

      We appreciate the reviewer’s concern about the framing of our conclusion, which was shared by reviewers 1 and 2. In response to these comments, we have revised our paper to more accurately reflect the observed data. Specifically, we focus on the specific sub-populations of voxels within the DN and dATN that show retinotopic responses, and we have removed references to explaining the overall pattern of activity across networks.

      (2) Given that positive pRF voxels in DMN positively correlate with dATN voxels and negative pRF voxels in DMN negatively correlate with dATN voxels, there is a concern that these results could be contributed to by imprecise brain network parcellations. E.g., could some of the positive pRF voxels in DMN be erroneously assigned to DMN and actually belong to one of the other task-positive networks? There is insufficient validation of network parcellation to put this worry to rest, especially since it depends on ICA, which has a degree of arbitrariness built in.

      We thank the reviewer for the opportunity to clarify our method for network definition.

      Precision functional mapping is a growing field with many methods for defining personalized functional networks for each individual. Because the NSD resting-state data is relatively high resolution, we chose an approach designed to improve the stability of voxel-wise network assignment: Multi-Session Hierarchical Bayesian Modeling approach (Kong et al. 2019; Du et al. 2024). This approach enhances stability of network assignment by including a group-based prior and accounting for both within- and across-subject variability. This approach is more stable than ICA, and, because this approach leverages a prior, there is less concern about arbitrary or idiosyncratic network definitions.

      However, it is still common for network assignments to have lower confidence around the borders between networks. Yet, we also do not think border misassignment is likely to explain the present results for two reasons: first, while DN and dATN nodes are sometimes adjacent, there are many regions where they are spatially distant, such as the IPS for dATN and the lateral temporal lobe for DN. Second, the DN pRFs do not appear to cluster selectively along DN–dATN borders, suggesting that they are not simply misassigned dATN voxels.(Fig. S3) Therefore, we think voxels on the edge of these networks are unlikely to drive the effects observed here (see Fig. S3).

      (3) The claim that retinotopic coding is intrinsic to the DN network is not supported by rigorous analysis and results. The analysis here has many arbitrary factors, including: the threshold of the 99th percentile of resting-state distribution; the designation of DN as "top-down" and dATN as "bottom-up"; the definition of "anti-matched" voxels instead of using randomly selected voxels; and the statistics being paired between matched and anti-matched voxels instead of using comparisons to baseline. Overall, I do not think that the result supports the conclusion that retinotopic coding in DN is intrinsic instead of being bottomup-driven, given the very high threshold (99%) used and the fact that many other networks could also send bottom-up input to DN. Furthermore, the idea that bottom-up inputs only occur when the dATN (or any other RSN)'s spontaneous BOLD activity is above a certain threshold is a huge and unvalidated assumption.

      The reviewer raises several interesting concerns about decisions in our event-detection analysis. Here, we clarify the rationale for several analytic choices:

      (1) The 99th-percentile threshold was chosen to identify sparse, high-amplitude events while minimizing contamination from smaller ongoing fluctuations.

      (2) The other reviewers also noted a concern with the top-down/bottom-up terminology. We have revised these terms to DN-driven and dATN-driven, which we think reflect our approach more accurately.

      (3) We used anti-matched rather than randomly selected voxels because the full event-by-voxel randomization procedure was computationally intractable at the network level.

      (4) We did not understand the reviewer’s contention about activation baseline, but we would welcome clarification.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor points

      (1) The reliability analysis (Figure 1D) notes that dATN negative pRF amplitude was not reliable above chance in 2 of 7 participants. This could be discussed more prominently as it suggests that negative pRFs may not be stable features in all networks, which tempers the generality of the sign distinction as a fundamental organizational property.

      We thank the reviewer for raising this point of clarification. It is true that 2/7 participants did not show reliable negative pRFs in the dATN. However, the majority of participants show stable negative pRFs, and even in these 2 participants, the negative result does not indicate that negative pRFs would not be stable in those individuals with additional data. 

      Based on the reviewer’s comment, we have added emphasis to this point, but we do not feel that this warrants greater discussion in the paper. 

      Both positive and negative pRF amplitude was reliable in the DN for all subjects. In the dATN, positive amplitude pRFs were reliable in all participants, and negative amplitude pRFs (which constituted a small proportion of the overall pRFs in this network) were reliable in 5/7 participants. For the remainder of the paper, we only consider positive pRFs in the dATN. Importantly, pRF center position was highly reproducible across runs of pRF data in the dATN and DN in all subjects (Fig. 1D).  Pg. 5

      (2) The paper would benefit from situating the findings more explicitly within the cortical gradient framework (Margulies et al., 2016), which predicts that DN regions have maximally abstract, transmodal codes. The present findings complicate this view productively and deserve to be "situated" within that ongoing debate.

      We agree that the gradient framework is interesting, and we have added discussion of Margulies to our paper. (Pg. 16)

      Relatedly, the DN is considered a transmodal hub for cortical processing, where disparate sensory and motor processes converge (59, 75) The DN’s position at the cortical apex implies connections with and influence over unimodal cortical areas. However, the mechanism for liking unimodal and transmodal networks had been unknown. Prior work posited that sensory coding in transmodal areas might serve this function (31, 35), and our data provide direct empirical support for this account: specific visually-responsive voxels provide an input/output interface linking perceptual and memory systems. This complements work delineating specific affective and effective subregions within the DN that link the DN to other brain areas (76). Thus, while the DN may be “distant from input” (28), these data suggest that it is not disengaged from sensory processing.

      (3) It would be informative to know whether the *proportion* of negative vs. positive pRFs differs between DN-A and DN-B, given their distinct functional roles.

      Despite the functional specialization of DN-A and DN-B, and the slightly higher mean proportion of negative pRFs in DN-A (61% vs 56%), we found no statistically significant difference in the proportion of negative pRFs across the two networks (t(6) = 0.888, p = 0.409).

      (4) Low N is inherent to the densely-sampled NSD design, and the within-subject consistency is a strength. Nevertheless, with 6 degrees of freedom, the precision of specific quantitative estimates (e.g., that 58.77% of DN pRFs are negative) is uncertain, and the authors should be cautious about the generalizability of these point estimates.

      The reviewer raises a concern about the inclusion of specific levels of decimal place in our statistical reporting. We do not think that this is a major issue with the paper, but we are willing to change if the reviewer feels strongly.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1C could use an explicit legend (I believe it is following the color convention from the bar plots in 1F?). Also, for consistency, it would be helpful to make all the colormaps in 1F correspond to the bars (i.e., change the dATN colormap to go white->green).

      We thank the reviewer for this suggestion, and we have added explicit labels to Fig. 1C

      (2) Providing a scatter plot, in which each dot is a voxel and the x and y axes are the pRF amplitude estimates in different runs, could help provide evidence that there are voxels with pRFs that have consistently negative amplitudes across runs. This would also go beyond the binary consistency analysis in Figure 1D to show that the magnitudes of the amplitude estimates are also consistent.

      We thank the reviewer for this suggestion. We feel that the binary consistency conveys sufficient information. Because the analysis is done using pairwise correlation, how the scatter plot would reflect the three-way consistency is not clear. 

      (3) For understanding how the overall correlation between DN and dATN could be driven by voxel populations with opposing effects (e.g., Figure 2B), it would be useful to show how the +pRF and -pRF voxels compare to other voxels within the DN. For example, are these the voxels with the strongest negative and positive correlations with dATN, or are there many other DN voxels (among the 90% that do not have pRFs) that also have similarly-strong dATN correlations?

      The reviewer offers a very interesting suggestion. Based on the reviewer’s suggestion, we have refocused our paper on the particular subpopulations of +/- pRFs in the DN, rather than on an explanation for the overall pattern of correlation between the DN and dATN. Because our revised framing focuses on the properties of these retinotopically defined voxel populations, rather than on explaining whole-network DN–dATN coupling, we have not added this additional analysis. We have revised the relevant text to avoid implying that these pRF subpopulations drive the overall network-level relationship.

      (4) Initially, the baseline comparison pRFs for the matched pRFs are labeled "random" pRFs, which seems misleading; these are closer to "mismatched"/"anti-matched" pRFs since they are selected from the 1/3 that are farthest away. Then the comparison switched to using the anti-matched pRFs that are the 10 very farthest away, though I didn't understand the rationale that "the large number of pRFs made the random matching procedure impractical" - in what way is the number of pRFs larger in this analysis? Having a more consistent baseline (e.g., just using the 10 anti-matched pRFs the whole time) would be easier to interpret.

      We thank the reviewer for this suggestion. We have compared the results between the randomly-sampled bottom ⅓ matched versus the 10 worst matched, and the pattern of results is identical (the effect is strongest in the 10 worst matched). Therefore, we include the bottom ⅓ matched in the main text as a more conservative test of this effect. We are happy to include this as a supplemental figure if the reviewer feels it is essential. 

      (5) In the past, I have only seen the terminology "bootstrapped" to refer to sampling with replacement from the data sample, producing samples/statistics that are centered on the observed data. Here (lines 704-708), the sampling is coming from the null distribution of randomly-chosen voxels, and therefore the term "bootstrapped" would not apply (and could just be replaced with "null").

      We have revised this terminology in the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract and Discussion should be significantly toned down. E.g., the claim that "These findings challenge the prevailing view of global DN-dATN antagonism" is not really supported by the data provided. The claim that "retinotopic coding underpins the dynamic coordination of perception and thought" is also unsupported by the presented data.

      We have revised the manuscript in light of this comment.

      (2) Line 233-235: The null statistical result cannot support the claim reached here. Correlation analysis or Bayesian statistics should be used.

      We have revised the manuscript in light of this comment.

      (3) Line 250-254: Comparison to baseline should be used, in addition to comparing matched and random voxels.

      We agree that baseline comparisons can be useful in event-triggered analyses. However, for the pRF-matching analysis discussed here, the critical question is whether shared visual-field preference influences resting-state functional coupling between DN and dATN voxels. For this question, we believe that the appropriate baseline is the coupling observed for pRFs that do not share visual-field preferences. We therefore compare retinotopically matched pRFs to randomly matched pRFs drawn from the same networks. 

      (4) Line 271: "not" is missing.

      We have revised the manuscript in light of this comment.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      Forbes et al. developed an integrated approach to identify cis-regulatory elements (CREs) in the large (3.6 Gbp) genome of the crustacean Parhyale hawaiensis, addressing the challenge of pinpointing these regions among large regions of non-coding sequences. They combined ATAC-seq chromatin accessibility profiling (both bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis). Without assembling congener genomes, they mapped reads with low stringency to the P. hawaiensis reference, identifying about 55k conserved islands that overlap ATAC peaks more than expected by chance. This dual filter was used to select CRE candidates for transgenic reporter validation, yielding 6 functional elements (out of 11 tested) driving ubiquitous, neuronal, or muscle-specific expression, a major advance for non-model systems with large genomes.

      Strengths:

      Forbes et al. generated high-quality ATAC data across multiple scales. Using bulk ATAC-seq (from whole embryos, developing and adult legs), they identified tens of thousands of open chromatin peaks across the assembled P. hawaiensis large genome. Moreover, using single-nucleus ATAC-seq from adult legs, they could resolve differentially accessible chromatin profiles across over 15 cell types previously identified by scRNA-seq, enabling cell-type-specific candidate selection.

      Furthermore, their innovative low-coverage comparative genomics method mapped 0.46-6.4% of congener reads to P. hawaiensis without genome assembly, revealing hundreds of thousands of conserved non-coding islands, including about 55k showing conservation in all four species, far exceeding random expectation.

      Using the developed approach, the authors could validate 6 (out of 11 candidates) reporter constructs, driving robust ubiquitous and tissue-specific expression, succeeding where prior promoter-only screening failed and providing immediately useful genetic tools for the Parhyale community.

      Weaknesses:

      The primary limitation is that functional CRE testing was performed only in P. hawaiensis. While conservation maps are valuable resources, the manuscript lacks functional validation in congener species, limiting claims about broad applicability across related genomes/species.

      The approach also failed to validate developmental CREs. None of the candidates from combined ATAC and conservation filtering drove reporter expression matching endogenous patterns. The authors appropriately hypothesize technical limits (low expression) or biological factors (long-range enhancers, shadow enhancers).

      Overall Assessment:

      Forbes et al. fully succeed with their integrated approach to (1) generate an ATAC-seq atlas plus functional CRE discovery and (2) innovative low-coverage sequencing for conservation mapping in the large 3.6 Gbp genome of Parhyale hawaiensis. Their combination of ATAC-seq chromatin accessibility profiling (bulk and single-nucleus) across embryonic and adult tissues with low-coverage genome sequencing of three congeneric species (P. aquilina, P. darvishi, P. plumicornis), without congener genome assembly, drastically shrank the CRE search space. Using this approach, the authors could validate six out of 11 candidate transgenic reporters (ubiquitous, neuronal, and muscle-specific), where prior promoter-only screening failed.

      The low-coverage mapping innovation cuts cost and labour while snATAC-seq provides cell-type resolution, making these resources valuable for building new genetic and imaging tools in Parhyale.

      This compelling method also has the potential to enable labs with limited resources to identify and characterize regulatory elements in more non-model organisms, advancing our understanding of their evolution while establishing a scalable pipeline for large-genome systems.

      We thank the reviewer for their comments and valuable feedback.

      Reviewer #1 (Recommendations for the authors):

      (1) Standardize terminology and introduce acronyms properly:

      I suggest standardizing technique names throughout the manuscript (e.g., ATAC-seq rather than ATACseq, RNA-seq rather than RNAseq, ChIP-seq rather than ChipSeq). Please also introduce technical terms with their full name at first use, such as 'Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq)' rather than just the acronym.

      Corrected.

      (2) Correct author name:

      On page 15, "Lo brutto" appears to be misspelt. Please review and fix throughout the text and the corresponding reference in the bibliography if this is the case.

      Corrected.

      (3) Revise the title to reflect dual contributions:

      The paper delivers two key advances: (a) ATAC-seq atlas plus functional CRE discovery in P. hawaiensis, and (b) innovative low-coverage sequencing for conservation mapping across congenerics. Current title highlights only the first. Consider modifying the title to better reflect both.

      The title reflects our overall objective, without highlighting one of the two approaches (advances) in particular. We would like to keep this concise title and invite readers to read about the two approaches in the abstract.

      (4) Emphasize the combined power of the approach in the Discussion and Conclusions:

      A significant innovation of this manuscript is integrating low-coverage comparative genomics with ATAC-seq to prioritize functional CREs in a large non-model genome. The abstract highlights this well, but the Discussion and Conclusions could better emphasize the power of this pipeline over ATAC-seq alone. The authors could also add 2-3 sentences quantifying cost savings versus traditional assemblies and reiterating how this complements ATAC-seq for efficient CRE prioritization in non-model species.

      We have added the following text in the Discussion to describe complementary contributions of ATAC-seq and sequence conservation to CRE discovery: "Previous efforts to identify cis-regulatory elements in Parhyale relied on reporter constructs carrying a few kb of sequences upstream of selected target genes, an approach that has worked well in animals and plants with relatively small genomes. As presented earlier, however, this approach was often unsuccessful in Parhyale: from tens of reporters tested, only four robust native drivers had so far been identified (refs). The present work adds 5 new drivers to that collection, including ones with ubiquitous, neuron- and muscle-specific activities. This result comes from combining information on genome-wide chromatin accessibility and evolutionary conservation profiles.

      We cannot at this point distinguish the relative contributions of chromatin profiling and sequence conservation to CRE discovery, because these sources of information were not tested separately. At minimum, we can state that (1) ATAC-seq profiles serve to identify robustly the promoters of candidate genes that are active in a particular cellular context (cell type and stage) and (2) coupling this information with sequence conservation narrows down candidate promoters and distant CREs by a factor of 4 to 10, since only a fraction of ATAC-seq peaks show sequence conservation (Figure 3D). This represents a great improvement in our ability to select candidates to test by transgenesis, the most labour-intensive step in the process."

      Further, we have added this text to explain the advantages and cost savings of our low coverage sequencing strategy: "Our strategy of mapping regions of sequence conservation by direct mapping of short sequence reads across species is much more accessible than conventional strategies that rely on genome assembly. The latter require much higher sequence coverage (> 50x) from multiple libraries, long-read sequencing or other scaffolding methods, and complex bioinformatic pipelines to assemble large genomes. Moreover, these approaches are often compromised by high levels of polymorphism found in natural populations. We estimate that our approach is 5- to 10-fold cheaper than assembly-based methods, even excluding labor costs."

      (5) Improve figure readability:

      The authors could improve figure readability by introducing a schematic representation of the specimens and more references in Figures 4, 5 and Supplementary Figure 5. A schematic representation would help non-experts in Parhyale understand what they are looking at. Some figures might also benefit from improved color contrast (e.g., Figure 3 has very similar orange/red colors; black dots on a dark grey background are hard to distinguish).

      We have added additional labels and explanations in the legends of Figures 4, 5 and Supplementary Figure 5, which we think will make the images more intelligible to the readers. In Figure 3 we modified the colouring in panels B and C to improve contrast.

      (6) Quantify reporter validation efficiencies:

      The authors should add a summary table/plot (e.g. n surviving, n fluorescent) or label the figures (n of specimens showing that pattern/n of specimens that do not show the pattern) with exact numbers to explicitly illustrate the observation. For example, Supplementary Table 3 contains excellent data for the putative developmental CREs tested, but the main text lacks equivalent quantification for successful reporters.

      This information is already provided in Table 3.

      (7) Discuss the rapid evolution of developmental CREs:

      The failure to validate developmental CREs using conserved candidates may also reflect the rapid evolution and turnover of developmental enhancers, which can erode detectable sequence conservation over these phylogenetic distances. As a result, functionally relevant elements may have been excluded during candidate selection. It may be worth discussing this possibility alongside the proposed long-range and shadow enhancers hypothesis.

      We added the phrase “or the rapid evolution of these enhancers leading to low sequence conservation" in the relevant part of the Discussion.

      Reviewer #2 (Public review):

      The manuscript by Forbes, Skafida, Karapidaki et al. concerns the in silico identification of cis-regulatory elements (CREs) in large genomes using chromatin accessibility (ATAC-seq) and sequence conservation (genomic DNA sequencing) data. They exemplify this method by applying it to identify novel CREs in Parhyale hawaiensis, which they validated using reporter constructs.

      The results are convincing and are well supported by the data and validations. Identified CREs are valuable for researchers interested in the regulation of the expression of genes they control.

      The methodology on the whole is also valid, as suggested by the results and previous publications on various taxa. Sequence conservation, as stated by the authors, was long used as a method to identify regions of non-coding DNA with functional and evolutionary constraints. The same applies to ATAC-seq data, which has also been used as a proxy for functional regions in different animals such as sea urchins and amphioxus. The methodology proposed is likely to be successfully used by researchers working on a variety of experimental organisms.

      The authors do not use existing genome assemblies and use short-read sequencing to identify conserved regions, and while it is not conceptually novel, such an approach is becoming more and more viable and useful considering the recent advances in next-generation sequencing technology and the decrease in price of short-read sequencing.

      We thank the reviewer for their comments and valuable feedback.

      Two major weaknesses are:

      (1) The novelty of the approach and its advantages should be more explicitly stated.

      (2) The authors do not discuss in depth the strength of using a combination of two methods rather than either of the two, especially considering that previously known CREs do not overlap with conserved sequences.

      We have added two paragraphs at the start of the Discussion to address the reviewer's comments 1 and 2 more explicitly (see response to reviewer 1, comment 4).

      Previously known CREs do include some conserved sequences, see Suppl. Figure 7.

      Reviewer #2 (Recommendations for the authors):

      In addition to addressing the two above-mentioned weaknesses, the authors should address the following minor issues:

      (1) It is difficult to refer to particular regions of text without line numbers.

      Spelling of ChIPseq is inconsistent in the Introduction.

      Spelling corrected. (Sorry for not including line numbering, we'll try to remember next time.)

      (2) "6.4% of reads from P. aquilina, 4.1% of reads from P. darvishi, and 0.46 % of reads from P. plumicornis could be mapped unambiguously to the P. hawaiensis genome" seems quite low for closely related species. Do the authors expect such low rates?

      Neutral nucleotide substitution rates in multicellular animals are in the order of 1 per site per 100 million years (e.g. https://pubmed.ncbi.nlm.nih.gov/12949132/) or a little lower (e.g. https://pubmed.ncbi.nlm.nih.gov/11792858/, https://pubmed.ncbi.nlm.nih.gov/34049492/). With the evolutionary times separating P. hawaiensis from P. aquilina/darvishi and P. plumicornis estimated at roughly 50 and 180 million years (2x25 and 2x90 million years, respectively), we expect a large fraction of neutrally evolving nucleotides in these genomes to have changed. We performed the read mapping using bowtie2, which requires a ~20 nt long perfect match with the reference sequence. We were therefore not surprised to obtain such low rates of read mapping. In fact, these low mapping rates (long divergence times) are important for islands of sequence conservation to stand out.

      (3) "Of these, 37% are found in introns, 54% in intergenic regions, and 1% overlap with promoters (TSS), marking regions that evolve at a lower rate than surrounding non-coding sequences". The authors explain in the Methods why they omit exons, but in the Results and Discussion, it is not stated. In addition, discussing the conservation with exons would be helpful, and the % in exons should be compared to non-coding regions.

      We have added "Of these, 8% are found in exons, likely reflecting conservation in protein-coding sequences".

      (4) "Two of the 7 reporters we tested, named neuro5 and neuro6," if I understood correctly, neuro5 and neuro6 are CREs, however, they are named quite ambiguously, and their names can be mistaken for gene names.

      Indeed, neuro5 and neur6 are the names of the CRE reporters. We have now added the names of the corresponding genes ("carrying CREs associated with the genes αTub and Cdk5α, respectively"). The gene names are also given in Table 3.

      (5) Why was single-end sequencing done for E24?

      We now explain this in the Methods: "Sequencing was carried out on an Illumina NextSeq 500 sequencer; we carried out single-end 76 bp sequencing for the first sample we generated (E24), and then switched to paired-end 76 bp sequencing for the other samples, because this leads to more specific read mapping."

      (6) Syntax related to in-line references should be double checked as the following sentences are broken by parentheses, e.g., "updated in (Almazán et al. 2022))".

      Corrected.

      (7) Could the authors discuss the P. aquilina genome size, which was estimated to be 3-times less than P. hawaiensis? Considering that in their phylogeny these two species are closest, it is quite surprising that they have such differing genome sizes. Do you expect it to be true? If yes, what could be the reason?

      As we explain in the manuscript, our estimates of genome size were obtained by dividing the total number of nucleotides sequenced by the estimated genome coverage, for each species. This method could overestimate genome sizes if there was a significant fraction of contaminating DNA in the preps, or a high degree of sequence variation that would prevent efficient mapping to BUSCO genes (both would underestimate genome coverage), but we find no evidence of this when we estimate the genome size of P. hawaiensis (see manuscript). We used the same method to estimate genome size in all four Parhyale species and have no reason to think that the method would be biased in one species and not in others. We therefore think that we have comparable estimates of genome size for the 4 species and the size difference is real.

      Variations in genome size can be driven by changes in the fraction of repetitive sequences found in a genome. We therefore checked the proportion of repetitive elements in each Parhyale genome using dnaPipeTE (https://github.com/clemgoub/dnaPipeTE). Based on this method (which likely underestimates the repetitive genome content) we find that the genomes of P. hawaiensis, P. aquiline, P. darvishi and P. plumicornis contain 31%, 22%, 18% and 39% of repetitive sequences, respectively. These figures do not fully account for the differences in genome size (particularly since P. darvishi appears to have even fewer repetitive sequences than P. aquiline). We therefore hesitate to add this very preliminary analysis to the manuscript.

      Of note, such rapid change in genome size is not unprecedented: in fruit flies genome size can vary more than 3-fold in species that have diverged over about 30 million years (https://elifesciences.org/articles/66405).

      (8) Wording "and found a genome coverage of 5.8x, corresponding to a genome size of 3.0 Gbp instead of 3.6 Gbp" is confusing and unclear as to what the authors exactly did here.

      We modified the sentence: "As a control, we followed the same procedure for P. hawaiensis, for which genome size is known (ref), and found a genome size of 3.0 Gbp instead of 3.6 Gbp (with a genome coverage of 5.8x)."

      (9) In the figures and supplementary figures, the genome browser screenshots should also include tracks of macs2 called peaks (those in narrowPeak format).

      Each ATACseq and sequence conservation track has its own set of peaks; we think that adding more tracks would overcrowd the figures. All the tracks (including called peaks) are provided as genome-browser-readable files in Supplementary Data files 1-3, so readers should be able to explore the data and reconstruct the figure panels without much effort.

      Reviewer #3 (Public review):

      Summary:

      Forbes et al. present a new approach for identifying cis-regulatory elements in large genomes. Using Parhyale hawaiensis, a crustacean with a large genome (~3.6 Gb, comparable in size to the human genome), the authors show that current methods for identifying cis-regulatory elements, effective in smaller genomes, are markedly inefficient in organisms with large genomes. To address this limitation, they combine bulk ATAC-seq and single-cell (sc) ATAC-seq to identify chromatin regions that are either ubiquitously accessible or specifically accessible in particular cell types. They further integrate comparative genomics across multiple Parhyale species (P. hawaiensis, P. aquilina, and P. darvishi), selected at appropriate phylogenetic distances (20-95 million years divergence), to pinpoint conserved open chromatin regions likely under functional constraint.

      Using this strategy, the authors predict a set of ubiquitous and cell-type-specific cis-regulatory elements. Importantly, they validate these predictions using rigorous transgenic reporter assays, convincingly demonstrating that their approach can successfully identify functional regulatory elements where previous methods had failed.

      Strengths:

      The approach introduced by Forbes et al. is conceptually straightforward, efficient, and readily transferable to other organisms. The validation experiments show not only that a substantial proportion of the predicted elements are functional, but also that the method is capable of identifying both ubiquitous and cell-type-specific regulatory elements. Given that the identification of regulatory regions remains a major bottleneck in understanding the molecular mechanisms underlying processes of development and regeneration, this work has the potential to make a significant impact in developmental and regeneration biology, particularly for studies involving non-model organisms with large genomes.

      An additional strength is the demonstration that only the genome of the focal species requires high-quality sequencing and assembly. In contrast, species used solely for comparative analysis can be sequenced at low coverage without assembly, substantially reducing costs and increasing the accessibility of the approach.

      Weaknesses:

      While the method is effective in identifying regulatory elements that are active ubiquitously or in differentiated cell types, it failed in detecting elements associated with developmentally regulated genes. This may be due to trivial reasons, such as a very low level of expression of the selected genes. However, as acknowledged by the authors, it may also indicate inherent challenges in identifying regulatory elements associated with developmentally dynamic gene regulation, compared to those associated with genes expressed in differentiated cell types.

      A second limitation, also acknowledged by the authors, is the absence of chromatin conformation capture data, which would help link distal regulatory elements to their target genes. This limitation may be particularly relevant for developmentally regulated genes, where long-range regulatory interactions may be critical.

      Addressing these limitations will be an important direction for future work. Nonetheless, the approach as presented in this manuscript represents a key contribution that sets the stage for further methodological advances in the identification of cis-regulatory elements in large genomes.

      Reviewer #3 (Recommendations for the authors):

      I have no specific comment for the authors. While in my opinion the study has two limitations (as described in the public review), these are clearly acknowledged and properly discussed in the manuscript.

      The manuscript is extremely well written. It has been a great pleasure to read it. Excellent job!

      Thank you!

    1. Author response:

      The following is the authors’ response to the original reviews.

      Overview of Revisions

      We thank the editors and reviewers for their constructive and insightful comments, which have substantially improved the manuscript. We have carefully addressed every point raised. The major revisions include:

      (1) Methods 2.1: completely reorganized to clarify the allotetraploid genome structure of Phragmites australis, the rationale for single-chromosome-anchored microsatellite markers, the maximum distinguishable alleles per ploidy level, and the conservative Ploidies(mydata) <- 4 setting in polysat.

      (2) Methods 2.3 / Results 3.3 / Discussion 4.4: clarified common garden sample sizes, added Cohen's d effect sizes, and acknowledged the correlational nature of the lineage-level comparisons.

      (3) Introduction: added a new paragraph on the eco-evolutionary significance of gene flow in mixed-ploidy systems, and another paragraph emphasizing the novelty of integrating SDMs with physiological and common garden experiments.

      (4) Discussion 4.1 and 4.4: reframed all "polyploidy-driven" language to "polyploidy-associated", explicitly acknowledging that ploidy is confounded with genetic background and was not experimentally manipulated.

      Public Reviews:

      Reviewer #1 (Public review):

      (R1-P1) Inadequate explanation of allele dosage for ploidy levels

      Inadequate explanation of allele dosage for ploidy levels, some of which do not match the allele counts expected for genome copy number.

      We appreciate this comment and have substantially revised Methods 2.1. The key clarifications are:

      (A) Marker specificity: All 42 microsatellite markers were aligned to the P. australis reference genome, and each marker mapped to a single unique chromosome (Table S2). This confirms that each marker amplifies a locus specific to one subgenome only. Therefore, in tetraploids each marker detects at most two alleles (the two homologous copies of that chromosome from one subgenome), while in octoploids (autopolyploid derivatives with four copies of the same chromosome) each marker detects up to four alleles.

      (B) Conservative ploidy setting: We set Ploidies(mydata) <- 4 for all samples in polysat because the exact ploidy of many samples could not be confidently assigned a priori. This uniform treatment is conservative: for actual tetraploids, the two unobserved "copies" are scored as null; for actual octoploids, all four detected alleles are accommodated.

      (C) Dosage estimation: Allele dosage was estimated from high-coverage sequencing read counts (mean >5,000× per locus per sample) using the SSRSeq V1.1 pipeline (Cui et al., 2022), which applies stutter correction, amplification bias correction, and ploidy-optimized dosage calling, not inferred from allele presence/absence alone.

      Methods 2.1, second paragraph onward

      Phragmites australis has a base allotetraploid genome. All 42 microsatellite markers used in this study were aligned to the P. australis reference genome, and each marker mapped to a single unique chromosome (Table S2), confirming that each marker amplifies from one subgenome only. Therefore, in tetraploids each marker detects at most two alleles (the two homologous copies of that chromosome from the target subgenome) while the homologous region from the other subgenome is not amplified (Saltonstall, 2003). In Asia, the prevalent octoploids are most likely autopolyploid derivatives of tetraploids, carrying four homologous copies of the same chromosome and thus capable of up to four distinguishable alleles per locus (Liu et al., 2022; Wang et al., 2024). Hexaploid individuals are rare and occur primarily in contact zones, likely originating from inter‑lineage hybridization (Wang et al., 2024).

      In practice, the 42 selected markers very rarely produced more than four alleles in any single individual (Table S2), consistent with a ploidy ceiling of octoploid. Because the exact ploidy of many samples could not be confidently assigned a priori (ploidy was inferred from a combination of chloroplast haplotype, geographic origin, and flow cytometry from prior studies; Lambertini et al., 2020; Liu et al., 2022), we consistently set Ploidies(mydata) <- 4 in the polysat R package (Clark & Jasieniuk, 2011), treating every individual as having four homologous copies. This uniform treatment is conservative: for an actual tetraploid (two copies per locus), the two unobserved "copies" are simply scored as null (missing data) in the dosage matrix; for an actual octoploid, all four detected alleles are accommodated. The allele dosage itself was estimated from high-coverage sequencing read counts (mean >5,000× per locus per sample) using the SSRSeq V1.1 pipeline (Cui et al., 2022), not inferred from allele counts alone.”

      (R1-P2) Common garden setup and sample sizes unclear

      The setup and sample sizes of the common garden experiments are very unclear. The numbers implied are extremely low to draw robust conclusions.

      We agree that the original description was insufficiently detailed. We have made three modifications:

      (A) Methods 2.3: clarified that each population contributed one rhizome segment planted in one pot (one biological replicate per population per site), and emphasized that the key inference rests on the direction and consistency of differences across four climatically distinct sites rather than on significance at any single site.

      (B) Results 3.3: added Cohen's d effect sizes (verified from the raw data), showing that where lineage differences are present, they are biologically substantial (d = 1.10–1.43 at three of four sites).

      (C) Discussion 4.4: added a paragraph acknowledging the limited number of populations per lineage and the need for future confirmation with a larger panel.

      Methods 2.3

      “(1) A previously published common garden experiment (Song et al., 2021) conducted in 2017 across Jinan (36.43°N, 117.45°E) and Panjin (41.20°N, 122.02°E), using CN (n = 11) and FEAU (n = 9) lineages, for which we determined the haplotype information of all samples; (2) A new common garden experiment established in 2021 across Qingdao (36.36°N, 120.69°E) and Shanghai (30.20°N, 121.29°E), with CN (n = 9) and FEAU (n = 8) lineages (Table S3). Each rhizome segment (2–3 buds per segment, one segment per population) was transplanted into an individual 20 L pot, yielding one biological replicate per population per site. Although the number of populations per lineage is modest, the key inference rests on the direction and consistency of lineage differences across four climatically distinct sites rather than on the statistical significance at any single site.”

      Results 3.3

      “Similarly, plant height was significantly greater in the FEAU lineage than in the CN lineage in Jinan, but not in the other common gardens (Figure 3C). Effect sizes for total biomass were large in Jinan (Cohen's d = 1.10), Panjin (d = 1.12), and Qingdao (d = 1.43), but negligible in Shanghai (d = 0.37), confirming that the lineage differences, where present, are biologically substantial.”

      Discussion 4.4

      “We also acknowledge that the common garden experiments, while replicated across four climatically distinct sites, involved a limited number of populations per lineage (9–11 CN and 8–9 FEAU), which constrains our ability to fully separate lineage-level effects from population-level variation. The consistent direction of biomass differences across three of four sites, supported by large effect sizes, nonetheless provides robust evidence for a lineage-level performance advantage that merits further confirmation with a larger, more geographically representative panel of populations.”

      (R1-P3) How allele dosage is determined

      Unclear how allele dosage is determined. Given it's so central to many analyses, it would be useful to see how this is done rather than use a citation.

      We agree that a self-contained description is warranted. We have rewritten the relevant paragraph in Methods 2.1 to describe the three core steps of the SSRSeq V1.1 pipeline (Cui et al., 2022): (i) stutter correction based on empirically estimated slip ratios; (ii) amplification bias correction across alleles of different repeat lengths; and (iii) ploidy-adjusted allele dosage calling. The pipeline's source code and full documentation are available at https://github.com/ccoo22/SSRseq_count.

      Methods 2.1

      “Microsatellite genotyping was performed using the SSRSeq V1.1 pipeline (Cui et al., 2022; https://github.com/ccoo22/SSRseq_count). Briefly, the pipeline takes the per-locus per-sample read count table generated from high-throughput sequencing and processes it through three core steps fully described in Cui et al. (2022): (i) stutter correction, which reallocates a fraction of reads from each allele to its adjacent repeat class based on empirically estimated slip ratios; (ii) amplification bias correction, which normalizes read counts across alleles of different repeat lengths using locus-specific bias coefficients; and (iii) allele dosage calling, which selects the maximum number of alleles consistent with the specified ploidy (four in this study) and assigns integer dosages (0–4) by comparing corrected read ratios to a ploidy-adjusted threshold optimized to minimize both allelic dropout and false positives. The final output is a genotype matrix with integer allele dosages for all samples and loci, which was used directly as input to the polysat R package for subsequent population genetic analyses.”

      Reviewer #2 (Public review):

      (R2-P1) Polyploidy has no causal evidence; confounded with genetic background

      First, no data support the claims that polyploidy has any causal effect. The ploidy levels are, in fact, completely confounded with other genetic differences, so it is not possible to eliminate genetic variation, independent of ploidy, as the causative factor. As the authors note, ploidy was not manipulated in the reported experiments. Thus, the focus on polyploidy in the introduction and elsewhere distracts from the novel and informative experiments that were conducted.

      We fully acknowledge this critical limitation and thank the reviewer for this important critique. We have revised the manuscript at four locations to reframe all claims from "polyploidy-driven" to "polyploidy-associated" and to explicitly state that ploidy is confounded with lineage identity and was not experimentally manipulated.

      Abstract (last sentence)

      “These results demonstrate that climate change interacts with intraspecific variation among polyploidy-associated lineages, manifested through differences in thermal tolerance, biomass production, and asymmetric gene flow, to drive potential lineage replacement within a native range”

      Introduction (polyploidy paragraph)

      “The octoploid FEAU lineage is distinguished from its tetraploid relatives not only by ploidy level but also by its distinct evolutionary history, genomic background, and geographic origin. Polyploidy has been shown in other systems to generate genetic novelty, alter gene expression, and enhance physiological stress tolerance (Bureš et al., 2024; Cheng et al., 2021; Kolář et al., 2017; Van de Peer et al., 2017), potentially pre-equipping polyploid lineages to occupy new geographical ranges and endure environmental shifts (Cheng et al., 2021; López-Jurado et al., 2019). The FEAU lineage's superior thermal tolerance and biomass are consistent with such polyploidy-associated effects, although ploidy is correlated with, rather than experimentally separable from, the broader genetic identity of each lineage.”

      Discussion 4.1 (title and opening paragraph)

      “Our findings demonstrate that the octoploid FEAU lineage of P. australis possesses greater heat tolerance and biomass production than the tetraploid CN lineage. Under a high emission scenario (SSP5-8.5), the projected suitable habitat for the FEAU lineage expands by 18.6%, while the CN lineage exhibits a much smaller relative increase. Several non-mutually-exclusive mechanisms could explain these lineage-level differences, including increased gene dosage from whole-genome duplication, divergent selection histories, and/or standing genetic variation in thermal tolerance loci unlinked to ploidy (Bures et al., 2024; Cheng et al., 2021; Van de Peer et al., 2017). Our data cannot fully partition these factors, but the strong association between lineage identity and both physiological performance and projected range dynamics highlights the importance of incorporating intraspecific lineage information into ecological forecasts, regardless of the ultimate causal mechanism.”

      Discussion 4.4 (limitations paragraph)

      “Crucially, ploidy was not experimentally manipulated in this study; it is inherently confounded with the distinct evolutionary history and genomic background of each lineage. While the observed thermal tolerance and biomass differences are consistently associated with the octoploid FEAU lineage, we cannot formally exclude the possibility that these traits are driven by genetic factors independent of ploidy per se. Future studies using experimental approaches that can partition ploidy effects from lineage-specific genetic effects, such as common gardens with synthetic polyploids or transcriptomic analyses comparing gene expression dosage responses, are needed to strengthen causal inference (Wei et al., 2020). Similarly, the common garden results should be interpreted as lineage-associated rather than ploidy-causal performance differences. The potential role of admixture in facilitating the adaptive introgression of heat tolerance alleles also warrants deeper investigation (Suarez-Gonzalez et al., 2018).”

      (R2-P2) SDMs treat lineages as homogeneous entities

      Second, the manuscript indicates that intraspecific variation is critical for the evolutionary potential of a species to respond to environmental change, but intraspecific variation is seldom considered in species distribution models... the manuscript performs species distribution modeling on a small number of sub-specific lineages, essentially treating them as homogeneous "species"—thus the analysis commits the same oversimplification that the manuscript highlights, but does so at a finer evolutionary scale than species.

      We acknowledge this important limitation and agree that it deserves explicit discussion. While disaggregating the species into three major genetic lineages is a step forward from species-as-monolith approaches, within-lineage variation in thermal tolerance, growth, and dispersal capacity is plausible given the broad geographic ranges of the CN and FEAU lineages. We have added a new paragraph in Discussion 4.4 to address this point.

      Discussion 4.4 (new paragraph)

      “We also recognize that our SDM approach, while disaggregating the species into three major genetic lineages, still treats each lineage as a homogeneous entity. This simplification parallels—albeit at a finer scale—the species-as-monolith assumption that we critique in the Introduction. Within-lineage variation in thermal tolerance, growth, and dispersal capacity is plausible, particularly given the broad geographic ranges of the CN and FEAU lineages. By modelling each lineage as a uniform group, our projections may overestimate the precision of range forecasts and underestimate the evolutionary potential of standing variation within lineages (Chardon et al., 2020). Future frameworks that incorporate trait variation at multiple hierarchical levels (population, lineage, ploidy) will be necessary to capture both the adaptive potential and the ecological constraints that shape species' responses to climate change.”

      (R2-P3) Asymmetric introgression and thermal tolerance lack context in Introduction and Discussion

      The title suggests that asymmetric introgression and thermal tolerance are the most important findings of the work. However, the introduction contains no explanation of the potential importance of gene flow (other than to say that asymmetric gene flow was suggested by some preliminary analyses), and the discussion offers only a limited explanation of either the potential mechanisms underlying the asymmetric gene flow or its importance for the long-term evolution of the species.

      We agree that the evolutionary significance of asymmetric gene flow was underdeveloped. We have added two substantial new passages:

      (A) Introduction: a new paragraph explaining the dual role of gene flow in climate adaptation: introgression of adaptive alleles vs. asymmetric introgression as a mechanism of gradual lineage replacement. This paragraph explicitly connects genome dosage differences (octoploid vs. tetraploid) to the natural directionality of backcrossing.

      (B) Discussion 4.2: a new paragraph extending the discussion of asymmetric introgression into its long-term evolutionary consequences, including the potential erosion of the CN lineage's genetic distinctiveness and the risk of losing cold-adapted alleles under future climate volatility.

      Introduction (new paragraph)

      “Gene flow between lineages of differing ploidy can play a dual role in climate adaptation. Introgression may introduce adaptive alleles (e.g., heat tolerance loci) into a recipient lineage, facilitating its persistence under warming (Suarez-Gonzalez et al., 2018). Conversely, if introgression is asymmetric, such that one lineage's genome is disproportionately represented in admixed populations, it can drive a gradual but systematic shift in genetic composition within the contact zone—effectively functioning as a mechanism of lineage replacement without requiring complete competitive exclusion (Bartolić et al., 2024; Zohren et al., 2016). In mixed-ploidy systems, genome dosage differences create a natural directionality in backcrossing: hybrids tend to backcross more frequently with the high-ploidy parent (Bartolić et al., 2024). In the present study, we test whether such a bias exists between the octoploid FEAU and tetraploid CN lineages and examine its consequences for future distribution under climate warming.”

      Discussion 4.2 (new paragraph)

      “From an evolutionary standpoint, asymmetric introgression can erode the genetic distinctiveness of the minority lineage (CN) while enriching the majority lineage (FEAU) with alleles that may have been locally adapted in the CN genomic background. This could reduce the species' overall evolutionary potential, even if the FEAU lineage itself thrives—because cold-adapted alleles from the CN lineage, which may be valuable under future climate volatility (including extreme cold events), risk being diluted or lost (Exposito-Alonso et al., 2022). The directionality of introgression is also not fixed; it could shift if environmental conditions alter hybrid fitness or if the demographic balance between lineages changes. Long-term genomic monitoring of the CN–FEAU contact zone will be essential to determine whether the asymmetric gene flow documented here represents a transient phase or a persistent trajectory toward genomic homogenization.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      (RE-1) Explain how ploidy level is inferred

      We invite the authors to clearly explain how the level of ploidy is being inferred (Reviewer #1).

      We agree that the rationale for ploidy assignment and its relationship to allele counts needed greater clarity. This has been addressed by the comprehensive revision of Methods 2.1 described in response to R1-P1 (Part 1, Reviewer #1 Public Reviews). The revised text now presents a complete step-by-step logical chain: (i) P. australis has an allotetraploid base genome; (ii) all 42 markers map to a single unique chromosome in the reference genome, confirming that each marker amplifies from only one subgenome; (iii) tetraploids therefore show at most two distinguishable alleles per locus, while octoploids (autopolyploid derivatives of tetraploids) show up to four; (iv) because many samples lacked independent ploidy confirmation, we uniformly set Ploidies(mydata) <- 4 in polysat as a conservative treatment that accommodates both tetraploids (two observed copies + two null) and octoploids (four observed copies).

      See the full revised text under R1-P1 (Part 1) above.

      (RE-2) Common garden results are correlational

      We note that the findings of the common garden experiment, although interesting, are mostly correlational (not causative) and rely on a relatively small sample size and confound lineage isolation and adaptive differentiation (both Reviewers).

      We fully acknowledge this limitation. Because all octoploids belong to the FEAU lineage and all tetraploids to CN, ploidy and lineage identity are inherently confounded. This has been addressed by the four-part revision described in response to R2-P1 (Part 1, Reviewer #2 Public Reviews). Specifically:

      The Abstract now frames the findings as "polyploidy-associated" rather than "rooted in polyploidy."

      The Introduction now explicitly states that ploidy is correlated with—but not experimentally separable from—the broader genetic identity of each lineage.

      The Discussion 4.1 title was changed to "Polyploidy-associated thermal tolerance" and the opening paragraph now presents multiple non-mutually-exclusive mechanisms rather than asserting a causal role for polyploidy.

      The Discussion 4.4 now includes an expanded limitations paragraph acknowledging that ploidy was not experimentally manipulated and that common garden results should be interpreted as lineage-associated rather than ploidy-causal.

      In addition, the Methods 2.3 and Discussion 4.4 revisions described in response to R1-P2 (Part 1) address the sample size concern by clarifying the experimental design and adding a dedicated acknowledgement of the limited population replication.

      See the full revised text under R2-P1 and R1-P2 (Part 1) above.

      Reviewer #1 (Recommendations for the authors):

      (R1-R1) "Large morphological traits"

      Line 85. Large morphological traits. Does this mean physically large? Or higher values of some trait.

      We agree the original phrasing was ambiguous. We have replaced "large morphological traits" with explicit trait descriptions.

      Introduction

      “The octoploid FEAU lineage exhibits greater shoot height, larger leaf size, and thicker stems (K. Chen et al., 1993; Guo et al., 2025; Liu et al., 2021b, 2026; Yin et al., 2024), along with stronger salt tolerance and higher thermal tolerance”

      (R1-R2) Why 2 allele copies expected for a tetraploid

      Line 119. Not clear why this allele copy number is expected. A tetraploid can have up to 4 unique alleles (e.g., ABCD), not two.

      This comment arises from the same conceptual gap addressed in R1-P1. The key point is that P. australis is an allotetraploid with two subgenomes, and our 42 markers each map to a single unique chromosome (one subgenome). Therefore, the marker only amplifies the two homologous copies from that subgenome, giving at most two distinguishable alleles. A true autotetraploid would indeed show up to four alleles—but that is not the genomic architecture of P. australis. The revised Methods 2.1 (see R1-P1 in Part 1) now explicitly explains this logic.

      Fully addressed by the Methods 2.1 revision in R1-P1.

      (R1-R3) Theoretical expectation of allele number vs. ploidy

      Line 127-129. As above, this is unclear and not what we expect theoretically. If there is a reasonable number of alleles, there should be a maximum of 4 for tets, 6 for hex, and 8 for octs.

      Same point as R1-R2. The reviewer's expectation (4 for tetraploids, 6 for hexaploids, 8 for octoploids) is correct for autopolyploids with markers that amplify all homologous copies. The discrepancy arises because P. australis is an allotetraploid and our markers are single-chromosome-anchored (each amplifying from only one subgenome). The revised Methods 2.1 now clarifies this distinction explicitly.

      Fully addressed by the Methods 2.1 revision in R1-P1.

      (R1-R4) Typo "makers"

      Line 136. Should be 'markers' not 'makers'

      We have performed a full-text search and corrected all instances of "makers" to "markers" in the manuscript.

      Full-text search and replace.

      (R1-R5) Reason for removing markers with >4 alleles

      Line 142. The reason for the removal of more than 4 alleles is not clear. What about hexaploids and octoploids? They can carry 6 or 8 alleles, respectively.

      We agree the original text did not adequately justify this quality-control step. In our study, the maximum expected distinguishable alleles (given the allotetraploid genome and single-chromosome-anchored markers) is two for tetraploids and four for octoploids. The observation of five or more alleles in multiple individuals is therefore diagnostic of multi-locus amplification (the marker amplifying more than one genomic locus), not of high ploidy. This is a quality-control filter, not a ploidy assignment criterion.

      Methods 2.1

      “During genotyping, eleven markers (including four of the five multi-mapping markers) were removed because more than ten samples exhibited more than four alleles per sample at these loci. Because the maximum number of distinguishable alleles expected under our ploidy model is two (tetraploid) to four (octoploid), the observation of five or more alleles in multiple individuals indicates that these markers amplify more than one genomic locus, rendering them unsuitable for dosage-based genotyping. This filtration is a quality-control step, not a ploidy assignment criterion.”

      (R1-R6) Unclear sample sizes in common garden

      Line 240. Unclear sample sizes. If these are the numbers, they are a very low level of replication expected for a common garden experiment.

      Same point as R1-P2 (Public Review). Please see the full response under R1-P2 in Part 1, where we have (A) clarified the experimental design in Methods 2.3, (B) added Cohen's d effect sizes in Results 3.3, and (C) acknowledged the sample size limitation in Discussion 4.4.

      Fully addressed by the three-part revision in R1-P2.

      Reviewer #2 (Recommendations for the authors):

      (R2-A1) De-emphasize polyploidy

      De-emphasize polyploidy, as it's not manipulated in the study and is entirely confounded with the genotypes of the distinct lineages, and the putative links between polyploidy and heat tolerance are circumstantial and lacking in a clear mechanism.

      We agree fully. This has been addressed comprehensively across four locations in the manuscript (Abstract, Introduction, Discussion 4.1, Discussion 4.4). See the full response under R2-P1 (Part 1) for the revised text at each location.

      Fully addressed by the four-part revision in R2-P1.

      (R2-A2) Emphasize the novelty of combining SDMs with experiments

      Emphasize the novelty of combining SDMs with experiments (or, if I'm not up on the literature and they are more common, explain how they have been used to make new insights).

      We appreciate this suggestion and agree that explicitly stating the novelty of our integrative approach strengthens the manuscript. We have added a new paragraph at the end of the Introduction.

      Introduction (end, before "Here, we integrate population genomics…")

      “Studies that combine species distribution models with physiological or common garden experiments remain surprisingly uncommon (but see López-Jurado et al., 2019). Such integration is essential for transforming correlative SDM projections into mechanistically grounded predictions. In the present study, we adopt this integrative approach: common garden experiments directly test growth performance under controlled conditions, heat-tolerance measurements identify the specific physiological thresholds (T<sub>crit</sub>, T<sub>50</sub>) underlying lineage-specific climate responses, and SDMs project how these experimentally documented differences translate into spatial dynamics under future warming. By linking experimental data with spatial forecasting, we move beyond correlative climate matching toward a trait-based understanding of how intraspecific variation shapes species' future distributions.”

      (R2-A3) Elaborate on the importance of gene flow

      Elaborate on the importance of gene flow and the potential connections between gene flow and evolving species (or lineage) geographical limits.

      This has been addressed by the two new paragraphs described under R2-P3 (Part 1)—one in the Introduction on the dual role of gene flow in climate adaptation (introgression of adaptive alleles vs. asymmetric introgression as a mechanism of lineage replacement), and one in Discussion 4.2 on the long-term evolutionary consequences of asymmetric introgression.

      Fully addressed by the two-part revision in R2-P3.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Many thanks to the three reviewers and the editors for their thoughtful comments and careful evaluation of our manuscript. These are fair, consistent and largely expected comments. On behalf of my co-authors, we provide this response to the public reviews to summarize the main issues raised and the corresponding revisions we have made in the revised manuscript.

      (1) The main consistent comment from all three referees was that our single-nucleus RNA-seq data should be further validated. The reviewers differ in the detail of exactly what they think should be validated, but collectively these comments referred to validation of: (1) the identified cell types, (2) pathways inferred from trajectory analysis, (3) differentially expressed genes between plucked and control conditions across the four sampled time points, and/or (4) inferred ligand–receptor pairs from the cell–cell communication analysis.

      We believe that we are on strong footing for some of these points because of extensive work we’ve done in the past in the cichlid fish model.

      In the references cited in the manuscript and highlighted below (References 1, 10, 11, 29, 30, 31), we tally 29 figures with 273 individual figure panels presenting histology, in situ hybridization, and immunohistochemistry featuring genes expressed in cichlid (replacement) teeth. Most of these genes are markers of dental competency and/or indicative of regenerative potential.

      In addition, in multiple of these papers, we use pharmacology to manipulate the role of key pathways (Hh, BMP, Wnt, Notch) in cichlid tooth development and replacement. Validation of cell types in the present study therefore draws on these published data in cichlids (and other vertebrates), as well as on an unbiased comparative approach, SAMap, which identifies homology between cichlid and mouse dental cell types based on shared gene expression.

      In short, experiments to validate cell types and pathways active in cichlid teeth have been published and are referenced herein. We recognized, however, that these references (some of which include Gareth Fraser as an author, when he was a postdoc in my group; for Reviewer 2) were cited primarily in the Introduction, rather than in the Rationale/Methods or Results sections. We have therefore clarified these connections in the revised manuscript (line 173-74).

      We have not validated nor analyzed functionally the ligand-receptor pairs we inferred from cell-cell communication analysis. This work is beyond the scope of the current paper, and we now state more clearly that these inferences represent hypotheses to be tested in future studies, although many of these ligand–receptor pairs have been noted in other tooth-related publications cited in the manuscript.

      (2) The biggest weakness of our manuscript, noted by referees, is that we do not provide serial histology to accompany our snRNA-seq time course after plucking. We previously described this as a limitation in the “Study limitations and future direction” section of the Discussion, but we have now strengthened this discussion. In particular, we more explicitly acknowledge that we do not directly document the histological progression of tissue responses across the plucking time course or the degree of tissue damage caused by the plucking paradigm at each sampled time point.

      In the “study limitations” section, we note both issues 1 and 2 and suggest that a spatial transcriptomics experiment across the timespan of plucking<>recovery would address simultaneously the desire to understand cellular context of plucking and cellular/spatial differences in plucked vs control cell-type gene expression.

      (3) Reviewers also asked about the presence and interpretation of stromal cells in our snRNA-seq data. In response, we re-examined the mesenchymal compartment and added additional analyses to better characterize stromal/mesenchymal populations and their inferred trajectories in the revised manuscript. This includes a revised Figure 4, revised text around Figure 4 and revised Supplementary Figures.

      (4) Multiple (minor) suggestions for clarification in text and figures have been adopted throughout the revised manuscript, figure legends, and supplemental materials.

      Overall, we do not anticipate that further reviewer engagement will be necessary, and we believe that editorial review of the revised manuscript should be sufficient.

      References cited in the manuscript, highlighted here:

      (1) Fraser, G. J. et al. An Ancient Gene Network Is Co-opted for Teeth on Old and New Jaws. PLoS Biol. 7, e1000031 (2009).

      (10) Fraser, G. J., Bloomquist, R. F. & Streelman, J. T. Common developmental pathways link tooth shape to regeneration. Dev. Biol. 377, 399–414 (2013).

      (11) Bloomquist, R. F. et al. Developmental plasticity of epithelial stem cells in tooth and taste bud renewal. Proc. Natl. Acad. Sci. 116, 17858–17866 (2019).

      (29) Streelman, J. T., Webb, J. F., Albertson, R. C. & Kocher, T. D. The cusp of evolution and development: a model of cichlid tooth shape diversity. Evol. Dev. 5, 600–608 (2003).

      (30) Fraser, G. J., Bloomquist, R. F. & Streelman, J. T. A periodic pattern generator for dental diversity. BMC Biol. 6, 32 (2008).

      (31) Bloomquist, R. F. et al. Coevolutionary patterning of teeth and taste buds. Proc. Natl. Acad. Sci. 112, (2015).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors used single-nucleus RNA sequencing (snRNA-seq) to investigate accelerated tooth replacement following tooth plucking in cichlid fish. They analyzed four stages of regeneration using elegant and well-designed approaches to characterize cellular trajectories and interactions within the dental epithelium and mesenchyme during the accelerated replacement process. Their analyses identified cell-type-specific gene expression profiles and intercellular signaling interactions associated with whole-tooth regeneration.

      Strengths:

      This is a highly interesting and thoughtfully executed study that provides compelling and convincing insights into the mechanisms underlying accelerated tooth regeneration.

      Weaknesses:

      The manuscript currently lacks experimental validation of the single-nucleus RNA-seq data.

      We thank Reviewer #1 for the thoughtful and positive assessment of our study, including the recognition that our snRNA-seq time course provides insight into cellular trajectories, cell-type-specific gene expression, and inferred intercellular signaling during accelerated tooth replacement in cichlid fish. We also appreciate the reviewer’s central concern that the manuscript would be strengthened by additional experimental validation of the single-nucleus RNA-seq data.

      As summarized above and discussed in more detail in our point-by-point responses below, we have clarified how the present cell-type annotations and pathway interpretations are supported by extensive prior experimental work in the cichlid tooth model, including histology, in situ hybridization, immunohistochemistry, and pharmacological perturbation of major developmental pathways. We have also added analyses demonstrating reproducibility across biological test subjects and consistency of representative differentially expressed genes between paired plucked and control samples. Finally, we have strengthened the Study Limitations section to more clearly state that future spatial transcriptomic, histological, and functional validation experiments will be important next steps.

      Reviewer #2 (Public review):

      Summary:

      Mubeen and colleagues studied the cellular basis of tooth regeneration in cichlid fish. Using an elegant tooth plunking strategy followed by single-nucleus RNA-sequencing, the authors were hoping to achieve an atlas of cellular and transcriptional changes that occur within and between cells during whole tooth replacement.

      Strengths:

      The major strengths of the methods and results are high novelty in the approach in a vertebrate with continuous tooth replacement, the temporal analysis of analyzing at plucking and three later time points, the thorough and sophisticated analysis of the snRNA-seq data, including the inference of trajectories and signaling events, and the robust signal of transcriptional differences induced by tooth plucking.

      Weaknesses:

      The major weaknesses of the methods and results are no validation of any of the inferred cell types, no functional tests of whether any of the changes in signaling pathways affect the plucking-induced tooth replacement process, and perhaps no clear takeaway message for biologists not necessarily interested in tooth replacement.

      Conclusion:

      The authors achieved their aims of identifying the changes in gene expression and cellular composition that occur during whole tooth replacement accelerated by plucking. Overall, the results support their conclusions, although some slight semantic qualifiers should probably be added (e.g., referring to "cell types" as "putative cell types").

      The work should have a high impact in the field of tooth and organ regeneration, and the novel methodological paradigm established here of accelerating tooth replacement three-fold by plucking has great promise for future follow-up studies to further study this process. The work could also have a strong impact through the computational methods used here to infer trajectories and signaling interactions. Specific pathways, genes, and cell types could be tested in other fish, such as zebrafish, to test function during tooth replacement.

      The work is unique and interdisciplinary, and also has significance by establishing that robust phenotypically plastic accelerations in regeneration rates occur upon tooth removal. There are very few studies like this one that combine genetic and environmental studies of regeneration. The result that three different species of cichlid fish that normally have very different tooth patterns all accelerate tooth replacement threefold upon tooth plucking also has significance in revealing a highly conserved plucking response.

      We thank Reviewer #2 for the careful and constructive evaluation of our manuscript and for highlighting the novelty of the cichlid tooth-plucking paradigm, the temporal design of the snRNA-seq experiment, and the computational analyses used to infer cellular trajectories and signaling interactions during accelerated tooth replacement. We also appreciate the reviewer’s comments regarding validation of inferred cell types and interpretation of signaling pathways.

      In response, we have revised the manuscript to clarify that our cell-type annotations are supported by marker-gene expression, previously published cichlid tooth studies, and an unbiased comparative approach, SAMap, which relates cichlid and mouse dental cell types based on shared gene-expression structure. We have also clarified that inferred ligand-receptor interactions represent computational hypotheses rather than functionally validated mechanisms. In addition, we revised Figure 6 and Figure 7A to improve the readability and interpretation of inferred signaling results, and we edited the relevant text and figure legends to make these results easier to follow. These points are addressed in greater detail in the point-by-point responses below.

      Reviewer #3 (Public review):

      Summary:

      This is an interesting paper. The process of tooth exfoliation and replacement in vertebrates remains an intriguing and fascinating subject of inquiry. As the scientists noted, there are no mammalian models that can be used to examine signaling pathways in real time.

      Strengths:

      This work integrates in vivo and high-resolution transcriptomics. The study confirms previous findings and emphasizes the need for additional research into the processes that drive the restoration of missing teeth for future therapeutic uses.

      Weaknesses:

      I disagree with the use of the phrase "plucking". Instead, the authors use tooth extraction or tooth removal, which is clinically more correct for the procedure they are doing.

      The inspiration for our ‘plucking’ experiment is work done in the hair follicle model (lines 73-74). Because cichlid teeth are so numerous, are very small, and lack dental roots, this is an accurate description of the procedure. We opt to retain the phrasing.

      The title is rather broad and appears to be more appropriate for a review than an original research work. I would advise specifying the species under research and/or the sort of damage model used in the transcriptome analysis.

      We opt to retain the title.

      It's uncertain whether the findings are exclusively based on regeneration. The presence of tooth remnants, as well as unintended harm to surrounding tissues, may have triggered repair mechanisms, thereby biasing the current data. How did the authors handle this issue? The oral cavity was under severe manipulation, increasing the inflammatory stimuli, a situation that does not take place in physiological exfoliation.

      In the revised manuscript, we have more clearly acknowledged that our plucking paradigm may induce tissue damage and repair-associated responses in addition to accelerated tooth replacement. We have strengthened the Study Limitations section to state that we do not directly document the histological progression of tissue responses across the time course or the degree of damage caused by plucking at each sampled time point. One caveat, however, is that bone remodeling and immune response is likely triggered on the ‘control’ side of the jaw also, just not to the same degree as after plucking.

      The authors indicated the use of microCT analysis; however, no such information appears in the main text. In fact, this manuscript lacks anatomical information. It is required to conduct histological examinations of the regenerated teeth at various time points.

      microCT data were included as a Supplemental Figure to demonstrate the dental formulae of our chosen species; but we did not characterize post-plucking recovery using this technique (see above summary and below point-by-point comments).

      Although the current findings confirm previously found and verified signaling pathways, the absence of functional data lends uniqueness to this work.

      In the revised manuscript, we also clarify that, while our transcriptomic analyses identify candidate cell states, pathways, and signaling interactions associated with accelerated replacement, the functional roles of these inferred pathways remain to be verified in future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Points:

      (1) Figure 1 should include representative H&E staining images comparing the left control side and the regenerated side at 7 days post-plucking. This would provide important histological context for the regeneration process and help readers better interpret the molecular findings.

      This would indeed be valuable information, but we did not carry out histology of paired control vs plucked jaws to accompany our pulse-chase and dissections for single-nucleus isolation. This comment is similar to that below about validation of what is happening on plucked vs. control jaw halves and is the first limitation we discuss in the “study limitations and future directions” section (from line 538).

      (2) Each tooth position consists of a functional tooth, a replacement tooth, and the dental (successional) lamina. On the control side, the successional lamina contains teeth at different developmental stages, analogous to the mammalian bud, cap, and bell stages. Can the snRNA-seq analysis distinguish among tooth families at different developmental stages, as well as the individual components within a single tooth unit? Clarification of this point would enhance the developmental interpretation of the dataset.

      No, our approach does not distinguish among teeth at different stages, nor among teeth in even vs odd positions that tend to be synchronized in replacement cycles. Theoretically, one could do this by dissecting individual teeth and pooling by tooth stage, but we did not.

      (3) Identifying successional lamina cells is critical, and the authors report putative SL cells within the VEE cluster. However, the stromal cells surrounding the successional lamina are also known to play important roles in tooth regeneration. Can the authors further annotate and characterize stromal cell populations in the snRNA-seq dataset? Additional analysis of these supporting cells would strengthen the conclusions regarding epithelial-mesenchymal interactions.

      We thank the reviewer for this insightful suggestion. In response, we further characterized mesenchymal subpopulations and included these new analyses in the revised manuscript (updated Figure 4 and Supplemental Figure 8). Specifically, pseudotime and CellRank analyses identified a mesenchymal subpopulation enriched for Twist1, Dnmt1, and Runx2, which we interpret as putative dental ectomesenchyme (DEM) based on the established roles of these genes in odontogenic mesenchymal development and differentiation, as well as their reported expression in mouse and human tooth single-cell transcriptomic studies. Notably, this putative dental ectomesenchymal population resides within the broader dental follicle compartment identified in our dataset (Figure 4A-C).

      To further assess supporting stromal populations, we examined the expression of established stromal marker genes, including Lum, Col6a3, Aspn, and Vegfc. These markers were broadly restricted to mesenchymal populations and showed strong enrichment overlapping the newly identified putative dental ectomesenchymal region (Supplemental Figure 8B). Consistent with these observations, differential expression analysis identified additional DEM-enriched genes that substantially overlap canonical stromal markers, including extracellular matrix-associated genes, supporting a close transcriptional relationship between the putative dental ectomesenchyme and the surrounding stromal mesenchymal compartment (Supplemental Figure 8C). Together, these findings refine the mesenchymal landscape surrounding the putative successional lamina and support the presence of a specialized stromal microenvironment associated with tooth regeneration.

      Consistent with this interpretation, our CellChat analysis identified significantly increased interactions between the dental ectomesenchyme and cycling ameloblast populations on the plucked side at Day 0 (Supplemental Figure 8D). These interactions were enriched for signaling pathways including SEMA4, EPHB, SLIT and SPP1, all of which have established roles in tissue remodeling, extracellular matrix organization, and regenerative processes. (Supplemental Figure 8E). Because Day 0 contained sufficient biological replicates and cell numbers for robust statistical comparison, we focused our interaction analyses on this time point. Collectively, these additional analyses provide a more comprehensive characterization of the stromal compartment and further support the conclusion that a specialized dental ectomesenchymal population actively participates in epithelial-mesenchymal communication during the earliest stages of tooth regeneration. So, in total, Figure 4 was revised, the text on lines 281-309 was revised, and Supplemental Figure 8 was added.

      (4) The manuscript currently lacks experimental validation of the single-nucleus RNA-seq data. The authors should validate the expression of major signature genes using RNAscope or immunostaining, ideally comparing regenerated samples with the left-side control. Such validation would significantly enhance the robustness of the conclusions.

      We did not validate up- or down-regulation of differentially expressed genes in intact tissue, owing in part to (1) the complexity of this experiment, (2) the fact that the majority of DEGs, or ‘major signature genes’ have been observed to be expressed in dentitions generally, and often by us in previous work on cichlid teeth, and the fact that (3) independent biological replicates were strongly consistent in the direction of effects (see below). In the “study limitations” section, we note this issue and suggest that a spatial transcriptomics experiment across the timespan of plucking<>recovery would address simultaneously the desire to understand cellular context of plucking and cellular/spatial differences in plucked vs control cell-type gene expression.

      Minor Points

      (1) In Figure 1, the color scheme used in the schematic drawing (Figure 1A) should match the corresponding structures shown in Figure 1B to improve clarity and consistency.

      We appreciate the reviewer’s thoughtful suggestion regarding the color consistency between the schematic (Figure 1A) and the fluorescence images (Figure 1B). However, the color scheme in the schematic (Figure 1A) was intentionally selected to maximize accessibility, particularly for readers with color vision deficiencies, and therefore differs from the magenta and green fluorescence channels used in Figure 1B. In the fluorescence images, the magenta and green colors reflect the native display colors used for the Alizarin Red and Calcein labeling channels in the pulse-chase experiment. Directly matching the schematic colors to the fluorescence images could reduce the visual contrast between key anatomical structures and compromise accessibility for some readers. We have therefore retained the current color scheme in Figure 1A while ensuring that the corresponding structures are clearly identified through consistent labels and annotations across both panels.

      (2) The abbreviation for successional lamina (SL) should be defined upon first use in the Introduction.

      We thank the reviewer for catching this omission. We have now defined the abbreviation “successional lamina (SL)” upon its first appearance in the Introduction.

      (3) Regarding biological replicates, the authors should provide data demonstrating the consistency and reproducibility across replicated samples.

      We thank the reviewer for this suggestion. To demonstrate the consistency and reproducibility across biological test subjects, we have added analyses summarizing sequencing quality metrics, test subject contributions, integrated clustering, and representative differential gene expression across individual samples (see Figure S4, panels C, D & E and Author response image 1). Panel A shows that nuclei from different biological test subjects are well integrated across clusters rather than segregating by sample origin. Finally, Panel B presents representative differentially expressed genes from multiple cell populations, demonstrating consistent expression differences between paired plucked and control samples across biological test subjects.

      Author response image 1.

      (A) UMAP embedding of dental nuclei. Each point represents a single nucleus, colored by test subject. (B) Representative differentially expressed genes show consistent expression differences between plucked and control samples across biological replicates. Paired boxplots of average gene expression for representative differentially expressed genes from multiple cell populations at Days 0, 1, 3, and 7. Each point represents one biological replicate (test subject), with paired plucked and control samples connected by dashed lines. The y-axis shows average gene expression, and the x-axis indicates the experimental condition. These representative examples illustrate the consistent direction of differential expression across biological replicates, supporting the reproducibility of the single-nucleus RNA-seq dataset.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 1: Can the panels to the right of panel B be labeled? It's not clear what these six images are showing, so giving them letters and explaining briefly in the legend what the point of each panel is would clarify. "Right, example of individually classified teeth" - can the authors elaborate on what each tooth is an example of (i.e., how each tooth shown was classified"?) For clarity, the graphs in panels C and D should have the y-axes labeled

      We thank the reviewer for this helpful suggestion. In response, we revised the Figure 1B legend to clarify the classification criteria used for dye incorporation analyses and to better describe the representative fluorescence images. Specifically, teeth positive for both Alizarin and Calcein were classified as pre-existing old teeth, whereas teeth positive only for Calcein were classified as newly formed teeth. We additionally clarified that the images to the right of panel B show representative individually classified teeth, with the top row representing pre-existing old teeth and the bottom row representing newly formed teeth. We also added y-axis labels to panels C and D to improve figure clarity and readability.

      (2) Figure 2 legend: should "the cell type" instead be "the putative cell type"? Without validation for all cell types, it seems adding some sort of qualifier is in order here. Can the authors comment further on examples of validation from other studies? For example, Gareth Fraser has published numerous studies that show Pitx2 expression marking dental epithelium in different fish, yet none of these older papers are cited.

      Identification and validation of cell types make use of multiple published datasets in cichlids (for markers matched to mouse), as well as an unbiased computational approach (SAMap) that draws homology between cichlid and mouse dental cell types, based on shared global patterns of gene expression. There is perhaps a philosophical debate to be had about the validity of ‘cell types,’ generally, but our data are validated using two methods. We edited the text in lines 167-177 to clarify, including citing references to our own work (these studies include Gareth Fraser as an author, when he was a postdoc with Streelman).

      (3) Figure 6 is extremely complicated. Can any portions of rows or columns in these tables be highlighted in the figure to help the reader follow the proposed signaling interactions highlighted in the text?

      We thank the reviewer for this helpful suggestion. To improve the readability of Figure 6 and better guide readers through the dynamic signaling patterns described in the text, we revised the figure by visually highlighting the key sender-receiver interaction regions discussed in the Results. Specifically, we annotated the interactions involving mesenchymal subpopulations and alveolar bone (OST) signaling toward CYC-AMB at Days 0 and 7, mesenchymal signaling toward NK/T cells at Day 1, and epithelial cross-talk centred around ES-2 at Day 3. These visual annotations allow readers to more readily identify the signaling interactions highlighted in the text and relate them to the corresponding regions of the interaction heatmaps.

      (4) In Figure 7A, what does the black font indicate (if grey is up in control and red is up in plucked)? I'd guess not up in either, which then makes it unclear whether the sets in black are different or why they are being presented.

      We thank the reviewer for pointing out this ambiguity. In Figure 7A, blue and red labels indicate signaling pathways identified by CellChat as condition-specific, with blue representing pathways detected only in the control condition and red representing pathways detected only in the plucked condition. In contrast, pathways shown in black represent signaling pathways detected in both conditions but exhibiting significant differences in inferred communication probability between conditions. Thus, the black labels denote shared signaling pathways whose activity differs significantly between control and plucked samples, rather than pathways unique to either condition. We have revised the figure legend to clarify this distinction and improve interpretability.

      Reviewer #3 (Recommendations for the authors):

      (1) I encourage the authors to offer information on the histological differences between teeth during physiological and accelerated replacement. I'm curious if the eruption's accelerated rate has any effect on the mineralization of those teeth.

      We did not examine the histology of individual teeth, and so can’t comment on differences in mineralization.

      (2) The findings section contains multiple sentences that should be moved under material and techniques.

      We expect the reviewer is referring to paragraph lines 104-114, which was a tricky paragraph to place in the manuscript. In the end, we believe it represents important context necessary to interpret findings (which could be missed if moved to ‘methods’) and so we’ve chosen to keep this paragraph in its place.

      (3) It would be useful to include a table showing sample distribution by experimental design.

      We thank the reviewer for this suggestion. Sample distributions across experimental conditions, time points, biological test subjects, and identified cell populations are already provided in Supplementary Table 1. To improve clarity and accessibility, we have revised the table legend to more explicitly describe the experimental design and sample annotations represented in the table.

      (4) The writers did a nice job with the graphics in Figure 8; however, the schematics in Figure C are difficult to follow and are not adequately discussed anywhere. Please note that this text may be of great interest to the dentistry community, including clinicians, and that a clear and succinct explanation of the schemes at the end would be quite beneficial.

      We thank the reviewer for this helpful suggestion. We have revised the Figure 8 legend to more clearly explain Panel C as a summary schematic of inferred cell–cell communication events associated with accelerated tooth replacement after plucking. The updated legend clarifies that the pathway labels in Figure 8C summarize results directly from Figure 7A: red pathway labels indicate plucked-only signaling events, corresponding to pathways shown as full red bars in Figure 7A, while black pathway labels indicate signaling interactions detected in both plucked and control conditions but showing significant differences in interaction probability between conditions. Panel C also includes a cell-type legend at the bottom to identify the relevant cell populations.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This manuscript provides an important contribution to the field of platelet biogenesis, and the convincing evidence will advance our understanding of signal transduction driving the development of late megakaryopoiesis and platelet reactivity that results in bleeding diathesis. The paper is noteworthy for analyzing two related, either singly or in combination, tyrosine phosphatases in this conditional, stage development gene knockout. Because SHP1 is a negative regulator and SHP2 is an activator, the synergistic effects found in the double knockout were surprising.

      We thank the reviewer for acknowledging the importance and novelty of our findings.

      Public Reviews:

      Reviewer #1 (Public review):

      Barré et al. investigated the role of Shp1 and Shp2 in megakaryocytes (MKs) and platelets by conditional knock-out of Shp1, Shp2, or both under the control of the Gp1ba promoter. Deletion of Shp1 and Shp2 in MKs and platelets was almost complete. The Shp1/Shp2 double knock-out mice displayed macrothrombocytopenia and increased bleeding, whereas the single knock-outs did not show significant defects. Platelet function was aberrant in DKOs, but not in single knock-outs, and so was ligand-induced signaling, particularly Syk phosphorylation.

      Megakaryocyte maturation was impaired in Shp1/Shp2 DKO mice. Ligand-induced signaling was impaired in Shp2 knock-out and DKO. Ex vivo formation of platelets and in vivo maturation of MKs were impaired in DKO mice. Pharmacological inhibitors of Shp1 and Shp2 had largely similar effects as observed in the single knock-outs. The authors conclude that Shp1 and Shp2 have synergistic functions in the MK/platelet lineage, and that Shp2 may be a potential therapeutic target in myeloproliferative neoplasms.

      Strengths:

      The data clearly show effects of the Shp1/Shp2 double knock-out on MKs and platelets.

      Weaknesses:

      There appears to be a discrepancy between the results with the Shp2 single knock-out and the Shp2 inhibitor: the Shp2 knock-out does not affect MKs and platelets, except Erk1/2 signaling, whereas the Shp2 inhibitors appear to affect MK function.

      This work is interesting and may have potential from a therapeutic point of view.

      Pharmacological effects do not always correlate with congenital anomalies arising for genetic defects. The Shp2 allosteric inhibitors used in our study only inhibit catalytically inactive Shp2, whereas targeted deletion of Ptpn11 results in a loss of total Shp2 expression, including catalytic and non-catalytic related functions, with developmental consequences. Further, Gp1ba-Cre+; Shp2fl/fl megakaryocytes express approximately 22% of normal Shp2 level, which likely also contributes to differences observed between pharmacological inhibition and genetic ablation of Shp2.

      We thank the reviewer for recognizing the therapeutic potential of our findings.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Barré et al. investigate the roles of the phosphatases Shp1 and Shp2 in the megakaryocyte and platelet lineage using genetic depletion in mice. By employing Gp1ba-Cre-based models, the study builds on the authors' previous work and addresses some limitations associated with earlier Pf4-Cre approaches. The authors report relatively mild alterations in megakaryocyte and platelet parameters in mice lacking either Shp1 or Shp2 alone, whereas combined deletion of both phosphatases results in macrothrombocytopenia, mild bleeding, and impaired GPVI-dependent platelet aggregation accompanied by reduced Syk phosphorylation. The functional platelet defects are linked to reduced expression of GPVI and integrin α2, while thrombocytopenia is associated with impaired megakaryocyte maturation, reduced ploidy, defective proplatelet formation, and altered TPO-dependent Ras/MAPK signaling. Similar effects on megakaryopoiesis are also observed in vitro following treatment with newly developed Shp2 inhibitors.

      Strengths and Weaknesses:

      The study addresses an important biological question and presents a substantial dataset that could contribute to a better understanding of Shp1 and Shp2 function in platelet biology. However, several aspects of data presentation and interpretation would benefit from additional clarification. In particular, while the authors conclude that single genetic deletion or pharmacological inhibition of Shp1 has a limited impact and that the major phenotypes are specific to combined Shp1/2 deletion or Shp2 inhibition, some of the data suggest more nuanced effects that may warrant further discussion.

      We thank the reviewer for raising this point. The manuscript is being revised accordingly, including highlighting the potential role of Shp1 in megakaryopoiesis and thrombopoiesis under steady-state and stressed conditions, requiring more detailed investigation.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Barré et al utilize the Gp1ba-Cre transgenic mouse model to build upon previous findings in a Pf4-Cre system to investigate the effects of individual and combined Shp1 and Shp2 deletion in megakaryocytes and platelets. They report decreased megakaryocyte maturation, macrothrombocytopenia, and increased bleeding primarily in association with the Shp1/Shp2 double-knockout condition. The authors further show that this phenotype appears to be driven primarily by Shp2 and implicate dysregulation of Mpl signaling and downstream Ras/MAPK pathways, including ERK1/2. Given the key role of these pathways in human diseases such as myeloproliferative neoplasms and the challenges associated with modulating such a central pathway, identification of a specific regulator of Mpl signaling poses intriguing questions for future studies on clinical applicability.

      We thank the reviewer for acknowledging the importance and novelty of our findings.

      Strengths:

      Overall, the experiments combine in vitro, in vivo, and ex vivo approaches and appear to have been carefully designed and carried out, with multiple technical and biological replicates where relevant. The authors make a compelling argument for using the Gp1baCre as opposed to the Pf4-Cre system and demonstrate both the dose- and stagedependent effects of Shp1 and Shp2 on megakaryopoiesis and thrombopoiesis. They find that Shp1 and Shp2 are required in late-stage megakaryocyte maturation and that even low levels of expression compared to baseline are likely sufficient to yield generally normal megakaryocytes. Their findings also lead to specific future directions, such as the mechanism by which Shp1 regulates megakaryopoiesis and thrombopoiesis that is distinct from TPO-mediated signaling.

      Weaknesses:

      While the experiments have been thoughtfully designed and carried out, there is limited background explanation on relatively complex or niche pathways/mechanisms, such as the relationship between P-selectin, CRP, and PAR4p; the interactions between SFK, Syk, GPVI, and CLEC-2; and TPO, MPL, ERK1/2, AKT, and STAT3, which, while likely intuitive to experts in their respective fields, may be less obvious to a reader approaching this manuscript with a global interest in megakaryopoiesis/thrombopoiesis and thus detract from the impact of the findings.

      We thank the reviewer for raising this point. The manuscript is being revised to better explain the rationale and molecular mechanisms linking these pathways and functions.

      With regard to the science itself, some of the conclusions feel premature based on the available data.

      (1) The section "Aberrant ITAM signaling in Shp1- and Shp2-deficient platelets" is challenging to follow for those not well-versed in ITAM signaling and associated pathways, and may take additional outside reading to follow the conclusion that Syk-dependent signaling is modulated downstream of GPVI and CLEC-2 based on lack of change in Src p-Tyr418, especially considering that Src p-Tyr418 was previously introduced as a measure of SFK rather than Syk. In the introduction, Shp1 is specifically mentioned as a negative regulator of the ITAM/Syk/phospholipase pathway. However, in Figure 4Ai and Bi, Syk phosphorylation/activation in Shp1 knockout cells did not appear to be different from Shp2 knockout cells, and is lower than the control, which is surprising for a negative regulator. It is also not clear why, in the section (Figure 4A-B), there is reduced Syk activation in Shp1 and Shp2 single knockout cells upon CLEC2 stimulation (but apparently not with CRP) when there was no difference in response to CLEC2 (but a difference in response to CRP) in the previous section (Figure 3A, C).

      We thank the reviewer for raising these important points. The manuscript is being revised accordingly, including clarifying the roles of SFKs, Shp1 and Shp2 in the ITAM-Syk-PLCγ2 signaling pathway.

      Briefly, SFKs are essential for phosphorylating ITAMs, allowing SH2-dependent docking of Syk. Reduced reactivity of Shp1/2 DKO platelets to CRP and collagen is likely due to downregulation of the ITAM-containing GPVI-FcR γ-chain complex and integrin α2 subunit, and concomitant reduction in Syk phosphorylation.

      However, the marginal albeit significant reduction in Syk phosphorylation downstream of CLEC-2 in Shp1 and Shp2 KO platelets was not determined and was insufficient to impact CLEC-2-mediated platelet aggregation under the conditions tested.

      Differences in the stoichiometry and docking of Syk to phosphorylated GPVI-FcR γ-chain and CLEC-2 likely contribute to the differences in platelet reactivity and Syk phosphorylation downstream of the two receptors in the absence of Shp1 and Shp2.

      (2) In the section "Reduced Tpo signaling in Shp1/2-deficient MKs," only Western blot data for (p)ERK1/2, AKT, and STAT3 are presented before concluding that decreased ERK1/2 activity is a mechanistic explanation for thrombocytopenia seen in the Shp1/2 doubleknockout condition. Such a statement would benefit from additional experiments, such as protein or transcriptional levels of ERK1/2 targets specifically relevant to megakaryopoiesis, such as ETS, FOS, and JUN, to assess the consequences of decreased phosphorylated ERK1/2.

      We thank the reviewers for these constructive comments. Further experiments are being planned to determine the biological and transcriptional consequences of reduced ERK1/2 phosphorylation during megakaryopoiesis and thrombopoiesis.

      (3) Suggesting that "inhibiting Shp2 will not have any bleeding consequence in patients" and that Shp2 may be a therapeutic target in myeloproliferative neoplasms when none of these studies have been carried out in a human model is a bold conclusion. There are no data presented on, for example, whether Shp2 inhibition can help reverse the MPL/JAK/STAT pathway in the setting of gain-of-function mutations specifically associated with myeloproliferative neoplasms.

      This conclusion is being tempered in the revised manuscript. Genetic- and pharmacological-based approaches will be used to establish the therapeutic potential of inhibiting Shp1 and Shp2 in mouse models of MPN, including Jak2 gain-of-function mice. Bleeding and thrombotic complications of inhibiting Shp1 and Shp2 will be explored as part of these studies.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Altogether, we feel that this is an important study for those in the fields of hematology or signal transduction. Your important study characterizes the roles in late megakaryopoiesis and platelet biogenesis of single or combined conditional deletion of two tyrosine phosphatases, Shp1 and Shp2. Strengths include technical advances in single and combined deletions, the somewhat surprising results of synergy between the two phosphatases, focusing on the critical stage of late megakaryopoiesis, and clinical implications in bleeding diathesis.

      Weaknesses are mostly minor, but the numerous points raised by reviewer 3 need to be addressed and typographical errors corrected. Further discussion should include the relevance or dissimilarity in megakaryopoiesis and platelet biogenesis between murine and human blood health and disease. Since SHP1 is a negative regulator and SHP2 is a positive activator, additional discussion about how they coordinate and fine-tune ("nuanced") signal transduction in TPO- or GPVI-induced signaling in an explicitly stated pathway.

      We invite you to respond to the critiques and submit a revised manuscript.

      Sincerely,

      Seth Corey, MD MPH

      We thank the editor for the positive evaluation of our study and for highlighting its relevance to the fields of haematology and signal transduction. We have carefully addressed all comments raised by Reviewer 3 and corrected typographical errors throughout the manuscript.

      As suggested, we expanded the Discussion to better address the relevance of our murine findings to human megakaryopoiesis and platelet biogenesis. While our study relies on mouse models, key components of TPO/MPL signaling and platelet production are conserved between mice and humans, although differences in megakaryocyte maturation dynamics and platelet biology are acknowledged and now discussed.

      We also clarified the coordinated roles of Shp1 and Shp2 in signaling. Although Shp1 generally acts as a negative regulator and Shp2 as a positive mediator of signal transduction, our results suggest that they function in a complementary manner to optimize signaling downstream of TPO/MPL and GPVI pathways, thereby ensuring appropriate regulation of late megakaryopoiesis, platelet production and activation.

      These additional considerations have been incorporated into the revised manuscript to provide a clearer conceptual framework for how Shp1 and Shp2 cooperate to regulate platelet biogenesis.

      Reviewer #1 (Recommendations for the authors):

      (1) The effects of the Shp1/Shp2 DKO are clear, but the effect of the Shp2 single knock-out is less clear on all parameters that were tested. The exception is ERK1/2 phosphorylation, which was reduced in the Shp2 knock-out as well as the Shp1/Shp2 DKO. Why do the authors conclude that Shp2 may be a potential therapeutic target, while the data show that knock-out of Shp1 and Shp2 is required for the observed effects?

      We agree that the most pronounced phenotypes were observed in the Shp1/Shp2 DKO. However, Shp2 single knock-out consistently reduced ERK1/2 phosphorylation, indicating that Shp2 contributes to MPL downstream signaling in megakaryocytes. The absence of a strong phenotype in Shp2 single knock-out may be due to residual Shp2 protein. However, given the established role of the Shp2–ERK pathway in megakaryopoiesis and the observation that pharmacological Shp2 inhibition significantly affected MK ploidy, proplatelet formation, and ERK1/2 phosphorylation, our data support a contribution of Shp2 to these processes and suggest it as a potential therapeutic target.

      (2) Inhibitors of Shp1 and Shp2 had largely similar effects as Shp1 and Shp2 single knock-outs, respectively. The effect of Shp2 knock-out on MK ploidy is not clear, cf. Figure 5Ai (no effect) and Figure 5Aii (reduction, which is not significant), whereas a clear and significant effect was reported for the Shp1/Shp2 DKO. In contrast, in Figure 7Ciii, the Shp2 inhibitors SHP099 and RMC-4550 clearly affect MK ploidy and the percentage of MKs forming proplatelets. The discrepancy between the effect of Shp2 knock-out and Shp2 inhibitors suggests that the inhibitors may affect other targets. The authors should consider using the Shp2 inhibitors on the Shp2 knock-out to prove or disprove that the effects of the Shp2 inhibitors are mediated exclusively by Shp2.

      Pharmacological inhibition does not necessarily phenocopy genetic deletion. The allosteric Shp2 inhibitors used in our study (SHP099 and RMC-4550) stabilize Shp2 in an inactive conformation and inhibit its catalytic activity, whereas Ptpn11 deletion results in complete loss of the Shp2 protein, including both catalytic and scaffolding functions. These mechanistic differences may lead to distinct biological outcomes and could explain the discrepancy observed between Shp2 knockout and inhibitor treatments.

      (3) Since the most profound effects were found in the Shp1/Shp2 DKO, it would be interesting to use combinations of the Shp1 and Shp2 pharmacological inhibitors to mimic the effect of the Shp1/Shp2 DKO.

      We thank the reviewers for these constructive comments. Further experiments are indeed being planned to use combinations of the Shp1 and Shp2 pharmacological inhibitors to mimic the effect of the Shp1/2 DKO.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Additional details on the strategy used to isolate megakaryocyte progenitors from mouse bone marrow would improve clarity, including sorting approach, gating strategy, and assessment of population purity.

      We thank the reviewer for this suggestion. We have now expanded the Methods section to provide a more detailed description of the strategy used to isolate megakaryocyte progenitors from mouse bone marrow.

      Briefly, bone marrow cells were first enriched for hematopoietic progenitors and stained with antibodies against lineage markers and megakaryocyte-associated markers. Megakaryocyte progenitors were then isolated by flow cytometric sorting based on established surface marker combinations, including c-Kit and CD41 expression. The gating strategy excluded lineage-positive cells and debris before selecting the progenitor population of interest.

      (2) Platelet GPVI expression appears reduced not only in Shp1/2 double-knockout mice but also, to some extent, in single Shp1- or Shp2-deficient models. A more detailed quantitative comparison and discussion would be helpful.

      We thank the reviewer for this observation. Although the most pronounced reduction in GPVI surface expression was observed in Shp1/Shp2 double knock-out platelets, minor variations may appear in the single knock-out models. To address this, we performed additional statistical analyses comparing WT platelets with each single knock-out genotype. These analyses did not reveal any significant statistical differences in GPVI expression between WT and either Shp1- or Shp2-deficient platelets, indicating that the apparent variations fall within the range of biological variability.

      (3) The aggregation traces shown in Figures 3A and 3B would benefit from clarification regarding their representativeness relative to the corresponding quantitative analyses.

      We thank the reviewer for this comment. The aggregation traces in Figures 3A and 3B represent experiments selected from independent replicates included in the quantitative analysis. The figure legends have been revised to clarify that these traces are representative of the experiments summarized in the quantification panels, which include data from multiple independent mice.

      (4) In several experiments, statistical significance may be influenced by differences in sample size across genotypes (e.g., Figures 2Ci, 3Ai, 3Di, and 6Ai). Using comparable numbers of replicates would strengthen the interpretation.

      We appreciate the reviewer’s attention to statistical rigour. The differences in sample size between genotypes reflect the availability of animals from the different breeding cohorts. Importantly, all statistical analyses were performed using appropriate tests that account for unequal sample sizes. The observed differences remain consistent across independent experiments.

      (5) The rationale for assessing only P-selectin exposure following CRP and PAR4p stimulation is not fully explained. Including integrin αIIbβ3 activation, or clarifying its exclusion, would provide a more complete assessment of platelet activation.

      We thank the reviewer for this suggestion. P-selectin exposure was used as a primary readout because it provides a robust measure of α-granule secretion downstream of GPVI and PAR signaling. Integrin αIIbβ3 activation was not assessed in these experiments because platelet aggregation assays were performed in parallel, which already provide a functional readout of integrin activation, as aggregation requires αIIbβ3 engagement. Nonetheless, we agree with the reviewer that direct measurement of integrin activation (e.g., fibrinogen binding) would provide complementary information and will be considered in future studies.

      (6) Figure 3Dii is described as an aggregation assay, although it appears to report P-selectin exposure; this distinction should be clarified.

      We thank the reviewer for identifying this inconsistency. Figure 3Dii reports indeed P-selectin exposure measured by flow cytometry, rather than platelet aggregation. We have corrected the description in the Results section.

      (7) The suggestion of compensatory extramedullary hematopoiesis based on splenomegaly would be strengthened by immunophenotypic analysis of splenic hematopoietic progenitor populations.

      We appreciate this important suggestion. In the current study, the evidence for possible compensatory extramedullary hematopoiesis is mainly based on the splenomegaly observed in Shp1/2 DKO mice. We agree that detailed immunophenotypic analysis of splenic hematopoietic progenitors would provide additional mechanistic insight; however, this was beyond the scope of the present study, which focuses on the intrinsic role of Shp1 and Shp2 in the megakaryocyte and platelet lineage. We have therefore revised the Discussion to present this interpretation more cautiously and to indicate that further studies will be required to determine whether splenic hematopoiesis contributes to compensatory platelet production in this model.

      (8) In Figure S3, differences in platelet recovery kinetics among genotypes appear evident. Clarification of the statistical tests used to assess these differences would be useful.

      We thank the reviewer for this comment. Platelet recovery kinetics were analyzed using two-way ANOVA with appropriate post hoc tests. No statistically significant differences between genotypes were observed. These details have been added to the Methods and figure legend for clarity.

      Reviewer #3 (Recommendations for the authors):

      Overall, the manuscript suffers from multiple typographical and grammatical errors that distract from the data being presented.

      We have carefully revised the manuscript to correct typographical and grammatical errors throughout, improving clarity and readability.

      (1) Figure S1: I believe this should be referenced in the first paragraph of the results section.

      We have now referenced the Supplemental Figure S1 in the first paragraph of the results section as suggested.

      (2) Figure 2A: Although the individual points for the replicates are informative, they do make it difficult to appreciate the SEM, and to my eye it appears that, for example, there may not be a difference between Shp2 and Shp1/2 or that there may be a difference between Shp1 and Shp1/2 in (ii), as Table S2 suggests. In other words, it seems that the increased MPV (as well as the leukocyte phenotype) may be driven by the knockout of Shp2; are there statistical analyses that could be performed to show that the increased MPV is specific to the double knockout?

      We thank the reviewer for this comment. Despite the slightly higher MPV observed in Shp2 single knockouts, statistical analysis using one-way ANOVA, which is appropriate for comparing means across multiple independent groups, and taking all individual data points into account, revealed no significant differences between Shp2 or Shp1 single KO and the Shp1/2 DKO.

      (3) Figure 2Bi: Is this missing a statistical significance bar, or was there no significant difference in cumulative bleeding time between the conditions? If the latter, this should be clarified in the main text (although the specific sentence regarding bleeding time only claims "mildly prolonged," the preceding sentence indicates "significant increase in bleeding").

      Thank you for this comment. There was no statistically significant difference in cumulative bleeding time between the groups. We have now modified the text accordingly to clarify this point and to indicate that, while bleeding time was not significantly different, blood loss was significantly increased in Shp1/2 DKO mice.

      (4) Figure 2Ci: What was the extent (statistically) of GPVI reduction in the Shp1 and Shp2 single knockout mice compared to the control? It seems that although there was no change in alpha2 expression in the single-knockout conditions, the contributions of Shp1 and Shp2 loss may be additive on GPVI (although I acknowledge that this is not necessarily borne out in Figure 3Ai).

      Thank you for this comment. After reanalyzing the data using an appropriate statistical test (one-way ANOVA followed by Tukey’s post hoc test), we found that GPVI expression is significantly reduced in both Shp1 and Shp2 single knockout platelets compared with controls. However, this reduction did not result in detectable functional consequences on platelet aggregation, as shown in Figure 3Ai.

      (5) Figure 3Ai: It seems that the individual replicates for the Shp1/2 double knockout cluster in two populations, extreme non-responders and arguably normal responders to CRP. Are there any biological or technical explanations for this?

      We thank the reviewer for this observation. We agree that the distribution of individual replicates in the Shp1/2 DKO group suggests the presence of two subpopulations, with some samples showing markedly impaired aggregation and others retaining near-normal responsiveness to CRP. While all experiments were performed under standardized conditions, subtle differences in platelet preparation, agonist sensitivity, or assay timing could also contribute to dispersion within this group. Importantly, despite this variability, the overall trend indicates a significant reduction in aggregation in the Shp1/2 DKO condition compared to controls, supporting a critical and partially redundant role for Shp1 and Shp2 in GPVI-mediated platelet activation.

      (6) "Aberrant functional responses of Shp1/2-deficient platelets": It may be helpful, in the last paragraph of this section, to briefly explain the relationship between P-selectin, CRP, and PAR4p. If short on space/words, the introduction likely does not need an explanation of platelet function and definitions of megakaryopoiesis and thrombopoiesis.

      We thank the reviewer for this suggestion. We have revised the last paragraph to clarify that P-selectin surface expression reflects α-granule secretion following platelet activation. We now specify that CRP activates platelets via GPVI signaling, whereas PAR-4 peptide signals through thrombin receptors, providing context for the differential responses observed in Shp1/2-deficient platelets.

      (7) "Aberrant ITAM signaling in Shp1- and Shp2-deficient platelets": Is there a cartoon figure panel that could be added to clarify how SFK (which, as an aside, is not defined as an acronym), Syk, GPVI, CLEC-2 receptor, Shp1, and Shp2 are interrelated? In addition to the comments left in the public review, I was perplexed by Figure 4Bi, as the band for the Shp1/2 double knockout condition appears to be stronger than the other 3 conditions, but this is not what is depicted in the bar graph on the right.

      We thank the reviewer for this helpful comment. We have now added a schematic cartoon (new Figure 8) to clarify the relationships between SFKs, Syk, GPVI, and the regulatory roles of Shp1 and Shp2. All acronyms, including SFK, are now defined at first mention to improve accessibility.

      Regarding Figure 4Bi, we appreciate this observation. The apparent discrepancy between the representative blot and the quantification reflects variability across experiments. The bar graph represents the average of independent replicates.

      (8) I would also recommend considering reshuffling the panels in Figure 4 so that the 2 assays measuring Syk phosphorylation and the 2 assays measuring Src phosphorylation are next to each other, as opposed to grouped by agonist. They should also be presented in the order of the text, which states that SFK activation was measured via Src before mentioning Syk (but the data are presented in reverse).

      We thank the reviewer for this suggestion. We have reorganized Figure 4 so that the panels measuring Src and Syk phosphorylation are presented together and, in the order, described in the text. The manuscript text has also been updated accordingly to match the revised figure layout.

      (9) GPVI overexpression experiments in these megakaryocytes or, conversely, Syk inhibition in control cells, to reverse or recapitulate the phenotype, respectively, may be additionally informative.

      We thank the reviewer for this suggestion. We agree that modulating GPVI or Syk activity could provide additional mechanistic insight. While these experiments were beyond the scope of the current study, we plan to explore GPVI overexpression and Syk inhibition in follow-up studies to further validate the pathway’s role in the observed phenotype.

      (10) "Reduced Tpo signaling in Shp1/2-deficient MKs": In addition to the comments left in the public review, I would suggest moving this section to after "Defective proplatelet formation and MK maturation in Shp1/2-deficient mice" so that the 2 sets of proplatelet and ploidy data are consecutively presented.

      We thank the reviewer for this helpful suggestion. We have now revised the manuscript accordingly by reorganizing both the text and figures. The ploidy and proplatelet formation data are now presented together in Figure 5, followed by the Tpo signaling data in Figure 6, improving the overall flow and clarity of the results section.

      (11) Figure 6Cii: Why does Shp1 add up to >100%?

      The reason the Shp1 bar exceeds 100% is due to how the data were quantified and normalized. Each segment represents the mean from separate experiments. Stacking these means can exceed 100% because the sum of averages is not equal to the average of the total.

      (12) Figure 7D: How do you reconcile these findings of impaired AKT phosphorylation with the addition of a Shp2 inhibitor but no change with Shp2 knockout (Figure 5C)? Would you attribute it to the residual Shp1 and Shp2 in the Cre-Lox MKs?

      Pharmacological effects do not always correlate with congenital anomalies arising for genetic defects. The Shp2 allosteric inhibitors used in our study only inhibit catalytically inactive Shp2, whereas targeted deletion of Ptpn11 results in a loss of total Shp2 expression, including catalytic and non-catalytic related functions, with developmental consequences. Further, Gp1ba-Cre+; Shp2fl/fl megakaryocytes express approximately 22% of normal Shp2 level, which likely also contributes to differences observed between pharmacological inhibition and genetic ablation of Shp2.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study provides valuable evidence that hilar mossy cells play important roles in maintaining the structural organization of the dentate gyrus and regulating the maturation of adult-born granule cells. The evidence for the structural reorganization and for the accelerated dendritic maturation of adult-born granule cells is convincing: it rests on converging anatomical, viral tract-tracing, retroviral birth-dating, and electrophysiological measurements, with appropriate controls for viral spread, off-target CA3 expression, and axonal degeneration. Support for the study's broader interpretive claim - that the dentate circuit functionally compensates for mossy cell loss - is incomplete. That claim rests on two null results obtained under baseline conditions (home-cage cFos and PTZ seizure metrics) in small cohorts, without behavioral assessment and without a stimulus-driven activity readout, and the manuscript does not engage with published work showing that mossy cells regulate neural stem cell activation and are required for stimulus-evoked neurogenic and behavioral responses.

      Thank you and we agree with all points. We specifically focused on the structural aspects of dentate rearrangement, the impact of mossy cell loss/silencing on dentate neurogenesis, and dentate function at the circuit level. Although our home-cage cFos and PTZ susceptibility assays are limited in terms of their sensitivity, these assays were chosen to address the role of mossy cells in controlling overall dentate activity levels and seizure susceptibility, and our data demonstrate no dramatic changes in overall activity levels or increased/decreased seizure susceptibility.

      Strengths:

      (1) The study is technically rigorous and employs multiple complementary approaches, including selective genetic manipulations, viral tracing, immunohistochemistry, retroviral labeling of adult-born neurons, electrophysiology, and anatomical analyses. The comparison between complete mossy cell ablation and chronic synaptic silencing is particularly powerful, allowing the authors to examine the significant role of mossy cells in structural and functional organization in the dentate gyrus.

      (2) One of the most notable findings is the identification of a previously unrecognized collapse of the inner molecular layer following extensive mossy cell ablation. This observation substantially expands current understanding of dentate gyrus structural plasticity. The demonstration that adult-born granule cells undergo accelerated dendritic maturation after both mossy cell loss and silencing also provides important insight into how mossy cells regulate adult neurogenesis.

      We were also surprised by the inner molecular layer (IML) collapse, as disease models that produce mossy cell loss often involve granule cell axon (mossy fiber) sprouting (and maintained IML thickness) rather than IML collapse. It is unclear whether axon sprouting, reduced degrees of mossy cell loss, or other signaling pathways drive the differences between our selective ablation and translational disease models. We agree that the differential effects of mossy cell ablation and silencing on adult neurogenesis and proximal spine formation highlight the remarkable plasticity in this circuit and provide insights into both the functional and structural circuit roles of mossy cell inputs.

      Weaknesses:

      (1) The functional significance of the observed structural remodeling remains incompletely addressed. Mossy cells have been strongly implicated in pattern separation, spatial information, and emotional behavior, yet no behavioral analyses were conducted. Consequently, it remains unclear whether the dramatic anatomical changes observed following mossy cell ablation translate into meaningful behavioral alterations.

      Our functional assays were primarily focused on the circuit (synaptic) level, with additional assessment of how mossy cell manipulations affected overall dentate activity levels (as reflected by cFos expression). Our limited behavioral analysis focused on seizures, based on prior foundational work on the roles of mossy cells in seizures/epilepsy. We tested the hypothesis that seizure susceptibility might be markedly changed in the near absence of mossy cells, using a “threshold” dose of PTZ that is just above that required to produce seizures, and which can produce dramatically enhanced seizures in hyperexcitable mice. Alternative seizure assays (dose-response curves, continuous monitoring, different seizure-inducing protocols), measurements of granule cell activity in response to environmental contingencies, and the assessments of the response of the dentate stem cell pool to neurogenesis-enhancing stimuli might absolutely produce further insights into how functional mossy cell inputs control stimulus-related dentate activation and/or neurogenesis. Our resubmitted manuscript will clarify that the preserved basal level of dentate gyrus activity after mossy cell loss does not preclude altered activity-dependent activation in other settings or behavioral/learning changes. This could be uncovered with additional behavioral testing or seizure modeling, and is something that we expect to address in future studies.

      (2) The conclusion that the dentate gyrus exhibits remarkable homeostatic compensation is reasonable but remains indirect. Although cFos expression and PTZ-induced seizure susceptibility are unchanged despite altered E:I balance, the mechanisms responsible for maintaining network stability are not investigated. Additional analyses of inhibitory circuit remodeling or compensatory synaptic adaptations would strengthen this conclusion.

      We believe that there are many potential mechanisms that could explain how the nearly complete loss of a major population of dentate neurons is not accompanied by dramatic changes in overall activity levels. Although a fully comprehensive functional assessment of dentate circuit elements is prohibitive, we will undertake what we believe to be the highest-yield analyses in this regard. We propose to stain tissue for inhibitory circuit markers such as VGAT and PV, to determine whether mossy cell loss alters the density or localization of inhibitory synapses as well as circuit elements involved in feed-forward inhibition. We also plan to perform additional electrophysiological experiments to directly assay whether changes in feed-forward inhibition, overall synaptic inhibition (sIPSCs) and/or tonic inhibition might accompany functional mossy cell loss. We will incorporate the outcomes from these additional assays into a revised manuscript. This will shed light on whether inhibitory circuit remodeling also contributes to compensation after mossy cell loss, and hopefully provide additional insights relevant to translational disease models that involve mossy cell loss.

      Reviewer #2 (Public review):

      Summary:

      The authors examine how hilar mossy cells (MCs) influence adult-born dentate granule cell (abDGC) maturation and dentate gyrus (DG) structural integrity. Using both MC ablation and chronic functional silencing, they find that lacking MC inputs accelerates early abDGC maturation without altering mature cellular or intrinsic properties. MC silencing specifically decreased inner molecular layer (IML) spine density, whereas MC ablation led to IML collapse and an increased E/I ratio. However, neither intervention altered overall network excitability (measured via c-Fos and seizure induction) or seizure thresholds. These results advance our understanding of DG circuit plasticity during neurodegeneration.

      Strengths:

      (1) The side-by-side comparison of ablation vs. silencing provides a clear distinction between structural synapse loss and functional inactivation.

      (2) The multi-level analysis spanning structural anatomy, single-cell physiology, and network-level assays yields a rich, comprehensive dataset.

      Thank you for these positive assessments of our study.

      Weaknesses:

      (1) Measuring composite E/I ratios without parsing isolated EPSCs and IPSCs limits direct evaluation of MC-driven excitatory inputs. Furthermore, electrical stimulation in the IML likely recruits local interneuron axons directly alongside MC fibers, complicating the attribution of these responses solely to feed-forward MC circuits.

      We fully expect electrical stimulation of the proximal molecular layer to directly recruit local interneuron axons in addition to feed-forward inhibition. Thus, our experimental design did not distinguish between directly stimulated and feed-forward inhibitory circuits, and we were only able to conclude that mossy cell loss caused circuit rearrangement without clearly attributing the differences specifically to feed-forward mechanisms. To provide additional insights into the underlying changes, we plan to examine both inhibitory circuit structure (using immunohistochemistry) and function using assays designed to distinguish between directly stimulated vs. feed-forward inhibitory mechanisms, which will be incorporated into the revised manuscript.

      (2) The dramatic structural reorganization and IML collapse observed following MC ablation make it difficult to attribute changes in the E/I ratio purely to functional synaptic remodeling rather than physical circuit distortion.

      We actually consider physical circuit remodeling after MC ablation to be the primary explanation for the E/I ratio changes, in that the proximal translocation of MEC synapses following mossy cell ablation allows them to be electrically stimulated in the proximal molecular layer. Thus, the altered E/I ratio of proximal synapses after MC ablation largely represents the fact that we are stimulating proximal MEC inputs rather than mossy cell inputs (which are now absent). Our MML terminal stain (VGlut2) and MEC viral labeling support this interpretation, which we will clarify in the results and discussion of this data.

      (3) Layer boundary shifts following MC ablation complicate the interpretation of site-specific spine density (Figure 4); without accounting for IML collapse, classifying spine loss purely by traditional layer boundaries rather than proximal vs. distal dendrites may obscure local structural changes.

      We initially kept the classic nomenclature (IML vs OML) to avoid confusion for readers, and defined “IML” vs “OML” spines based on proximity to the inner and outer edges of the molecular layer (the innermost and outermost 40 µm; see Methods). Thus, in the setting of IML collapse after mossy cell ablation, these spines almost certainly occurred in regions innervated by the MEC (formerly “MML”). To avoid obscuring this aspect of the data, we will clarify this in the Results and Figure 4, making the proximal vs. distal designations clear.

      (4) The convulsive dosing protocol used for the seizure threshold test lacks the sensitivity required to reveal subtle changes in excitability.

      Our PTZ dose (40 mg/kg i.p.) is just above a dose (30 mg/kg i.p.) that almost never causes seizures in healthy mice in our hands, making it potentially able to detect seizure resistance. This 40 mg/kg dose causes short, limited seizures with a relatively consistent latency, and in other (unrelated) experiments, mice with genetic hyperexcitability have dramatically increased seizure duration and accelerated seizure onset (and sometimes mortality) at this dose, indicating that it is sensitive to at least some forms of increased seizure susceptibility. That stated, we agree that this single-dose PTZ protocol could miss subtle changes in dentate excitability or seizure susceptibility. These could be unmasked by a more detailed dose-response analysis or by other induced seizure assays; we will clarify this limitation in our manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study uses the analysis of connectomic and transcriptomic datasets to survey the anatomy and connectivity of neurosecretory cells in the Drosophila brain. While the connectivity analyses are convincing, the anatomical and functional data provided to verify cell type identity and paracrine signaling is incomplete. Once these aspects are improved, this study would be of interest to neuroscientists working on hormonal signaling in Drosophila and other animals.

      We thank the editor and reviewers for their assessment of our manuscript. We hope that the additional results in the revised manuscript addresses all of the concerns.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by McKim et al seeks to provide a comprehensive description of the connectivity of neurosecretory cells (NSCs) using a high-resolution electron microscopy dataset of the fly brain and several single-cell RNA seq transcriptomic datasets from the brain and peripheral tissues of the fly. They use connectomic analyses to identify discrete functional subgroups of NSCs and describe both the broad architecture of the synaptic inputs to these subgroups as well as some of the specific inputs including from chemosensory pathways. They then demonstrate that NSCs have very few traditional presynapses consistent with their known function as providing paracrine release of neuropeptides. Acknowledging that EM datasets can't account for paracrine release, the authors use several scRNAseq datasets to explore signaling between NSCs and characterize widespread patterns of neuropeptide receptor expression across the brain and several body tissues. The thoroughness of this study allows it to largely achieve it's goal and provides a useful resource for anyone studying neurohormonal signaling.

      Strengths:

      The strengths of this study are the thorough nature of the approach and the integration of several large-scale datasets to address short-comings of individual datasets. The study also acknowledges the limitations that are inherent to studying hormonal signaling and provides interpretations within the context of these limitations.

      We thank this reviewer for the thorough assessment and highlighting the strengths of our manuscript. Based on comments from the other reviewer, we now include additional analyses of NSCs from two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome. Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      Weaknesses:

      Overall, the framing of this paper needs to be shifted from statements of what was done to what was found. Each subsection, and the narrative within each, is framed on topics such as "synaptic output pathways from NSC" when there are clear and impactful findings such as "NSCs have sparse synaptic output". Framing the manuscript in this way allows the reader to identify broad takeaways that are applicable to other model system. Otherwise, the manuscript risks being encyclopedic in nature. An overall synthesis of the results would help provide the larger context within which this study falls.

      We agree with the reviewer and have modified the subsection titles to highlight the main findings within those sections.

      We have also included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

      The cartoon schematic in Figure 5A (which is adapted from a 2020 review) has an error. This schematic depicts uniglomerular projection neurons of the antennal lobe projecting directly to the lateral horn (without synapsing in the mushroom bodies) and multiglomerular projection neurons projecting to the mushroom bodies and then lateral horn. This should be reversed (uniglomerular PNs synapse in the calyx and then further project to the LH and multiglomerular PNs project along the mlACT directly to the LH) and is nicely depicted in a Strutz et al 2014 publication in eLife.

      We thank the reviewer for spotting this error. We have now modified the schematic as suggested.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to provide a comprehensive description of the neurosecretory network in the adult Drosophila brain. They sought to assign and verify the types of 80 neurosecretory cells (NSCs) found in the publicly available FlyWire female brain connectome. They then describe the organization of synaptic inputs and outputs across NSC types and outline circuits by which olfaction may regulate NSCs, and by which Corazon-producing NSCs may regulate flight behavior. Leveraging existing transcriptomic data, they also describe the hormone and receptor expressions in the NSCs and suggest putative paracrine signaling between NSCs. Taken together, these analyses provide a framework for future experiments, which may demonstrate whether and how NSCs, and the circuits to which they belong, may shape physiological function or animal behavior.

      Strengths:

      This study uses the FlyWire female brain connectome (Dorkenwald et al. 2023) to assign putative cell types to the 80 neurosecretory cells (NSCs) based on clustering of synaptic connectivity and morphological features. The authors then verify type assignments for selected populations by matching cluster sizes to anatomical localization and cell counts using immunohistochemistry of neuropeptide expression and markers with known co-expression.

      The authors compare their findings to previous work describing the synaptic connectivity of the neurosecretory network in larval Drosophila (Huckesfeld et al., 2021), finding that there are some differences between these developmental stages. Direct comparisons between adults and larvae are made possible through direct comparison in Table 1, as well as the authors' choice to adopt similar (or equivalent) analyses and data visualizations in the present paper's figures.

      The authors extract core themes in NSC synaptic connectivity that speak to their function: different NSC types are downstream of shared presynaptic outputs, suggesting the possibility of joint or coordinated activation, depending on upstream activity. NSCs receive some but not all modalities of sensory input. NSCs have more synaptic inputs than outputs, suggesting they predominantly influence neuronal and whole-body physiology through paracrine and endocrine signaling.

      The authors outline synaptic pathways by which olfactory inputs may influence NSC activity and by which Corazonin-releasing NSCs may regulate flight. These analyses provide a basis for future experiments, which may demonstrate whether and how such circuits shape physiological function or animal behavior.

      The authors extract expression patterns of neuropeptides and receptors across NSC cell types from existing transcriptomic data (Davie et al., 2018) and present the hypothesis that NSCs could be interconnected via paracrine signaling. The authors also catalog hormone receptor expression across tissues, drawing from the Fly Cell Atlas (Li et al., 2022).

      We thank this reviewer for the thorough assessment and for highlighting the strengths of our manuscript. Based on comments from the other reviewer, we now include additional analyses of NSCs from two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome. Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      Weaknesses:

      The clustering of NSCs by their presynaptic inputs and morphological features, along with corroboration with their anatomical locations, distinguished some, but not all cell types. The authors attempt to distinguish cell types using additional methodologies: immunohistochemistry (Figure 2), retrograde trans-synaptic labeling, and characterization of dense core vesicle characteristics in the FlyWire dataset (Figure 1, Supplement 1). However, these corroborating experiments often lacked experimental replicates, were not rigorously quantified, and/or were presented as singular images from individual animals or even individual cells of interest. The assignments of DH44 and DMS types remain particularly unconvincing.

      We thank the reviewer for this comment. We would like to clarify that all immunohistochemical images presented in this manuscript are representative images based on at least 5 independent samples. We have now clarified this in the methods.

      Additionally, we show DH44 > retro-Tango signal across five samples (new Figure 2 Supplement 3) to highlight the consistency of retrograde trans-synaptic labeling. We also show the neurons providing inputs to putative m-NSC<sup>DH44</sup> and putative m-NSC<sup>DMS</sup> in both FAFB and maleCNS connectomes (new Figure 2 Supplement 2B-C). In both the FAFB and maleCNS datasets, we see a group of neurons (marked by black arrows) providing inputs to m-NSC<sup>DMS</sup> but not m-NSC<sup>DH44</sup>. Importantly, these input neurons are not labelled in DH44 > retro-Tango samples, lending further support to our assignment of DH44 and DMS cell types.

      The electron micrographs showing dense core vesicle (DCV) characteristics (new Figure 2 Supplement 2E-G) are also representative images based on examination of multiple neurons. However, we agree with the reviewer that a rigorous quantification would be useful to showcase the differences between DCVs from NSC subtypes. Therefore, we have now performed a quantitative analysis of the DCVs in putative m-NSC<sup>DH44</sup> (n=6), putative m-NSC<sup>DMS</sup> (n=6) and descending neurons (n=2) known to express DMS across three datasets (FlyWire, BANC and maleCNS connectomes). For consistency, we examined the cross section of each cell where the diameter of nuclei was the largest. We quantified the mean gray value of at least 50 DCVs per cell. The individual who performed these analyses was blind to the neuron identity. Our analysis (new Figure 2 Supplement 2H-J) shows that mean gray values of putative m-NSC<sup>DMS</sup> and DMS descending neurons in FAFB and maleCNS are not significantly different, whereas the mean gray values of m-NSC<sup>DH44</sup> are significantly higher. This analysis agrees with our initial DH44 and DMS NSC subtype assignments. Nonetheless, given the similarity in morphology and synaptic connectivity of DH44 and DMS neurons, we have included the limitation on cell type assignment in the absence of molecular markers in the connectome datasets.

      The authors present connectivity diagrams for visualization of putative paracrine signaling between NSCs based on their peptide and receptor expression patterns. These transcriptomic data alone are inadequate for drawing these conclusions, and these connectivity diagrams are untested hypotheses rather than results. The authors do discuss this in the Discussion section.

      We agree with the reviewer that the novel paracrine pathways presented are untested hypotheses. However, there is a very high likelihood that a given NSC subtype can signal to another NSC subtype using a neuropeptide if its receptor is expressed in the target NSC. This is due to the fact that all NSC axons are part of the same nerve bundle (nervi corpora cardiaca) which exits the brain. The axons of different NSCs form release sites that are extremely close to each other. While the release sites in NSCs cannot be visualized in adult Drosophila connectomes (since these regions were not included in the sample prep), these have been mapped in the larvae and shown to be in close proximity (Hückesfeld et al., 2021: https://doi.org/10.7554/eLife.65745). Neuropeptides from these release sites can easily diffuse via the hemolymph to peripheral tissues (e.g. fat body and ovaries) that are much further away from the release sites on neighboring NSCs. We believe that neuropeptide receptors are expressed in NSCs near these release sites where they can receive inputs, not just from the adjacent NSCs, but also from other sources such as the gut enteroendocrine cells. Hence, neuropeptide diffusion is not a limiting factor preventing paracrine signaling between NSCs, and receptor expression is a good indicator for putative paracrine signaling. Consistent with this, several pathways highlighted in the plot (CRZ to CAPA, DH44 to Hugin and Hugin to DH44) have been anatomically and/or functionally validated previously (Zandawala et al., 2021: https://doi.org/10.1371/journal.pgen.1009425; King et al., 2017: https://doi.org/10.1016/j.cub.2017.05.089; Mizuno et al., 2021: https://doi.org/10.1111/dgd.12733). Additionally, a similar analysis was also employed to depict putative interactions between NSCs in larval Drosophila (Hückesfeld et al., 2021). We have now modified the caption for this figure to explicitly state these connections are putative. We hope that the putative pathways presented here will inspire future functional studies, and have also highlighted this outstanding question in the summary Figure 10.

      Reviewer #3 (Public review):

      Summary:

      The manuscript presents an ambitious and comprehensive synaptic connectome of neurosecretory cells (NSC) in the Drosophila brain, which highlights the neural circuits underlying hormonal regulation of physiology and behaviour. The authors use EM-based connectomics, retrograde tracing, and previously characterised single-cell transcriptomic data. The goal was to map the inputs to and outputs from NSCs, revealing novel interactions between sensory, motor, and neurosecretory systems. The results are of great value for the field of neuroendocrinology, with implications for understanding how hormonal signals integrate with brain function to coordinate physiology.

      The manuscript is well-written and provides novel insights into the neurosecretory connectome in the adult Drosophila brain. Some, additional behavioural experiments will significantly strengthen the conclusions.

      Strengths:

      (1) Rigorous anatomical analysis

      (2) Novel insights on the wiring logic of the neurosecretory cells.

      We thank this reviewer for the thorough assessment and highlighting the strengths of our manuscript.

      Weaknesses:

      (1) Functional validation of findings would greatly improve the manuscript.

      We agree with this reviewer that assessing the functional output from NSCs would improve the manuscript. Given that we currently lack genetic tools to measure hormone levels and that behaviors and physiology are modulated by NSCs on slow timescales, it is difficult to assess the immediate functional impact of the sensory inputs to NSC using approaches such as optogenetics. However, since l-NSC<sup>CRZ</sup> are the only known cell type that provide output to descending neurons, we have functionally tested this output pathway using different behavioral assays (new Figure 8 and Supplements). Our analysis identifies a novel role for l-NSC<sup>CRZ</sup> and DNg27 neurons in female reproduction (based on the number of eggs laid).

      Recommendations for the authors:

      Reviewing Editor Comments:

      You will see that the reviewers found your work interesting and valuable, but had some suggestions for how revision could improve the manuscript. A common thread in the reviews is that functional speculations about the extracted circuits and paracrine signaling would benefit from revision, and would fit better in the Discussion, not Results section. Caveats could be more explicitly stated and language asserting functionality could be tempered. The reviewers were unanimous in their desire for a summary diagram or model.

      We thank the editor for these suggestions to improve the manuscript. We have now functionally validated some output pathways from l-NSC<sup>CRZ</sup>. We have also toned down the language regarding functionality where appropriate. Finally, we included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

      Reviewer #2 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      The authors present connectomic analyses for NSCs identified in the FlyWire dataset. All of their connectomic findings would be strengthened by executing these same analyses in the freely available female hemibrain connectome (Scheffer et al. 2020; Plaza et al. 2022), thereby effectively increasing their sample size from one whole brain to three hemispheres. It is unclear why the authors chose only to focus on the FlyWire dataset.

      We thank the reviewer for this suggestion. We had performed a preliminary analysis using the hemibrain dataset. However, out of the 80 endocrine cells that we found in FlyWire, the hemibrain dataset lacks both the NSC subtypes in the SEZ (SEZ-NSC<sup>CAPA</sup> and SEZ-NSC<sup>Hugin</sup>) as well as l-NSC subtypes in the other hemisphere (l-NSC<sup>ITP</sup>, l-NSC<sup>DH31</sup>, l-NSC<sup>CRZ</sup>). In addition, a majority of the input synapses for all NSC are in the SEZ region which allowed us to classify the different NSC subtypes in FlyWire. Since this information is missing in the hemibrain dataset, we are unable to classify the m-NSC into the different subtypes (not shown). Therefore, we cannot perform a comprehensive analysis of input and output pathways of different NSC subtypes using the hemibrain dataset. To address this concern, we have repeated several analyses with two new recent datasets – the brain and nerve cord (BANC) connectome and the male central nervous system (maleCNS) connectome (Table 1, new Figure 1 Supplement 1, new Figure 2 Supplement 2, new Figure 3 Supplement 3, new Figure 6 Supplement 1, new Figure 7 Supplement 3). Our original conclusions based on the FlyWire connectome remain unchanged, further validating our analyses.

      The authors initially map assign NSC types based on anatomical locations and clustering of presynaptic connections and morphological features. Due to matching cell counts and similar soma locations, they find that DMS and DH44 types cannot be easily distinguished. The authors attempt to assign cell types to these two populations using two methods, neither of which are convincing as executed:

      (1) The authors attempt to distinguish the identities of the two populations by anatomically comparing presynaptic inputs in FlyWire to those observed with light microscopy using retrograde trans-synaptic labeling. Due to the lack of a genetic driver line for the DMS population, the authors could complete this only for the DH44 population. The authors present only one animal, at inadequate magnification to see the absence of distinguishing presynaptic neurons. The results would be strengthened by the presentation and quantification of multiple samples; without more than one sample, it is not possible to know how robust this finding is in this genetic driver line. The authors might also consider taking advantage of the widely-used template brain (Bogovic, 2020) to align their light micrographs of presynaptic inputs from the retrograde tracing, with the presynaptic skeletons from FlyWire and compare in a more quantitative and precise manner. The authors might also consider taking a similar approach using anterograde tracing (Talay et al. 2017) to label postsynaptic outputs. Given that postsynaptic outputs are fewer, so long as there are identifiable, distinct postsynaptic partners, it may be easier to distinguish the two populations with anterograde tracing.

      We thank the reviewer for this comment. We would like to clarify that all immunohistochemical images presented in this manuscript are representative images based on at least 5 independent samples. We have now clarified this in the methods.

      Additionally, we show DH44 > retro-Tango signal across five samples (new Figure 2 Supplement 3) to highlight the consistency of retrograde trans-synaptic labeling. We also provide a magnified image in this figure to highlight the absence of presynaptic neurons that distinguish m-NSC<sup>DH44</sup> and m-NSC<sup>DMS</sup>.

      We also show the neurons providing inputs to putative m-NSC<sup>DH44</sup> and putative m-NSC<sup>DMS</sup> in both FAFB and maleCNS connectomes (new Figure 2 Supplement 2B-C). In both the FAFB and maleCNS datasets, we see a group of neurons (marked by black arrows) providing inputs to mNSC<sup>DMS</sup> but not m-NSC<sup>DH44</sup>. Importantly, these input neurons are not labelled in DH44 > retroTango samples, lending further support to our assignment of DH44 and DMS cell types.

      As per this reviewer’s suggestion, we also aligned our retrograde tracing light micrographs to a template brain (Author response image 1). However, we were unable to quantitatively compare neurons in our light micrographs with neuronal skeletons from the connectome. This is because retroTango labels several neurons in the SEZ which obscures morphology of individual neurons needed for such comparisons. Additional experiments, where retro-Tango output is restricted to sparse populations of neurons using a Flp-out strategy, are needed to perform such quantitative analyses. These experiments are beyond the scope of this study since we now provide additional lines of evidence for cell assignments.

      Author response image 1.

      DH44 > retro-Tango presynaptic signal aligned to JRC2018 unisex template brain

      We appreciate the suggestion to use the anterograde tracing tool trans-Tango to distinguish mNSC<sup>DH44</sup> and m-NSC<sup>DMS</sup>. There is very little synaptic output from m-NSC<sup>DMS</sup> and m-NSC<sup>DH44</sup> based on the FlyWire connectome. There is no synaptic output from both of these cell types if we use a threshold of 5 synapses for significant connections (new Figure 7). Using a threshold of 2 synapses for significant synaptic connections, 3 neurons are downstream of m-NSC<sup>DH441</sup> and 5 neurons are downstream of m-NSC<sup>DMS</sup> (not shown). Since these postsynaptic neurons are not bilaterally paired (we do not anticipate unilateral pathways), we don’t think that these connections are significant. Consistent with our analysis with the FlyWire connectome, we did not observe any significant post-synaptic signal with DH44 > trans-Tango (Author response image 2) even using flies raised at 21ºC which increases the synaptic strength during development. Since we do not have a GAL4 driver to specifically target m-NSC<sup>DMS</sup> , we could not perform similar trans-Tango analysis of m-NSC<sup>DMS</sup>.

      Author response image 2.

      DH44 > trans-Tango (left) and w<sup>1118</sup> > trans-Tango (right; control). Presynaptic neurons are labelled in green and post-synaptic neurons are in red. Representative images based on 5 samples.

      Our connectome analyses revealed that putative m-NSC<sup>DMS</sup> receive direct synaptic inputs from enteric neurons but m-NSC<sup>DH44</sup> do not. We used this information to perform another trans-Tango analysis using Gr43a-Gal4 which labels a subpopulation of enteric neurons (Miyamoto and Amrein, 2013: https://doi.org/10.4161/fly.27241) (Author response image 3).

      Author response image 3.

      Initiating trans-Tango from Gr43a neurons (green) does not label any postsynaptic neurons (magenta) in the pars intercerebralis (white arrow head), including those labelled by the DMS antibody (cyan).

      Unfortunately, initiating trans-Tango from Gr43a neurons did not label any post-synaptic neurons in the pars intercerebralis where m-NSC<sup>DH44</sup> and m-NSC<sup>DMS</sup> soma are located. This could be due to a) low trans-Tango sensitivity or b) m-NSC<sup>DMS</sup> are downstream from other enteric neurons not captured by Gr43a-GAL4. In the absence of other broad enteric neuron drivers, we are unable to perform additional analyses.

      (2) The authors attempt to assign cell types by qualitatively assessing the darkness of dense core vesicles in these two populations. However, there is a presentation of only single planar images through three selected cells (a DMS-expressing descending neuron, DMS-expressing NSC, and DH44-expressing NSC) without any quantitative analyses of vesicle characteristics within or across NSC cell types. It is not possible for the reader to assess whether the darker vesicles constitute a real trend, or if these images are hand-selected to support their point. This piece of evidence would be more convincing if the authors demonstrate consistent vesicle characteristics within NSC type and differences across type. Moreover, such analysis of dense core vesicle features in cell types with distinct and known peptide expression would be broadly interesting.

      Given that NSC type assignment is a major contribution of the present paper, it is critical that the authors are clear about the remaining uncertainty in assigning cell types, so as not to propagate false certainty into future work.

      This comment has been addressed above, and we refer the reviewer to the new Figure 2 Supplement 2E-J.

      The authors suggest larger peptide release capacity from CAPA-producing NSCs based on their larger morphological features (Figure 1, Supplement 2), which is more speculative than certain. In Figure 1, Supplement 1 the authors demonstrate the capacity to visualize vesicles number and size in individual NSCs. Rather than speculate over larger peptide release capacity based on cell size, the authors could quantify these vesicle features, which are surely a better indication of peptide release capacities.

      We thank the reviewer for this comment. We agree that number of dense core vesicles within these and other neurons would be a better indicator of their peptide release capacity. We are performing these analyses on a brain-wide scale as part of another project. Therefore, we have removed the following speculative statement from the present manuscript:

      “But given their location, large size, and presumed large release capacity, we speculate that SEZ-NSC<sup>CAPA</sup> participate in global modulation of post-feeding physiology.”

      The authors provide an analysis of NSCs' synaptic inputs and outputs, but never mention whether NSCs are synaptically connected to each other. If connected, it would be very sensible to provide some analysis of synaptic connectivity between NSCs. If they are not connected, the authors should explicitly mention this in the main text, as it is relevant to the overall aim of this study.

      All NSCs are classified as endocrine cells in the FlyWire connectome. Hence, as shown in new Figures 3B and 7B, NSCs do not provide output to any endocrine cells (NSCs) using a threshold of 5 synapses for significant connections. Similarly, we do not see any synaptic connectivity between NSCs in the BANC dataset (new Figure 7 Supplement 3A-B). We do observe sparse connectivity between NSCs in the maleCNS dataset with a threshold of 5 synapses (new Figure 7 Supplement 3C-D), as well as in the Flywire connectome when the threshold is reduced to 2 synapses (new Figure 7 Supplement 2). However, we refrain from emphasizing on these connections because additional validation is required to rule out false positives in synapse predictions. Dense-core vesicles in NSCs can frequently be mistaken for synaptic T-bars during the prediction (unpublished observation).

      Although unlikely, NSCs could also form synapses with each other near their release sites and outside the brain volumes captured in all three datasets examined in this study. This limitation has been included in the discussion.

      There is no substitute for a good circuit wiring diagram; the motifs that are extracted in Figure 3H might be better appreciated if the reader was first presented with a well-formatted complete circuit diagram, which may then foreshadow the points made in the main text and in Figure 3H.

      We appreciate this suggestion. We now include a circuit diagram (new Figure 3G) to highlight the connectivity between NSCs and their presynaptic partners. The proportion plot (old Figure 3G) has now been moved to new Figure 3 Supplement 5.

      The authors provide extensive bar graphs showing synaptic input body IDs in Figure 3 Supplement 2, however they don't complete the same analysis for synaptic outputs (likely due to low numbers). Even so, it would be useful to the reader if the body IDs and cell types for both synaptic inputs and outputs were documented in a supplemental table. Providing such an inventory is aligned with the goals of this study.

      Only l-NSC<sup>unknown</sup> and l-NSC<sup>CRZ</sup> provide synaptic outputs in the FlyWire connectome (new Figure 7B-F). We have included bar graphs showing output from both these cell types at a single-cell level (new Figure 7G). Additionally, we have annotated all the NSC subtypes in FlyWire and BANC on Codex. Further exploring the inputs and outputs of NSC subtypes can be done interactively on Codex. For example, the search command “{upstream_union} cell_type == SEZ_NSC_CAPA” will retrieve all the neurons providing inputs to SEZ-NSC<sup>CAPA</sup>. As a quick search, this is more convenient than pasting individual body IDs from a supplementary table into Codex. All code outputs (csv files) containing this information are also available on Zenodo.

      In describing possible paracrine signaling, the authors write "Given the proximity of NSC axon terminations, it is extremely likely that a hormone released from a given NSC will influence the activity of other NSC types if its receptor is expressed in those cells." In the absence of functional experiments and/or information about spatial localization and/or peptide diffusion and the proximity of receptors to release sites, the expression patterns alone are insufficient to support this conclusion. Thus, the authors might consider removing the circular connectivity plots in Figure 7C, and Figure 7, Supplement 1A-H, and instead emphasize what can be concluded with certainty from the transcriptomic data (which are expression patterns of the hormones and receptors across NSCs and other tissue types). The authors might instead speculate over potential paracrine signaling between NSCs in the Discussion. Given that the authors describe paracrine signaling between NSCs as 'putative' in the abstract and main text, the authors will likely agree the legend for Figure 7 is misleading.

      This comment has been addressed above. We agree with the reviewer and have modified the figure legends (new Figure 9 and Figure 9 Supplement 1) to emphasize that the connections are putative.

      Should the authors keep these connectivity diagrams, it is important to reconsider their threshold wherein 50% of cells in a cluster must express a given hormone for it to be considered present in their analysis. It is entirely conceivable that there is real heterogeneity in hormone expression within the cluster, so it is surprising the authors have applied this artificial criterion.

      We thank the reviewer for presenting us with this option.

      We also apologize for the oversight in explaining our thresholding carefully. To minimize false positives, neuropeptides were subjected to a two-step filtering process. First, only those expressed in at least 50% of the cells within a given cluster were retained. Second, a composite expression score was calculated for each neuropeptide by multiplying its average expression by its percent detection. These values were normalized to the maximum observed signal across the dataset, and only hormone-cluster pairs maintaining a relative score of 0.25 or higher were included in the final analysis. This stringent filtering approach was implemented to focus the analysis on dominant neuropeptides and to exclude contamination from ambient RNA, which is common for neuropeptides (Allen et al., 2020: https://doi.org/10.7554/eLife.54074). To account for lower abundance of receptor transcripts, we used a more permissive threshold for receptors by retaining those expressed in at least 5% of the cells within a cluster. Unlike the neuropeptides, no secondary relative-score filtering was applied to the receptors to ensure that biologically relevant signaling targets were not prematurely excluded due to low transcript density. We have now revised our methods to explain these details.

      Importantly, we used this thresholding criteria to align previous anatomical studies with our single-cell expression analysis and filter out neuropeptides that likely represent contamination: 1) Transcript for leucokinin (Lk) neuropeptide is expressed in l-NSC<sup>ITP</sup> (but previous studies have not been able to detect this peptide in l-NSC<sup>ITP</sup> (Zandawala et al., 2018: https://doi.org/10.1371/journal.pgen.1007767). 2) Hugin is not expressed in m-NSC<sup>DMS</sup> (Oh et al., 2021: https://doi.org/10.1016/j.neuron.2021.04.028). 3) Ilp2 is not expressed in SEZNSC<sup>CAPA</sup> and m-NSC<sup>DMS</sup> (this study and various others). 4) ITP is not expressed in any m-NSC (Gera et al., 2025: https://doi.org/10.7554/eLife.97043.3). Based on these and other examples, we feel that our stringent criteria recover putative pathways that are likely functional while filtering obvious false positive. Nonetheless, we agree with the reviewer that some of these NSCs could represent heterogeneous clusters as has been shown recently for m-NSC<sup>DILP</sup> (Held, Bisen, Zandawala et al., 2025: https://doi.org/10.7554/eLife.99548.3). This heterogeneity could result in some authentic connections to drop out. However, our goal for this analysis was to not identify all the putative paracrine connections, but rather the strongest ones with the hope that it can inspire future functional studies.

      The data shown in Figure 2 would be easier to interpret and therefore more convincing with better use of insets, appropriate overlays of multiple markers, higher image magnifications, and quantification across samples. Specifically: In Figure 2C, authors show mCherry expression but it isn't clear where these cell bodies are located with respect to the image in Figure 2B. This is also true for Figure 2D. Insets in Figure 2B that correspond with regions shown in 2C and 2D would be helpful. In Figure 2, the authors do not provide cell counts across samples for all markers. Thus, it isn't clear how consistent these cell counts are across samples. In Figure 2E, authors claim that no Gr64f-positive cells innervate the NCC, yet there is clearly a GFP signal in the NCC region in the merged image. The authors should provide an additional marker or a higher magnification image to convince the reader that these projections are not in the NCC region.

      We thank the reviewer for these suggestions. To improve clarity, we have made the following changes:

      Figures 2C and 2D are based on different samples than the one shown in Figure 2B. But we have added dashed boxes in Figure 2B to indicate the regions shown in Figure 2C and 2D.

      Included sample sizes in Figure 2A and 2C. The rest of the panels are representative images based on at least 5 samples. This has been included in the methods.

      The cell count has been provided for m-NSC<sup>DILP</sup> for both markers in Figure 2A. The cell counts for m-NSC<sup>DMS</sup>, labelled using mCherry alone, has been provided in Figure 2C. We did not perform cell counting when using the membrane GFP reporter as it is difficult to accurately count overlapping cells (see dashed box labelled C in Figure 2B).

      We also provide a new supplementary file (new Figure 2 Supplement 1) showing cell counts for m-NSC<sup>DILP</sup> using different markers. Based on this, we can confidently conclude that adult Drosophila typically have more than 14 m-NSC<sup>DILP</sup>.

      We have corrected a typo in our label for Figure 2E: it should be Gr64a instead of Gr64f.

      We have modified the Figure 2E inset to show that the four pairs of Gr64a > myrGFP expressing corazonin cells do not project via the NCC (labelled with an arrow). We have also identified the four pairs of Gr64a neurons (Author response image 4 left panel) in the FlyWire connectome, which shows that these neurons do not exit the brain via the NCC.

      Author response image 4.

      Corazonin-expressing Gr64a neurons in the FlyWire connectome (left) and a light micrograph (right, same as in Figure 2E) showing Gr64a neurons (green) and corazonin neurons (magenta).

      Recommendations for improving the writing and presentation:

      Throughout the paper, the authors provide scant or, at times, no citations. Inadequate citation is as much an issue in the introduction as it is in the results and discussion sections. As such, the authors often do not provide a well-supported premise for the present work and/or do not place their findings and interpretations into the context of existing literature. Related, there is a predominance of references to the work of the authors themselves, often in place of citing earlier foundational work. Citations are nearly exclusive to the Drosophila literature, with the exception of the second paragraph of the introduction. This paper would be greatly improved with references to a broader literature.

      We have now added additional references to give credit to foundational work where appropriate. We have also included citations to non-Drosophila literature for more general statements in the introduction and discussion; however, we refrain from citing such studies in the results section to keep it focused.

      Figure 1 Supplement 1 is referenced after Figure 2 in the text. The authors might consider reassigning it as a supplement to Figure 2, which also uses imaging methodologies to distinguish NSC cell types.

      We agree with the reviewer and have reassigned the figures accordingly.

      Figure 3G is difficult to interpret, and its figure legend is brief and inadequate.

      As suggested by this reviewer, we have replaced this panel with a circuit diagram. The proportion plot (old Figure 3G) has now been moved to Figure 3 Supplement 5, and we have expanded the figure legend.

      The bar graphs in Figure 3 Supplements 2 and 3 would best benefit the reader if the x-axis labels are not simply body IDs, but also cell types or instances (if assigned in FlyWire).

      We appreciate this suggestion. Cell types are routinely updated on Codex while the root IDs remain static for v783. Therefore, we chose root IDs for these plots as they can be used to query Codex easily and reliably. We now provide all raw data as csv files on Zenodo used to make these plots. This includes cell types and other classifications.

      Minor corrections to the text and figures:

      Table 1 compares the observed numbers of NSC types in adult flies to those in larvae and those expected based on previous literature. The authors should cite the previous studies that support each of the expected or larval numbers, either within the table or in the table legend. It would also be appreciated if the expected numbers were cited in the main text.

      References for NSC numbers in larvae and expected numbers in adults are now included in Table 1.

      In describing the author's approach to analyzing synaptic connectivity by cosine similarity, authors cite their own previous work rather than the foundational study describing this approach or earlier studies that use it.

      We have now also cited Schlegel et al., 2021 (https://doi.org/10.7554/eLife.66018) who used a similar approach in the olfactory system.

      Reviewer #3 (Recommendations for the authors):

      (1) The observation that most gustatory inputs to NSCs are indirect (particularly for feedingrelated NSCs) is very interesting but lacks functional validation. I suggest that the authors conduct behavioural assays where specific sensory inputs are activated or silenced while monitoring outputs from NSCs. This could include optogenetics to stimulate or inhibit sensory neurons, or alternative feeding assays.

      We thank the reviewer for this insightful suggestion. We agree that the functional validation of gustatory-to-NSC pathways is a highly compelling direction for future research. However, we believe that behavioral assays, as suggested, pose significant interpretive challenges for the following reasons:

      NSCs primarily function by releasing hormones into the systemic circulation. Unlike classical neurotransmission, hormonal modulation typically operates on much slower timescales (minutes to hours). Consequently, acute activation of sensory inputs is unlikely to elicit immediate, quantifiable behavioral changes that can be specifically attributed to NSC activity.

      Most NSC classes are known to influence multiple physiological and behavioral processes simultaneously. Attributing a specific behavioral phenotype to a single NSC class following sensory stimulation would be confounded by these overlapping roles.

      Activating or silencing taste neurons will directly impact feeding behavior through canonical motor circuits, independent of the neuroendocrine system. In such a paradigm, it would be nearly impossible to isolate the specific "indirect" contribution of the NSCs to the observed behavior.

      While we agree that functional connectivity, such as optogenetic activation of taste neurons paired with calcium imaging (e.g., GCaMP) in NSCs, would be the ideal way to validate these inputs, we consider these extensive physiological experiments to be beyond the scope of this anatomical and connectomic study.

      Nonetheless, to address this important question, we now use a recently developed approach (Bates et al., 2026: https://doi.org/10.1101/2025.07.31.667571) based on linear dynamical modeling to estimate the influence of various sensory neurons (gustatory, olfactory, enteric, hygrosensory, etc.) on different NSC classes. Our analysis (new Figure 6 and Figure 6 Supplement 1) reveals that contents of consumed food (detected by enteric neurons) have a stronger influence on NSCs compared to inputs from external taste receptors.

      (2) Descending neurons appear to play a crucial role in regulating both motor and endocrine output. However, their functional contribution is only inferred from the connectomic data. The authors could perform functional activity manipulations (silencing or activating) of these descending neurons (for instance dMS descending neurons) to explore their role in behaviour. This could be tested with simple behavioural assays such as feeding or reproduction (i.e egg laying).

      We believe that there might be some confusion. DMS descending neurons (DNp32 cell type) used for dense-core vesicle comparisons with m-NSC<sup>DMS</sup> and m-NSC<sup>DH44</sup> (new Figure 2 Supplement 2) are different from the descending neurons (DNg27 cell type) that receive inputs from l-NSC<sup>CRZ</sup> (new Figure 7). We have indicated the cell type of DMS descending neurons in the text to clarify this. We have also functionally tested DNg27 (instead of DMS descending neurons suggested by the reviewer) using optogenetic and chemogenetic approaches for effects in feeding, food preference, starvation survival, egg laying and flight (new Figure 8 and Figure 8 Supplement 1). While we expected DNg27 to influence flight based on our connectome analyses, we do not see any phenotype in our free flight setup following DNg27 activation (new Figure 8 Supplement 1). However, this could be due to the split GAL4 driver used being very weak (new Figure 8 Supplement 2). This is also supported by the egg-laying assay where DNg27 inactivation only produces a phenotype after day 8 (Figure 8). Since we currently don’t have access to another driver to specifically target DNg27, we are unable to validate our results in the free flight setup using an independent driver.

      (3) The authors describe a sparse olfactory input pathway to NSCs, with emphasis on odours playing major roles. However, the physiological consequences of these connections are not explored in detail. Authors should use ORN/AL stimulation (e.g., using optogenetics) to explore how odour sensory pathways affect hormonal secretion in NSCs.

      We acknowledge the reviewer’s interest in the physiological consequences of the olfactory to NSC pathways identified in our study. While we agree that exploring how specific odors modulate neuroendocrine output is a logical next step, we believe that such experiments are currently unfeasible due to significant technical and biological constraints as highlighted above for taste neurons. Hence, we calculated the influence of olfactory receptor neuron activation on different NSC classes using an approach based on linear dynamical modeling (new Figure 6 and Figure 6 Supplement 1). Our analysis reveals that smell has a weaker influence on NSC compared to taste.

      (4) The authors present a large amount of nice yet complex data, which can be difficult to navigate through and is sometimes hard to follow. Consider adding more schematic diagrams to summarize the key pathways and interactions between NSC types and their inputs/outputs.

      We thank the reviewer for this suggestion. We have now included a figure (new Figure 10) which summarizes the main findings from our manuscript and places them within the larger context of neuroendocrine signaling in adult Drosophila in relation to other studies.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Mast cells have previously been reported to play an important role in bacterial immune defense and act protectively in sepsis. However, many of these findings were based on studies using Kit mutant mice. In this study, the authors conducted a detailed investigation using mast cell-deficient Cpa3 Cre-Master mice. As a result, the authors found that the Cpa3 Cre-Master mice exhibited responses similar to wildtype mice in terms of bacterial immune defense. This suggests that the observed phenotype is not due to mast cell-dependent bacterial immune defense, but rather is associated with dysbiosis of the gut microbiota.

      Strengths:

      Mast cells have long been reported to play an important role in the protective response against sepsis, and their function in infection defense has been demonstrated. However, Kit mutant mice have been reported to exhibit impaired peristalsis, and several mast cell-specific genetically modified mouse lines have since been developed and examined in detail. This study presents an important finding by logically demonstrating that the exacerbation of sepsis in Kit mice is due to alterations in the gut microbiota, and that the phenotype previously thought to be mast cell-dependent was, in fact, not.

      In addition, the experiments were carefully designed using mice with matched genetic backgrounds. These findings underscore the importance of microbiota composition in interpreting immune phenotypes and highlight the need for cohousing controls in mutant mouse studies.

      A major strength of this work is the robustness of the CLP data, generated over eight years by three independent researchers across two institutions with large sample sizes, lending strong support to the conclusions.

      Weaknesses:

      The study assesses only a limited subset of gut bacterial species, leaving the extent to which E. coli expansion contributes to the observed phenotype unclear.

      We now performed 16S rRNA sequencing of cecal samples isolated from Kit<sup>W/Wv</sup> and Cpa3<sup>Cre/+</sup> mice and their respective littermates. Results are display in a new Figure 4. Our comparative analysis of the cecal microbial communities in Kit<sup>W/Wv</sup> and Kit<sup>+/+</sup> mice (Figure 4A+B) confirmed the expansion of E. coli (Enterobacteriaceae) that we had observed by CFU counts (Figure 3D). Furthermore, it revealed a dysbiotic shift marked by increased abundance of Peptostreptococcaceae, Verrucomicrobiaceae, Coriobacteriaceae, and Erysipelotrichaceae in KitW/Wv mice.

      None of these changes was observed when comparing the cecal microbiomes of Cpa3Cre/+ and Cpa3+/+ mice (Figure 4C+D), indicating that the compositional shift in Kit<sup>W/Wv</sup> mice is due to the deficiency in Kit but not mast cells. Of note, as stated on page 14, the microbial changes that we observed in Kit<sup>W/Wv</sup> mice resemble dysbiotic patterns reported in chronic intestinal inflammation, experimental colitis, and impaired barrier function. These new findings fully align with and further support our earlier conclusion that Kit<sup>W/Wv</sup> mice harbour pro-pathogenic microbiota.

      The new results are display in a new Figure 4, and described on pages 9-10 and discussed on pages 13-14.

      Moreover, in the cohousing experiments, there is no evidence provided to confirm successful microbiota normalization between groups.

      It is correct that we have no direct data to confirm microbiota normalization between groups after co-housing. We note, however, that co-housing is a generally accepted method for microbiota equalization or conversion (Caruso et al., Cell Rep. 2019, Ridaura et al., Science 2013, and reviewed in Moore et al., Clin. Transl. Immunol. 2016). In any case, Kit<sup>W/Wv</sup> mutants were made resistant to CLP by co-housing. Similar microbiota sequencing results between groups, while useful, would again only be correlative.

      A more detailed analysis of the microbial composition would be necessary to strengthen the reliability of the findings.

      See above the new data from 16S rRNA sequencing.

      It is also important to note that Cpa3-deficient mice exhibit not only mast cell depletion but also defects in basophils and T cells. These additional immunological alterations may counterbalance one another, potentially masking phenotypic changes and complicating interpretation.

      Regarding basophils in Cpa3<sup>Cre/+</sup> mice, compared to wild-type mice, basophils are reduced to about 40% of normal (Feyerabend et al., Immunity 2011). In Kit<sup>W/Wv</sup> mice, compared to wild-type mice, basophils are reduced to about 10% of normal. To our knowledge, there has been no phenotype reported in which a reduction in basophils compensates for the loss for mast cells. Given that Kit<sup>W/Wv</sup> mice have about threefold lower numbers of basophils and are highly susceptible to sepsis, there is no evidence that a reduction in basophils is protective in mast cell-deficient mice. On the contrary, mice that were normal for mast cells but had their basophils depleted were more susceptible to sepsis (Piliponsky et al., Nat. Immunol. 2019). Hence, basophils appear to be protective, and their reduction increases susceptibility. In light of these data and considerations, there is no evidence for a reduction in basophils to counterbalance the loss of mast cells in Cpa3<sup>Cre/+</sup> mice.

      Regarding T cells, there is no evidence, and there are no reports, that Cpa3<sup>Cre/+</sup> mice have defects in T cells (Feyerabend et al., Immunity 2011, Feyerabend et al., Cell Metabolism 2016). Cpa3 is weakly and transiently expressed early in the T cell lineage (Feyerabend et al., Immunity 2009; for expression levels in T cells versus mast cells, see Author response image 1). In summary, in contrast to the reviewer's claim, there are no known defects in T cell development or T cell functions in Cpa3<sup>Cre/+</sup> mice. We think the reviewer needs to provide published evidence for his/her claim that Cpa3-deficient mice exhibit defects in T cells. We as authors are also obliged to support our claims scientifically, and rightfully so.

      Author response image 1.

      Generated from the Immgen database. Shown are RNAseq gene expression levels of diverse T-cell and mast cell populations.

      Furthermore, it remains to be determined whether the altered gut microbiota observed in KitW/Wv mice is a consequence of impaired intestinal motility, whether a similar phenotype is observed in KitW-sh/W-sh mice, and whether comparable results occur in SCF-deficient models. Addressing these questions would provide greater clarity on the contribution of mast cells versus secondary factors in the observed phenotypes.

      The purpose of our study was to verify or refute the key claim dating back to two 1996 Nature papers that mast cells play important roles against sepsis. We demonstrate here that this is not the case because mice without mast cells (Cpa3<sup>Cre/+</sup> mice) were as resistant to sepsis as wild-type mice. Hence, mast cells are not involved in the immunity against sepsis, and 'secondary factors' are not involved in this simple experiment (both groups of mice, wild-type and Cpa3<sup>Cre/+</sup> mice, were on the identical genetic background). Second, Kit<sup>W/Wv</sup> mice are also as resistant to sepsis as wild-type mice when confronted with the identical intestinal slurry. Therefore, Kit<sup>W/Wv</sup> mice have no immune deficit in response to sepsis. Hence, in our view, the underlying immunological question regarding the role of mast cells in sepsis has been conclusively addressed and answered by our data. We have changed the title to emphasize this central question.

      The reviewer now asks us to delve even deeper into Kit biology and in particular intestinal pathophysiology in this and other Kit or steel mutants. While we share his/her interest in such questions, we fully disagree with the statement that 'addressing these questions would provide greater clarity on the contribution of mast cells versus secondary factors in the observed phenotypes.' We do not intend to enter the field of gut physiology or its link to microbiota, all the more because any results would not affect the central conclusion of our manuscript.

      Given that KitW/Wv mice exhibit impaired peristalsis, is the observed increase in E. coli a consequence of this dysfunction?

      See above

      Previous studies with BMMC reconstitution experiments have indicated that mast cells are a source of TNF - how does this align with the current findings?

      It does not align well. It is possible that cultured and transplanted mast cells (BMMC) produce TNF. Given that we did not find a reduction in TNF levels in the peritoneal lavage or serum in mice without mast cells undergoing sepsis, under physiological conditions mast cell-derived TNF does not seem to have a measurable impact on total TNF levels.

      Reviewer #2 (Public review):

      Summary:

      This study presents a useful finding that the high susceptibility to CLP sepsis of Kitmutant mice is not due to mast cell deficiency, but to dysbiosis.

      However, the present data are insufficient and incomplete to support the conclusion, and would benefit from more rigorous approaches. With the mechanism part strengthened, this paper would be of interest to researchers on mast cell biology and mucosal immunology.

      We disagree with the view that our data are insufficient and incomplete. Our results demonstrate that mice lacking mast cells (Cpa3<sup>Cre/+</sup> mice) are as resistant to sepsis as wild-type mice, demonstrating that mast cells do not play a detectable role in immunity against sepsis. Additionally, we show that Kit<sup>W/Wv</sup> mice exhibit the same resistance to sepsis as wild-type mice when confronted with the identical intestinal slurry. This finding demonstrates that Kit<sup>W/Wv</sup> mice have no immune deficit in response to sepsis. These central data are both sufficient and complete, given that our data fully address the potential role of mast cells in sepsis. Our study aimed to investigate the role of mast cells in sepsis, not to examine the mechanisms of dysbiosis or associated pathological phenotypes in Kit-mutant controls. We have changed the title to make this point.

      Recommendations:

      (1) The authors showed that E. coli increases in the cecum of Kit-mutant mice, which causes high CLP susceptibility. However, they did not provide any evidence E. coli is responsible for the high susceptibility.

      We showed that E. coli CFUs were increased in the cecum of Kit-mutant mice, but we did not state that this causes CLP susceptibility. We wrote: 'Hence, Kit<sup>W/Wv</sup> microbiota contains high levels of E. coli, which may underlie the observed pathogenicity'. We demonstrated that intestinal slurry from Kit<sup>W/Wv</sup> mice is more pathogenic compared to intestinal slurry from wild-type mice. However, we did not search for or identify the bacterial species that causes this increased pathogenicity because we were addressing the role of mast cells in sepsis. We demonstrate an association of pathogenicity in sepsis experiments with cecal content of pathogenic bacteria (see also the new data on 16S rRNA sequencing). The same argument could be made for each bacterial species identified but this would be very complex experiments (both microbiologically and immunologically) given requirements for bacterial isolation, titration, and considerations of synergism. We therefore refrain from this undertaking.

      In the Figure 3 experiments, the authors administered the same number of cecal bacteria and did not show the number of E. coli after the administration.

      The samples were split and one aliquot was analysed by microbiology and the other aliquot was injected intraperitoneally. Fig. 3d shows the colony-forming units (for Lactobacilli and E coli) from aliquots of cecal slurry used in the intraperitoneal injection experiments shown in Fig. 3a-c. Hence, our data show the colony-forming units that were injected into the mice. It is unclear to us why this is not the key information rather than 'the number of E. coli after the administration'.

      The authors should provide evidence showing that depletion of E. coli decreases susceptibility.

      See response to point 1 above.

      (2) The author should provide direct evidence of dysbiosis by, for example, shotgun sequencing of cecal and fecal contents.

      We performed 16S rRNA sequencing of cecal contents and observed a dysbiotic shift towards an increase of Peptostreptococcaceae, Verrucomicrobiaceae, Coriobacteriaceae, Enterobacteriaceae, and Erysipelotrichaceae in Kit<sup>W/Wv</sup> mice compared to Kit<sup>+/+</sup> controls. None of these changes was observed when comparing the cecal microbiomes of Cpa3<sup>Cre/+</sup> and Cpa3<sup>+/+</sup> mice, indicating that the compositional shift in Kit<sup>W/Wv</sup> mice is due to deficiency in Kit but not mast cells. Of note, as stated on page 14, the microbial changes we observed in Kit<sup>W/Wv</sup> mice resemble dysbiotic patterns reported in chronic intestinal inflammation, experimental colitis, and impaired barrier function.

      These new findings fully align and further support with our earlier conclusion that Kit<sup>W/Wv</sup> mice harbour pro-pathogenic microbiota.

      The new results are display in a new Figure 4, and described on pages 10-11 and discussed on pages 13-14.

      (3) In case the authors find dysbiosis, they should analyze the mechanisms by which Kit mutation causes dysbiosis.

      We have no intention to further explore Kit biology and in particular the intestinal pathophysiology caused by the Kit mutation because any results would not affect the central conclusion of our manuscript (see title). The review process and the revision shall center on making the core of a paper as conclusive as possible, and not widen a paper by requests 'tangential to the main conclusion' (Kaelin Jr. Nature 2017).

      References:

      Caruso, R., Ono, M., Bunker, M. E., Núñez, G. & Inohara, N. Dynamic and Asymmetric Changes of the Microbial Communities after Cohousing in Laboratory Mice. Cell Rep. 27, 3401-3412.e3 (2019).

      Feyerabend, T. B. et al. Deletion of Notch1 Converts Pro-T Cells to Dendritic Cells and Promotes Thymic B Cells by Cell-Extrinsic and Cell-Intrinsic Mechanisms. Immunity 30, 67–79 (2009).

      Feyerabend, T. B. et al. Cre-Mediated Cell Ablation Contests Mast Cell Contribution in Models of Antibody- and T Cell-Mediated Autoimmunity. Immunity 35, 832–844 (2011).

      Feyerabend, T. B., Gutierrez, D. A. & Rodewald, H.-R. Of Mouse Models of Mast Cell Deficiency and Metabolic Syndrome. Cell Metab 24, 1–2 (2016).

      Kaelin Jr, W. G. Publish houses of brick, not mansions of straw. Nature 545, 387– 387 (2017).

      Moore, R. J. & Stanley, D. Experimental design considerations in microbiota/inflammation studies. Clin. Transl. Immunol. 5, e92 (2016).

      Piliponsky, A. M. et al. Basophil-derived tumor necrosis factor can enhance survival in a sepsis model in mice. Nat. Immunol. 20, 129–140 (2019).

      Ridaura, V. K. et al. Gut Microbiota from Twins Discordant for Obesity Modulate Metabolism in Mice. Science 341, 1241214 (2013).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for improved or additional experiments, data, or analyses:

      (1) The study examines only a limited range of gut bacterial species, making it difficult to determine the specific contribution of E. coli expansion to the observed phenotype. A more comprehensive microbial profiling (e.g., 16S rRNA sequencing or metagenomics) would significantly strengthen the conclusions (e.g., Cpa3-mast cell deficient mice, Kit WWv, Kit W-sh/W-sh mice).

      As mentioned above, we now performed 16S rRNA sequencing of cecal contents from Kit<sup>W/Wv</sup> and Cpa3<sup>Cre/+</sup> mice and their respective littermates. Addressing microbiota in Kit<sup>W-sh/W-sh</sup> mice would not add information relevant for our paper.

      (2) The role of impaired peristalsis in KitW/Wv mice as a contributor to microbial dysbiosis and increased E. coli burden should be further explored. Complementary studies using KitW-sh/W-sh or SCF-deficient mice could clarify whether the observed microbiota changes are unique to the W/Wv model.

      We changed the title of the manuscript to emphasize our (unchanged) focus on the immunological role of mast cells in protecting against bacterial sepsis. The responses of Cpa3<sup>Cre</sup> mice clearly ruled out a role of mast cells to these infectious conditions. Our observation of altered microbiota in Kit<sup>W/Wv</sup> mice is consistent with their increased CLP susceptibility. We cite the known peristalsis deficit of Kit<sup>W/Wv</sup> mice as a possible explanation for the microbiota alterations. In the future, other investigators may find it interesting to elucidate the link between Kit mutations and dysbiosis. As stated further above, in our view, these additional questions and potential data have no bearings on the conclusions of our paper.

      (3) In the cohousing experiments, no data are provided to confirm whether microbiota normalization was achieved between groups. Including microbial composition data pre- and post-cohousing would improve the reliability of the interpretation.

      Cohousing made the susceptibility of Kit<sup>W/Wv</sup> and Kit<sup>+/+</sup> mice comparable. Detailed analysis of the extent of microbiota normalization would only make sense to ultimately determine specific taxa or combinations thereof that are responsible for the increased susceptibility of Kit<sup>W/Wv</sup> mice, a question that was never the goal of this study.

      (4) The use of Cpa3 Cre/+ mice introduces potential confounders, as these mice also have defects in basophils and T cells. Functional validation or additional models (e.g., Mas-TRECK or Mcpt5-Cre mice) could help isolate the mast cell-specific effects.

      See our detailed explanation above (Reviewer #1 Public review). It is incorrect to claim that Cpa3<sup>Cre/+</sup> mice have defects in T cells. The reviewer needs to provide published evidence for his/her claim that Cpa3-deficient mice exhibit defects in T cells. We as authors are also obliged to support our claims scientifically, and rightfully so.

      (5) Clarification is needed regarding the role of mast cell-derived TNF. Given previous reports using BMMC reconstitution that implicate mast cells as a source of TNF, reconciling these findings with the current study's results would strengthen interpretation.

      We also disagree here. Experiments in normal unmanipulated mice are inevitably superior to Kit mutants after BMMC reconstitution which is an artificial system. The transplanted cells do not mature normally and don’t settle in their natural niches. We mentioned in the discussion that there is conflicting literature derived from different models and mice.

      But we do not share the expectation that results obtained with Kit-independent models, that differ from previous studies using Kit mutants, necessarily require reconciliation. The experiments are simply incomparable and the most physiologically relevant experiment will pave the way. Research on mast cell functions based on BMMC-reconstituted Kit mutants has meanwhile proven to be unreliable.

      (6) A clearer delineation between mast cell-dependent and microbiota-mediated mechanisms in the discussion sections would enhance readability and impact.

      We restructured the discussion and distinguished between Kit-dependent and mast cell-dependent phenotypes. We also discussed in detail the observed microbiota differences and their influences for the outcomes in the different sepsis experiments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Weaknesses:

      (1) The cutoffs the authors used to define "conditionally essential" mutants are not reported. The results also lack validation for lethality using a titratable system. It would be ideal to validate several genes in each dataset to determine cutoffs (i.e. 5-fold decrease in insertion mutants) for conditional lethality. It was not done (or described) here.

      We acknowledge that independent validation using targeted mutants would further strengthen the assignment of synthetic lethality. As the primary aim of this study was the genome-wide identification of genetic interactions associated with loss of fitness in several mutant backgrounds, such validation was beyond the scope of the current work. Our experiments identified hundreds of lethal combinations and we have six datasets; therefore, validation of these interactions is not feasible and is indeed not common for a publication using TraDIS to generate leads for the community to follow up on. However, we already validated some of the hits in our original submission and compared our TraDIS data to some known synthetic-lethal interactions. We have also revised the manuscript to describe all other loci as candidate synthetic-lethal interactions and have highlighted the need for future validation studies in the Discussion.

      Regarding the reviewer’s query on thresholds, candidate synthetic-lethal interactions were identified using the tradis_essentiality.R script within the BioTraDIS analytical framework independently on each library: the six mutant backgrounds (DbamB, DbamC, DbamE, DsurA, Dskp, DdegP) and two E. coli BW25113 WT reference sets (an "internal" WT replicate sequenced as part of this study, and an "external" WT dataset from a previous study). This classifies each gene as essential, ambiguous, or non-essential for each library based on the bimodal distribution of insertion indices. Synthetic-lethal gene lists were then built by comparing essentiality classifications between each mutant and the WT sets, which were then flagged as shared/not shared with the internal or external WT essential gene lists in Supplementary Table 1. Therefore, a gene was treated as synthetic-lethal in a given mutant when it was called essential in that mutant but not shared with the WT essentiality call. We have clarified this point on line 923 in the Methods section as follows:

      “We ran the tradis_essentiality.R script within the BioTraDIS package independently on each library: the six mutant backgrounds (DbamB, DbamC, DbamE, DsurA, Dskp, DdegP) and two E. coli BW25113 WT reference sets (an "internal" WT replicate sequenced as part of this study, and an "external" WT dataset from a previous study[95]). This classifies each gene as essential, ambiguous, or non-essential for each library based on the bimodal distribution of insertion indices [30, 34]. Synthetic-lethal gene lists were then built by comparing essentiality classifications between each mutant and the WT sets, which were then flagged as shared/not shared with the internal or external WT essential gene lists in Supplementary Table 1. Therefore, a gene was treated as synthetic-lethal in a given mutant when it was called essential in that mutant but not shared with the WT essentiality call.”

      (2) Also, two mutations that both make the cells sick could provide an additive effect (i.e. dapF and BamB), which doesn't necessarily mean the pathways are linked. The authors should revise their wording. They have not shown genetic linkage in some cases.

      We revised the text to address this on line 693. However, the bamC mutant demonstrates no significant fitness cost under any of the conditions tested in the manuscript. Therefore, if this is simply an additive effect then it is not clear how this occurs, especially in the case of the dapF, bamC double mutant, and we offer an alternative explanation in the Discussion based on interpretation of the literature.

      (3) Mutations throughout the manuscript are not complemented. It would be ideal to add complementation data to show the gene-phenotype relationship is specific.

      We thank the reviewers for highlighting this and have complemented the experiments for the bamB-DNA replication link observation as described in response to reviewer 3.

      (4) Also, I would argue the term "conditionally essential genes" should be replaced with "synthetically lethal". Strains were compared in the same conditions but with different genetic backgrounds.

      We take the reviewer’s point and revised the text throughout.

      Reviewer #2 (Public Review):

      Weaknesses:

      (1) An important control in any genetic interaction study is to do complementation tests to demonstrate that the phenotype observed is indeed due to the missing gene under analysis. Although the Keio library was designed to avoid polar effects, it is impossible to predict other undesirable effects of the deletions (hitting of a non-annotated sRNA or RNA stability effects, for example). Thus, before one can safely conclude that a proposed genetic interaction is real, complementation tests should be carried out. This seems particularly important in the case of a new and surprising interaction, such as that between bamB and DNA replication and repair genes.

      We thank the reviewers for highlighting this and have provided the complementation experiments for the bamB-DNA replication link observations.

      (2) Why not include the suppressor interactions in the work? There are probably plenty, and in principle, they should be as informative as the conditional essential (or synthetic lethal) ones. The only one highlighted in the paper is that between bamB and diaA, since it nicely fits with the synthetic lethal effects with initiation inhibitors seqA and hda. Even if the authors cannot make sense of the suppressor interactions, their inclusion in the paper should make the dataset richer and more valuable to the community.

      Due to the nature of the BioTraDIS pipeline, we focused on gene essentiality and so only picked up potential genes that are essential in the parent but become non-essential in the mutants. This misses observations such as that made for diaA, which we hypothesised and checked manually. The data are publicly available for readers to use for their own studies and we have included some notes in Supplementary Table 1 to explain the filtering process along with another tab including the filtered essential gene lists.

      (3) The enrichment analysis in Figure 2B deserves some clarification. What is the meaning of gene ratio? How can single genes of a pathway yield an enrichment signal? Why weren’t seqA and hda included in the DNA replication class in 2B?

      We thank the reviewer for highlighting this point and realise we did not include a section on this analysis in the Methods section. As such we have included a section on line 935. KEGG pathway enrichment analysis was performed on the conditionally essential gene sets for each mutant background using the enrichKEGG function from the clusterProfiler R package [PMCID: PMC3339379], with the whole E. coli K-12 BW25113 genome used as the background gene set. Gene ratio is defined as the proportion of genes within a given conditionally essential gene set that are annotated to a specific KEGG pathway. Enrichment significance was assessed using a hypergeometric test comparing pathway representation within each query gene set to the whole-genome background.

      SeqA and Hda were not included in the DNA replication enrichment category because the KEGG enrichment analysis was based on existing KEGG pathway annotations, in which these genes are not assigned to the DNA replication pathway despite their well-established roles in replication initiation control. Considering the revision of the results regarding DNA replication, we feel this does not warrant further changes.

      (4) The writing puts too much emphasis on demonstrating that bam lipoproteins and chaperones are specialized instead of fully redundant. However, I have the impression this is a long-settled conclusion in the field, as the manuscript itself describes at several points when reviewing the literature.

      We revised the manuscript throughout to reduce this emphasis.

      Reviewer #3 (Public Review):

      In this work, Bryant, et al. investigate genetic interactions between non-essential members of the outer membrane protein biogenesis pathway and other genes in the genome using a transposon-directed insertion sequencing (TraDIS) approach in E. coli K-12. The authors identify interactions with other components of the envelope including LPS, peptidoglycan, and enterobacterial common antigen biogenesis, and they tie these interactions to specific members of the outer membrane biogenesis pathway. Although many of these interactions are known and have been previously investigated in the field, the study provides several synthetic phenotypes that could be useful for further investigations.

      The strengths of the paper include their unbiased, TraDIS approach, and follow up on the interactions they observe. The interactions with genes of unknown function also are of interest as they may suggest experiments to find the functions of these genes. The largest weakness of this paper is the use of a gene deletion allele for bamB that is known to be polar leading to decreased expression of an essential gene. This largely invalidates all results related to DNA replication. In addition, it is a weakness that the paper does not adequately address its place in the field through discussion of existing results on the interactions they investigate.

      The bamB mutant used here has been widely used in several previous studies (Cox et al., 2017, Gunasinghe et al., 2018, Psonis et al., 2019, Storek et al., 2019, Ranava et al. 2021, Steenhuis et al., 2021, Thewasano et al., 2023) with no concern raised and so we appreciate the reviewer’s expertise here and that they highlighted this issue for us to address.

      We thank the reviewer for highlighting this issue, as we have now completed complementation experiments for the CRISPRi depletion experiments and found that expression of bamB from a pBAD plasmid does not complement the DbamB strain in which seqA or hda is depleted, but expression of der in this system does complement the phenotype. Therefore, we have revised the title and the text to remove discussion of this potential link to DNA replication. We have included the new results and revised the existing DNA replication related figures as new figures S6-S8 and included a brief discussion of this polar effect in lines 283-314. We are very grateful to the reviewer.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      This study aimed to determine whether bacterial translation inhibitors affect mitochondria through the same mechanisms. Using mitoribosome profiling, the authors found that most antibiotics, except telithromycin, act similarly in both systems. These insights could help in the development of antibiotics with reduced mitochondrial toxicity.

      They also identified potential novel mitochondrial translation events, proposing new initiation sites for MT-ND1 and MT-ND5. These insights not only challenge existing annotations but also open new avenues for research on mitochondrial function.

      Strengths:

      Ribosome profiling is a state-of-the-art method for monitoring the translatome at very high resolution. Using mitoribosome profiling, the authors convincingly demonstrate that most of the analyzed antibiotics act in the same way on both bacterial and mitochondrial ribosomes, except for telithromycin. Additionally, the authors report possible alternative translation events, raising new questions about the mechanisms behind mitochondrial initiation and start codon recognition in mammals.

      Weaknesses:

      All the weaknesses I previously highlighted were adequately addressed.

      We thank the reviewer for carefully considering our revision.

      Reviewer #3 (Public review):

      Summary:

      Recently, the off-target activity of antibiotics on human mitoribosome has been paid more attention in the mitochondrial field. Hafner et al applied mitoribosome profilling to study the effect of antibiotics on protein translation in mitochondria as there are similarities between bacterial ribosome and mitoribosome. The authors conclude that some antibiotics act on mitochondrial translation initiation by the same mechanism as in bacteria. On the other hand, the authors showed that chloramphenicol, linezolid and telithromycin trap mitochondrial translation in a context-dependent manner. More interesting, during deep analysis of 5' end of ORF, the authors reported the alternative start codon for ND1 and ND5 proteins instead of previously known one. This is a novel finding in the field and it also provide another application of the technique to further study on mitochondrial translation.

      Strengths:

      This is the first study which applied mitoribosome profiling method to analyze mutiple antibiotics treatment cells. The mitoribosome profiling method had been optimized carefully and has been suggested to be a novel method to study translation events in mitochondria. The manuscript is constructive and well-written.

      Weaknesses:

      This is a novel and interesting study, however, most of conclusion comes from mitoribosome profiling analysis, as the result, the manuscript lacks the cellular biochemical data to provide more evidence and support the findings.

      Comments on revisions:

      The authors addressed most of my concerns and comments, although there is still no biochemical assay which should be performed to support mitoribsome profiling data.

      The author also carefully investigated the structure of complex I, however, I am surprised that the author chose to analyse a low-resolution structure (3.7 A). Recently, there are more high-resolution structures of mammalian complex I published (7R41, 7V2C, 7QSM, 9I4I). Furthermore, the authors should not only respond to the reviewers but also (somehow) discuss these points in the manuscript.

      We thank the reviewer for suggesting additional structural analyses. Of the suggested additional structures to look at, only 9I4I from Nguyen et al. 2026 was of human-derived complex I. Other structures were derived from species that utilized alternative codons at the 3’ end of these mRNAs, preventing comparison. However, for 9I4I, the authors were able to fit density for every amino acid of all mitochondrially-encoded proteins in complex I, except the first two residues of ND1 and ND5, which, as our manuscript suggest, are not translated.

      In addition to structures, we also attempted to identify additional publicly available mass spectrometry data for complex I. However, data which we identified either derived their proteins from bovine tissue, which utilizes different codons at the initiation sites, or did not capture the N-terminal region of the proteins. Therefore, we did not include this analysis.

    1. Author response:

      The following is the authors’ response to the previous reviews

      (1) Interpretation of LC3-II accumulation and phenocopying

      Therefore, LC3-II accumulation alone is insufficient to support phenocopying in my view.

      We agree with this assessment. Upon reconsideration, we concluded that LC3-II accumulation alone does not justify the use of the term "phenocopying." We have therefore removed this language from the manuscript.

      (2) Strength of conclusions regarding autophagosome biogenesis

      As presented, the findings support a correlative relationship rather than a defined role in autophagosome biogenesis.

      We agree that our original wording overstated the strength of the conclusions. To better reflect the data, we revised the text to state that our findings "expound upon" rather than "elucidate" the role of these membranes in autophagosome biogenesis.

      (3) Title wording

      The title states that ATG2A ‘engages’ Rab1A- and ARFGAP1-positive membranes during autophagosome formation... A more descriptive term, such as ‘associates,’ would more accurately reflect the data.

      We appreciate this suggestion. We revised the title to avoid implying a causal dependency. The title now states that ATG2A interacts with Rab1A- and ARFGAP1-positive membranes, emphasizing the membrane association observed in our study rather than a direct interaction with the proteins themselves.

      (4) ARFGAP1 knockdown phenotype

      The authors claim: ‘siRNA against ARFGAP1 had very little effect’ but the quantification and blots show actually no effect.

      We agree that the original wording was imprecise. The sentence has been revised to state:

      “siRNA against ARFGAP1 had no effect on flux.”

      This more accurately reflects both the data and our original interpretation.

      (5) Interpretation of ARFGAP knockdown experiments

      Conclusions drawn from KD experiments in Fig. S2 should be interpreted with caution, as knockdown efficiency is very low, particularly for ARFGAP1/3 in the triple knockdown.

      We agree with this caution and have revised the text accordingly. The manuscript now states:

      “Knockdown of ARFGAP2, ARFGAP3 and ARFGAP1-3 marginally increased autophagic flux (Fig. S2B,C), suggesting either no role or a minor role as negative regulators of autophagy.”

      This wording more appropriately reflects the limitations of the experiment.

      (6) Discussion of ERGIC/ERES remodeling literature

      It would strengthen the manuscript to discuss previous studies reporting ERES and ERGIC remodeling and formation of ERES-ERGIC contact sites (PMID: 34561617; PMID: 28754694).

      We appreciate the suggestion. These studies were already cited and discussed in the original submission, and therefore no additional changes were required.

      (7) Figure readability

      The font size in Figure 1A and Supplementary Figure S1G is too small for comfortable reading.

      We agree and have enlarged the labels in both figures to improve readability.

      (8) Clarification of starvation conditions in figure legends

      In Figures 2A-C and Figure 4, it is unclear how the cells were treated. Were they starved in EBSS?

      We have updated the corresponding figure legends to explicitly state the starvation conditions used in these experiments.

      (9) Interpretation of ARFGAP1 knockdown and LC3 lipidation

      In Figure 2A, ARFGAP1 knockdown appears to reduce LC3 lipidation without affecting Halo-LC3 cleavage.

      We do not observe a reproducible reduction in LC3 lipidation following ARFGAP1 knockdown and therefore do not believe this conclusion is supported by the data. No changes were made in response to this comment.

      (10) Clarification of protein-protein interaction statement

      The phrase ‘but protein-protein interactions appear to be limited to RAB1’ would benefit from clarification.

      We agree and have adopted the suggested wording. The manuscript now states:

      “but stable protein-protein interactions appear to be limited to RAB1.”

      We thank the reviewers again for their constructive feedback and for helping us improve the clarity and accuracy of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      King and colleagues generated a mouse with a point mutation in IL21R and investigated the influence on IL-21-mediated T and B cell activation and differentiation. They found that mutant mice show a reduced T and B cell response, with CD4 T cell differentiation into T follicular helper cells being primarily affected.

      Strengths:

      The authors combined in vitro and in vivo analysis, including bone-marrow chimeric mice.

      Weaknesses:

      The effect of the IL21R EINS mutant does not specifically affect STAT1, as clearly shown in Figure 1 H, I. Particularly at lower doses of IL21, which may be more relevant in vivo, the effects are very similar. A second key weakness is the very small Tfh response, a not very clear PD-1 and CXCR5 staining to identify Tfh, and a lack of a steady-state (prior to immunisation) comparison of Tfh numbers in the different mouse strains. The latter makes it impossible to know what fraction of the response is antigen-specific.

      Reviewer #2 (Public review):

      Summary:

      In the manuscript, "An IL-21R hypomorph circumvents functional redundancy to define STAT1 signaling in germinal center responses," Cecile King and colleagues identify a cytoplasmic site of the IL-21 receptor that differentially regulates STAT1 and STAT3 activation upon IL-21 stimulation. They further examine the immunological consequences of this site-specific alteration on Tfh differentiation and Tfh-dependent humoral immunity, raising important questions about how geneknockout models may obscure nuanced functional roles of signaling molecules.

      Strengths:

      The study convincingly highlights a non-redundant role for STAT1 downstream of IL-21-IL-21R signaling in the Tfh differentiation pathway. This conclusion is supported by in vitro analyses of STAT1 and STAT3 activation in CD4 T cells stimulated with IL-21 or IL-6; by in vivo assessments of Tfh and germinal center B cell responses in WT and IL21R-EINS mutant mice, including bonemarrow chimera systems; and by investigating the expression of Tfh-related molecules in WT versus IL21R-EINS CD4 T cells.

      Weaknesses:

      Although the experiments were carefully executed with appropriate controls, a key question remains unresolved: whether the Tfh differentiation defect in IL21R-EINS mice is directly attributable to reduced STAT1 activation. Rescue experiments that restore STAT1 signaling in IL21R-EINS TCR-transgenic CD4 T cells would provide strong evidence linking the mutation to impaired STAT1 activation and, consequently, defective Tfh differentiation. Without such evidence, it remains formally possible that additional, uncharacterized mutations introduced during ENU mutagenesis contribute to the phenotypes observed, particularly given the discrepancies between IL21R knockout and IL21R-EINS mutant mice.

      We agree that further experiments are needed to definitively show that the effect is attributable to reduced STAT1 activation alone. Rescue experiments will be a focus of future experiments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure1

      I would recommend changing the conclusion in Line 141 to 'potentially less affected' rather than unaffected as there is a clear and very consistent effect on pSTAT3 in every IL21 dose tested. Also, it seems that much more IL21 is required to induce STAT1 phosphorylation, which may explain the increased effect of the EINS mutant on this signalling pathway. Similarly, for the STAT5 data in Figure S1, there is very little phosphorylation beyond baseline phosphorylation (unstimulated), but there is a clear and consistent reduction in pSTAT5 in the EINS mutants.

      Neither pSTAT3 nor pSTAT5 were significantly different between WT and IL21rEINS cells in either the percentage or MFI. We have edited line 141 to state “the levels of phosphorylated STAT3 in CD4+ T cells were significantly less affected by the Il21r<sup>EINS</sup> mutation (Fig. 1I).”

      A minor point: Why is the MFI in IL21R-/- mice at 300 in panel C and at 200 in panel E? How representative is the reduced baseline pSTAT5 in IL21r-/- mice? 

      This is likely due to machine voltage during acquisition in a different experiment.

      Collectively, I would suggest concluding from that data that the EINS mutation affects IL21R signaling, which results in reduced STAT1, 3, and 5 phosphorylation, with pSTAT1 being most strongly reduced, particularly at high IL21 concentration. 

      Please also see response to above comment. Since neither pSTAT3 nor pSTAT5 were significantly different between WT and IL21rEINS cells, the data does not support that conclusion.

      All subsequent data therefore do not investigate the effect of the EINS mutation on STAT1, but on overall reduced IL21R signalling. This needs to be considered when interpreting the data. For example, the text in lines 171, 172, and 174 should be adjusted as the effect is neither only dependent on STAT1 nor is STAT3 signalling intact.

      Please also see responses above. It is possible that a different method for detection of phosphorylated STAT3 and STAT5 could have looked more closely into the effect at very low concentrations of IL-21. However, our findings using Westen Blot and flow cytometry only observed a significant difference in STAT1 activation. We have edited our sentence on line 333 to state “response in the presence of an IL-21 receptor mutant that predominantly affects IL-21 activation of STAT1”.

      A key signalling pathway downstream of IL21R is AKT and S6 phosphorylation. It would be important to also investigate the effect of the EINS mutation on these pathways.

      We agree and this will be a focus of future experiments.

      (2) Figure 2

      PNA or BCL-6 staining would be preferable to identify GC in Figure 1A, but the flow cytometry data in Figure 3 are convincing, so this is not absolutely necessary.

      Also, no conclusions can be made here about STAT1 specifically, and there could be other reasons why Tfh are slightly and temporarily reduced in EINS mice.

      (3) Figure 3

      Line 200: Please explain what is meant by 'despite an expansion of the IgG1 FAs B cell population on day5, the percentages of EINS IgG1* GC B cells were significantly lower (Fig. 3F). I cannot see any expansion of IgG1 FAS B cells, nor can I see a specific effect on day 7. Both total GC B cells and IgG1 GC B cells are similarly affected throughout the response. Some data points may not reach statistical significance, but the trend is very clear.

      Figure 3E shows the percentage of GC B cells increasing from day 3 to day 7 in WT and from day 3 to day 5 in IL21rEINS, with significant differences between WT and IL21rEINS on days 5 and 7. In Figure 3F he percentage of IgG1+ GC B cells increase from day 3 to day 14 in both genotypes, with a significant difference between IL21rEINS on day 7. We have edited to manuscript to state “Despite an increase in the IgG1<sup>+</sup> FAS<sup>+</sup> B cell population from day 3 in response to immunization, the percentages of Il21r<sup>EINS</sup> IgG1+ GC B cells were significantly lower relative to WT cells 7 days after SRBC immunisation (Fig. 3F).”

      (4) Figure 4

      How many days after SRBC immunization was the analysis done?

      The data show an intrinsic role of IL21R signalling to Tfh development, which may include a role for STAT1. The absence of any effect on GC B cells is somewhat surprising. Chimeric and irradiated mice sometimes mount poor immune responses, and GC B cell numbers are very low. What is the frequency of GC B cells in non-immunized mice? This would be important to know if the mice responded at all, and if they did, if the 'baseline' of GC activity differs.

      As stated in the figure legend for Figure 4 – on day 7 “Mixed BM chimaeras were reconstituted with equal ratios of WT CD45.1+ BM cells and Il21rEINS CD45.2+ BM cells. 8 weeks after transfer, the mice were immunized with SRBC and analysed 7days later.

      (5) Figure 7

      Please highlight that while only IL21R-/- mice showed a significant difference in the frequency of Tfh, a similar trend was observed in WT and IL21Reins mice. The data spread is smaller in the IL21R-deficient mice, facilitating statistical significance. As throughout the manuscript, this is not a STAT1 IL-21R mutant; it is a mutant with reduced IL21R signalling. In fact, the finding that IL-6 does not compensate for the EINS mutation may suggest that STAT1 plays a minor role in the biological effects observed.

      We can only report on the statistical significance of the data we have.

      (6) Other comments

      Figure 1 F/G. I think the Y axis should read pSTAT1 and pSTAT3, respectively.

      Thank you, we have corrected the graph accordingly.

      Figure S1A. Please change the order of WT, IL21EINS, and IL21R-/- to match the main figures (IL21R-/- last). Currently, A, B, and D have a different order, but C is like the main figures.

      The figure panels are aligned to show media, then either IL-2 or IL-6 and then IL21.

      Please provide complete flow cytometry gating strategies for all figures.

      Flow cytometry dating for P-STAT1 and p-STAT3 is shown in figure 1, for Tfh cells and Tfr cells in Figure 2, 4 and 5. Please also see supplementary figures for T cell gating and methods for detailed description of antibodies and dilutions used for immunostaining.

      Reviewer #2 (Recommendations for the authors):

      Line 332 requires revision.

      We have edited the final sentence to state” Taken together these findings demonstrate that, despite the strong ability of IL-6 to activate STAT1, IL-6 is ineffective at fully compensating for the germinal centre response in the presence of an IL-21 receptor mutant that predominantly affects IL-21 activation of STAT1.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully considered all comments and have revised the manuscript to address the key points raised. We have also updated the author list to include Kinga Niedobecka, who performed the additional flow cytometric validation of the engineered THP1 cell lines included in the revised manuscript. In particular, we have strengthened the validation of the THP1-CD1c system, clarified and better signposted the characterisation of CD1c-autoreactive T-cells using existing data, and refined the explanation of the mechanisms underlying enhanced responses to Mtb-infected cells. Some of the suggestions represent significant additional experimental work beyond the scope of this manuscript, and in these instances we have amended the text to clarify interpretation and limitations.

      eLife Assessment

      The study investigates how CD1c-restricted T cells respond to Mtb-infected APCs, leading to increased cytokine production and cytotoxic activity that may help control Mtb infection. While the work is important and will interest researchers in the field, the supporting evidence is incomplete and could be strengthened by additional experiments. Experiments would: (i) evaluate THP1-CD1c cells to determine whether MHC surface expression is reduced or entirely abolished, (ii) enhance confidence in the purity of the CD1c-specific T cell population isolated from blood, and (iii) suggest what additional signal THP1-CD1c cells treated with Mtb express that is absent from the untreated cells.

      (i) evaluate THP1-CD1c cells to determine whether MHC surface expression is reduced or entirely abolished

      We thank the Editor for highlighting this important point. We agree that it is essential to establish whether conventional MHC-mediated antigen presentation could contribute to the observed T-cell responses. To address this directly, we repeated and extended our flow cytometric validation of the engineered THP1 system. These data are now presented in an expanded Fig. 1A and include assessment of CD1c, classical MHC class I, MHC class II, β2m, CD1b and HLA-E across WT THP1, THP1-KO and THP1-CD1c cells. Our THP1-KO system is based on CRISPR-mediated knockout of both β2microglobulin (β2m) and the Class II transactivator (CIITA). Loss of β2m removes surface expression of β2m-dependent molecules, including classical MHC class I and endogenous CD1 proteins, while CIITA knockout prevents MHC class II expression. In this new analysis, WT THP1 cells expressed β2m and classical MHC class I, with low detectable MHC class II and HLA-E. In contrast, THP1-KO cells lacked detectable β2m, MHC class I, MHC class II, HLA-E, CD1b and CD1c. Importantly, THP1-CD1c cells retained robust CD1c expression through the CD1c-β2m fusion construct, while MHC class I, MHC class II, CD1b and HLA-E remained undetectable by flow cytometry.

      These extended validation data support the conclusion that residual MHC expression does not account for the observed T-cell responses, which are instead dependent on CD1c expression. We have revised the relevant section of the Results to incorporate these data and to clarify that the engineered THP1-CD1c APC system provides robust CD1c expression in the absence of detectable surface MHC-I or MHC-II (revised manuscript, page 5-6, lines 111-122; Fig. 1A and Fig. 1 legend).

      (ii) enhance confidence in the purity of the CD1c-specific T-cell population isolated from blood

      We agree that confidence in the specificity and purity of the CD1cautoreactive T-cell populations is essential. The relevant data were included in the original manuscript, but we recognise that they were not signposted clearly enough. We have therefore revised the Results to describe the enrichment, sorting, post-expansion validation and functional specificity of the T-cell lines more explicitly on page 8, lines 176-191.

      CD1c-autoreactive T-cell lines were generated from two independent donors using two complementary strategies. One line was generated by expansion with THP1-CD1c APCs followed by CD1c-endo tetramer-guided sorting and expansion. A second line was generated by direct enrichment using CD1c-endo streptamers, followed by CD1c-endo dextramer sorting and expansion. The gating strategy and post-sort validation are shown in Fig. S4. Importantly, after expansion, the enriched cells stained strongly with CD1c-endo tetramers, whereas unstained and irrelevant tetramer controls showed no detectable staining. In the main figure, both donor-derived lines are shown to be strongly CD1c-endo tetramer-positive, with post-expansion tetramer positivity of 97100% (Fig. 3A and 3C). Both lines were αβTCR+CD4+ and lacked detectable γδTCR or CD8 expression (Fig. 3B and 3D).

      We also highlight the functional validation of specificity. Both T-cell lines were activated by THP1-CD1c APCs but not THP1-KO APCs, as assessed by CD69 and CD25 upregulation (Fig. 3E). Importantly, the specificity of these cells was further supported by TCR transfer experiments. TCRs cloned from one of the CD1c-endo tetramer-positive T-cell lines were expressed in Jurkat reporter cells and conferred CD1c-endo tetramer binding, activation in response to plate-bound CD1c-endo protein, and enhanced activation in response to Mtb-infected THP1-CD1c APCs (Fig. 5B-D, page 10, lines 225244). This provides independent confirmation that the enriched T-cell line contained CD1c-reactive TCRs capable of mediating CD1c-dependent recognition.

      Together, these data support that the T-cell populations used in the functional assays are highly enriched CD1c-specific T-cell lines rather than mixed or nonspecific populations.

      (iii) suggest what additional signal THP1-CD1c cells treated with Mtb express that is absent from the untreated cells.

      We agree that identifying the additional signal provided by Mtb-treated THP1-CD1c cells is an important mechanistic question. We have now revised the Discussion to clarify our interpretation and to more explicitly outline the likely mechanisms (revised manuscript, page 16-17, lines 380-400).

      Our data suggest that the enhanced response to Mtb-infected THP1-CD1c cells is unlikely to be explained simply by increased CD1c expression, generic APC activation, or soluble cytokine release. CD1c expression was maintained but not increased on THP1-CD1c cells after Mtb infection, and stimulation with TLR2 or TLR4 agonists did not reproduce the enhanced cytotoxicity observed after Mtb infection. In addition, Mtb-treated THP1-CD1c cells alone produced IL-8 and RANTES, but not the broader cytokine profile observed in T-cell co-cultures. Together, these data suggest that Mtb exposure provides an additional CD1c-dependent activating signal.

      We now discuss that this signal is most likely an altered CD1c-presented lipid repertoire on Mtb-exposed APCs. Possible mechanisms include presentation of Mtb-derived lipids, infection-induced accumulation of host-derived stimulatory “stress lipids”, presentation of bacterial and mammalian shared lipids, or altered lipid processing and trafficking during infection. These possibilities are consistent with prior studies showing enhanced responses of autoreactive CD1-restricted T-cells to microbial stimulation and our TCR transfer experiments seemingly support a CD1c-TCR-dependent recognition mechanism. However, because we have not directly identified the lipid ligands presented by CD1c on Mtb-infected APCs, we now state this as a mechanistic hypothesis rather than a conclusion, and a key outstanding question.

      We have revised the Discussion (page 16-17, lines 380-410) to make this limitation explicit. Future studies will require isolation of CD1c molecules from Mtb-infected cells and then lipidomic analysis and mass spectrometry to define the CD1c-associated lipid species.

      Reviewer #1 (Public review):

      Strengths:

      (1) This study asks an important question. The single-cell transcription analysis suggests the inherent cytotoxic program of lipid-CD1c cells and provides insights into their phenotypic and potential functional profiles. Function experiments suggest that these autoreactive T-cells can react to Mtb infection, adding to the paradigm of infection control by these non-conventional T-cell populations.

      We thank the reviewer for this positive assessment of the importance of the study and for recognising the value of the single-cell transcriptional analysis and functional experiments. We are pleased that the reviewer agrees that our findings provide insight into the cytotoxic effector programme of CD1c-autoreactive T-cells and their potential contribution to immune responses during Mtb infection.

      Weaknesses:

      (2) The study lacks sufficient rigor; conclusions may be strengthened with the incorporation of more controls, and some deeper characterization of the THP1 system and the CD1c-specific T-cells isolated from blood. Crucial conclusions are drawn from the cell mixing experiments involving the engineered THP-1 system and CD1c-lipidspecific T-cells from blood. These cells need more in-depth characterization. The expression of MHC-I/II is clearly reduced in THP1-CD1c cells. However, it is important to ensure that it is completely abolished, since a residual expression can skew the result with activation of conventional T-cells in the blood or low levels of conventional T-cells that may be present in the CD1c-tetra/multimer sorted T-cells

      We agree that this is an important point and have addressed it by adding new experimental controls and by clarifying the validation of the CD1c-autoreactive T-cell lines.

      First, we repeated flow cytometric validation of the existing markers and extended the panel to assess additional surface molecules across WT THP1, THP1-KO and THP1CD1c cells. The revised Fig. 1A therefore includes repeat staining for β2m, classical MHC class I, MHC class II and CD1c, together with newly added staining for HLA-E and CD1b.

      The THP1-KO system is based on CRISPR-mediated knockout of both β2-microglobulin (β2m) and the Class II transactivator (CIITA). Loss of β2m removes surface expression of β2m-dependent molecules, including classical MHC class I and endogenous CD1 proteins, while CIITA knockout prevents MHC class II expression. The repeated analyses confirmed the original staining pattern, while the additional HLA-E and CD1b stains further extended validation of the system. WT THP1 cells expressed β2m and classical MHC class I, with low detectable MHC class II and HLA-E. In contrast, THP1-KO cells lacked detectable β2m, MHC class I, MHC class II, HLA-E, CD1b and CD1c.

      Importantly, THP1-CD1c cells retained robust CD1c expression through the CD1c-β2m fusion construct, while MHC class I, MHC class II, HLA-E and CD1b remained undetectable by flow cytometry. We have revised the Results to describe these new validation experiments more clearly (revised manuscript, page 5-6, lines 111-122, Fig. 1A and Fig. 1 legend).

      Second, we have strengthened the description of the purity and specificity of the CD1c-autoreactive T-cell lines. These lines were generated using two complementary approaches, namely expansion with THP1-CD1c APCs followed by CD1c-endo tetramer-guided sorting, and direct enrichment using CD1c-endo streptamers followed by CD1c-endo dextramer sorting and expansion. The gating strategy and post-sort validation are shown in Fig. S4. After expansion, the enriched cells stained strongly with CD1c-endo tetramers, whereas unstained and irrelevant tetramer controls showed no detectable staining. Both donor-derived lines were strongly CD1c-endo tetramer positive, with post-expansion tetramer positivity of 97 to 100%, and both were αβTCR+CD4+ with no detectable γδTCR or CD8 expression (Fig. 3A-D). Functionally, both lines responded to THP1-CD1c APCs but not parental THP1-KO APCs, as assessed by CD69 and CD25 upregulation (Fig. 3E). We have revised the Results to signpost these data more clearly (revised manuscript, page 8, lines 176–191).

      Finally, TCR transfer experiments provide independent confirmation of CD1c-specific recognition. TCRs cloned from one of the CD1c-endo tetramer-positive T-cell lines conferred CD1c-endo tetramer binding and CD1c-dependent activation when expressed in Jurkat reporter cells (Fig. 5B-D, revised manuscript, page 10, lines 225244). Together, the absence of detectable MHC-I/MHC-II expression in the engineered APC system, the high CD1c-endo tetramer enrichment of the T-cell lines, the lack of activation against THP1-KO cells, and the TCR transfer experiments support the conclusion that the observed responses are driven by CD1c-dependent recognition rather than residual conventional MHC-mediated activation.

      (3) Figure 2: The immunohistochemistry appears to be shown only for one biopsy; it may be worth quantifying the immunohistochemistry of all five.

      We thank the reviewer for this helpful suggestion. We agree that quantitative analysis of CD1c immunohistochemistry across all biopsies would be valuable. We examined CD1c staining across all five TB lung biopsies and observed a consistent spatial pattern, with CD1c staining generally low or infrequent in central granulomatous regions and more apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich areas.

      However, because these were diagnostic human biopsy samples with substantial variation in tissue size, architecture, granuloma representation and inflammatory composition, we do not think that simple bulk quantification of CD1c-positive area across biopsies would be robust or biologically interpretable. In particular, quantification would be strongly affected by whether a section captured granuloma centre, peripheral inflammatory regions, lymphoid aggregates, or uninvolved lung tissue. We have therefore retained the IHC as representative spatial evidence of CD1c expression in TB lung tissue, rather than presenting it as a quantitative comparison across anatomical compartments.

      We have revised the Results to make this clearer, stating that CD1c expression was observed across the biopsies analysed but was spatially heterogeneous, with staining most apparent away from the granuloma centre and in lymphoid/inflammatory regions (revised manuscript, page 7, lines 144-151). We have also tempered the interpretation in the Discussion to avoid overstatement and now highlight systematic quantitative spatial analysis of larger tissue cohorts as an important future direction (revised manuscript, page 18, lines 427-430).

      (4) The expression of CD1 molecules goes up during the differentiation of MoDC, and Mtb infection prevents or dampens the upregulation. Does Mtb infection downregulate the CD1 expression of mature DCs? Can the effect of Mtb on the expression of CD1a,b,c molecules be investigated using CD1c-expressing DCs from blood? What could be the reason THP-1 cells do not downregulate CD1 molecules upon Mtb infection, and how about the expression of CD1a and b?

      We agree that the distinction between impaired CD1 upregulation during MoDC differentiation and active downregulation of CD1 expression on already differentiated CD1-expressing DCs is important.

      In the revised manuscript, we have clarified that our primary cell data assess the effect of Mtb infection on differentiated MoDCs that already express CD1 molecules, rather than only examining failure of CD1 induction during differentiation. Specifically, we analysed a published RNA-sequencing dataset from differentiated human MoDCs infected with live Mtb and observed reduced expression of group 1 CD1 genes, including CD1A, CD1B and CD1C, at 48 hours after infection. We then validated this experimentally at the protein level by flow cytometry, showing reduced CD1c expression on primary MoDCs after live Mtb infection. These data support the conclusion that Mtb infection can reduce CD1 expression on CD1c-expressing primary DCs.

      We agree that analysis of freshly isolated blood CD1c+ DCs would be valuable. However, these cells are rare in peripheral blood and are technically challenging to isolate in sufficient numbers for live Mtb infection assays and downstream flow cytometric or functional analysis. For this reason, we used MoDCs as a tractable primary human DC model to assess infection-induced changes in CD1 expression. We now acknowledge in the revised Discussion that validation in primary blood-derived CD1c+ DCs would be an important future direction.

      We have also clarified why CD1c expression is not downregulated in the engineered THP1-CD1c system. In primary DCs, Mtb-mediated suppression of CD1c has been linked to host regulatory mechanisms, including post-transcriptional regulation by miRNAs such as miR-381-3p, which targets the 3′ UTR of endogenous CD1c transcripts. In contrast, CD1c expression in our THP1-CD1c cells is driven by a lentiviral CD1c-β2m fusion construct under a heterologous promoter and expressed from a cDNA lacking the native untranslated regions. Therefore, CD1c in this system is not expected to be regulated in the same way as endogenous CD1c in primary DCs.

      This is a deliberate feature of the model. It allows us to assess CD1c-dependent T-cell responses to Mtb-infected APCs without the confounding effect of infection-induced CD1c loss. THP1-CD1c cells do not express endogenous CD1a or CD1b because the parental THP1-KO cells lack β2m-dependent endogenous CD1 surface expression, and only CD1c is reintroduced through the CD1c-β2m fusion construct. We have revised the Results and Discussion to clarify these points (revised manuscript, pages 7- 8, lines 165174 and pages 17- 18, lines 411-430).

      (5) Figure 3: (F) What does the X-axis read for the no infection group? The value for MOI = 0 should be incorporated for the infected T-cell group.

      We agree that the original presentation could be clearer. The uninfected condition corresponds to MOI = 0, whereas the remaining points represent THP1-CD1c APCs exposed to increasing amounts of UV-killed Mtb. We have retained the figure layout but revised the figure legend to clarify that the MOI values on the x-axis apply only to the Mtb-treated conditions, and that the no-infection/no-treatment control represents MOI = 0 (Fig. 3 legend).

      (6) Figure 4: In the lysis assay, THP1-CD1c cells (uninfected and infected) incubated alone should be incorporated.

      We agree that APC-only controls are essential for interpreting the lysis assay, and we apologise that this was not sufficiently clear in the original manuscript. THP1-CD1c cells cultured alone, both uninfected and Mtb-infected, were included in all assays and used to define baseline target-cell viability for each matched condition.

      The data in Fig. 4 are presented as specific lysis to isolate the effect of T-cells on target cell viability. Specifically, THP1 viability in APC-only wells was used as the baseline and subtracted from the corresponding T-cell co-culture condition within the same experiment. Thus, lysis of uninfected THP1-CD1c cells was calculated relative to uninfected THP1-CD1c cells cultured alone, and lysis of Mtb-infected THP1-CD1c cells was calculated relative to Mtb-infected THP1-CD1c cells cultured alone. This presentation allows the T-cell-mediated effect to be visualised while accounting for baseline viability differences.

      We have revised the Methods and Fig. 4 legend to make this calculation more explicit (revised manuscript, page 25, lines 608-614; Fig. 4 legend).

      (7) A quantitative brief on the single cell TCR sequencing - including how many T-cells were sequenced and the frequency of different clone including EM1 and EM2 - should be shown.

      We agree that the single-cell TCR sequencing data required clearer quantitative description. We have expanded the Results and Fig. 5 legend to include the number of single cells analysed and the frequency of the dominant clonotypes. Single CD1c-endo tetramer-positive T-cells were sorted into individual wells for targeted TCR sequencing. After filtering and manual curation, 11 single cells yielded productive paired αβ TCR sequences. The repertoire was oligoclonal, with two dominant productive clonotypes accounting for 10 of 11 paired TCRs. EM1 was detected in 6 of 11 cells and EM2 was detected in 4 of 11 cells. These data support the selection of EM1 and EM2 for TCR-transfer experiments and clarify that they were dominant clonotypes within the CD1c-endo tetramer-positive T-cell line rather than arbitrarily selected TCRs. We have revised the Results and Fig. 5 legend accordingly (revised manuscript, page 10, lines 225-235, Fig. 5 legend).

      Reviewer #1 (Recommendations for the authors):

      (8) Perform an experiment to assess activation of T-cells expressing EM1 or EM2, upon mixing with CD1c-expressing dendritic cells isolated from human blood, with and without Mtb infection.

      We agree that testing EM1 and EM2 TCRs against primary dendritic cells is an important question. However, in the specific context of Mtb infection, the proposed experiment is difficult to interpret because Mtb downregulates CD1c expression on primary dendritic cells. This is supported by previous studies showing that Mtb and BCG suppress CD1c expression on DCs [1,2], and by our own data showing reduced CD1 group 1 transcript expression in Mtb-infected MoDCs and reduced CD1c protein expression on primary MoDCs following Mtb infection (Fig. 2B and 2C). Therefore, mixing EM1- or EM2-expressing Jurkat T-cells with Mtb-infected primary CD1c-expressing DCs would introduce a major confounder: reduced T-cell activation could reflect loss of CD1c expression rather than absence of a CD1c-dependent Mtb-induced activating signal. This is precisely why we used the engineered THP1-CD1c system, in which CD1c expression is preserved during Mtb infection (Fig. 2D). This model allowed us to test whether Mtb infection enhances CD1c-TCR-dependent activation without the confounding effect of infection-induced CD1c loss.

      Using this controlled system, we show that EM1 and EM2 TCRs confer CD1c-endo tetramer binding, activation in response to plate-bound CD1c-endo protein, activation in response to THP1-CD1c but not THP1-KO APCs, and enhanced activation in response to Mtb-infected THP1-CD1c APCs (Fig. 5B-D). These data support the conclusion that the enhanced response to Mtb-infected APCs is mediated through CD1c recognition by the TCR.

      We have revised the Results and Discussion to clarify this rationale and to explain why the engineered THP1-CD1c system was necessary for these experiments (revised manuscript, page 10, lines 242-244; page 17-18, lines 411-430).

      (9) Conduct an experiment to assess T-cell cytotoxicity expressing EM1 or EM2, in the presence and absence of Mtb infection.

      We agree this would be a valuable experiment. As outlined in our response to 8, EM1 and EM2 were cloned into Jurkat T-cells to test TCR-dependent CD1c recognition and activation, not cytotoxic effector function. Jurkat T-cells are not cytotoxic effector cells<sup>3</sup>, so they are not suitable for target-cell killing assays.

      Cytotoxicity was instead assessed using the original CD1c-autoreactive T-cell lines (Figs. 3F-G and 4D-E). Testing EM1- or EM2-mediated killing would require engineering and validating primary human T-cells expressing these TCRs, which is a substantial additional workflow. We have clarified in the revised manuscript that the EM1/EM2 experiments demonstrate TCR-dependent recognition, while cytotoxicity was assessed using the CD1c-autoreactive T-cell lines (revised manuscript, page 10, lines 234–244).

      (10) A list of primers used for TCR sequencing should be provided.

      We have now provided the primer sequences used for targeted single-cell TCR sequencing in a new supplementary table (Table S1). We have also revised the Methods to provide additional detail on the single-cell TCR sequencing workflow, including CD1c-endo tetramer-guided single-cell sorting, oligo-dT reverse transcription, universal cDNA amplification, targeted amplification of TCR variable regions using TRAC-, TRBC-, TRGC- and TRDC-specific primers, well-specific 8-bp barcoding, size selection, library preparation and MiSeq sequencing. In addition, we now cite the SMART-seq2 protocol on which the approach was based (Picelli et al., 2014) (revised manuscript, page 20, lines 491-503; new Table S1).

      Reviewer #2 (Public review):

      Strengths:

      (1) The study is designed well and has developed many exciting tools to generate specific information.

      We thank the reviewer for this positive assessment of the study design and for recognising the value of the experimental tools developed in this work.

      Weaknesses:

      (2) The study has weaknesses in two important parameters - novelty and relevance in controlling TB. Further, the results could be better presented and discussed to allow easy understanding of the experimental design

      We accept that the novelty and relevance to TB control could be made clearer in the manuscript. However, we believe the study makes several important and previously unreported contributions, and we have revised the Introduction, Results and Discussion to improve the clarity of the experimental design and to state the conceptual advance more explicitly.

      First, to our knowledge, this is the first study to demonstrate that human CD1c-autoreactive T-cells respond more strongly to Mtb-infected CD1c+ APCs than to uninfected CD1c+ APCs. Previous work has shown that CD1c-autoreactive T-cells exist in human blood and can respond to CD1c-expressing cells in the absence of exogenous antigen. However, their role during infection has remained unclear. Our findings extend the field beyond the established steady-state, autoimmune and tumour contexts of CD1c autoreactivity by identifying Mtb-infected APCs as a biologically relevant setting in which these cells acquire enhanced effector activity. We propose that CD1cautoreactive T cells may not simply represent autoreactive bystanders that become pathogenic in disease, but instead form an evolutionarily conserved arm of lipid immune surveillance that can detect infection-associated changes in antigen presentation. Given the long-standing selective pressure imposed by microbial infection throughout human evolution, it is plausible that protection against infection represents a central physiological function of these cells, with their roles in autoimmunity and cancer reflecting the same capacity to sense altered self-lipid landscapes in other settings. Our data provide initial functional evidence supporting this model.

      Second, the study provides functional evidence that these cells are not simply activated by infected APCs, but can mediate effector functions relevant to antimicrobial immunity. CD1c-autoreactive T-cells showed enhanced activation, cytokine production and cytotoxicity in response to Mtb-infected APCs, and led to reduced Mtb burden under in vitro conditions. These findings are directly relevant to TB immunity because cytotoxic T-cell pathways and antimicrobial molecules such as granulysin have been implicated in control of intracellular Mtb.

      Third, the study links these functional observations to the ex vivo biology of human CD1c-autoreactive T-cells. Single-cell transcriptomic profiling demonstrates that these cells are enriched for cytotoxic effector-memory programmes and express molecules associated with target-cell killing and antimicrobial activity. This provides an independent, unbiased cellular basis for the functional assays and strengthens the conclusion that CD1c-autoreactive T-cells represent a plausible effector population in anti-mycobacterial immunity.

      We agree that the experimental design needed clearer presentation. In the revised manuscript, we have improved signposting of the stepwise logic of the study: (1) defining and validating the THP1-CD1c APC system, (2) demonstrating CD1cautoreactive T-cell enrichment and specificity, (3) testing responses to UV-killed and live Mtb, (4) confirming TCR-dependent CD1c recognition using EM1 and EM2 TCR transfer, and (5) integrating these functional data with single-cell transcriptomic profiling of ex vivo CD1c-autoreactive T-cells. These revisions aim to make the experimental design easier to follow and to clarify how each section supports the overall conclusion.

      We have revised the Introduction and Discussion accordingly to more clearly state the novelty and TB relevance of the work (revised manuscript, page 4-5, lines 83-104; page 14, lines 323-329).

      (3) At several places, UV-killed or live Mtb were used. What is the rationale behind that?

      We have now added a Methods statement explaining that UV-killed Mtb was used for controlled exposure to defined amounts of Mtb-derived antigen, particularly in dose-response cytotoxicity and cytokine-release assays, whereas live Mtb was used to assess T-cell activation, target-cell lysis and relative bacterial burden during APC infection with proliferating Mtb. We have also ensured that the figure legends clearly specify whether UV-killed or live Mtb was used in each experiment (revised manuscript, page 22, lines 534-540).

      (4) Why use irradiated THP1-CD1c cells for activating T-cells?

      Irradiated THP1-CD1c cells were used only during the T-cell expansion phase to provide sustained CD1c-mediated stimulation while preventing proliferation of the THP1 APCs. This was necessary because the expansion cultures lasted up to 12 days, during which non-irradiated THP1 cells would continue to divide and could overgrow the T-cell culture. Irradiation therefore allowed THP1-CD1c cells to function as APCs while maintaining controlled culture conditions and enabling selective expansion of CD1c-reactive T-cells. We have clarified this rationale in the Methods (revised manuscript, page 19-20, lines 472-474).

      (5) While functional assays identified only CD4+ cells as CD1c-restricted, scRNAseq shows that both CD4+ and CD8+ cells exhibit this phenotype

      We agree with the reviewer’s observation and have clarified this point in the revised Discussion. The functional assays were performed using CD1c-autoreactive Tcell lines generated from two donors. Both lines were CD4+αβTCR+, reflecting the outcome of the enrichment, sorting and expansion process used to generate sufficient T-cells for functional assays. These lines therefore provide mechanistic evidence that CD4+ CD1c-autoreactive T-cells can recognise CD1c+ APCs and respond more strongly to Mtb-infected APCs, but they are not intended to represent the full diversity of the CD1c-autoreactive T-cell compartment.

      By contrast, the single-cell RNA-seq analysis was designed to provide a broader ex vivo assessment of CD1c-endo-binding T-cells without relying on prolonged in vitro expansion. This revealed that CD1c-autoreactive T-cells include both CD4+ and CD8+ populations, with enrichment of cytotoxic effector-memory programmes. We therefore interpret the functional and single-cell datasets as complementary: the functional assays provide mechanistic validation using tractable CD1c-reactive T-cell lines, while the single-cell data demonstrate that the broader ex vivo CD1c-autoreactive compartment is phenotypically diverse and includes both CD4+ and CD8+ cytotoxic populations.

      We have revised the Discussion to clarify this point and to emphasise that combining in vitro functional assays with ex vivo single-cell profiling allowed us to capture both mechanistic activity and broader cellular diversity (revised manuscript, page 14-15, lines 341-349).

      (6) Identifying the specific lipid antigen presented by CD1c could add greater value to the study.

      We agree that identifying the specific CD1c-presented lipid antigen(s) would add important mechanistic insight. As noted in the response to the Editor, our data suggest that the enhanced response to Mtb-infected THP1-CD1c APCs is most likely due to altered CD1c-associated lipid presentation. However, defining these lipid species would require isolation of CD1c from infected APCs followed by specialised mass spectrometry-based lipidomics, which is a substantial additional workflow. We have revised the Discussion to state this limitation clearly and to highlight lipid identification as an important next step (revised manuscript, page 16-17, lines 380-410).

      (7) Since autoreactivity was independent of exogenous antigen, the cytotoxic activity should also be independent of exogeneous antigens? What additional signal a THP1-CD1c cells treated with UV-killed Mtb express that is absent from the untreated cells?

      CD1c-autoreactive T-cells likely recognise self-lipids presented by CD1c, but our data show that this response is enhanced after Mtb exposure. We interpret this as evidence that Mtb alters the quality or abundance of CD1c-associated lipid ligands, potentially through Mtb-derived lipids or infection-induced changes in host lipid metabolism. We have revised the Discussion to clarify that the precise lipid ligand(s) remain unidentified and will require future CD1c-lipidomic analysis (revised manuscript, page 16-17, lines 380-410).

      (8) The relative Mtb growth assay is confusing. CD1c cells with Mtb infection triggers massive lytic response, as shown in Figure 4. Under similar conditions, in Figure 6, the authors report a significant decline in Mtb growth in these cells. The problem is that with the kind of lytic response observed, a lot more Mtb could be present extracellularly and would evade killing. How do we reconcile the two observations?

      In the Mtb lux assay, extracellular bacteria were removed by washing after the initial infection step, so the starting bacterial population measured in the co-culture assay is expected to be predominantly cell-associated. We also recognise that the luminescence readout measures total viable lux-expressing Mtb under the assay conditions and does not distinguish intracellular from extracellular bacteria at later time points.

      The cytotoxicity observed in Fig. 4 reflects enhanced but incomplete lysis of infected APCs. Therefore, although T-cell-mediated lysis could release some bacteria from infected target T-cells, the reduced luminescence observed in Fig. 6C indicates a lower net viable Mtb burden under these co-culture conditions. We interpret this as the combined outcome of CD1c-autoreactive T-cell effector activity, including cytotoxicity and antimicrobial mediators such as granulysin and cytokines, rather than as direct evidence of selective intracellular bacterial killing.

      We have revised the Results and Methods to clarify that the Mtb lux assay measures relative viable bacterial burden/luminescence under in vitro co-culture conditions. We have changed the discussion to avoid over-interpreting this assay as distinguishing intracellular from extracellular Mtb killing (revised manuscript, page 11-12, lines 266274, page 26, lines 633-637).

      Reviewer #2 (Recommendations for the authors):

      (9) Nearly 40-50% of the samples did not respond to THP1-CD1 stimulation. What contributes to this diversity?

      We agree that there is clear donor-to-donor variability in the response to THP1-CD1c stimulation [4]. Approximately one-third of donors did not show detectable expansion under these assay conditions. This likely reflects differences in the precursor frequency and TCR repertoire composition of CD1c-autoreactive T-cells between donors, together with variation in activation state and responsiveness during short-term in vitro expansion. Apparent non-response may also reflect low-frequency CD1c-reactive populations that are present but fall below the detection threshold. We have revised the Discussion to acknowledge donor heterogeneity as an expected feature of primary human CD1c-autoreactive T-cell responses (revised manuscript, page 14-15, lines 341-349).

      (10) For lung biopsy staining, how is CD1 expression in healthy tissue or some unrelated inflammatory condition?

      The purpose of the lung biopsy staining was to determine whether CD1c-expressing cells are present in human TB lung tissue and to assess their spatial relationship to granulomatous inflammation, rather than to perform a formal comparison between healthy, non-TB inflammatory and TB lung tissue.

      Across the TB biopsies analysed, CD1c staining was spatially heterogeneous. CD1c expression was generally low or infrequent in central granuloma regions and in tissue regions remote from granulomatous inflammation, whereas staining was more apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich regions. We have revised the Results to clarify that these data are presented as representative spatial observations within TB lung tissue, rather than as a quantitative comparison with healthy or unrelated inflammatory tissue.

      We agree that comparison with healthy lung and non-TB inflammatory lung tissue would provide useful additional context, particularly for distinguishing TB-associated changes from more general inflammatory induction of CD1c. We now acknowledge this as an important future direction (revised manuscript, page 18, lines 425-430). We have also revised the Results to clarify the spatial pattern of CD1c staining within TB lung tissue (revised manuscript, page 7, lines 144-151).

      (11) What was the rationale for using UV-killed or live Mtb for different experiments?

      This point is addressed in our response to 3 above.

      Reviewer #3 (Public review):

      Strengths:

      (1) The manuscript is well written, and the novelty, impact, and limitations of this study are precisely highlighted by the authors.

      We thank the reviewer for this positive assessment of the manuscript, particularly their recognition of the study’s novelty, impact and balanced discussion of its limitations.

      (2) Lipid antigen identification and direct lipid identification via lipidomics/MS of CD1c-bound lipids from Mtb-infected APCs would clarify whether the enhancement arises from altered self-lipids or subtle Mtb lipids

      We agree that direct identification of CD1c-bound lipids from Mtb-infected APCs would provide important mechanistic insight and help determine whether enhanced activation reflects altered self-lipids, Mtb-derived lipids, or shared lipid species. As noted above, this would require isolation of CD1c from infected APCs followed by specialised mass spectrometry-based lipidomic analysis, which represents a substantial additional workflow. We have revised the Discussion to state this limitation clearly and to highlight CD1c-lipidomic analysis as an important next step (revised manuscript, page 16-17, lines 380-410).

      Reviewer #3 (Recommendations for the authors):

      (3) Figure 2Ai-vi, lines 134-136. The authors should include the data from central granuloma staining to solidify their claim of the presence of CD1c expression remote from the centre of TB granulomas.

      Central granuloma regions are included in Fig. 2A, including panels showing staining within granulomatous tissue where CD1c expression is low or infrequent compared with distal inflammatory and lymphoid/B-cell follicle-rich regions. We agree that this spatial distinction was not sufficiently clear in the original text and figure legend.

      We have therefore revised the Results and Fig. 2 legend to more explicitly guide the reader through the central versus distal regions shown in Fig. 2A. The revised text now states that CD1c expression was observed across lung biopsies from all five TB patients, but was spatially heterogeneous, with staining most apparent in distal inflammatory tissue and lymphoid/B-cell follicle-rich areas, and generally low or infrequent in central granuloma regions (revised manuscript, page 7, lines 144-151; Fig. 2 legend).

      (4) Figure 2D, lines 149-151. The authors should clarify whether the CD1c resistance to downregulation is model-specific to THP1-CD1c-APCs or an overexpression artefact

      As described in our response to Reviewer 1, 5, we agree that this point required clearer explanation. We have clarified in the revised Results and Discussion that the preservation of CD1c expression in THP1-CD1c APCs likely reflects the engineered nature of this system, rather than a general feature of endogenous CD1c regulation during Mtb infection.

      Specifically, in primary MoDCs, Mtb infection reduces CD1c expression at both transcript and protein levels (Fig. 2B and 2C). In contrast, CD1c in THP1-CD1c APCs is expressed from a heterologous CD1c-β2m fusion construct rather than from the endogenous CD1C locus. Therefore, its resistance to downregulation is likely model-specific and related to the expression system. We now state this in the Results and Discussion, and explain that this feature allows CD1c-dependent T-cell responses to Mtb-infected APCs to be assessed without the confounding effect of infection-induced CD1c loss (revised manuscript, page 7-8, lines 165-174; page 17-18, lines 411-430).

      (5) Figure 6C. The relative Mtb burden is measured through luminescence. While this correlates closely with CFUs, confirmation with plating is better evidence.

      We have previously shown close correlation between luminescence in our system and CFUs<sup>5</sup> (Bielecka mBio 2017, PMID: 28174307). Perhaps controversially, we propose that luminescence is a better readout of total Mtb load. Luminescence captures all metabolically active Mtb, whilst CFUs may be confounded by clumping of bacteria, for example, giving an underestimate. However, we agree with the reviewer that luminescence is an indirect measure of bacterial burden and that CFU plating would provide additional confirmatory evidence. We have therefore revised the Results, Discussion and Methods to describe the assay more cautiously as a measure of relative viable Mtb burden/luminescence, and we now acknowledge the absence of CFU confirmation as a limitation of the study (revised manuscript, page 11-12, lines 266-274; page 17, lines 400-403; page 26, lines 633-637).

      (6) Figure 7. The authors show that CD1c-autoreactive T-cells exhibit cytotoxic effector memory phenotype. While the sc-RNAseq subsampling is robust, the number of sample donors being 2 might create a potential bias.

      We agree that the use of two donors for the single-cell RNA-seq analysis is a limitation and could introduce donor-specific bias. We have now stated this more explicitly in the Discussion. Importantly, the scRNA-seq data are not used alone to define function, but rather provide an ex vivo phenotypic framework that complements the functional assays showing CD1c-dependent activation, cytokine production, cytotoxicity and reduced relative Mtb burden.

      The subsampling analysis supports the robustness of the transcriptional patterns within this dataset, but we agree that larger donor cohorts will be required to determine how consistently these cytotoxic effector-memory programmes are represented across the broader human CD1c-autoreactive T-cell compartment. We have revised the Discussion accordingly (revised manuscript, page 17, lines 403-410).

      (7) The authors should on how it might compare with non-autoreactive CD1c-restricted T-cells.

      We agree that it is important to place these findings in the context of non-autoreactive, antigen-specific CD1c-restricted T-cells. Previous studies have shown that CD1c can present microbial lipid antigens, such as mycobacterial lipids, to T-cells with defined antigen specificity. In contrast, the CD1c-autoreactive T-cells studied here recognise endogenous ligands and appear to respond to infection through changes in the CD1c-presented lipid repertoire rather than through recognition of a single defined foreign antigen.

      Our findings suggest that autoreactive CD1c-restricted T-cells may provide a complementary mode of immune surveillance, capable of sensing infection-induced changes in lipid presentation, whereas non-autoreactive CD1c-restricted T-cells may respond more directly to specific microbial lipid antigens. A head-to-head comparison would clearly be very interesting, but an extensive new piece of work beyond the scope to the current study. We have expanded the Discussion to more clearly highlight this distinction, and that direct comparison is required (revised manuscript, page 16-17, lines 380-410).

      Concluding remarks

      In summary, we have addressed the key concerns raised by the reviewers by strengthening validation of the experimental system, improving characterisation of T cell populations, and clarifying mechanistic interpretation. We believe these revisions significantly improve the clarity and rigour of the manuscript. We accept that identification of the CD1-presented lipids is an important next step that will give significant mechanistic insight, but is beyond the scope of the current work.

      References:

      (1) Wen, Q. et al. MiR-381-3p Regulates the Antigen-Presenting Capability of Dendritic Cells and Represses Antituberculosis Cellular Immune Responses by Targeting CD1c. J Immunol 197, 580-589 (2016). https://doi.org/10.4049/jimmunol.1500481

      (2) Gagliardi, M. C. et al. Bacillus Calmette-Guerin shares with virulent Mycobacterium tuberculosis the capacity to subvert monocyte differentiation into dendritic cell: implications for its efficacy as a vaccine preventing tuberculosis. Vaccine 22, 3848-3857 (2004). https://doi.org/10.1016/j.vaccine.2004.07.009

      (3) Grailer, J. et al. A Novel Cell-based Luciferase Reporter Platform for the Development and Characterization of T-Cell Redirecting Therapies and Vaccine Development. J Immunother 46, 96-106 (2023). https://doi.org/10.1097/CJI.0000000000000453

      (4) Guo, T. et al. A Subset of Human Autoreactive CD1c-Restricted T Cells Preferentially Expresses TRBV4-1(+) TCRs. J Immunol 200, 500-511 (2018). https://doi.org/10.4049/jimmunol.1700677

      (5) Bielecka, M. K. et al. A Bioengineered Three-Dimensional Cell Culture Platform Integrated with Microfluidics To Address Antimicrobial Resistance in Tuberculosis. mBio 8 (2017). https://doi.org/10.1128/mBio.02073-16

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study asks whether synapses formed by the same broad neuronal class (excitatory pyramidal neurons, PN) adapt their presynaptic organization in a cortex-specific manner, comparing the prefrontal cortex (PFC) with the primary somatosensory cortex (S1). The authors combine sophisticated electrophysiology (paired recordings and extracellular minimal stimulation), pharmacological perturbations of presynaptic Ca<sup>2+</sup>-secretion coupling, bouton Ca<sup>2+</sup> imaging, and mechanistic modeling. Across two prominent excitatory connections (Layer 5 (L5) PN-L5PN and L2/3-L5PN), they provide convergent evidence that mature PFC synapses operate with looser Ca<sup>2+</sup> channel-release sensor coupling than their S1 counterparts.

      Overall, the study provides an appealing mechanistic link between synaptic nano/micro-architecture and cortical-area specialization. The idea that PFC synapses retain a more "plasticity-favoring" presynaptic state, while the primary sensory cortex emphasizes reliability and timing precision, is potentially impactful for how we think about circuit computation and plasticity across cortical hierarchies.

      Strengths:

      A major strength is the multi-pronged experimental strategy. The paper first establishes robust, area-dependent differences in synaptic efficacy, reliability, timing, and short-term plasticity (facilitation prevailing in PFC versus depression in S1), using both paired recordings and minimal extracellular stimulation paradigms. The coupling interpretation is then directly supported by differential sensitivity to EGTA (and appropriate positive-control effects of fast chelators). Finally, volume-averaged calcium signals are reported to be similar across areas, arguing against trivial explanations based on gross differences in calcium influx, and the modeling provides a quantitative framework for interpreting the observed chelator effects.

      Weaknesses:

      Limitations are minor and concern interpretation/clarity rather than core results. Some key inferences rely on indirect readouts (chelator sensitivity, fluctuation analysis-derived parameters, bouton-averaged calcium signals), each of which carries assumptions and potential confounds that should be discussed more explicitly. In particular, the repatching paradigm for the paired-recording EGTA experiment, though very impressive, and the limited number of extracellular calcium conditions used for fluctuation analysis (three concentrations), can influence quantitative estimates and the confidence intervals around them.

      We would like to thank the reviewer for his/her overall positive assessment of our manuscript and the constructive advice, which helped us to improve our manuscript. We discussed the limitations, assumptions and potential confounding factors in more detail. We addressed them pointwise in the recommendations for the authors.

      Reviewer #2 (Public review):

      Schwarze et al. investigated whether synaptic efficacy is brain-region specific. To this end, they compared synaptic connections established by layer 5 (L5) neocortical pyramidal cells and between L5 and L2/3 pyramidal cells. In order to identify the mechanism of this brain region specificity, the authors employed several experimental approaches, including paired electrophysiological recordings, extracellular stimulation, low- and high-affinity intracellular calcium chelators (EGTA and BAPTA), multiple probability fluctuation analysis (MPFA), and intracellular measurements of calcium transients as well as computational modelling. The findings of the present study indicate that synaptic connections in the primary somatosensory cortex (S1) are significantly stronger and more reliable than those in the prefrontal cortex (PFC).

      The study is timely, and the topic is of significant interest to the neuroscience community. Despite the extensive research that has been carried out on the neuroanatomy and receptor distribution of different brain regions, comparatively little attention has been paid to differences in synaptic physiology. The authors' approach is characterised by its elegance and comprehensive nature, and the conclusions drawn are compelling. Nevertheless, there are a number of unresolved issues.

      First, we would like to thank the reviewer for his/her detailed survey of our work, which was very helpful in improving our manuscript. We are happy about the overall positive evaluation and the constructive comments. To fully clarify all points, we performed new experiments and analyses, in particular we determined EGTA sensitivity in PFC and S1 from the same animal and we performed MPFA with an additional extracellular Ca<sup>2+</sup> concentration. We extended the discussion on the examined cell types. Overall, we carefully revised the manuscript to address all points. Please see below our point-wise response.

      Major points:

      (1) The authors state that data from the S1 cortex were obtained in a previous study. In the context of an explicitly comparative study (PFC vs. S1cortex), it would have been advantageous for the authors to perform a subset of experiments in which both cortices were obtained from a single animal. This is a feasible undertaking, given the spatial separation of the PFC and S1 cortex.

      This is only true for the paired recordings from L5PN-L5PN connections in S1, which were obtained in a previous study and partially reanalyzed. All recordings from L2/3-L5PN connections in S1 and PFC as well as the paired recordings on L5PN-L5PN synapses in PFC were obtained in the present study. To make this clearer, we have added Table 1. This lists which data and associated figures are from this study and which are from previous studies (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019).

      Our experiments are lengthy and therefore it is challenging to achieve two successful recordings within the lifetime of acute brain slices. For this reason, the previous version of the manuscript did not include recordings from PFC and S1 of the same animal. We have now measured EGTA effects in L2/3-L5PNs from S1 and PFC of the same animal. Two new recordings were added to the EGTA-AM plots in Figure 2C-F and in the results section. Example recordings are shown in Figure S3A and B, as is the comparison of EGTA-AM effects in L2/3-L5PNs from PFC and S1 in Figure S3C, with data points from the same animal marked.

      “On the other hand, EGTA significantly reduced EPSC amplitudes only in PFC (0.54, 0.47-0.67; 55% of control) but not in S1 (0.85, 0.81-1.04; 100% of control). For a better comparison of EGTA effects some recordings were performed in PFC and S1 derived from the same animal to rule out interindividual effects (see example recording in Figure S3A-C).”

      (2) Figure 1A is somewhat misleading because it could suggest that the authors have performed dual recordings in identified PFC pyramidal cells.

      We thank the reviewier for this helpful note. We added “L2/3 or L5” to the stimulation panel of Figure 1A to illustrate that we stimulated either extracellularly in L2/3 or L5PNs directly via the patch pipette.

      (3) PFC and S1 cortex in rodents differ markedly in their morphological organisation. For example, in all sensory cortices, layer 4 is very pronounced; however, in the PFC of rodent,s no clear layer 4 can be found. On the other hand, PFC shows a clear separation of layers 2 and 3, which is not visible inthe S1 cortex. Furthermore, PFC pyramidal cells in layers 2, 3, and 5 exhibit significant heterogeneity, diverging considerably from those found in layers 5a and 5b of S1 cortex. Thus, there is no clear correlation between L5 pyramidal cells in the PFC and the S1 cortex. In order to achieve a meaningful comparison of the data obtained in PFC and S1 cortex, it is necessary for the authors to determine whether the record is from similar pyramidal cell populations.

      (3) In addition, PFC pyramidal cells in layer 2, 3 and 5 are highly heterogeneous and differ markedly from those in layer 5a and 5b of S1 cortex. To achieve a meaningful comparison of the data obtained in the PFC and the S1 cortex, the authors need to determine whether the record from similar pyramidal cell populations.

      We apologize for having not been precise about the specific location and type of pyramidal neurons in the original manuscript. Extracellular stimulation in PFC and S1 was always performed in layer 2, where the first large cell bodies, relative to the pia mater, are located within a cortical column. Therefore, we assume that the same cell populations were stimulated in both brain regions (van Aerde & Feldmeyer, Cereb. Cortex 2015; Oberlaender et al., Cereb. Cortex 2012; Lefort et al, Neuron 2009). To stick with the standard terminology, we refer to it in the manuscript as upper layer 2/3. We kept stimulation intensity as low as possible to ensure that only a few presynaptic cells within the target region were activated.

      In S1, paired recordings were obtained from pyramidal neurons in layer 5A following the procedures and criteria described in detail in our previous work (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019; Bornschein et al., Science 2025). Briefly, these criteria are as follows: close proximity to layer 4 and the barrels as well as the PPR of 0.78 (0.69-0.90), which is consistent with depression dominating in L5A (Frick et al., Cereb. Cortex 2008; Bornschein et al., Front. Syn. Neurosci. 2019) and different from L5B with PPR ≥ 1 (Lefort & Petersen, Cereb. Cortex 2017). Within layer 5A we did not attempt to further differentiate between pyramidal neuron types. For recordings in S1 with extracellular stimulation we focused on the same locations as for the paired recordings. We extended the corresponding section in Materials and Methods of the revised manuscript.

      We agree that there is no clear layer 4 in PFC, making the distinction between layer 2/3 and layer 5 less clear. Layer 2/3 and layer 5 have approximately the same diameter (van Aerde & Feldmeyer, Cereb. Cortex 2015). Based on this, we performed recordings in upper layer 5 of PFC. We did neither morphologically nor electrophysiologically differentiate between pyramidal neuron cell types. It should be noted that within a cortical area (S1 or PFC), we did not find a difference between glutamatergic synapses from L2/3 onto L5PNs and L5PN-to-L5PN synapses, neither with regard to the EGTA-sensitivity of release nor with regard to the release probability. In particular, we found homogeneous results and similar variability in both, the examined connections in the PFC and in S1, with no discernible clustering in the data that would suggest stimulation of different cell populations. These findings suggest that excitatory inputs to L5PNs exhibit similar properties (PPR, p<sub>N</sub>, CD) irrespective of whether they originate in L2/3 or in neighbouring PNs in L5A. However, we do see significant differences between synapses in the different cortical areas S1 and PFC. Thus, intra-area specific differences in morphology and spiking patterns among pyramidal neurons appear to not be reflected on the level of their synapses.

      The Reviewer probably refers to such differences and heterogeneity in morphology and spiking patterns of pyramidal neurons. If he/she has more specific differences in mind, it would be helpful if references for the significant heterogeneity could be given.

      Please also note that the type of experiments we perform with paired recordings and long-lasting patch-clamp measurements is not suitable for analyzing population differences among pyramidal neuron types.

      We refer to the problem of pyramidal neuron heterogeneity in the revised manuscript in the discussion.

      “Patch-clamp recordings from L5PNs located in the upper layer 5 (L5A in S1) were established according to the criteria described in detail in our previous work on this connection in S1 (Bornschein et al., 2019b; Bornschein et al., 2025). Presynaptic neurons were stimulated extracellularly in upper layer 2/3 (L2/3-L5PN connections) straight above the patched L5PN or in on-cell mode in L5A right next to the postsynaptic cell (L5PN-L5PN connections; Figure 1).“

      “We did neither morphologically nor based on spiking patterns differentiate further between PN subtypes within a given layer. However, within a cortical area (S1 or PFC) we did not find a difference between glutamatergic synapses from L2/3 onto L5PNs and L5PN to L5PN synapses, neither with regard to the EGTA sensitivity of release nor with regard to p<sub>N</sub>. In particular, we found homogeneous results and similar variability in both, the examined connections in the PFC and in S1, with no discernible clustering in the data that would indicate stimulation of different cell populations. These findings suggest that excitatory inputs to L5PNs exhibit similar properties (PPR, p<sub>N</sub>, CD) irrespective of whether they originate in L2/3 or in neighboring PNs in L5A. However, we do see significant differences between synapses in the different cortical areas S1 and PFC. Thus, intra-area specific differences in morphology and spiking patterns among PNs appear to be not reflected on the level of their synapses.“

      (4) For the S1 cortex, in rats it has been found that L5 synaptic connection between pairs of L5a pyramidal cells and pairs of L5b pyramidal cells differ markedly with respect to mean EPSP amplitude, latency and coefficient of variation (cv, a surrogate measure for the synaptic release probability) (cf. Markram et al., 1997; Frick et al., 2008). It is therefore likely that PFC and S1 pre- and postsynaptic pyramidal cells are not only morphologically and electrophysiological distinct but also with respect to their synaptic properties. At least, the authors need to discuss these confounding issues and preferentially address them experimentally. For example, it would be helpful to demonstrate that paired recordings were made from the same pyramidal cell types, perhaps by documenting their morphology and/or firing patterns. In addition, they should discuss the marked difference in EPSP amplitude and putative release probability between their data and the earlier studies.

      We agree that Markram et al. (J. Physiol. 1997) and Frick et al. (Cereb. Cortex 2008) provided highly valuable insights into synaptic transmission between pyramidal neurons in S1. We referred to their work in detail in our previous work on developmental changes in the presynaptic organization of transmitter release in L5APN synapses in S1 (Bornschein et al., Cell Rep. 2019). Both studies were performed in young rats and EPSPs were measured, whereas we worked in mice and recorded EPSCs. This impedes a direct comparison of amplitudes.

      Markram et al. (J. Physiol. 1997) recorded in 2-week-old rats from thick tufted PNs, corresponding to L5BPNs. Given the longer lifespan and slower development of rats compared to mice, this likely reflects a maturation state that corresponds better to our previous measurements in 8 to 10-day-old mice. Markram et al. found small failure rate (median 7%), which is similar to what we found in our previous study for young L5APN synapses (low failure rates and high p<sub>N</sub>; Bornschein et al., Cell Rep. 2019).

      The study by Frick et al. (Cereb. Cortex 2008) is closer to our present study and to the mature age window in our previous study, although they also recorded from rats but from L5APN-L5APN pairs in almost 3-week-old animals in S1. Again, EPSPs rather than EPSCs were recorded, impeding a direct comparison of amplitudes.

      Both studies concluded, based on the synaptic failure rate and CV analysis of EPSP amplitude, that the synapses they investigated operate with high release probability. This is fully in line with our findings. Of note, we found no significant difference between L2/3-L5PN and L5PN-L5PN synapses within a given area, indicating that varibality on the synaptic level between PNs of a given area is not pronounced.

      In order to further substantiate this, we determined the relative variability in median EPSC amplitudes to test whether there is a higher variability of recorded cell types in PFC compared to S1. The relative MAD (median absolute deviation) of EPSC amplitudes was 0.46 in PFC and 0.50 in S1. The similarity in these values argues against higher cell-type variability in PFC compared to S1. We have discussed the results of these studies in relation to our own findings.

      “Two other previous studies on L5APN (Frick et al., 2008) and L5BPN (Markram et al., 1997) connections concluded that these synapses operate with high release probability, which nicely agrees with our previous (Bornschein et al., 2019b) and current results. It is remarkable that we did not even detect any differences between the L2/3-L5PN and L5PN-L5PN synapses within a given cortical area. Overall these results from different studies (Markram et al., 1997; Reyes and Sakmann, 1999; Frick et al., 2008; Bornschein et al., 2019b; Bornschein et al., 2019a) may indicate that variability on the synaptic level between PNs of a given area is not pronounced. In order to further substantiate this, we determined the relative variability in median EPSC amplitudes to test whether there is a higher variability of recorded cell types in PFC compared to S1. The relative MAD (median absolute deviation) of EPSC amplitudes was 0.46 in PFC and 0.50 in S1. The similarity in these values argues against higher cell-type variability in PFC compared to S1.

      (5) In order to perform multiple probability fluctuation analysis (MPFA), a parabolic fit with a mere three points is inadequate, particularly because 2 mM and 5 mM Ca<sup>2+</sup> are close to the peak of the variance-to-mean parabola, and only 1 mM Ca<sup>2+</sup> is on its initial linear part. A more meaningful result would have been obtained with an additional Ca<sup>2+</sup> concentration between 1.0 and 2.0 mM, as these are closer to the physiological range. In this context, the authors should have quoted the more recent and more detailed paper by the Silver group (Saviane and Silver, 2006; Lanore and Silver, 2016) and not just the Clements and Silver review paper.

      We used only three Ca<sup>2+</sup> concentrations for MPFA as these resulted in a low (<0.5), a medium (~0.5) and a large (>0.5) p<sub>N</sub> condition, thereby clearly determining a parabola. Also Saviane and Silver (Nature 2006) performed MPFA with three extracellular Ca<sup>2+</sup> concentrations (1, 2, and 8 mM). We have now cited this paper, as well as the more recent work by Lanore and Silver (Neuromethods 2016), in relation to the MPFA method. To verify the reliability of MPFA, the determined parameters were compared with values estimated based on the EPSC amplitudes (EPSC = N p<sub>N</sub> q; PFC, 6 pA; S1, 48 pA) and failure rates (F = (1-p<sub>N</sub>)^N; PFC, 0.25, S1, 0.0001). The estimated values did in fact match those of the MPFA (EPSCs in PFC: 8 pA, 5-15 pA, and S1: 29 pA, 18-53 pA; failure rates in PFC: 0.16, 0.08-0.28, and S1: 0, 0-0.03; see original manuscript.

      To further support this, we have now conducted additional experiments using four Ca<sup>2+</sup> concentrations. The results are consistent with those from the experiments using three concentrations. The additional Ca<sup>2+</sup> concentration of 1.5 mM did not improve the parabolic fit, as it yielded p<sub>N</sub> values very close to those determined with 2 mM Ca<sup>2+</sup>. Therefore, the additional experiments are shown in Author response image 1. The novel p<sub>N</sub> data are included in the summary of p<sub>N</sub> values (now n=6) in the results section and in Figure 3F.

      Author response image 1.

      MPFA with four different extracellular Ca<sup>2+</sup> concentrations. (A) MPFA of EPSC amplitudes recorded at the indicated [Ca<sup>2+</sup>]<sub>e</sub> from L2/3-L5PNs in PFC. Top: Individual EPSCs (grey, average in black) recorded from L5PNs after extracellular stimulation in L2/3. Middle: Plot of EPSC amplitudes over time. Bottom: Corresponding mean-variance plot fitted with a parabola estimating the quantal parameters of release. p<sub>N</sub> is for 2 mM [Ca<sup>2+</sup>]<sub>e</sub>. Recordings were made in the presence of 10 µM Bicuculline, 0.25 mM Kynurenic acid and 50 µM 2-Amino-5-phosphonovaleriansäure. (B) As in (A), but for a L2/3-L5PN connection in S1. (C) Summary of determined p<sub>N</sub> values in PFC and S1 (dots represent individual experiments; P=0.485, Mann-Whitney U rank-sum test).

      Additionally, we performed bootstrap analyses with 10,000 replicates which were generated with replacement from the original data sample. Distributions of bootstrap 25% trimmed means showed a clear separation for N and q between PFC and S1 but not for p<sub>N</sub>. The bootstrap results are shown in Author response image 2.

      Author response image 2.

      Bootstrap analysis (A-C) Distribution of bootstrap 25% trimmed means of the quantal parameters p<sub>N</sub> (A), N (B) and q (C) in PFC (orange), S1 (blue) and S1 with gDGG (light blue). 10,000 bootstrap replicates were generated with replacement from the original data sample obtained by MPFA in L5PN-L5PN connections (cf. Figure 3D). (D) Same as in (A) but for MPFA in L2/3-L5PN connections.

      (6) Methods: The authors should clarify whether their paired recordings from L5 pyramidal cells involved whole-cell recordings from both pre- and postsynaptic neurons. From Figure 1B, it appears as if the presynaptic neurons were not recorded in whole cell mode but rather stimulated in cell-attached mode. This is also reflected in the artefact visible in the current trace recorded in the postsynaptic neuron. The authors should explicitly state their methodological approach and mention how reliable the timing of the presynaptic action potential was under these circumstances. The same holds true for the extracellular stimulation protocol. A significantly more detailed description of the experimental protocol is necessary here.

      In the paired recordings, presynaptic cells were stimulated in the cell-attached mode. For presynaptic EGTA application the whole-cell configuration was established after re-patching to allow buffer perfusion of the presynaptic L5PN. This is described in the methods section of the original version of the manuscript. We extended this description as follows:

      “In paired recordings, presynaptic L5PNs were stimulated in on-cell configuration (200-500 mV, 1-2 ms). In the chelator wash-in experiments, presynaptic neurons were repatched with a pipette solution supplemented with 10 mM EGTA (K-gluconate concentration was reduced to 135 mM to adjust osmolarity) and whole-cell configuration was established to allow EGTA perfusion of the presynaptic neuron.”

      The amplitudes were determined by fitting a product of two exponential functions to the baseline-subtracted currents, which allows for independent adjustment of the time constants of the rising and falling phases and minimizes noise effects (cf. Bornschein et al., J. Physiol. 2013). Synaptic delays were determined from the onset of stimulation to the fitted EPSC onset. We added this more detailed explanation to the methods section. in the timing of the presynaptic action potential was similar in recordings from PFC and S1. This applies to the paired-recordings with on-cell stimulation of presynaptic neurons as well as to the extracellular stimulation experiments.

      “Synaptic responses were determined by fitting a product of two exponential functions to the baseline-subtracted currents, which allows for independent adjustment of the time constants of the rising and falling phases and minimizes noise effects (Bornschein et al., 2013). Synaptic delays were determined from the onset of stimulation to the fitted onset of the EPSC. PPRs were calculated by dividing the second amplitude of two consecutive EPSCs by the first.”

      (7) Methods: The authors use Student's t-test for data comparison. The authors should verify that the data distribution was indeed normal, e.g. by using a Shapiro-Wilk test. If this is not the case, non-parametric tests should be used.

      We typically used non-parametric tests as stated in the figure legends of the corresponding figures. We have now explained the abbreviations for the Mann-Whitney U test (MWU) and the Wilcoxon signed-rank test (WSR) in the figure legends. A paired t-test was used only in Figure 5F after testing for normal distribution with the Shapiro-Wilk test. This is described in the methods section of the original version of the manuscript. Additionally, results of the Shapiro-Wilk test were now included in the figure legends.

      “Normality was tested using the Shapiro-Wilk test. Normally distributed data were compared with the t-test (two groups) or a one-way ANOVA (more than two groups). Non-normally distributed or small samples of data were compared with the Mann-Whitney U rank-sum test (MWU; two groups) or a Kruskal-Wallis ANOVA on ranks (more than two groups). (…). To compare pre- and post-treatment data the paired t-test or the Wilcoxon signed-rank test (WSR) was used, depending on the distribution of the data.”

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Max Schwarze and colleagues examined the coupling distance between presynaptic Ca<sup>2+</sup> channels and the vesicular release sensor at neocortical synapses in mice. They propose that Ca<sup>2+</sup> channel-release sensor coupling differs across cortical areas, with relatively loose (microdomain) coupling in prefrontal cortex (PFC) and tighter (nanodomain) coupling in primary somatosensory cortex (S1) for comparable pyramidal-neuron synapse types. To test this, they combine paired recordings and minimal stimulation with chelator manipulations (EGTA/BAPTA), mean-variance/MPFA-style analyses, presynaptic Ca<sup>2+</sup> imaging, and computational modeling. They conclude that presynaptic coupling organization is area-specific in the mature cortex and contributes to regional differences in synaptic timing, reliability, and short-term plasticity.

      Strengths:

      This study tackles an important question and is strengthened by a cohesive body of evidence assembled from multiple complementary approaches. A major asset is the inclusion of high-value datasets, particularly the paired recordings between L5 pyramidal neurons and the systematic assessment of EGTA sensitivity, which provide a solid functional foundation for the authors' central claims. The work is further distinguished by its genuinely multimodal design: combining electrophysiology with presynaptic calcium imaging (and integrating these observations with quantitative analyses and modeling) offers a more mechanistic view of neurotransmitter release than any single method could provide. Overall, the direct, within-framework comparison of presynaptic release-control mechanisms across cortical areas for comparable synapse types is compelling and gives the conclusions a level of robustness and interpretability that is often difficult to achieve in studies of cortical synaptic diversity.

      Weaknesses:

      Several aspects would benefit from clearer explanation, stronger integration with the existing literature, and a more explicit discussion of limitations and potential confounds. Without these additions, some conclusions remain speculative. Throughout the manuscript, the authors also often imply that different measurements reflect the same underlying synapse population. This is unlikely to be strictly true across all experiments and makes it difficult to integrate results from the various approaches into a single, unified set of functional synaptic properties. In addition, some statements-particularly those linking coupling mode to "higher-order neocortical functions"-appear broader than what is directly supported by the experiments and should be tempered or more precisely scoped.

      Below, I list several topics that could help better frame the main findings of the present study and clarify how it relates to previously published work.

      We would like to thank the reviewer for the comprehensive and detailed assessment of our manuscript and his/her overall positive evaluation. We have addressed all of the reviewer's points. We expanded the model description and discussion, and slightly toned down our conclusion.

      (1) The authors use EGTA sensitivity of EPSCs (together with additional metrics) to argue that S1 and PFC synapses differ in Ca<sup>2+</sup> channel-release sensor coupling. While this is a plausible interpretation, EGTA effects are not uniquely determined by coupling distance and can also reflect differences in Ca<sup>2+</sup> entry kinetics, action potential waveform, endogenous buffering/extrusion, or release-sensor/vesicle state. The authors use a constrained modeling approach, but the rationale for the different constraint sets is not fully clear from the current description. It would be helpful to expand and clarify the Methods section to explain how these constraints were defined, justified, and applied (and how alternative constraint choices would affect the results). In this context, the Abstract's broader claim that the study "reveals microdomain coupling as a presynaptic structure-function correlate of higher-order neocortical functions" appears overstated. Given the well-known diversity of cortical synapses even within a single region (e.g., synapses onto different interneuron subclasses or different PN cell types, extracortical sources like thalamus), the authors should clarify the intended scope: is the conclusion meant to apply broadly across synapse classes in S1 and PFC, or only to the specific connection type(s) examined here?

      We would like to thank the reviewer from pointing out that our description fell a bit short, in particular with respect to the interpretation of the EGTA effects. We addressed the points as follows in the revised manuscript: We discussed the interpretation of EGTA effects in more detail. We toned down the concluding statement in the last sentence of the Abstract.

      “Differences in the sensitivity of release to low to moderate concentrations of EGTA (≤ 30 mM) are a standard indicator of differences in the coupling distance (e.g. Adler et al., 1991; Bucurenciu et al., 2008; reviewed in Eggermann et al., 2012; Vyleta and Jonas, 2014; Kusch et al., 2018; Bornschein et al., 2019b). p<sub>N</sub> is determined by the size of the Ca<sup>2+</sup> signal at the release sensor and the binding kinetics and affinity of the sensor. The former in turn is determined by the details of the Ca<sup>2+</sup> influx and the diffusional coupling distance between the VGCCs and the sensor. The similarity of Ca<sup>2+</sup> signals between synapses in PFC and S1 (Figure 4) indicates that Ca<sup>2+</sup> influx is similar between boutons, although more subtle differences in the influx kinetics may have remained undetected in these volume-averaged signals. Regarding sensor affinity, results in a previous study indicate that differences in EGTA sensitivity show differences in coupling rather than sensor affinity even if k<sub>on</sub> of the sensor and its affinity should differ as much as ten-fold, which appears to be an unlikely scenario given that even the two major isoforms of Synaptotagmin that trigger synchronous release differ by less than a factor of three to four in their affinity (Bollmann et al., 2000; Schneggenburger and Neher, 2000; Bornschein et al., 2025). Finally, the increase in the PPR induced by the application of Cd<sup>2+</sup> further supports our conclusion of microdomain coupling in the PFC synapses (Scimemi and Diamond, 2012).”

      “They suggest that microdomain coupling in pyramidal neuron synapses could be a presynaptic structure-function correlate of higher order neocortical functions.”

      (2) The chelator logic is sound in principle, but the Discussion should more explicitly acknowledge standard caveats and alternative explanations. The authors partly address this by including presynaptic Ca<sup>2+</sup> imaging and modeling, yet it would help to explain more clearly how the combination of (i) chelator sensitivity, (ii) presynaptic Ca<sup>2+</sup> signals, and (iii) model constraints rules out-or substantially reduces the likelihood of-changes in AP waveform, Ca<sup>2+</sup> influx kinetics, buffering/extrusion, or sensor/vesicle state as the primary drivers. In addition, recent hypotheses emphasizing vesicle priming and/or release-site occupancy as contributors to apparent EGTA sensitivity should be discussed as a complementary or alternative interpretation.

      Please see above the first part of the discussion to point one.

      (3) A substantial portion of the S1 comparison appears to rely on previously published datasets. This should be made unambiguous in the Results and Methods, and it would be helpful to summarize this clearly (e.g., in a table indicating which figures/analyses use new data versus reanalysis of published data). If this information is already present, it should be highlighted more prominently.

      Please excuse us for not having made it clearer which data had already been published. Only the paired recordings from L5PN-L5PN connections in S1 were obtained in previous studies and partially reanalyzed. Paired recordings on the same synapses in PFC as well as all recordings from L2/3-L5PN connections in PFC and S1 were obtained in the present study. At your suggestion, we have added Table 1 highlighting which data and associated figures are from this study and which were acquired in previous studies (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019).

      (4) The modeling is informative, but the choice of a specific VGCC-release-site geometry and channel arrangement is not sufficiently justified. The manuscript adopts a particular spatial configuration, yet the rationale for selecting this geometry, rather than other plausible architectures discussed in the literature, is not clearly explained, nor is it meaningfully revisited in the Discussion. The authors should justify why the same organization is assumed across two distinct cortical areas and, ideally, include (or at a minimum discuss) a sensitivity analysis showing how key inferences (e.g., coupling distance and channel number) depend on the assumed geometry.

      We extended the discussion of why a ring-like structure of VGCCs was assumed in the model.

      “The microdomain was assumed to be formed by a ring-like structure of VGCCs around a vesicle (Figure 5D). This topography was chosen because such a microdomain was found to best predict the experimental data of transmitter release from PNs in young S1 (Bornschein et al., 2019b). Other previously described distributions of VGCCs suitable to reproduce release data cover random distributions of VGCCs (Scimemi and Diamond, 2012), VGCC clusters (Meinrenken et al., 2002; Nakamura et al., 2015), and exclusion zones (Keller et al., 2015). In the early S1, all of these models predicted a higher EGTA sensitivity of the microdomain, however, these models provided a poorer fit to the full set of the experimental data than the ring-like structure (Bornschein et al., 2019b). Since the experimental data from PNs in the mature PFC were similar to those in young S1, these other microdomain models were not tested explicitly here.”

      (5) The calcium imaging data are valuable, but given the diversity of synapses within each cortical layer, it is not clear that imaged boutons can be confidently assigned to the specific connection types being interrogated electrophysiologically. A substantial fraction of boutons likely corresponds to different postsynaptic targets (including interneurons and distinct pyramidal-cell classes), and this heterogeneity could complicate interpretation. This limitation should be discussed explicitly

      Excitatory pyramidal cells make up 80-85% of cortical neurons, with the highest density in layer 5 (Keller et al., Front. Neuroanat. 2018). In the somatosensory cortex, inhibitory synapses account for only about 10% (Santuy et al., Brain Struct. Funct. 2018). We imaged a large number of presynaptic boutons within layer 5 (about 10 boutons per cell, in total 85 boutons in PFC and 100 boutons in S1, numbers of boutons were now included in Figure 4). In this respect, the impact of inhibitory synapses is minor. Since connectivity between neighboring PNs in layer 5A is high (Feldmeyer, Front. Neuroanat. 2012), we assume that a large proportion of the imaged boutons target neighbouring L5PNs. We added a sentence on potential postsynaptic targets in the results section.

      “The imaged presynaptic boutons most likely connect to neighboring pyramidal cells, as connectivity between L5PNs in layer 5A is high (Feldmeyer, 2012). Nevertheless, a small proportion of other postsynaptic targets, such as interneurons, cannot be ruled out.”

      (6) In unitary connections, the authors assess EGTA effects alongside other functional parameters (strength, delay, short-term plasticity), which is a major strength. However, for L2/3 to L5 connections, it appears that EGTA sensitivity was tested primarily using extracellular stimulation. Given anatomical and circuit differences between PFC and S1, extracellular stimulation may recruit different synapse populations across regions, potentially confounding regional comparisons of EGTA sensitivity. This limitation should be acknowledged explicitly. While I am not requesting technically demanding L2/3↔L5 paired recordings in S1, the possibility that different synapse identities are being sampled should be treated as a meaningful source of uncertainty. The Discussion would also benefit from placing the magnitude of EGTA effects in the context of prior "loose coupling" literature, where comparatively large EGTA effects have been reported in some systems. In addition, the reported difference between adult PFC EGTA effects and S1 inhibition appears small (on the order of <10%) and should be interpreted cautiously, especially given that PFC and S1 mature on different timelines and P21-P26 is unlikely to reflect a mature PFC circuit state. The adult cohort (P90-P100) is therefore important, but the age mismatch complicates PFC-S1 comparisons; ideally, S1 should be assessed at matched ages, or this limitation should be discussed explicitly. Finally, for statistical robustness, in panel D of Figure 2, were the comparisons corrected for multiple testing to control Type I error?

      EGTA sensitivity was examined in PFC and in S1, for two connections in each region - using paired recordings for L5PN-L5PN connections and using extracellular stimulation for L2/3-L5PN connections. The L5PN-L5PN data from S1 were collected in an earlier study (Bornschein et al., Cell Rep. 2019; Bornschein et al., Front. Syn. Neurosci. 2019), have now been reanalyzed for the test period between 20 and 30 min, and included in Figure 2B for the sake of consistency (see Table 1). To emphasize this point, despite stimulating different input synapses with different stimulation methods, we obtained similar results in the respective brain regions. This suggests that the EGTA sensitivity observed in the investigated PFC connections is not a solely synapse-specific property.

      It is difficult to compare the absolute EGTA sensitivities from different synapses from different publications, since EGTA effects do not depend exclusively on the coupling distance, as the reviewer also noted in point 1. They are, among other factors, influenced by the Ca<sup>2+</sup> sensitivity of the release machinery, which differs between our Syt1-expressing cortical synapses and Syt2-expressing synapses in other brain regions (Schneggenburger et al., Nature, 2000; Bollmann et al., Science, 2000; Bornschein et al., Science 2025), such as the calyx of Held or the cerebellar basket to Purkinje cell synapse. Furthermore, direct patching and loading of the presynaptic bouton with EGTA - as feasible at the calyx of Held and other large synapses - results in higher effective EGTA concentrations compared to somatic loading of presynaptic terminals, despite identical pipette concentrations. The buffer-AM method introduces additional uncertainty regarding the effective intra-bouton EGTA concentration, since the loading efficacy has to be estimated. Thus, although differences in EGTA sensitivity primarily show differences in coupling distances, the comparison of absolute values between different publications is difficult. Consistently, data-constrained models are used to estimate the coupling topography and to compare these topographies rather than comparing the absolute EGTA effects (e.g. Buccurenciu et al., Neuron, 2008; Vyleta and Jonas, Science, 2014; Bornschein et al., Cell Rep., 2019; Chen et al., Neuron, 2024; Bornschein et al., Science, 2025).

      We include a note on this in the discussion.

      “Thus, although the absolute EGTA sensitivity is influenced by different factors, which necessitates data-constrained models for quantitative comparisons, the general sensitivity of release to EGTA indicates loose coupling.”

      In mouse neocortex postnatal maturation in S1 and PFC follows the same time course. Kroon et al. (Sci. Rep. 2019) reported that maturation of dendritic morphology and intrinsic properties of pyramidal neurons occurs within the first two weeks after birth, now cited in the discussion. Therefore, it is unlikely that the EGTA effect in PFC is due to a delayed maturation. The difference in EGTA sensitivity between PFC at P90-100 and S1 at P21-26 is indeed small but significant (P=0.009, Mann-Whitney-U rank sum test).

      “Since postnatal development follows the same time-course in mouse PFC and S1 and occurs predominantly within the first two weeks after birth, (…) (Kroon et al., 2019).”

      Thank you for the advice concerning statistical robustness. We replaced the Mann-Whitney-U rank sum test in Figure 2D by a one-way ANOVA and performed a Holm-Sidak post-hoc test correcting for multiple comparisons. Similarly, ANOVA was used to compare more than two groups in Figures 1K and 2F. We changed the corresponding P values and tests in the figure legends and added the performed post-hoc tests in the methods section.

      “For multiple comparisons post-hoc testing was performed with the Holm-Sidak (one-way ANOVA) or Dunn´s method (ANOVA on ranks).”

      (7) Alterations in initial release probability are often associated with changes in short-term plasticity. In the present manuscript, the authors report similar initial release probability at PFC and S1 synapses, yet observe differences in short-term plasticity profiles. The mechanistic basis for this apparent dissociation is not addressed and should be discussed explicitly, including potential explanations.

      Various other factors besides p<sub>N</sub> can influence short-term plasticity, that are the coupling distance, the number of occupied release sites (N<sub>occ</sub>), the replenishment of N<sub>occ</sub> or the recruitment of newly formed N<sub>occ</sub> as well as the expression of endogenous Ca<sup>2+</sup> buffers (Blatow et al., Neuron 2003; Felmy et al., Neuron 2003; Matveev et al., Biophys. J. 2004; Neher, Cell Calcium 1998; Regehr, CSH Perp. Biol. 2012) or fascilitation sensors (Turecek & Regehr, J. Neurosci. 2018; Shin et al., eLife 2025). Traditionally, p<sub>N</sub> had been assumed to have a major impact on short-term plasticity (STP) which is indeed the case at low replenishment rates (e.g. Feldmeyer and Radnikow, J. Physiol. 2009; Zucker and Regehr, Ann. Rev. Physiol. 2002). But at several synapses very fast replenishment rates have been described driving a progressive overfilling of the initial RRP and increasing N<sub>occ</sub> above baseline levels (Brachtendorf et al., Front. Cell. Neurosci. 2015; Doussau et al., eLife 2017; Miki et al., Neuron 2016; Valera et al., J. Neurosci. 2012) making replenishment the stronger determinant of STP.

      Additionally, the size and organization of sub-pools from which vesicle recruitment and release occurs affects the speed and reliability of vesicular release. In our previous study on L5PN-L5PN connections in S1 we found that developmental tightening of CDs was associated with an increase in PPR without altering p<sub>N</sub> (Bornschein et al., Cell Rep. 2019). We could show that the maturation of a replenishment pool during postnatal development increases vesicle recruitment and reliability thereby affecting STP (Bornschein et al., Front. Syn. Neurosci. 2019).

      We discussed this in the revised manuscript.

      “Classically, p<sub>N</sub> was considered as the major determinant of short-term plasticity (e.g. reviewed in Zucker and Regehr, 2002; Feldmeyer and Radnikow, 2009). More recently other factors, including the number of occupied release sites, their replenishment or an increase in their occupancy, or the expression of endogenous Ca<sup>2+</sup> buffers have been considered as more important determinants of short-term plasticity (Rozov et al., 2001; Blatow et al., 2003; Felmy et al., 2003; Matveev et al., 2004; Bornschein et al., 2013; Miki et al., 2016; Doussau et al., 2017; Jackman and Regehr, 2017; Neher and Brose, 2018). (…)”

      Short-term plasticity changes during postnatal development at different cortical PN connections without alterations in p<sub>N</sub> (Reyes and Sakmann, 1999; Bornschein et al., 2019a). For L5PN-to-L5PN connections these differences were found to result from the maturation of an intermediate replenishment vesicle pool (Bornschein et al., 2019).

      (8) There are multiple instances where the text appears to cite non-existent or misnumbered figure panels (e.g., references to "Figure 4G-I / 4J" when the relevant material appears elsewhere). These should be corrected throughout, as they currently reduce readability and confidence.

      We apologize for the misnumbering which originated from a previous version of this manuscript. Figure 4G-J is actually Figure 5A-D. We corrected the references to Figure 5 in the methods section.

      (9) The Methods describe P21-P26 animals, whereas the Results include older cohorts (e.g., P90-P100) and additional regions (e.g., mPFC). The Methods should be updated so that all cohorts and regions analyzed in the Results are fully described.

      Thank you for thoroughly reading the methods. We added the missing cohort (P90-100) and brain region (mPFC) to the method section.

      “C57BL/6J mice at P21-26 and P90-100 of either sex were decapitated under deep Isoflurane (Curamed) inhalation anaesthesia. (…) Coronal neocortical slices (150-250 μm thick) were cut from the lateral PFC, medial PFC (mPFC) or S1 region (Figure 1A) with a vibratome (HM 650 V, Microm).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      These are mostly points for discussion; there is no need for additional experiments.

      (1) Discuss potential effects of (re-)patching on presynaptic physiology and the "control" time course in Figure 2A.

      The repatching strategy is a strength, but it also introduces opportunities for physiological drift (dialysis effects, changes in access resistance, altered excitability, and the switch to presynaptic stimulation in whole-cell after repatching). Please discuss (and, if already quantified, briefly report) why EPSC amplitudes appear to increase in the control time course after repatching (as seen in Figure 2A). Even a short explanation (e.g., run-up after whole-cell access, recovery from on-cell stimulation, washout of endogenous buffering, improved spike waveform reliability, etc.) plus reassurance that baseline stationarity criteria were met would strengthen confidence in the repatching-based inference.

      Baseline recordings were usually performed with the presynaptic neuron in cell-attached mode. In this configuration intracellular ion concentration, second messenger systems as well as cell-specific resting membrane potential stay essentially unaffected. When switching to whole-cell mode after re-patching of the presynaptic cell, the pipette solution determines the intracellular environment, which may also affect second messenger systems and mobile endogenous buffers will be washed out. Although layer 5 pyramidal neurons do not express relevant concentrations of these mobile buffers (Helmchen et al., Biophys. J., 1996; Tran and Stricker, Biophys. J., 2018; Bornschein et al., Cell Rep., 2019), this might have contributed to the moderate and temporary run-up after whole-cell access to presynaptic cells.

      In postsynaptic neurons, we also routinely controlled the stability of R<sub>s</sub> and I<sub>leak</sub> during our prolonged measurements. In the controls, we found a median initial increase in relative EPSC amplitudes to 1.17 (1.02-1.38) between 0 and 10 min after the presynaptic whole-cell access was established, which correlated with a temporary decline in R<sub>s</sub> in 3 out of 6 recordings. EPSC amplitudes returned to their baseline values after 20 min at the latest (1.02, 0.75-1.20). We added potential reasons for the temporary amplitude increase in the controls in the results section.

      “In control recordings a temporary initial increase in EPSC amplitudes after whole-cell access to the presynaptic neuron was evident. Although L5PNs do not express large concentrations of mobile buffers (Helmchen et al., 1996; Tran and Stricker, 2018; Bornschein et al., 2019b), their wash-out might have contributed to the temporary run-up. On the other hand, run-up correlated with a temporary decline in R<sub>s</sub> in some recordings (3 out of 6). Run-up effects normalized after 20 min at the latest (1.02, 0.75-1.20).”

      (2) Clarify interpretation/robustness of MPFA-derived "N" given the large range, and discuss uncertainty from using only three [Ca<sup>2+</sup>]e conditions for L2/3-L5PN MPFA.

      The manuscript reports that "N" is markedly smaller in PFC than S1 (median ~2.1 vs 8), but the S1 range is very broad (3-19).

      (a) Briefly discuss whether/why such a wide N range is expected and how it should be interpreted (binomial "N" vs anatomical release sites; sensitivity to CV assumptions; potential dependence on connection geometry, bouton number, dendritic filtering, etc.).

      (b) Add a short statement on uncertainty/identifiability when fitting MPFA with only three conditions for L2/3-L5PN (e.g., whether confidence intervals/bootstraps were examined; how stable q and N are to small changes in the variance estimates). Even a qualitative note would help readers judge how much weight to put on the absolute N estimates versus the overall cross-area trend.

      (a) In our previous study on this connection we determined a wide range for N (8, 3-19) even though we used four extracellular Ca<sup>2+</sup> concentrations in MPFA (Bornschein et al., Cell Rep. 2019). The binomial parameter N can be considered to represent the number of release sites, including empty release sites (see Brachtendorf et al., Front. Cell. Neurosci. 2025). The range of release sites is likely to reflect the variability in the number of anatomical synaptic contacts ranging from 1 to 6 for these synapses (Frick et al., Cereb. Cortex 2008). Since 1 to 3 active zones/release sites per synaptic contact appear to be typical for small cortical synapses (e.g. Xu-Friedman et al., J. Neurosci. 2001), this results in a wide range of 1-18 release sites per connection.

      In our previous study (Bornschein et al., Cell Rep. 2019), we also investigated the effects of different values of CV1 and CV2 by repeating the MPFA fitting procedures for different combinations of CV1 and CV2 ranging from 0.1 to 1 each. We have quantified a deviation of ≤10% in the estimates of vesicular release probability from the typically used CV values of 0.3 across a wide range of CV value combinations (see also Schmidt et al., Curr. Biol. 2013). We have added a note regarding CV sensitivity in the methods section.

      “For CV assumptions that deviate from the standard value of 0.3, deviations in the calculated p<sub>N</sub> values of less than 10% are to be expected (Schmidt et al., 2013; Bornschein et al., 2019b).”

      (b) The three Ca<sup>2+</sup> concentrations we used for MPFA resulted in a low (<0.5), a medium (~0.5) and a large (>0.5) p<sub>N</sub> condition. With this, a parabola is uniquely determined by three parameters. To further ensure the reliability of the parameters determined by MPFA, we compared them to values estimated from EPSC amplitudes (EPSC = N p<sub>N</sub> q; PFC, 6 pA; S1, 48 pA) and failure rates (F = (1-p<sub>N</sub>)^N; PFC, 0.25, S1, 0.0001), which yielded values similar to those from MPFA (EPSCs in PFC: 8 pA, 5-15 pA, and S1: 29 pA, 18-53 pA; failure rates in PFC: 0.16, 0.08-0.28, and S1: 0, 0-0.03; see original manuscript).

      To further support this, we have now conducted additional experiments using four Ca<sup>2+</sup> concentrations. The results are consistent with those from the experiments using three concentrations. The additional experiments are shown for review purposes in Figure R1. The novel p<sub>N</sub> data are included in the summary of p<sub>N</sub> values (now n=6) in the results section and in Figure 3F.

      Additionally, we performed bootstrap analyses with 10,000 replicates which were generated with replacement from the original data sample. Distributions of bootstrap 25% trimmed means showed a clear separation for N and q between PFC and S1 but not for p<sub>N</sub>. The bootstrap results are shown for review purposes in Author response image 2 (cf. point 5 of Reviewer#2).

      (3) Broaden the discussion, e.g. by linking to nanodomain/microdomain coupling as a general strategy for stimulus encoding, including sensory periphery examples.

      The work will resonate beyond the cortex if the authors explicitly connect their findings to broader principles: how the spatial coupling regime shapes the transfer function between Ca<sup>2+</sup> entry and vesicle fusion, thereby tuning reliability, timing, and dynamic range. Requested addition: Please consider adding a short subsection discussing analogous implementations in the sensory periphery, especially ribbon synapses of cochlear inner hair cells and rod photoreceptors, where nanodomain coupling has been discussed as a key determinant of encoding and release dynamics. Also, citing relevant work such as that by Scimemi and Diamond 2012 would further strengthen the paper.

      We agree that the work by Scimemi and Diamond (J.Neurosci. 2012) is important and we cited and discussed their work in several of our previous publications. We now also included the paper in the revised version of the present manuscript. As requested, we also included a discussion on findings from ribbon type synapses and also from the neuromuscular junction.

      “Nanodomain coupling was also found in the peripheral nervous system, in particular at retinal (Singer and Diamond, 2003; Jarsky et al., 2010) and auditory (Moser and Beutner, 2000; Brandt et al., 2005) ribbon-type synapses and at the neuromuscular junction (Harlow et al., 2001; Shahrezaei et al., 2006). These synapses have highly specialized properties and appear to be optimized for very reliable transmission and, in the case of ribbon synapses, also for high-frequency coding of sensory information (reviewed in Matthews and Fuchs, 2010; Eggermann et al., 2012). Thus, it appears that synapses in the sensory pathways, in particular those engaged in reliable high-frequency coding of sensory information, both in the periphery and in the lower processing stages of the CNS, up to primary sensory cortices, operate with nanodomain coupling. In the executing motor pathway, the neuromuscular junction uses nanodomain coupling and, as recent results from our group suggest, also PNs in the primary motor cortex (Yarim et al., in preparation). It is tempting to speculate that complete loops from or to the primary cortices to their peripheral target organs operate with nanodomains. Microdomain coupling, on the other hand, appears to come into play only if integration of information from multiple sources and plasticity are the main focus, as at certain synapses in PFC (this study) or hippocampus (Vyleta and Jonas, 2014).”

      “The microdomain was assumed to be formed by a ring-like structure of VGCCs around a vesicle (Figure 5D). This topography was chosen because such a microdomain was found to best predict the experimental data of transmitter release from PNs in young S1 (Bornschein et al., 2019b). Other previously described distributions of VGCCs suitable to reproduce release data cover random distributions of VGCCs (Scimemi and Diamond, 2012), VGCC clusters (Meinrenken et al., 2002; Nakamura et al., 2015; Rebola et al., 2019), and exclusion zones (Keller et al., 2015; Rebola et al., 2019). In the early S1, all of these models predicted a higher EGTA sensitivity of the microdomain, however, these models provided a poorer fit to the full set of the experimental data than the ring-like structure (Bornschein et al., 2019b). Since the experimental data from PNs in the mature PFC were similar to those in young S1, these other microdomain models were not tested explicitly here.”

      “(…) Finally, the increase in the PPR induced by the application of Cd<sup>2+</sup> further supports our conclusion of microdomain coupling in the PFC synapses (Scimemi and Diamond, 2012).”

      (4) Address limitations of basal/resting Ca<sup>2+</sup> estimates and make explicit that measured Ca<sup>2+</sup> signals are volume-averaged (not microdomain) readouts.

      (a) The reported basal [Ca<sup>2+</sup>]i values are in the ~tens of nM range. Given the stated in vitro KD for Fluo-5F in the authors' pipette solution (439 nM), the resting estimates are far below KD; this does not invalidate the approach, but it does warrant a brief discussion of sensitivity/uncertainty (influence of Rmin estimation, background subtraction, and how errors propagate into basal [Ca<sup>2+</sup>]i). Repeating experiments is not necessary-just clearer framing of limitations.

      (b) Please also emphasize more prominently (ideally in Results and/or Discussion) that the bouton signals are volume averaged and therefore do not directly report calcium microdomains at active zones or nanodomains at release sensors. The Methods already state this point; echoing it in the main text would prevent over-interpretation by readers.

      (a) We agree that Fluo5F is less suitable for determining absolute basal calcium levels. In a previous study (Bornschein et al., Science 2025) we determined the basal Ca<sup>2+</sup> concentration with OGB1 (K<sub>D</sub>=166 nM; basal [Ca<sup>2+</sup>]<sub>i</sub>=44 nM, 24-58 nM, n=43 boutons from 10 cells) and observed no significant difference to basal [Ca<sup>2+</sup>]<sub>i</sub> values determined with Fluo5F despite the K<sub>D</sub> of 439 nM (31 nM, 16-54 nM, 14 boutons from 3 cells; P=0.204, MWU; data not published). We added this limitation to the results section and swapped Figure panels 4E and F for confluence. The calibration curve of Fluo5F as well as the comparison to basal [Ca<sup>2+</sup>]<sub>i</sub> values determined with OGB1have been included in Figure S4.

      “The quantification of absolute basal [Ca<sup>2+</sup>]<sub>i</sub> was limited by the K<sub>D</sub> of Fluo5F (439 nM), which slightly underestimated basal [Ca<sup>2+</sup>]<sub>i</sub> values in comparison to quantification with OGB1 (K<sub>D</sub>=166 nM, Figure S4). Nevertheless, relative comparison of basal [Ca<sup>2+</sup>]<sub>i</sub> yielded no significant differences between PFC (30 nM, 21-34 nM) and S1 (22 nM, 13-38 nM; Figure 4F).”

      (b) In the results section, we have now emphasized that volume-averaged Ca<sup>2+</sup> signals were measured.

      “We performed dual-dye two-photon Ca<sup>2+</sup> imaging (Sabatini et al., 2002) to quantify volume-averaged Ca<sup>2+</sup> signals at presumed presynaptic boutons located on axon collaterals of L5PNs in PFC and in S1.”

      Reviewer #2 (Recommendations for the authors):

      (1) For a meaningful comparison, recordings from the PFC and the S1 cortex of the same animals should be undertaken. Additionally, I suggest performing additional experiments regarding the different cell types of L5 pyramidal cells in layer 5a.

      We performed new experiments to determine EGTA sensitivity in PFC and S1 from the same animal. The results from these experiments agree with the previous results. They are included in the results section , in Figure 2C-F and in Figure S3A-C.

      Additionally, we extended the discussion on the examined cell types. For S1 cortex we refer in more detail to our previous work, where we described in depth where and under consideration of which criteria our recordings were established and that based on these criteria we recorded from pyramidal neurons in layer 5A in S1 (Bornschein et al., Cell Rep. 2019; Bornschein et al. Front. Synapt. Neurosci. 2019; Bornschein et al., Science 2025). Within layer 5A, we did not attempt to further distinguish between types of pyramidal neurons. We include this in the methods section.

      “Patch-clamp recordings from L5PNs located in the upper layer 5 (L5A in S1) were established according to the criteria described in detail in our previous work on this connection in S1 (Bornschein et al., 2019b; Bornschein et al., 2025). Presynaptic neurons were stimulated extracellularly in upper layer 2/3 (L2/3-L5PN connections) straight above the patched L5PN or in on-cell mode in L5A right next to the postsynaptic cell (L5PN-L5PN connections; Figure 1).”

      For the recordings in PFC and heterogeneity in pyramidal neuron types we refer to our detailed response to the point 3 of Reviewer 2. There we also discuss that the heterogeneity in morphology and spiking patterns is probably not reflected on the synaptic level. We would also like to emphasize that the type of experiments we perform with paired recordings and long-lasting patch-clamp measurements is not suitable to differentiate between subpopulations of pyramidal neurons. This would require successful recordings form several tens of different pyramidal neurons, which is not feasible in our type of experiment. We discuss this limitation of the discussion.

      (2) The authors need to comment in depth on their MPFA data, and if feasibl,e perform additional experiments.

      Concerning the robustness of quantification of synaptic parameters by MPFA, we refer to our comments on point 2b of the recommendations for the authors to Reviewer 1. Additionally, we performed new MPFA experiments with four extracellular Ca<sup>2+</sup> concentrations that agree with our results with three Ca<sup>2+</sup> concentrations.

      (3) The statistical analysis should be revised and a test for the normality of data distribution should be implemented.

      A test for normal distribution (Shapiro-Wilk test) has already been described in the methods section in the previous version of this manuscript.

      (3) Figure 1A is somewhat misleading because it could suggest that the authors have performed dual recordings in identified PFC pyramidal cells.

      We added “L2/3 or L5” to the stimulation panel of Figure 1A to illustrate that we stimulated either extracellularly in L2/3 or L5PNs directly via the patch pipette.

      (4) Is the relative variance of the mean EPSC amplitude and latency between connections larger in the PFC connections than in S1 cortex? This could indicate a variability in cell types.

      The relative variance of EPSC amplitudes calculated as median absolute deviation (MAD) was 0.46 in PFC and 0.50 in S1 arguing against differences in the variability in cell types. The larger variability in delays expressed as SD<sub>Delay</sub> is the result of the larger coupling distance in PFC compared to S1 (Bullmann et al., J. Neurosci. 2024). Consequently, also the relative MAD is larger (0.89) in PFC compared to S1 (0.14) and is therefore not able to detect differences in the variability of recorded cell types.

      (5) Reyes and Sakmann (1999) have previously described differences for L2/3-L5b and L5b-L5b synaptic connections in S1 cortex at different developmental stages. This paper needs to be cited as it is highly relevant to this study.

      Reyes and Sakmann (J. Neurosci. 1999) reported layer-specific differences in short-term plasticity in young sensorimotor cortex which disappeared as maturation progressed and short-term plasticity increased. In a previous study (Bornschein et al., Front. Syn. Neurosci. 2019) we also described a developmentally driven increase in short-term plasticity caused by the maturation of vesicle pools. In the present study we used mature animals and would therefore not expect layer-specific differences neither in S1 nor in PFC since the time course of postnatal maturation was described to be comparable in both neocortical circuits (Kroon et al., Sci. Rep. 2019).

      We discussed this paper in the context of developmental changes in short-term plasticity.

      “Short-term plasticity changes during postnatal development at different cortical PN connections without alterations in p<sub>N</sub> (Reyes and Sakmann, 1999; Bornschein et al., 2019a). For L5PN-to-L5PN connections these differences were found to result from the maturation of an intermediate replenishment vesicle pool (Bornschein et al., 2019a). Such pool maturation may also underlie the elimination of layer-specific differences in short-term plasticity between L2/3-L5B and L5B-L5B synaptic connections that were evident in young rats but eliminated during the first weeks of postnatal development (Reyes and Sakmann, 1999). Since postnatal development follows the same time course in mouse PFC and S1 and occurs predominantly within the first two weeks after birth, significant layer-specific differences in PN synapses are unlikely in both areas in our experimental time window (Kroon et al., 2019). Consistently, we found similar PPRs at L2/3-L5PN synapses and L5PN-L5PN synapses in both areas, with facilitation in PFC and depression in S1, irrespective of the presynaptic PN synapse type.

      (6) Regarding the point of loose or tight Ca<sup>2+</sup> channel coupling: Could some of the differences result from differences in the presynaptic Ca<sup>2+</sup> channel complement? Please comment.

      This can be excluded. Ca<sub>v</sub>2.1 and Ca<sub>v</sub>2.2 are the main channels gating release at PN synapses. The gating kinetics of these channels are very similar and they only differ somewhat in their peak current amplitude (Bornschein et al., Cell Rep. 2019, Figure 4). Since the number of open channels is a fit parameter in our simulations there would only be an effect on the estimate of the number of channels gating release but not for the estimate of the coupling distance. This is all the more true since the EGTA effect depends on the diffusion distance rather than on the gating kinetics. These considerations will also hold for Ca<sub>v</sub>2.3 channels, which have slower closing kinetics, but anyway play only a very minor role for triggering release.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study is a valuable contribution to the evidence base. However, the evidence provided is incomplete as the study results only partially support the study conclusions. Addressing the methodological and reporting issues raised by the peer reviewers and properly aligning the claim made for providing a tool for early warning with the study analysis/results would improve the study quality and usefulness of its findings.

      We are deeply encouraged by the editors’ recognition of this study as a valuable contribution to the evidence base. We fully concur with the eLife assessment that the manuscript, in its original form, required substantive methodological recalibration and rhetorical refinement to ensure the claims were strictly supported by the data. We accept these profound critiques unreservedly and have undertaken a comprehensive, ground-up revision of this study.

      This revision was guided by three overarching principles: (1) systematic decoupling of epidemiological confounders, (2) rigorous elimination of COVID-19-era surveillance bias, and (3) precise alignment of our claims with the actual scope of our predictive framework. Most importantly, we have fundamentally recalibrated our core claim: we have entirely removed all assertions of providing a “ready-to-use early warning tool.” Instead, we accurately reposition our contribution as developing a “robust, climate-informed predictive surveillance framework” that establishes an evidence-based foundation for understanding non-stationary climate-influenza dynamics in a subtropical urban setting.

      To ensure our analytical results definitively support this recalibrated conclusion, we executed a comprehensive remodeling of our entire dataset, rebuilding both the DLNM and LSTM networks. The major structural upgrades include:

      (1) Shift to Positivity Rate to Eliminate Testing Bias: To directly address the reviewers’ incisive concern regarding testing volume bias (i.e., raw case counts were artificially inflated or suppressed by massive fluctuations in PCR testing during the COVID-19 pandemic), we have fundamentally replaced “influenza positive case counts” with the “influenza test positivity rates” as our primary outcome metric. This relative metric mathematically standardizes the denominator, elegantly isolating intrinsic viral transmissibility from the severe artificial fluctuations of healthcare-seeking behaviors and diagnostic intensity.

      (2) Integration of New Epidemiological Covariates: We significantly enhanced our control for critical confounders (as suggested by the reviewers) by incorporating new, highly relevant covariates into our models. These include:

      - Weekly detection volumes: To capture residual testing capacity fluctuations.

      - Mask-wearing stringency indices (OWID): To computationally decouple the artificial suppression of cases caused by strict non-pharmaceutical interventions (NPIs).

      - Proportion of the non-local/transient population: To explicitly account for external viral importation risks.

      - Day of the Week (DOW, weekday vs. weekend): To reflect societal mobility and administrative patterns.

      Simultaneously, we removed redundant variables (e.g., raw COVID-19 case numbers) to minimize noise.

      (3) Comprehensive Re-optimization, Retraining, and Sensitivity Analyses: Following this massive feature engineering, we re-executed the Bayesian hyperparameter tuning process and completely retrained the LSTM networks. We also conducted rigorous sensitivity analyses (detailed in the new Table 2 and Supplementary files) by systematically ablating variables like testing volume and mask-wearing indices, definitively proving the necessity of these covariates in our framework. All corresponding codes, main figures (Figures 1– 6), supplementary figures, and statistical tables have been entirely updated.

      (4) Streamlining the Manuscript: We acknowledge the reviewers’ observation that the original methodology was overly meandering. In this revision, we have ruthlessly streamlined the text. We condensed routine laboratory protocols, reorganized the Methods to logically present the “Study area” prior to the “Study design,” and thoroughly rewrote the Discussion to weave our limitations directly into the interpretation of the results. This ensures conciseness, logical flow, and a sharp focus on the central narrative, tailored for the broad and rigorous readership of eLife.

      We believe these deep methodological revisions have profoundly elevated the scientific rigor, interpretability, and integrity of our study. Below, we provide detailed, point-by-point responses outlining how each specific comment was addressed.

      Reviewer #1 (Public review):

      A major concern is that the model is trained in the midst of the COVID-19 pandemic and its associated restrictions and validated on 2023 data. The situation before, during, and after COVID is fluid, and one may not be representative of the other. The situation in 2023 may also not have been normal and reflective of 2024 onward, both in terms of the amount of testing (and positives) and measures taken to prevent the spread of these types of infections. A further worry is that the retrospective prospective split occurred in October 2020, right in the first year of COVID, so it will be impossible to compare both cohorts to assess whether grouping them is sensible.

      We deeply appreciate this astute epidemiological critique. You have precisely identified the most formidable methodological challenge in pandemic-era time series modeling: the profound non-stationarity of influenza dynamics spanning the 2018–2023 timeline, driven by intensive Non-Pharmaceutical Interventions (NPIs), pandemic-related behavioral shifts, volatile testing volumes, and subsequent immunity debt. We fully agree that 2023 represented an atypical post-restriction “rebound” year and is unlikely to be straightforwardly representative of a stabilized post-2024 epidemiological steady state.

      First, we wish to clarify a minor but crucial methodological detail regarding the October 2020 split. You understandably raised concerns about comparing “both cohorts.” We must emphasize that there is no change in the cohort or the data collection methodology. The surveillance system, sentinel hospitals, and diagnostic protocols (managed by Putian CDC) remained identical and uninterrupted from 2018 to 2023. Although the original manuscript labelled the January 2018 – October 2020 and October 2020 – December 2023 phases as “retrospective” and “prospective” cohorts respectively, this distinction reflected only the administrative timing of ethical approval. It does not indicate any changes in sentinel hospital locations, ILI case definitions, specimen collection procedures, RT-PCR diagnostic protocols, or laboratory quality control. The patient population and clinical criteria are completely homogeneous. We have revised the “Ethics statement” subsection to eliminate this semantic confusion.

      Upon receiving this insightful feedback, our team conducted extensive mathematical evaluation on how best to address this timeline heterogeneity. We initially considered formal stratified temporal analyses (e.g., splitting models into strict Pre-COVID 2018-2019, COVID-disruption 2020-2022, and Post-COVID 2023 periods) and implementing a rolling-window validation scheme. However, after careful evaluation, we concluded that strict time-slicing of this particular dataset would introduce its own substantial methodological problems more severe than the heterogeneity it sought to resolve:

      (1) Statistical power constraints on for stratified Non-linear Lag analysis: The DLNM architecture requires continuous, robust longitudinal data to stably estimate two-dimensional exposure–lag–response surfaces with natural cubic splines, which typically demand approximately 16–25 effective degrees of freedom across the joint exposure × lag space. Stratifying our six-year dataset into the three intervals list above would yield:

      - Pre-COVID stratum (2018–2019): ~730 daily observations and only ~200 influenza B positive events, insufficient to stably identify cross-basis surfaces, with confidence intervals expected to widen to non-informative ranges.

      - COVID-disruption stratum (2020–2022): influenza circulation was substantially suppressed (though, importantly, not eliminated, a point we return to below), reducing signal density and risking that estimated surfaces reflect suppression dynamics rather than climate–transmission relationships.

      - Post-restriction stratum (2023): a single year cannot independently support a DLNM with meaningful lag structure given the typical 0–14-day lag window we examine.

      (2) The “training-domain contamination” problem in rolling-window validation: We carefully examined whether expanding-window rolling validation (e.g., train 2018–2020, validate 2021; train 2018–2021, validate 2022; …) would resolve the non-stationarity concern. We concluded it would not, for a subtle but consequential reason: each such window straddles the abrupt NPI transitions, meaning that within any single training window, the model is exposed to a nonstationary mixture of regimes without any explicit signal indicating which regime each observation belongs to. Implementing a rolling window in this specific context forces the algorithm to repeatedly train on fragmented, incomplete phases of this epidemiological cycle. The model is therefore likely to learn a confounded representation in which the climate signal is partially absorbed into the implicit regime-shift signal, a pathology that is in some respects more difficult to diagnose than that of unified modelling. By employing a single, continuous 88% training block (2018–2022), we structurally guarantee that the LSTM’s memory cell is exposed to the complete sequence of regime shifts, from pre-pandemic natural baseline, through extreme suppression, and ending right at the brink of the NPI relaxation. Reserving the entirely unseen 2023 “rebound year” as a chronological hold-out thereby serves as the ultimate extreme stress-test of the network’s capacity to dynamically synthesize NPI relaxation signals and climate variables to forecast a historically unprecedented surge.

      (3) Architectural mismatch between time-slicing and the LSTM’s design principle: A core motivation for adopting an LSTM rather than period-stratified statistical models was precisely to leverage its long-term dependency memory mechanism, which is designed to allow a unified architecture to learn how predictive relationships are modulated by time-varying contextual conditions. Presegmenting the data into NPI-defined strata forecloses this principal architectural advantage and would, in effect, reduce the analysis to a series of disconnected period-specific models, a design for which the LSTM’s complexity provides no benefit over simpler approaches.

      (4) Tension with the surveillance-bias correction: As we discuss in our response to Public review para 2 below, a separate and equally important critique from you concerns surveillance bias from temporally varying testing intensity. Stratified analysis would compound this problem, because each stratum carries a distinct testing-intensity profile (notably the surveillance surge during 2020–2022), and stratum-specific models cannot leverage cross-period testing-volume normalization. A unified model with explicit testing-volume covariates is protective against bias than period-stratified alternatives.

      Our Methodological Solution: Covariate-Driven Adaptive Learning

      Instead of artificially fracturing the timeline, we substantially restructured the analytical framework to teach the model how to contextualize the pandemic-era disruption while preserving longitudinal continuity:

      - First, we shifted the predictive target to Influenza Positivity Rates (laboratory confirmed cases ÷ ILI specimens tested): The positivity rate is the WHO recommended sentinel surveillance metric and intrinsically corrects for the substantial fluctuations in testing volumes, healthcare-seeking behaviour, and surveillance intensity that characterized the pre-pandemic, pandemic, and post-restriction periods, rendering the outcome metric far more comparable across the timeline.

      - Second, we integrated explicit Epidemiological Context Covariates: We structurally upgraded the LSTM network by feeding it vital time-varying covariates alongside meteorological data. Specifically, we incorporated the Our World in Data (OWID) mask-wearing stringency indices, weekly detection volumes, day of the week (DOW, distinguish between weekdays and weekends to capture administrative reporting patterns) and the proportion of non-local residents among tested patients (to account for population mobility and viral importation risks).

      - Third, we performed a covariate-ablation sensitivity analysis: To prove the model actively utilizes these contextual signals, we performed targeted ablation studies (new Table 2). Removing mask-wearing stringency indices increased the 2023 influenza A forecast MAE to 0.012 and the influenza B forecast MAE to 0.003 (compared with MAEs of 0.009 and 0.002, respectively, obtained when the covariate was retained), indicating a decline in the network’s predictive accuracy. Removing the weekly testing volume had an even more catastrophic impact on Influenza B, surging the MAE by 100% and SMAPE by 199.6%. This provides empirical evidence that the unified network does not passively average across regimes, but dynamically leverages NPI and surveillance signals to adapt to non-stationarity. We have added a paragraph to the revised Discussion interpreting these findings as suggestive (though not definitive) evidence that the unified-with covariates architecture is contextualizing rather than averaging across regimes.

      By explicitly providing the LSTM with these explicit contextual parameters, the network autonomously learned the “regime shifts.” It learned that when the mask-wearing index is high, transmission is dampened despite favorable meteorological conditions.

      Re-interpreting the 2023 Validation

      Guided by your critique, we fully agree that 2023 was an anomalous “rebound” year and not a steady-state reflection of a post-2024 “normal.” We have completely abandoned the framing that our model represents a steady-state tool for the future. Instead, we now explicitly frame the 2023 validation as an “extreme epidemiological stress test.” The revised framing rests on three explicit acknowledgements:

      - The 2023 validation tests forecasting performance during a transitional, post-restriction rebound year, and is best understood as an extreme stress test of the framework’s adaptive capacity, not as evidence of long-term predictive validity under stabilized future conditions.

      - Performance metrics observed in 2023 should not be naively extrapolated to 2024 and beyond.

      - Continuous prospective recalibration as 2024–2025 data accumulate, ideally combined with adaptive learning approaches capable of detecting regime shifts in real time, will be essential before any operational deployment.

      The fact that our updated LSTM network accurately forecasted the explosive, atypical 2023 viral rebound, despite being trained heavily on the suppressed 2020-2022 data, demonstrates its robust capacity to synthesize climate variables and NPI relaxation signals.

      We have substantially rewritten the Discussion section to honestly acknowledge this limitation and contextualize the 2023 results, stating that continuous recalibration will be essential for future forecasting.

      We invite you to review the recently included sentences in the relevant subsections as specified above.

      “A major methodological strength of this study lies in its robust, uninterrupted longitudinal data collection framework spanning January 1, 2018, to December 31, 2023. While the analytical timeline encompasses a “retrospective” phase (January 1, 2018 – October 13, 2020) prior to formal ethical approval, and a “prospective” phase thereafter, we emphasize that this distinction represents a purely administrative demarcation regarding the timing of ethical approval. It does not reflect any shift in demographic cohorts, sentinel hospital locations, or data collection methodologies. Importantly, the historical data (2018– 2020) were not subjected to the recall biases or misclassification risks typical of traditional retrospective chart reviews. Rather, they were systematically extracted from a continuously operating, highly standardized public health sentinel surveillance network. From the inception of data collection through the end of 2023, the local CDC maintained absolute uniformity in clinical influenza-like illness (ILI) definitions, nasopharyngeal swabbing procedures, and real-time reverse transcription polymerase chain reaction (RT-PCR) diagnostic assays. Consequently, the pre-2020 data possess the high-fidelity characteristics of a strict prospective cohort, ensuring unparalleled longitudinal consistency and mitigating temporal measurement bias across the entire pre-pandemic, pandemic, and postrestriction timeline.” (Methods, page 8-9)

      Secondly, the interpretation of our framework’s predictive performance during the 2023 validation period requires careful epidemiological and methodological contextualization. The year of 2023 represented an anomalous, post-restriction “rebound” period characterized by rapid NPI relaxation and the release of accumulated population-level immunity debt, resulting in an atypical influenza surge that exceeded pre-pandemic peaks. The framework’s high accuracy across this period should therefore be interpreted as evidence of algorithmic agility and adaptive capacity during a highly volatile transitional phase, rather than as definitive proof of long-term predictive validity under a stabilized post-2024 epidemiological regime. Methodologically, while our strict chronological OOT data partitioning prevented temporal information leakage, a critical requirement for LSTM integrity, the reliance on a single, fixed chronological split point (December 31, 2022) intrinsically limits our evaluation to one specific structural break. This fixed-split approach may not exhaustively probe the DLNM-LSTM framework’s resilience against all forms of future epidemiological non-stationarity. Consequently, naive extrapolation of the reported 2023 performance metrics to future surveillance years should be avoided absent prospective recalibration. Future studies should consider employing expanding-window or rolling-origin cross-validation frameworks to provide a more continuous characterization of algorithmic robustness. Continuous integration of accumulating 2024 and 2025 data, combined with adaptive learning architectures capable of detecting regime shifts in real time, will be essential before any operational deployment of this, or similar forecasting frameworks, for routine public health surveillance.” (Discussion, page 44-45)

      We believe this integrated, covariate-based approach maintains the mathematical integrity of the time-series analysis while fully addressing your valid concerns regarding epidemiological non-stationarity.

      The outcome of interest is the number of confirmed influenza cases. This is not only a function of weather, but also of the amount of testing. The amount of testing is also a function of historical patterns. This poses the real risk that the model confirms historical opinions through increased testing in those higher-risk periods. Of course, the models could also be run to see how meteorological factors affect testing and the percentage of positive tests. The results only deal with the number of positive (only the overall number of tests is noted briefly), which means there is no way to assess how reasonable and/or variable these other measures are. This is especially concerning as there was massive testing for respiratory viruses during COVID in many places, possibly including China.

      We are exceptionally grateful for this incisive methodological observation. You have accurately identified a fundamental validity threat, “surveillance intensity bias”, that inherently constrains much of the existing climate–infectious disease literature. We fully concur with your assessment that raw case counts are jointly determined by underlying viral transmission and dynamic testing intensity. Furthermore, we recognize your highly valid concern regarding the risk of “circular reasoning,” whereby meteorological factors might simply trigger higher clinical suspicion and testing rates rather than genuine transmission events—a bias severely exacerbated by the massive respiratory testing surges during the COVID-19 pandemic. To systematically dismantle this threat and address your specific recommendations, we executed a ground-up restructuring of our analytical framework, implementing four complementary strategies:

      (1) Primary outcome redefined as Positivity Rates. Throughout the entire revised manuscript, both the DLNM and LSTM pipelines have been completely re-analyzed using influenza positivity rates (laboratory-confirmed cases ÷ total ILI specimens tested) as the primary outcome, rather than raw case counts. The positivity rate is the WHOrecommended metric for sentinel surveillance precisely because it mathematically standardizes the denominator, normalizing the raw testing volume variability. Its adoption effectively neutralizes the circular-reasoning concern raised by you. All predictive models, Figures 4, 5, and 6, and all primary metrics in Table 2 now reflect positivity-based estimates.

      (2) Weekly detection volume integrated as a dynamic LSTM covariate. Even after positivity-rate normalization, residual testing-intensity effects can persist (e.g., if testing patterns shift among demographic subgroups with systematically different positivity profiles). To capture this, our revised LSTM network incorporated weekly detection volumes as an explicit dynamic input feature. As detailed in our covariate-ablation sensitivity analysis (Table 2), removing the testing-volume covariate drastically degraded the forecasting accuracy for 2023, increasing the Mean Absolute Error (MAE) by 22.2% for influenza A and an astounding 100% for influenza B. This indicates that the model’s predictions are not driven solely by meteorological inputs but are appropriately and dynamically conditioned on the surveillance context.

      (3) SHAP quantification of surveillance bias. By incorporating testing volume into the LSTM, we made the surveillance-bias concern concretely visible and quantifiable. As shown in our new SHAP analysis (Figure 6E-H), weekly_detection emerged as the second most impactful predictor of the positivity rate for influenza B, and, when examining the influenza A results, weekly_detection likewise ranked fifth among the most impactful predictors. This proves that the deep learning algorithm autonomously recognized the profound impact of testing intensity and actively utilized it to dynamically adjust and calibrate its epidemiological forecasts.

      (4) Decoupling “Circular Reasoning” via DLNM Sensitivity Analysis. To definitively prove that our model does not merely “confirm historical opinions through increased testing,” we conducted a targeted DLNM sensitivity analysis (detailed in the new Additional file 4 and Figures S3–S4). We compared DLNM exposure-response curves predicting positivity rates with and without adjusting for weekly testing volumes. The results were striking: the non-linear exposure-response curves and extreme-weather lag patterns remained highly consistent across both models. This empirical stability proves that the identified meteorological drivers represent intrinsic biological/environmental triggers of viral transmission, independent of fluctuating surveillance intensity.

      We are deeply grateful that this critique prompted such substantial methodological refinement. The revised analyses are vastly more robust and epidemiologically interpretable than the original case-count-based version. We invite you to review the recently included sentences in the relevant subsections as specified above.

      “To rigorously address the inherent confounding effects of “surveillance intensity bias”, where fluctuations in raw case counts may merely reflect transient surges in clinical testing capacity rather than true community transmission, the primary outcome metric for all DLNM modeling was mathematically defined as the influenza positive rate, with meteorological factors and weekly detection volume serving as independent variables. The formula for calculating the daily influenza positivity rate is as follows:

      where R represents the daily influenza positivity rate, I represents the number of daily influenza positive cases, and N denotes the total number of daily influenza tests performed. By adopting this WHO-recommended surveillance metric, our our analytical framework explicitly standardizes the epidemiological denominator. This mathematical normalization effectively neutralizes the severe surveillance intensity bias caused by dramatic testing volume surges during the COVID-19 pandemic, ensuring that our models capture intrinsic viral transmissibility rather than artificial fluctuations in healthcare-seeking behavior or diagnostic capacity.” (Methods, page 16)

      “Furthermore, to evaluate the robustness of our findings against testing intensity, we conducted a targeted sensitivity analysis by reconstructing the DLNM models without the “weekly detection volumes” covariate. We then compared the non-linear cumulative risks and extreme weather lag effects between these ablated models and the original fully adjusted models. This allowed us to determine whether the identified climate-transmission associations were stable and biologically intrinsic, or merely artifacts of weather-correlated testing behaviors (Supplementary Information Additional file 4).” (Methods, page 17-18)

      “To assess the contribution of pandemic-related confounding variables, we performed targeted sensitivity analyses on the LSTM networks. Specifically, we sequentially removed covariates of mask-wearing stringency indices and weekly detection volumes from the input features while maintaining identical Bayesian-optimized hyperparameters. The performance of these ablated models was evaluated using MAE, RMSE, MAPE, and SMAPE, allowing us to quantify the exact necessity of incorporating testing and behavioral covariates in forecasting models during periods of epidemiological non-stationarity.” (Methods, page 22)

      “Crucially, this LSTM stage explicitly incorporates weekly detection volumes, maskwearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 40)

      “This study has several limitations that should be considered when interpreting the findings. Firstly, our primary analysis relies on influenza surveillance data collected from seven sentinel hospitals in Putian, which inherently captures only a fraction of all influenza cases occurring in the broader community. Although employing positivity rates as the primary outcome substantially mitigates the surveillance bias inherent in count-based analyses, residual selection effects may persist if testing patterns shift differentially across demographic subgroups with systematically divergent positivity profiles. Our inclusion of weekly detection volumes as an explicit LSTM covariate, complemented by parallel DLNM sensitivity analyses validating the independence of meteorological effects from testing volumes, were designed to characterize and partially account for this residual bias, but cannot fully eliminate it. Consequently, positivity rates likely underestimate the true burden of community influenza infection. Although the surveillance infrastructure in Putian remained uniform and uninterrupted throughout the 2018–2023 timeline, mitigating measurement bias and temporal confounding risks, the transferability of our findings to other global subtropical regions with different socioeconomic structures, healthcare systems, or population behaviors requires cautious, region-specific calibration. Accordingly, our framework should be strictly interpreted as a high-fidelity tool designed to forecast the observable public health surveillance signal, which is the most operationally relevant target for public health agencies, rather than for estimating unobserved, absolute community disease burden.” (Discussion, page 43-44)

      “This limitation was particularly exacerbated by the profound epidemiological disruptions during the COVID-19 pandemic, where raw numbers of confirmed cases became heavily confounded by surveillance intensity (i.e., fluctuating testing volumes) rather than solely reflecting underlying viral transmission. Our adoption of influenza positivity rates as the primary modeling endpoint and incorporation of weekly detection volumes, face-covering stringency, and non-local population proportion as dynamic covariates, rigorously mitigated these aggregate-level biases and linked our methodological design directly to the forecasting outcomes. As unequivocally demonstrated by our covariate-ablation sensitivity analyses (Table 2), failing to account for mask mandates and testing volumes leads to severe, mathematically predictable deviations in absolute forecasting accuracy. Furthermore, while daily case counts in a single city can occasionally be small, sporadic, and driven by external importations, our incorporation of non-local population proportion effectively adjusted for these localized importation risks. Consequently, although our findings characterize population-level associations between meteorological factors and influenza activity, they should not be interpreted as evidence of micro-level causal mechanisms at the individual patient level. Ultimately, the DLNM-LSTM framework’s robust performance across the non-stationary transition out of NPI policies highlights the absolute necessity of integrating behavioral and virological baseline metrics into future climate-driven predictive surveillance systems.” (Discussion, page 45-46)

      We are deeply grateful that this critique prompted such substantial methodological refinement. The revised analyses are vastly more robust and epidemiologically interpretable than the original case-count-based version.

      (1) Although the authors note a correlation between influenza and the weather factors. The authors do not discuss some of the high correlations between weather factors (e.g., solar radiation and UV index). Because of the many weather factors, those plots are hard to parse.

      We sincerely appreciate your constructive feedback regarding the visual clarity of the correlation plots and the methodological implications of highly correlated meteorological variables. We agree that the original 11 × 11 scatterplot matrix was visually overwhelming and could obscure the statistical implications of highly correlated features (such as solar radiation and the UV index, Pearson’s r > 0.9).

      To address this, we have taken two specific revisions:

      (1) Improved Visualization (Revised Supplementary Figure S2): To make complex relationships easier to parse, we have completely redesigned Supplementary Figure S2. It is now logically partitioned into two distinct visual components:

      - Panel A features a high-contrast Pearson correlation heatmap, where color intensity and statistical significance asterisks allow for rapid, intuitive identification of highly correlated pairs (such as solar radiation and UV index).

      - Panel B retains the pairwise scatterplots (lower-left) and density distribution curves (diagonal) to facilitate the detailed visual inspection of non-linear trends, data skewness, and potential anomalies.

      (2) Methodological Defense on Potential Collinearity: We did not arbitrarily eliminate these highly correlated variables because our two-stage modeling approach inherently mitigates multicollinearity risks associated with potential collinearity:

      - For the DLNM analysis: To prevent coefficient instability caused by multicollinearity, all DLNM lagged analyses strictly employed univariate exposure-response models for meteorological factors (i.e., evaluating the relationship between a single meteorological factor and influenza positivity rates at a time, while controlling for long-term temporal trends and including other covariates as fixed terms). Consequently, the independent effect sizes and lag structures derived from the DLNMs are completely unaffected by inter-variable correlations.

      - For the LSTM network: Unlike traditional multiple linear regression (or ARIMA) where potential collinearity inflates standard errors and destabilizes coefficients, Deep Learning architectures (LSTM) are natively robust to redundant features. The network’s non-linear activation functions and gating mechanisms naturally weight overlapping signals during the optimization process, effectively using redundant variables as a form of algorithmic regularization without compromising predictive stability.

      We have expanded the Statistical Analysis and Discussion sections of the manuscript to explicitly articulate this rationale, ensuring maximum methodological transparency.

      We invite you to review the recently included sentences in the relevant subsections as specified above.

      “Prior to analytical modeling, we assessed the pairwise associations and potential collinearity among all meteorological variables using a comprehensive correlation heatmap and scatterplot matrix (Supplementary Figure S2). Notably, certain variables, such as solar radiation and the UV index, exhibited high positive correlations. To avoid coefficient instability typically caused by multicollinearity in regression models, all DLNM analyses adopted a univariate approach for meteorological factors, sequentially evaluating the nonlinear and lagged effects of individual meteorological predictors while adjusting for time trends and including other fixed covariates. Conversely, all variables were retained during the LSTM forecasting phase, as the non-linear gating architecture of recurrent neural networks inherently exhibits robust regularization against potential collinearity among input features.” (Results, page 26)

      “Furthermore, extensive environmental inputs inevitably introduce severe collinearity, such as the strongly correlated solar radiation and UV index. While traditional multivariate models are highly vulnerable to such overlapping variances, the recurrent, weighted representation learned by the LSTM is comparatively tolerant of such redundancy, allowing broader covariate integration than in previous efforts.” (Discussion, page 40)

      We hope that our responses and revisions will meet your expectations and demonstrate our dedication to improving the scientific quality of this study.

      (2) The authors do not actually compare the results of both methods and what the LSTM adds.

      We sincerely appreciate your perceptive critique. We agree that our initial manuscript lacked a sufficiently rigorous, side-by-side comparison, and more importantly, it failed to adequately articulate why the LSTM architecture succeeds where classical statistical baselines fail.

      To rectify this, we have comprehensively overhauled the comparative analysis. We formulated our baseline as a multivariate ARIMA model, supplying it with the exact same multidimensional covariate matrix as the LSTM (including meteorological variables, mask-wearing indices, and weekly testing volumes), focusing on three dimensions:

      (1) Direct Quantitative Comparison: Following the restructuring of our outcome variable to the Influenza Positivity Rate, we re-evaluated both models using identical training (2018–2022) and validation (2023) datasets. The LSTM consistently and substantially outperformed the ARIMA model across all metrics. For Influenza A, the LSTM achieved an MAE of 0.009 and RMSE of 0.035 (vs. ARIMA: MAE 0.136, RMSE 0.238). For Influenza B, the LSTM yielded an MAE of 0.002 and RMSE of 0.011 (vs. ARIMA: MAE 0.049, RMSE 0.057). We have updated Table 2 and the corresponding Results section to explicitly present these side-by-side comparisons.

      (2) Visualizing “Where” the LSTM Outperforms: We have updated Supplementary Figure S3 to map the ARIMA predictions for the 2023 positivity rate, allowing for a direct visual comparison with the LSTM predictions in Figure 6 (A, B). The visualizations explicitly reveal the ARIMA model’s fundamental limitation: it tends to predict relatively flat or conservatively smoothed values, failing entirely to capture the extreme, explosive non-linear peaks of the 2023 viral rebound. Conversely, the LSTM network accurately tracks these sudden epidemic phase transitions.

      (3) Articulating “What the LSTM Adds” (Discussion Expansion): We have significantly expanded the Discussion section to intellectually articulate why the LSTM succeeds where ARIMA fails. To ensure a strictly fair methodological comparison, we formulated our baseline as a multivariate ARIMA model (ARIMAX), supplying it with the exact same multidimensional covariate matrix as the LSTM (including meteorological variables, mask-wearing indices, and weekly testing volumes). Therefore, the LSTM’s superior performance is not due to information asymmetry (i.e., it did not “see” more variables), but stems directly from its algorithmic architecture. The consistent superiority of the LSTM over both the linear sequential baseline (ARIMA) and the non-linear non-sequential baseline (XGBoost) isolates the recurrent gated architecture itself as the source of the predictive gain. This advantage arises from four architectural properties intrinsic to recurrent gated networks yet absent in both tree ensembles and linear autoregressive models:

      a) Sensitivity to temporal ordering, which decision-tree splits and linear regressors cannot natively encode;

      b) Gated propagation of long-range dependencies through forget–input–output mechanisms;

      c) Explicit accommodation of serial autocorrelation, violated by the i.i.d. assumptions underlying gradient-boosted trees;

      d) Paradigmatic comparability with sequential statistical models: the joint failure of ARIMA (linear, sequential) and XGBoost (non-linear, non-sequential) isolates recurrent sequential memory, rather than non-linearity per se, as the critical feature for forecasting under pandemic-era non-stationarity.

      We emphasize that this interpretation applies specifically to the present non-stationary epidemiological forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain competitive performance across many structured prediction domains.

      We invite you to review the recently revised table, figures and sentences in the relevant subsections as specified above.

      “The LSTM networks accurately captured both the timing and magnitude of these nonlinear epidemic surges, including the two outbreak peaks of influenza A during February-March and November-December of 2023, as well as the peak of influenza B in November-December, demonstrating good predictive performance. The predictive performance of the LSTM networks was quantitatively assessed using metrics such as MAE, RMSE, MAPE, and SMAPE. For influenza A, the MAE was 0.009, RMSE was 0.035, MAPE was 0.158, and SMAPE was 0.521; for influenza B, the MAE was 0.002, RMSE was 0.011, MAPE was 0.17, and SMAPE was 0.484 (Table 2).” (Results, page 31)

      “To benchmark the predictive value added by the LSTM architecture, we constructed two baseline models using identical training (2018–2022) and validation (2023) positivity-rate datasets, inclusive of all contextual covariates: the multivariate ARIMA model and the XGBoost gradient-boosting model. The LSTM model demonstrated decisive superiority over both baselines across all evaluated metrics (Supplementary Table S3). For influenza A, the LSTM achieved an MAE of 0.009 and SMAPE of 0.521, compared with the ARIMA’s MAE of 0.136 (SMAPE 1.212) and XGBoost’s MAE of 0.138 (SMAPE 1.081), representing approximately 15-fold error reductions relative to both benchmarks. For influenza B, the LSTM yielded an MAE of 0.002 (SMAPE 0.484), compared with 0.049 (SMAPE 0.810) for ARIMA and 0.070 (SMAPE 0.737) for XGBoost, approximately 25- to 35-fold reductions. Notably, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, suggesting that the LSTM’s advantage derives not from non-linear modeling capacity per se, but from architectural properties specific to recurrent sequential processing.

      Beyond global error metrics, visual comparison of forecasting trajectories (Figure 6; Supplementary Figures S5 and S6) reveals critical behavioral disparities. For influenza A, both the linear ARIMA model and the non-linear XGBoost model generated conservatively smoothed forecasts that entirely failed to capture the sudden, explosive peaks of the 2023 post-restriction viral rebound. For influenza B, ARIMA produced continuous spurious fluctuations during non-epidemic periods (when true positivity was near zero), likely overreacting to covariate variations, while XGBoost generated largely flat trajectories that missed the mid-year outbreak peak. In stark contrast, the LSTM’s recurrent gating mechanisms successfully filtered out covariate noise during low-transmission periods while accurately tracking extreme epidemiological phase transitions, demonstrating a qualitative advantage in handling non-stationary regime shifts.” (Results, page 33-34)

      “To rigorously isolate the predictive contribution attributable to the LSTM’s recurrent architecture, we benchmarked it against two covariate-matched baselines representing distinct methodological paradigms: the multivariate ARIMA model and the XGBoost gradient-boosting ensemble. This design controls simultaneously for linearity (ARIMA→LSTM contrast) and for non-linearity without recurrent memory (XGBoost→LSTM contrast), allowing us to attribute observed performance gains to specific architectural inductive biases rather than to informational asymmetry or model non-linearity in general. The LSTM substantially outperformed both baselines on the 2023 validation window. For influenza A, it achieved an MAE of 0.009, compared with 0.136 for ARIMA and 0.138 for XGBoost, approximately 15-fold reductions. For influenza B, the LSTM yielded an MAE of 0.002, versus 0.049 for ARIMA and 0.070 for XGBoost, 25- to 35-fold reductions. Critically, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, and both systematically under-predicted the explosive 2023 post-NPI rebound for influenza A while generating flat or spurious trajectories for influenza B (Supplementary Figures S5–S6). That parallel failure indicates the LSTM’s advantage under pandemic-era non-stationarity derives not from non-linearity per se, but from four architectural properties intrinsic to recurrent gated networks yet absent in tree ensembles: (i) threshold-based, sequence-insensitive splits fail to encode present–past dynamics; (ii) XGBoost lacks forget–input–output gates to propagate and re-weight historical states across time lags; (iii) gradient-boosted trees assume near-independence, contradicting the pronounced temporal autocorrelation in epidemiological time series; (iv) ARIMA and LSTM are sequential models with different functional forms, whereas XGBoost is non-sequential. The joint failure of ARIMA and XGBoost, despite differing non-linear treatment, isolates recurrent sequential memory as the critical architectural feature for forecasting under non-stationarity. We emphasize that this interpretation applies specifically to the present non-stationary influenza forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain state-of-the-art performance across many structured prediction domains. Rather, it highlights that for surveillance time series exhibiting pronounced temporal dependencies and abrupt regime shifts, such as the 2023 post-NPI rebound, explicit sequential memory becomes functionally essential. This mechanistic reading, together with the LSTM’s comparative edge over previously reported ARIMA-based (Li et al. 2024) and LSTM-based influenza prediction models (Zhu et al. 2022), positions our framework as a substantive methodological advance in predictive modeling for climate-sensitive diseases.” (Discussion, page 41-42)

      Results

      “The MAE for the covariate-adjusted ARIMA model of influenza A was 0.136, RMSE was 0.238, MAPE was 1.115, and SMAPE was 1.212. For influenza B, the MAE was 0.049, RMSE was 0.057, MAPE was 1.426, and SMAPE was 0.810. Overall, the predictive performance of the ARIMA model was substantially inferior to that of the Bayesian optimized LSTM. As illustrated in Figure S5, the linear model struggled significantly with the non-stationary dynamics of the 2023 viral rebound. It either failed entirely to capture the extreme, explosive non-linear peaks (as seen in Influenza A) or generated continuous spurious predictions during zero-case periods due to mechanical linear reactions to covariate inputs (as seen in Influenza B).” (Supplementary information, page 16)”

      Results

      “The XGBoost model produced substantially higher forecasting errors than the LSTM across both influenza subtypes (Table S3; Figure S6). For Influenza A, XGBoost yielded MAE = 0.1379, RMSE = 0.2161, MAPE = 1.8195, and SMAPE = 1.0805, errors of a magnitude broadly comparable to the multivariate ARIMA baseline (MAE = 0.136) and approximately 15-fold higher than the LSTM (MAE = 0.009). For Influenza B, XGBoost achieved MAE = 0.0704, RMSE = 0.0739, MAPE = 1.9172, and SMAPE = 0.7371, exceeding both the ARIMA benchmark (MAE = 0.049) and the LSTM (MAE = 0.002) by approximately 35-fold relative to the latter.” (Supplementary information, page 19)

      We believe these additions provide a much more rigorous and academically satisfying comparative analysis.

      (3) The methods are long and meandering. They could be cleaned up and shortened. E.g., there is no need for 30 lines on PCR testing; the study area should come before the study design. The authors discuss similar elements in multiple places; this whole section can be shortened considerably without affecting the content.

      We sincerely appreciate the reviewer’s editorial guidance. We agree that the initial Methods section was somewhat disjointed and unnecessarily verbose, reflecting iterations from previous drafts. We have completely restructured and aggressively streamlined this section to ensure a logical and concise flow:

      - Structural Reorganization: As suggested, we have moved the “Study area” section to the very beginning of the Methods, providing the geographical and climatic context before detailing the study design and cohort.

      - Consolidation of Redundancies: We have merged the fragmented descriptions regarding ethical approvals, data anonymization, and sentinel hospital protocols into a single, cohesive “Study design and cohort” subsection.

      - Condensing Laboratory Protocols: We completely agree that 30 lines on standard PCR testing in the main text are unnecessary. We have condensed the “Sample collection and pathogen typing” section into a brief, 4-line summary explicitly stating the use of commercial assays (Da’an Gene Co., Ltd) and strict adherence to China CDC guidelines. The highly technical nuances (e.g., RNA extraction integrity, spectrophotometer ratios, and precise PCR amplification thresholds) have been relocated to the Supplementary Information (Additional file 1) for interested readers.

      We invite you to review the recent revisions in the Methods subsection as specified above.

      “Sample collection and pathogen typing

      Respiratory specimens (nasopharyngeal swabs) were collected from ILI patients during their initial visit, prior to treatment, and stored at 4°C in viral transport medium. All samples were delivered to the Putian CDC laboratory within 24 hours of collection. Nucleic acid extraction and one-step real-time fluorescent RT-PCR for influenza A/B subtyping (H1N1, H3N2, Victoria, and Yamagata lineages) were executed within 24 hours upon sample arrival. All assays utilized commercial diagnostic kits (supplied by Da’an Gene Co., Ltd., Guangzhou, China) and were processed in a Biosafety Level 2 (BSL-2) laboratory, strictly adhering to the manufacturer’s instructions and China CDC’s standardized protocols. Detailed laboratory procedures, including RNA integrity parameters and PCR amplification thresholds, are comprehensively documented in the Supplementary Information (Additional file 1).” (Methods, page 13)

      These revisions have significantly improved the readability of the manuscript without sacrificing methodological transparency.

      (4) How reliable is the "Our Word in Data" website for subnational coverage of restrictions? Some of the authors are from Putian and should be able to confirm the accuracy for both studied areas.

      We are very grateful for this insightful question. You astutely identify a common challenge in geospatial epidemiology in China: the systemic lack of publicly accessible, standardized, daily non-pharmaceutical intervention (NPI) datasets at the municipal (subnational) level. Local CDC policy records are typically maintained as internal administrative documents without standardized time-series data interfaces.

      Given this limitation, we utilized the national Our World in Data (OWID) stringency index as a proxy. To directly address the reviewer's excellent point regarding local verification, the co-authors of this study—who are frontline epidemiologists stationed at the Putian CDC and who directed the local pandemic response, including conducted a rigorous, retrospective cross-validation of the OWID index against local realities.

      Our local experts confirmed a high degree of fidelity between the OWID 0–4 scale and the actual policies enforced in Putian and Sanming:

      - 2020 (Initial Outbreak): Both cities enforced “mandatory face coverings outside the home at all times” (OWID Level 4), aligning perfectly with the dataset.

      - 2021–2022 (Normalized Control): Policies shifted to “required in all shared/public spaces” (OWID Level 3), which accurately reflects the local mandates required for public transit, schools, and commercial venues.

      - 2023 (Post-Pandemic Shift): Following the national policy pivot, mandates were downgraded to “recommended” (OWID Level 1), perfectly mirroring local ground truths.

      Therefore, while OWID provides a national-level index, our local CDC authors have empirically verified that its temporal variations accurately capture the intensity of behavioral restrictions experienced by the populations in our specific study areas. We have incorporated a concise statement regarding this expert validation into the revised Methods section.

      “While OWID stringency indices represent national-level policy, publicly accessible and standardized daily NPI datasets at the municipal level are currently unavailable in China. To ensure the spatial validity of these indices, co-authors from Putian CDC, who actively managed the local epidemic response, conducted a rigorous cross-validation. Our local public health experts confirmed that temporal fluctuations of OWID indices (ranging from Level 4 strict mandates in 2020 to Level 1 recommendations in 2023) exhibited high fidelity with the actual, on-the-ground enforcement of NPIs in both Putian and Sanming. Thus, OWID stringency indices serve as a highly reliable contextual proxy for local social contact restrictions.” (Methods, page 15)

      We hope that our responses and the additional statement will meet your expectations.

      (5) Figure 2A is hard to parse; it would make more sense to plot these as line plots (y=count, x=month).

      We completely agree with you. The original heatmap visualization obscured the temporal dynamics of the distinct influenza subtypes. We have entirely redrawn Figure 2A as multi-line plots mapping the monthly incidence trajectories of Influenza A (H1N1, H3N2) and Influenza B (Victoria, Yamagata) across the six-year study period. This new visualization (included in the revised manuscript) vastly improves readability and explicitly highlights the distinct phase shifts and interruptions caused by the pandemic.

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3 is hard to parse. The flu counts are repeated in every figure, which makes them look very similar. I would recommend that the authors pick 1 or 2 as subpanels for the main paper and put the rest in the supplement.

      We sincerely appreciate your feedback. We completely agree that the original 3x3 square layout severely compressed the x-axis, making the daily temporal fluctuations difficult to parse and causing the flu curves to look visually redundant.

      To resolve this visual clutter without losing valuable environmental context in the main text, we adopted a highly effective structural solution simultaneously suggested by Reviewer #2. We have completely redrawn Figure 3 into an 8x1 vertically stacked format with a single, shared continuous x-axis across the entire study period. This extended aspect ratio vastly expands the timeline, dramatically revealing the highly distinct, daily microfluctuations of each meteorological factor alongside the epidemiological curves. We believe this new layout completely resolves the parsing difficulty you rightly pointed out. Given that a substantial portion of our readership relies heavily on the main text figures for immediate epidemiological context, retaining these cleanly formatted panels in the main manuscript maximizes the paper’s scientific impact. We hope you find this redesigned visualization satisfactory.

      (2) Using a thousand-separator throughout will make the manuscript more readable.

      We completely agree. We have meticulously applied thousand-separators to all relevant numerical values (e.g., 20,488; 17,333) throughout the revised manuscript to enhance readability.

      (3) Line 556 "quantified calculated", pick 1 word.

      We sincerely apologize for this typographical oversight resulting from the drafting process. However, the original sentence that led to the duplicated phrasing you highlighted has been removed, as we had already undertaken a comprehensive revision of the relevant material in that subsection in response to your earlier remarks. We invite you to review the newly substituted paragraph below.

      “Prior to analytical modeling, we assessed the pairwise associations and potential collinearity among all meteorological variables using a comprehensive correlation heatmap and scatterplot matrix (Supplementary Figure S2). Notably, certain variables, such as solar radiation and the UV index, exhibited high positive correlations. To avoid coefficient instability typically caused by multicollinearity in regression models, all DLNM analyses adopted a univariate approach for meteorological factors, sequentially evaluating the nonlinear and lagged effects of individual meteorological predictors while adjusting for time trends and including other fixed covariates. Conversely, all variables were retained during the LSTM forecasting phase, as the non-linear gating architecture of recurrent neural networks inherently exhibits robust regularization against potential collinearity among input features.” (Results, page 26)

      Reviewer #2 (Public review):

      Summary:

      The study aimed to assess the associations between meteorological drivers and influenza is important although not new. The authors used only 6 years of surveillance data and deep learning models, combining distributed lag non-linear models (DLNM) with Bayesian optimized LSTM neural networks for predictive modeling. The key interest in this area is to explore the subtropical locations, where influenza is less common and circulates year round. The authors further claimed that such an association could be able to provide an early warning in the community. In this direction, the current manuscript has several scopes of improvements and clarification of the claims, as I list here.

      Strengths:

      Study design based on a prospective cohort to analyse the data for retrospective outcomes.

      We sincerely thank you for the careful and constructive evaluation of our manuscript, and in particular for recognising the value of our prospective surveillance design and the importance of investigating influenza–meteorological associations in subtropical settings where year-round circulation patterns differ substantively from those in temperate regions. We are grateful that you have identified four specific dimensions in which the manuscript can be strengthened, rationale clarity, methodological/data-integration transparency, validation reporting, and the calibration of the “early warning” claim. We address each of these four points in detail below, and we have undertaken substantive revisions to the manuscript in response.

      Weaknesses:

      (1) The rationale of the study is not clearly stated.

      We sincerely thank you for this incisive observation. We agree that the original Introduction did not adequately articulate the study’s rationale, specifically, the causal chain linking public-health need, existing methodological limitations, and the incremental contribution of our integrated DLNM-plus-LSTM framework. The Introduction has been substantively rewritten to make this rationale explicit, structured around four logical pillars:

      - Disease burden grounding. We have added quantitative evidence on the global burden of seasonal influenza, such as annual mortality estimates, drawing on solid epidemiological sources, to establish the public-health magnitude that motivates the study.

      - Subtropical-specific knowledge gap. We articulated the distinctive epidemiological challenges of subtropical influenza transmission, including year-round circulation patterns, complex non-linear meteorological associations, and lag-structured exposure-response relationships, that fundamentally differentiate subtropical contexts from temperate epidemiological settings where most existing research has been conducted. This articulation directly motivates our adoption of distributed lag non-linear models (DLNM) as the appropriate analytical framework for capturing these complex non-linear and lag-structured associations.

      - Methodological gap and incremental contribution. We now position our integrated framework against three specific gaps in the existing literature: (a) studies using DLNM alone characterize lag-distributed exposure–response relationships but lack forecasting capability; (b) studies using LSTM alone provide forecasts but typically do not incorporate Bayesian hyperparameter optimization, do not stratify by influenza subtype, and do not account for COVID-19-era non-pharmaceutical interventions; (c) no existing study, to our knowledge, integrates DLNM-based mechanistic interpretation with Bayesian-optimized, subtype-specific LSTM forecasting in a subtropical Chinese setting under pandemic-perturbed surveillance conditions. Our study is positioned to fill this specific gap.

      - Adequacy of the six-year data window. We additionally address your implicit concern regarding study duration. While six years (2018–2023) is shorter than some long-horizon influenza time-series studies, this window was deliberately selected because it brackets a uniquely informative epidemiological transition: two prepandemic baseline years (2018–2019), three Non-Pharmaceutical Interventions (NPI)-suppressed years (2020–2022), and one post-suppression rebound year (2023). This structure allows the model to learn from a structural break that a longer but earlier-only series could not provide. We have made this argument explicit in the revised Introduction.

      - The dual-model rationale. We clarify that our core rationale is to bridge this gap. By utilizing DLNM to uncover the underlying environmental biological triggers and subsequently employing the Bayesian-optimized LSTM network, uniquely upgraded to ingest epidemiological context (mask-wearing stringency index, testing volumes, day and week, viral importation), we provide a comprehensive framework that achieves both mechanistic insight and operational forecasting agility. We have substantially rewritten the Introduction to reflect this explicit storyline.

      We close the revised Introduction by explicitly framing the study’s translational endpoint: providing a methodologically integrated framework (DLNM for mechanistic interpretation; Bayesian-optimised LSTM for forecasting) to support climate-informed influenza preparedness in subtropical settings. We have, however, calibrated the language used to describe this endpoint (see our response to Weakness #4) to avoid overstating the study’s operational readiness as an early-warning tool.

      We have substantially rewritten the Introduction to reflect this sharpened rationale. We invite you to read the whole section of the Introduction in the revised manuscript.

      (2) Several issues with methodological and data integration should be clarified.

      We sincerely appreciate your identification of methodological and data integration ambiguities in the original manuscript. We have substantively addressed this concern through three categories of clarifying revisions: (i) explicit articulation of the covariate framework, (ii) clarification of the analytical relationship between DLNM and LSTM components, and (iii) detailed specification of data sources and quality control procedures:

      - Clarification 1: Comprehensive covariate framework specification. We have explicitly articulated the complete covariate framework integrated into both DLNM and LSTM analyses. Beyond the meteorological factors (mean temperature, maximum temperature, minimum temperature, diurnal temperature range, relative humidity, atmospheric pressure, precipitation, sunshine duration), our analytical framework systematically incorporates: (a) influenza positivity rates as the primary outcome variable (replacing raw case counts to mitigate surveillance intensity bias, as detailed in our response to Reviewer 1’s Public Review Comment 2); (b) weekly testing volumes as an explicit covariate to control for residual surveillance-intensity variations; (c) mask-wearing stringency indices to capture pandemic-era public health intervention effects; (d) day-of-week indicators distinguishing weekdays from weekends to control for healthcare-seeking behavioral cycles; and (e) nonlocal population proportion to account for population mobility-related transmission dynamics.

      - Clarification 2: Articulation of the DLNM-LSTM analytical relationship. We have explicitly clarified the complementary analytical roles of DLNM and LSTM within our integrated framework, addressing potential confusion regarding whether these methods serve redundant or complementary functions.

      - Clarification 3: Data source specification and quality control documentation. We have substantively expanded the data source specification and quality control documentation to ensure full methodological transparency.

      We have revised the “Study design” subsection in the Methods to transparently outline how these multifaceted data streams were temporally aligned and fed into the dual-model architecture. We invite you to review these rewritten paragraphs.

      “The study spanned January 1, 2018, to December 31, 2023, integrating four categories of data sources: (i) ILI and laboratory-confirmed cases from seven influenza sentinel hospitals across Putian’s urban and rural areas, ensuring representative coverage of diverse healthcare-seeking populations; (ii) daily meteorological data; (iii) COVID-19 associated public health intervention indicators (mask-wearing stringency indices) recorded from January 2020 onwards; and (iv) demographic mobility indicators (non-local population proportion) obtained from ILI consultation records. To ensure consistency between meteorological measurements and influenza incidence records across all data sources, we applied rigorous quality control and pre-processing procedures, including: temporal alignment of all data streams to a unified daily resolution; missing value imputation using temporally adjacent observations for sporadic gaps (<5% of records); cross-validation of laboratory-confirmed cases against ILI consultation records to identify and resolve coding inconsistencies; and standardization of meteorological measurements against the regional monitoring network’s established calibration protocols. Final datasets underwent independent verification by two co-investigators to ensure analytical reliability.” (Methods, page 11)

      “DLNM was first constructed to screen meteorological factors and other covariates with substantial influence on influenza seasonality. Subsequently, an LSTM neural network was developed within the same covariate system. The integrated DLNM–LSTM framework was employed as methodologically complementary rather than redundant components, leveraging the distinctive strengths of each approach to address different analytical objectives within a unified investigation. Specifically, DLNM models characterize the non-linear exposure–lag–response relationships between meteorological factors and influenza risk, providing biologically interpretable insights into the temporal structure of weather-influenza associations and identifying meteorological factors with statistically and clinically significant effects on influenza dynamics. Building upon this DLNM-derived foundation, the LSTM network constructs a time-series forecasting tool within an identical covariate framework, evaluating predictive capability for influenza transmission trends. Beyond meteorological factors, our DLNM-LSTM framework systematically incorporated influenza positivity rates as the primary outcome variable, weekly detection volumes, mask-wearing stringency index (indicator of NPIs during the COVID-19 pandemic), day of the week (DOW, distinguishing weekdays from weekends), and non-local population proportion (defined as the ratio of the number of individuals whose reported residential district at the time of testing lies outside Putian city to the total number of tests) as covariates within both DLNM and LSTM modeling pipelines. This comprehensive covariate framework ensures that observed meteorological associations are estimated after controlling for surveillance intensity, public health intervention status, behavioral healthcare-seeking cycles, and population mobility patterns.” (Methods, page 11-12)

      (3) Validation of the models is not presented clearly.

      We sincerely appreciate your identification of insufficient clarity in the validation framework presentation. We acknowledge that the original manuscript inadequately articulated the multi-tiered validation architecture underlying our analytical framework. We have substantively expanded the validation framework documentation through three categories of clarifications: (i) explicit articulation of the three-tier data partitioning architecture, (ii) detailed specification of validation procedures across each tier, and (iii) systematic enumeration of validation evidence supporting each analytical conclusion.

      - Clarification 1: Three-tier data partitioning architecture. Our validation framework employs a rigorously designed three-tier data partitioning architecture that addresses different validation objectives at each tier.

      Tier 1: Internal training and validation (Putian, 2018-2022). We allocated the 2018-2022 Putian surveillance data as the primary training set, within which 10% of samples were further randomly partitioned as an internal validation subset for hyperparameter tuning and overfitting monitoring during LSTM training. Early stopping mechanisms were implemented to terminate training when internal validation loss plateaued, preventing overfitting to training-specific patterns.

      Tier 2: Internal testing (Putian, 2023). We allocated the 2023 Putian surveillance data as the internal testing set, providing a temporally independent assessment of model predictive performance on data not utilized during training or hyperparameter optimization. The chronological partitioning preserves time series modeling validity by ensuring that all training data temporally precede testing data, avoiding data leakage that could artificially inflate performance estimates.

      Tier 3: External validation (Sanming, 2023). We obtained surveillance data from Sanming city for the period January 1, 2023, to December 31, 2023, matching the temporal coverage of the Putian internal testing set. This external validation set provides geographically independent assessment of model transferability across subtropical Chinese contexts, evaluating whether the Putian-derived model architecture generalizes to a different subtropical city sharing comparable climatic characteristics, influenza seasonality patterns, and public health intervention frameworks.

      - Clarification 2: Unified model framework across all validation tiers. A critical methodological feature of our validation framework is that the identical Bayesian-optimized LSTM architecture trained on Putian 2018-2022 data was applied without modification across all three validation tiers. This unified framework approach is methodologically essential because: (a) it tests genuine model transferability rather than evaluating differently-tuned models at each tier, which would conflate validation with re-optimization; (b) it enables direct performance comparison across internal testing and external validation, isolating the marginal performance degradation attributable to geographic transfer; and (c) it aligns with operational deployment scenarios where a trained model must be applied to new contexts without re-training.

      - Clarification 3: Validation evidence enumeration. The validation evidence supporting our analytical conclusions encompasses four complementary dimensions:

      Predictive performance metrics: Across both influenza A and B, the Bayesian-optimized LSTM achieved low error metrics on the internal testing set (influenza A: MAE = 0.009, RMSE = 0.035, MAPE = 0.158, SMAPE = 0.521; influenza B: MAE = 0.002, RMSE = 0.011, MAPE = 0.170, SMAPE = 0.484), substantially outperforming ARIMA benchmark models (Supplementary Figure S3).

      External validation: The Putian-derived LSTM successfully generalized to Sanming external validation data, with performance metrics maintaining comparable magnitudes to internal testing performance, substantiating model transferability across subtropical contexts.

      Sensitivity analyses: We conducted systematic sensitivity analyses across both DLNM and LSTM components. DLNM sensitivity analyses (Supplementary Figures S4-S5) demonstrate substantial concordance in cumulative risk patterns and lag-specific extreme condition responses across models with and without weekly testing volume adjustment, substantiating robustness of meteorological associations. LSTM sensitivity analyses (New Table 2) demonstrate that systematic covariate exclusion produces predictable and biologically plausible performance degradation patterns rather than artificially robust performance, confirming the absence of overfitting characteristics.

      Interpretability verification: SHAP interpretability analysis (Figure 6E-H) substantiates that the LSTM autonomously identified epidemiologically plausible feature importance hierarchies, providing independent verification that model predictions reflect genuine biological signal recognition rather than data artifacts.

      We invite you to review the corresponding manuscript clarifications.

      “The LSTM network was trained on data from Putian corresponding to the four categories described above, with the time period 2018-2022, and Putian’s 2023 data serving as the internal validation set. To rigorously evaluate model transferability beyond the training context, data from Sanming city, a mountainous subtropical city exhibiting comparable climatic characteristics, influenza seasonality, and public health intervention frameworks to Putian, were acquired for the period January 1, 2023, to December 31, 2023, temporally aligned with the Putian internal validation set. This Sanming dataset constituted our external validation set, facilitating a geographically independent assessment of model generalization within subtropical Chinese environments. The predictive performance of the LSTM algorithm for influenza A/B prevalence in 2023 Putian data was benchmarked against a parallel multivariate ARIMA model with exogenous variables. To ensure a fair methodological comparison, this baseline model was supplied with the exact same meteorological and epidemiological covariate matrix as the LSTM.” (Methods, page 12-13)

      “To respect the temporal dependence inherent in LSTM architectures and avoid data leakage, we adopted a strict chronological out-of-time (OOT) validation strategy. Time-series data from January 1, 2018, to December 31, 2022 (88.26% of the Putian dataset) were used for model training, with 10% reserved during Bayesian optimization as an internal validation subset for convergence monitoring and hyperparameter tuning only. Data from January 1, 2023, to December 31, 2023 (11.74%) were held out as a chronologically internal validation set for final performance evaluation, covering a complete annual cycle. External validation was conducted using concurrent 2023 data from Sanming city to assess spatial generalizability. To prevent distributional leakage, all normalization parameters were derived exclusively from the training set and consistently applied to the validation sets, with predictions subsequently transformed back to the original scale. First, we performed data normalization, a crucial step to ensure that training and test set data are compared on a unified scale. We normalized the training and test set data separately within the range [0, 1]. For the test set normalization, we used the maximum and minimum values from the training set as boundaries. This approach ensured consistency between the normalized test set data and the training set data. After the algorithm conducted predictions on the test set data, we performed denormalization to convert the predicted results back to the original data scale and rounded them to integers. These steps ensured that the final prediction results accurately and objectively reflected the LSTM’s performance in real-world scenarios and provided reliable data for subsequent calculation of evaluation metrics.” (Methods, page 18-19)

      (4) The claim for providing tools for 'early warning' was not validated by analysis and results.

      We are grateful for this incisive critique, which identifies a critical mismatch between our research achievements and the terminology employed in the original manuscript. Upon careful re-examination, we acknowledge unreservedly that the original manuscript’s use of “early warning” terminology overstated our actual research contribution. Our research has constructed and validated a methodologically rigorous LSTM-based influenza forecasting framework demonstrating strong predictive performance and external transferability; however, this constitutes a forecasting framework foundation rather than a fully-validated operational early warning tool ready for direct public health implementation.

      We recognize that genuine early warning tools require additional validation dimensions that our current research does not yet comprehensively address. In response to this important critique, we have implemented three categories of substantive corrections:

      - Correction 1: Comprehensive terminology revision throughout the manuscript. We have systematically revised “early warning” terminology throughout the manuscript, replacing it with more accurate descriptors that precisely characterize our actual research contribution. Specifically: “early warning system” has been revised to “forecasting framework” or “forecasting model”; “early warning tool” has been revised to “predictive modeling foundation”; and “early warning capability” has been revised to “predictive capability supporting future early warning system development”. These terminological refinements ensure that manuscript claims precisely correspond to demonstrated research achievements.

      - Correction 2: Manuscript title revision. We have correspondingly revised the manuscript title to remove “early warning” terminology and accurately reflect the study’s actual contributions: Revised title: “Meteorological Drivers of Influenza A and B Positivity in a Subtropical Chinese City: A Six-Year Surveillance Study Integrating Distributed Lag Non-Linear Models and Deep Learning”. This revised title precisely articulates the study’s actual scope: characterization of meteorological drivers (DLNM contribution), focus on positivity rates (methodological refinement addressing surveillance bias), specification of subtropical context (geographic scope), six-year temporal coverage (data scope), and integration of DLNM and deep learning (methodological framework).

      - Correction 3: Explicit articulation of forecasting framework versus operational early warning tool distinction. We have explicitly articulated the distinction between our current achievements and operational early warning tool requirements in both the Discussion and Conclusion sections, framing future research directions for operational early warning system development.

      We invite you to review the corresponding manuscript clarifications.

      “Importantly, however, this framework should be regarded as a methodological foundation for future operational developments rather than as a deployable early-warning system: routine use in public health practice would require prospective recalibration, integration with operational surveillance infrastructure, and additional validation beyond the scope of the present study.” (Discussion, page 43)

      “In conclusion, this study elucidates the distinct, non-linear meteorological drivers of influenza A and B transmission in a subtropical Chinese urban setting through an integrated dual-stage DLNM-LSTM framework. By adopting influenza positivity rates as the primary outcome and integrating socio-behavioral covariates, including mask-wearing stringency indices, weekly detection volumes, and non-local population proportion indicators, our approach mitigates surveillance-related biases and accommodates pandemic-era nonstationarity. It shows lower forecast error than a covariate-matched ARIMA baseline and provides preliminary evidence of portability within southeastern subtropical China, combining the interpretability of distributed lag modeling with the flexibility of deep learning. The framework offers an interpretable, climate-informed methodological foundation for future operational surveillance developments in subtropical settings. (Discussion, page 47)

      Reviewer #2 (Recommendations for the authors):

      (1) The title is not data-driven in different contexts, including 'early warning'; I was expecting substantial analyses in this direction to assess the 'early warning' in the manuscript. But I hardly found them in the text, merely utter as the implication of understanding the associations between meteorological drivers and influenza in advance. I suggest either revising the title or clarifying the claim by providing significant evidence and its impact on the epidemic onset and intensity. Further, revise 'Subtropical China' as 'a Subtropical Chinese city', as the former one is not accounted under this study.

      We completely agree with this constructive feedback. As detailed in our response to your Public Review weakness (4), we acknowledge that claiming an operational “early warning” system requires extensive real-world feasibility and threshold validations that exceed the scope of our current time-series analysis. We understand that achieving early warning capabilities necessitates the completion of at least three additional validation dimensions, as listed below.

      Dimension 1 - Threshold determination: Early warning systems require explicit thresholds defined by integrating predictions with established epidemiological thresholds, typically via ROC analysis to balance sensitivity and specificity for trigger activation; our framework provides predictions but does not define thresholds.

      Dimension 2 - Deployment validation: Validation of warning timeliness, false alarm control, and integration with existing CDC surveillance architectures is required; our work does not assess these operational dimensions.

      Dimension 3 - Robustness across heterogeneous scenarios: Early warning tools need systematic robustness testing across diverse social environments, public health policy contexts, extreme meteorological events, and co-circulation of emerging pathogens; our results cover 2018–2023 Putian-Sanming but not broader operational scenarios.

      We have modified the title to “a Subtropical Chinese city” and systematically revised “early warning” terminology to “forecasting” throughout the manuscript. These revisions ensure accurate representation of the study’s scope and contribution.

      (2) There are several studies establishing the potential association between influenza and the climatic drivers in several locations across the globe. I couldn't find sufficient text on establishing the rationale of this study from the perspective of existing literature. This should clearly be uttered in the introduction section itself.

      We are grateful for this critique, which echoes your Public Review weakness (1). We fully agree that the original Introduction lacked a cohesive narrative connecting the existing literature to our specific methodological innovations.

      To address this, we have comprehensively rewritten the Introduction section. The revised text now systematically establishes our rationale through a clear logical progression:

      - Acknowledging existing studies on climatic drivers but highlighting the unique challenge of non-linear, year-round influenza transmission in subtropical regions.

      - Identifying the methodological gap: existing models either use DLNM purely for retrospective explanation (lacking prediction) or employ deep learning (LSTM) purely for prediction (lacking epidemiological interpretability).

      - Highlighting the critical failure of current literature to mathematically adjust for the profound non-stationarity and surveillance intensity biases introduced by the COVID-19 pandemic (fluctuating testing volumes and NPIs).

      - Introducing our dual-stage solution: integrating DLNM and LSTM to forecast the Influenza Positivity Rate (rather than raw cases), structurally augmented with masking and testing volume covariates.

      We believe this robust literature review now unequivocally establishes the necessity and novelty of our study.

      (3) In connection with the above point, why the authors required the prospective cohort to assess a historical outcome should be highlighted clearly, which is one of the selling points of the study.

      You astutely highlight one of the core methodological strengths of our study design, and we appreciate the opportunity to emphasize this “selling point.”

      As noted in our response to Reviewer #1 regarding timeline splits, the phrase “prospective cohort to assess historical outcomes” reflects the administrative timeline of our study, but mathematically, the data stream is a continuous, longitudinal ecological surveillance.

      The critical selling point here, which we have now explicitly highlighted in the revised Methods section, is that our “historical” data (2018–2020) was not collected via traditional, unstructured retrospective chart reviews. Instead, it was derived from an already operational, highly standardized public health sentinel surveillance system. Because this system utilized identical clinical case definitions, swabbing protocols, and RT-PCR diagnostic assays continuously from 2018 through 2023, the historical data inherently possesses the high fidelity, standardized quality, and lack of recall bias typically reserved for strict prospective cohorts. We have modified the Ethics statement subsection to explicitly underscore this epidemiological advantage.

      “A major methodological strength of this study lies in its robust, uninterrupted longitudinal data collection framework spanning January 1, 2018, to December 31, 2023. While the analytical timeline encompasses a “retrospective” phase (January 1, 2018 – October 13, 2020) prior to formal ethical approval, and a “prospective” phase thereafter, we emphasize that this distinction represents a purely administrative demarcation regarding the timing of ethical approval. It does not reflect any shift in demographic cohorts, sentinel hospital locations, or data collection methodologies. Importantly, the historical data (2018–2020) were not subjected to the recall biases or misclassification risks typical of traditional retrospective chart reviews. Rather, they were systematically extracted from a continuously operating, highly standardized public health sentinel surveillance network. From the inception of data collection through the end of 2023, the local CDC maintained absolute uniformity in clinical influenza-like illness (ILI) definitions, nasopharyngeal swabbing procedures, and real-time reverse transcription polymerase chain reaction (RT-PCR) diagnostic assays. Consequently, the pre-2020 data possess the high-fidelity characteristics of a strict prospective cohort, ensuring unparalleled longitudinal consistency and mitigating temporal measurement bias across the entire pre-pandemic, pandemic, and postrestriction timeline.” (Methods, page 8-9)

      (4) The authors retrieved the daily data on the cases, which is usually small in number for most of the time, can often be driven by the importation (by population mobility with risk of infections) for particularly in a small location like a city.

      We deeply appreciate this incisive epidemiological observation. We fully agree that in a municipal-scale study, relying solely on daily absolute case counts presents significant mathematical and epidemiological vulnerabilities: absolute numbers can be small, highly stochastic, and susceptible to sudden spikes driven by imported cases rather than indigenous climate-driven transmission.

      Driven directly by your comment, we have implemented two fundamental, structural upgrades to our study design:

      - Shift to Positivity Rates: As detailed in our previous responses, we have entirely abandoned daily absolute case counts. Our DLNM and LSTM models now strictly utilize daily influenza positivity rates (positive cases ÷ total daily tested samples) for influenza A and B as the primary outcome. Positivity rates inherently smooth out the stochastic noise of small daily counts and provide a robust, normalized metric of true transmission intensity.

      - Explicit Modeling of Importation Risk: To directly address the risk of importation via population mobility, we have integrated a novel covariate into our LSTM network: the daily proportion of non-local population/residents tested (labeled as outsiders_proportion in our SHAP analysis). By explicitly feeding this mobility proxy into the deep learning algorithm, the model is now mathematically equipped to contextualize and partial out the influence of imported infections when forecasting local transmission trends.

      (5) The authors have not considered this extrinsic factor in the account and not even discussed it.

      We apologize for previously neglecting this critical extrinsic factor. As outlined in our response to Recommendation 4, we have now explicitly operationalized this extrinsic factor by incorporating daily non-local population proportion (outsiders_proportion) as a dynamic input feature in our revised LSTM architecture.

      Furthermore, to empirically assess the magnitude of this importation risk, we conducted a retrospective analysis of the demographic data spanning our six-year study period. Our descriptive statistics reveal that days where the non-local population accounted for >50% of the daily tested cohort represented less than 1% of the total study days.

      This empirical finding allows us to draw two important conclusions: First, while importation undoubtedly occurs (and is now accounted for by our LSTM covariate), indigenous transmission remains the overwhelmingly dominant driver of the observed epidemic curves in Putian. Second, massive importation shocks are rare enough that they do not systematically skew the overarching climate-disease associations identified by our models. We have thoroughly integrated both the methodological adjustment and this empirical discussion into the revised Methods and Discussion sections.

      We invite you to review the corresponding manuscript clarifications.

      “The study spanned January 1, 2018, to December 31, 2023, integrating four categories of data sources: (i) ILI and laboratory-confirmed cases from seven influenza sentinel hospitals across Putian’s urban and rural areas, ensuring representative coverage of diverse healthcare-seeking populations; (ii) daily meteorological data; (iii) COVID-19 associated public health intervention indicators (mask-wearing stringency indices) recorded from January 2020 onwards; and (iv) demographic mobility indicators (non-local population proportion) obtained from ILI consultation records.” (Methods, page 11)

      “Beyond meteorological factors, our DLNM-LSTM framework systematically incorporated influenza positivity rates as the primary outcome variable, weekly detection volumes, mask-wearing stringency indices (indicator of NPIs during the COVID-19 pandemic), day of the week (DOW, distinguishing weekdays from weekends), and non-local population proportion (defined as the ratio of the number of individuals whose reported residential district at the time of testing lies outside Putian city to the total number of tests) as covariates within both DLNM and LSTM modeling pipelines. This comprehensive covariate framework ensures that observed meteorological associations are estimated after controlling for surveillance intensity, public health intervention status, behavioral healthcare-seeking cycles, and population mobility patterns.” (Methods, page 12)

      Multivariate DLNMs efficiently expose transparent, lag-resolved main-effect surfaces for each meteorological variable, but cannot accommodate high-dimensional interactions among meteorological, autoregressive, and socio-behavioral factors without parameter inflation and severe multicollinearity. The LSTM stage was therefore not intended to replace DLNM inference, but to complement it by learning joint non-linear structure across concurrent covariates.

      Crucially, this LSTM stage explicitly incorporates weekly detection volumes, maskwearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 39-40)

      (6) Further, how could such a small number of cases (which can be sporadic) define the epidemic onset and its uncertainty?

      You are absolutely correct: defining an epidemic onset using a small, sporadic number of absolute daily cases introduces severe statistical uncertainty and false-positive onset triggers. This specific methodological vulnerability was a primary catalyst for our decision to fundamentally pivot our analytical framework from absolute cases to Influenza Positivity Rates.

      Unlike absolute counts, where a jump from 1 to 5 sporadic cases might artificially trigger an “onset” definition, positivity rates provide a continuous, normalized epidemiological signal. By assessing the proportion of positive tests against the total testing denominator, positivity rates mathematically stabilize the variance caused by sporadic daily testing. Consequently, an upward trajectory in positivity rates provides a highly reliable, low uncertainty signal of true epidemic onset and acceleration. The exceptional validation metrics of our revised LSTM model (e.g., MAE of 0.009 for Influenza A positivity rate) demonstrate that utilizing this normalized metric virtually eliminates the noise and uncertainty associated with sporadic small-number counts.

      (7) Figure 3 presents the time series of the cases. I wonder whether the data for these factors and outcomes are daily or aggregated by week/month? I suggest representing it in 9x1 format with a single x-axis to compare, instead of 3x3 format. Authors can refer similar plot in https://doi.org/10.1371/journal.pcbi.1012311 in Figure 1.

      We are extremely grateful for this specific and highly constructive visualization suggestion. To answer your query: the data plotted for both the meteorological factors and the influenza outcomes are indeed daily observations.

      We fully agree that the original 3x3 format severely compromised the readability of this daily data. Following your excellent advice and referencing the suggested literature, we have entirely redesigned Figure 3 into an 8x1 vertically stacked format with a single shared continuous x-axis. This structural upgrade has completely transformed the figure, eliminating the horizontal compression and elegantly exposing the fine-grained, daily temporal alignments between climatic extremes and viral surges. The newly rendered Figure 3 is now much more intuitive and analytically valuable. Due to space constraints and given that the revised Figure 3 has been included in our response to a similar comment from Reviewer #1, we will not reproduce the figure in this response. We invite you to review the revised manuscript or refer to our response to Reviewer #1’s recommendation 1 for Figure 3.

      (8) "Additionally, we plotted the loss function curves for the network on the training and validation sets to monitor LSTM convergence and the risk of overfitting." The authors validated the model with predefined training and validation sets. I suggest providing more details on the techniques and the length of the sets. How are these considerations safe for the assumptions and limitations of the models?

      We sincerely thank you for this crucial request for methodological transparency. You are absolutely correct that the techniques used for data partitioning are fundamental to the safety and validity of time-series modeling assumptions. To strictly respect the temporal dependencies of LSTM networks and avoid future-to-past data leakage, we avoided standard random train/test splitting. Instead, we implemented strict Chronological Out-of-Time (OOT) Three-Tier Validation Architecture.

      We have extensively expanded the Methods and Discussion sections to detail this partitioning scheme, our rationale, and its associated limitations:

      - The Three-Tier Validation Architecture and Length of Sets:

      Tier 1 — Training and Internal Cross-Validation (2018–2022, ~88%): Used for initial model fitting. Within this phase, 10% of the samples were held out during Bayesian optimization as an internal validation fold strictly for monitoring convergence, controlling overfitting (via early stopping), and guiding the hyperparameter search. This internal fold never contributed to the final reported performance metrics.

      Tier 2 — Internal Hold-out Test Set (2023, ~12%): A continuous 365-day block reserved as a strictly chronological hold-out, used solely for final performance evaluation. No information from this period influenced training or hyperparameter tuning.

      Tier 3 — External Independent Validation (Sanming 2023): The full surveillance time series from a geographically distinct subtropical city, providing the strongest evidence of cross-location generalizability.

      - Justification of the 88% / 12% Partition Ratio:

      This specific ratio was deliberately chosen to balance two competing epidemiological and computational considerations:

      Sufficient Training Memory: The model required a multi-year training continuum (2018–2022) to autonomously learn the structural breaks and complex non-stationarities introduced by the COVID-19 pandemic and strict NPIs.

      Epidemiological Gold Standard for Testing: Allocating exactly one year (2023) for testing is the epidemiological gold standard for seasonal infectious diseases. A full 365-day cycle ensures that model performance is evaluated across all seasonal phases (spring peaks, summer lulls, winter rebounds) rather than a biased, partial-year fragment.

      - Methodological Safeguards Against Information Leakage (Safe Assumptions):

      To ensure these partitions were safe for the model’s assumptions, we implemented strict safeguards:

      No Temporal Leakage: Tier 2 strictly follows Tier 1 in calendar time, preserving the sequential integrity assumed by LSTM architectures.

      No Distributional Leakage: All feature scaling and normalization parameters (means, standard deviations, min-max ranges) were derived exclusively from Tier 1 (Training) and applied unchanged to Tiers 2 and 3.

      - Explicit Acknowledgement of Limitations:

      We honestly acknowledge that while fixed chronological partitioning is the methodological standard for LSTM forecasting, a single fixed split point (Dec 31, 2022) fundamentally tests only one structural break. It does not exhaustively probe all potential future non-stationarities. Alternative strategies, such as expanding-window or rolling origin cross-validation, could offer additional robustness characterizations. We have integrated this crucial point into the Limitations section of the Discussion.

      We invite you to review the corresponding manuscript clarifications.

      “To respect the temporal dependence inherent in LSTM architectures and avoid data leakage, we adopted a strict chronological out-of-time (OOT) validation strategy. Timeseries data from January 1, 2018, to December 31, 2022 (88.26% of the Putian dataset) were used for model training, with 10% reserved during Bayesian optimization as an internal validation subset for convergence monitoring and hyperparameter tuning only. Data from January 1, 2023, to December 31, 2023 (11.74%) were held out as a chronologically internal validation set for final performance evaluation, covering a complete annual cycle. External validation was conducted using concurrent 2023 data from Sanming city to assess spatial generalizability. To prevent distributional leakage, all normalization parameters were derived exclusively from the training set and consistently applied to the validation sets, with predictions subsequently transformed back to the original scale.” (Methods, page 18)

      “Secondly, the interpretation of our framework’s predictive performance during the 2023 validation period requires careful epidemiological and methodological contextualization. The year of 2023 represented an anomalous, post-restriction “rebound” period characterized by rapid NPI relaxation and the release of accumulated population-level immunity debt, resulting in an atypical influenza surge that exceeded pre-pandemic peaks. The framework’s high accuracy across this period should therefore be interpreted as evidence of algorithmic agility and adaptive capacity during a highly volatile transitional phase, rather than as definitive proof of long-term predictive validity under a stabilized post-2024 epidemiological regime. Methodologically, while our strict chronological OOT data partitioning prevented temporal information leakage, a critical requirement for LSTM integrity, the reiance on a single, fixed chronological split point (December 31, 2022) intrinsically limits our evaluation to one specific structural break. This fixed-split approach may not exhaustively probe the DLNM-LSTM framework’s resilience against all forms of future epidemiological non-stationarity. Consequently, naive extrapolation of the reported 2023 performance metrics to future surveillance years should be avoided absent prospective recalibration. Future studies should consider employing expanding-window or rolling-origin cross-validation frameworks to provide a more continuous characterization of algorithmic robustness. Continuous integration of accumulating 2024 and 2025 data, combined with adaptive learning architectures capable of detecting regime shifts in real time, will be essential before any operational deployment of this, or similar forecasting frameworks, for routine public health surveillance.” (Discussion, page 44-45)

      (9) The authors considered the DLNM analysis without considering the potential interactions among meteorological factors, which can't be avoided in real-world environmental contexts. I would suggest constructing such models by incorporating more reasonable interaction terms to reflect the complex relationships between these variables. Although the impact of COVID-19 was considered on the outcome of influenza directly.

      We sincerely thank you for raising this vital conceptual point. We completely agree that meteorological factors exhibit physically real interactions under real-world environmental conditions (e.g., the synergistic effect of extreme heat and high humidity on viral viability and aerosol dynamics). We welcome the opportunity to clarify how our analytical framework structurally addresses this precise complexity without compromising mathematical stability.

      Rather than forcing interaction terms into a single statistical model, we designed our dual-stage architecture specifically to create a functional division of labor between the DLNM and LSTM frameworks:

      - Methodological Constraints on Explicit DLNM Interaction Terms: While conceptually appealing, explicitly incorporating pairwise interaction terms across eight meteorological variables within the DLNM stage would generate 28 two-way interaction cross-bases, each with its own non-linear and lag-distributed spline structure. This parameter explosion leads to the “curse of dimensionality,” causing: (a) severe multicollinearity given the strong baseline correlations among weather variables; (b) profound instability of the cross-basis estimates and inflated standard errors; (c) a massive risk of overfitting; and (d) the complete loss of visual interpretability, which is the principal value proposition of DLNM. Consequently, as is standard practice in environmental epidemiology (Gasparrini et al., 2010), we restricted our DLNM stage to isolating interpretable, lag-distributed main-effect exposure–response surfaces.

      - The LSTM Stage as the Engine for Complex Interactions: This is precisely where the deep learning architecture provides its unique methodological value. The LSTM network does not require analysts to manually pre-specify rigid interaction terms. Instead, its multi-layer, non-linear gating architecture is intrinsically capable of autonomously extracting and representing arbitrary, high-dimensional interactions among all input variables simultaneously.

      Therefore, complex meteorological interactions are absolutely not ignored in our study; rather, they are absorbed into and resolved by the LSTM stage to maximize predictive accuracy, while the DLNM stage provides the lag-resolved, interpretable backbone for individual main effects. Our multi-stage variable integration ensures that the limitations of traditional statistical models do not artificially bottleneck the deep learning framework’s capacity to synthesize real-world complexities.

      We have now added explicit paragraphs to both the Methods and Discussion sections articulating this functional division of labor, ensuring maximum methodological transparency regarding how meteorological interactions are accommodated within our framework.

      We invite you to review the corresponding manuscript clarifications.

      “In the first stage, DLNMs characterize subtype-specific, non-linear, and lag-distributed associations between meteorological variables and influenza A and B positivity. In the second stage, a Bayesian-optimized LSTM network integrating meteorological, autoregressive, and socio-behavioral covariates is used for short-horizon forecasting, benchmarked against a covariate-matched multivariate ARIMA model and evaluated in an independent subtropical city (Sanming) as a preliminary test of model portability. Influenza positivity rate is used as the primary modeling endpoint to mitigate testing-related surveillance bias. Our aim is to provide an interpretable, climate-informed forecasting approach for subtropical influenza that can serve as a methodological foundation for future operational surveillance developments.” (Introduction, page 7-8)

      “DLNM was first constructed to screen meteorological factors and other covariates with substantial influence on influenza seasonality. Subsequently, an LSTM neural network was developed within the same covariate system. The integrated DLNM–LSTM framework was employed as methodologically complementary rather than redundant components, leveraging the distinctive strengths of each approach to address different analytical objectives within a unified investigation. Specifically, DLNM models characterize the nonlinear exposure–lag–response relationships between meteorological factors and influenza risk, providing biologically interpretable insights into the temporal structure of weather influenza associations and identifying meteorological factors with statistically and clinically significant effects on influenza dynamics.” (Methods, page 11)

      “Building on the significant non-linear and lagged effects through DLNM analysis of real world environmental exposures, this study further constructed multi-factor influenza A and B prediction LSTM networks. Leveraging a recurrent architecture with non-linear gating mechanisms, these networks are able to automatically capture and represent complex, high dimensional interactions among meteorological variables without the need for manual prespecification. This functional division of labor between the DLNM and LSTM models enhances predictive performance while preserving the interpretability and inferential stability established in the DLNM stage.” (Methods, page 18)

      “A central methodological feature of our framework is the deliberate division of labor between the DLNM and LSTM components. Multivariate DLNMs efficiently expose transparent, lag-resolved main-effect surfaces for each meteorological variable, but cannot accommodate high-dimensional interactions among meteorological, autoregressive, and socio-behavioral factors without parameter inflation and severe multicollinearity. The LSTM stage was therefore not intended to replace DLNM inference, but to complement it by learning joint non-linear structure across concurrent covariates. Furthermore, extensive environmental inputs inevitably introduce severe collinearity, such as the strongly correlated solar radiation and UV index. While traditional multivariate models are highly vulnerable to such overlapping variances, the recurrent, weighted representation learned by the LSTM is comparatively tolerant of such redundancy, allowing broader covariate integration than in previous efforts (Zhu et al. 2022). Crucially, this LSTM stage explicitly incorporates weekly detection volumes, mask-wearing stringency indices, non-local population proportion, and DOW effects alongside meteorological inputs, and uses influenza positivity rates rather than absolute case counts as the modeling endpoint. Together, these design choices are intended to mitigate, rather than fully eliminate, the surveillance-related biases that can distort count-based forecasting during periods of fluctuating testing intensity.” (Discussion, page 39-40)

      (10) In context with the above points, although COVID-19-related variables are included, important confounding factors such as population mobility, vaccination coverage, and school calendar (e.g., school openings/closings) are not adequately considered. Additionally, there is a potential risk of overfitting due to an imbalanced data split-too much data is allocated to the training set, while the validation set is relatively small on the other hand.

      We sincerely appreciate your comprehensive evaluation regarding confounding control and the risk of overfitting. These are highly pertinent methodological concerns, and we have implemented multiple refinements and empirical justifications to address each of them systematically.

      - Comprehensive Control of Confounding Factors

      To address the omitted confounders you rightfully identified, we have substantially expanded our covariate framework in the revised models:

      Population Mobility: We introduced the proportion of the migrant/non-local population as a new quantitative covariate (computed as the ratio of tested individuals reporting non-Putian residential addresses). This explicitly captures the extrinsic transmission pressure and viral importation risk exerted by mobile populations.

      Social/School Routines: To capture cyclical social contact patterns and surveillance reporting dynamics, we incorporated the Day-of-the-Week (DOW) indicator (distinguishing weekdays from weekends).

      Vaccination Coverage & School Calendar (Limitations): We honestly acknowledge that highly granular, municipal-level daily vaccination registry data and official macro-school holiday timelines were unavailable for integration into our daily time-series framework. However, for essential epidemiological context, the overall influenza vaccination coverage in mainland China during the 2018–2023 study window is historically estimated at a mere 2% to 3% of the general population, substantially lower than the 40–60% coverage typically observed in high-income temperate countries. Given this exceedingly low baseline, the population-level confounding contribution of vaccination on our predictive accuracy is expected to be minimal in absolute magnitude. We have now explicitly addressed this specific regional epidemiological context in the Discussion section.

      - Justification of the Data Split Ratio (88% vs. 12%)

      Regarding the perceived imbalance in the data partition, the chronological 88% / 12% split was not arbitrary; it was designed to balance two competing epidemiological necessities:

      Preserving Training Memory: The 88% training block (2018–2022) was strictly required for the LSTM to autonomously learn the multi-year seasonal cycles, the structural breaks induced by COVID-19 NPIs, and the suppressed-regime dynamics.

      The Epidemiological Gold Standard: The 12% testing block equates exactly to the 2023 calendar year (365 days). In seasonal infectious disease forecasting, evaluating performance across a complete, unbroken annual cycle is the epidemiological gold standard, ensuring the model is tested across all phases (spring peaks, summer lulls, winter rebounds) rather than a biased, partial-year fragment.

      - Safeguards Against Overfitting & Empirical Proof (Sensitivity Analysis)

      To definitively mitigate and disprove the risk of overfitting, we implemented a Chronological Three-Tier Validation Architecture:

      During the training phase, we randomly partitioned a 10% internal validation subset exclusively for hyperparameter tuning (via Bayesian Hyperopt) and implementing early stopping to terminate training the moment validation loss plateaued.

      We utilized the full 2023 surveillance data from an entirely distinct city (Sanming) as an external independent validation set. The fact that our model generalized excellently to Sanming is the strongest empirical proof against localized overfitting.

      Finally, to further substantiate model robustness, we conducted a controlled covariateablation sensitivity analysis (new Table 2). If a deep learning model is severely overfitted (i.e., memorizing noise), removing covariates often yields chaotic or random performance changes. However, when we explicitly removed the “mask-wearing stringency indices” or the “weekly detection volumes”, our model’s performance degraded in a predictable, biologically plausible manner (e.g., MAE increased by 33.3% to 100% across subtypes).

      This structurally proves that our LSTM architecture is not achieving artificially robust performance through overfitting, but exhibits appropriate sensitivity to the exact epidemiological features driving true viral transmission.

      We have extensively documented these justifications, safeguards, and limitations in the revised Methods, Results, and Discussion sections.

      We invite you to review the newly added limitation statement paragraph.

      “Thirdly, beyond the surveillance coverage limitation noted above, the aggregate-level nature of the available data further constrained our ability to adjust for individual-level confounders, including personal vaccination status, detailed comorbidities, healthcare seeking behavior, and socioeconomic status. This limitation was particularly exacerbated by the profound epidemiological disruptions during the COVID-19 pandemic, where raw numbers of confirmed cases became heavily confounded by surveillance intensity (i.e., fluctuating testing volumes) rather than solely reflecting underlying viral transmission. Our adoption of influenza positivity rates as the primary modeling endpoint and incorporation of weekly detection volumes, face-covering stringency, and non-local population proportion as dynamic covariates, rigorously mitigated these aggregate-level biases and linked our methodological design directly to the forecasting outcomes. As unequivocally demonstrated by our covariate-ablation sensitivity analyses (Table 2), failing to account for mask mandates and testing volumes leads to severe, mathematically predictable deviations in absolute forecasting accuracy. Furthermore, while daily case counts in a single city can occasionally be small, sporadic, and driven by external importations, our incorporation of non-local population proportion effectively adjusted for these localized importation risks. Consequently, although our findings characterize population-level associations between meteorological factors and influenza activity, they should not be interpreted as evidence of micro-level causal mechanisms at the individual patient level. Ultimately, the DLNMLSTM framework’s robust performance across the non-stationary transition out of NPI policies highlights the absolute necessity of integrating behavioral and virological baseline metrics into future climate-driven predictive surveillance systems.” (Discussion, page 45-46)

      (11) The authors should justify why the baseline model selection was made by comparing the LSTM model only with ARIMA? How the outcomes could be sensitive to other commonly used machine learning methods, such as Random Forest or XGBoost, etc, as a benchmark for their performance.

      We are deeply grateful for this methodologically incisive suggestion, which prompted us to substantially strengthen the manuscript’s benchmarking framework. Recognising the methodological importance of this recommendation, our team initiated the construction of an eXtreme Gradient Boosting (XGBoost) benchmark model in parallel with the Round 1 revision submission, and has now completed its full validation and interpretive analysis. XGBoost, one of the leading non-deep-learning machine-learning frameworks for structured predictive tasks, was implemented using the identical covariate matrix employed in the LSTM. The complete methodology, results, and interpretive analysis have been incorporated as Supplementary Additional File 6, with corresponding updates to the Methods (Study Design and Model Construction), Results, and Discussion sections of the main manuscript.

      Methodological design. The XGBoost model was implemented in R (xgboost package, version 3.2.1.1) and received a covariate set strictly matched to the LSTM: eight meteorological variables, weekly testing volumes, the mask-wearing stringency index, a day-of-week indicator, and the proportion of non-local (migrant) population. Critically, no additional lag features, sliding-window statistics, or autoregressive terms were introduced, ensuring that XGBoost’s information set was identical to the contemporaneous covariate stream consumed by the LSTM at each time step. This design deliberately isolates the predictive contribution of architectural inductive bias, the LSTM’s intrinsic sequential memory versus XGBoost’s memory-free tree-partitioning geometry, from any confounding due to hand-engineered temporal features. The training (2018 – 2022) and validation (2023) partitions and the evaluation metrics (MAE, RMSE, MAPE, SMAPE) followed the LSTM and ARIMA protocols precisely.

      Empirical findings.

      A complete performance matrix including MAE, RMSE, MAPE, and SMAPE for all three models is presented in Supplementary Table S3 (Additional File 6), which confirms the same rank-ordering (LSTM ≪ ARIMA ≈ XGBoost) across all four error metrics.

      Despite receiving identical inputs, XGBoost yielded errors approximately 15-fold higher than the LSTM for Influenza A and 35-fold higher for Influenza B, with errors broadly comparable in magnitude to the ARIMA baseline. Visual inspection of the 2023 forecasting trajectories (Supplementary Figure S6) revealed that XGBoost systematically under-predicted the explosive post-NPI Influenza A rebound in late 2023 and generated largely flat forecasts for Influenza B that failed to reproduce the mid-year peak.

      Interpretation. These findings are mechanistically informative and reinforce, rather than merely confirm, the manuscript’s central methodological claim. Four architectural properties of XGBoost account for the observed performance gap:

      Insensitivity to temporal ordering. Decision-tree splits operate on feature-dimensional thresholds and cannot structurally encode “the present depends on the past” in the manner natively accommodated by recurrent architectures.

      Absence of gated recurrence. XGBoost lacks any mechanism analogous to the LSTM’s forget-input-output gating for propagating, filtering, and dynamically re-weighting historical states, and therefore cannot represent long-range non-linear temporal dependencies.

      Violated independence assumption. Tree ensembles implicitly assume independent and identically distributed (i.i.d.) observations, an assumption fundamentally at odds with the strong serial autocorrelation of epidemiological time series.

      Non-comparable modelling paradigms. Both ARIMA and LSTM are explicit sequential models within a common paradigm, whereas XGBoost belongs to a fundamentally non-sequential family. The LSTM–ARIMA contrast therefore isolates the specific contribution of deep sequential learning within a paradigmatically comparable framework, while the LSTM–XGBoost contrast demonstrates that even a state-of-the-art memory-free non-linear learner cannot substitute for a genuinely sequential architecture under epidemiologically non-stationary conditions.

      The consistent superiority of the LSTM over both a linear statistical benchmark (ARIMA) and a non-linear memory-free machine-learning benchmark (XGBoost), all receiving identical inputs and evaluated on the identical validation window, isolates the recurrent gated architecture itself as the source of the predictive gain. We are indebted to the Reviewer for prompting this three-way benchmark analysis, which has materially strengthened the methodological rigour and interpretive depth of the manuscript.

      This additional benchmark analysis, though completed subsequent to the Round 1 submission, has now been fully integrated into the Round 2 revision at the appropriate positions across the Abstract, Importance Statement, Methods, Results, Discussion, and Conclusion, together with the new Additional File 6. We are grateful to you for prompting this three-way benchmark, which we believe has materially strengthened the methodological rigour and interpretive depth of the manuscript.

      We invite you to review the newly added Discussion paragraphs.

      “The optimized LSTM algorithm demonstrated strong predictive performance, with loss curves for both subtype-specific networks exhibiting favourable convergence, consistent with prior work (Du et al. 2023). To rigorously isolate the predictive contribution attributable to the LSTM’s recurrent architecture, we benchmarked it against two covariate-matched baselines representing distinct methodological paradigms: a multivariate ARIMA model and an XGBoost gradient-boosting ensemble. This design controls simultaneously for linearity (ARIMA→LSTM contrast) and for non-linearity without recurrent memory (XGBoost→LSTM contrast), allowing us to attribute observed performance gains to specific architectural inductive biases rather than to informational asymmetry or model non-linearity in general. The LSTM substantially outperformed both baselines on the 2023 validation window. For influenza A, it achieved an MAE of 0.009, compared with 0.136 for ARIMA and 0.138 for XGBoost, approximately 15-fold reductions. For influenza B, the LSTM yielded an MAE of 0.002, versus 0.049 for ARIMA and 0.070 for XGBoost, 25- to 35-fold reductions. Critically, ARIMA and XGBoost produced errors of comparable magnitude despite their disparate assumptions regarding linearity, and both systematically under-predicted the explosive 2023 post-NPI rebound for influenza A while generating flat or spurious trajectories for influenza B (Supplementary Figures S5–S6). That parallel failure indicates the LSTM’s advantage under pandemic-era non-stationarity derives not from non-linearity per se, but from four architectural properties intrinsic to recurrent gated networks yet absent in tree ensembles: (i) threshold-based, sequence-insensitive splits fail to encode present–past dynamics; (ii) XGBoost lacks forget–input–output gates to propagate and re-weight historical states across time lags; (iii) gradient-boosted trees assume near-independence, contradicting the pronounced temporal autocorrelation in epidemiological time series; (iv) ARIMA and LSTM are sequential models with different functional forms, whereas XGBoost is non-sequential. The joint failure of ARIMA and XGBoost, despite differing non-linear treatment, isolates recurrent sequential memory as the critical architectural feature for forecasting under non-stationarity. We emphasize that this interpretation applies specifically to the present non-stationary influenza forecasting task and does not constitute a general dismissal of gradient-boosted ensembles, which retain state-of-the-art performance across many structured prediction domains. Rather, it highlights that for surveillance time series exhibiting pronounced temporal dependencies and abrupt regime shifts, such as the 2023 post-NPI rebound, explicit sequential memory becomes functionally essential. This mechanistic reading, together with the LSTM’s comparative edge over previously reported ARIMA-based (Li et al. 2024) and LSTM-based influenza prediction models (Zhu et al. 2022), positions our framework as a substantive methodological advance in predictive modeling for climate-sensitive diseases.” (Discussion, page 40-42)

      “Fourthly, our benchmarking strategy was deliberately structured as a paradigmatically layered three-way comparison (linear-sequential ARIMA, non-linear non-sequential XGBoost, non-linear sequential LSTM) rather than an exhaustive algorithmic survey. This design prioritized methodological clarity, isolating the contribution of recurrent sequential inductive bias, over horizontal coverage. Nonetheless, our evaluation does not extend to Transformer-based attention architectures, nor to hybrid ensemble strategies (e.g., LSTM–XGBoost stacking or multi-model Bayesian model averaging) that may offer complementary strengths. Systematic benchmarking against these emerging architectures, alongside prospective recalibration on additional subtropical surveillance streams and operational stress-testing under real-time data latency, will be essential next steps as the framework evolves toward operational deployment.”(Discussion, page 46)

      (12) I was expecting the statement of generalizability in terms of locations with the link to the results of the study. I suggest including such statements in the text along with limitations.

      We sincerely thank you for this highly constructive suggestion. We fully agree that an explicit, results-linked generalizability statement is crucial for appropriately interpreting and safely deploying our forecasting framework. To address this, we have substantially expanded the Limitations subsection within the Discussion to articulate a rigorous, three tier generalizability framework directly linked to our study findings:

      - Empirically Validated Generalization: The successful external validation of our framework in Sanming, a geographically distinct, inland prefecture-level city, constitutes direct empirical evidence of cross-location generalizability within the subtropical south eastern Chinese context. As detailed in our results, the framework retained high operational forecasting accuracy (e.g., Influenza A MAE of 0.015, SMAPE of 0.610) without any retraining on Sanming’s local data, proving that the meteorological exposure–response patterns and LSTM architecture are transferable across similar subtropical climates under analogous public health intervention frameworks.

      - Plausible Inferential Generalization: Beyond the directly validated Putian–Sanming pair, the framework’s architecture is plausibly generalizable to other subtropical cities across adjacent regions (e.g., Guangdong, eastern Guangxi) that share comparable monsoon climates, year-round multi-peak influenza circulation patterns, and similar sentinel surveillance infrastructures.

      - Boundaries of Transferability (Non-generalizability): We explicitly state the boundaries where our specific model parameters should not be directly extrapolated without rigorous local recalibration. These include: (a) regions located in subtropical latitudes but characterized by fundamentally distinct climate regimes (e.g., monsoonal systems with characteristics that are markedly incomparable), where the temperature–humidity coupling structures and seasonality with multiple winter peaks differ qualitatively from those in the study setting; and (b) regions with substantially different public-health intervention landscapes, such as settings with high population-level vaccination coverage or school-closure-based mitigation strategies.

      By explicitly defining these boundaries in the revised Discussion, we ensure transparent communication of the model’s appropriate application scope and methodological limitations.

      We invite you to review the newly added limitation statement paragraph.

      “Furthermore, the Bayesian-optimized LSTM architecture was directly applied, without re-training or hyperparameter adjustment, to both the Putian internal test set and the Sanming external validation set. This unified modeling framework ensures that performance differences between the two evaluation contexts primarily reflect geographic transferability rather than model re-optimization.” (Methods, page 21)

      Finally, it is imperative to explicitly delineate the generalizability of our forecasting framework within geographic and epidemiological contexts. The successful external validation in Sanming provides direct empirical evidence that our LSTM architecture and the identified meteorological thresholds are robustly transferable across the subtropical southeastern Chinese context, sharing comparable monsoon climates, year-round influenza circulation, and standardized public health intervention frameworks. However, we explicitly define the boundaries of this transferability. The specific meteorological coefficients, lag structure, and predictive parameters derived in this study should not be indiscriminately extrapolated to other regions, even within subtropical latitudes, where climatic regimes (e.g., monsoon regimes with markedly non-comparable characteristics) or public health contexts (e.g., higher baseline influenza vaccination coverage or distinct non-pharmaceutical intervention strategies) differ substantially. In such disparate settings, directly applying our pretrained model may yield substantial systematic biases. While the underlying modeling framework remains methodologically transferable, its parameterization requires rigorous local recalibration using region-specific surveillance data. Acknowledging these boundaries ensures that the framework can be deployed more safely and appropriately to realize targeted, climate-sensitive infectious disease forecasts.” (Discussion, page 44-47)”

      (13) The flow of the manuscript should be revised for general readers, for example author spent a lot of text on limitations in general without linking them to the outcomes and study design. The English language has a huge scope for improvement. I suggest paying attention to the presentation of the text for the English language use and continuity.

      We sincerely appreciate your candid feedback on the manuscript’s readability. We recognize that in our previous draft, integrating diverse critiques from multiple rounds of peer review inadvertently resulted in disjointed narrative flows, particularly in the Methods and Limitations sections.

      To address this, we have undertaken a massive structural overhaul and linguistic refinement of the entire manuscript:

      - Logical Flow: The Introduction has been sharply refocused; the Methods section has been streamlined (e.g., moving lengthy PCR protocols to the Supplement); and the Results have been reorganized with clear subheadings.

      - Contextualized Limitations: We completely rewrote the Discussion section. Rather than listing generic limitations, we have deeply anchored them to our specific study design and outcomes. For instance, we now explicitly discuss how our shift to predicting the “Positivity Rate” directly mitigates the “Surveillance Bias” limitation, and how the 2023 “rebound” limits steady-state generalizations.

      - Language Polish: The manuscript has undergone comprehensive editing by a native English-speaking academic expert to ensure grammatical precision, sophisticated vocabulary, and seamless continuity.

      We believe these revisions have dramatically elevated the clarity and scholarly tone of the text.

    1. Author response:

      Response to Reviewer 1:

      We thank the reviewer for their valuable and constructive suggestions.

      (1) The categorization of morphological traits was performed by researchers who are familiar with the taxonomy and morphology of this taxa. We acknowledge that this may introduce some degree of subjectivity, and we will provide photographs of other morphological traits for each species in the Supplementary data.

      (2) In our original analysis, we treated the PCAmix values derived from multiple discrete traits as continuous variables and used the lm () function to test the correlation between these values and net diversification fates (lines 352-357). We agree that a phylogenetic comparative approach would be more appropriate. We plan to re-analyze the data using either glm () or phylogenetic generalized least squares (PGLS) to properly account for phylogenetic relationships. In lines 331-333, we would clarify that this part refers to the HiSSE analysis based on discrete traits, which is independent of the linear regression analysis mentioned above. Nevertheless, we will ensure that both analyses are clearly distinguished and properly described in the revised Methods section. We will update the relevant sections accordingly.

      (3) We will carefully re-examine the entire manuscript and add appropriate measures of uncertainty.

      (4) We agree with the reviewer that two hypotheses are not mutually exclusive. In the revised manuscript, we will rephrase this paragraph to present both possibilities more neutrally.

      (5) We have observed mating behaviors in this group and found that ASE function occurs after the male has successfully grasped the female using its legs. However, we acknowledge that our study did not include direct experiments to quantitatively test the effect of ASE complexity on mating success. Therefore, our discussion in this section is indeed somewhat speculative.

      Response to Reviewer 2:

      We thank the reviewer for their encouraging and constructive comments. We agree that our study lacks direct experimental evidence to explicitly demonstrate the grasping and anti-grasping functions of the various male and female traits in Pseudovelia, and that our interpretations currently rely on comparisons with functional studies in more distantly related taxa. In the revised manuscript, we will explicitly state this limitation and refer to these traits as “putative” grasping or anti-grasping traits throughout the text where appropriate.

      We will also carefully address all minor editorial suggestions, including clarifying terminology, correcting typos, and improving figure legends.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Ramirez Carbo et al. use the powerful M. xanthus spore morphogenesis model to address fundamental mechanisms in coordinated peptidoglycan remodeling and degradation. As peptidoglycan is an essential macromolecule and difficult to study in vivo, the authors use indirect but important methodology. The authors first identify two lytic transglycosylase (Ltg) enzymes necessary for spore morphogenesis using mutant phenotypic studies. They characterize these mutants for their role in coordinating spore morphogenesis induced either in fruiting bodies (starvation-dependent) or in liquid-rich media conditions (chemical-dependent). They conclude from these phenotypic and epistatic analyses that LtgA is necessary for morphogenesis during chemical-induced sporulation, and LtgB appears to be necessary to coordinate LtgA activity by interfering with LtgA function. Under starvation-induced sporulation, the absence of LtgB interferes with the building of fruiting bodies. LtgA does not appear to play a primary role in promoting aggregation into fruiting bodies, nor in degradation of peptidoglycan as assayed by loss of signal in anti-PG immunofluorescence. The authors demonstrate that the purified periplasmic domain of LtgA is highly active in degrading purified PG sacculi in vitro, while that of LtgB is highly reduced (relative to LtgA or lysozyme). The authors use photoactivated mCherry Lyt fusions and PALM to track the fusion protein mobility, which they state correlates with activity as immobilization results from PG binding. They demonstrate that in vegetative cells, a greater proportion of LtgA-PAmCh is more immobile (more active) than LtgB-PAmCh, but that directly after chemical-induction of sporulation, LtgB-PAmCh becomes more immobile (active). These analyses in the partner mutant backgrounds suggest that LtgA-PAmCh is more immobile (less active) in the absence of LtgB, but the reverse is not observed. Finally, the authors demonstrate that overexpression of LtgA in vegetative conditions leads to cell rounding, likely because of uncontrolled PG degradation, while overexpression of LtgB displays no phenotype.

      Strengths:

      This paper capitalizes on a novel spore morphogenesis mechanism to define proteins and mechanisms involved in peptidoglycan reorganization. The authors use the powerful PALM microscopy technique to assess Ltg activity in vivo by assaying for immobility as a proxy for PG binding. The authors elucidate a novel mechanism by which two Ltg's function together- with one (LtgB) seeming to regulate the activity of the other (the primary Ltg).

      Despite some weaknesses, there is no question that this study provides important insight into mechanisms of peptidoglycan remodeling- a difficult but highly impactful area of study with implications for the development of novel therapeutics and the discovery of mechanisms of fundamental bacterial physiology.

      Weaknesses:

      In many places, the authors do not adequately justify interpretations of their assays, leading to some apparently unjustified conclusions. Many of these are minor and may just require citations to demonstrate that the interpretations are justified by previous studies (detailed in recommendations below), but two bigger concerns are as follows:

      (1) It is not clear how the muropeptides listed in Figure 1 were assigned, and it is missing in the methods. In the sporulating conditions, the spectra look like combinations of multiple peaks, and the data, as stated, is not convincing to the non-specialist eye.

      We thank the reviewer for raising this point. We've expanded the Methods section to give a fuller account of how muropeptides were identified. In particular, we now describe the chromatographic separation and the assignment process, which relies on comparison with published data, fragmentation patterns, retention times, and accurate mass values. We acknowledge that the chromatograms show several peaks; nevertheless, only those muropeptides that we could clearly identify by MS/MS analysis were annotated in Figure 1. This clarification is now made explicit in the figure legend in the revised version of the manuscript.

      (2) The observation that the lytB mutant prevents appropriate aggregation into fruiting bodies does not allow the interpretation that the absence of LtgB prevents PG morphogenesis in the starvation-induced sporulation pathway, per se. It is more likely that in the LtgB mutant, the morphogenesis program is not even triggered. This is because signaling proteins and regulators (specifically, C-signal accumulation/activated FruA), which are dependent on increased cell-cell signaling in the fruiting body, do not accumulate appropriately in shallow aggregates. C-signal/FruA are necessary to trigger the sporulation program in FBs. BTW: A hypothesis to explain the indirect effect of ltgB absence on aggregation could be that UDP-precursors are not regulated appropriately (unregulated LtyA (LtgA [sic])??), so polysaccharides necessary for motility are not properly produced.

      Along these lines, fruiting body formation does not equal sporulation, and even "darkened" fruiting bodies can be misleading, as some mutants form polysacchariderich fruiting bodies (that appear dark under certain light conditions in the stereomicroscope) but do not sporulate efficiently. The wording in the text suggests that the authors assume that sporulation levels are normal because fruiting bodies are produced (see specific comments for details).

      We deeply appreciate this question. Seeking the answer, we repeated the fruiting body assay and found that both the ΔltgA and ΔltgB mutants formed dark aggregates that were comparable to wild-type fruiting bodies. However, these “fruiting body-like” aggregates did not contain sonication-resistant spores. Thus, regardless of the signals, either glycerol or starvation, sporulation requires both LtgA and LtgB. We have corrected the mistakes in the first submission.

      (3) The authors repeatedly state that production of spore coat polysaccharides likely affects the PG IP staining (see below), but this is not well justified. A citation is needed if this has already been directly shown, or the language needs to be softened.

      We agree with the reviewer. We have softened our language as “However, we cannot exclude the possibility that the polysaccharide spore coats (Voelz & Dworkin, 1962) hinder antibody access to PG.”

      (4) Better justification for the immobility of Ltg proteins in vivo as an assay for activity may be required. If this is well known in the field, it should be explicitly stated. The authors address this better in the discussion - but still state it is a correlation.

      We elaborated the justification, “Thus, when diffusive enzymes bind to PG, their mobility decreases (Lee et al., 2016; Zhang et al., 2023). For instance, DacB, another PG hydrolase, reduces its single-particle mobility in the conditions where its activity is activated (Zhang et al., 2023). By tracking single fluorescently-labeled enzyme particles, we can approximate their PG-binding in different physiological conditions and genetic backgrounds (Ramirez Carbo et al., 2024, Zhang et al., 2023, Ramírez Carbó & Nan, 2026).”

      We further discussed the correlation in discussion, “The simultaneous occurrence of reduced LtgA mobility and PG degradation during glycerol-induced sporulation indicates that the molecular dynamics of LtgA accurately mirrors its enzymatic activity. Such correlation between decreased particle mobility and increased enzymatic activity applies to many other PG-related enzymes, including multiple PG polymerases in E. coli and the endopeptidase DacB in M. xanthus (Lee et al., 2016, Zhang et al., 2023, Yang et al., 2021).”

      Reviewer #2 (Public review):

      Summary:

      The authors' initial goal was to demonstrate loss of PG during the slow sporulation process of Myxococcus xanthus, with examination of the PG degradation products in order to implicate possible enzymes involved. Upon finding a predominance of LGT products, they examined sporulation in strains lacking each of the 14 candidate LTGs encoded in the genome, leading to the identification of two sporulation-linked LTGs. An extensive characterization of the roles played by these LTGs. One LTG is responsible for the slow sporulation PG degradation, while another is required for the rapid sporulation process. Interestingly, the "slow" LTG seems to provide an important regulatory brake on the rapid enzyme. Single-molecule fluorescent tracking of these enzymes was used to develop a model for their interaction with PG that mimics their observed activity. The rate of PG synthesis activity was also shown to impact the rate of PG degradation, suggesting potential interplay between the synthetic and degradative enzymes.

      Strengths:

      The genetic analysis to identify sporulation-linked LTGs and their effects on growth, sporulation, and spore properties was well done and productive. The fluorescence microscopy to track LTG mobility, presumably tied to activity, produced a convincing argument about the mechanism of regulation of one LTG by another.

      Weaknesses:

      While the impact of LTGs on sporulation was clearly demonstrated, the PG analysis that resulted from the study of LTGs raised some important unanswered questions. The analyses suggest that the PG is degraded to quite small fragments, which would normally be lost during the purification of PG. How these small fragments were thus detected is unclear, and this suggests a more complex story concerning PG metabolism during sporulation. An anti-PG antibody is used to quantify PG in the spores, but it is not made clear what the specificity of this antibody is, and thus whether it would recognize the LTG -altered PG of the spore. The authors suggest a "new mechanism of sporulation" when they have actually simply identified an important factor (PG degradation by LTGs) within a complex "process of sporulation".

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Details on places in the text that could be improved:

      (1) Line 77: more appropriate "homologs of the sporulation genes"

      Corrected.

      (2) Line 137-139. Where are the 14 KO mutant data? Only showing the ones with phenotypes? Needs a citation if published elsewhere.

      The phenotypes of other mutants are shown in the new Fig. S1.

      (3) Line 144 and later: why "ORF"? Are they not homologous to Ltg's??. Also, the orf designation is not in the figure, so it is hard to follow the data the authors are presenting.

      Following the reviewer’s recommendation, we deleted “ORF” and presented the ORF designation in the figure legend.

      (4) Figure 2B - add the y-axis legend to the figure (also S1B).

      Added

      (5) Line 156: What does "same below" mean?

      Deleted “same below”.

      (6) Line 161: emtA is not indicated on Figure 2D - can you reword or add to the legend??

      Changed to “MltE, an LTG encoded by Escherichia coli emtA”

      (7) Line 163: "opposite roles" could be a bit better defined... do you mean ltgA induces rounding and ltgB delays rounding???

      We agree with the reviewer. “opposite” was replaced by “different”.

      (8) Line 172: But did the ltgA mutant make the same number of viable spores (not just FBs)? Can't conclude "the slow sporulation pathway only requires ltgB" if ltgA mutant was not tested.

      We agree with the reviewer. In the revised manuscript, we quantified the starvation-induced spores and found that both LtgA and LtgB are required for forming mature spores. The data are shown in Figure 3 and supplement figures.

      (9) Line 202: might be appropriate to indicate the degree of homology (% identity over protein length).

      Added.

      (10) Line 178, 181, and thereafter: DIC microscopy.

      Corrected.

      (11) Line 182/183 and 190: What is the basis for the conclusion that "polysaccharides sustained unflattened structures"??

      We changed the description to “likely due to the deposition of spore coat polysaccharides that sustained unflattened cell structures (Wartel et al., 2013; Holkenbrink et al., 2014).”

      (12) Line 185/186: What is the basis for the conclusion that ltgB only contained a small amount of PG? If it is based on reduced intensity, it is not obvious from the images presented. Quantitative analysis would better support this statement.

      We changed the description to “Sacculi from the ∆ltgB pseudospores still contained PG but showed lower fluorescence intensity”. We also updated the figure to show the typical fluorescence intensities.

      (13) Figure legend 3 line 245/6: It is not totally clear what the difference is, in that the white arrows are pointing to in ltgA vs ltgB mutant. Is it the proportion of large spherical objects in ltgB? What is the significance of the puncta in A vs. B?

      We updated the description in the text as “While the remaining PG sacculi of these pseudospores were largely spherical, they lost integrity during purification, with many sacculi displaying irregular shapes in the fluorescence channel (Figure 4A).” In the figure legend, we changed the description to White arrows point to the sacculi in irregular shapes.”

      (14) Line 190-192. Again, not really seeing what the authors are identifying to conclude "irregular shapes and ruptures" could authors pin-point more specifically and quantify such structures?

      We updated the description in the text as “While the remaining PG sacculi of these pseudospores were largely spherical, they lost integrity during purification, with many sacculi displaying irregular shapes in the fluorescence channel (Figure 4A).”

      (15) Line 194: if the authors want to make such a strong conclusion, they really should demonstrate this. Anti-spore coat antibodies are available in the field. Or take out these statements.

      We took this statement out.

      (16) Line 210: either they lack PG, OR the antibody can't gain access? How to conclude both? Might be more accurate to say "although we can't rule out that the polysaccharide spore coat prevents access to PG"

      Following the reviewer’s recommendation, we changed the description to “These fruiting body spores lacked PG-specific fluorescence (Fig. 3A), consistent with their markedly reduced PG content (Fig. 1). However, we cannot exclude the possibility that the polysaccharide spore coats (Voelz & Dworkin, 1962) hinder antibody access to PG.

      (17) Line 217-19: This is often stated in the literature, but full glycerol-induced spore maturation requires much more than 2 hours (note the authors are using O/N glycerol induction in Figure 1). And the long starvation-induced sporulation process is likely due to the differential start of sporulation, because the cells don't all enter the aggregate at the same time.... so they are not triggered to induce sporulation at the same time.

      We agree with the reviewer on the first statement and changed “two hours” to “four hours”. For the second statement, we do not completely agree. Much research showed that spore development in fruiting bodies is well synchronized, for example, in (Dworkin & Voelz, 1962), starvation-induced spores did not at 48 h.

      (18) Line 222: What is the evidence that it is really variation in production from the van promoter? Could it also just be due to differences in the cell cycle in the population? The authors may be correct, but should be less definitive about those conclusions since they haven't measured LtgA levels directly.

      We agree with the reviewer. We changed the description to, “Upon induction with 200 μM vanillate, cells exhibited heterogeneous morphology, likely resulting from variations in LtgA expression or differences in cell cycle stages within the population.” Following this suggestion, we also changed the description on murA overexpression, “Similar to the cells that overexpressed LtgA (Fig. 3B), this heterogeneity likely reflects variable murA induction or the unsynchronized growth stages within the population.”

      (19) Line 224: Where is the over-expressed LtgA data in Figure 2A? Is it that the authors are referring to over-expressed murA data, and the point is that overexpression of this kind of gene can lead to veg cell rounding? If the latter, this should be specifically stated in the text.

      We apologize for this mistake. The data were shown in Fig. 3B, rather than Fig. 2A, 3B.

      (20) Line 224: Did the authors test that LtgB is stably overproduced in veg cells? It could be that LtgB is turned over while LtgA is not. From the purified protein blot in Figure 3C, it does look like LtgB contains a degradation product (which is perhaps inhibiting the LtgB activity).

      (21) Line 229: AgmT? Do the authors mean Ltg?

      Corrected.

      (22) Line 238: Is "rate" the right word? Technically, kinetics haven't been measured. Could state "LtgB less active" or "less efficient"?

      Changed.

      (23) Line 255: "ltgB shows a slight increase in expression during slow sporulation". Do the authors mean over the entire dev time course or specifically during the sporulation phase (which will be different timing for DZ2 vs DK1622)?

      We clarified the description as “Consistent with its role in PG degradation, in a microarray-based transcriptome analysis, ltgA transcription was found to increase about twofold during rapid sporulation (4 h) but remain unchanged during slow sporulation (96 h). in the closely related DK1622 strain (Muller et al., 2010). Conversely, ltgB expression gradually rises during slow sporulation, reaching 1.8 times the vegetative level at 96 h, while remaining stable during rapid sporulation (Muller et al., 2010, Munoz-Dorado et al., 2019).”

      (24) Line 278: This needs to be corrected. The authors have shown that production of fruiting bodies is not affected by the fusions. (also in Fig. S1 legend). This is really not the same as sporulation efficiency.

      We changed the description to “the PAmCherry tags did not affect the formation of either glycerol-induced spores or starvation-induced fruiting bodies”.

      (25) Figure S1A: panel A: Is there a deg product partially cut off at the bottom of the gel? (or is this a non-specific cross-reactive band). What is the predicted molecular mass for both proteins with the PAmCherry fusion? What conditions were these lysates generated from: veg prior to glycerol induction? I understand there are probably no antibodies available to LtgA or B, but it is important to note that it is not possible to know if there is simultaneously wt LtgA or B produced (by cleavage and degradation of the mCh fusion). Panel B: Were the differences in l/w between wt and the fusion strains tested to see if there really were no significant differences? Please state in the text (it looks like the data variance is higher in the fusion strains relative to the wt at 1 hr).

      The bands at the bottom of the gel are the running front that appear in both lanes. They are not mCherry because there estimated molecular weight is much lower than that of mCherry (26.3 kDa). To avoid confusion, we cut these bands from the figure. The expression of both fusion proteins was detected from vegetative cells, which was clarified in the legend. The predicted molecular weights of them were provided in the legend too.

      (26) Figure S1C: The ability to make fb is not the same as the production of spores. The authors should test the number of viable spores produced under starvation conditions, if they want to state starvation-induced sporulation is not affected by the fusions.

      We agree with the reviewer. We quantified starvation-induced sporulation in the revised manuscript.

      (27) Line 399: Figure S2 looks at fb formation in the absence of vanillate, not sporulation.

      We changed the description to “cells grown without vanillate progressed normally through glycerol-induced sporulation and starvation-induced fruiting body formation”.

      We also quantified starvation-induced sporulation in the revised manuscript.

      (28) Line 281: How many fold is the ltgA transcript reduced compared to the ltgB? (i.e., If it is 1.2 fold reduced that may not be as worth mentioning as if it was 5-10 fold reduced).

      ltgA transcription in vegetative cells was detected in a microarray (Muller et al., 2010) but not reported in RNAseq (Munoz-Dorado et al., 2019), significantly different from that of ltgB. We pointed this out in the revised manuscript.

      (29) Figure 4B legend line 358. Define D (should it be italicized?).

      Corrected.

      (30) Line 363: Please define how significance was calculated. Is this the p-value?

      We deleted the word “significant”.

      (31) Line 298: How do the authors know immobility is from binding to PG? Provide a reference if this is well-known.

      The rationale and references have been mentioned at the beginning of this section.

      (32) Line 302 and thereafter: suggest "2.62 × 10-2 (plus minus) 2.0 × 10-3 μm2/s" is presented as "2.62 (plus minus) 0.20 × 10-2 μm2/s" for easier reading; switch to past tense (Were not are).

      Changed following the reviewer’s recommendation.

      (33) Line 310: "suggest" not "indicate", because binding of PG was directly tested.

      Corrected.

      (34) Line 341: "confirming it restricts access" seems very strong wording. Suggest: may compete with.

      Changed.

      (35) Line 392: SOME fb are larger- many are significantly smaller.

      Because fruiting body sizes do not reflect sporulation efficiency, we removed this description.

      (36) Figure S2 legend. Leaky expression; or on fruiting body formation.

      Corrected.

      (37) Line 436: stationary phase cells decrease PG synthesis- does this increase PG degradation? Perhaps it does lead to the death phase, which is striking in M. xanthus....

      How do cells die, either through death phase or under antibiotic stresses, is not well understood (Baquero & Levin, 2021) Very likely, cell death is due to the accumulation of oxidative damages (Kohanski et al., 2007) and cell lysis could be a byproduct of cell death, when cells lose control of the enzymes that break PG. While cell death is a great topic to investigate, it is beyond the scope of this study.

      (38) Line 507: washed.

      Corrected.

      (39) Line 524 (514 [sic]) and thereafter: sacculi.

      Corrected.

      (40) Line 587: reference for cell lysis procedure?

      The procedures of cell lysis, column loading and elusion were described in details “…cells were harvested by centrifugation at 6,000 × g for 20 min and lysed by sonication in buffer A (20 mM Tris-HCl pH 8.0, 200 mM NaCl), (Nan et al., 2010, Nan et al., 2006). Proteins were loaded to an NGC™ Chromatography System (BIO-RAD) and 5-ml HisTrap™ columns (Cytiva) and eluted by buffer B (20 mM Tris-HCl pH 8.0, 200 mM NaCl, 500 mM immidazole) (Pogue et al., 2018, Nan et al., 2010).”

      (41) Line 269: by microscopy (or "under THE microscope").

      Corrected.

      (42) Line 271: THE cell/PG.

      “The” added.

      (43) Line 603: OR not and.

      Corrected.

      (44) Line 497 (and elsewhere): Is it really CFU? If determined by OD, then not technically CFU because some cells will not grow into colonies.

      We agree with the reviewer. Cell concentrations were determined by OD. “CFU” was deleted.

      (45) Line 506: washed.

      Same as recommendation (38). Corrected.

      (46) Line 521: min.

      Corrected.

      (47) Line 532: How were peaks assigned?

      We described peak assignment in details in the revised manuscript, “Muropeptides were assigned based on: (i) accurate mass matching to theoretical monoisotopic masses of expected M. xanthus PG building blocks (Bui et al., 2009, White et al., 1968) and (ii) comparison of retention times with those reported in previous analysis with similar PG compositions.”

      Reviewer #2 (Recommendations for the authors):

      (1) If almost all the muropeptides detected in spores are anhydro products of LTGs, then it might be expected that these are all very small peptidoglycan fragments in the spores. If the anhydro units were at the ends of short PG chains, then muramidase digestion would release similar amounts of non-anhydro products, but none are detected. So, is muramidase digestion doing anything to the PG derived from the spores? Is muramidase digestion required to observe the spore muropeptide pattern? A control sample in which muramidase digestion is omitted would answer these questions.

      Muramidase digestion is essential for solubilizing PG into different muropeptides for UPLC analysis. In each sample, both anhydro and non-anhydro products were detected. In our original submission, we pointed out that “The two spore types showed similar profiles of a discernible presence of muropeptides that resembled those found in vegetative cells, albeit in significantly reduced quantities (Fig. 1).” We clarified the PG analysis. “The purification procedure yields only sedimentable PG, as all soluble fragments are removed during the washing steps. The resulting sacculi were then digested with muramidase, and the solubilized muropeptides were analyzed by UPLC (see Materials and Methods). The chromatograms in Fig. 1 reflect the muropeptides released specifically from the sedimented sacculus fraction.” Because muramidase release polysaccharides or disaccharides, so it’s digestion does not release nonanhydro GlcNAc species, which is why we did not see equal amounts of anhydro and non-anhydro products. In our case, over 90% of the polysaccharides and disaccharides contain Anhydro-MurNAc. The dominance of anhydro products indicates that LTGs cut very frequently on glycan chains.

      (2) This also raises the question of how these very small peptidoglycan fragments are even retained in the spores. They would be expected to be lost during spore purification or during PG purification prior to muramidase digestion. How do you even purify sacculi when the spores have no PG chains? One could theorize that the polysaccharide coats hold everything in, but then the PG would be protected from muramidase digestion.

      We thank the reviewer for this question. Our analysis suggests that M. xanthus spores do not lack PG entirely but retain a residual PG mesh that is still crosslinked. Such crosslinked material sediments during PG purification and remains accessible to muramidase, which hydrolyses internal glycosidic bonds and releases anhydro muropeptides. Non-crosslinked fragments would indeed be washed away during purification steps, so the detected muropeptides reflect the structure of this residual sacculus (Fig. 1).

      (3) What is the anti-PG antibody recognizing, the glycan backbone, the peptide side chain, or both? Can the antibody recognize the very short anhydro-containing disaccharides proposed to be predominant in spores? If not, then the PG quantification using the antibody is not accurate.

      The structures recognized by the anti-PG antibodies are unknown. Our samples do not contain small degradation products, which was clarified in the revised manuscript, “To answer this question, we purified cell sacculi and used immunofluorescence and an anti-PG serum (de Pedro et al., 1997) to visualize the remaining PG.” We used immunofluorescence to display the PG scaffolds remained in each sample and we did not perform any quantitative analysis based on the images.

      (4) Lines 325-327: This is a speculative conclusion and should be stated as such, i.e., "may control the pace.." This conclusion could be somewhat more strongly stated at the end of the next section, around lines 345-350.

      Moved following the reviewer’s recommendation.

      (5) Lines 403-404: This first sentence of the discussion seems completely dissociated from the topic of the paper; it should be deleted.

      Deleted.

      (6) Lines 404-405. I am not convinced that these findings "elucidate a new mechanism of sporulation." The study was undertaken because PG degradation was already tied to sporulation in a previous study. Furthermore, I am not sure that PG degradation is a "mechanism of sporulation." It is clearly an important step in sporulation of this species, but is it the driving "mechanism"?

      We tuned down our statement as “Our findings demonstrate that M. xanthus, a nonfirmicute bacterium, relies on PG degradation to change cell shape during sporulation”.

      (7) Lines 510-516 describe PG purification from vegetative cells. How was this process modified for spores?

      We apologize for the confusion. The process was clarified as “For PG analysis, samples were processed as previously described for Gram-negative bacteria (Alvarez et al., 2016; Desmarais et al., 2013). Vegetative cells were harvested at mid-stationary phase by centrifugation (30 min, 8,000 g). Vegetative cells and purified spores (as described in the previous section) were resuspended…”.

      Alvarez, L., Hernandez, S.B., de Pedro, M.A., and Cava, F. (2016) Ultra-Sensitive, High-Resolution Liquid Chromatography Methods for the High-Throughput Quantitative Analysis of Bacterial Cell Wall Chemistry and Structure. Methods Mol Biol 1440: 1127.

      Baquero, F., and Levin, B.R. (2021) Proximate and ultimate causes of the bactericidal action of antibiotics. Nat Rev Microbiol 19: 123-132.

      Bui, N.K., Gray, J., Schwarz, H., Schumann, P., Blanot, D., and Vollmer, W. (2009) The peptidoglycan sacculus of Myxococcus xanthus has unusual structural features and is degraded during glycerol-induced myxospore development. J Bacteriol 191: 494505.

      de Pedro, M.A., Quintela, J.C., Holtje, J.V., and Schwarz, H. (1997) Murein segregation in Escherichia coli. J Bacteriol 179: 2823-2834.

      Desmarais, S.M., De Pedro, M.A., Cava, F., and Huang, K.C. (2013) Peptidoglycan at its peaks: how chromatographic analyses can reveal bacterial cell wall structure and assembly. Mol Microbiol 89: 1-13.

      Dworkin, M., and Voelz, H. (1962) The formation and germination of microcysts in Myxococcus xanthus. J Gen Microbiol 28: 81-85.

      Holkenbrink, C., Hoiczyk, E., Kahnt, J., and Higgs, P.I. (2014) Synthesis and assembly of a novel glycan layer in Myxococcus xanthus spores. J Biol Chem 289: 32364-32378.

      Kohanski, M.A., Dwyer, D.J., Hayete, B., Lawrence, C.A., and Collins, J.J. (2007) A common mechanism of cellular death induced by bactericidal antibiotics. Cell 130: 797-810.

      Lee, T.K., Meng, K., Shi, H., and Huang, K.C. (2016) Single-molecule imaging reveals modulation of cell wall synthesis dynamics in live bacterial cells. Nature communications 7: 13170.

      Muller, F.D., Treuner-Lange, A., Heider, J., Huntley, S.M., and Higgs, P.I. (2010) Global transcriptome analysis of spore formation in Myxococcus xanthus reveals a locus necessary for cell diberentiation. BMC Genomics 11: 264.

      Munoz-Dorado, J., Moraleda-Munoz, A., Marcos-Torres, F.J., Contreras-Moreno, F.J., MartinCuadrado, A.B., Schrader, J.M., Higgs, P.I., and Perez, J. (2019) Transcriptome dynamics of the Myxococcus xanthus multicellular developmental program. Elife 8.

      Nan, B., Liu, X., Zhou, Y., Liu, J., Zhang, L., Wen, J., Zhang, X., Su, X.D., and Wang, Y.P. (2010) From signal perception to signal transduction: ligand-induced dimeric switch of DctB sensory domain in solution. Mol Microbiol 75: 1484-1494.

      Nan, B., Zhou, Y., Liang, Y.H., Wen, J., Ma, Q., Zhang, S., Wang, Y., and Su, X.D. (2006) Purification and preliminary X-ray crystallographic analysis of the ligand-binding domain of Sinorhizobium meliloti DctB. Biochim Biophys Acta 1764: 839-841.

      Pogue, C.B., Zhou, T., and Nan, B. (2018) PlpA, a PilZ-like protein, regulates directed motility of the bacterium Myxococcus xanthus. Mol Microbiol 107: 214-228.

      Ramirez Carbo, C.A., Faromiki, O.G., and Nan, B. (2024) A lytic transglycosylase connects bacterial focal adhesion complexes to the peptidoglycan cell wall. Elife 13.

      Ramírez Carbó, C.A., and Nan, B. (2026) Using Single-Particle Fluorescence Microscopy to Quantify Substrate Binding of Peptidoglycan-Modification Enzymes. Bio-protocol 16: e5696.

      Voelz, H., and Dworkin, M. (1962) Fine structure of Myxococcus xanthus during morphogenesis. J Bacteriol 84: 943-952.

      Wartel, M., Ducret, A., Thutupalli, S., Czerwinski, F., Le Gall, A.V., Mauriello, E.M., Bergam, P., Brun, Y.V., Shaevitz, J., and Mignot, T. (2013) A versatile class of cell surface directional motors gives rise to gliding motility and sporulation in Myxococcus xanthus. PLoS Biol 11: e1001728.

      White, D., Dworkin, M., and Tipper, D.J. (1968) Peptidoglycan of Myxococcus xanthus: structure and relation to morphogenesis. J Bacteriol 95: 2186-2197.

      Yang, X., McQuillen, R., Lyu, Z., Phillips-Mason, P., De La Cruz, A., McCausland, J.W., Liang, H., DeMeester, K.E., Santiago, C.C., Grimes, C.L., de Boer, P., and Xiao, J. (2021) A two-track model for the spatiotemporal coordination of bacterial septal cell wall synthesis revealed by single-molecule imaging of FtsW. Nat Microbiol 6: 584-593.

      Zhang, H., Venkatesan, S., Ng, E., and Nan, B. (2023) Coordinated peptidoglycan synthases and hydrolases stabilize the bacterial cell wall. Nature communications 14: 5357.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The revised manuscript includes several useful additions, and I appreciate the efforts to clarify parts of the analysis. The dataset remains valuable. However, several key issues raised previously are not yet fully resolved and continue to limit the clarity of the main conclusions.

      (1) I appreciate that the authors guide the reader to the relevant regions in the analysis of chromosome fusions (Fig. 2b). However, these subtelomeric regions are not clearly visualized, making it difficult to compare fused and unfused profiles, even though the conclusions rely largely on visual inspection of them. A more direct comparison between fused and unfused ends, together with quantitative summaries (e.g., binned Red1 enrichment and comparisons with internal regions), would make this experiment more convincing.

      Thank you for this suggestion. Figure 2 – figure supplement 1 now shows Red1 enrichment in 20-kb bins tiling in from the fusion points to more clearly show that there is no significant difference in Red1 enrichment between fused and unfused chromosomes. These data are consistent with the model that Red1 enrichment is not affected by the presence of telomeres and imply that Red1 under-enrichment near telomeres is primarily encoded in cis.

      (2) The SK1/S288c comparison (Fig. 2c) is an excellent approach, but is currently presented just as profiles, which again requires substantial effort from the reader to extract the relevant information. A systematic analysis across all informative chromosome ends-for example, comparing Red1 levels in syntenic regions using binned log2 fold-change-would more directly test the proposed in cis effect (L168) and clarify the contribution and range of Y'-associated effects. Other factors (e.g. distance from chromosome ends) could also be assessed within this framework.

      Thank you. Figure 2 – figure supplement 2 now shows the profiles placed in register using peak distribution. This analysis demonstrates that registered S288c and SK1 profiles have the same enrichment of Red1, indicating that there are no detectable long-range effects of Y’ elements or other telomere-associated sequences on the neighboring axis binding sites. Figure 2 – figure supplement 3a further quantifies the effect of registering profiles and separates the data based whether Y’ elements are present. These data are consistent with the interpretation that the presence of Y’ elements primarily affects the average axis protein enrichment profiles by displacing strong axis protein binding sites towards the chromosome interior. Our analyses also indicate that this effect is not limited to the Y’ elements as other telomere-associated sequences have a very similar effect on axis protein distribution near chromosome ends.

      Related to this, it is unclear if Y' elements themselves exhibit lower Red1 binding than the genome average. Providing the mean Red1 signal per Y' element would clarify this point and may also aid interpretation of the relationship between coding density and Red1 enrichment.

      Figure 1 – figure supplement 4a and Figure 2 – figure supplement 3b now show that the mean Red1 enrichment on Y’ elements is on average lower than in the rest of the genome. However, as shown in Figure 2 – figure supplement 3b, this effect is not unique to Y’ elements as other telomere-associated sequences show a very similar level of depletion.

      (3) The Dot1-Sir3 section is now simpler. However, I still find it difficult to follow the underlying rationale. In particular, it is unclear why a Dot1 function dependent on H3K79 methylation is introduced, given that the data in the previous section suggest H3K79 methylation is dispensable for subtelomeric Red1 depletion. A clearer statement of the authors' working model would be helpful.

      We apologize for this confusion. We restructured this section in an attempt to clarify the link between Dot1 activity and Sir3.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Raghavan and his colleagues sought to identify cis-acting elements and/or protein factors that limit meiotic crossover at chromosome ends. This limitation is important for avoiding chromosome rearrangements and preventing chromosome mis-segregation.

      By comparing protein axis recruitment in SK1 and S288C background, which differ in their number and distribution of Y' elements, the authors show that Y' element have a limited impact on axis protein enrichment. Genetic analyses coupled with ChIP experiments revealed that the differential binding of the Red1 protein in subtelomeric regions requires the methyltransferase Dot1. Interestingly, the lack of Red1 depletion in subtelomeric regions in this mutant does not impact DSB formation. Another surprising finding is that deleting DOT1 has no effect on Red1 loading in the absence of the silencing factor Sir3. Unlike Dot1, Sir3 directly impacts DSB formation, probably by limiting promoter access to Spo11. As now clearly stated in the abstract and the discussion, this explains only a small part of the low levels of DSBs forming in subtelomeric regions and the main mechanisms suppressing crossover close to the ends of chromosomes remain to be deciphered.

      Strengths:

      This work provides intriguing observations, such as the impact of Dot1 and Sir3 on Red1 loading and the uncoupling of Red1 loading and DSB induction in subtelomeric regions.

      The separation of axis protein deposition and DSB induction observed in the absence of Dot1 is interesting because it rules out the possibility that the binding pattern of these proteins is sufficient to explain the low level of DSB in subtelomeric regions.

      The demonstration that Sir3 suppresses the induction of DSBs by limiting the openness of promoters in subtelomeric regions is convincing.

      Weaknesses:

      The section examining the impact of Dot1 and Sir3 remains complex, which is partly inherent to the intricate relationship between Dot1 and Sir3. However, the authors conclude that Dot1 acts independently of its catalytic activity based on the phenotype of the H3K79R mutant phenotype. Although this is possible it is not fully demonstrated as the H3K79R mutant may exhibit its own phenotype independently of Dot1. Unless the authors test the impact of the catalytic dead mutant Dot1-G401R on axis protein enrichment at subtelomeres they cannot claim that Dot1 act independently of its catalytic activity.

      Thank you. We softened the relevant statements and do not invoke Dot1 catalytic activity.

      Sir3's impact on DSB induction is compelling, yet it only accounts for a small proportion of DSB depletion in subtelomeric regions. Thus, the main mechanisms suppressing crossover close to the ends of chromosomes remain to be deciphered.

      We explicitly state the fact that further regulation remains to be discovered in the abstract, results, and discussion.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors used high-speed atomic force microscopy (HS-AFM) to study the impact of VBIT-4 on VDAC1 oligomerization in real time at nanoscale resolution. Toward this end, they adsorbed POPC:POPE:cholesterol membranes reconstituted with or without VDAC1 on mica. This revealed that the addition of VBIT-4 produced small perforations in the bilayer that were independent of VDAC1. In the absence of VBIT-4, VDAC1 showed the characteristic honeycomb topography that the authors described in a previous study (Reference 17). To quantitatively assess whether VBIT-4 affects VDAC1 organization, they analyzed protein compaction within clusters using inter-protein distance measurements. This analysis revealed no significant difference in VDAC1 organization between control conditions, 1 uM and 10 uM VBIT-4, supporting a model in which VBIT-4 primarily perturbs the lipid matrix rather than VDAC1 assemblies. This conclusion is based on the assumption that VDAC channels retain some lateral mobility in bilayers adsorbed onto mica. Do the authors have evidence that this is indeed the case? Did they also perform HS-AFM on VDAC1-containing membranes treated with VBIT-4 prior to adsorption onto mica?

      We thank the reviewer for this important question regarding the lateral mobility of VDAC1 in supported lipid bilayers (SLBs).

      It is well established that membrane proteins can retain lateral mobility in SLBs formed on mica. This is enabled by the presence of a thin interstitial water layer (typically on the order of 10–20 Å) between the substrate and the bilayer, which reduces frictional coupling and preserves membrane fluidity. This property is a key advantage of SLB systems and has been extensively described [1].

      Furthermore, HS-AFM studies provide experimental evidence supporting such mobility. For example, Casuso et al. [2] showed that OmpF trimers—whose extracellular domain is larger than that of VDAC1— exhibit measurable lateral diffusion in supported membranes. This indicates that even relatively bulky membrane proteins are not immobilized by the mica support.

      In the specific case of VDAC1, our data provide direct evidence of mobility. As shown in Figure 4J of Ref. 17, individual VDAC1 pores display lateral displacements within the membrane, despite the presence of strong protein–protein interactions. This observation indicates that VDAC1 is not rigidly immobilized upon adsorption.

      Regarding the reviewer’s question about VBIT-4 treatment prior to membrane adsorption onto mica, we also performed experiments in which VDAC1-containing proteoliposomes were pre-incubated with VBIT4 before deposition onto mica and formation of supported lipid bilayers. These samples were not imaged at sufficiently high resolution to allow the same quantitative analysis of VDAC1 cluster organization as performed for the experiments shown in the main text. However, at the resolution obtained, we did not observe obvious large-scale changes in membrane organization compared with samples in which VBIT-4 was added after supported bilayer formation.

      Finally, in light of VDAC1 mobility observed under our experimental conditions and the well-established properties of SLBs, the mica support does not constitute a major limiting factor for lateral diffusion.

      Reviewer #2 (Public review):

      (1) The main limitation is that the conclusion that VBIT-4 does not affect VDAC1 oligomerization is strongest for the specific readouts used here: atomic force microscopy measurements of cluster compaction, VDAC1 channel properties, and simulated assembly behavior. These are direct and informative measurements, but they are not identical to the chemical cross-linking readouts used in much of the prior VBIT-4 literature. Readers should therefore distinguish between VDAC1 cluster organization in membranes, as measured here, and cross-linking-defined VDAC1 proximity.

      We agree with the reviewer that AFM-defined VDAC1 cluster organization and cross-linking-defined VDAC1 proximity are related but non-equivalent readouts, and we have revised the manuscript to make this distinction explicit. However, this distinction also highlights an important limitation in interpreting changes in cross-linking efficiency as direct evidence of altered VDAC1 oligomerization. Chemical cross-linking primarily reports the proximity and accessibility of reactive residues and does not directly provide information on the number, size, stability, or supramolecular organization of VDAC1 assemblies. Moreover, VDAC1 organization is strongly influenced by the lipid environment, as shown in giant proteoliposomes with reconstituted VDAC1 by fluorescence correlation spectroscopy [3] and AFM [4], and changes in protein spacing or orientation within dynamic lipid–protein clusters could alter cross-linking efficiency without disrupting the assemblies themselves.

      Cross-linking can provide valuable information when interpreted in the context of independently defined oligomeric structures or interfaces, as recently illustrated by Takeda et al. for yeast Por1 [5]. However, in most studies reporting inhibition of VDAC1 oligomerization by VBIT-4, the evidence relies primarily on SDSPAGE analysis of chemically cross-linked species, frequently quantified as changes in VDAC1 dimers. In contrast, our HS-AFM measurements directly assess the spatial organization and compaction of VDAC1 assemblies in lipid membranes, while our simulations independently assess assembly behavior. Although these approaches do not measure cross-linking efficiency, neither reveals measurable disruption of VDAC1 assemblies by VBIT-4 under the conditions tested. Our observations are therefore difficult to reconcile with the interpretation that VBIT-4 inhibits VDAC1 assembly formation. Instead, we propose that previously reported changes in cross-linking efficiency could reflect changes in protein proximity, orientation, or residue accessibility within dynamic lipid–protein clusters, particularly given the membrane-perturbing properties of VBIT-4 demonstrated here, rather than disruption of VDAC1 assemblies themselves.

      We have therefore revised the Discussion to acknowledge that these approaches probe distinct aspects of VDAC1 organization, while also clarifying that changes in cross-linking efficiency alone cannot be interpreted as direct evidence that VBIT-4 inhibits VDAC1 assembly formation.

      “Cross-linking—the assay most commonly used to monitor VDAC1 “oligomerization”—does not report on oligomer number or stability but rather on the proximity of proteins within these adaptable clusters in MOM. In contrast, HS-AFM directly measures the spatial organization of VDAC1 assemblies in lipid membranes, while molecular dynamics simulations provide an independent description of assembly behavior. These approaches therefore probe related, but non-equivalent, aspects of VDAC1 organization. The organization of VDAC1 in the MOM is extremely sensitive to lipid composition [17]; any hydrophobic compound that perturbs membrane properties may therefore influence cross-linking efficiency through changes in protein spacing, orientation, or residue accessibility, without necessarily altering the overall organization of VDAC1 assemblies. Accordingly, although our results do not directly address cross-linking efficiency, AFM quantification (Supplementary Figure 2) and molecular dynamics simulations (Supplementary Figure 8) consistently show that VBIT-4 does not measurably alter VDAC1 cluster compaction or prevent assembly formation under the conditions examined.”

      (2) A second limitation is the uncertainty around effective VBIT-4 concentration. Because VBIT-4 is poorly soluble, aggregation-prone, pH-dependent, membrane-partitioning, and storage-sensitive, nominal added concentration may differ substantially from the concentration of active compound available in each assay. This complicates comparisons across the different in vitro, simulation, cellular, and previously published assays.

      We thank the reviewer for this important comment. We agree that the nominal concentration of VBIT-4 does not necessarily reflect the effective concentration of active compound available in solution or within lipid membranes. We have therefore expanded the Discussion to explicitly distinguish nominal from effective concentration and to emphasize that, because of VBIT-4's poor solubility, aggregation, membrane partitioning, pH-dependent protonation, and limited stability during storage, the effective concentration cannot be readily determined or compared across experimental systems. “As a consequence, the nominal concentration of VBIT-4 added to an experiment is unlikely to correspond to the effective concentration of active compound available in solution or within lipid membranes. The effective concentration is expected to vary substantially with pH, storage conditions, formulation, and membrane composition, complicating direct comparisons between different assays and across studies. “

      Importantly, to facilitate comparison with the existing literature, we deliberately used the same nominal concentration range as previous studies investigating VBIT-4. Thus, although the effective membrane concentration is inherently uncertain, this limitation applies equally to previous studies using VBIT-4 and represents an intrinsic limitation of the compound rather than of our experimental approach. Measuring the effective membrane concentration is currently not feasible and is beyond the scope of the present study. We further note that this intrinsic uncertainty likely contributes to the variability observed between published studies.

      (3) The coarse-grained simulations provide a coherent mechanistic framework for membrane partitioning, aggregation, and defect formation. However, the VBIT-4 coarse-grained model is newly parameterized and is used to support a quantitative partitioning argument. The manuscript would be easier to interpret if the coarse-grained-derived partition coefficient were reported with uncertainty, convergence information, and protonation state, and compared with a matched all-atom octanol-water partition estimate from the same atomistic model used to build the coarse-grained mapping. This matters because the partitioning argument is used quantitatively to relate micromolar aqueous VBIT-4 to millimolar concentrations in the bilayer.

      Following the Reviewer’s comments, we have expanded the Methods section to add further detail and references on the transfer free-energy calculations and include the requested details. The protonation state used throughout is neutral VBIT, as alchemical free energy calculations of charged solutes require additional corrections, and reliable reference logP values for charged species are scarce given that most empirical predictors are parameterised for neutral molecules; this is the standard Martini pathway for nonbonded term validation.

      The calculated octanol/water logP for neutral VBIT, averaged over three independent replicates, is 3.52 ± 0.01 (replicate values: 3.54, 3.53, 3.51). Convergence and overlap diagnostics confirmed well-sampled simulations across all lambda windows; the corresponding forward/backward convergence plots and MBAR overlap matrices are provided in the Supporting Information.

      Regarding the suggestion to compare against a matched atomistic free-energy calculation: we chose not to pursue this route, as atomistic MD-based logP estimates are not a more reliable reference than empirical predictors for this purpose. Benchmark studies have shown that empirical consensus methods generally outperform atomistic free-energy calculations when compared against experiment, while being substantially less computationally demanding (See [6,7]). We therefore benchmark against a consensus of five established empirical logP predictors (iLOGP, XLOGP3, WLOGP, MLOGP, SILICOS-IT via SwissADME), yielding a consensus logP of 3.42 ± 0.63 for neutral VBIT, in good agreement with our CG estimate of 3.52 ± 0.01.

      We replaced the following manuscript text in the Methods section:

      “These choices were validated by estimating CG octanol/water partitioning free energies, which were compared to predictors obtained via SwissADME[80] (iLOGP[81], XLOGP3[82], WLOGP[83], MLOGP[84], SILICOS-IT). The calculated partitioning free energies were obtained by thermodynamic integration as described elsewhere[77].”

      by the following:

      “These choices were validated by calculating CG octanol/water partitioning free energies for neutral VBIT and comparing them against a consensus of reference values obtained by theoretical predictors. Rather than comparing our Martini logP measurements against atomistic molecular dynamics calculations, we benchmark against established logP prediction methods. For equilibrium octanol/water partitioning, empirical predictors have generally demonstrated accuracy comparable to, or better than, atomistic free energy calculations when evaluated against experiment [6]. We further use a consensus of five empirical models, as consensus predictions have been shown to outperform individual predictors [7]. Reference logP values for neutral VBIT were obtained using the prediction methods available through SwissADME [8] (iLOGP [9], XLOGP3 [10], WLOGP [11], MLOGP [12], SILICOS-IT), yielding predicted logP values of 3.66, 3.85, 4.06, 2.25, and 3.28, respectively. The average of these predictions was used as a consensus estimate, giving a logP value of 3.42 ± 0.63.”

      “The calculated CG partitioning free energies were obtained for neutral VBIT by thermodynamic integration as described elsewhere [13]. In short, the solute is alchemically decoupled from each solvent environment independently across 12 lambda windows, gradually turning off all non-bonded interactions between the solute and its surroundings. The free energy change along this path corresponds to the solvation free energy in that solvent, and taking the difference between octanol and water (ΔG_octanol − ΔG_water) yields the transfer free energy, which is converted to a partition coefficient. Three independent replicates were run, yielding octanol-water logP values of 3.54, 3.53, and 3.51, with an average of 3.52 ± 0.01. Convergence and overlap analysis confirmed that all lambda windows were well-sampled and that the free energy estimates were statistically reliable as required per the guidelines for the analysis of free energy calculations [14]; the corresponding forward/backward convergence plots and MBAR overlap matrices are provided in the Supplementary Figures 11 and 12.”

      (4) Finally, the cellular data strongly support VDAC1-independent cytotoxicity, but the lower-dose mitochondrial functional phenotypes were not directly compared between wild-type and VDAC1-knockout backgrounds. VDAC1 independence is therefore more directly established for cytotoxicity than for the lower-dose mitochondrial phenotypes.

      We agree that our cellular data cannot confirm that the mitochondrial effect of VBIT-4 is independent of VDAC1. However, a previous study by Belosludtsev et al. shows a similar decrease in membrane potential upon VBIT-4 treatment due to inhibition of electron transport chain complexes [15]. Following the Reviewer’s advice, we narrowed the wording accordingly in the Results and Discussion sections.

      Overall, this work provides a valuable and timely reassessment of VBIT-4, and its central conclusion will be useful for researchers interpreting studies that use this compound as a probe of VDAC1 function.

      Suggestions for authors:

      (1) Soften categorical statements such as "VBIT-4 does not alter VDAC1 oligomerization" by specifying the tested readouts: VDAC1 cluster compaction, channel properties, and simulated assembly behavior under the conditions used here. A matched cross-linking experiment under the authors' own VBIT-4 handling and concentration conditions could be useful, but is not essential; the essential point is to make clear that cross-linking-defined VDAC1 proximity and AFM/simulation-defined membrane cluster organization are related but non-equivalent readouts.

      We modified the manuscript to soften the tone. In the discussion, we now emphasize that cross-linking and HS-AFM/simulation probe related but non-equivalent aspects of VDAC1 organization, and that differences in cross-linking efficiency previously reported could be due to alteration of protein spacing or orientation.

      These findings demonstrate that VBIT-4 acts by perturbing lipid bilayers rather than through detectable direct modulation of VDAC1,

      Change title: VBIT-4 Does Not Alter VDAC1 Oligomerization

      To “VBIT-4 Does Not Measurably Alter VDAC1 Cluster Organization or Assembly Behaviour”

      This indicates that VBIT-4 neither prevents nor disrupts VDAC oligomerization.

      To “This indicates that VBIT-4 does not measurably prevent or disrupt VDAC1 assembly under the simulated conditions.”

      “Cross-linking—the assay most commonly used to monitor VDAC1 “oligomerization”—does not report on oligomer number or stability but rather on the proximity of proteins within these adaptable clusters in MOM. In contrast, HS-AFM directly measures the spatial organization of VDAC1 assemblies in lipid membranes, while molecular dynamics simulations provide an independent description of assembly behavior. These approaches therefore probe related, but non-equivalent, aspects of VDAC1 organization. The organization of VDAC1 in the MOM is extremely sensitive to lipid composition [17]; any hydrophobic compound that perturbs membrane properties may therefore influence cross-linking efficiency through changes in protein spacing, orientation, or residue accessibility, without necessarily altering the overall organization of VDAC1 assemblies. Accordingly, although our results do not directly address cross-linking efficiency, AFM quantification (Supplementary Figure 2) and molecular dynamics simulations (Supplementary Figure 8) consistently show that VBIT-4 does not measurably alter VDAC1 cluster compaction or prevent assembly formation under the conditions examined.”

      (2) More explicitly distinguish nominal added VBIT-4 concentration from effective available concentration, given the solubility, aggregation, pH-dependence, membrane partitioning, and storagesensitivity observations.

      We modified the Discussion (see public review)

      (3) Report the coarse-grained-derived octanol/water partition coefficient or transfer free energy numerically, with uncertainty, convergence information, and protonation state. Consider providing the corresponding all-atom octanol/water transfer free energy or partition coefficient for the same protonation state(s).

      We answered this comment and added two Supplemental figures 11 and 12. (see public review)

      (4) Either repeat the oxygen consumption rate, TMRM, and Rhod-2 assays in VDAC1-knockout cells, or narrow the wording so that only cytotoxicity is described as directly shown to be VDAC1-independent.

      Thank you for pointing this out. Following the Reviewer’s advice, we narrowed the wording accordingly in the Results (suppression of “This demonstrates that the cytotoxicity is due to a loss of membrane integrity.”) and Discussion sections.

      We removed the direct link to VDAC1: At concentrations below 10 μM, VBIT-4 decreased mitochondrial calcium, respiration, and membrane potential in HeLa cells without affecting mitochondrial mass. They align with reports that VBIT-4 also accumulates in the mitochondrial inner membrane, where it inhibits respiratory complexes I, III, and IV and decreases mitochondrial membrane potential [15,16].

      (5) Consider moving the storage-stability observation into the main text, given its likely importance for interpreting variability in the broader VBIT-4 literature. It would also be useful to include clearer information on stock age, storage temperature, freeze-thaw history, solvent conditions, and whether precipitation or turbidity was observed.

      The Supplemental Figure 9C was moved the main text as new Figure 6, and additional information about storage conditions is added to the Methods section.

      (6) Minor correction: the parenthetical "10^3.5 = 3.2" should be corrected. Since 10^3.5 is approximately 3,162, the intended statement appears to be that a 1 µM aqueous concentration corresponds to approximately 3.2 mM in the bilayer.

      Thank you for pointing it out, it is corrected.

      References:

      (1) Castellana, E. T. & Cremer, P. S. Solid supported lipid bilayers: From biophysical studies to sensor design. Surf. Sci. Rep. 61, 429–444 (2006).

      (2) Casuso, I. et al. Characterization of the motion of membrane proteins using high-speed atomic force microscopy. Nat. Nanotechnol. 7, 525–529 (2012).

      (3) Betaneli, V., Petrov, E. P. & Schwille, P. The role of lipids in VDAC oligomerization. Biophys. J. 102, 523–531 (2012).

      (4) Lafargue, E. et al. Lipid composition of the membrane governs the oligomeric organization of VDAC1. 2024.06.26.597124 Preprint at https://doi.org/10.1101/2024.06.26.597124 (2024).

      (5) Takeda, H. et al. Oligomer-based functions of mitochondrial porin. Nat. Commun. 16, (2025).

      (6) Işık, M. et al. Assessing the accuracy of octanol–water partition coefficient predictions in the SAMPL6 Part II log P Challenge. J. Comput. Aided Mol. Des. 34, 335–370 (2020).

      (7) Calculation of molecular lipophilicity: State‐of‐the‐art and comparison of log P methods on more than 96,000 compounds - Mannhold - 2009 - Journal of Pharmaceutical Sciences - Wiley Online Library. https://onlinelibrary.wiley.com/doi/10.1002/jps.21494.

      (8) Daina, A., Michielin, O. & Zoete, V. SwissADME: a free web tool to evaluate pharmacokinetics, druglikeness and medicinal chemistry friendliness of small molecules. Sci. Rep. 7, 42717 (2017).

      (9) Daina, A., Michielin, O. & Zoete, V. iLOGP: a simple, robust, and efficient description of noctanol/water partition coefficient for drug design using the GB/SA approach. J. Chem. Inf. Model. 54, 3284–3301 (2014).

      (10) Cheng, T. et al. Computation of octanol-water partition coefficients by guiding an additive model with knowledge. J. Chem. Inf. Model. 47, 2140–2148 (2007).

      (11) Wildman, S. A. & Crippen, G. M. Prediction of Physicochemical Parameters by Atomic Contributions. J. Chem. Inf. Comput. Sci. 39, 868–873 (1999).

      (12) Lipinski, C. A., Lombardo, F., Dominy, B. W. & Feeney, P. J. Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Adv. Drug Deliv. Rev. 46, 3–26 (2001).

      (13) Souza, P. C. T. et al. Protein-ligand binding with the coarse-grained Martini model. Nat. Commun. 11, 3714 (2020).

      (14) Klimovich, P. V., Shirts, M. R. & Mobley, D. L. Guidelines for the analysis of free energy calculations. J. Comput. Aided Mol. Des. 29, 397–411 (2015).

      (15) Belosludtsev, K. N. et al. Effect of VBIT-4 on the functional activity of isolated mitochondria and cell viability. Biochim. Biophys. Acta Biomembr. 1866, 184329 (2024).

      (16) Belosludtsev, K. N. et al. Pharmacological and Genetic Suppression of VDAC1 Alleviates the Development of Mitochondrial Dysfunction in Endothelial and Fibroblast Cell Cultures upon Hyperglycemic Conditions. Antioxid. Basel Switz. 12, 1459 (2023).

    1. Author response:

      We sincerely thank the editors and the three reviewers for their thorough, highly constructive, and positive evaluation of our manuscript. We are gratified by the reviewers’ recognition of the genetic rigor of our study and the compelling nature of our findings regarding the dose-dependent role of Bcl11b in virtual memory CD8 T cell differentiation.

      We agree that the reviewers have raised fair and addressable points that will undoubtedly strengthen the final manuscript. Below, we outline our planned revisions to address the primary themes raised in the public reviews:

      (1) Genomic Targets and Bcl11b Occupancy (Reviewers #1 & #3)

      To address the request for direct genomic targets of Bcl11b, we will incorporate our existing Bcl11b ChIP-seq data from Bcl11b haploinsufficient and control peripheral naïve CD8+ T cells. We will provide comparative analyses demonstrating that Bcl11b ChIP-seq read density per region is highly concordant with population-matched ATAC-seq accessibility profiles, highlighting that direct Bcl11b occupancy closely mirrors the accessible chromatin landscape, which itself we have assayed thoroughly in thymic precursors. Furthermore, we will integrate this ChIP-seq analysis to cross-reference our bulk and pseudobulk differentially expressed gene lists to better define target overlaps.

      (2) Phenotypic Definitions, Cytokine Independence, and Quantification (Reviewers #2 & #3)

      We appreciate the reviewers’ suggestions to further solidify the phenotypic definitions of our populations. In our revision, we will:

      Perform targeted flow cytometry utilizing our Bcl11b haploinsufficient models to provide explicit quantification of CD5 expression across mature thymic (DP to CD8SP) and peripheral CD8 T cell populations.

      Include supplementary flow cytometry panels for CD122 alongside our standard CD44, CD62L, and CD49d gating strategies used throughout the manuscript in order to comprehensively lock down the T<sub>VM</sub> vs. T<sub>CM</sub> phenotypic definitions.

      Provide absolute cell counts (rather than just relative frequencies) for neonatal CD8 populations to explicitly confirm absolute expansion.

      Expand our evaluation of cytokine-independence by analyzing ImmGen-derived cytokine-response signatures against our scRNA-seq datasets.

      (3) Single-Cell Transcriptomic Alignments (Reviewer #3)

      To contextualize our findings within the broader literature, we will score our scRNA-seq datasets against the derived T<sub>VM</sub> transcriptional signatures recently published by Zhang et al. (2024). While exact cluster-to-cluster matching may be limited by differences in experimental models (e.g., steady-state ontogeny versus influenza infection), this alignment will allow us to demonstrate where our newly minted thymic T<sub>VM</sub>(/precursor) cells map along the established peripheral T<sub>VM</sub>-state continuum.

      Additionally, we will generate targeted split-violin visualizations of T<sub>VM</sub> module scores specifically within scRNA-seq Cluster 8. This will visually clarify the transcriptomic shifts driven by Bcl11b haploinsufficiency within this specific cluster, supplementing the DEG/GSEA tables currently provided.

      (4) Conceptual Clarifications and Discussion Expansions (Reviewers #2 & #3)

      We will expand our discussion to address several excellent conceptual points raised by the reviewers:

      Negative Selection vs. Fate Diversion: We will clarify why attenuated Bcl11b alters TCR-induced gene programs without triggering negative selection, emphasizing that Bcl11b dose reduction drives a portional failure of specific transcriptional repression rather than a general deregulation of global TCR-dependent signaling.

      Haploinsufficiency vs. Knockout: We will explicitly contrast our haploinsufficient (<2 fold reduction) virtual memory phenotype with the innate-like T (and ex-T) cell phenotypes previously reported in complete Bcl11b loss-of-function models.

      Mitochondrial Dynamics: We will refine our text regarding mitochondrial biology to more clearly distinguish between compensatory nuclear transcription (the mitochondrial gene module) and physical organelle performance (mitochondrial membrane potential).

      We look forward to submitting the fully revised manuscript and a detailed point-by-point response in the near future.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We appreciate the additional clarifications suggested by the reviewers, and we will include these in the final Version of Record.


      The following is the authors’ response to the original reviews.

      Reviewing Editor Comments:

      The reviewers are very enthusiastic about this study, but have pointed out a central issue: are "space-time attractors" really attractors?

      The reviewers would be willing to increase the assessment of significance if the comments are properly addressed, and in particular, the issue about space-time attractors.

      We thank the editors and reviewers for the feedback on our manuscript and have revised the paper to address their questions and concerns. This document includes (i) an overview of the major changes to the paper, and (ii) point-by-point responses to the reviewers. We have also attached a version of the revised paper that highlights substantial changes to the text.

      Briefly, the reviewers asked for improved intuition about the STA representation and dynamics, and its relationship to attractor networks. To address their questions, we have restructured the paper. It now starts by introducing the STA, which has a representation and connectivity that are both handcrafted. The revised manuscript characterises the resulting dynamics and fixed points in more detail, both empirically and analytically. We then introduce a new model that directly optimises the fixed points of a neural network to represent an explicit plan of the future. The optimal weights for inferring such representations resemble the STA connectivity empirically. Finally, we analyse our unconstrained recurrent neural network, which learns both optimal representations and connectivity. As also shown in the original paper, this network learns to implement an algorithm that closely resembles an STA. Together, these results show that attractor networks can infer PFC-like representations of the future, and this is an efficient solution to dynamic planning problems known to depend on PFC.

      RE1: Expanded theory of STA dynamics

      We have now formalised how the STA relates to a formulation of planning as an inference process over future trajectories, which has been previously proposed in cognitive science and reinforcement learning. We show in the revised paper that the STA dynamics resemble an algorithm for approximate inference in the corresponding probabilistic graphical model. This allows us to characterise the fixed points of the algorithm analytically and relate them directly to a well-established cognitive theory of planning. These analyses help bridge the gap between neural implementation and cognitive computation. They shine new light on previous results in the paper while also providing more intuition for the STA dynamics.

      We have also included a new model that directly optimises the fixed points of an attractor network to resemble a posterior distribution over future locations from planning-as-inference. This analysis complements the handcrafted STA, where we impose both the representation and connectivity, and the RNN, where both the representation and connectivity are learned. The new model imposes (i) an explicit spacetime representation, and (ii) the multiplicative structure of a message passing algorithm. We then train the weights associated with the forward and backward messages by gradient descent on the KL divergence between (i) the true posterior marginals and (ii) the approximate distribution over future locations implied by the network representation at the fixed point. Supplementary Figure S2 of the revised manuscript shows that the optimal weights reflect the transition structure of the environment, similar to the handcrafted STA model and the task-optimised RNN. This makes the connection between attractor dynamics and planning-as-inference more explicit by showing that the fixed points of an attractor network can be optimised directly for planning.

      RE2: Improved characterisation of fixed points

      We have clarified how and why the STA is an attractor network. Attractor networks are defined by the existence of stable fixed points. In ring and grid attractors, there is a continuum of such fixed points in the absence of structured inputs (but often with tonic excitation). In contrast, the STA has a discrete set of input-dependent fixed points. We show explicitly in the revised manuscript how these fixed points depend on the reward inputs to the network, and also how they relate to planning-as-inference.

      We are not claiming that the STA is exactly equivalent to continuous ring and grid attractors. Instead, we want to convey the intuition that the connectivity of the STA constrains the possible fixed points to be plausible trajectories through space and time. The reward inputs determine which of these possibilities is an actual fixed point in a given planning problem. This is not unlike ring attractors in the presence of strong visual inputs. The connectivity enforces a single bump of activity, and the visual input ‘yokes’ the bump to an appropriate orientation. These similarities and differences are highlighted in the revised paper.

      Finally, we have added a new Figure 3 to the main text that characterises the STA fixed points empirically. This figure:

      (a) Shows the evolution of the STA dynamics and convergence to different fixed points in different environments (panels A-B).

      (b) Shows that the network can converge to different fixed points on different trials in the same environment. This happens when there are multiple equally good paths to a goal (panels B-D).

      (c) Shows that other fixed points also exist that correspond to longer trajectories, but the dynamics of the network bias it towards representations of shorter paths. The STA reliably converges to fixed points representing longer trajectories if it is initialised within their basin of attraction (panel F).

      Updated main text:

      “Unlike ring and grid attractors, the fixed points of the spacetime attractor depend on tonic inputs. However, the connectivity constrains the fixed points to represent continuous trajectories for any combination of inputs. In this section, we show this empirically. Later, we will see that such connectivity is optimal for planning-as-inference.

      To compute a plan, it is necessary to know which states will be rewarding in the future. This reward information is provided as an input to the STA and enables fast adaptation without rewiring the synaptic connections. It alters the fixed points of the recurrent dynamics to only include trajectories that are also associated with high cumulative reward (Figure 3; Methods).”

      RE3: Ground truth rewards as an input to the network

      Both reviewers asked about the external input to the STA that specifies the reward available at different states in the future. In reinforcement learning and cognitive science, ‘planning’ is usually defined as the problem of computing a trajectory that maximises cumulative future reward, given an initial state and a reward function (e.g. Mattar & Lengyel, 2022). This is similar to many real-life situations, where we have a known but distant goal (win a game of chess, finish our paper before a deadline, …). When such a reward function is known, it remains challenging to determine the sequence of actions to get there. This has been the topic of much previous work in neuroscience, including (i) the successor representation, which combines a trial-specific reward function with stable transition statistics; and (ii) different types of sequential search, which use a known reward function to evaluate different possible future trajectories.

      To highlight the importance of planning, even when the reward function is known, Figure 4 of the revised manuscript shows that the STA performs better than a greedy baseline that acts according to the immediate reward input instead of planning to maximise cumulative reward. Planning is therefore distinct from learning or inferring a reward function, which is itself a major open question in cognitive science. While undoubtedly interesting, a solution to this problem is beyond the scope of our paper. That is why we decided to simply provide ground truth rewards as an input to the STA. We have clarified the distinction between planning and ‘reward learning’ in the revised paper, and the supplementary material now includes a discussion of where the reward input to the STA could come from.

      Author response image 1.

      All performance quantifications in Figure 4 now include an additional ‘greedy’ baseline (grey bars). This is an agent that acts according to the immediate future reward. The performance improvement of the STA over this baseline highlights the importance of planning to maximise cumulative reward.

      Updated main text:

      “There are several possible sources of reward input to a spacetime attractor (Supplementary Note). We focus on planning under a known reward function and therefore assume access to ground-truth rewards.”

      Reviewer #1 (Public review):

      Summary:

      This work builds a theory to implement planning trajectories towards a goal in a known environment, inspired by analyses of prefrontal neural recordings. Unlike standard neural architectures for this task, such as value-based learning and successor representations, their proposed theory is able to adapt to novel goal locations within a trial. The key to the theory is that future times are represented by orthogonal groups of neurons. The recurrent connectivity between groups of neurons selective to specific future times and locations reflects the learned knowledge of the task. Finally, the authors show that standard networks trained on the task approximate their proposed theory.

      Strengths

      The structure of the work is clear, and the presentation of the results is very well written, which is particularly noticeable given the consequential amount of results presented. The authors are able to link their theory with experimental findings in neural recordings. The reverse-engineering of trained recurrent neural networks is very thorough, by analyzing both dynamics and connectivity. The assumptions and predictions of their model are clearly stated.

      We appreciate the encouraging comments and hope our revised manuscript addresses the reviewer’s questions.

      Weaknesses

      (1.1) It is unclear whether their proposed theory, "space-time attractors", actually is an attractor network. The authors used recurrent neural networks with very few timesteps, and long single neuron time constants with respect to the task time scales. Attractor networks, as the ones the authors cite, refer to networks that generate nontrivial patterns of activity through recurrent interactions, after long periods of time.

      See RE1 & RE2 for a comprehensive response to this question. Briefly, we show in the revised manuscript how the fixed points of the STA dynamics relate to planning-as-inference, and we clarify the similarities and differences between the STA and other attractor networks in the main text. We show in the new Figure 3 that (i) representations of future paths are stable over long periods of time, and (ii) multiple fixed points can exist when there are multiple paths to the goal. It is also worth noting that the RNN representation in Figure 6H remains stable for 75 time constants and recovers from perturbations. This is substantially longer than during training, where ‘planning’ lasted up to 14 network time constants, and it suggests that the network representation is a stable fixed point.

      (1.2) The authors gloss over how the reward inputs are calculated. Computing these reward inputs should be part of the planning process, and the authors are implicitly leaving this problem aside. How does the reward input, which includes future time and location, depend on the actions that have not yet been taken by the agent? It feels like most of the planning computation is already provided by these reward inputs at the beginning of the trial. It could be that the network is only learning to process the planned sequence of actions present in the inputs.

      See RE3 for a comprehensive response to this question. Briefly, ‘planning’ is often defined as the problem of computing a trajectory that maximises future reward, given a reward function, initial state, and transition function. The reward function provided to the agent indicates which future states it would be desirable to reach, but not how to reach them. To make this point clearer, we show in Author response image 1 that the representations computed by the STA generate better behaviour than an agent acting greedily according to the reward function specified by the inputs. This highlights the importance of considering distant goals when choosing immediate actions.

      Reviewer #1 (Recommendations for the authors):

      The text is very nicely written, and I appreciated the way in which methods are presented, with a clear structure and a logical chaining of the different sections. My comments and suggestions refer mostly to the methods and the RNN implementation. Please find below a list of issues.

      Relatively major:

      (1.3) All the equations of the dynamics should be written in discrete time and not in continuous time. There is no notion of "iteration" in continuous time, so it is currently very hard to understand how the RNN works, and what the different epochs are ("the RNN performed 10 network iterations..."?).

      We have rewritten all equations in discrete time and clarified the notion of ‘iterations’.

      (1.4) This is pointed out in the public review, but the authors insist on making an analogy between the spacetime attractor implementation of planning and attractor networks. It seems to me that these two types of models are very different. What defines attractor networks (such as grid- or ring-attractor networks) is that recurrent connections internally generate stable states of activity for long periods of time, in the absence of inputs. Nothing like that is shown here. Robust input-driven trajectories are neither necessary nor sufficient for showing that an RNN is an attractor network. The fact that, given the inputs, networks are run for very short periods of time in this work seems to indicate that this is a very different type of network compared to the attractor networks mentioned previously.

      See RE1 & RE2 for a comprehensive response to this question. Briefly, it is correct that the fixed points of the STA depend on the inputs, which is different from canonical ring and grid attractors.

      We show that the fixed points of the recurrent STA dynamics are reward-maximising paths when conditioned on those inputs. Briefly, we now (i) analytically characterise the input-dependence of the fixed points and show how they relate to planning-as-inference; and (ii) show empirically that STA representations remain stable for long periods of time (Figure 3). The revised paper also clarifies the similarities and differences to previous attractor models.

      Minor

      (1.5) It would be nice to show more clearly what the inputs are in a given trial, and how they change over time in the RNN (specifying the planning and execution phases). A supplementary figure may help.

      How is the information about walls provided to the network exactly? More generally, it would be nice to clearly indicate in the Methods what all the inputs "x" to the RNN are, and how they change over time (during planning, during execution, and how they change as the environment steps are updated).

      We have made a new Supplementary Figure S3 that illustrates the inputs to and outputs from the different models. Briefly, the information about the walls is provided to the RNN as a binary vector x<sub>w</sub> ∈ ℝ <sup>2𝑁</sup>. The elements of this vector indicate for each of the N states whether there is a wall (i) to the right of, and (ii) above it. In each trial, a subset of these is present, and a subset is absent. All ‘present’ walls are assigned a value of +1 in x<sub>w</sub> , and all absent walls are assigned a value of 0 in x<sub>w</sub> . We have also clarified this in the revised Methods.

      (1.6) It would help to clarify, at least in the methods, the shape of all the matrices and vectors that are trained.

      We have clarified the shapes of all matrices and vectors in the Methods.

      (1.7) N in the methods is not defined (I think N = 16, the total number of locations on the grid).

      N is indeed the total number of locations in the state space. This is 16 for almost all analyses in the paper, which involve planning on a 4x4 grid. The updated manuscript includes a few analyses in larger environments, where N is larger. We have clarified this in the Methods.

      Reviewer #2 (Public review):

      This well-written manuscript proposes to use attractors in space and time (STA) as a mechanistic explanation for planning in the prefrontal cortex. The main conceptual hypothesis is that planning is implemented as attractor dynamics in a representation that encodes states at each time step jointly. Depending on inputs, the network relaxes to a trajectory that already contains future states that will be visited at each time step, rather than computing a scalar value at each point in time and space like other classical approaches from RL. The authors compare this approach to implementations such as TD learning and successor representation, and further show that trained recurrent neural networks on specific tasks involving planning develop structured subspaces resembling the ones postulated in STA.

      The idea of treating attracting trajectories unfolding in time as the computational substrate for planning is very interesting and potentially important. The explicit construction of a state x time representational space and its implementation via recurrent dynamics are appealing and convincing in the idealized tasks considered. I found the manuscript to be refreshingly explicit regarding several of the assumptions and limitations of the models, for example, the fact that certain advantages can be viewed as properties of the state space itself and not necessarily of a fundamentally new planning mechanism.

      Overall, the manuscript presents a cool attractor model that extends in time and explores its performance in a subset of illustrative tasks involving planning. My doubts concern mostly the interpretation and scope of the claims made in the manuscript. Here are a few comments where I detail my questions/concerns:

      We appreciate the enthusiasm about the manuscript and its potential importance. We address the remaining questions and concerns below.

      (2.1) The authors nicely discuss that much of the difference between STA and classical TD or SR agents is "in some sense a property of the state space rather than the decision making algorithm," and that TD and SR could in principle be implemented in a comparable space x time representation. This is fair, but it also suggests that the central contribution of the manuscript lies primarily in the representational factorization (state x time tiling) and its dynamical implementation via attractors, rather than in a fundamentally new planning algorithm or theory, mechanistic or not. I think theory should be distinguished from mechanism, and it would therefore help the reader to describe the conceptual advancement more as a novel mechanism or implementation than a novel (mechanistic) theory for decision/planning.

      We respectfully disagree that ‘theory’ has to be distinguished from ‘mechanism’. We do agree that ‘computation’ and ‘mechanism’ can often be distinguished. However, we think theories can live at either of these (and other) levels of explanation. What we propose is indeed a potential mechanism for planning that combines recently characterised prefrontal spacetime representations with attractor dynamics to infer desirable ‘plans’. As we show in the revised manuscript, this mechanism resembles the computation of ‘planning-as-inference’, which has previously been proposed in cognitive science (e.g. Botvinick & Toussaint, 2012). Our theory is therefore not about the computation – it is about the mechanism. The title “A mechanistic theory of planning…” is meant to clarify what level of description our paper addresses.

      As an example of the importance of mechanistic theories, the computation of angular velocity integration can be implemented in many different ways. Seminal work by Skaggs et al. (1994) and others in the 1990s showed how it can be implemented in neural networks, inspired by experimental data. These theories paved the way for detailed experimental characterisations of the fruit fly head direction circuit more than two decades later (Turner-Evans et al., 2017; Kim et al, 2017; and others). Inspired by this and other success stories, we think an important role of theoretical neuroscience is to develop theories about neural mechanisms that can be tested in future experiments!

      (2.2) Related to my previous point, I think it would be helpful to position STA more explicitly relative to computational/theoretical literature in which attractor networks encode temporally ordered patterns (so effectively including future times). For example, classical extensions of Hopfield networks with asymmetric connectivity implement retrieval of sequences and ordered transitions between patterns (Sompolinsky & Kanter, 1986). More recently, sequential attractors and limit-cycle dynamics have been constructed in structured recurrent networks by the Morrison group (Parmelee et al., 2021). These works do not implement an explicit discretized state x future-time tiling as in STA and do not specifically discuss the usage for planning. However, they do provide concrete precedents for attractor dynamics over temporally structured trajectories in terms of mechanism. It would be useful to discuss this literature and clarify a little what's new mechanistically in the view of the authors.

      We agree that this is not the first use of attractor networks to represent or compute sequences. Instead, we show that a combination of spacetime representations with attractor dynamics is sufficient to compute plans in dynamic problems known to depend on prefrontal cortex. As the reviewer points out, the primary difference from most previous work lies in the fact that the entire sequence is encoded in a single fixed point of the STA dynamics. This differs from e.g. Sompolinsky & Kanter, where the population encodes one element at a time and generates sequences as limit cycles. The instantaneous encoding of an entire sequence in the STA is what enables planning through parallel message passing rather than sequential search. This is highlighted in the main text of the revised manuscript, which also includes a supplementary discussion of the similarities and differences between the STA and related work on sequences in attractor networks.

      Updated main text:

      “Entorhinal grid cells are also embed a world model in their connectivity (McNaughton et al., 2006), but they only encode a single location at a time (Vollan et al., 2026). Such networks can generate sequences, but the individual elements are represented one by one (Sompolinsky and Kanter, 1986; Kleinfeld, 1986; Widloski et al., 2025). The spacetime attractor suggests that circuit principles in prefrontal cortex resemble other cortical areas that use structural knowledge to infer features of the world. The major difference is that PFC instantaneously represents many points in time, which generalises known circuit principles to complex planning.”

      (2.3) A central claim of the manuscript is that space-time trajectories are attractors of the STA dynamics. The manuscript does provide empirical evidence consistent with attractor-like behavior. However, it is not explicitly shown whether trajectory representations persist in the absence of sustained external inputs. So it's not clear to me whether the trajectories should be interpreted as intrinsic attractors of the recurrent system, which can be selected by delivering transient inputs, or whether they must be stabilized by a specific continuous external drive. It would be useful if the author could clarify/discuss this point.

      We show in the revised paper that the fixed points of the STA dynamics take the form r <sub>δ</sub> = e<sup>R<sub>δ</sub></sup> ◦ (Ar<sub>δ−1</sub>) ◦ (A<sup>T</sup> r<sub>δ+ 1</sub>) (Methods). Here, r δ is the activity of neurons representing expected locations in δ actions; R δ is the reward function in δ actions; and A is the environment adjacency matrix. These fixed points depend on the reward inputs through the first term. In the absence of reward inputs, the fixed points are ‘diffusive’, while still respecting the transition structure of the environment. In the presence of reward inputs, they concentrate probability mass on trajectories with high expected reward. We have clarified these properties in the main text and introduced a new Figure 3 that characterises the fixed points of the STA in more detail. See also RE1 and RE2.

      (2.4) As far as I understand it, reward information is provided as input to specific populations encoding future time steps, and that's essential for rapid adaptation without rewiring connectivity. How such future-time-specific reward inputs would be generated and routed to distinct neural populations isn't entirely clear to me. Since this seems to be an essential component of the model, I think it would be important to discuss more deeply the source and plausibility of these reward signals related to different timesteps.

      See RE3 for a comprehensive response to this question. Briefly, ‘planning’ is often defined as the problem of computing a trajectory that maximises future reward, given a reward function, initial state, and transition function. The reward function provided to the agent indicates which future states it would be desirable to reach, but not how to reach them (see Author response image 1). We agree that the challenge of estimating future reward is an interesting question, but it is beyond the scope of this paper. We have clarified this distinction in the main text and added a supplementary discussion that speculates about where reward information could originate in biological circuits.

      (2.5) The authors note that vanilla STA scales linearly with planning horizon, and discuss potentially hierarchical extensions for longer horizons. They acknowledge that learning abstractions remains an open challenge, yet the examples of planning in the manuscript are restricted to very short temporal horizons and limited branching complexity. It is not obvious to me in what cases the current implementation and interpretation of STA remains viable (for example, in terms of relaxation iterations) as the horizon and branching factor increase. Relatively simple planning can be managed by simpler, less costly models/algorithms, whereas complex planning is a lot harder to deal with, and it's something that a mechanistic "theory" should address. In the context of the claims of the paper in its present form, I think this is possibly the most important conceptual and practical limitation in the manuscript.

      It is correct that planning gets increasingly challenging with planning depth. In the absence of noise, the STA scales to sequences of up to 12-13 actions – and even longer if the minimum path length is known a priori. Performance gets progressively worse when recurrent activity and parameter noise increase. The revised paper includes a new Supplementary Figure S1A-C that shows how the STA planning ability depends on planning depth for different levels of noise.

      We do not consider planning depth to be a major limitation of the work, since humans are rarely thought to plan much more than 6 steps into the future at a single level of abstraction (e.g. van Opheusden et al., 2023). Instead, we believe that hierarchical planning is used to infer trajectories to distant goals (Eckstein & Collins, 2020). To illustrate this point, we have now implemented a proof-of-principle hierarchical STA in Supplementary Figure S1D-E. This simulation shows how an ‘abstract plan’ inferred by one STA can be treated as a goal to infer a more ‘detailed plan’ in a second STA. In principle, this enables the system to compute plans that are arbitrarily long, provided they can be broken down into chunks smaller than the limits imposed by the analyses in Supplementary Figure S1A-C.

      Finally, RNNs learn an STA-like algorithm when trained on dynamic planning problems with a planning depth of 6. It is therefore not clear to us whether simpler and less costly algorithms can be easily implemented in the dynamics of recurrent networks.

      (2.6) The RNN analyses show that trained networks develop structured subspaces aligned with future time indices and exhibit perturbation behavior consistent with attractor-like dynamics. The manuscript also explicitly notes differences between the trained RNN and the handcrafted STA (e.g., long-range couplings between subspaces and differences in behavior of lower-value trajectories under perturbation), which I much appreciated. My doubt is on the specificity of this result, as trained RNNs on fixed-horizon tasks can develop latent dimensions correlated with temporal progress within a trial or time-to-goal. I think it would help the reader to clarify whether the results demonstrate that STA-like computations emerge in RNNs trained on planning tasks, or that RNNs generally develop some kind of structured spacetime representations when tasks involve future timesteps and some degree of flexibility in the decisions.

      An important point to note is that the subspaces we identify do not encode time-to-goal, since they are all active at the very beginning of the trial. We also show that RNNs trained on simpler static tasks do not learn the same algorithm (Supplementary Figure S7) and do not generalise to dynamic problems (Figure 5F). Finally, other algorithms are capable of solving the dynamic problems we study (e.g. the ‘value agent’ in Figure 5B-E). We therefore do not think it is trivial that RNNs learn an STA-like algorithm.

      We do think that ‘structured spacetime representations’ generally emerge in RNNs trained on tasks that involve flexible behaviour in changing environments – in some sense that is the claim we are trying to make. It is known that spacetime representations are optimal for structured sequence memory tasks (e.g. Whittington et al., 2025; Dorrell et al., 2026), and we think this is for exactly the same reason. In sequence working memory, the reward function changes in time – for each action, the reward is only non-zero at the corresponding sequence element. However, the adjacency matrix is uniform for sequence memory – any sequence element can follow any other sequence element – so there is no need for planning. We are therefore not claiming that spacetime representations only emerge in the specific planning task we consider here. Instead, we expand the set of problems solvable by such representations to also include adaptive planning known to depend on prefrontal cortex. We have made this more explicit in the revised manuscript.

      Updated main text:

      “Together, our analyses show that RNNs trained on a dynamic planning task learn to approximate a spacetime attractor. This was also true across variations in model architecture (Methods; Figure S10; Figure S11). These results extend previous findings that explicit spacetime representations are optimal for sequence memory (Supplementary Note; Whittington et al., 2023; Dorrell et al., 2026; Wang et al., 2025). Additionally, RNNs with too few hidden units to learn a spacetime attractor failed to solve the task (Figure S12), suggesting that other solutions are not readily learned by gradient descent.”

      A few more minor points, mainly concerning clarity:

      (2.7) The main dynamical equation combines a log-domain recurrent term, a floor operation, and a log-sum-exp normalization step, followed by exponentiation. The intuition/logic behind this specific formulation could be clarified for the reader. For example it would be helpful to explain why the recurrent input appears inside a log, and also whether/how these operations relate to any multiplicative constraint.

      The specific form of these equations comes from the intuition that the STA approximates planning as an inference process over future trajectories. We have clarified this in the revised manuscript, which explicitly shows how these equations relate to planning-as-inference as formulated previously (e.g. Botvinick & Toussaint, 2012; Levine, 2017).

      (2.8) While the computational cost of successor representation in an expanded NT x NT representation is discussed, the corresponding scaling of STA in terms of number of units and connections (as a function, for example, of the planning horizon) isn't clear to me. Perhaps the authors could compare costs more explicitly.

      The memory cost of a spacetime-SR would be (NT)^2 and the computational cost (NT)^3 (it is possible that both of these could be reduced by taking advantage of the structured nature of the spacetime successor matrix, but that is beyond the scope of this work). The memory cost of the STA is NT, and the computational cost is (NT)^2 (each iteration of the network dynamics requires the calculation of T matrix-vector products of size NxN, and the number of steps to convergence is approximately linear in T). We have included this comparison in the Supplementary Discussion of the revised paper.

      (2.9) In the RNN analyses, structured subspaces aligned with future time indices are shown. I couldn't find a quantification of how much variance is captured by the subspaces, relative to other latent dimensions. Adding it would help get a feeling for the strength of the alignment.

      We have added a new Supplementary Figure S6 to the revised manuscript, which quantifies the variance explained by the future-coding subspaces over the course of a trial. The variance explained by the K dimensions encoded by these subspaces is substantially higher than a random baseline, and it approaches the upper bound given by the top K PCs. Interestingly, the future-coding subspaces all explain a lot of variance early in the execution period. During later stages of execution, only the ‘immediate future’ subspaces explain substantial variance. This suggests that the RNN only maintains information in subspaces that represent times before the end of the trial.

      References

      Botvinick, Matthew, and Marc Toussaint. "Planning as inference." Trends in cognitive sciences 16.10 (2012): 485-488.

      Dorrell, William, et al. "An Efficient Computing Theory of Prefrontal Structured Working Memory Representations." bioRxiv (2026): 2026-02.

      Eckstein, Maria K., and Anne GE Collins. "Computational evidence for hierarchically structured reinforcement learning in humans." Proceedings of the National Academy of Sciences 117.47 (2020): 29381-29389.

      Kim, Sung Soo, et al. "Ring attractor dynamics in the Drosophila central brain." Science 356.6340 (2017): 849-853.

      Levine, Sergey. "Reinforcement learning and control as probabilistic inference: Tutorial and review." arXiv preprint arXiv:1805.00909 (2018).

      Mattar, Marcelo G., and Máté Lengyel. "Planning in the brain." Neuron 110.6 (2022): 914-934.

      Skaggs, William, et al. "A model of the neural basis of the rat's sense of direction." Advances in neural information processing systems 7 (1994).

      Turner-Evans, Daniel, et al. "Angular velocity integration in a fly heading circuit." Elife 6 (2017): e23496.

      Van Opheusden, Bas, et al. "Expertise increases planning depth in human gameplay." Nature 618.7967 (2023): 1000-1005.

      Whittington, James CR, et al. "A tale of two algorithms: Structured slots explain prefrontal sequence memory and are unified with hippocampal cognitive maps." Neuron 113.2 (2025): 321-333.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Overall, this is an interesting and well-written manuscript on a fascinating question in a "charismatic" model system.

      Strengths:

      (1) The Introduction is concise, though it might be helpful to the non-specialist reader to learn a bit more about what is known about the social control of somatic growth across diverse species (including humans), which would help to make this work more generally interesting.

      (2) The experiment is well-designed.

      (3) The data collected are comprehensive.

      (4) The complementary analysis of both feeding and aggression/submission data with and without known social roles is a neat idea and compelling!

      Thank you for the positive feedback!

      Here, we investigate phenotypic plasticity associated with the adoption of social roles in the clown anemonefish, with strategic growth being just one aspect of that plasticity. Strategic growth, also known as social control of growth, is a fascinating form of adaptive phenotypic plasticity, whereby individuals modify their growth and size in response to fine-scale changes in social conditions (Buston & Clutton-Brock, 2022). In cooperative breeding systems with high reproductive skew, particularly fishes and mammals (possibly including humans), individuals have been shown to i) increase growth/size on the acquisition of dominant status (Dengler-Crish & Catania, 2007; Johnston et al., 2021; Thorley et al., 2018; Van Schaik & Van Hooff, 1996; Walker & McCormick, 2009), ii) increase growth/size when paired with size matched reproductive rivals (Huchard et al., 2016; Reed et al., 2019; this study), and iii) decrease growth/size to avoid conflict (Buston, 2003; Heg et al., 2004; Wong et al., 2007). While strategic growth is fascinating and clearly occurring in this study, we show coordinated changes of multiple aspects of the phenotype as fish adopt social roles. Therefore, we deliberately framed the Introduction broadly to avoid biasing the reader toward viewing growth as the sole or main driver.

      Weaknesses:

      (1) I was surprised that the HPA/stress axis was not considered here at all. Wouldn't we expect that subordinates have increased stress axis activation, which in turn could inhibit their growth and aggressive behavior?

      We also expected to see the HPA/stress axis activated in subordinates, which is why we carried out a targeted exploration of genes known to play a role in this axis. We did not find any genes that were significantly differentially expressed. We believe that there could be two explanations for this. First, from a methodological perspective, it could be due to our use of a whole-body RNA-seq, which may have masked this signal. Alternatively, the stress axis might play a more complex role than just acting as a simple on/off switch for reduced growth. Its activation may peak when competition over size is at its highest (during week one) or, conversely, it may peak later and help maintain reduced growth once hierarchies are firmly established (particularly after the dominant individual reaches its maximum size). To understand the role of the stress axis, future studies should observe how its activation varies over time. We acknowledge that the absence of a stress‑axis signal and its potential explanations were not clearly discussed in the original manuscript. In the revised version, we have addressed this in the Discussion at lines 564-567 and Methods at lines 795-797 and 802-805.

      Discussion lines 564-567 and 577-580:

      “These include appetite regulation (orexigenic and anorexigenic signaling), metabolic pathways (e.g., glycolysis, lactic fermentation, TCA cycle, fatty acid β-oxidation), growth-regulating pathways (GH/IGF, insulin/PI3K-AKT, mTOR, and Hippo), and potential molecular signatures to varying levels of social stress.”

      Methods lines 795-797 and 802-805:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      (2) To what extent are growth, food intake, agonistic behavior, and/or gene expression patterns coordinated across P1 vs P2 pairs? The lack of such an analysis seems like a missed opportunity.

      We had a similar thought. Specifically, we were interested in testing the hypothesis that the final size ratio of pairs, which is indicative of the amount of conflict remaining, would predict gene expression. We examined gene expression within pairs to test for coordinated changes and repeated the analysis, accounting for the pair size ratio. In both cases, we found no clear or consistent pattern within pairs. We have included these analyses in the revised manuscript, Methods lines 781-7946-804 and 807-814, and Supplementary Materials lines 144-155 and 164-172 along with three new Supplementary Figures: Fig. S8, S9, S10.

      Supplementary Materials lines 144-155 and 164-172:

      “We have identified genes associated with growth and ossification that showed strong positive correlations with body size. For this set of genes, we tested the hypothesis that the final size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of remaining conflict (Wong et al., 2007), would predict variation observed in gene expression within social position (Fig. 5B). Visual inspection of the size ratio annotation in the heatmap of these genes (Supplementary Fig. S8), together with a PCA of paired individuals (P1 and P2) based on the same row Z-score values, was used to assess whether remaining conflict explained expression variation (PERMANOVA: p = 0.858; Supplementary Fig. S10A). These results suggest that size ratio between pairs was not associated with gene expression variation within social position.”

      “We have identified genes associated with appetite regulation and metabolism, which were significantly downregulated in P1 individuals compared to P2 and S fish. As above, we tested the hypothesis that the final size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of remaining conflict (Wong et al. 2016), would predict variation observed in gene expression within social position (Fig. 6). We found no significant effect (PERMANOVA: p = 0.844; Supplementary Fig. S9 and S10B), indicating that size ratio does not explain gene expression differences within social positions.”

      Methods lines 781-794 and 807-814:

      “In the heatmap, samples (columns) were a priori grouped by social position (P1, P2, S) and subsequently ordered within each group by gene expression similarity, reflected by the column dendrogram.

      Additionally, we tested the hypothesis that size ratio of pairs (P2 SL / P1 SL), which is indicative of the amount of conflict remaining (Wong et al., 2016), would predict variation in gene expression. This was assessed by visual inspection of the same heatmap with size ratio annotation using Complex Heatmaps (Gu et al., 2016), together with a PCA of the expression of these genes of paired individuals (P1, P2) only. We also performed a PERMANOVA analysis using the same row Z-score values (adonis2 function in vegan package: Oksanen et al., 2017) accounting for social position and genetic background (clutch ID), using restricted permutations to account for the non-independence of P1 and P2 within pairs was carried out. Solitary individuals were excluded from this analysis because they lack size ratio data. We found that size ratio between pairs was not associated with gene expression variation within social position (see Supplementary Materials; Supplementary Fig. S8, S10A).”

      “In the heatmap, samples (columns) and genes (rows) are grouped a priori by social position (P1, P2, S) and pathway (APT, TCA, GLY), respectively, with dendrograms reflecting expression similarity within each group. For the candidate gene heatmap, we repeated the PCA and PERMANOVA exploration and associated heatmap visualization for pairs only (as described above), to test whether final size ratio (P2 SL / P1 SL) predicted gene expression variation within social positions. We found that size ratio between pairs was not associated with gene expression variation within social position (see Supplementary Materials; Supplementary Fig. S9, S10B).”

      (3) What was the rationale for using whole bodies for the transcriptome analysis? Given the hypotheses, the forebrain or hypothalamus and certain other organ systems (e.g.,liver, gonads, skin, etc.) would have been obvious candidate tissues here. I realize that cost is always a consideration, but maybe a focus on the fore-/midbrain could have been prioritized.

      We decided to use whole-body samples for this initial transcriptomic analysis to capture a broad view of gene-expression differences while keeping sequencing costs and sample requirements manageable. We agree with the reviewer that future work should explore specific tissues sampled from individuals at multiple time points to disentangle transcriptomic differences across tissue types. In our revised manuscript we explicitly state the limitations and outline future steps in the Discussion at lines 553-564, as well as in the Methods section we now explicitly state the rationale behind whole-body RNA-seq and acknowledge the limitations of this approach in lines 688-692.

      Discussion line 5536-564:

      “In this study, we used whole-body transcriptomics, which revealed overarching expression patterns, however, we acknowledge that this approach limited our ability to detect finer-scale signals. Obviously, this approach cannot resolve tissue-specific gene expression changes (Lu et al., 2020; Roux et al., 2023; Yin et al., 2023) which are critical for a complete understanding social role differentiation, such as adjustments of behavior, appetite, and growth (a list we consider non-exhaustive). To disentangle gene expression signatures associated with socially induced phenotypes such as the strategic up- and downregulation of growth, future studies should include sampling points aligned with the onset of these physiological shifts. Moreover, to understand underlying pathways involved, tissue-specific transcriptomics will be essential. Sampling multiple tissues across multiple time points would disentangle nuanced regulatory processes underlying coordinated functional shifts leading to different social phenotypes of strategic growth.”

      Methods lines 688-692:

      “To initially explore overall gene expression patterns associated with coordinated changes during the emergence of social roles, whole-body RNA-seq was performed to keep both sequencing cost and sample requirements manageable. We acknowledge the limitations of this approach and consider this study a hypothesis-generating tool that lays the foundation for future studies.”

      (4) Given the preceding point, why was a fold-change threshold used for assessing DEGs (supplementary Figure 3)? There is no biological justification to ever use a fold-change threshold, especially in bulk RNA-seq analysis. This is particularly true here, where wholebodies were used for RNA-seq analysis, which is a bit unusual. Relatively small cell populations (such as hypothalamic neurons that regulate growth or food intake) may show substantial gene expression variation across social types, yet will be masked by the masses of other cells in the whole body sample. However, gene expression may still vary significantly, albeit the fold-difference may be small. I therefore suggest a reanalysis that omits any fold-change threshold.

      We thank the reviewer for this important point, and agree that an arbitrary fold‑change cutoff is inappropriate/unnecessary. It should be noted that this fold-change cutoff was only used in this single figure, and all other analyses used p-values from the entire dataset. We have removed the fold‑change threshold cutoff and corrected the Figure (previously Supplementary Figure 3) now Supplementary Fig. S5, and all corresponding text.

      (5) Why is the analysis of color (hue, saturation) buried in the supplementary materials? Based on the hypotheses that motivated the study, color seems just as relevant as food intake, growth, and agonistic behavior, so even if the results are negative, they should be presented in the main paper.

      We agree that color can be an important social signal, so we included color measurements in our experimental design. However, after careful consideration of the color results, we decided that our experimental timing and husbandry changes introduced multiple confounding factors, preventing us from drawing confident conclusions. Specifically, our fish were ≈1 month old at the transfer from larval to experimental tanks and had already begun to deepen their orange hue, before our experiment. (In the wild, they would settle at one to two weeks of age, prior to the deepening of the orange hue). Once individuals attain a certain hue, it seems that color development can be halted, but not reversed. The transfer also involved changes in lighting, tank background, and diet, factors known to strongly affect coloration (Maytin et. al., 2018). Our results show a uniform shift in orange hue and saturation across social groups, suggesting that these confounding factors might have dominated changes in hue.

      For transparency, we report the color data in the Supplementary Materials, but we caution against drawing any strong conclusions. In the revised Supplementary Materials document, we added lines 231-236 recommending that future work should involve a targeted experiment to robustly test for the effect of the adoption of social roles on coloration or the effect of coloration on the adoption of social roles.

      Supplementary Materials lines 231-236:

      “Together, these considerations suggest that timing, environmental uniformity, and future reproductive potential may limit the expression or detection of socially mediated color plasticity under laboratory conditions. Future studies should carefully consider experimental design to minimize confounding factors and robustly test the effects of the adoption of social roles on coloration and the effects of coloration on the adoption of social roles.”

      (6) The Discussion is sometimes difficult to follow. The authors may want to consider including a conceptual graphic that integrates the different aspects of growth and satiety regulation, etc., into a work-in-progress model of sorts, which would also facilitate clearer hypotheses for future research.

      Thank you for flagging that parts of the Discussion are a bit difficult to follow. In the revised manuscript, we worked to improve readability of the Discussion. We also appreciate the suggestion of including a conceptual schematic. For this manuscript, we refrained from adding such a “work-in-progress” schematic, as we felt that our whole-body RNA-seq approach and single gene expression sampling time point substantially limited our ability to make predictions.

      Reviewer #2 (Public review):

      In this manuscript, the authors test growth, behavior, and gene expression in pairs of clownfish as they establish social dominance hierarchies, examining patterns of gene expression in these pairs after dominance has been established. The authors show solid evidence that emerging dominant clownfish show increased growth, aggression, and food consumption compared to their submissive or solitary counterparts, eventually adopting distinct gene expression profiles.

      Major Comments:

      (1) The Introduction is comprehensive, but it could be condensed. Likewise, the discussion could be condensed. There is considerable redundancy between the methods, the results,and the legend in Figure 1. The authors should consolidate and remove the redundancy.

      Thank you for flagging that parts of the manuscript could be condensed, we will work on this as we revise the manuscript.

      (2) For Figure 3, the authors are showing PC2 and PC3; why is PC1 not shown? There is so much overlap between the three groups in PC2 vs PC3; it seems unlikely that researchers could conclusively identify any individual as belonging to a group based on the expression profile. The ovals shown do not capture all the points within each of the groups, and particularly the grey S oval seems misaligned with the datapoints shown.

      We understand the concern raised by the reviewer about the overlap among points in the PCA. We have explored PC1-PC3 and found that PC2 and PC3 showed the clearest, statistically significant clustering by social position, while PC1 did not capture any variation due to social position. We have explored whether other factors might be masking differences, such as genetic relatedness, tank effects, total read count per sample, and found that none of these factors explained sample clustering. Regarding the ellipses shown around the points, they were not intended to capture all points, but rather they show the estimated 95% multivariate t-distribution for that given social group. We revised the figure legend (Fig. 3) to clearly reflect this. . In addition, for transparency we have revised the Results section (lines 279-285) and Methods section (lines 754-765) to clarify that we have performed PCAs on (1) all genes, (2) top 50%, (3) top 25%, (4) top 5% most variable genes, and all three pairwise PC comparisons (PC1 and PC2, and PC1 and PC3) which we show in the added Supplementary Fig. 3S.

      Results lines 279-285:

      “PCAs were performed on all genes, and the top 50%, top 10%, and top 5% of the most variable genes, all of which showed consistent clustering across all pairwise PC comparisons (PC1 vs PC2; PC2 vs PC3; PC1 vs PC3; Supplementary Fig. S3). Of all the examined PCAs, PC2 and PC3 with all genes showed the clearest clustering by social position (p = 0.003), and revealed overall no significant difference in gene expression between P2 and S individuals, with P1 individuals exhibiting more distinct clustering (Fig. 3A; replicate‑annotated version of Fig. 3A in Supplementary Fig. S4; for all other PCAs see Supplementary Fig. S3).”

      Methods lines 754-765:

      “To test for overall GE patterns between social positions and genotypes, data were rlog-normalized and the effect of genotype and social position were compared through a Principal Component Analysis (PCA), followed by PERMANOVA analysis using Euclidean distances in the vegan package (Oksanen et al., 2017). To assess whether the strength of clustering by social position differed depending on the genes included, additional PCAs were performed on four gene subsets: (1) all genes, (2) the top 50%, (3) the top 10%, and (4) the top 5% of the most variable genes. For each subset, the first three principal components (PC1 vs PC2, PC1 vs PC3, PC2 vs PC3) were visualized to identify axes capturing variation associated with social position. To test for overall gene expression differences between social position and genotype, separate PERMANOVAs were performed for each PCA using the adonis2 in the vegan package (Oksanen et al., 2017).”

      (3) The authors indicate that the 15 replicates exhibiting the greatest size difference between P1 and P2 were selected for gene profiling. Does this mean that each of the P1and P2 were pairs with each other? Have the authors tried examining the gene expression patterns in a paired manner? E.g., for the pairs that showed the greatest size differences,do they also show the greatest differences in gene expression? Do the P1s show the most extreme differences from P2s that also show the most extreme P2 differences? Perhaps lines on Figure 3A connecting datapoints from the P1 and P2 pairs would be informative.

      Yes, “15 replicates exhibiting the greatest size difference between P1 and P2 were selected for gene profiling” refers to pairs of P1 and P2, we made sure this is clearly stated in the revised Results (lines 266-269) and Methods (lines 686-688). Yes, we have explored gene expression data considering the size difference between pairs, and found that it showed no clear differences in gene expression patterns (see our response and manuscript edits above under Reviewer #1 point 2). We also added a version of the main text Fig. 3A which clearly labels replicates within the PCA as Supplementary Fig. S4.

      Results lines 266-269:

      “To test the prediction that whole-body gene expression patterns will vary with social roles once clear social positions have emerged (P1, P2 and S), we performed gene expression profiling on 45 individual fish (experimental groups containing paired individuals and corresponding solitaries; N = 15 samples per social position; Fig. 1D).”

      Methods line 686-688:

      “Fish from these 15 replicates (n=45 samples; 15 paired individuals and their corresponding solitaries) were individually homogenized to allow equal RNA extraction from all tissues.”

      (4) For the specific target pathways that are up- and downregulated in the different backgrounds, I recommend that the authors include boxplots (or heatmaps) showing the actual expression values for these targets. Figure 6 shows a heatmap for appetite-related genes, and it would be great to see a similar graph for the metabolism and glycolytic genes; it would also be informative to see similar graphs for hormonal and sexual maturation pathways as well.

      We have explored genes across a broad set of metabolic pathways (glycolysis, TCA cycle, lactic fermentation, PDH complex, cholesterol biosynthesis, fatty-acid synthesis, and beta-oxidation) and show all metabolic genes that showed significant differential expression between P1, P2, and S in Figure 6. Overall, very few metabolism-associated genes were significantly differentially expressed, which is why we decided to combine appetite-regulation and metabolism-associated genes into a single figure (Figure 6). In our revised version of the manuscript, we have modified Fig. 6 to clearly indicate which pathways the genes belong to and modified Methods section lines 795-797 and 802-805.

      Methods lines 795-797 and 802-805:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      We also examined hormonal pathways (glucocorticoid and thyroid signaling), but did not find genes in these pathways that were significantly differentially expressed. Finally, we would like to clarify that our samples consist of two-month-old juvenile individuals that are sexually immature —under ideal conditions, clown anemonefish can mature in one to two years, but they can also remain sexually immature for a decade or more (Buston 2004; Buston & García, 2007) — which is why we did not observe distinct molecular signatures of sexual maturation. We recognize that the sentence at line 520 was misleading, as we did not identify any gene expression signature that we could confidently associate with signs of sexual maturation. We have revised the sentence in the Discussion (used to be line 520, now line 543-544) as well as in the Methods section we added lines 601-606.

      Discussion line 543-544:

      “In our current study, individuals within pairs were ultimately progressing towards rank 1 dominant female and rank 2 subordinate male roles.”

      Methods lines 601-606:

      “All fish used in this experiment were sexually immature juveniles (one-month-old at the beginning, two-month-old at the end), as clown anemonefish reach sexual maturity between the ages of one and two years old. Also, under ideal conditions, individuals can remain sexually immature for a decade or more (Buston 2004b; Buston & García, 2007). Therefore, in this study, the emergence of social roles and associated phenotypes do not reflect changes associated with sexual maturation.”

      (5) Particularly given that there is a relatively small number of genes enriched in the different rank conditions, I did not understand the need to do the WGCNA module analysis. I thought that an analysis of GO terms across the dataset would have been more meaningful than the GO term analysis shown in Figure 4, which considers only genes assigned to the "brown WGCNA module". This should be simplified or clarified.

      To clarify, GO enrichment analysis does not establish correlations with traits, it only describes which functions or pathways are over-represented in a given gene set. That is why we began by using WGCNA to define gene sets (modules) that are correlated to phenotypes. Our primary rationale for WGCNA was to identify modules of co-expressed genes that show significant statistical correlation with the phenotypes of interest (social role: P1, P2, S; growth; and food intake). Pairwise differential expression analysis (Figure 3B) identified a few hundred significantly differentially expressed genes, but those tests treat genes independently and are not able to help us link coordinated changes of co-expressed genes to phenotypes of interest. Because WGCNA is blind to traits, it first identifies groups of co-expressed genes, which can help resolve gene expression patterns.

      We therefore ran WGCNA on the rlog-transformed dataset to identify modules of co-expressed genes that show significant correlation with phenotypes of interest. For every module that showed such a correlation, we performed GO enrichment and carefully evaluated the resulting GO enrichment trees (see Supplementary Figs. S6, S7). The brown module was highlighted in the main text because it was one of the modules with a significant correlation to growth, and its associated GO enrichment showed clear growth-related signals that were not identified in the pairwise differential expression analysis results.

      In our revised manuscript we have clarified the rationale for the analytical approaches in the Results section at lines 297-303 and Methods section lines 767-772 and 775-779.

      Results lines 297-303:

      “...a weighted gene co-expression network analysis (WGCNA; Langfelder & Horvath, 2008) was conducted on the rlog-transformed gene expression (GE) dataset. Unlike pairwise DEG analysis, which treats genes independently, WGCNA identifies groups of co-expressed genes (modules) and tests whether modules of eigengene expression are correlated with variation in phenotypes of interest. For modules showing significant correlations with phenotypes, gene ontology (GO) analyses were performed using Fisher`s exact tests to identify over-represented pathways within each module.”

      Methods lines 767-772 and 775-779:

      “To test whether gene expression patterns are correlated with observed phenotypic changes in growth and appetite across social positions, a Weighted Gene Correlation Network Analysis (WGCNA) was conducted on the rlog-transformed GE dataset (Langfelder & Horvath, 2008). WGCNA identifies groups of co-expressed genes (modules) and correlates eigengene expression of each module with phenotypes of interest to determine modules of genes whose expression correlates with trait variation.”

      “For modules whose eigengene expression correlated with phenotypes of interest, gene ontology (GO) enrichment analysis was subsequently performed using Fisher's exact tests (presence/absence in a module) to identify over-represented pathways within each module across GO divisions of Biological Processes (BP), Molecular Functions (MF), and Cellular Components (CC) (Wright et al., 2015).”

      (6) The authors say that they have identified coordinated changes in behaviors and the"underlying gene expression, leading to the emergence" of social roles. This is a little bit misleading, since the gene expression analysis occurred well after the behavioral and phenotypic differences emerged. Presumably, the hormonal and genetic shifts that actually caused the behavioral and phenotypic difference occurred during the weeks during which the experiment was underway, and earlier capture of the transcriptome would presumably reveal different patterns, and ones that would be considered more causative.The authors acknowledge this in 434-435, but it could be emphasized further.

      We appreciate the reviewer raising this point. In the updated version of the manuscript, we have revised wording to convey that food intake, agonistic behavior, size and growth, and gene expression are all changing continuously, in response to each other and in response to social feedback (Introduction lines 133-136). An underappreciated aspect of this system (and likely many other systems) is that phenotype (including transcriptome) influences the outcome of social interactions, and the outcome of social interactions influences the phenotype (including the transcriptome). Earlier capture of the transcriptome would reveal different levels of gene expression, reflecting the state of the system at that moment in time.

      Introduction lines 134-137:

      “Underlying this cascade is an underappreciated aspect of this system (and likely many other systems) where phenotype (including gene expression) influences the outcome of social interactions, and the outcome of social interactions influences the phenotype (including gene expression).”

      (7) The authors have measured a number of differences between the different dominance classes of fish. All these differences were measured relative to the other classes, but in my view, the Solitary group was the closest to a baseline control. So I'm not sure that it is fair to say that "P2 and S individuals showed consistent downregulation of these genes and pathways" (line 401). I encourage the authors to emphasize the differences in gene expression from the "perspective" of the P1 individuals compared to the baseline of P2and S individuals. Line 474 says that "P2 fish showed significant upregulation" of a number of pathways. It should be very clear what that is compared to (compared to P1, presumably?)

      We agree with the reviewer that solitary individuals are the most intuitive baseline. Indeed, the experimental design included solitary fish because we expected they would serve as a useful control. Without social restraint, we anticipated they would show unrestricted growth, feeding, behavior, and associated gene‑expression patterns, similar to dominants.

      We initially ran analyses using solitaries as the baseline, but after examining the results, which showed subordinate‑like characteristics for the solitary individuals, we concluded that solitary individuals are not an ecologically appropriate control for this context. Removing juveniles from a social context and housing them in isolation may be stressful and can affect physiology and behavior in ways that do not reflect a natural baseline. From a life‑history standpoint, solitary living is not the typical state for A. percula.

      For these reasons, we reanalysed the dataset using the dominant (P1) as the reference to enable more ecologically meaningful comparisons (this choice was somewhat arbitrary, subordinates could also have been used as the reference). Given that gene expression is relative, we interpret results from both the dominant (P1) and subordinate (P2) perspectives in the Discussion to provide a complete view. We have clarified wording throughout the manuscript to make it clear that everything is relative (Introduction lines: 159-162 and 167-169; Methods lines: 622-625 and 726-728) as well as revised language throughout to make sure comparisons are clear.

      Introduction lines 159-162 and 167-169:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to lack of social interactions), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely dependent on social context with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses.”

      Methods lines 622-625 and 726-728:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely social context-dependent with no intrinsic baseline, in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses.”

      (8) Along the same lines, the authors say in line 514 that subordinates and solitaries strategically downregulate their growth. I'm not convinced that this is the case: I would consider this growth trajectory to be the default and the baseline. I would interpret that under certain social conditions, a P1 dominant pattern of growth, behavior, and gene expression is allowed to emerge.

      We respectfully disagree with the idea that a single baseline/reference growth trajectory exists for any individual of this species. Growth of individuals is entirely social context-dependent: neither fast nor slow growth represents an inherent baseline. When two size‑matched juveniles meet and compete to establish dominance, accelerated growth is the expected trajectory. By contrast, juveniles joining an existing hierarchy are expected to exhibit reduced growth, which minimizes conflict and facilitates their social integration. Unlike species that show non socially mediated growth trajectories, clown anemonefish do not have a context‑independent growth rate, rather, individuals constantly readjust their growth according to their immediate social environment (Buston 2003).

      Therefore, growth trajectories must be considered from the perspective of all group members, because they emerge from interactions among individuals rather than reflecting an intrinsic baseline. In this study, we were interested in the establishment of dominance hierarchy and how individuals adjust their phenotypes during this process. By experimentally pairing size‑matched rivals, both individuals are initially expected to pursue the dominant trajectory, and thus neither individual represents a default state. Instead, the outcome reflects a social decision, after which both individuals reinforce their emerging social roles through coordinated changes. See our response and manuscript edits above under Reviewer #2, point 7 (public reviews).

      Reviewer #3 (Public review):

      Summary:

      The authors tested the hypothesis that interactions among size- and age-matched rivals will lead to the emergence of social roles, accompanied by divergence in four aspects of individual phenotypes: growth, feeding behavior, fighting behaviors, and gene expression in clownfish.

      Strengths:

      The data on growth, feeding rate, and fighting behaviors support the authors' claims.

      Thank you for the positive feedback!

      Weaknesses:

      Gene analysis conducted in this study is not sufficient to clarify how the relevant genes actually regulate growth and behavior.

      The information obtained from whole-body gene expression analysis is very limited.Various gene expression is associated with the regulation of fighting behaviors, food intake, growth, and metabolism, and these genes are regulated differently across tissues, even within a single individual. Gene expression analysis should be performed separately for each tissue.

      We understand the reviewer’s concern about whole‑body transcriptomes and agree that tissue‑specific sampling would provide greater resolution of the mechanisms linking gene expression to growth, agonistic behaviors, and food intake. For this initial study, however, we deliberately chose whole‑body samples to capture a broad, unbiased view of gene expression differences while keeping sequencing costs and sample requirements manageable. We explicitly acknowledge the resulting interpretational limits in the Discussion (lines 497; 553–5647), and suggest in the last paragraph that the patterns reported here should be used to build on in future studies exploring targeted, tissue‑specific hypotheses. See our response and manuscript edits above under Reviewer #1, point 3 (public reviews).

      Clownfish undergo sex change depending on social status and body size, as the authors mention in the manuscript. Numerous gene expressions are affected by sex change. It is unclear how this issue was addressed.

      We thank the reviewer for raising this point. Sex change and sexual maturation can indeed drive major transcriptional shifts in clown anemonefish, but our experiment did not encompass such a life‑history transition. All individuals in this experiment were juveniles (≈1 month old at the start, ≈2 months old at the end) and were sexually immature at these ages. Clown anemonefish reach sexual maturation around one to two years under ideal conditions, can delay sexual maturation for years under normal conditions (Buston 2004; Buston & García, 2007), and sex change in the genus Amphiprion is known to take over ~5 months (Moyer & Nakazono, 1978). Accordingly, individuals in this study were not sexually mature, and sex change was not biologically plausible over the five-week experimental period of our study. We recognize that the sentence at line 520 may be misleading, as we did not identify any gene expression signature that we could confidently associate with signs of sexual maturation. During the revisions we made sure that it is clearly stated that the fish in this study were sexually immature. See our response and manuscript edits above under Reviewer #2, point 4 (public reviews).

      Recommendations for the authors:

      Reviewing Editor Comments:

      While appreciating the work presented in the manuscript, we note a few common concerns that the authors could prioritise in their revisions:

      (1) The transcriptomics

      The authors have used whole-body RNA-seq and so are constrained in the kind of mechanistic or tissue-specific insight it can really provide. There are also questions about how far the authors can go in linking these data to causation, given that sampling happened after the phenotypes had already diverged.

      We therefore recommend tempering the causal language, clarifying the rationale for their analytical choices (WGCNA, fold-change threshold...), and more explicitly discussing the limitations of the whole-body RNA-seq and what can (and cannot) be concluded from it.

      We thank the editors for these recommendations. We have addressed each point as follows:

      Tempering causal language: We have revised our manuscript and tempered with language throughout to remove wording of “influence” or “cause” and instead used “associated with”, “correlated with”, or “suggesting”, as appropriate. We made changes to our Abstract (lines 39-41), Introduction (lines 115-118, 121-123), and Discussion (lines 431-434, 571-574), and Methods sections (lines 795-797). Whereas other parts of the manuscript already used non-causal language such as Discussion lines 333, 355, 358, 381, 437, etc.

      Abstract lines 39-41:

      “Here we identify associations between changes in gene expression, growth, and feeding behavior regulation that reinforce social role differentiation during dominance hierarchy formation in clownfish.”

      Introduction lines 115-118, 121-123:

      “Gene expression profiling offers a powerful approach to uncover coordinated changes in gene expression across multiple pathways, providing insight into associations between the development of social role-specific phenotypes and their underlying proximate mechanisms.”

      “Its well-annotated genome (Lehmann et al., 2019) enables transcriptomic analyses that can help characterize the molecular patterns associated with these processes.”

      Discussion lines 431-434, 571-574:

      “To explore underlying GE differences, we examined genes known to be associated with appetite regulation and metabolism in A. ocellaris (Herrera et al., 2025), and found downregulation of these genes in P1 individuals compared to P2 and S.”

      “Together, these approaches provide a powerful framework for uncovering the dynamic, context-dependent mechanisms that regulate strategic growth and coordinated changes associated with the establishment and maintenance of dominance hierarchies in social vertebrates.”

      Methods lines 795-797:

      “To further investigate GE differences, we explored genes associated with growth (including thyroid signaling), appetite regulation, metabolism, and stress (corticoid) pathways across social positions.”

      WGCNA rationale: We have revised our Results and Methods sections to explicitly define why we used WGCNA and GO enrichment analysis, and how these analyses provided more information than simple pairwise differential gene expression analysis. See our response and manuscript edits above under Reviewer #2, point 5, (public reviews).

      Fold-change threshold: In our revised manuscript, we have removed the arbitrary imposed fold-change threshold cutoff, which was only used in that single figure. See our response and manuscript edits above under to Reviewer #1, point 4, (public reviews).

      Whole-body RNA-seq limitations: We have revised the Methods and Discussion sections, and see our response and manuscript edits above under Reviewer #1, point 3 (public reviews).

      (2) Framing and interpretation

      Several of the reviewers' comments flag that the manuscript overstates coordination or "strategic" regulation, or where what's being treated as the baseline/derived isn't clear (for example, whether P2/S are actively downregulating or whether P1 represents the divergent trajectory). We recommend revisiting any wording that implies stronger mechanistic inferences than the data actually support (and defining more clearly what is meant by baseline and socially induced state).

      We thank the editors for these comments and feedback. We have revised the manuscript to clearly state that our system does not have the traditional baseline/socially induced states, rather it is all social context-dependent. In particular, we have adjusted the wording throughout the manuscript to make it clear that everything is relative, and we clarified that our choice of P1 as the reference group was arbitrary. We have made revisions in the Introduction (lines 159-162 and 167-169) and Methods (lines 622-625 and 726-728). See our response and manuscript edits above under Reviewer #2, points 7 and 8, (public reviews).

      We have also revised the manuscript to temper with wording that could be interpreted as implying stronger mechanistic or directional conclusions than our data allow. We adjusted language throughout to remove wording of “influence” or “cause” and instead used “associated with”, “correlated with”, or “suggesting”, as appropriate. See our response and manuscript edits above under Reviewing Editor Comments, (1) the transcriptomics (recommendations to authors).

      Reviewer #1 (Recommendations for the authors):

      (1) The reader would benefit from a brief overview of the physiology and molecular basis of somatic growth, regulation of food intake, and aggressive behavior, especially in teleost fishes. This would also help the authors with formulating hypotheses that are a bit more explicit when it comes to the transcriptomic part of the study.

      We thank the reviewer for this suggestion, however, as another reviewer raised concerns about the length of the Introduction as is and we felt that adding detailed paragraphs on the physiology and molecular basis of somatic growth, food intake regulation, and aggressive behavior would significantly increase the length of our Introduction.

      As a compromise, we have revised the relevant sections of the Introduction (lines 109-115) to point readers to key review papers and primary literature where the molecular basis of these traits is covered in detail. We believe this approach balances the need for mechanistic context with manuscript length constraints, while also allowing readers with specific interests to follow up with the relevant literature.

      Introduction lines 109-115:

      “Most studies, often conducted without relevant social context of an individual or in species entirely lacking dominance hierarchies, have examined single pathways to uncover variation in coloration (Salis et al., 2019, 2021), appetite regulation (see review for teleosts: Volkoff, 2019), behavior (Bender et al., 2006; Renn et al., 2008; Santema et al., 2013; Solomon-Lane et al., 2022; see review for teleosts: St-Cyr & Aubin-Horth, 2009), and growth (Beckman, 2011; Lu et al., 2020; see reviews for teleosts: Reinecke, 2010; Zhou et al., 2024), leaving the broader molecular shifts associated with social role adoption unresolved.”

      (2) I certainly agree with that statement that "Understanding the proximate mechanisms that facilitate phenotypic adjustments is key to disentangling whether phenotypes are the cause or consequence of social rank" (line 90f.), though maybe the authors can elaborate a bit, as this relationship is quite dynamic and obviously goes both ways.

      We agree that this relationship is dynamic and bidirectional, and we have revised the Introduction to reflect this more explicitly. In particular, we added text noting that, in social vertebrates, phenotype, including gene expression, can both influence and be influenced by social interactions, often through continuous feedback loops rather than a clear directional relationship (Introduction lines 93-96). We also revised the experimental framing in the Introduction and Methods (Introduction lines 159-162 and 167-16972; Methods lines 622-625 and 726-728) to emphasize that phenotypes are socially context-dependent in A. percula and that the statistical reference group was chosen for analytical clarity rather than as a true biological baseline.

      Introduction lines 93-96, 159-162 and 167-169:

      “This is further complicated by the fact that, in social vertebrates, phenotype (including gene expression) can both influence and be influenced by social interactions, often through continuous feedback loops rather than a clear directional relationship.”

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to lack of social interactions), individuals took on varying social roles as dominant and subordinate members.

      “Given that all phenotypes are entirely dependent on social context with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses. Given that all phenotypes are entirely social context-dependent with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analysis.”

      Methods lines 622-625 and 726-728:

      “This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role, however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members. This experimental design allowed individuals equal opportunity to attempt taking on the dominant social role; however, through continuous social interactions (or in the case of solitaries, due to the lack of it), individuals took on varying social roles as dominant and subordinate members.”

      “Given that all phenotypes are entirely social context-dependent with no intrinsic baseline, in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analyses. Given that all phenotypes are entirely social context-dependent with no intrinsic baseline in this species, we arbitrarily chose P1 (dominant) as the statistical reference group for all downstream analysis.”

      (3) Much of the information in the last paragraph of the Introduction (lines 151ff.) is best presented in the Methods section.

      We respectfully disagree with this suggestion. The eLife journal format does not include a standalone Methods section preceding the Results, so we expect most readers not to consult the Methods before reading the Results. We believe it is important to briefly orient the reader to the experimental approach and the ideas tested, thus providing some context before they encounter the Results and Discussion.

      (4) What count (TPM or similar) and abundance (above count threshold in fraction of samples) thresholds were used for the transcriptome analysis? Maybe I missed it, but how many genes were in the analysis?

      We thank the reviewer for pointing this out. We have updated our Methods section (lines 714-719) to explicitly include filtering steps used as well as included the total number of genes that were used in downstream analysis.

      Methods lines 714-719:

      “The read count file was then imported into R version 4.3.1 (R Core Team, 2021), size factors were estimated for each sample using the median ratio method in DESeq2 (Love et al., 2014) to account for differences in sequencing depth across samples. No outlier samples were identified, and all samples (n=45) were retained for downstream analyses. Genes with a mean raw count <10 were then removed, retaining 24,840 genes for downstream analyses.”

      (5) Were all these genes used in PCA, and if so, why? Would it not make more sense to only use the 50% or 25% most variable genes (which would likely enhance the separation of social types)? Also, did the authors inspect higher-order PCs to see whether any of them separate the social types or separate samples according to some other variable(e.g., size, hue, feeding, any technical factors, etc.)?

      We explored PCAs using four gene subsets: all genes, the top 50%, top 10%, and top 5% most variable genes, and all pairwise PC comparisons (PC1 vs PC2, PC1 vs PC3, PC2 vs PC3; Supplementary Fig. S3). Of all these examined PCAs, PC2 and PC3, with all genes, showed the clearest, statistically significant clustering by social position, and more stringent filtering did not strengthen this signal. We therefore retained all genes in the primary analysis to avoid imposing arbitrary filtering thresholds that could exclude biologically relevant low-variance genes. See our response and manuscript edits above under Reviewer #2, point 2, (public reviews).

      Regarding the inspection of higher-order PCs for potential confounding variables, we examined whether genetic background (clutch ID), tank identity, and sequencing depth explained clustering patterns across PC1–PC3 and found no evidence that any of these variables (PCA for sequencing depth not shown). See our response and manuscript edits above under Reviewer #2, point 2, (public reviews).

      We did not explicitly explore whether body size at week 5 drove any of the observed separation among social positions. Rather, we investigated whether size ratio, an indicator of the amount of remaining conflict within a pair, could drive gene expression variation within social positions. See our response and manuscript edits above under Reviewer #1, point 2, (public review).

      Regarding orange hue, we do not believe that orange hue at week 5 (the time point at which whole-body samples were collected for gene expression profiling) represents a meaningful socially mediated result. See our response and manuscript edits above under Reviewer #1, point 5, (public reviews).

      Considering food intake, we chose not to explore this as a potential driver of PC variation because food intake could not be reliably assigned to all individuals at week 5. For a subset of pairs, individuals could not be confidently distinguished in videos, and we did not want to make assumptions that could introduce biases into the analysis.

      (6) Figures 5/6: Based on the gene expression shown in the heatmaps, I cannot see how the samples would cluster so cleanly by social type. Consider a bootstrapping analysis and provide bootstrap values and "confident" nodes. Also, what does "matrix" in the legend refer to? I assume some measure of gene expression level, maybe z-scored?

      We thank the reviewer for flagging this point. We acknowledge that the samples do not cluster freely by social position in the heatmaps and we have revised our Methods (see our response and manuscript edits above under Reviewer #1, point 2, public reviews) and Figure captions to clearly reflect this. To clarify, samples in Figures 5 and 6 were not ordered by unsupervised hierarchical clustering; rather, they were first grouped a priori by social position and then ordered within each group by similarity. We chose this presentation because it best illustrates the gene expression patterns associated with social position, which is the primary focus of the manuscript. We have updated the figure legends of Figures 5 and 6, to clarify that the color scale previously denoted as matrix in the heatmap represents row-scaled Z-scores of normalized gene expression values.

      To assess whether social position explains a gene expression variation in an unsupervised framework, we performed a PCA followed by PERMANOVA (using the adonis2 function in R, 999 permutations, Euclidean distance; see Author response image 1). Both analyses used the same row Z-score-scaled expression values as shown in Figures 5 and 6. Results showed a significant effect of social position on gene expression (PERMANOVA: A: social position p <0.001, B: social position p < 0.001; marginal clutch ID significance p=0.015), which was primarily driven by dominant individuals (P1) being significantly different from both subordinate (P2) and solitary (S) individuals. To include these into our manuscript.

      Author response image 1.

      Principal component analysis (PCA) of gene expression based on heatmaps of A) growth, B) appetite, and metabolism genes. PCA was performed on gene sets of the main text (growth: Fig 5B and appetite and metabolism: Fig. 6), using row Z-score scaled expression values, consistent with the heatmap scaling. Each point represents one individual. Ellipses represent 95% confidence intervals around each social position group. PERMANOVA results (adonis2; shown in the bottom right corners).

      (7) I applaud the authors for considering genetic/relatedness effects in their experimental design and analysis, but I am confused by the microsatellite vs. SNP analyses: why even use microsatellites in this day and age? Were only the SNP data used as the source of genetic information? This should be clarified.

      In our experiment, paired individuals were initially assigned a temporary rank based on their size at the start of the experiment; however, the initially larger individual (often only by a tenth of a mm) does not necessarily emerge as the dominant (P1). To avoid any assumptions regarding the identity of individuals in given social positions at the end of the experiment, we needed to verify that individuals within pairs were correctly identified. For the 45 individuals included in the gene expression dataset, this was done by calling SNPs directly from the TagSeq data. For the remaining 36 individuals not included in the gene expression dataset, we opted for microsatellite genotyping, as a validated panel with established markers was already available from a previous study (Rueger et al., 2025), making it a reliable and cost-effective solution. We have clarified the text in the Methods (lines 630-636) and Supplementary Materials (lines 42-49 and 65-67).

      Methods lines 630-636:

      “This step was necessary to avoid assumptions regarding social roles and fish identity. At the beginning of the experiment, individuals within pairs were provisionally assigned a social rank based on initial body size; as size differences were negligible (often less than 0.1mm), the initially larger individual does not necessarily emerge as the dominant (P1). To correct this assumption, identities were verified at the end of the experiment using either microsatellite genotyping or SNPs called from TagSeq data, depending on whether individuals were included in the gene expression dataset (see Supplementary Materials).”

      Supplementary Materials lines 42-49 and 65-67:

      “This step was necessary to avoid assumptions regarding social roles and fish identity, and the microsatellite method allowed for a reliable and cost-effective solution as a panel with established markers was already available from a previous study (Rueger et al., 2025). At the beginning of the experiment, individuals within pairs were provisionally assigned a social rank based on initial body size; as size differences were negligible (often less than 0.1 mm), the initially larger individual does not necessarily emerge as the dominant (P1). To correct this assumption, identities were verified at the end of the experiment.”

      “Similarly, for the remaining 45 individuals, we applied a necessary correction step to avoid assumptions regarding social role and fish identity. For these 45 individuals, we called SNPs from our TagSeq data to assign clutch identity.”

      Reviewer #2 (Recommendations for the authors):

      (1) Line 520 indicates that individuals showed early gene expression signatures of sexual maturation, but I did not see where those results were presented.

      We have addressed this recommendation, the claim “individuals showed early gene expression signatures of sexual maturation” was incorrect and has been removed from the revised manuscript (Discussion lines 543-544). We have also updated the Methods section (lines 601-606) to explicitly clarify that all individuals were sexually immature juveniles throughout the experiment. See our response and manuscript edits above under Reviewer #2, point 4, (public reviews).

      (2) The paragraph starting at line 409 refers to GE profiles. I was confused about what that was.

      Do the authors mean GO profiles?

      We did not find any modules using GO enrichment analysis which showed strong enrichment of appetite- and metabolism-related GO terms. Therefore, we looked for differentially expressed genes in the rlog-normalized gene expression dataset that are known to be associated with appetite regulation and metabolism. We have revised the Discussion (lines 431-434) and Methods (lines 812-815) to explicitly state what we are referring to and avoid confusion.

      Discussion lines 431-434:

      “In our study, P1 individuals showed increased food intake compared to P2 individuals. To explore underlying GE differences, we examined genes known to be associated with appetite regulation and metabolism in A. ocellaris (Herrera et al., 2025), and found downregulation of these genes in P1 individuals compared to P2 and S.”

      Method lines 802-805:

      “A complete list of retrieved A. percula gene IDs were then filtered against the whole-body GE dataset, and pathways containing significant genes associated with social position were reported (Supplementary Table S2; appetite and metabolic genes shown).”

      (3) I saw several typos, e.g., Vulcano in Supplementary Figure 3.

      Thank you for flagging this. We have addressed typos such as Supplementary Fig. S4 (used to be Supplementary Fig S3) See our response and manuscript edits above under Reviewer #1, point 4 (public reviews).

      (4) The personal observations cited in 495 should be more explicit. Which author made these observations, over how long, in how many instances?

      We thank the reviewer for this comment. We have updated the text to explicitly name the authors who made these observations, to clarify that they were made during two independent long-term field studies, and to note that this pattern was observed in 8 or more instances across the two field studies (Discussion lines 516-520).

      Discussion lines 516-520:

      “Similar patterns have been anecdotally observed in the wild, where juvenile clownfish remained small for extended periods (over four months) following the loss of a dominant partner, only initiating changes in social role towards dominant characteristics upon the arrival of a new group member (personal observations of 8+ instances during two long-term independent studies by Pete Buston and Lili Vizer).”

      References:

      Buston, P. (2003). Forcible eviction and prevention of recruitment in the clown anemonefish. Behavioral Ecology, 14(4), 576–582. https://doi.org/10.1093/beheco/arg036

      Buston, Peter M. (2004). Territory inheritance in clownfish. Proceedings of the Royal Society B: Biological Sciences, 271(SUPPL. 4), 252–254.

      Buston, P. M., & García, M. B. (2007). An extraordinary life span estimate for the clown anemonefish Amphiprion percula. Journal of Fish Biology, 70(6), 1710–1719. https://doi.org/10.1111/j.1095-8649.2007.01445.x

      Buston, P., & Clutton-Brock, Tim. (2022). Strategic growth in social vertebrates (WITH REVIEWER COMMENTS). Trends in Ecology & Evolution, 37(8), 694–705. https://doi.org/10.1016/j.tree.2022.03.010

      Dengler-Crish, C. M., & Catania, K. C. (2007). Phenotypic plasticity in female naked mole-rats after removal from reproductive suppression. THE JOURNAL OF EXPERIMENTAL BIOLOGY.

      Heg, D, Bender, N, & Hamilton, I. (2004). Strategic growth decisions in helper cichlids. Proceedings of the Royal Society of London. Series B: Biological Sciences, 271(suppl_6). https://doi.org/10.1098/rsbl.2004.0232

      Huchard, E, English, S, Bell, M B. V., Thavarajah, N, & Clutton-Brock, T. (2016). Competitive growth in a cooperative mammal. Nature, 533(7604), 532–534. https://doi.org/10.1038/nature17986

      Johnston, R A., Vullioud, P, Thorley, J, Kirveslahti, H., Shen, L., Mukherjee, S., Karner, C. M., Clutton-Brock, T, & Tung, J (2021). Morphological and genomic shifts in mole-rat ‘queens’ increase fecundity but reduce skeletal integrity. eLife, 10, e65760. https://doi.org/10.7554/eLife.65760

      Maytin, Alexander K., Davies, Sarah W., Smith, Gabriella E., Mullen, Sean P., & Buston, Peter M. (2018). De novo transcriptome assembly of the clown anemonefish (Amphiprion percula): A new resource to study the evolution of fish color. Frontiers in Marine Science, 5(AUG), 1–11.

      Moyer, J. T., & Nakazono, A. (1978). Protandrous Hermaphroditism in Six Species of the Anemonefish Genus Amphiprion in Japan (No. 2). The Ichthyological Society of Japan. https://doi.org/10.11369/jji1950.25.101

      Reed, C., Branconi, R., Majoris, J., Johnson, C., & Buston, P. (2019). Competitive growth in a social fish. Biology Letters, 15(2), 20180737. https://doi.org/10.1098/rsbl.2018.0737

      Rueger, Theresa, Bhardwaj, Anjali Kristina, Turner, Emily, Barbasch, Tina Adria, Trumble, Isabela, Dent, Brianne, & Buston, Peter Michael. (2022). Vertebrate growth plasticity in response to variation in a mutualistic interaction. Scientific Reports, 12(1), 11238.

      Thorley, J, Katlein, N, Goddard, K, Zöttl, M, & Clutton-Brock, T. (2018). Reproduction triggers adaptive increases in body size in female mole-rats. Proceedings of the Royal Society B: Biological Sciences, 285(1880), 20180897. https://doi.org/10.1098/rspb.2018.0897

      Van Schaik, C P., & Van Hooff, J A. R. A. M. (1996). Toward an understanding of the orangutan’s social system. In Linda F. Marchant, Toshisada Nishida, & William C. McGrew (Eds.), Great Ape Societies (pp. 3–15). Cambridge University Press. https://doi.org/10.1017/CBO9780511752414.003

      Walker, S P. W., & McCormick, M I. (2009). Sexual selection explains sex-specific growth plasticity and positive allometry for sexual size dimorphism in a reef fish. Proceedings of the Royal Society B: Biological Sciences, 276(1671), 3335–3343. https://doi.org/10.1098/rspb.2009.0767

      Wong, M. Y. L., Buston, P. M., Munday, Philip L., & Jones, Geoffrey P. (2007). The threat of punishment enforces peaceful cooperation and stabilizes queues in a coral-reef fish. Proceedings of the Royal Society B: Biological Sciences, 274(1613), 1093–1099. https://doi.org/10.1098/rspb.2006.0284

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      (1) This article purports to show that ML-SA8, a synthetic activator of the lysosomal TRPML1 channel, results in AMPK activation and glucose uptake in hepatocytes, and that this action has therapeutic potential for metabolic disease. The final figure shows that glucose levels are improved in db/db mice, although it is not entirely clear whether this is due to an effect on the liver, on other tissues, or on glucose production or uptake. The earlier figures try to make the case that SA8 causes activation and GLUT4 translocation and glucose uptake in liver cells; however, these data are not convincing. GLUT4 is expressed at such low levels in liver that it is likely not physiologically important. The authors use a fluorescent glucose analog to measure glucose uptake, and this molecule has been shown to enter cells largely by fluid phase endocytosis. Overall, this reviewer finds the premise misguided and the data unconvincing.

      Thank you for the critical comments and constructive suggestions. We have carefully considered all the concerns raised and provide our point-by-point responses below.

      (2) The initial figures show phosphorylation of AMPK on Thr172, but no downstream effects are shown. Usually, to convincingly show that AMPK activity is increased, it would be appropriate to immunoblot phospho-ACC or some other substrate. This is minor.

      Suggestion was taken! We will investigate the effect of SA8 on AMPK downstream effectors i.e. ACC activation by Western blotting. Ie p-ACC (Ser79) / total ACC.

      (3) Lines 135-148: GLUT4 is not expressed at levels that are significant for physiology in liver cells, and its function in liver is not particularly relevant. The authors cite references 38-40 to support that it may be expressed at low levels in liver, but no knockout studies have been done to show that this expression is physiologically important.

      We thank the reviewer for this critical comment. We agree that GLUT4 is not the predominant hepatic glucose transporter, but it is expressed at a relatively low level in the liver compared to other tissues.

      Regarding the physiological role of GLUT4, it has been shown to mediate glucose uptake in hepatic stellate and sinusoidal endothelial cells (Tang and Chen 2010, Karim, Liaskou et al. 2014). Furthermore, ischemia‑reperfusion (IR) significantly upregulated the expression of GLUT4 in the liver, rather than GLUT2, and the increased GLUT4 localized to the membrane peripheries of hepatocytes and enhanced glucose uptake, which in turn led to marked glycogen deposition (Kim, Jung et al. 2014, Kurabayashi, Furihata et al. 2022). These studies establish a clear physiological role for GLUT4 in the liver.

      Notably, Ranalletta, et. al. (2005) showed that GLUT4-null mice exhibit compensatory alterations in hepatic glucose and lipid metabolism, including increased hepatic glucose uptake and triglycerides conversion (Ranalletta, Jiang et al. 2005), indicating that GLUT4 ablation influences liver metabolism. Nevertheless, these existing evidences including the GLUT4 expression data and the functional changes observed in GLUT4 null mice—supports the relevance of GLUT4 in hepatic glucose metabolism. However, it is necessary to perform the liver-specific GLUT4 KO studies to clarify the role of GLUT4 in the liver. (We will incorporate these points and limitations into the revised Discussion section).

      (4) Figure 1e is not convincing. No controls are included to show the specificity of the antibody for immunofluorescent staining. No intracellular GLUT4 is visible in the unstimulated samples.

      We understood the reviewer’s concern about the specificity of GLUT4 antibody for immunofluorescent staining. The GLUT4 antibody (Abcam, ab33780) employed in our study has been extensively validated in previous studies for both Western blotting (Xie, Liu et al. 2024, Amanollahi, Holman et al. 2025, Ando, Takeda et al. 2025) and immunofluorescence (see also Johansson, Mannerås-Holm et al. 2013, Xiao, Zhang et al. 2025). in addition, we confirmed its specificity in our system by Western blot (Suppl. Fig.4), which showed a single band at ~45–55 kDa. Collectively, the combination of published validations, our own data supports the specificity of GLUT4 immunofluorescence detection.

      About the intracellular GLUT4 signal in Fig 1e. We apologize for the unclear GLUT4 signal in our original Fig. 1e. This was due to an inadvertently short exposure for the control condition. To improve this, we have now acquired new images with uniformly increased exposure time for all groups. As shown in the new Fig. 1e, GLUT4 is now clearly detected in control cells and mainly in cytosol, and ML-SA8 treatment obviously increases its accumulation at the plasma membrane.

      (5) In Figure 1f, again, the data are not convincing. The bands seem too sharp for GLUT4, which has 12 membrane-spanning domains as well as an N-linked glycosylation, so that it usually runs as a smear.

      We understood the reviewer’s concern. As an N-glycosylated membrane protein, GLUT4 may exhibit broader or diffuse migration patterns on immunoblotting. Nevertheless, the final band pattern is influenced by multiple factors, e.g. antibody specificity, sample preparation, electrophoresis conditions, and detection condition. Of note, multiple independent studies have shown endogenous GLUT4 as a relatively “sharp” immunoreactive band around 55–60 kDa (Gurley, Ilkayeva et al. 2016, Habtemichael, Li et al. 2021, Wu, Yu et al. 2024) (see also Ando et al., 2025; Amanollahi et al., 2025; Xie et al., 2024), which is very consistent with our results. Thus, we are confident that our GLUT4 band is specific and reliable.

      (6) Figure 1h. Data are not convincing. 2-NBDG is not a valid approach to measure glucose uptake. 2-NBDG enters cells largely via fluid phase endocytosis, and its accumulation is independent of known GLUT inhibitors such as cytochalasin B (Yazdani et al., MBoC 2022; PMID: 35921166; see also PMID: 42287154). The idea that such a bulky derivative of glucose could enter the transporter channel is not compatible with known structural data.

      We thank the reviewer for raising this point. We acknowledge that the uptake mechanism of 2-NBDG is controversial. As the reviewer noted, some studies have reported that it enters cells largely by endocytosis in certain cell types, and this remains a subject of ongoing discussion. Nevertheless, 2-NBDG continues to be widely employed as a glucose uptake tracer in this field, including several recent high-profile studies (Nobs, Kolodziejczyk et al. 2023, Xiong, Helm et al. 2023, Wu, Lv et al. 2025) as following.

      (1) Nobs, S.P., et al., Lung dendritic-cell metabolism underlies susceptibility to viral infection in diabetes. Nature, 2023. 624(7992): p. 645-652.

      (2) Xiong, L., et al., Nutrition impact on ILC3 maintenance and function centers on a cell-intrinsic CD71-iron axis. Nat Immunol, 2023. 24(10): p. 1671-1684.

      (3) Wu, Y., et al., Dalbergia odorifera T.C. Chen leaf extract promotes microglial energy expenditure to phagocytize neutrophils after cerebral ischemia-reperfusion. Phytomedicine, 2025. 149: p. 157508.

      In the current study, given the consistency of our results with parallel functional assays, the use of 2‑NBDG is justified in this context.

      (7) Supplementary Figure 5 uses 2-NBDG glucose uptake again. This reviewer is not convinced that the data reflect transporter-mediated glucose uptake, as suggested by the authors.

      Please see response to #6.

      (8) As well, although palmitate treatment of cells can cause an insulin-resistant-like phenotype in some cell types, this is not characterized in the present work.

      We understand the reviewer’s concern regarding the characterization of PA-induced insulin resistance model.

      First, this PA-induced insulin resistant hepatic model is well-established and validated in several literatures (Lee, Cho et al. 2010, Zhang, Cai et al. 2020, Malik, Inamdar et al. 2024), and it has been used for T2DM natural and synthetic drug screening (Faria, Calixto et al. 2025).

      Second, we have characterized this model in our system. As shown in Suppl. Fig. 5a, b, insulin (100 nM, 0.5 h) induced an increase of glucose uptake in HepG2 cells measured by 2-NBDG (Yamada, Nakata et al. 2000). In contrast, in PA-treated HepG2 cells, this effect was almost completely blocked, suggesting that PA-treated HepG2 cells are less sensitive to insulin. Overall, this model has been well validated and is suitable for the purposes of our study.

      (9) Finally, as noted, one would not expect hepatocytes to exhibit insulin-responsive glucose transport. Glycogen synthesis is the main insulin-regulated step that might be affected.

      As stated in the Responses#9, PA-induced hepatic insulin-resistance has become a widely accepted in vitro model for investigating therapeutic strategies (Lee, Cho et al. 2010, Zhang, Cai et al. 2020, Malik, Inamdar et al. 2024, Faria, Calixto et al. 2025).

      (10) The data in Figures 2b,c,f,g,k,l are not convincing. Again, 2-NBDG is used.

      Please see response to #6.

      (11) For the glucose consumption measurements in other panels of Figure 2, the methods section states that cells were cultured in 10 mM glucose. What volume was used? It is difficult to believe that a monolayer of cells would consume very much of the glucose that is present in the culture medium. Data are shown as a percent of controls, and look reasonable, but it would be helpful to include absolute as well as relative units.

      We appreciate the reviewer’s comments and apologize for any confusion about the method.

      First, we have revised the methods section to clarify the assay procedure “Following the manufacturer’s protocol, 2.5 μL of sample (medium/standard) was mixed with 250 μL of working solution a 96-well plate. The mixture was then incubated at 37 °C for 10 min and the absorbance was measured…” (line 421-423).

      Second, in response to the suggestion to include both absolute and relative units, we will provide the data with absolute value for reviewer’s reference. In the main figures, we have retained the normalized data as this format allows direct comparison of treatment effects across independent experiments.

      (12) In Figure 2, in experiments using the TRPML1 KO cells, no panel is shown to demonstrate knockout. The authors cite a previous paper for the construction of these cells, but the control immunoblot should still be shown here.

      We thank the reviewer for raising this point. The TRPML1 knockout cell line used in our study was originally generated and provided by Prof. Haoxing Xu’s laboratory, this cell line has been validated in several published literature (Wang, Gao et al. 2015, Zhang, Cheng et al. 2016). We understand the reviewer’s concern, so we will further validate this TRPML1 KO cell line.

      (13) In Figure 3, controls are missing in the BAPTA experiment in Figure 3a (only SA8-treated cells were treated with BAPTA and with EGTA). Again, it would be helpful to have p-ACC or some other readout of AMPK activity, and not just AMPK phosphorylation. 2NBDG is again used in this figure.

      Suggestion taken! We will add the controls including BAPTA-AM and EGTA only data. p-ACC/total ACC will also be measured. About the 2-NBDG, please see Responses#6.

      (14) Line 212-213 the text states "considering our finding that TRPML1-mediated Ca2+ release is essential for AMPK activation." This has not been shown. The work uses chelators and does not necessarily indicate a role for TRPML1. The drug may be specific, as suggested by the authors, but the way this phrase is worded is too strong. As well, AMPK was shown to be phosphorylated, but full activation towards its various substrates has not been shown.

      We thank the reviewer for raising this concern. We fully agree that the Ca<sup>2+</sup> chelators experiment alone could not specially attribute the effect to TRPML1.

      In fact, we have performed experiment to address the TRPML1-dependent mechanism in the original submission. As shown in Fig. 2i-l, In TRPML1 KO HAP1 cells (Qi, Xing et al. 2021), ML-SA8-induced AMPK phosphorylation and cellular glucose uptake were almost completely abolished compared to wild-type (WT) HAP1 cells. Moreover, pharmacological inhibition of TRPML1 with a TRPML1 specific synthetic inhibitor-ML-SI5 completely abolished ML-SA8-triggered AMPK phosphorylation (Fig. 2d, e). In addition, in IR-HepG2 model, ML-SI5 could substantially inhibited ML-SA8-induced cellular glucose uptake (Fig. 2f-h).

      Accordingly, we have also revised the statement to” considering our finding that TRPML1-mediated Ca<sup>2+</sup> release is necessary for AMPK activation” to make the sentence more rigorous.

      (15) Figure 4cd suggests that GLUT4 expression is increased by 2 or 3-fold in the liver of DB+SA8-treated mice, compared to controls. This may be the case, but its abundance is still likely ~1000-fold less in liver compared to skeletal muscle or adipose tissue. This reviewer is still not convinced that this is physiologically relevant. The images in Supplementary Figure 8 suggest a larger increase, but it remains uncertain whether the staining really represents GLUT4.

      We understood the reviewer’s concern. About the physiological importance of GLUT4, please see Responses #3. About the specificity of GLUT4 antibody, please see Responses#4.

      (16) Data showing that blood glucose and HbA1c are reduced in SA8-treated mice are reasonable, and GTTs and ITTs are shown. Unfortunately, there are no insulin concentrations, and it remains uncertain whether glucose production is reduced or uptake is increased (or if both effects are present).

      We will measure the insulin concentrations.

      (17) In the discussion, the authors again state that GLUT4 is present in the liver and that it regulates hepatic glucose homeostasis, and they cite reference 63. This review article does not argue that GLUT4 acts in the liver to regulate hepatic glucose homeostasis, but that its actions in muscle and fat have secondary effects on the liver.

      Sorry for the oversights. We have supplementary more precise references (Rossetti, Stenbit et al. 1997, Kurabayashi, Furihata et al. 2022, Fan, Jiao et al. 2023, Jiang, Luo et al. 2024).

      Reviewer #2 (Public review):

      The manuscript contains interesting studies suggesting that pharmacological activation of TRPML1 could be useful to treat T2D by increasing glucose uptake via activation of AMPK. Preclinical studies suggest the inhibitor improved blood glucose in Db/Db mice. Ex vivo studies in cell lines examine both pharmacologic and genetic manipulations, both to activate and to inactivate TRPML1, and the results consistently suggest that TRPML1 activates AMPK and increases glucose uptake.

      Strengths:

      The manuscript is well written, and the studies are carefully performed.

      (18) All mechanistic studies were performed in transformed cell lines; conclusions would be stronger if performed in primary cells. The in vivo studies were only performed in male mice. Performing metabolic studies in both sexes is standard practice now. Whether the findings would extend to females was not tested and remains uncertain. Some controls are missing, such as plasma membrane loading controls for fractionation studies. The GLUT4 staining was performed after fixation and permeabilization, yet control cells appear to be devoid of intracellular (and all) staining, a confusing result that doesn't reflect the expected biology.

      Thank you for your support and the constructive suggestions! We will answer these questions in the following point-to-point responses.

      Reviewer #3 (Public review):

      (19) Zhu et al. present a proof-of-concept for targeting the lysosomal calcium channel MCOLN1/TRPML1endolysosomal ion channels to restore type 2 diabetes mellitus (T2DM). Using synthetic TRPML1 agonists (ML-SA8) and genetic manipulation, the authors demonstrate that TRPML1 stimulation triggers localized lysosomal calcium release. This calcium efflux sequentially activates CaMKKβ and phosphorylates AMPK at Thr172 in various cell models, including palmitic acid-induced insulin-resistant HepG2 cells. This signaling pathway promotes GLUT4 translocation to the plasma membrane and increases intracellular glucose uptake. When administered daily to diabetic db/db mice over six weeks, ML-SA8 lowers fasting and random blood glucose, improves oral glucose and insulin tolerance tests, reduces hepatic steatosis, and lowers serum ALT and AST levels.

      Strengths:

      Based on the TFEB-independent pathway activated by TRPML1 and the experimental approaches described by Medina's group (PMID: 31822666), the authors use a combination of pharmacological and genetic tools to dissect such an intracellular signaling pathway. Additionally, the animal experiments show consistent phenotypic improvements across independent metabolic parameters. The ability of ML-SA8 to restore glycogen deposition and clear hepatic lipid accumulation in db/db mice without causing weight loss or overt toxicity provides a strong rationale for exploring lysosomal targets in metabolic disease.

      Thank you for the support!

      (21) The authors focus almost exclusively on hepatic GLUT4 to explain the observed glucose disposal. However, other glucose transporter isoforms such as GLUT2 dominate basal glucose transport. While the authors show increased AMPK phosphorylation in skeletal muscle and adipose tissue, they do not measure GLUT4 translocation or glucose uptake in these primary disposal organs. As a result, attributing systemic glycemic recovery primarily to hepatic GLUT4 translocation overlooks the major physiological roles of peripheral tissues.

      We thank the reviewer for this critical comment and we agree with the reviewer. Our initial focus on hepatic GLUT4 was driven by our primary interest in liver metabolism and the fact that we observed a consistent and robust effect of ML‑SA8 on hepatic GLUT4 translocation, AMPK activation and glucose uptake. We acknowledge that the possible contribution of GLUT2 to these effects cannot be excluded. Also, ML‑SA8 may exert similar effects in other tissues, such as skeletal muscle and adipose tissue as we found out that GLUT4 levels were upregulated in these tissues (Suppl.Fig.8b, c). We therefore agree that the systemic glycemic recovery should be attributed to the multi‑tissue effects of ML‑SA8, rather than a liver-restricted phenomenon. Accordingly, we will revise the Discussion section to incorporate this. Note that the precise mechanisms of ML-SAs on GLUT2 and in peripheral tissues warrant further investigation.

      (22) In both HepG2 cells and mouse liver tissues, ML-SA8 treatment increases total GLUT4 protein expression in addition to plasma membrane localization. Because total protein pools expand, the enrichment of GLUT4 in plasma membrane fractions cannot be cleanly attributed to acute vesicular translocation alone. The manuscript does not explain the timescale or mechanism behind this rapid total protein upregulation, leaving a mechanistic gap between acute ion channel gating and protein expression.

      Great suggestion. we will re-calculated the enrichment of GLUT4 in PM with total GLUT4 protein to clarify.

      (23) While the in vitro specificity of ML-SA8 is well-controlled, the systemic animal experiments lack a specific rescue or knockout control. Small-molecule agonists administered intraperitoneally over six weeks can exert off-target effects. Without demonstrating that co-administering the TRPML1 inhibitor ML-SI5 blunts the therapeutic effect in vivo, or showing that ML-SA8 lacks efficacy in TRPML1-null mice, the definitive link between in vivo glycemic recovery and TRPML1 activation remains incomplete.

      We understand the reviewer’s concern about the specificity of ML-SAs for TRPML1. In fact, the specificities of ML-SAs and ML-SIs have been rigorously validated in previous studies using TRPML1 knockout (KO) cells (Sahoo, Gu et al. 2017, Yu, Zhang et al. 2020) and further confirmed in the atomic-resolution co-structures (Schmiege, Fine et al. 2017, Schmiege, Fine et al. 2021).

      In the current study we provide additional evidence supporting this specificity. We show that ML-SA8-induced AMPK phosphorylation and glucose uptake were abolished in TRPML1 knockout cells (Fig. 2i-l) and blocked by specific inhibitor ML-SI5 (Fig. 2d, e, f, g). In addition, ML‑SA5 has been show to lack efficacy in TRPML1‑null mice in several studies (Yu, Zhang et al. 2020, Zhang, Wang et al. 2024, Xing, Wang et al. 2025), further reinforcing the target specificity of this class of compounds. Hence, multiple lines of evidence support the on-target effect of ML-SA8: (1) in vitro blockade by ML-SI5 and TRPML1 KO; (2) consistent effects across structurally distinct TRPML1 agonists; (3) dose‑dependent responses in vivo; and (4) absence of efficacy in TRPML1‑null mice.

      We agree with the reviewer that rescue experiment or in vivo knockout validation would provide the most definitive proof of target specificity. However, breeding TRPML1‑null mice or performing extensive dose-finding studies for ML-SI5 in vivo would require substantial time and resources which probably fall beyond the scope of the current study. We have therefore acknowledge this limitation in the Discussion section and noted that future validation with ML-SI5 co-administration or TRPML1-null mice is warranted. We hope the reviewer finds our response acceptable.

      Amanollahi, R., S. L. Holman, A. S. Meakin, M. Padhee, K. J. Botting-Lawford, S. Zhang, S. M. MacLaughlin, D. O. Kleemann, S. K. Walker, J. M. Kelly, S. R. Rudiger, I. C. McMillen, M. D. Wiese, M. C. Lock and J. L. Morrison (2025). "In Vitro Embryo Culture Impacts Heart Mitochondria in Male Adolescent Sheep." J Dev Biol 13(2).

      Ando, T., R. Takeda, R. Kano, T. Kusano, Y. Nonaka, Y. Kano and D. Hoshino (2025). "Effects of pyruvate administration on mRNA expression of inflammatory cytokines in adipose tissue and whole-body glucose metabolism in male mice." Physiol Rep 13(15): e70362.

      Casimiro, I., N. D. Stull, S. A. Tersey and R. G. Mirmira (2021). "Phenotypic sexual dimorphism in response to dietary fat manipulation in C57BL/6J mice." J Diabetes Complications 35(2): 107795.

      Fan, X., G. Jiao, T. Pang, T. Wen, Z. He, J. Han, F. Zhang and W. Chen (2023). "Ameliorative effects of mangiferin derivative TPX on insulin resistance via PI3K/AKT and AMPK signaling pathways in human HepG2 and HL-7702 hepatocytes." Phytomedicine 114: 154740.

      Faria, B. Q., P. S. Calixto, G. Picheth, L. M. Ferreira, F. G. M. Rego, J. F. C. Guerra and M. H. M. Sari (2025). "Palmitate-induced hepatic insulin resistance as an in vitro model for natural and synthetic drug screening: A scoping review of therapeutic candidates and mechanisms." Chem Biol Interact 420: 111717.

      Gurley, J. M., O. Ilkayeva, R. M. Jackson, B. A. Griesel, P. White, S. Matsuzaki, R. Qaisar, H. Van Remmen, K. M. Humphries, C. B. Newgard and A. L. Olson (2016). "Enhanced GLUT4-Dependent Glucose Transport Relieves Nutrient Stress in Obese Mice Through Changes in Lipid and Amino Acid Metabolism." Diabetes 65(12): 3585-3597.

      Habtemichael, E. N., D. T. Li, J. P. Camporez, X. O. Westergaard, C. I. Sales, X. Liu, F. López-Giráldez, S. G. DeVries, H. Li, D. M. Ruiz, K. Y. Wang, B. S. Sayal, S. González Zapata, P. Dann, S. N. Brown, S. Hirabara, D. F. Vatner, L. Goedeke, W. Philbrick, G. I. Shulman and J. S. Bogan (2021). "Insulin-stimulated endoproteolytic TUG cleavage links energy expenditure with glucose uptake." Nat Metab 3(3): 378-393.

      Jiang, Y., P. Luo, Y. Cao, D. Peng, S. Huo, J. Guo, M. Wang, W. Shi, C. Zhang, S. Li, L. Lin and J. Lv (2024). "The role of STAT3/VAV3 in glucolipid metabolism during the development of HFD-induced MAFLD." Int J Biol Sci 20(6): 2027-2043.

      Johansson, J., L. Mannerås-Holm, R. Shao, A. Olsson, M. Lönn, H. Billig and E. Stener-Victorin (2013). "Electrical vs manual acupuncture stimulation in a rat model of polycystic ovary syndrome: different effects on muscle and fat tissue insulin signaling." PLoS One 8(1): e54357.

      Karim, S., E. Liaskou, J. Fear, A. Garg, G. Reynolds, L. Claridge, D. H. Adams, P. N. Newsome and P. F. Lalor (2014). "Dysregulated hepatic expression of glucose transporters in chronic disease: contribution of semicarbazide-sensitive amine oxidase to hepatic glucose uptake." Am J Physiol Gastrointest Liver Physiol 307(12): G1180-1190.

      Kim, S., J. Jung, H. Kim, R. W. Heo, C. O. Yi, J. E. Lee, B. T. Jeon, W. H. Kim, J. R. Hahm and G. S. Roh (2014). "Exendin-4 Improves Nonalcoholic Fatty Liver Disease by Regulating Glucose Transporter 4 Expression in ob/ob Mice." Korean J Physiol Pharmacol 18(4): 333-339.

      Kim, S. J., A. Gajbhiye, A. R. Lyu, T. H. Kim, S. A. Shin, H. C. Kwon, Y. H. Park and M. J. Park (2023). "Sex differences in hearing impairment due to diet-induced obesity in CBA/Ca mice." Biol Sex Differ 14(1): 10.

      Kurabayashi, A., K. Furihata, W. Iwashita, C. Tanaka, H. Fukuhara, K. Inoue, M. Furihata and Y. Kakinuma (2022). "Murine remote ischemic preconditioning upregulates preferentially hepatic glucose transporter-4 via its plasma membrane translocation, leading to accumulating glycogen in the liver." Life Sci 290: 120261.

      Lee, J. Y., H. K. Cho and Y. H. Kwon (2010). "Palmitate induces insulin resistance without significant intracellular triglyceride accumulation in HepG2 cells." Metabolism 59(7): 927-934.

      Luo, J., H. Alkhalidy, Z. Jia and D. Liu (2024). "Sulforaphane Ameliorates High-Fat-Diet-Induced Metabolic Abnormalities in Young and Middle-Aged Obese Male Mice." Foods 13(7).

      Malik, S., S. Inamdar, J. Acharya, P. Goel and S. Ghaskadbi (2024). "Characterization of palmitic acid toxicity induced insulin resistance in HepG2 cells." Toxicol In Vitro 97: 105802.

      Nguyen-Phuong, T., S. Seo, B. K. Cho, J. H. Lee, J. Jang and C. G. Park (2023). "Determination of progressive stages of type 2 diabetes in a 45% high-fat diet-fed C57BL/6J mouse model is achieved by utilizing both fasting blood glucose levels and a 2-hour oral glucose tolerance test." PLoS One 18(11): e0293888.

      Nobs, S. P., A. A. Kolodziejczyk, L. Adler, N. Horesh, C. Botscharnikow, E. Herzog, G. Mohapatra, S. Hejndorf, R. J. Hodgetts, I. Spivak, L. Schorr, L. Fluhr, D. Kviatcovsky, A. Zacharia, S. Njuki, D. Barasch, N. Stettner, M. Dori-Bachash, A. Harmelin, A. Brandis, T. Mehlman, A. Erez, Y. He, S. Ferrini, J. Puschhof, H. Shapiro, M. Kopf, A. Moussaieff, S. K. Abdeen and E. Elinav (2023). "Lung dendritic-cell metabolism underlies susceptibility to viral infection in diabetes." Nature 624(7992): 645-652.

      Qi, J., Y. Xing, Y. Liu, M. M. Wang, X. Wei, Z. Sui, L. Ding, Y. Zhang, C. Lu, Y. H. Fei, N. Liu, R. Chen, M. Wu, L. Wang, Z. Zhong, T. Wang, Y. Liu, Y. Wang, J. Liu, H. Xu, F. Guo and W. Wang (2021). "MCOLN1/TRPML1 finely controls oncogenic autophagy in cancer by mediating zinc influx." Autophagy 17(12): 4401-4422.

      Racine, K. C., L. Iglesias-Carres, J. A. Herring, K. L. Wieland, P. N. Ellsworth, J. S. Tessem, M. G. Ferruzzi, C. D. Kay and A. P. Neilson (2024). "The high-fat diet and low-dose streptozotocin type-2 diabetes model induces hyperinsulinemia and insulin resistance in male but not female C57BL/6J mice." Nutr Res 131: 135-146.

      Ranalletta, M., H. Jiang, J. Li, T. S. Tsao, A. E. Stenbit, M. Yokoyama, E. B. Katz and M. J. Charron (2005). "Altered hepatic and muscle substrate utilization provoked by GLUT4 ablation." Diabetes 54(4): 935-943.

      Rossetti, L., A. E. Stenbit, W. Chen, M. Hu, N. Barzilai, E. B. Katz and M. J. Charron (1997). "Peripheral but not hepatic insulin resistance in mice with one disrupted allele of the glucose transporter type 4 (GLUT4) gene." J Clin Invest 100(7): 1831-1839.

      Sahoo, N., M. Gu, X. Zhang, N. Raval, J. Yang, M. Bekier, R. Calvo, S. Patnaik, W. Wang, G. King, M. Samie, Q. Gao, S. Sahoo, S. Sundaresan, T. M. Keeley, Y. Wang, J. Marugan, M. Ferrer, L. C. Samuelson, J. L. Merchant and H. Xu (2017). "Gastric Acid Secretion from Parietal Cells Is Mediated by a Ca2+ Efflux Channel in the Tubulovesicle." Developmental Cell 41(3): 262-273.e266.

      Schmiege, P., M. Fine, G. Blobel and X. Li (2017). "Human TRPML1 channel structures in open and closed conformations." Nature 550(7676): 366-370.

      Schmiege, P., M. Fine and X. Li (2021). "Atomic insights into ML-SI3 mediated human TRPML1 inhibition." Structure 29(11): 1295-1302 e1293.

      Tang, Y. and A. Chen (2010). "Curcumin prevents leptin raising glucose levels in hepatic stellate cells by blocking translocation of glucose transporter-4 and increasing glucokinase." Br J Pharmacol 161(5): 1137-1149.

      Tukhovskaya, E. A., E. R. Shaykhutdinova, I. A. Pakhomova, G. A. Slashcheva, N. A. Goryacheva, E. S. Sadovnikova, E. A. Rasskazova, V. A. Kazakov, I. A. Dyachenko, A. A. Frolova, A. N. Brovkin, V. E. Kaluzhsky, M. Y. Beburov and A. N. Murashev (2022). "AICAR Improves Outcomes of Metabolic Syndrome and Type 2 Diabetes Induced by High-Fat Diet in C57Bl/6 Male Mice." Int J Mol Sci 23(24).

      Wang, W., Q. Gao, M. Yang, X. Zhang, L. Yu, M. Lawas, X. Li, M. Bryant-Genevier, N. T. Southall, J. Marugan, M. Ferrer and H. Xu (2015). "Up-regulation of lysosomal TRPML1 channels is essential for lysosomal adaptation to nutrient starvation." Proc Natl Acad Sci U S A 112(11): E1373-1381.

      Wu, D., H. C. Yu, H. N. Cha, S. Park, Y. Lee, S. J. Yoon, S. Y. Park, B. H. Park and E. J. Bae (2024). "PAK4 phosphorylates and inhibits AMPKα to control glucose uptake." Nat Commun 15(1): 6858.

      Wu, Y., W. Lv, S. Xiong, G. Cao, L. Fu, W. Liu, F. Shao, Y. Mei and Y. Lv (2025). "Dalbergia odorifera T.C. Chen leaf extract promotes microglial energy expenditure to phagocytize neutrophils after cerebral ischemia-reperfusion." Phytomedicine 149: 157508.

      Xiao, B., W. Zhang, N. Ji and Q. Chen (2025). "Knockdown of CCNB1 alleviates high glucose-triggered trophoblast dysfunction during gestational diabetes via Wnt/β-catenin signaling pathway." Open Med (Wars) 20(1): 20241119.

      Xie, Y., X. Liu, W. Liu, L. R. Carr, L. P. Lee, N. Imai, E. A. Ortlund and D. E. Cohen (2024). "Activity and phosphatidylcholine transfer protein interactions of skeletal muscle thioesterase Them2 enable hepatic steatosis and insulin resistance." J Biol Chem 300(11): 107855.

      Xing, Y., M. M. Wang, F. Zhang, T. Xin, X. Wang, R. Chen, Z. Sui, Y. Dong, D. Xu, X. Qian, Q. Lu, Q. Li, W. Cai, M. Hu, Y. Wang, J. L. Cao, D. Cui, J. Qi and W. Wang (2025). "Lysosomes finely control macrophage inflammatory function via regulating the release of lysosomal Fe(2+) through TRPML1 channel." Nat Commun 16(1): 985.

      Xiong, L., E. Y. Helm, J. W. Dean, N. Sun, F. R. Jimenez-Rondan and L. Zhou (2023). "Nutrition impact on ILC3 maintenance and function centers on a cell-intrinsic CD71-iron axis." Nat Immunol 24(10): 1671-1684.

      Yamada, K., M. Nakata, N. Horimoto, M. Saito, H. Matsuoka and N. Inagaki (2000). "Measurement of glucose uptake and intracellular calcium concentration in single, living pancreatic beta-cells." J Biol Chem 275(29): 22278-22283.

      Yu, L., X. Zhang, Y. Yang, D. Li, K. Tang, Z. Zhao, W. He, C. Wang, N. Sahoo, K. Converso-Baran, C. S. Davis, S. V. Brooks, A. Bigot, R. Calvo, N. J. Martinez, N. Southall, X. Hu, J. Marugan, M. Ferrer and H. Xu (2020). "Small-molecule activation of lysosomal TRP channels ameliorates Duchenne muscular dystrophy in mouse models." Sci Adv 6(6): eaaz2736.

      Zhang, G., X. Cai, L. He, D. Qin, H. Li and X. Fan (2020). "Skimmin Improves Insulin Resistance via Regulating the Metabolism of Glucose: In Vitro and In Vivo Models." Front Pharmacol 11: 540.

      Zhang, H., Y. Wang, R. Wang, X. Zhang and H. Chen (2024). "TRPML1 agonist ML-SA5 mitigates uranium-induced nephrotoxicity via promoting lysosomal exocytosis." Biomed Pharmacother 181: 117728.

      Zhang, X., X. Cheng, L. Yu, J. Yang, R. Calvo, S. Patnaik, X. Hu, Q. Gao, M. Yang, M. Lawas, M. Delling, J. Marugan, M. Ferrer and H. Xu (2016). "MCOLN1 is a ROS sensor in lysosomes that regulates autophagy." Nat Commun 7: 12109.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers for their thoughtful and constructive feedback. In response, we substantially revised the manuscript and clarified the rationale for using in vitro reconstitution assays and iPSC-derived neurons to determine how tau hyperphosphorylation alters its interaction with microtubules and its regulation of intracellular transport. We have also more clearly articulated how these findings relate to neurodegeneration and discussed the limitations of the model systems used.

      The primary concern of the reviewers was the justification for using COS7 cell lysates in reconstitution assays and iPSC-derived neurons as model systems. We have revised the manuscript to clarify that these experimental systems provided a means to isolate and examine how AD-related tau hyperphosphorylation alters tau-microtubule interactions and the regulation of intracellular transport. COS7 cells were selected because they are widely used for expression of mammalian proteins, including kinesins and tau. Human iPSC-derived neurons were chosen for their amenability to TIRF microscopy and transfection-based experiments, as well as the ability to compare CRISPR-generated tau knockout (MAPT-KO) neurons with their isogeneic control counterparts. Accordingly, we have revised the language throughout the manuscript to more clearly define the study’s objectives and emphasize that these systems were intentionally chosen as robust, well-controlled platforms for addressing specific mechanistic questions. We agree that they do not fully recapitulate AD pathology and that more representative models, such as mature, aged neurons or patient-derived neurons would be better suited to studying disease progression, and have included this limitation in the Discussion. However, because the central objective of this study was to dissect the mechanistic consequences of tau hyperphosphorylation on microtubule interactions and intracellular transport, we believe that these experimental approaches are well suited to address the questions asked.

      We also more explicitly addressed how background levels of phosphorylation may contribute to the effects observed with the pseudo-phosphorylation model of AD-related tau perturbations. We’ve addressed this by citing recent studies (Fan et al., 2025; Siahaan et al., 2026; Moretto et al., 2026) that quantitatively assess phosphorylation across expression systems and clarified how our experimental design, which directly compares WT, AP and E14 tau, effectively minimize uncertainty arising from background phosphorylation. While some degree of background phosphorylation is likely to be present, any resulting effects would be expected to occur consistently across all tau phospho-variants. We now discuss the limitation of our study that we did not directly quantify phosphorylation levels in cells.

      The reviewers also expressed concern about the potential influence of endogenous microtubule-associated proteins present in lysates and differences in tau occupancy on microtubules contributing to motility outcomes. To address this, we included additional analyses correlating tau intensity along microtubules with kinesin motility. We also expanded the Discussion to consider how tau competes with other MAPs for microtubule binding and how phosphorylation-dependent changes in tau–microtubule interactions may alter the MAP landscape. Consequently, the transport phenotypes observed with different tau phospho-variants may reflect both direct effects of tau and indirect effects arising from changes in MAP occupancy and competition on the microtubule lattice.

      We provide detailed, point-by-point responses to each reviewer comment below. We appreciate the thoughtful feedback from reviewers and are confident that the revisions, which include clearer language, strengthened justification of the experimental approaches, and additional supporting analyses, have substantially improved the clarity, rationale, and overall impact of the study.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work by Beaudet and colleagues aims at exploring the effect of phosphorylation on the formation of tau envelopes and consequently on axonal transport, both in vitro on reconstituted microtubules and in human excitatory neurons derived from IPSCs.

      The authors found that a relatively widely used construct in which 14 serine or threonine residues, often hyperphosphorylated in Alzheimer's disease, are mutated to alanines (phosphodeficient), increases the density of tau envelopes compared to wildtype tau, whereas a phosphomimetic (same residues mutated to glutamic acid) reduces envelope density both in vitro and in human excitatory neurons derived from IPSCs.

      By analysing the trafficking of different kinesins (KIF1a and KIF5C), they observed different effects of tau phosphorylation status on the movement of these two motors.

      They then analyse transport of lysosomes by employing live imaging of lysotracker in human excitatory neurons derived from IPSCs transfected with wildtype, phosphodeficient or phosphomimetic tau, observing that phosphodeficient tau seems to reduce transport of lysosomes while phosphomimetic increases transport compared to wildtype tau.

      Strengths:

      (1) The work aims to study a novel and underexplored topic in the tau field, tau envelopes, and investigate their relevance to Alzheimer's disease pathology.

      (2) Experiments are well conducted and of high quality.

      Weaknesses:

      Relying only on in vitro reconstituted microtubules and human neurons derived from IPSCs leaves some doubts about the relevance of these results for Alzheimer's disease, considering the embryonic state of IPSCs-derived neurons.

      We agree with the reviewer that iPSC-derived neurons represent an immature state compared with the neurons most affected in Alzheimer’s disease. However, iPSC-derived neurons and in vitro reconstitution are robust experimental approaches that provide insight into (1) the effects of hyperphosphorylation on tau’s cooperative microtubules association and envelope formation, (2) how tau hyperphosphorylation affects the motility of kinesin motors that are sensitive to regulation by tau, and (3) how tau hyperphosphorylation alters the bi-directional transport of endogenous degradative organelles such as lysosomes. Our studies reveal the molecular effects of how hyperphosphorylation influences tau’s role in regulating intracellular transport and we believe that these findings will help to inform future studies examining how tau-related dysfunction first influences axonal transport, which would be expected to alter axonal health and homeostasis prior to the more severe pathological effects observed at later disease stages.

      We have included a paragraph under the subheading ‘Limitations of this study’ in the Discussion section to better contextualize our findings within the broader effort to understand tauopathies and Alzheimer’s disease. We clarify the limitations of using in vitro reconstitution and iPSC model systems on pages 20 and 21.

      Reviewer #2 (Public review):

      This manuscript examines how disease-associated hyperphosphorylation disrupts tau's role as a cooperative microtubule-binding regulator of intracellular transport. Using in vitro reconstitution assays and live-cell imaging in iPSC-derived neurons, the authors employ phosphomutant tau constructs (E14 to mimic hyperphosphorylation, AP to prevent phosphorylation) at 14 disease-associated residues to isolate phosphorylation effects independent of expression system-dependent PTM heterogeneity. The results show that hyperphosphorylated tau fails to form cooperative envelope-like structures on microtubules, instead binding diffusely and dissociating rapidly. In contrast, wild-type and phospho-resistant tau form cohesive envelopes that regulate motor protein access. At the single-molecule level, hyperphosphorylation reduces KIF5C inhibition while maintaining or enhancing KIF1A inhibition through altered processivity and detachment rates. In live neurons, hyperphosphorylated tau phenocopies tau knockout conditions, weakening tau-mediated inhibition of lysosome transport and increasing processive motility. The authors quantify tau binding using Gaussian mixture model-based image analysis and measure tau kinetics via FRAP, demonstrating that hyperphosphorylation-induced loss of cooperative binding correlates with dysregulated organelle transport. These findings establish a mechanism by which phosphorylation-driven disruption of tau's gatekeeper function on microtubules compromises axonal transport prior to aggregation in tauopathies. The paper provides interesting new knowledge for the field, but there are outstanding concerns that could be further addressed by the authors to strengthen and clarify the current manuscript:

      (1) Lack of Phosphatase-Treated Control and Explicit WT Phosphorylation Quantification

      Wild-type tau expressed in insect and mammalian cells is known to be phosphorylated by endogenous kinases (eg, GSK3, CDK5, MARK). The manuscript acknowledges this in the Discussion but provides no phosphatase-treated lysate control or quantification of endogenous phosphorylation on WT tau via phospho-specific Western blots. This leaves ambiguity about whether observed differences between WT and E14 reflect purely the introduced mutations or confounding baseline differences in phosphostate content.

      Tau contains ~85 putative phosphorylation sites and is modified by several kinases in cells. Studies by Siahaan et al. (2026) and Fan et al. (2025) provide detailed insight into tau phosphorylation heterogeneity, its role in protecting the microtubule lattice from severing enzymes, and the implications of phosphorylation patterns for aggregate formation. We reference these papers and include detailed description of these findings when initially establishing our justification for using pseudo-phosphorylation model.

      We used a pseudo-phosphorylation approach to test the effects of phosphorylation of specific residues in the proline-rich region and the pseudo-repeat domain in the C-terminus, which together with the microtubule-binding repeats, establish the minimal regions required for tau’s cooperative microtubule binding (Tan et al., 2019). This system enabled us to dissect the effects of tau phosphorylation without the added complexities of heterogeneity and multiple isoforms of tau that would otherwise be endogenously expressed. We’ve clarified these points in the revised manuscript (Pages 6, 7, 17, and 18).

      Background phosphorylation in the different phospho-variants used might contribute to the observed changes in tau’s MT interactions and regulation of transport. However, based on our results and the significance in the changes between the different phospho-variants, even if there is some basal level of phosphorylation, the results indicate that the effects of the pseudo-phosphorylation sites are strong enough to make observable changes above the basal levels of phosphorylation (see p. 6 of the revised manuscript).

      Disease-associated phosphorylation is likely more heterogeneous and dynamic than the pseudo-phosphorylation mutants used here, and phosphorylation at different sites may differentially regulate tau function (see p. 21 of the revised manuscript).

      (2) Limited Normalization of Motor Effects to Measured Tau Lattice Occupancy

      Although kinesin trajectories are classified inside vs. outside tau envelopes (inherently normalizing to local tau density), motor parameters are not systematically reported as functions of tau fluorescence intensity across all constructs. Co-purifying MAPs or microtubule-modifying enzymes in cell lysates is not quantified or excluded, leaving residual uncertainty about tau-specificity of observed motor inhibition. This should be at least acknowledged in the results section.

      As noted by the reviewer, it is challenging to compare conditions where the occupancy of tau on microtubules is dissimilar across conditions. To address this point, we performed a Spearman’s correlation analysis to compare how tau intensity affects kinesin dynamics along microtubules (Fig S3G). On page 12, our results show that kinesin dynamics are generally reduced in regions of high tau occupancy. However, in regions of comparable higher intensities, KIF5C is less inhibited by E14 tau, whereas KIF1A is less inhibited by AP tau.

      On pages 12 and 13, we acknowledge that while effects from other MAPs or motor proteins could potentially affect kinesin motility, we would expect that any effect from residual lysate components would be similar across tau phospho-variants.

      (3) Insufficient Citation of Prior Neuronal Tau Envelope Evidence

      In the Introduction, the authors state, "it was an open question if tau forms envelopes in neurons," but this understates existing evidence. Tan et al. (2019) report tau neuronal staining consistent with envelope formation, while Siahaan et al. (2021) provide more direct evidence in non-neuronal cells. The framing should acknowledge and integrate these prior findings.

      We agree with the reviewer that evidence from several studies using reconstitution systems, fixed neurons, and live cultured cells provides evidence of tau envelope formation in neurons. Specifically, tau envelopes have been observed along taxol-stabilized or GMPCPP-capped GDP microtubules in vitro (e.g., Dixit et al., 2008; Monroy et al., 2018; Tan et al., 2019; Siahaan et al., 2019), in 4% PFA-fixed and Triton X-100–extracted DIV7 mouse hippocampal neurons (Tan et al., 2019), and in live, non-neuronal U-2 OS cells following taxol treatment (Siahaan et al., 2022) or elevated pH (Siahaan et al., 2024). To our knowledge, our study is the first to demonstrate tau envelope formation in live neuronal cells under normal cell culture conditions. We revised the introduction (see pages 3 and 4) to more precisely position our findings within the context of prior studies.

      (4) Unclear Wording on Expression System-Dependent Phosphorylation

      The sentence "The phosphostate of tau is strongly dependent on the expression system" requires rewording. It is ambiguous whether this refers to the final phosphostate achieved after expression or the inherent phosphorylating capacity of each system. Clearer language would strengthen the methodological justification.

      On pages 6 and 7, we clarify the rationale for using COS7 cells to express GFP-tau and elaborate on recent papers demonstrating how different expression systems used to study tau (e.g., bacterial, insect, mammalian) produce tau with variable phosphorylation patterns (Siahaan et al., 2026; Fan et al., 2025).

      (5) Insufficient Quantification of Motor and Lysosome Transport Effect Magnitudes in Results Section

      The data on molecular motor motility and lysosome transport are densely described. The magnitude of effects (fold-changes, percentage differences) should be explicitly stated in the Results section when first presenting findings to orient readers to biological significance. For example, effect magnitudes for lysosome run lengths, velocities, and directional bias should be quantified in text, not left to figure inspection.

      We now incorporate the relevant quantifications in the text.

      (6) Incomplete Discussion of Projection Domain Necessity for Envelope Formation

      The Discussion states the projection domain is "a critical regulator of both tau-tau and tau-microtubule interactions," but does not engage with prior domain dissection work. Tan et al. (2019) found that the entire projection domain is not necessary for envelope formation in vitro. The authors should discuss which projection domain regions are specifically regulated by phosphorylation vs. required for cooperativity, providing a more nuanced interpretation than implied by their current framing.

      Tan et al. (2019) demonstrated that part of the proline-rich region (residues 198–244) within the N-terminal projection domain and the pseudo-repeat region within the C-terminus, together with the microtubule-binding repeats are the minimal region required to maintain tau’s ability to form cooperative envelopes along microtubules. We revised the text to better incorporate this previous work into the discussion and place our findings within this context. Our work demonstrates how phosphorylation within the proline-rich region and pseudo-repeats are important regulators of tau–tau cooperativity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It is unclear how the method implemented by the authors to identify tau envelopes works exactly (GMM and BIC) and how appropriate it is. It does not appear similar to what others have done in the literature on tau envelopes. Moreover, when checking the intensity plots along microtubules, both in Figures 1C and 2A, one is left to wonder if the observed differences are not simply caused by the different thresholds. Indeed, the intensity profiles do not seem greatly different between conditions in Figure 1C, whereas it is evident that the threshold is quite different, with it being lower for AP tau and higher for E14 tau, which explains the differences in how many envelopes are detected. The authors could also try to quantify in a different way (e.g. even just a threshold based on Average+SD) to see if the results remain the same?

      We used a GMM/BIC approach to avoid biased comparisons of tau envelope formation on microtubules across phospho-conditions. In TIRF assays, intensity signals are inherently inconsistent, making it difficult to directly compare fluorescence intensity signals on different microtubules across regions within the same field of view. Additionally, tau distribution between microtubules and in solution varied between conditions (e.g., background tau signal is elevated in E14 conditions compared to WT or AP tau). Given these challenges, we quantified and compared tau intensities on a per-microtubule basis, which produced more robust results. While we initially attempted the reviewer’s suggested approach of using average + SD, per-microtubule variability in minimum and maximum signal, along with differences in local background, prevented the ability to set a threshold that reliably captured intensity differences along microtubules across and within replicates. We now more clearly explain why we chose this approach (see page 7).

      (2) The authors should discuss the possibility that the presence of a GFP tag at the N terminus of tau could affect the formation of envelopes, given the importance of this region. Also, they refer to tau GFP in some points of the text and other times to GFP tau. As it would seem they have always used tau tagged at the N terminus, they should refer to GFP-tau in order to avoid confusion in the position of the tag.

      We agree with the reviewer that the position of the N-terminal GFP could influence the projection domain, However, all tau constructs carry the GFP tag at the same position and differ only in their phospho-site mutations. The correct nomenclature for “GFP-tau” is now consistently used throughout.

      (3) Previous work (Tan et al., 2019) has shown that different isoforms of tau have different propensities to form tau envelopes. The authors should specify in each figure which isoform of tau they are expressing.

      The tau isoform used throughout this study is 4R0N. The tau isoform is clearly identified in the revised text.

      (4) Figure 1 C-E: It would be interesting to see the size of envelopes quantified, also.

      We now include a comparison of the mean envelope width for each phospho-variant (Fig 1F).

      (5) Figure 2F: As the FRAP experiment is not on tau envelopes but generally on axonal tau, this needs to be clearly stated to highlight how this limits the link between the different FRAP dynamics and the behaviour of tau envelopes.

      We changed the text to indicate that we perform FRAP on axonal tau and not specifically tau envelopes.

      (6) The authors should discuss whether they expect Kif1a and Kif5c to be responsible for transporting lysotracker-positive vesicles in neurons? This does not seem to be the case based on a quick literature search. If these are not the motors responsible for the transport of lysosomes, why do the authors decide to look at the transport of these organelles and not others? Also, what is the rationale for studying the transport of lysosomes, an organelle that is mainly transported retrogradely, after identifying defects in kinesin transport? The authors could either study in vitro the effect of tau phosphorylation on the movement of a kinesin more directly linked to lysosomes (e.g. KIF5B, KIF1B) or study the transport of some other organelle which is mediated by KIF1A and KIF5C.

      We revised the text to clarify this point. Several studies have shown that kinesin-1 and -3 are strongly inhibited by tau, whereas kinesin-2 and dynein are less sensitive (Hoeprich et al., 2017; Chaudhary et al., 2018; and others). Within this context, we asked how phosphorylation alters tau’s inhibitory effects on motors that are most sensitive to tau. The in vitro reconstitution assays were not intended to isolate the effects of tau on lysosome-specific motors. Rather, they were used to determine how tau phosphorylation affects representative kinesin-1 and -3 motors that drive a substantial fraction of anterograde axonal transport and are among the most sensitive to tau-mediated regulation.

      We next examined lysosome transport using LysoTracker to investigate how tau phosphorylation influences bidirectional cargo transport. Lysosomes are transported by teams of kinesin-1, -2, -3, and dynein, making them a useful model for assessing the consequences of tau regulation in a more physiological context. Current models of bidirectional transport proposed that cargo movement emerges from tug-of-war, which is a result of a balance of forces generated by opposing motors. Under this assumption, strong inhibition of kinesin by tau would be expected to reduce anterograde transport and/or enhance retrograde transport by shifting this balance towards dynein. We have clarified throughout the manuscript that our goal was to determine how tau phosphorylation affects bidirectional transport and to interpret these findings within this context. Because defects in degradative pathways are thought to contribute to neurodegeneration, these experiments may also provide insight into how tau hyperphosphorylation disrupts lysosome function during disease.

      Although KIF5C and KIF1A are not the primary kinesin homologs responsible for lysosome transport, we expect that other kinesin-1 and kinesin-3 motors respond similarly to tau. The in vitro findings provide mechanistic insight into how tau phosphorylation could alter motor function and ultimately contribute to changes in lysosome trafficking and distribution within axons. On page 19, we further clarify that the magnitude of tau-mediated regulation is likely to vary among kinesin family members due to differences in their intrinsic motor properties, and that the effects on lysosome transport are therefore expected to be more nuanced than those observed for individual motors in vitro.

      (7) In the trafficking experiments with lysotracker in human excitatory neurons, there seems to be a large fraction of anterogradely transported lysotracker-positive organelles. Based on a quick search, it would appear this occurs frequently in IPSC-derived neurons, but it's not the case in primary neurons (see, for example, Kulkarni et al., 2022). Given that IPSCs-derived neurons maintain an immature embryonic maturation status (as correctly stated by the authors when mentioning that they express mainly 3R tau) and that neuronal maturation influences transport in primary neurons (e.g. Moutaux et al., 2018), the authors should discuss these aspects highlighting the possible limitations, especially considering the claim of importance of their results for Alzheimer's disease, a pathology that hits neurons at full maturation stages. Alternatively, they could perform a similar experiment in murine neurons at mature stages.

      See response to comments from reviewer 1 under ‘weaknesses’.

      (8) In the context of the previous point, the immature phenotype of IPSCs could explain the apparent discrepancy between the results obtained by these authors and previously published work (Hallinan et al., 2019), which found that mature hippocampal neurons expressing E14 tau had reduced transport of lysosomes. Moreover, these authors also described patches of higher intensity of tau along the axons formed by E14 tau compared to WT tau, which are closely reminiscent of tau envelopes. The authors should discuss these discrepancies.

      We agree that there are discrepancies between our findings and those reported by Hallinan et al. (2019). In that study, E14 tau was shown to misfold in cultured mouse hippocampal neurons, forming MC1-positive axonal aggregates that impair lysosome transport. In contrast, we observe nearly the opposite effect: E14 tau remains diffusely distributed in axons and produces a phenotype resembling tau knockout conditions, with enhanced lysosome transport.

      These differences may stem from methodological factors, including the neuronal models used (murine hippocampal cultures versus human iPSC-derived neurons), fixation and immunolabeling compared with live-cell imaging, differences in neuronal maturity (DIV), and the presence of endogenous tau versus our knockout-and-rescue approach. Importantly, tau aggregation may reflect later stages of disease progression, where aggregates physically clog axons leading to obstructed axonal transport rather than tau acting as a regulatory “roadblock” to specific motor proteins. We now cite this paper and compare our results with this study and discuss these discrepancies and their potential implications in the ‘Limitations of this study’.

      (9) Figures 4 and 5 are quite hard to read. Perhaps the distinction between proximal, mid and distal axon, although valuable, could be moved to the supplementary, maintaining an overall average, or the most significant of the 3 in the main figures to improve readability?

      We made substantial revisions to figures 4 and 5 and the associated analysis. Because the effects of tau on lysosomal transport were largely consistent across proximal, mid, and distal axonal regions, we combined these datasets and report overall transport trends within the axon (from ~50 µm distal to the AIS to ~50 µm proximal to the growth cone). The region-specific analyses and figures showing lysosome motility in each axonal segment have been moved to Supplementary Figure S4.

      (10) In the discussion, the authors write an entire paragraph on how their results are important to stress the importance of the N terminus of tau in the formation of tau envelopes. This is based on the fact that most of the residues mutated in the phosphomimetic and phosphodeficient constructs are located in the N-terminal projection domain. However, some of these residues are located in the C terminus of tau, which also appears to have a role in tau envelope formation (Tan et al., 2019). The experiments presented do not discriminate the phosphorylation of which of the 14 residues is important to mediate the effects. Hence, I feel this paragraph needs to be toned down or removed entirely.

      See response to comment 6 from reviewer 1.

      (11) The authors make a point of using tau produced in mammalian cells in the experiments performed in vitro, stressing the advancement compared to previous work that used tau produced in bacteria or insect cells. Although this is certainly closer to physiological conditions, the production is done in cancerous kidney cells, so I feel the author should highlight that neurons might drive a distinct phosphorylation pattern. Could recombinant tau be produced in neuroblastoma cells?

      In the revised manuscript, we’ve addressed this comment in the “Limitations of this study” page 21. We used COS-7 cells, which are not cancerous but immortalized fibroblast-like cells derived from African green monkey kidney obtained from ATCC. These cells were chosen because of their widespread use for protein expression and their high transfection efficiency. We agree that the physiology of COS-7 lysates is not directly comparable to that of neurons. Our intention was to convey that proteins expressed in mammalian systems undergo post-translational modifications and are produced by cellular machinery that more closely resembles neuronal systems than bacterial or insect expression platforms. Although neuronal cell lines such as neuroblastoma cells may appear more physiologically relevant, they are often difficult to transfect (Alabdullah et al., 2019). This can create practical challenges in equalizing protein concentrations and obtaining sufficient amounts of overexpressed tau from lysates. While methods exist to improve transfection efficiency, there is no literature that we found stating that these cells would yield protein expression characteristics more comparable to neurons than COS-7 cells. A more comprehensive evaluation of alternative expression systems would require a substantially deeper literature search or systematic characterization of multiple cell lines, which falls beyond the scope of this manuscript. Therefore, we relied on the robust COS-7 expression system and will clarify this rationale in the revised manuscript.

      Alabdullah AA, Al-Abdulaziz B, Alsalem H, et al. Estimating transfection efficiency in differentiated and undifferentiated neural cells. BMC Res Notes. 2019;12(1):225

    1. Author response:

      The following is the authors’ response to the original reviews.

      As requested by all three reviewers we have added a new figure which applies our new end-to-end sorters on real openly available data. This demonstrates that our modular algorithms, optimized to work on simulated data, also performs well in real world cases. In addition, to demonstrate that our results are not due overfitting on simulated data on a specific probe geometry (Neuropixels 1.0), we added a Supplementary Figure demonstrating that our positive results for the end-to-end spike sorters are observed for numerous probe geometries (Neuropixels 2.0, SiNAPS, tetrode and Cambridge NeuroTech). We hope that the extended applications will convince the reviewers and readers of the robustness of our results.

      Reviewer #1 (Public review):

      Weaknesses:

      The reviewer identifies several weaknesses:

      (1) The main concern is the limited support for the claim that ’Lupin’ and individual modules’ outperform existing spike sorters.

      (2) Evidence is primarily from a single benchmark based on an intentionally simplified simulation. While the authors discuss the trade-offs between simulated and real data, the current evaluation does not provide enough diversity to justify claims of superiority.

      (3) While improving individual modules that run in a serial fashion could aid overall spike sorting performance, acknowledging that some end-to-end sorters work in an iterative fashion across multiple of these modules would be fair. Perhaps the optimal spike sorter is not a serial set of modules.

      (4) There is also a risk of benchmark overfitting. A modular approach makes it easy to select components that excel on specific benchmarks (or a specific project’s data characteristics) without generalizing.

      We would like to thank the reviewer for the comments and the valuable feedback. We revised our paper to answer the major concerns that were raised by the reviewer. Regarding the claims about Lupin (1,2), we modified the manuscript in two directions: (i) we attempted to stress that the goal of the paper is not the introduction of the new Lupin sorter per se, but rather presenting and highlighting the modularity of the sorting components framework. In doing so, we also toned down our claims of superiority; (ii) we added simulations on a diverse range of probes and included three real experimental datasets, obtained with different probes, in the results. Regarding the iterative aspect of some sorters – point (3) – we added a paragraph to the discussion highlighting that Kilosort4, unlike previous versions of KiloSort, [9] does not iterate over modules anymore, and thus it can be regarded as a serial algorithm. In addition, although all sorters we present are serial, our proposed framework does not prevent iterative schemes: for example, one could add a re-clustering step after a first template-matching pass. In fact, the modular framework makes this task even easier than before.

      Finally, related to point (4), we added some comments on the problem of overfitting with modular benchmarks in the Discussion, but we think that the risks have been mitigated with the addition of more end-to-end examples with various probe types and experimental data in the revised manuscript.

      The reviewer also points toward some possible ways to strengthen this work:

      (1) Evaluate on multiple simulation regimes, consider adding at least one biophysically detailed simulation, benchmark on multiple probe-geometries with neurons also clustered in different depth profiles (as this will affect drift solutions), and provide real-data validation. Even without full ground truth, real-data can be evaluated with expert curation, functional validation (e.g., refractory violations, quality metrics, unit waveform consistency), agreement across sorters, and consistency across time.

      (2) Related to real-data applicability, it is also important to acknowledge that modulatory approaches can enable overfitting to the needs of individual projects. Without real-data benchmarking (or benchmark diversity), it is unclear how the framework will guide users towards generalizable ’best practices’ rather than optimized configurations that work for their specific conditions.

      In response to these comments, the manuscript has been extended both with more simulated ground truth recordings and experimental data. For additional ground truth, we generated recordings with various probe geometries, to check that the results observed for our modular pipeline could generalize (see Supplementary Figure). Regarding real experiments, we chose three open dataset of various types (chronic and acute implantations, IMEC and Cambridge Neurotech devices) and compared the results of several sorters at a macroscopic level relying on high-level automatic curation tools [7, 3, 6] to quantify how many “good”, “oversplit” and “noise” units are found. This is now a new Figure 8 in our manuscript. We believe that the results from all these datasets demonstrate that Lupin is, as is claimed in the paper, on par with the most popular spike sorting algorithm, Kilosort4 [9].

      Regarding overfitting, we do not believe this is a direct consequence of the modular approach introduced in this article, but rather a general potential “risk” in the spike sorting field. Each spike sorter exposes a large array of parameters that users can tweak to attempt to optimize outcomes on specific datasets, but this end-to-end fine-tuning is hard to control and quantify. We believe that the modular benchmarks introduced in this paper may enable a finer, more controlled, and quantifiable parameter exploration. As an example, very high firing rates such as those observed in the cerebellum might require different parameters for peak detection than for neurons in the cortex. To test this, one could use our generation framework to mimic key macroscopic features of the system you are studying, and benchmark the peak detection step to find optimal parameters. Overall, we do not see this optimization strategy as a problem. Extracellular electrophysiology is so diverse based on brain regions, species, conditions, tasks, etc. that generalizable “best practices” might not exist.

      Reviewer #1 (Recommendations for the authors):

      (1) Tone down or further support the Lupin and specific modules’ superiority claims.

      This has been modified in the manuscript, and we added a final figure to discuss how Lupin is on par with Kilosort on real data, but with no claim of superiority.

      (2) Add benchmark diversity, this would test generalization and mitigate benchmark overfitting. Specifically: add more probe-geometries, and allow for different depth-profiles of clusters of neurons.

      We thank the reviewer for the suggestion, and indeed, we added some more benchmarks to convince the readers that the results observed can be generalized properly. More specifically, we extended the duration of the recordings to 30 min, and included additional benchmark datasets with four different probe geometries (Neuropixels 2.0, a Cambridge Neurotech probe, a SiNAPS probe and a tetrode).

      (3) Clarify how (x,y) positions of neurons are distributed around the shanks.

      This has been clarified in the Methods section. The (x,y) positions are generated uniformly within a rectangle covering the probe boundaries, plus a 20µm margin. Regarding the depth, z positions (distance from the probe) are drawn uniformly from the range [5,40] µm.

      (4) Add at least one real-data benchmark. Some suggestions for evaluation are: stability of firing rates over time, agreement across sorters, quality metrics, functional validation, expert curation.

      As suggested by almost all reviewers, we added some real-data benchmarks (see last figure). Since experimental data don’t have ground truth, we used automatic curation tools as a proxy for “goodness” of the results. We used automatic labels from Bombcell [3] and UnitRefine, which label units as good, multi-unit activity (MUA) and noise, and SLAy [7] for automatic merging, which correlates with the amount of putative oversplits. We felt it was fairer to use external curation tools instead of creating our own methods to assess quality.

      (5) Clarify recommendations for users facing drift. While your statement of ’not having drift is ideal’ is true, the reality is that many recordings have drift. Some practical solutions would be useful. For example, when to trust results and how to report drift sensitivity.

      The reviewer is right, and we added a sentence to clarify when motion correction methods should be used, in our opinions.

      (6) It would help to see what types of signals are most often missed, for example: low-SNR units, drifting units, bursty units, and show which modules affect which of these issues.

      This is already shown in Figures 4 of the manuscript, at the clustering level. These Figures show that cells with low firing rates and/or low SNR are most likely to be missed by all sorters. Our ground truth simulator does not include a bursting mechanism yet. We think that this would be an interesting aspect to simulate and we plan to include bursting units, with bursty spike trains and waveform modulation, in future releases. We thank the reviewer for the suggestion.

      (7) To overcome the issue of overfitting to specific datasets rather than generalization, it would help if a ’default’ or ’recommended starting point’ for users were described in more detail.

      We overcome the issue of overfitting by adding other artificial ground truth recordings (see Supplementary Figure S1), and also real world dataset (see Figure 8). In all these simulations, Lupin is used with default parameters, and this is, we believe, a good starting point. Of course, for very special needs (animal species, brain structures, ...), one might need to adapt parameters, but so as for any other sorters, and such an exploration of the parameter space is out of the scope of the current manuscript. This has been added in the Discussion.

      (8) A lot of the plots have ticks / labels too small to read in 100%, or show quite low-resolution. For example, Figure 3 and Figure 5. Please homogenize across all figures.

      The figures have been regenerated and homogenized as suggested.

      (9) Consider archiving the GitHub version used to generate the figures on Zenodo (DOI) for posterity.

      This has been done for the current state of the manuscript at https://zenodo.org/records/20407862 and the code is available at https://github.com/SpikeInterface/sorting_components_benchmark_paper

      (10) To make the manuscript more reader-friendly, I recommend adding graphics representing the different methods. For example, in Figure 3 one could add schematics of the two peak-detection methods.

      We thank the reviewer for this suggestion. However, this project and our manuscript is not introducing these methods, only re-implementing them in a modular framework. Hence we feel it is out of the scope of this manuscript to produce schematics of the many methods discussed in the paper readers should view the original sources to find out more information.

      (11) Discuss what is meant by KS-like clustering, as I was under the impression that KS4 is also iterative. This may be on the template-matching side, but it’s difficult to know where you draw the border between iterative clustering and (iterative) template-matching. Potentially, we would want to see these processes as one module together, as many sorters work in an iterative fashion across these steps.

      By KS-like clustering, we meant that the code is a direct port, in Python, of the clustering algorithm implemented in KiloSort4. However, we cannot guarantee that this is the exact same clustering method, because of the way the clustering of KiloSort is interleaved with some others steps part of the algorithm. To be more specific, Kilosort has its own special way of performing matched filtering, with a custom grid of templates generated internally with a higher resolution compared to the recording channels positions. These templates are then used to estimate a putative position of each spikes, and these positions are the ones that are used to start the clustering. Thus, the term KS-like comes from the fact that we do not reproduce this exact same mechanism. In SpikeInterface matched filtering is performed using an equivalent approach, but positions are not estimated in the exact same manner. Since the clustering code is equivalent, the results might differ. Regardless of these details, the clustering is not iterative. In former versions of KiloSort, the core algorithm used ideas from k-SVD algorithms, often used in the Machine Learning community. In such an algorithm, the goal is to learn a sparse dictionary of templates to reconstruct the signals, and indeed, there was some iterations between optimizing the templates and the spike times. However, this is no longer the case in KiloSort4. The algorithm works in a serial manner, following the global strategy mentioned in our paper.

      (12) Rather than showing results for one specific simulation, it would be more convincing if we saw the average result of multiple simulations. We don’t need to see individual neurons necessarily (e.g., Figure 3B and C, but this applies throughout the manuscript).

      While we tend to agree with the reviewer than averaging over multiple datasets might be more informative (as we did in the final figure, for end-to-end sorter comparison), we want to say that given the fact that we are using ground-truth recordings that are randomly generated, as long as we do not change the macroscopic properties of the recordings, results on various instances of the noise will be very similar, and averages might not be as informative as one could hope for. One option would be to vary the parameter space of these ground truth recordings, but then there are so many parameters (noise levels, firing rates, distributions of the cells, ...) than averaging everything, and/or even choosing what should be primarily studied is an open question on its own. However, to ease the redibility of the figures, individual neurons were removed from the plots.

      (13) Legends are often incomplete. For example, describe what individual dots are and what lines are in the different figures, even if it seems obvious.

      Legends have been updated.

      (14) Figure 7 would benefit from average thick curves for each model (optionally with individual thin lines for each instance, or error bars). Individual neurons can be left out. Figure 7E (and likewise Figure 5E) would benefit from having x labels to indicate precise labels, so the color attribution is reserved for specific models. It’s quite an intense ’search’ game to understand these figures.

      All the figures have been regenerated, and for the sake of readability, we removed the scatter plots for the individual neurons, focusing only on the averaged lines.

      (15) While I appreciate the effort is gigantic, it may be better for the reader to conclude that.

      This has been rephrased.

      Reviewer #2 (Public review):

      Major comments

      The model ground truth data used in the paper does not need to be a perfect match to experimental data to provide useful benchmarking. However, as with all measurements of spike sorting accuracy, extrapolation to experimental data can be complicated. Users of these tools will need to assess how well the simulated data matches their recordings.

      We agree with the reviewer that extrapolating our results to real data is difficult, and the same point was raised by the other referees. We have now extended the results to include three experimental datasets from different probes. Due to the lack of ground truth, we used automatic curation tools (Bombcell [3] and UnitRefine for labeling, SLAy for merging [7]) as a proxy for performance and showed that Lupin is on par with Kilosort4 on all datasets. We hope this gives users some idea of how well the new sorters will work on their data.

      Reviewer #2 (Recommendations for the authors):

      (1) Any comparison to experimental data would be welcome. Is it possible to add firing rate, amplitude, and template similarity distributions from measured recordings to the panels in Figure 2D?

      Experimental data is very diverse, depending on brain region, species, recording technology, etc. Instead of extending the comparison between simulated and experimental data, we rely on new Figure 8 to showcase the applicability and performance of the presented methods on real recordings.

      (2) Do yields of units passing basic quality metrics look different with the new sorter vs. established sorters? Another option that would help establish the range of applicability of the results would be a different model ground truth system, e.g., the hybrid ground truth data used by the authors in reference 7.

      As it can be seen on synthetic recordings (e.g Figure 7A-B), all sorters behave similarly with respect to firing rates, signal to noise ratios. To help establish the range of applicability, we added an extra Figure 8 in the manuscript, as described above.

      (3) Interpolation errors are likely larger for NP1.0 probes - which have 40 um vertical pitch - than NP2.0 probes, with 15 um pitch. If it’s possible to include even a small-scale comparison of the results from the 2.0 geometry, that would be very valuable for readers trying to decide what probe type to use.

      In the revised manuscript, we added new ground truth benchmarks with a NP2, and other, layouts to show that the key results do not depend on the geometry of the probe (see Supplementary Figure 1). But to address more specifically the point raised by the reviewer, we note that in the case of applying Lupin to NP2 layouts, there is only a loss of ≈ 7.5% of well-detected units between static and motion-corrected cases (see Supplementary Figure 1). Where when we apply Lupin to NP1 layouts (see Figure 7) there is a decrease of 20% of well-detected units. So clearly, a smaller pitch and the columnar arrangement of NP2 seems to help motion-correction methods, and thus reduce the failures due to motion. This has been added in the Discussion.

      About the text:

      (1) In panel E of Figures 5 and 7: Adding the name of the sorter algorithm under the bar charts would be helpful. It is encoded by the color of the outline of each box, but I found that cue rather subtle.

      The figures have been regenerated, with increased ticks and label fonts. We found that adding the names made the Figures too dense, and acronyms would not simplify the figures. In the end we decided not to add them.

      (2) Is the number of features (K) used for clustering already mentioned in the main text or methods? I couldn’t find it.

      We thanks the reviewer for pointing out this problem, and the answer (K = 5) has been added in the manuscript, in the methods section.

      (3) Equation 2 in Methods, defining accuracy, appears to be incorrect. I believe the correct equation is: accuracy = TPij/(Ni + Nj - TPij).

      The reviewer is right, and this has been corrected

      (4) In the abstract, line 18: Component based spike sorters => component based spike sorter.

      This has been corrected

      (5) Line 169 "we can artificially boost the signal-to-noise..." I don’t really see anything artificial about leveraging the extra information in neighboring sites. Maybe just remove that adjective?

      We removed “artificial” from the sentence.

      (6) Line 196-197, describing the result in Figure 3: Especially since 3A is a log plot, it would be helpful to add a percentage to the spikes missed on the high end. These are probably pretty unusual cases, under 0.1%?

      This has been commented in the text, but both because this number is only an approximation (the problem of pairing peaks between detected and ground truth is slightly ill-defined), and because it depends on the particular seeds and parameters of the artificial ground-truth, we avoided numerical values.

      (7) Line 567: "number of dimensions from M to K 5" => "number of dimensions from M to K[5]" That is, is the 5 meant to be a reference? Or is 5 the number of dimensions retained to describe the temporal waveform?

      5 is the value of K, and this has been corrected in the manuscript.

      (8) Line 585-590: Does this description of how to handle borders between groups of sites also apply to the KS-clustering method?

      For the KS-clustering methods, we used the original method implemented in KiloSort. In fact, KiloSort aggregates all the clusters found during the clustering steps, launched per bins, but duplicated templates (based on their shapes only) are removed before matching (and those are the cells at the borders found numerous times). After the template-matching step, once cells have been “populated” with theirs spikes, templates are once again removed and/or merged. However, as explained in the methods, currently this step has not been implemented in our framework. To be more explicit, during template matching, KiloSort stores the features of all discovered spikes, in a space where the effects of nearby spikes have been subtracted, to get a clearer picture of these features. In this “denoised” space, the clustering algorithms is launched again on all spikes (decimated), to assign labels.

      (9) Line 687: "In fact, a major main with Kilosort lies after the template matching step..." => "In fact, a major difference with Kilosort lies after the template matching step...".

      This has been corrected.

      Reviewer #3 (Public review):

      Major comments

      (1) The simulator itself has to be improved and extended. Right now, it simply generates, for every unit, a mother waveform from a sum of exponentials, scales that over channels, and then adds up multiple instantiations of every unit on every channel, along with noise. This is not a biophysical simulator: it is an ad hoc procedure, and the sentence "we firmly believe that.." (lines 482-483) does not make the procedure convincing. To make the simulator credible, the authors should: (1) use a set of biophysical equations, with multi-compartmental modeling of currents and return currents; (2) use noised data from extracellular recordings; or (3) some combination thereof.

      The reviewer is right when pointing out that the current ground truth generator is not “biophysical”, and this is why in the manuscript we used the terms “biophysically plausible”. However, we decided to changed this phrase to “phenomenological” in order to avoid confusion. We believe that our generation tool has the key ingredients to challenge (and also demonstrate limits of) modern spike sorters, which is the goal of the proposed simulator. Spike sorters performing well on such phenomenological simulated data should be a necessary, but not sufficient, condition to convince experimentalists that they will work well on real data. In previous papers [5, 4], we used MEArec [1], which relies on biophysical modeling of the reconstructed neurons to simulate extracellular potentials to generate ground-truth recordings. However, such biophysical simulations have two shortcomings: i) simulations are very slow and resource-hungry; ii) it is not guaranteed that the superior simulation environment (multicompartment modeling) translates to simulated data that are more similar to experimental data. In fact, the Kilosort4 paper [9] shows that the action potentials generated by MEArec [1] have an almost doubled duration compared to realistic data, which requires further ad-hoc parametrization. We believe this discrepancy can be due to the fact that virtually all multi-compartment models are built from in vitro slices, not in vivo recordings. Given these limitations, we decided to rely on a simpler but better controllable model for generating templates. Despite not being “biophysical”, the generation model can replicate, to some extent, the variability in waveforms (using different parameters for waveforms widths and spatial decays) and allows us to have a ground-truth model for drifting as well, that we can use to assess interpolation errors. Nevertheless, we agree with the reviewer that some aspects of the simulator can be improved, such as the structure of the noise in the data. We are currently working on the Spikeinterface side to improve the generation module so that it supports temporally correlated noise (spatial correlation is already supported and used in this manuscript).

      (2) The simulated dataset has to be extended in time. Maybe I missed something, but 500 units over 10 minutes, with some units having firing rates as low as 0.1 spikes/s, corresponds to some of the units firing an expected 60 spikes. This is clearly too short, and does not replicate the standard situation in extracellular experiments.

      We extended the simulated recording to an half an hour duration. However, as it can be seen in the Figures, this does not affect the main results of the paper.

      (3) The simulated dataset has to be extended in space. The choice of using NeuroPixels 1.0 geometry is a poor one. Many labs use other monolithic electrode arrays (MEAs, silicon probes, other rigid arrays); tetrodes remain a major tool, and flexible probes (polyimide, mesh) are evolving. Assessing algorithms over a single spatial architecture is likely to lead to local maxima in performance and potentially erroneous conclusions.

      We believe the that Neuropixels 1.0 geometry is a good choice: the NP1.0 paper it is the most highly cited paper about a high-density electrophysiology probe, suggested that it is currently the most widely used probe in the world. However, the reviewer is correct that there is a risk of overfitting. To demonstrate generalizability of our algorithm, we have added results obtained with NP2.0 layout, a Cambridge NeuroTech layout, a SINAPs probe layout and tetrodes. The results demonstrate that the Lupin sorter is not overfitted.

      (4) The existing spike sorters evaluated are not completely described. Some sorters (e.g., SpyKING Circus and KS4) were described in previous publications, but it is unclear whether the implementation that was used for the present tests is exactly the same as those previously published. More importantly, some of the sorters evaluated (e.g., TDC, TDC2, SpyKING Circus 2) were never described in a peerreviewed paper. This does not mean that they cannot be evaluated - but if they are, they must be described in full. Relying on the fact that the code is open source cannot replace a complete and accurate scientific description.

      The reviewer is rising a valid point, but we think it is beyond the scope of this paper to fully describe every component of every spike sorter mentioned in the paper. We made the deliberate choice to describe the sorters as chains of components in order to demonstrate the flexibility of our modular approach. Full details of every components can be found in the online documentation. In order to ensure the manuscript felt more complete and concrete, we revised the descriptions in our Methods section, in order to give some more details.

      (5) Related to the above, all relevant code should be made available online in permanent repositories, not only in author-controlled ones.

      We are not sure of what exactly is suggested by the reviewer. In order to clarify the situation, we pushed the notebooks and all the code needed to reproduce the Figures of the paper in a Zenodo archive https://zenodo.org/records/19695406

      (6) It is unclear why SpyKING Circus 2 and TDC2 are evaluated - these could potentially be described as straw men. I recommend reorganizing the manuscript so that after every module is evaluated separately based on a limited ground truth dataset, a single "best" sorter would be constructed, and then tested extensively (and compared to the de facto state of the art). Such reorganization would both demonstrate the utility of a modular approach and clarify the general usefulness of the outcome.

      Although we agree that these two sorters might be seen as straw men, we decided to keep them in the paper for various reasons. This has been clarified in the manuscript, but the primary reason is an historical one: the developers of these sorters decided to unite their efforts while designing new tools and algorithms, which led to the initial work on the modular framework described here. Because SpyKING CIRCUS and TriDesClous had to evolve for maintenance, it was decided to try to write a common “grammar” that would allow these two spike sorters to be described in the same framework. The reviewer is right in the fact that once the foundations were stabilized, most of the development efforts were put to Lupin, that was built as the best combination of all the expertise gained on the two aforementioned sorters. A second reason to keep them is that, once again, we decided that Lupin should not be the main focus of the paper. Of course, this is a nice illustration of what the modular framework can do, but we do not want to push it per se. What matters most is the methodology, especially since we can not claim, here, that Lupin would be the best spike sorter regardless of data types, probe geometries, .... We extended the paper with other probe geometries (see added Supplementary Figure 1) and real world data (see Figure 8), and we observed that Lupin was on par, and/or slightly better than KiloSort 4 with respect to number of False Positives for examples. But ultimately, the paper is really about the development of a common ecosystem such that all tools can be improved upon, at the community level.

      (7) The new algorithms developed, for example, clustering and template matching, have to be described in more detail, and demonstrated graphically on simple datasets. This can be done in supplementary material if the authors prefer not to extend the manuscript too much.

      As suggest by the reviewer, we tried to extend the methods of the clustering and the template matching steps, bearing in mind that some of them have already been published in detail elsewhere. We really want to underline that the central point of the paper is the modularity of the architecture, not so much the low-level details. To populate our framework and demonstrate its generalization, we implemented some key algorithms. But describing with schematics and in depth every individual methods is something that is not even done in papers focused on the spike sorters themselves. Later, the reviewer complains that the paper sounds like a technical report. Delving into more details would only amplify this problem.

      (8) This reviewer finds the description and interpretation of the results to be inadequate. As an example, focusing on Figure 5: The results in Figure 5A have to be supplemented and summarized as a scalar point estimate (e.g., median accuracy), an estimate of dispersion (e.g., using MAD, IQR, or SD), evaluated over multiple runs, and compared using statistical tests between tools and conditions (e.g., using a multi-dimensional analysis of variance, a mixed effect model, etc.). The results in Figure 5D must have an indication of dispersion. Any conclusions based on the numerical experiments must be based on these metrics and statistical evaluations.

      To try and simply the plots and their interpretation, we decided to remove the scatter plots of the individual neurons in all Figures, to really focus on the core trends. We believe that the results, such as the dependence on accuracy as a function of SNR, are difficult to capture with summary statistics, and that the results are best understood by looking at the plots we have made. If we wanted to make strong claims about one algorithm being more suitable for a specific task, then these summary statistics would be suitable. Instead, we are trying to demonstrate the general utility of the components framework, and we optimize Lupin simply to maximize the accuracy of each step.

      (9) The entire MS would benefit from expert proofreading; there are many language errors, mostly in indefinite articles and grammatical numbers.

      The manuscript has been intensively proofread by native english speakers.

      Reviewer #3 (Recommendations for the authors):

      (1) Lines 14-15: "...a... sorters...": either "...sorter..." or "...a... sorter...".

      This has been corrected

      (2) Line 30: "peak detection" or "event detection"?

      We prefer the term “peak detection", since this is exactly what the algorithm are looking for: spatio-temporal extrema in the signals

      (3) Line 30: the purely sequential structure of the modules is very limiting and generally incorrect. Template matching may replace peak detection, as is the case in many real-time hardware implementations. The feedforward process is limiting, and many sorters use feedback or multiple loops. It is unclear whether and how the proposed framework supports such structures.

      The reviewer raises a good point that we did not explain well in the manuscript. Since our framework is modular, with each component independent of the others, a developer has freedom to create “iterative" sorters. E.g. they could loop over a pair of steps until a criteria is met. In fact, this is one of the advantages of creating modular components. The examples we show in the manuscript are sequential feedforward structures, leading to this confusion. We have clarified the point in the text. We show sequential sorters because to our knowledge there are currently no iterative sorters in wide use. Older versions of Kilosort were iterative, but KiloSort4 is not. Modularity also allows us to isolate one component for a specific task. Hence, for a real-time implementation, we can pre-compute templates using the peak-detection and clustering components. Then these steps would be skipped for “online” sorting, which would only use the template matching component. The reviewer states that “template matching may replace peak detection”. Indeed, in Lupin, SC2 and TDC2, template matching does replace peak detection – the initial peak detection is only used to construct templates for downstream matching. We have clarified this point in the text.

      (4) Lines 62-63: Are all of these sorters supported by the spike Interface framework? Please include a table of which are and which are not. The same for every module.

      All the sorters listed are indeed supported by Spike Interface, and this has been added in the text.

      (5) Lines 84, 98, and elsewhere: TriDesClous 2 is mentioned. What about TriDesClous - is there such a sorter, and if yes, what is the scientific reference?

      TriDesClous (https://github.com/tridesclous/tridesclous) is a spike sorting pipeline that has not been properly published with a DOI, but that has been developed by the first author of the manuscript and has been used by many papers [8].

      (6) Line 84: Lupin - suggest reorganizing the manuscript around this sorter and evaluating it on multiple datasets.

      The paper is not about Lupin, and we tried to rewrite the manuscript in order to make this point more explicit. The paper is intended to be a proof of concept of the benefits that can be obtained thanks to a modular approach. This is why we do not want to reorganize the paper on Lupin itself.

      (7) Line 104, Figure 1, and elsewhere: What is the advantage of evaluating three sorters, if two are predicted to be worse than the third/state of the art? Suggest to reorganize the MS: (a) describe a proper simulator; (b) describe every individual modules: mention the existing algorithms and elaborate + demonstrate the new algorithms; (c) evaluate every individual module - on properly realistic dataset; (d) evaluate existing complete spike sorters + the proposed best combination - on the same ground truth dataset; (e) compare the best two sorters on multiple datasets. In addition, may identify the WEAKEST link in each sorter and demonstrate the improvement by replacing ("upgrading") that link alone.

      As clarified in the introduction and in the “End-to-end evaluation..." section, we decided to keep three sorters for various reasons. The first is historical: SpyKING CIRCUS 2 and TridesClous 2 were the first two sorters that motivated the creation of the sortingcomponents framework. Because both authors realized that they had so much code in common, they decided to unite their efforts while rewriting them and share some common building blocks. The second reason is that the paper is not about Lupin, per se, but more about the general philosophy of the modular architecture presented here. We want to push forward the idea, in the community, that we should share tools, ideas and algorithms in order to enhance the analysis pipelines. Keeping several sorters, even if sub-optimal, is a way to showcase the flexibility of the framework, and this is why we did not re-organize the manuscript as suggested by the reviewer.

      (8) Line 121: "can drastically cut the time.." - provide quantitative support. In general, avoid superlatives and unsupported statements.

      Spikeinterface has been primarily design to ease the comparison between spike sorting pipelines [2]. Thus all comparison metrics such as agreement matrices, false positives, false negatives, ... are available out of the box when using these Benchmark objects. It would be hard to provide a quantitative support for such a speedup since it will depend on the algorithm, but it allows developer to simply focus on the core implementation while benefiting for free of the whole ecosystem that will launch benchmarks and compare it, with appropriate metrics validated by a large community. Since this validation process, on its own, can be quite complex depending on the processing step, we truly believe the development gain is important, despite the fact that it might be hard to quantify. However, we rewrote the sentence in order to avoid superlatives.

      (9) Line 129: "powerful and fast way to generate artificial" - again, avoid statements with superlatives and lacking quantitative support.

      We rewrote the sentence to avoid superlatives.

      (10) Line 129: "powerful and fast way to generate artificial" - the assumptions made in constructing the artificial data critically and strongly affect the conclusions of the benchmark processes. For instance, it is well known that about 10% of the spikes in the cortex have a positive extrema, but the generator is limited to negative spikes. Also, the generator produces spatially-displaced and scaled versions of the waveform generated by the putative soma, but extracellular waveforms almost never behave that way. Even if the simulator is improved, it will always remain a simulator, and therefore the caveat should always be kept in mind.

      The reviewer is right: our simulator has some limitations compared to biological data. But the fact is that, even on synthetic data, there is still plenty of room for improvements of current spike sorting pipelines before even getting to real data, where ground truth are unknown. We believe that such ground truth simulator is a necessary, but not sufficient, condition to validate sorting algorithms. Other options such as hybrid recordings would also suffer from the same flaws, and might be even more questionable. In the revised version of the manuscript, we also added tests on real data to compare more qualitatively Lupin and KiloSort. The fact that results are in line with what is observed on synthetic data gives us confidence in our observations.

      (11) The authors should add a limitations section to the MS.

      Some limits have been more extensively discussed in the Discussion of the paper, with respect to the feedforward architecture, the validity of the ground truth data.

      (12) Line 129: Can the benchmark object be used with other data (e.g., existing)? The methods indicate that this is the case - please demonstrate.

      We think that a proper demonstration would be out of the scope of this paper, but indeed, as long as the user can provide a recording alongside with a sorting (exhaustive or not), then the Benchmark objects can be used. This allows the use of hybrid recordings, and/or manually curated datasets where users would be able to provide a ground truth. Of course, depending if the ground truth is exhaustive or not (if one knows the activity of all the neurons in the recordings), the metrics might not be the same. But everything is built-in in the object.

      (13) Line 144: 10 minutes is too short and nearly irrelevant. In particular, many units spike at rates much lower than 0.1 spikes/s, and in 10 minutes would emit less than a score of spikes (e.g., 0.01 spikes/s would accumulate, on average, 6 spikes..).

      We regenerated all the figures in the paper with 30 min long recordings, and the results remain similar to the 10 minute recordings.

      (14) Line 151: firing rates are not independent of the waveform as implicitly assumed; this should be accounted for, at least in the options in the simulator. Same for library waveforms - the firing rates should be a parameter that is optionally provided along with the waveform library.

      We are not sure what is meant by the reviewer. We believe that the point is that some cell types, with particular waveforms, have particular firing rates, such as fast-spiking interneurons. This could be dealt with in the current simulator, since users can provide, on a per cell basis, some particular values to generate the waveforms of the neurons. One could clearly imagine having some particular cell types with dedicated waveforms and firing rate parameters. However, for the sake of simplicity in the paper, we chose not to use such granularity. What the simulator can not do, at the moment, is to perform amplitude modulation of the templates as function of bursts for example. But this could easily be implemented, and should be part of future works to consolidate the generator.

      (15) Figure 2A, D: It is unclear why the specific zig-zag motion was simulated. Please rationalize, or use a motion from a real dataset. In particular, simulate (a) breathing-induced micro-motions, (b) gradual drift, and/or (c) a step jump.

      We used a zig-zag plus a Brownian motion to cover a rather broad range of continuous motion. Of course, as pointed out by the reviewer, the heterogeneity of real drifts is large, and again, we believe it is out of the scope of the manuscript to cover them all. In previous works, we explored how motion correction methods were working as function of the drifts [5], and the conclusion was that this simple continuous drift was already challenging enough to make the spike sorters fail. Adding discontinuities, as often encountered in experiment, would only make things worse. Finally, we added real world data (Figure 8) from three randomly picked dataset with heterogeneous probe geometries. They all come with some drift, that might be representative of what is typically dealt with.

      (16) Line 181: "more computationally demanding" - quantification?

      We forgot to add the reference to panel 3D, and this has been corrected in the manuscript.

      (17) Line 192: "(Figure 3A" - add ")"

      This has been corrected

      (18) Lines 195-6: "... not that missing a few spikes at peak detection will not have a large impact..." - this statement should be quantified.

      We reformulated the sentence, but the point here is simply that template-matching based algorithms use the template-matching step exactly for this purpose: to label spikes that would have been ignored/missed by the peak detection and clustering steps. The fact that template-matching based pipelines such as KiloSort, SpyKING-CIRCUS, ... outperforms clustering-based solution when detecting spikes [4] is a clear support for our sentence.

      (19) Lines 195-6: "... not that missing a few spikes at peak detection will not have a large impact..." - this statement, if correct, exposes a key weakness of the purely modular approach. For instance, assume that module 2.1 performance can be 0.5 and module 2.2 performance is 0.9, and that will have zero impact on the overall performance of a sorter that has module 4.1 as its fourth module, yielding an overall performance of 0.7. But when module 4.2 is used, module 2.1 performance of 0.5 results in an overall performance of 0.5, whereas module 2.2 performance of 0.9 translates to an overall performance of 0.9. The point should be clear now: evaluating every module in isolation cannot fully predict the behavior of the full system - even if a purely feedforward, single iteration (no loop) architecture is assumed. This requires algorithmic support in the proposed framework, and at the very least explicit discussion.

      The reviewer is right, benchmarking every module in isolation cannot fully predict the behavior of the full system. However, the fact that Lupin, built as an optimal combination of the components and can be on par, if not better in some situations, than KiloSort 4 on various artificial and real dataset makes a compelling argument in favor of assembling the best algorithmic pieces one after the other. In fully assembled spike sorting pipelines, the failures at one stage might be compensated by some algorithmic optimizations latter on. But still, being able to identify these failures, and eventually correct them at the appropriate level, i.e. as soon as they appear might be beneficial for the development of the tool, and to ease their readability/maintenance. This has been added in the discussion.

      (20) Line 203: "we hope that future efforts": This statement is strange for two reasons: (a) if the authors hope for it, why not simply do it? (b) if the authors cannot do it / it is out of scope, the proper place to mention their hopes is in the Discussion section.

      This has been moved to the Discussion section

      (21) Line 232: "motion is only partially compensated for" - so what is the utility of the compensation? This becomes clear later, but the statement is obscure at this point - please clarify.

      This has been clarified.

      (22) Figure 4A, right: The fact that the lines do not reach unity even for high FRs suggests that the FRs are NOT the key limiting factor. Please provide a scalar measure and a 2D analysis of the accuracy as a function of FRs and SNRs, separately for static and motion-corrected.

      We are not sure that we understand what is being requested by the reviewer. A scalar measure with a 2D analysis of the accuracy as a function of FRs and SNRs would mean, per case (static and motion corrected) at least 4 panels, thus a total of 8 panels. This seems like an overly dense figures. Further, a scalar measure is unlikely to provide more insight into the analysis.

      (23) Line 275: "kriging method" - describe.

      For a full description we refer readers to the original paper describing the method [9] and other work on motion correction [5]. We have added a brief description of the method in the text in the "Motion interpolation reduces spike sorting performance" section.

      (24) Lines 309-310: Where are the templates from in this case - true or estimated? If the latter, say it.

      This has been clarified.

      (25) Line 317: This is the point where I almost gave up - the MS up to this stage seemed very much like a technical report and is not organized in a clear, results-oriented manner. Even for a Methods-oriented paper, one expects to learn the key results, but those are lost in this MS.

      While we agree that this part of the manuscript is rather technical, this is something that we believe is at the core of some scientific questions that have not been yet properly addressed by most of the spike sorting pipelines. The idea of performing template-matching, i.e. to seek for spatio-temporal pattern in the signals while at the same time compensating only partially the motion has never been properly benchmarks. While template-matching is clearly a game-changer in static recordings, i.e. without motion, the extra-addition of motion might limit its performances. We firmly believe that our modular approach can be used to isolate such core questions from the whole integrated pipelines, and guide both users and developers.

      (26) Figure 6A, second and third panels: The fourth (rightmost) panel shows that there are differences between the "true" and "interpolated" templates, but this cannot be seen in the two middle panels - probably because too much information is overloaded on every channel. My suggestion is to show a reduced number of channels, but for each of those, show all five waveforms (corresponding to different spatial positions of the centers of mass) in a NON-overlapping display.

      The figure has been redrawn, as requested by the reviewer.

      (27) Figure 6 and the entire issue of the failure in motion correction + lines 331-332: It is not 100% convincing that the failure is not at the DREDGE level - please demonstrate that the motion is estimated perfectly.

      We checked that the DREDGE algorithm is correctly estimating the motion. To convince the reader this is the case, this has now been shown in Figure 2, panel D, where the simulated motion is displayed on top of the motion estimated by DREDGE. As shown, both are very similar.

      (28) Figure 6 and the entire issue of the failure in motion correction: If motion is estimated perfectly and interpolation is done properly, then is the failure an outcome of discrete spatial sampling? In other words, if the movement in space over time (in the simulator) were limited to discrete steps that correspond exactly to the positions of the electrodes, would the motion correction allow perfect performance? Stated differently: if the electrodes were not 15 or 20 micro-meters apart but rather 1 micro-meter apart, would the difference (e.g, between Figure 4A/4B, or between Figure 5A and Fig. 5B) be reduced? Please check and report.

      The reviewer is raising an interesting point, and indeed, this is something that we are planning to investigate. We felt it would have been too technical for the scope of this paper, which is aimed at showcasing the modular framework and its possibilities. But we are planning to explore such failures with ultra-dense probes as has been done in the DREDGE paper [10]. We also have the intuition that errors are originating from failures of discrete spatial sampling: hence the larger the sampling, the more pronounced the errors.

      (29) Figure 6B: It is impossible to discern which line is which - this underscores the general point of summarizing every CDF by a scalar with measures of dispersion (i.e., descriptive statistics) + statistical testing (i.e., quantitative statistics).

      While we agree with the reviewer that the plots are dense, they convey the global message of the paper and are the same that were used in [9]. To ease the comparison, we decided to stick to this representation.

      (30) Figure 6D: add error bars.

      There are no error bars, since there was only a single run in this figure, i.e. on a single dataset.

      (31) Line 318, line 328: So what is the result here - that the interpolation idea is not useful? If yes, what is the source of the failure? Is it discrete/poor spatial sampling as suggested above? Something else?

      The result is that interpolation, rather than motion estimation, is the bottleneck in recoding with motion. A careful analysis using data from dense probes, as already stated above, would be the focus of further work. But we suspect that errors results mostly from bad interpolation due to discrete spatial sampling.

      (32) Line 334: up to this point, there are no quantified results - and actually, no clear results at all.

      We rewrote the section accordingly, to make the key observations more striking. The goal of this section is not to propose quantified results and say which interpolation methods is the best (a full paper should be devoted to this question), but rather to show the possibilities offered by the modular framework, looking at questions that have been not yet properly addressed in spike sorting algorithms.

      (33) Line 345: "This is likely due" - provide support? Example?

      We rephrased the sentence accordingly.

      (34) Line 349: "this is mostly because both clustering..." - but feature detection was evaluated together with clustering, so a conclusion specifically about clustering is unwarranted. This is a general point - can the proposed software/modular framework evaluate feature extraction separately from clustering? If yes, please demonstrate.

      The reviewer is right about the fact that feature extraction was evaluated together with clustering, however, we want to stress that it is exactly the same method used by all the clustering algorithms developed in the paper. Thus, even if the feature detection method might not be optimal, it is not introducing any biases. Currently, almost all spike sorting algorithms these days are using PCA (or truncated SVD) as a feature detection method, on nearby channels where peaks are detected, and this is why we decided to only focus on this method.

      (35) Lines 358-374: If SpyKING-CIRCUS 2 and TriDesClous do not provide any advantage, why include them in the MS? Many other sorters could be included as well. In other words, what do we LEARN from the failure of these sorters to compete with KS4 and Lupin? If nothing, please remove them. If something, please state the conclusions clearly.

      Both these software have advantages, but we decided to keep them as they offer direct illustrations than implementation of fully integrated pipelines, other than Lupin, are possible within the proposed framework. TriDesClous2 is fast, and SpyKingCircus2 (in line with Spyking Circus [11]) is mostly tailored for in-vitro data, and thus can scale for more than thousands of channels while some other algorithms might not [9]. Of course, listing all the pros and cons of each software would be impossible within the scope of the paper, but we decided to keep them for illustrative purpose.

      (36) Line 369: "the main advantages of these sorters lie in their modularity" - modularity is an advantage only if it is useful for something - if the sorters are modular but perform poorly, the point is nullified.

      We agree with the reviewer, but we would like to point out that here, none of the modular sorters shown in the paper are performing poorly, and all have pros and cons as said in the previous point.

      We believe that this is important to show that modularity can offer some options both for users and developers with respect to probe geometries, animal species, data types, ...

      (37) Line 371: "on our dataset" - this is a very important limitation. See general comments, and discuss in the Limitations section of the Discussion.

      We extended the paper by adding more datasets, for different probe layouts, and also real datasets with meta comparison of state of the art spike sorters (see Supplementary Figure S1).

      (38) Figure 7C: the difference between the cyan (KS-like) and the orange (KS4) lines makes the KS-like irrelevant. Please improve or remove.

      The reviewer is right, and we removed the KS-like pipeline, since it at the stage of the manuscript it is not yet exactly like KiloSort 4.

      (39) Line 379: "each individual algorithmic steps" - should be "step".

      This has been corrected

      (40) Line 380: Is Lupin available for download and usage as a separate, standalone package - i.e., not as part of the spike interface framework? If not, please make it available. And if yes, indicate this clearly.

      Lupin is part of the SpikeInterface project, and thus can not be installed in a standalone mode, outside of the SpikeInterface framework

      (41) Line 385: "our... gigantic...effort" - avoid superlatives, especially to self.

      This has been removed

      References

      (1) A. P. Buccino and G. T. Einevoll. Mearec: a fast and customizable testbench simulator for ground-truth extracellular spiking activity. Neuroinformatics, pages 1–20, 2020.

      (2) A. P. Buccino, C. L. Hurwitz, S. Garcia, J. Magland, J. H. Siegle, R. Hurwitz, and M. H. Hennig. Spikeinterface, a unified framework for spike sorting. Elife, 9:e61834, 2020.

      (3) J. M. J. Fabre, E. H. v. Beest, A. J. Peters, M. Carandini, and K. D. Harris. Bombcell: automated curation and cell classification of spike-sorted electrophysiology data.

      (4) S. Garcia, A. P. Buccino, and P. Yger. How do spike collisions affect spike sorting performance? Eneuro, 9(5), 2022.

      (5) S. Garcia, C. Windolf, J. Boussard, B. Dichter, A. P. Buccino, and P. Yger. A Modular Implementation to Handle and Benchmark Drift Correction for High-Density Extracellular Recordings. eNeuro, 11(2):ENEURO.0229–23.2023, Feb. 2024.

      (6) A. Jain, R. Greene, C. Halcrow, J. A. Swann, A. Kleinjohann, F. Spurio, S. Graff, A. Pan-Vazquez, B. Kampa, J. Gall, S. Grün, O. Winter, A. Buccino, M. H. Hennig, and S. Musall. UnitRefine: A community toolbox for automated spike sorting curation.

      (7) S. Koukuntla, T. DeWeese, A. Cheng, R. Mildren, A. Lawrence, A. R. Graves, K. E. Cullen, J. Colonell, T. D. Harris, and A. S. Charles. SLAy-ing oversplitting errors in high-density electrophysiology spike sorting. bioRxiv: The Preprint Server for Biology, page 2025.06.20.660590, 2025.

      (8) J. Magland, J. J. Jun, E. Lovero, A. J. Morley, C. L. Hurwitz, A. P. Buccino, S. Garcia, and A. H. Barnett. Spikeforest, reproducible web-facing ground-truth validation of automated neural spike sorters. Elife, 9:e55167, 2020.

      (9) M. Pachitariu, S. Sridhar, J. Pennington, and C. Stringer. Spike sorting with Kilosort4. Nature Methods, 21(5):914–921, May 2024. Publisher: Nature Publishing Group.

      (10) C. Windolf, H. Yu, A. C. Paulk, D. Meszéna, W. Muñoz, J. Boussard, R. Hardstone, I. Caprara, M. Jamali, Y. Kfir, D. Xu, J. E. Chung, K. K. Sellers, Z. Ye, J. Shaker, A. Lebedeva, R. T. Raghavan, E. Trautmann, M. Melin, J. Couto, S. Garcia, B. Coughlin, M. Elmaleh, D. Christianson, J. D. W. Greenlee, C. Horváth, R. Fiáth, I. Ulbert, M. A. Long, J. A. Movshon, M. N. Shadlen, M. M. Churchland, A. K. Churchland, N. A. Steinmetz, E. F. Chang, J. S. Schweitzer, Z. M. Williams, S. S. Cash, L. Paninski, and E. Varol. DREDge: robust motion correction for high-density extracellular recordings across species. Nature Methods, 22(4):788–800, Apr. 2025.

      (11) P. Yger, G. L. Spampinato, E. Esposito, B. Lefebvre, S. Deny, C. Gardella, M. Stimberg, F. Jetter, G. Zeck, S. Picaud, et al. A spike sorting toolbox for up to thousands of electrodes validated with ground truth recordings in vitro and in vivo. Elife, 7:e34518, 2018.

    1. Author response:

      We thank the editors and reviewers for their thoughtful feedback on our manuscript. We are encouraged that they recognized the importance accounting for media consideration when modeling in vitro disease phenotypes, as well as the value of this dataset as a resource for the field. We appreciate the points raised regarding biological replicates and data normalization methods, and we plan to address these fully in our formal response and in revisions to the manuscript. Below, we provide preliminary responses to several comments and indicate how we anticipate addressing them in the revised manuscript.

      (1) Reviewer 1 & 3: Unclear definition of iPSC lines/clones used in data generation.

      We thank the reviewers for pointing out this ambiguity in the Methods. The project was completed with multiple donor iPSC lines, with at least 2 clones generated from each line, and each experiment was performed using at least three independent iPSC RPE lines. To minimize confounding variability, we took several precautions where possible: each experiment was performed within the same culture plate (coated with Matrigel from the same lot number), using RPE seeded at the same time to ensure comparable maturity, and RPE of the same passage number were used across multiple experiments to reduce de-differentiation or senescence effects. We agree that RPE derived from different iPSC differentiation batches can vary. For this reason, all iPSC RPE used in this study were generated from a single differentiation attempt. Because the goal of this project was to isolate the impact of nutrient composition on RPE phenotype, rather than to characterize variability arising from clonal or donor differences, we did not stratify our analysis by clone or donor. For imaging-based assays, multiple fields were selected at random from each well to ensure representative sampling. TEM analysis of sub-RPE deposits was performed with n=3 independent filters per medium condition, with three panoramic sections imaged per filter. We will add these details, including the number of clones used per iPSC line, to the revised Methods to clarify experimental unit and level of replication for each assay and will include sample number in legends.

      (2) Reviewer 2: In the Seahorse studies provided in Figure 3. basal readings for OCR are abnormally low compared to Oligomycin treatment and background, suggesting difficulties with the assay. Findings should be taken with caution.

      We thank the reviewer for this careful reading. The apparent discrepancy likely arises from comparing the raw OCR trace (left) versus the background-subtracted “Basal respiration” bar graph (right) in Figure 3A. In the raw trace, basal OCR (~45-65 pmol/min) is appropriately higher than both the oligomycin-treated (~35-45 pmol/min) and background (~30-45 pmol/min) rates, as expected. The bar-graph value is smaller (~7-25 pmol/min) only because the non-mitochondrial rate has been subtracted out, whereas the oligomycin plateau in the raw trace has not. These raw basal values fall within Agilent's recommended starting range for the XFe96 platform (~20-160 pmol/min), consistent with the modest basal energy demand of quiescent, differentiated RPE rather than an assay problem. A recent survey of 530 published Cell Mito Stress Tests [1] found that 17% report the implausible result of maximal OCR below basal OCR, and higher basal rate can compromise the FCCP-stimulated maximal rate. We titrated our cell numbers before this assay to ensure a clear FCCP response. The substantially increased maximal rates in all six media conditions indicate a technically sound assay. We will clarify this calculation in the revised Methods.

      (3) Reviewer 3: The metabolic analyses also require additional methodological clarification. For intracellular metabolomics, the culture format, cellular biomass, extraction volume, pooling strategy, and normalization method are not reported sufficiently. Normalization of extracellular measurements to unspent medium accounts for differences in starting metabolite abundance but not for differences in cell number or biomass. Similarly, normalization of intracellular signals to medium 1 does not correct for differences in the amount of cellular material extracted.

      We agree that these methodological details require clarification. Intracellular metabolomics was performed on RPE lysates collected from 12-well plates (n=3 independent wells/RPE lines per medium), each extracted separately without pooling. RPE were scraped directly into a fixed volume of 300 µL chilled 80% methanol per well, regardless of the medium condition; 10 µL of the resulting lysate was dried together with an internal standard (nicotinamide-D4), reconstituted in 100 µL of mobile phase, and 5 µL was injected for LC-MS/MS analysis. Media were processed in parallel using an identical workflow: 50 µL of conditioned media was collected at 24 and 48 hours, of which 10 µL was mixed with 40 µL cold methanol, and 10 µL of the resulting supernatant was dried with internal standard, reconstituted in 100 µL mobile phase, and 5 µL injected for analysis.

      Biomass data, including average nuclei count and total protein content reported in Supp. Fig. 2D, were obtained from RPE cultured in parallel under identical conditions. Because nuclei count and protein content did not consistently agree with one another across the six media, metabolite intensities were not normalized to either protein content or cell number. Instead, the intensity of each metabolite was instead normalized to Medium 1 to allow relative comparison across conditions, avoiding an additional, potentially skewed layer of correction from an imperfect biomass metric. We acknowledge this as a limitation of the method: since RPE size and biomass differ across media, our fold-changes reflect metabolite pool per well rather than per cell, which could over- or under-represent true per-cell differences in media that yield especially large or small RPE. We will clarify these details in the Methods and Discussion.

      (4) Reviewer 1: Since a major purpose of the manuscript is to highlight how cell culture conditions influence RPE biology and metabolism, it would be helpful to also report whether Mycoplasma testing was performed and confirmed to be negative across all cell lines.

      We thank the reviewer for the suggestion to include this information. To confirm, Mycoplasma testing was performed on all cell lines used in this study with negative results. We will add this information in the revised Methods.

      (5) Reviewer 3: public availability of the underlying metabolomics data would be important for a study intended to serve as a community resource.

      The metabolomics data has been deposited to UCSD Center for Computational Mass Spectrometry (CCMS) repository (Dataset: MSV000095024) and will be made publicly accessible upon publication.

      References:

      (1) Ransy C, Boissan M, Hammad N, Bouaboud A, Issad T, De Dieuleveult M, Miotto B, Ye M, Pasmant E, Bouillaud F. Extracellular flux analyses indicate low ATP yield and require refinement for accurate determination of maximal oxygen consumption rate. Sci Rep. 2026 Jun 10;16(1):18344.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Thach et al. report on the structure and function of trimethylamine N-oxide demethylase (TDM). They identify a novel complex assembly composed of multiple TDM monomers and obtain high-resolution structural information for the catalytic site, including an analysis of its metal composition, which leads them to propose a mechanism for the catalytic reaction.

      In addition, the authors describe a novel substrate channel within the TDM complex that connects the N-terminal Zn<sup>2+</sup>-dependent TMAO demethylation domain with the C-terminal tetrahydrofolate (THF)-binding domain. This continuous intramolecular tunnel appears highly optimized for shuttling formaldehyde (HCHO), based on its negative electrostatic properties and restricted width. The authors propose that this channel facilitates the safe transfer of HCHO, enabling its efficient conversion to methylenetetrahydrofolate (MTHF) at the C-terminal domain as a microbial detoxification strategy. Experimental data that shows an involvement of TDM in the reaction of HCHO with THF is less convincing.

      Strengths:

      The authors provide convincing high-resolution cryo-EM structural evidence (up to 2 Å) revealing an intriguing complex composed of two full monomers and two half-domains. They further present evidence for the metal ion bound at the active site and articulate a hypothesis for the catalytic cycle. Substantial effort is devoted to optimizing and characterizing enzyme activity, including detailed kinetic analyses across a range of pH values, temperatures, and substrate concentrations. Furthermore, the authors validate their structural insights through functional analysis of active-site point mutants.

      In addition, the authors identify a continuous channel for formaldehyde (HCHO) passage within the structure and support this interpretation through molecular dynamics simulations. These analyses suggest an exciting mechanism of specific, dynamic, and gated channelling of HCHO. This finding is particularly appealing, as it implies the existence of a unique, completely enclosed conduit that may be of broad interest, including potential applications in bioengineering.

      Weaknesses:

      Although the idea of an enclosed channel for HCHO is compelling, the experimental evidence supporting enzymatic assistance in the reaction of HCHO with THF is less convincing. The linear regression analysis shown in Figure 1C demonstrates a THF concentration-dependent decrease in HCHO; however, it is well established that HCHO and THF can react spontaneously in a non-enzymatic manner, raising the possibility that the observed effect does not require enzymatic involvement. I appreciate the authors' clarification that the data in Figure 1 were not intended to demonstrate enzymatic channelling or catalytic involvement in the HCHO-THF reaction, and that the assay does not distinguish between changes in HCHO production and downstream consumption. However, the statement "these findings show that TDM carries out two linked reactions: TMAO demethylation at one active site, and the HCHO produced can condense with THF at the C-terminal domain, connecting TMAO breakdown to one-carbon metabolism" (page 2) still implies a mechanistic and functional coupling that is not supported by the presented data and appears inconsistent with the authors' clarification. In light of this, I recommend revising this statement to avoid implying mechanistic or functional coupling between the two reactions unless additional experimental evidence is provided.

      We thank the reviewer for this clarification. We have revised as per recommendation (page 2).

      “Overall, these findings suggest that TDM-mediated TMAO demethylation generates HCHO, which can subsequently react with THF, potentially linking TMAO breakdown to one-carbon metabolism.”

      Overall, the authors were successful in advancing our structural and functional understanding of the TDM complex. They suggest an interesting oligomeric complex composition which should be investigated with additional biophysical techniques.

      Additionally, they provide an intriguing hypothesis for a new type of substrate channelling. Additional kinetic experiments focusing on HCHO and THF turnover by enzymatic proximity effects would strengthen this potentially fundamental finding. If this channelling mechanism can be supported by stronger experimental evidence, it would substantially advance our understanding and knowledge of biologic conduits and enable future efforts in the design of artificial cascade catalysis systems with high conversion rate and efficiency, as well as detoxification pathways.

      Reviewer #2 (Public review):

      Summary:

      The manuscript reports a cryo-EM structure of TMAO demethylase from Paracoccus sp. This is an important enzyme in the metabolism of trimethylamine oxide (TMAO) and trimethylamine (TMA) in human gut microbiota, so new information about this enzyme would certainly be of interest.

      Strengths:

      The cryo-EM structure for this enzyme is new and provides new insights into the function of the different protein domains, and a channel for formaldehyde between the two domains.

      Weaknesses:

      (1) The proposed catalytic mechanism in this manuscript does not make sense. Previous mechanistic studies on the Methylocella silvestris TMAO demethylase (FEBS Journal 2016, 283, 3979-3993, reference 7) reported that, as well as a Zn2+ cofactor, there was a dependence upon non-heme Fe2+, and proposed a catalytic mechanism involving deoxygenation to form TMA and an iron(IV)-oxo species, followed by oxidative demethylation to form DMA and formaldehyde.

      In this work, the authors do not mention the previously proposed mechanism, but instead just say that elemental analysis "excluded iron". This is alarming, since the previous work has a key role for non-heme iron in the mechanism. The elemental analysis here gives a Zn content of about 0.5 mol/mol protein (and no Fe), whereas the Methylocella TMAO demethylase was reported to contain 0.97 mol Zn/mol protein, and 0.35-0.38 mol Fe/mol protein. It does, therefore, appear that their enzyme is depleted in Zn, and the absence of Fe impacts on the mechanism, as explained below.

      The proposed catalytic mechanism in this manuscript, I am sorry to say, does not make sense, for several reasons:

      (i) Demethylation to form formaldehyde is not a hydrolytic process; it is an oxidative process (normally accomplished by either cytochrome P450 or non-heme iron-dependent oxygenase). The authors propose that a zinc (II) hydroxide attacks the methyl group, which (a) is unprecedented, (b) even if it were possible, would generate methanol, not formaldehyde.

      (ii) The amine oxide is proposed to deoxygenate, with hydroxide appearing on the Zn - unfortunately, amine oxide deoxygenation is a reductive process, for which a reducing agent is needed, and Zn2+ is not a redox-active metal ion;

      (iii) The authors say "forming a tetrahedral intermediate, as described for metalloprotease," but zinc metalloproteases attack an amide carbonyl to form an oxyanion intermediate, whereas in this mechanism, there is no carbonyl to attack, so this statement is just wrong.

      So on several counts the proposed mechanism cannot be correct. Some redox cofactor is needed in order to carry out amine oxide deoxygenation, and Zn2+ cannot fulfil that role. Fe2+ could do, which is why the previously proposed mechanism involving an iron(IV)-oxo intermediate is feasible. But the authors claim that their enzyme has no Fe. If so then there must be some other redox cofactor present. Therefore, the authors need to re-analyse their enzyme carefully and look either for Fe or for some other redox-active metal ion, and then provide convincing experimental evidence for a feasible catalytic mechanism. As it stands the proposed catalytic mechanism is unacceptable.

      Revised version. The authors have essentially not changed the proposed mechanism. They have removed the reference to zinc metalloproteases, but still propose a mechanism mediated only by Zn2+. As explained above, attack by zinc (II) hydroxide is unprecedented and would generate methanol, not formaldehyde, and amine deoxygenation is a reductive process that cannot be fulfilled by Zn2+. So the proposed mechanism is still not feasible at all. The authors now say that "oxidative chemistry....remains unresolved", I'm sorry, but that is not acceptable.

      I have urged the authors to re-examine the metal content of their enzyme, In the Supporting Information (Figure S5) they give ICPMS data that indicates a Zn stoichiometry of 0.5 mol Zn/mol protein, and Fe is not detected. Have the authors analysed for other redox active metals? The authors say that there is no evidence for any other metal binding site, but there is only 50% occupancy of Zn in their protein, so could there be a different metal ion present in place of Zn in the other 50% of the protein, that accounts for the observed activity?

      Since there is clearly a major discrepancy here, the onus is on the authors to explain the discrepancy, rather than just returning with the same data. For example, they could treat the enzyme with EDTA to remove all metals (and check the treated enzyme by ICPMS), and then add different metal ions to test activity with different metals (could even titrate with different molar equivalents of metal ions). They could then test a range of different redox-active metal ions.

      We have re-examined our data and repeated experiments on the reviewer's opinion. We have repeated the IC-PMS several times with different preps, including full scans (data presented). Our enzyme is active, but no iron signal is detected. Moreover, the experimentally determined structure does not support the presence of a non-heme iron-binding site (Bugg TDH, Ramaswamy S., doi:10.1016/j.cbpa.2007.12.007, and other papers). More detailed response in the recommendations to authors.

      (2) Given the metal content reported here, it is important to be able to compare the specific activity of the enzyme reported here with earlier preparations. The authors have now done this in the revised version.

      (3) The consumption of formaldehyde to form methylene-THF is potentially interesting, but the authors say "HCHO levels decreased in the presence of THF", which could potentially be due to enzyme inhibition by THF. Is there evidence that this is a time-dependent and protein-dependent reaction? Not yet addressed.

      We thank the reviewer for this important point. At present, we have not performed detailed time-dependent or protein-dependent analyses to determine whether the observed decrease in HCHO levels in the presence of THF reflects enhanced downstream consumption or indirect effects, such as inhibition of TDM activity by THF. We acknowledge that further kinetic and protein-dependence studies will be important directions for future work.

      Also in Figure 1C, HCHO reduction (%) is not very helpful, because we don't know what concentration of formaldehyde is formed under these conditions; it would be better to quote in units of concentration, rather than %. This point has been addressed by the authors in the revised version.

      (4) Has this particular TMAO demethylase been reported before? It's not clear which Paracoccus strain the enzyme is from; the Experimental Section just says "Paracoccus sp.", which is not very precise. There has been published work on the Paracoccus PS1 enzyme, is that the strain used? Details about the strain are needed, and the accession for the protein sequence. Addressed in the revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      As noted above, there is still a major problem with the proposed mechanism not being feasible, and remaining questions about the presence or absence of a redox-active metal ion in their enzyme. They should:

      (1) Re-examine for other metal ions (apart from Zn and Fe) using ICPMS. The redox metal ion could, in theory, be some other transition metal ion.

      (2) Seek evidence for the role of metal ions in the activity of this enzyme, for example, by treating with EDTA to remove metal ions, and then adding different metal ions, to correlate activity with a particular metal ion.

      We thank the reviewer for these valuable suggestions regarding the metal identity and its functional role in TDM activity. We also acknowledge the reviewer’s concerns regarding the catalytic mechanism. In response, we have substantially revised Scheme 1 and the associated Discussion text to focus on the observed interactions of the substrate (TMAO) and products (DMA and HCHO) within the Zn<sup>2+</sup>-containing active site, rather than proposing a detailed catalytic mechanism that is not fully supported by the current data.

      (1) We repeated the ICP–MS analysis using a wide-range full-scan survey to examine the presence of additional metal-associated isotopes beyond Zn and Fe. The corresponding experimental details have been added to the revised ICP–MS Methods section (Page 9). Full-scan ICP–MS profiling of purified TDM detected Zn as the predominant associated metal species (Author response image 1A). In contrast, signals corresponding to Fe and other transition metals were either undetectable or present only at trace levels comparable to, or lower than, those observed in the digested HNO<sub>3</sub> solution control. To further validate this observation, we performed targeted ICP–MS quantification for both Zn and Fe on the same purified samples. These measurements confirmed that Fe was below the detection threshold, whereas Zn was consistently detected at an approximate ratio of 0.5 Zn<sup>2+</sup> per protein monomer (Author response image 1B, C, Figure S5).

      The observed 0.5 Zn<sup>2+</sup>-to-protein stoichiometry is consistent with the previously discussed 2 full-length + 2 half-domain (2+2½) assembly. In this complex, only the intact core domains retain the complete metal-binding motif, whereas the truncated half-domains lack the Zn<sup>2+</sup>-binding region. Consequently, only two metal-binding sites are expected per assembled complex, in agreement with the ICP–MS measurements. We additionally note that Zn<sup>2+</sup> was not intentionally supplemented during purification. Based on the current cryo-EM and biochemical data, both metal-binding sites in the full-length subunits appear similarly occupied, with no evidence for asymmetric metal loading.

      Importantly, the previously published FEBS Journal model proposed an Fe<sup>2+</sup>-binding site based on metal analysis and homology modeling rather than direct experimental determination. The residues implicated in Fe<sup>2+</sup> binding are well resolved in our experimental maps and do not define a metal-coordination environment compatible with a second mononuclear metal-binding site. Consistent with the ICP–MS results, the experimental structure provides no evidence of a second metal-binding site. Nevertheless, the enzyme remains catalytically active under these conditions.

      (2) We attempted metal depletion experiments using EDTA treatment to evaluate the functional role of the bound metal ion. However, removal of metal ions resulted in rapid protein aggregation, preventing subsequent activity measurements. These observations suggest that the bound Zn<sup>2+</sup> ion plays an important role in maintaining the structural integrity and stability of the TDM complex. While metal reconstitution experiments would be informative, the aggregation observed following metal depletion precluded a meaningful assessment of alternative metal ions in the current study.

      Author response image 1.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      They use cultures of insulinoma MIN6 cells that form spheroids in a micro-patterned PEG-hydrogel to measure Ca <sup>2+</sup> oscillations in multiple cells simultaneously.

      Strengths:

      They demonstrate that insulinoma spheroids are formed in multi-well plates and that Ca <sup>2+</sup> imaging can be performed on them.

      Weaknesses:

      The type of equipment and multi-wells used for the experiments are very specialized to be used as a common tool. Insulinoma cells are tumoral cell lines that divide, unlike primary beta cells. Pancreatic islets are very different from this preparation, as they are highly heterogeneous, whereas these cells all respond equally. It would be good to see the same technique applied to primary cells.

      MIN6 cells do not respond to glucose and other secretagogues in the same way as primary cells, and they cycle, depending on the phase of the cycle to which they are exposed.

      The authors should report the number of cells per spheroid and the number of cells that are alive and dead.

      I would like to examine the effects of calcium channel blockers on calcium transients, and the use of pregnenolone is already described in the literature, but remains less well known.

      MIN6 cells secrete much insulin, because detecting the hormone in ELISAs requires too many primary cells. The authors should discuss the model in greater detail and compare it with primary beta cells. Also, they take 3 mM glucose as the basal concentration, which is low.

      We thank the reviewer for their valuable comments, which have helped us to significantly enhance the quality of our manuscript. We have carefully considered these comments and have revised the manuscript accordingly. A point-by-point rebuttal is provided below.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript contains numerous typos, including the combination of numbers and units without a space.

      All spacing inconsistencies between numerical values and unit symbols (e.g. mM, μM, µL, and Hz) have been corrected throughout the text and images. In addition, the following typographical and grammatical errors have been addressed:

      - Abstract: "the frequency of Ca <sup>2+</sup> oscillations correlate" → correlates

      - Figure 2A caption: "200uL" → 200 µL

      - Figure 5A caption: "glimepirde" → glimepiride

      - Figure 5 title: "KATP-antagonists induces" → induce

      - Figure 7D caption: "concentrations" → concentration

      - Figure 7 caption: "Tukey’s post-hoc testm" → Tukey’s post-hoc test

      - Results section header: "increases insulin secretions" → insulin secretion

      - Discussion: "we obtained an EC50 values 7.4 ± 0.4 mM" → "we obtained EC50 values of 7.4 ± 0.4 mM"

      - Discussion: "a EC50 value" → an EC50 value

      - Discussion: "concentration dependence profiles that matches" → match

      - Discussion: "PS-induced activation TRPM3" → activation of TRPM3

      - Conclusion: "found the Islets of Langerhans" → found in the Islets of Langerhans

      - Acknowledgements: "grant agreement No. 955643" → agreement No. 955643

      (2) I suggest showing the experiments in primary beta cells because of the many differences from insulinoma cells, as the most important result is a better culture technique for calcium imaging.

      We thank the reviewer for this valuable suggestion. We did indeed attempt to perform comparable experiments using the Cellartis<sup>®</sup> hiPS Beta Cell Media Kit (Takara, cat. no. Y10108) as a more physiologically relevant cell model. To minimize cellular stress during the transition, we adjusted our protocol by seeding the differentiated cells into the micropatterned plates 24 hours prior to imaging, thereby maintaining the recommended culture conditions for as long as feasible. However, during calcium imaging, the cells failed to respond to either the elevated glucose stimulus or the positive control, suggesting that the cells did not survive the transfer to our plate format with sufficient viability to mount a functional response.

      We acknowledge that the use of primary beta cells or hiPS-derived beta cells would strengthen the physiological relevance of the platform. Nevertheless, we would like to emphasize that the primary objective of this study is to demonstrate the feasibility of high-throughput calcium imaging in 3D cell culture using our optimized protocol: a proof-of-concept that is, by design, independent of the specific cell model employed. Adapting the protocol to accommodate more sensitive or terminally differentiated cell types is a meaningful avenue for future work, but falls outside the scope of the current manuscript. We have added a brief note to the Discussion section to explicitly acknowledge this limitation and to identify hiPS-derived beta cell compatibility as a priority for subsequent optimization.

      (3) These kinds of cultures in three dimensions are interesting, but it has been shown that it is even better to have the liquid flow, simulating blood flow; this can at least be discussed.

      We thank the reviewer for this insightful comment. We agree that the introduction of perfusion-based flow represents a meaningful improvement over static 3D culture systems. We have incorporated this point into the Discussion section, where we now explicitly acknowledge the potential benefits of dynamic culture conditions, including improved viability, insulin secretory function, and morphological integrity of 3D β-cell tissues, and identify the integration of perfusion-induced flow as a potentially meaningful path to explore for future development of the platform.

      Reviewer #2 (Public review):

      Summary:

      The study by Robben et al., show 3D beta-cell spheroid platform, a valuable tool allowing high-throughput monitoring of cytoplasmic Ca concentrations and insulin secretion, with Ca signals comparable to those recorded in primary islets. The authors demonstrate a solid method to culturing MIN6 cells in a 3D culture system, recording Ca signals in a high-throughput format and characterizing these Ca signals using pharmacological tools, including TRPM3 channel and K-ATP channel modulators. This highlights the utility of the 3D beta-cell spheroid for screening new ion channel modulators in beta-cells of the pancreas.

      Strengths:

      - The study shows that the MIN-6-based 3D beta-cell model is better to study Ca-signaling and insulin secretion compared to 2D culture of single MIN-6 cells.

      - The method allows imaging of Ca signaling in many spheroids in parallel followed by collecting medium to measure insulin release and correlate both effects.

      - The authors demonstrate that this system is suitable for screening new pharmacological modulators and used as an agonist of the ATP-sensitive potassium channel (diazoxide) and the agonist and antagonist of the TRPM3 channel.

      Weaknesses:

      - The study is based on only one cell line, the MIN6 insulinoma cells, which may not fully mimic the pancreatic beta-cells within the islet.

      - The authors show only spheroids cultured overnight. A long-term culture is missing to assess beta-cell viability long term function.

      - The authors tested their platform using only two compounds. Testing a larger compound library is necessary to make a clear conclusion about the suitability of the platform for high-throughput screening.

      We thank the reviewer for their valuable comments, which have helped us to significantly enhance the quality of our manuscript. We have carefully considered these comments and have revised the manuscript accordingly. A point-by-point rebuttal is provided below.

      Reviewer #2 (Recommendations for the authors):

      Major Points

      (1) In this study, only 2 pharmacological compounds (for TRPM3, and K channel) were tested. Testing a larger compound library would be necessary to fully demonstrate its suitability for high-throughput screening applications. If this is not possible at the moment, including data on additional pharmacological compounds, e.g., modulators of voltage-gated Ca channels, which are key regulators of Ca signaling in beta cells of the pancreas would strengthen the study. (e.g., use voltage gated Ca channel blocker such as verapamil or nimodipine).

      We thank the reviewer for this constructive suggestion. We would like to clarify that our pharmacological characterization was not limited to two compounds. In total, six compounds spanning two distinct ion channel targets were evaluated: the K-ATP channel modulators diazoxide, glimepiride, tolbutamide and nateglinide, and the TRPM3 modulators pregnenolone sulphate and isosakuranetin. This panel includes both agonists and antagonists across two mechanistically distinct targets, and we believe this is sufficient to demonstrate the platform's suitability for high-throughput compound screening in the context of a proof-of-concept study.

      We nonetheless agree with the reviewer that extending the compound panel to include modulators of additional ion channel classes, such as voltage-gated Ca <sup>2+</sup> channel blockers like verapamil or nimodipine, would further demonstrate the versatility of the platform. We have added a statement to the Discussion explicitly identifying this as a valuable direction for future work.

      (2) Testing another beta-cell line (e.g., INS-1 cells) would strengthen the manuscript.

      We thank the reviewer for this suggestion. We agree that validating the platform using an additional β-cell line, such as INS-1 cells, would further broaden the applicability of the approach. We did indeed attempt experiments with INS-1 cells; however, the results were inconclusive due to cell quality issues at the time of testing, and we were unable to generate reliable data suitable for inclusion in the manuscript.

      We would also like to emphasize that the primary aim of this study was to demonstrate the methodology of the high-throughput screening platform, rather than to provide a comprehensive cross-cell-line validation. In this context, the use of the well-established MIN6 β-cell line is sufficient to serve as a proof-of-principle demonstration of the platform's capabilities.

      Nonetheless, we consider a systematic evaluation of INS-1 cells on this platform as an important and natural next step and have included this explicitly as a future perspective in the Discussion.

      (3) I find the presentation of the results and analysis of Ca oscillation frequency and area under the curve excellent. Could you please provide more details on the analysis method used to quantify the frequency of glucose-induced Ca oscillation. If a custom script was used, sharing this information with the scientific community would be great.

      We thank the reviewer for this positive feedback. We confirm that the Ca <sup>2+</sup> oscillation analysis was performed using a custom Python script (v3.10.11), the key steps of which are described in the Data Analysis section of the experimental procedures.

      In line with our commitment to open and reproducible science, the script will be made publicly available upon acceptance of the manuscript, allowing the broader scientific community to apply, adapt, and build upon the analysis pipeline.

      (4) Please move the Supplementary Figure to the main Figure 3. This will allow a direct comparison between Ca signals in spheroids and in single cell (2D cultures) under identical conditions.

      We thank the reviewer for this suggestion. We have partially incorporated the supplementary figure into the main manuscript. The mean Ca <sup>2+</sup> response of 2D-cultured MIN6 cells at 20 mM glucose has been added as Figure 3D, enabling direct visual comparison with the spheroid data under identical stimulation conditions. The remaining panels of the supplementary figure — showing representative Ca <sup>2+</sup> traces across multiple glucose concentrations and the corresponding dose-response curves for peak frequency and area under the peaks in 2D monolayers — have been retained in the supplementary information, as their inclusion in the main figure would substantially increase its complexity. The figure legend has been updated accordingly.

      (5) A direct comparison of insulin secretion between 3D cultured spheroids and 2D cultures should also be shown.

      We thank the reviewer for this suggestion. We attempted to include a direct comparison of insulin secretion between 3D spheroids and 2D MIN6 monolayers; however, the 2D measurements proved unreliable for quantitative comparison. Insulin values in the 2D condition consistently exceeded the upper detection limit of the ELISA, and inter-well variability was too high to draw meaningful conclusions. We therefore chose not to include this comparison and instead present the Ca <sup>2+</sup> imaging data in Figure 3D as a functional readout enabling direct comparison between the two culture formats under identical stimulation conditions.

      (6) The authors should further discuss the remaining effects of pregnenolone sulphate on insulin secretion.

      We thank the reviewer for this comment. We would like to clarify that in our experimental setup, pregnenolone sulphate and isosakuranetin were applied simultaneously rather than sequentially. As a result, the incomplete inhibition of PS-induced Ca <sup>2+</sup> oscillations and insulin secretion observed in the presence of isosakuranetin may in part reflect a kinetic offset between the two compounds, whereby PS-induced TRPM3 activation and downstream signalling may have been initiated prior to the establishment of effective TRPM3 blockade by isosakuranetin. In addition, as noted in the Discussion, TRPM3 may not be the only molecular target of PS in these spheroids, and alternative signalling pathways may contribute to the residual insulin secretion observed in the presence of the antagonist. We have added a brief clarification to the Discussion to explicitly acknowledge the potential influence of this kinetic limitation on the interpretation of these results.

      (7) Is it possible to collect 3D cultured spheroids after each experiment to measure for example intracellular insulin content or protein levels by Western blot.

      Physical recovery of spheroids from the PEG hydrogel plates for downstream biochemical analysis, such as intracellular insulin content measurements or Western blot, would indeed be a meaningful addition to the platform's capabilities. We would like to note that the firm attachment of spheroids to the glass substrate, while essential for maintaining spheroid positioning during the extensive washing and liquid handling steps, does present a practical challenge for post-experimental recovery. Although we have successfully extracted spheroids of other cell types from comparable plate formats, reliable recovery of the MIN6 β-cell spheroids without compromising their structural integrity has not yet been achieved. We therefore identify the optimization of spheroid recovery as a valuable direction for future development of the platform.

      (8) Ca signals in response to the application of glucose appears more robust in spheroids compared to single MIN6 cells. What are the possible mechanisms underlying this difference. Whole RNA-seq experiments would be one approach to identify differentially expressed genes in 2D versus 3D culture (this is maybe a whole project by itself). An alternative is to look by RT-qPCR analysis for key β-cell markers and genes encoding ion channel and ion channel subunits.

      We thank the reviewer for this thoughtful comment. The more robust Ca <sup>2+</sup> signals observed in 3D spheroids compared to 2D monolayer cultures likely reflect several interconnected factors. First, and importantly, it should be noted that Ca <sup>2+</sup> measurements in 2D monolayer cultures typically represent an averaged signal across a large population of cells, which tends to obscure individual oscillatory events and reduce the apparent amplitude and regularity of Ca <sup>2+</sup> responses. In contrast, our 3D spheroid platform enables Ca <sup>2+</sup> measurements at the level of individual spheroids, allowing discrete oscillatory peaks to be resolved with much greater fidelity. Beyond this methodological distinction, the 3D architecture also promotes enhanced cell-to-cell communication, better recapitulation of in vivo β-cell coupling, and a more physiologically relevant microenvironment, all of which are likely to contribute to the improved oscillatory Ca <sup>2+</sup> dynamics observed.

      We agree with the reviewer that elucidating the transcriptional underpinnings of these differences, through whole RNA-seq or targeted RT-qPCR analysis of key β-cell markers and genes encoding ion channel subunits, would be highly informative and represents an elegant approach to understanding the molecular basis of the observed functional improvements. As the reviewer rightly acknowledges, however, such experiments constitute a substantial research effort in their own right. We have added a statement to the Discussion identifying this as a valuable direction for future investigation, alongside the other platform development priorities already outlined.

      (9) A more detailed discussion of the limitations of the platform and potential strategies to further improve this system would strengthen the manuscript.

      We thank the reviewer for this constructive suggestion. In response, we have expanded the Discussion to provide a more comprehensive overview of the current limitations of the platform and the strategies we envision for future development. Specifically, the revised Discussion now addresses the following points:

      First, the platform is currently optimized for the MIN6 insulinoma cell line, which differs from primary pancreatic beta cells in several important respects, including glucose sensitivity and secretory capacity. Initial attempts to adapt the protocol to hiPS-derived beta cells were unsuccessful, likely due to insufficient cellular viability following transfer to the micropatterned plate format. Optimizing the platform for use with primary beta cells or hiPS-derived beta cells is therefore identified as a priority for future development.

      Second, while the current study demonstrates proof-of-concept pharmacological characterization using six compounds across two mechanistically distinct ion channel targets, extending the compound panel to include modulators of additional ion channel classes, such as voltage-gated Ca <sup>2+</sup> channel blockers, as well as validation using alternative insulinoma cell lines such as INS-1, would further demonstrate the versatility and generalizability of the platform.

      Third, the introduction of perfusion-induced flow, which has been shown to improve viability, insulin secretory function, and morphological integrity of 3D beta-cell tissues under dynamic culture conditions, is identified as an additional avenue for optimization.

      Fourth, while the current platform operates in a 96-well plate format, which already provides a substantial throughput of up to 1824 individual spheroid measurements per plate, adaptation to higher density plate formats such as 384-well or 1536-well plates would be a necessary step towards true high-throughput screening compatible with industrial drug discovery pipelines. Miniaturization of the hydrogel design and adaptation of the molding procedure to accommodate these formats therefore represents an important direction for future development.

      Minor points:

      (10) Please clarify the glucose concentration at which the MIN6 cells the cultured 3D beta cell spheroids were maintained overnight prior to the experiments.

      Prior to the experiments, both 2D MIN6 cells and 3D MIN6 β-cell spheroids were maintained overnight in standard high-glucose DMEM (25 mM glucose), consistent with widely used MIN6 culture protocols. This information has been added to the Methods section.

      (11) In Figure 3A, is the presented Ca trace derived from one single spheroid? Showing representative Ca traces (e.g., 5 i traces per condition) would show the reproducibility of the recordings.

      We thank the reviewer for this suggestion. The trace shown in Figure 3A is indeed derived from a single representative spheroid. To address the concern regarding reproducibility, we refer the reviewer to the updated Figure 3C, which now displays the individual Ca <sup>2+</sup> response traces of all spheroids within a single well in response to 20 mM glucose stimulation. Grey lines represent individual spheroids, while the black line denotes the mean response. This panel illustrates not only the reproducibility of the oscillatory response across the imaged population but also highlights an important consequence of inter-spheroid variability: because individual spheroids oscillate asynchronously, their peaks cancel out when averaged, resulting in a mean trace that appears relatively flat and lacks the oscillatory features visible in individual recordings. We note that this same cancellation effect likely underlies the comparably flat mean response observed for 2D-cultured MIN6 cells in Figure 3D. Additionally, the difference in the initial response profile between 2D and 3D cultures may in part reflect the geometry of the hydrogel microenvironment. In 2D cultures, the glucose stimulus equilibrates rapidly and near-uniformly across the culture plane, potentially driving a more synchronized initial response and the early peak visible in Figure 3D. In contrast, the PEG hydrogel surrounding the 3D spheroids may act as a diffusion barrier, causing the stimulus to reach individual spheroids with variable delay and thereby further desynchronizing response onsets across the population. The figure legend has been updated accordingly.

      (12) In Figure 3A, the potassium application bar looks a bit shifted. Please check and correct if necessary.

      We thank the reviewer for carefully examining the figure. The apparent shift in the potassium application bar is not an error. In these experiments, glucose was added after an initial 10-minute baseline period, and the observed delay in the Ca <sup>2+</sup> response reflects two contributing factors. First, the glucose solution was pipetted at the top of the well, and diffusion to the level of the spheroids introduces a short lag before the stimulus reaches the cells. Second, the spheroid shown in Figure 3A was located at the outer edge of the well, where mixing is slower and the stimulus arrives with additional delay compared to centrally positioned spheroids. Together, these factors account for the offset between the start of the glucose application bar and the onset of the visible Ca <sup>2+</sup> response.

      (13) Is the system also compatible with the ratiometric Ca imaging dye Fura-2?

      The platform is indeed compatible with ratiometric Ca <sup>2+</sup> imaging using Fura-2. The glass-bottom plate format is a prerequisite for Fura-2 imaging due to the requirement for UV excitation at 340/380 nm, and the transparency of the PEG-based hydrogel ensures that the optical properties of the platform are fully compatible with this approach. Furthermore, the µCELL FDSS fluorescence plate imager used in this study supports dual-excitation ratiometric imaging, making it instrumentally compatible with Fura-2 without any additional hardware modifications.

      (14) In Figure 4 legend: 100 µm diazoxide should be corrected to 100 µM diazoxide.

      This has been addressed in the revised manuscript

      (15) Please comment on the cost of spheroid generation compared with conventional 2D MIN6 cultures.

      We thank the reviewer for this relevant question. The cost of spheroid generation using our platform is largely comparable to conventional 2D MIN6 cultures, with the primary additional expense being the specialized micropatterned PEG-based hydrogel plates. All other aspects of the workflow, including cell culture reagents, imaging consumables, and instrumentation, remain identical. It should be noted that providing a precise cost comparison is difficult, as the price of the hydrogel plates represents the dominant variable cost and is subject to change depending on production scale and supplier agreements. Nevertheless, we consider the platform to be cost-effective relative to alternative 3D culture systems, which often require more complex fabrication procedures or proprietary consumables.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study is to develop high-throughput screening assays utilizing homogeneous 3D cell cultures that more accurately replicate the intricate architecture and cellular communication found in tissues. The authors have chosen pancreatic islet β-cells as a model system to evaluate agents that modulate insulin release, which is particularly relevant given the increasing prevalence of diabetes mellitus-a significant global health concern. Moreover, the incorporation of human-based 3D spheroids, organoids, or organ-on-chip technologies into drug discovery protocols is essential for enhancing clinical translation, as candidate compounds identified using animal models have often demonstrated limited success in clinical settings.

      Strengths:

      This study was thoughtfully planned and skillfully carried out. The use of micropatterned hydrogels to observe 19 spheroids at once is an ingenious aspect, which has been effectively validated with Ca microfluorography. Overall, I found this investigation to be exceptionally well-executed and free from notable flaws, as the results clearly back up the conclusions. Additionally, the developed method achieved the proposed aims, providing a high-throughput format with 3D cultures. I believe this study deserves publication.

      Weaknesses:

      For an HTS assay, authors should incorporate the Z-factor.

      We thank the reviewer for their valuable comment, which we have directly addressed in the revised version, as outlined below.

      Reviewer #3 (Recommendations for the authors):

      (1) The study is very well performed, but to support the claim of suitability for HTS, the Z-factor of the method should be reported.

      We thank the reviewer for this helpful suggestion. In response, we have now included the Z′-factor in the manuscript to explicitly quantify assay performance and robustness. Specifically, the Z′-factor (0.6) has been added and discussed in the Results section, incorporated into the Discussion to contextualize assay suitability for high-throughput screening, and included in the Methods under “Statistical Analysis,” where its calculation is described.

    1. Author response:

      We thank the Reviewing Editor and the reviewers for their highly constructive feedback and their positive assessment of our study. We plan to submit a revised manuscript that addresses these critiques through textual revisions, contextualization of our data, and explicit discussion of the study's limitations.

      To address the comments from Reviewers 1 and 3 regarding the sensing mechanism and cue specificity, we agree that the precise sensor for diacetyl remains an open question. While we cannot specifically pinpoint the mechanism with our current data, we hypothesize that this response relies on a non-canonical sensing mechanism, given our multiple negative results for canonical diacetyl receptors (odr-10, sri-14), signalling pathways, and cilia-defective mutants (daf-19; daf-12). We will revise the text to suggest the mechanism could be cilia-independent or cell-autonomous. Additionally, we will explicitly frame the investigation of AWA-ablated worms, diacetyl derivatives such as acetoin, and additional volatile food cues as important future directions to establish the generalizability of this response.

      Regarding the contrasting survival phenotypes observed by Park et al.24 raised by Reviewer 1, our manuscript currently discusses how chronic odour exposure represses the longevity benefits of dietary restriction48, hypothesizing that this may stem from age-dependent olfactory decline49 and the confounding variable of olfactory learning8. To make this connection clearer, we will explicitly cite Park et al. in this section to directly link their findings with our hypothesis that repeated odour exposure without a nutritional reward extinguishes its efficacy as an anticipatory cue.

      In response to Reviewer 2’s feedback on our metabolic profiling, we will ensure it is clear in the text that we highlighted the lipid species most prominently affected by diacetyl, explicitly noting that triglycerides were not significantly altered. We will also better direct readers to Supplementary Table 1, which contains the comprehensive lists of all metabolites and lipids detected in our metabolomics and lipidomics, and quantifies how they are affected by the treatments in our study. As noted in our manuscript, both DHAP and Gro3P were successfully detected in our metabolomics platform but were not significantly altered by diacetyl exposure. Glycerol, however, was measured via a commercial enzymatic assay because it was not detected by our specific LC-MS platform. We will acknowledge that independently measuring the effects on cellular redox states and utilizing GC-MS for broader metabolic profiling are valuable future directions.

      Finally, to address Reviewer 3’s queries regarding experimental readouts and nhr-49, we will clarify our rationale for using thrashing as our primary readout for hyperosmotic stress. Because diacetyl exposure triggers an acute induction of the DHAP-glycerol shunt, we specifically chose an acute behavioural readout to match this rapid timeline. Regarding nhr-49, we will add a new point to the discussion proposing that the distinct phenotypes may come down to expression thresholds. We hypothesize that while nhr-49 RNAi partially reduces gpdh-1 expression, this residual level of expression might still be sufficient to allow development and survival during sustained hyperosmotic stress."

      References:

      (8) Choi, J. I., Yoon, K., Kalichamy, S. S., Yoon, S.-S. & Lee, J. I. A natural odor attraction between lactic acid bacteria and the nematode Caenorhabditis elegans. ISME J. 10, 558–567 (2016).

      (24) Park, S. et al. Diacetyl odor shortens longevity conferred by food deprivation in C. elegans via downregulation of DAF‐16/FOXO. Aging Cell 20, e13300 (2021).

      (48) Zhang, B., Jun, H., Wu, J., Liu, J. & Xu, X. Z. S. Olfactory perception of food abundance regulates dietary restriction-mediated longevity via a brain-to-gut signal. Nat. Aging 1, 255–268 (2021).

      (49) Suryawinata, N. et al. Dietary E. coli promotes age-dependent chemotaxis decline in C. elegans. Sci. Rep. 14, 5529 (2024).

    1. Author response:

      We thank the editors for the eLife Assessment and the reviewers for their thorough and constructive evaluation of our manuscript.

      We are glad that the evidence for reduced metacognitive sensitivity in relation to the compulsive hypersensitivity dimension was considered solid. In hindsight, we agree that the evidence for the abstraction findings is currently incomplete, and we plan to address this with additional analyses as outlined below (to clarify the robustness of the abstraction metric and its association with symptom dimensions).

      Below, we outline how we plan to address the main points raised by the reviewers, grouped thematically, given the overlap across the three reviews.

      A full, detailed point-by-point response accompanied by the corresponding analyses will follow in our formal revision response.

      (1) Exclusion rate and lack of relevant sensitivity analysis

      In the revision, we will report a detailed comparison of included versus excluded participants on demographic and symptom variables, and we will conduct the originally preregistered sensitivity analysis to estimate abstraction and metacognition metrics, including the excluded sample, and compare these metrics with psychopathological variables, rather than omitting this analysis.

      (2) Abstraction metric and its deviation from the preregistered definition

      We recognise that our primary measure of abstraction differs from the preregistered metric (the proportion of blocks better fit by the Abstract RL model), and that this change appears to affect the pattern of results. In the revision, we will report split-half reliability for the abstraction metric, present the bootstrapped analysis of the preregistered metric, and provide a clearer justification for the switch, making the source of the discrepancy (metric versus inference method) transparent.

      (3) Generalisability of the transdiagnostic factor structure

      We acknowledge that our claim of consistency between samples of the transdiagnostic dimensions requires more support and more careful framing. In the revision, we will provide full methodological detail on both factor analyses (extraction method, criteria for the number of factors, etc.), and we will moderate our interpretation, especially in the discussion, to reflect the actual strength of correspondence rather than describing the structure as broadly consistent.

      (4) Parameter and model recovery

      We will extend our recovery analyses to include a model recovery/confusion analysis between the Abstract RL and Feature RL models, and we will investigate the source of the relatively low learning-rate recovery (e.g., by testing recovery stratified by block length and without injected noise), reporting these results in the revision.

      (5) Documentation of the attention-check procedure

      We will provide all details on the attention-check items and criteria used to identify inattentive participants, including a reference where applicable, and we will standardise terminology (e.g., infrequency vs. inattention items) throughout the manuscript and supplementary information.

      (6) Other methodological and presentational clarifications

      We will review the manuscript, methods, and supplementary materials to resolve the remaining methodological ambiguities and presentational issues raised by the reviewers. This includes, among other points, clarifying whether the reported regression coefficients are standardised and correcting labelling inconsistencies, missing references/DOIs, and other minor textual issues throughout.

    1. Author response:

      We thank the editors and all three reviewers for their careful and constructive evaluation of our manuscript. We recognize that a single concern, the possibility that cells classified as CD56<sup>dim</sup>CD16<sup>dim</sup> after co-culture represent activated CD56<sup>dim</sup>CD16<sup>bright</sup> cells that have shed CD16 rather than a pre-existing subset, underlies the majority of the comments. We therefore address this concern first, in a central response, and then respond to each reviewer and editor comment in turn. Where a comment relates to this shared concern, we point to the central response rather than repeating the argument.

      The analytical and presentational revisions described below are complete: the statistical analyses have been re-run with corrections for multiple comparisons, the Discussion has been rewritten, the figures and supplemental tables have been renumbered and corrected, and existing data on pre-stimulation receptor expression and NKp30 have been incorporated. These will appear in the revised manuscript. The new experiments described will be completed within approximately six to eight weeks and provided with the revised manuscript.

      Central are CD56<sup>dim</sup>CD16<sup>dim</sup> cells a pre-existing subset, or activated CD56<sup>dim</sup>CD16<sup>bright</sup> cells that have shed CD16?

      We agree with the premise that CD16 is rapidly shed by ADAM17 upon activation, and that classifying subsets by post-assay CD16 expression alone cannot, on its own, distinguish a pre-existing subset from activation-induced conversion. For this reason, our conclusion does not rest on post-assay classification. The evidence below, from experiments already in the manuscript, argues against activation-induced conversion, and we will strengthen it with the expanded sorted-subset experiments described at the end.

      Cells sorted before target exposure establish the advantage independently of any during-assay shedding (Fig. 2, unchanged in the revised manuscript).

      The most direct evidence comes from subsets purified before the assay. In Fig. 2, NK cells were sorted into CD56<sup>dim</sup>CD16<sup>dim</sup> and CD56<sup>dim</sup>CD16<sup>bright</sup> populations before any exposure to target cells, and their cytolytic function was measured as specific lysis of autologous HIV-infected T cells. Purified CD16<sup>dim</sup> cells lysed infected targets approximately twice as efficiently as purified CD16<sup>bright</sup> cells across the effector-to-target range, reaching 77.58% versus 39.18% at 1:1. The difference was significant at 1:4 (p = 0.0008), 1:2 and 1:1 (both p < 0.0001); at the lowest ratio tested, 1:8, specific lysis was low in both subsets and the difference did not reach significance (p = 0.2121; Supplemental Table 4). The additional donors described below will allow this comparison to be made across a larger data set, including at the lowest ratios where specific lysis is low in both subsets.

      Because the subsets are defined by sorting before target contact, and because the readout is direct target lysis rather than post-assay CD16 gating, this advantage cannot arise from activation-induced CD16 shedding during the assay. Public reviewer 2 and the peer reviewer both identified this experiment as the strongest evidence in the manuscript. Its interpretation is secure; what it requires is additional donors for statistical robustness, which we provide in the planned expansion below.

      The two subsets respond to ADAM17 inhibition in opposite directions (Figs. 7 and 9, now Figures 6 and 8).

      If CD56<sup>dim</sup>CD16<sup>dim</sup> cells were simply CD56<sup>dim</sup>CD16<sup>bright</sup> cells that had shed CD16, the two would be one population sampled at different points along a shedding continuum, and inhibiting ADAM17 would move them in the same direction. Instead, ADAM17 inhibition moves them in opposite directions. In the antibody-dependent degranulation assay (Fig. 7B, now Figure 6B), ADAM17 inhibition increased CD56<sup>dim</sup>CD16<sup>bright</sup> degranulation but decreased CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation across VRC01 concentrations. The same opposition is seen when serial degranulation is resolved by the number of degranulation events per cell (Fig. 9, now Figure 8): ADAM17 inhibition increased multiple degranulation events in CD56<sup>dim</sup>CD16<sup>bright</sup> cells, with cells undergoing three events rising from 0.79% to 5.12%, while in CD56<sup>dim</sup>CD16<sup>dim</sup> cells it reduced them, with three events falling from 15.16% to 2.24% and the non-degranulating fraction rising from 61.56% to 92.13%.

      A single population would not be expected to respond to the same perturbation in opposite directions, and these observations are difficult to reconcile with the CD16<sup>dim</sup> cells being activated CD16<sup>bright</sup> cells; rather, they point to two functionally distinct subsets with opposite dependence on ADAM17 activity. This is consistent with our model, in which CD16<sup>dim</sup> cells use ADAM17-mediated shedding to detach and serially re-engage, whereas CD16<sup>bright</sup> cells are hindered by the loss of CD16.

      The CD16<sup>dim</sup> degranulation advantage is driven by NKG2D through a mechanism separable from ADAM17 (Fig. 8, now Figure 7).

      The change in subset frequency on exposure to VRC01-treated infected cells (Fig. 8B, now Figure 7B) is abolished by anti-NKG2D even though VRC01 and ADAM17 remain present, indicating that this frequency shift is driven by NKG2D-dependent activation rather than by antibody-CD16 engagement alone.

      The degranulation data in Fig. 8A (now Figure 7A) show that the two perturbations act differently in the two subsets. In CD56<sup>dim</sup>CD16<sup>dim</sup> cells, both reduce the response, and the combination reduces it further than either alone: from 12.60% under vehicle to 5.24% with anti-NKG2D (p < 0.0001), 5.05% with ADAM17 inhibition (p < 0.0001), and 2.03% with both (p < 0.0001 versus vehicle; p < 0.0001 versus anti-NKG2D alone; p = 0.0001 versus ADAM17 inhibition alone). If anti-NKG2D acted only by removing the activation trigger for ADAM17, that is, if NKG2D and ADAM17 lay on a single linear pathway, blocking the pathway at two points would not be expected to add to the effect of either alone. The further reduction therefore indicates that NKG2D and ADAM17 contribute through separable mechanisms.

      In CD56<sup>dim</sup>CD16<sup>bright</sup> cells the two perturbations act in opposite directions. Anti-NKG2D reduced degranulation from 2.58% to 0.92% (p = 0.0263), whereas ADAM17 inhibition increased it to 3.99% (p = 0.0679). NKG2D therefore supports the response of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while ADAM17 activity constrains it, the reverse of the pattern in CD56<sup>dim</sup>CD16<sup>dim</sup> cells, where ADAM17 activity is required. Two populations differing only in the extent to which they have shed CD16 would not be expected to respond to the same two perturbations in opposite ways.

      Supporting evidence: pre-sorted subsets are stable and differ before stimulation.

      Two further observations support a pre-existing subset. First, we have directly tracked the fate of each subset sorted before target exposure. NK cells were sorted into CD16<sup>bright</sup> and CD16<sup>dim</sup> subsets, exposed to HIV-infected cells for one hour, and reanalyzed for CD16 expression. One hour is the point at which we observe the highest frequency of degranulating cells in both subsets (Figure 3—Figure Supplement 1A in the revised manuscript), and therefore the point at which activation-induced shedding would be most likely to be detected. Sorted CD16<sup>bright</sup> cells remained predominantly CD16<sup>bright</sup> (approximately 62%); of those that lost CD16, most became CD16<sup>negative</sup> (approximately 35%) rather than CD16<sup>dim</sup> (approximately 4%). Sorted CD16<sup>dim</sup> cells likewise shifted predominantly to a CD16<sup>negative</sup> phenotype (approximately 75%). Activation-induced CD16 shedding therefore directs cells of both subsets toward the CD16<sup>negative</sup> gate rather than generating the CD16<sup>dim</sup> population from CD16<sup>bright</sup> cells. These data are presented in Author response image 1.

      Author response image 1.

      Phenotype of sorted CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup> NK cells after exposure to HIV-infected T-cells. NK cells were sorted into CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup> subsets, exposed to purified autologous productively HIV-1<sup>SHM-1</sup>-infected T cells for 1 hour at a 1:1 effector-to-target cell ratio, and reanalyzed for CD16 expression. Bars show the percentage of each sorted NK cell population (CD56<sup>dim</sup>CD16<sup>bright</sup> and CD56<sup>dim</sup>CD16<sup>dim</sup>) falling into the CD16<sup>bright</sup>, CD16<sup>dim</sup>, and CD16<sup>negative</sup> gates after exposure, as the mean ± standard deviation (SD) of three replicates. One hour is when the highest frequency of degranulating cells is observed in both subsets.

      Second, the subsets differ before any stimulation. CD56<sup>dim</sup>CD16<sup>dim</sup> cells express higher NKG2D than CD56<sup>dim</sup>CD16<sup>bright</sup> cells before any target-cell contact. In the no-target condition, NKG2D was 1182 gMFI higher on CD56<sup>dim</sup>CD16<sup>dim</sup> cells (p < 0.0001), a difference of approximately 1.6-fold; across all conditions tested the subset means were 2521.5 versus 1772.2 gMFI, or 1.42-fold (two-way ANOVA: subset F(1, 20) = 1897, p < 0.0001, 70.97% of the total variation; Fig. 6, now Figure 5; Supplemental Table 18). This difference is specific to NKG2D: NKp46, measured on the same cells in the same wells, did not differ between the subsets in the no-target condition (mean difference 45.67 gMFI, p = 0.0668), and was higher on CD56<sup>dim</sup>CD16<sup>bright</sup> cells when targets were present. The subsets therefore differ in NKG2D density before activation, and not in activating receptor density generally.

      Planned strengthening.

      To place this beyond doubt, we will expand the sorted-subset experiments, performing the specific-lysis assay (Fig. 2) and the antibody-dependent degranulation assay on subsets purified before target exposure across additional donors, together with uninfected-target controls. If cell yields from the sort permit, we will also perform the serial degranulation assay on sorted subsets; because the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56<sup>dim</sup> NK cells and the serial degranulation assay requires four sequential labelling and washing steps, we cannot commit to this in advance of the sort. These experiments require sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      We note that performing every functional assay in this study on sorted subsets is not feasible within the scope of this revision. Sorting the CD56<sup>dim</sup>CD16<sup>dim</sup> subset, which constitutes fewer than 5% of CD56<sup>dim</sup> NK cells, from a sufficient number of donors to repeat the full panel of assays would require resources beyond those currently available to us. We have therefore prioritized the specific-lysis and antibody-dependent degranulation assays, which bear most directly on the concern raised by the reviewers and the editor, and will extend the approach to the remaining assays as resources allow.

      Public Reviews:

      Reviewer #1 (Public review):

      Overall organization

      Overall, the manuscript includes many data in nine figures plus supplemental figures, and would benefit from some focusing of the results.

      We agree, and we have reduced the main figures from nine to eight. The NKG2D ligand histograms, previously Figure 5A, have been removed for the reason given in our response to the peer reviewer, Recommendation 4. The degranulation against wild-type and ΔVpr-infected targets, previously Figure 5B, and the degranulation and killing frequency measurements in mixed populations, previously Figure 3, have been moved to the supplementary material. The inhibitory receptor analyses have been reduced from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detail retained in the supplemental tables. Each Results section now opens with a statement of the principal finding before the supporting data, and the Discussion synthesizes what the findings mean rather than restating them.

      Figure 1 and Figure 1—figure supplement 2

      The observation that CD16dim NK cells responded more strongly by degranulation to K562 cells and HIV-1-infected cells could be due to the shedding of CD16 following activation. In other words, more strongly activated NK cells express higher levels of CD107 but also shed CD16, resulting in higher CD107 expression in CD16low NK cells. The authors should investigate this, for example by performing the degranulation assays shown in Figure 1 in the presence and absence of an ADAM17 inhibitor.

      We thank the reviewer for raising this important point, which we recognize is shared by all three reviewers and the editors, and which we address in full in the central response above. We note for clarity that the K562 data are not in Figure 1; they are presented in Figure 1—figure supplement 2, both in the reviewed preprint and in the revised manuscript. They were included to reproduce a previously established finding (Amand et al., Front Immunol 2017;8:699), and were not intended as a central experimental claim; the mechanistic focus of this study is the response to HIV-infected cells. Notably, that same study addressed the question by sorting the CD56<sup>dim</sup> subsets before stimulation rather than gating after, and our study applies the same approach and extends it to the HIV-infected setting. Four lines of evidence argue against the interpretation that CD16<sup>dim</sup> degranulation reflects activation-induced CD16 shedding of CD16<sup>bright</sup> cells:

      (1) In cells sorted before target exposure, purified CD16<sup>dim</sup> cells lyse HIV-infected targets approximately twice as efficiently as purified CD16<sup>bright</sup> cells across the effector-to-target range, significantly so at 1:4 and above (Fig. 2; Supplemental Table 4); because the subsets are defined before any activation and the readout is direct target lysis rather than post-assay CD16 gating, this advantage cannot arise from shedding during the assay.

      (2) The two subsets respond to ADAM17 inhibition in opposite directions, both in the magnitude of degranulation (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8): ADAM17 inhibition increased degranulation in CD16<sup>bright</sup> cells but decreased it in CD16<sup>dim</sup> cells.

      (3) Blocking NKG2D together with ADAM17 reduced CD16<sup>dim</sup> degranulation below either treatment alone (Fig. 8A, now Figure 7A), indicating that the CD16<sup>dim</sup> advantage is driven by NKG2D through a mechanism separable from ADAM17-mediated shedding.

      (4) The subsets also differ before any stimulation, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells in the no-target condition (p < 0.0001) and no corresponding difference in NKp46 under the same condition (Fig. 6, now Figure 5).

      We will further strengthen these findings by expanding the sorted-subset experiments across additional donors, as described in the central response.

      Figure 2

      The authors sorted CD16dim and bright NK cells for these experiments and observed higher lysis of HIV-1-infected CD4+ T cells. Important controls should be included in these experiments - how strong was the lysis of HIV-1-uninfected CD4+ T cells by these different NK cell subsets? It also appears that the results shown were derived using NK cells from one donor, and "representative of two independent sort experiments performed with separate donors, each yielding similar results". Why are the authors now showing the respective data? One or two experiments appear too few to come to these conclusions. To support the broad conclusions drawn by the reviewers, the experiments should be performed in a larger number of individuals.

      We thank the reviewer for these constructive points.

      (1) Uninfected-target control. We agree this is an important control and will include lysis of uninfected autologous CD4 T cells by the sorted CD16<sup>dim</sup> and CD16<sup>bright</sup> subsets in Figure 2, confirming that the observed lysis is specific to HIV-infected targets. We note that the corresponding CD107a degranulation controls against uninfected targets are presented in Figure 1—Figure Supplement 4C of the revised manuscript.

      (2) Number of donors and presentation of data. We agree that the conclusions require more than the representative donor shown. As described in the central response, we will expand these sorted-subset experiments to a larger number of individuals and will present the data from all donors rather than a single representative experiment. Because they require cell sorting and primary-cell work, these experiments will be completed within approximately six to eight weeks and provided with the revised manuscript.

      Figures 3 and 4

      It appears that experiments were performed again using bulk NK cell populations, and superior degranulation and killing frequencies by CD16dim NK cells might reflect different levels of activation again, as described above for Figure 1. The same applies to Figure 4 - lower degranulation events in CD16bright NK cells are consistent with lower activation of these cells, resulting in less CD16 downregulation. Also, it is not clear to the reviewer why CD107a expression and killing frequencies decrease with higher effector-to-target ratios (Figure 3).

      (1) Activation-induced shedding in bulk experiments (Figs. 3 and 4, now Figure 2—figure supplement 1 and Figure 3). We agree that these figures use bulk NK cell populations gated by CD16, and we address the underlying shedding concern in full in the central response. The concern that the lower serial degranulation of CD16<sup>bright</sup> cells in the direct-killing assay simply reflects lower activation and therefore less shedding is addressed directly by our ADAM17-inhibition data. At 0 µg/mL VRC01, that is, in the absence of antibody, ADAM17 inhibition already affects the two subsets differently rather than in the same direction (Figs. 7B and 7C, now Figures 6B and 6C), as would be expected if they were one population differing only in activation level. This differential response is also seen across the antibody-dependent conditions in Figs. 7B and 9 (now Figures 6B and 8). As described in the central response, we will additionally repeat the specific-lysis and antibody-dependent degranulation measurements on subsets purified before target exposure across additional donors, and the serial degranulation assay as well if cell yields from the sort permit.

      (2) Decrease in CD107a and killing frequency at higher effector-to-target ratios (Fig. 3, now Figure 2—figure supplement 1). This reflects the nature of the readout. CD107a mobilization is measured per effector cell, as the percentage of NK cells that degranulate, and is therefore maximized when targets are in excess. At low effector-to-target ratios, nearly every NK cell can encounter and engage a target, yielding a high percentage of CD107a-positive cells; at high ratios, targets become limiting, so a large fraction of NK cells never contact a target and remain unstimulated, and the rapid destruction of the limited target pool further reduces the stimulus available to the remaining cells. This lowers the measured per-effector degranulation frequency even as the absolute number of targets killed is maintained, and it is distinct from a lysis assay, which measures the fate of the target population and accordingly rises with increasing effector-to-target ratio. The same per-effector readout behavior applies to the degranulation data shown in Figs. 5C and 6C (now Figures 4A and 4B, and Figure 5C). This explanation has been added to the revised Discussion.

      Pages 19-25

      It would be helpful if the authors could provide some conclusions regarding their findings - it is very difficult for the reader to follow the many reported frequencies and p-values. What does this actually mean? Overall, the results appear to follow prior observations that licensed (KIR3DL+) NK cells respond more strongly than unlicensed (KIR3DL1neg) NK cells. The consistent observation within these different subanalyses that CD16dim NK cells degranulate more than CD16bright NK cells is probably the result of activation-induced CD16 downregulation in these assays, as mentioned above. Providing two-way ANOVA analysis results for these very many observations would furthermore require, in the opinion of the reviewer, adjustments for multiple comparisons.

      We thank the reviewer, and we have addressed this in three ways.

      (1) Readability. We agree that these sections were difficult to follow as presented. We have rewritten them, opening each with a statement of the principal finding before the supporting statistics, and reducing the inhibitory receptor section from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detail retained in the supplemental tables. The Discussion now synthesizes what the findings mean.

      (2) CD16<sup>dim</sup> degranulation in these subanalyses. The consistent observation that CD16<sup>dim</sup> cells degranulate more than CD16<sup>bright</sup> cells across these subanalyses is addressed in full in the central response, where several lines of evidence, including subsets sorted before target exposure (Fig. 2) and the opposite responses of the two subsets to ADAM17 inhibition (Figs. 7B and 9, now Figures 6B and 8), argue against activation-induced CD16 downregulation as the explanation.

      (3) Multiple comparisons. We agree, and we have re-analyzed these comparisons, applying the post-hoc test matched to each comparison structure: Dunnett's where every subset is compared against a single designated subset, Tukey's where all pairwise comparisons are of interest, and Šidák’s where a prespecified subset of comparisons is of interest. Adjusted p-values are reported throughout, and the design and post-hoc test used for each figure and panel are given in a new supplemental table. The streamlining described above has also reduced the number of comparisons reported in the main text.

      Figures 5 and 6

      These figures demonstrate that NK cell-mediated activation by HIV-1-infected cells depends on NKG2D ligands and can be inhibited by blocking this interaction - this is consistent with data presented by the Barker group and others previously, and does not provide new information.

      We agree that the dependence of NK-cell recognition of HIV-infected cells on NKG2D and its ligands is established, including in our own earlier work (Ward et al., PLoS Pathog 2009;5(10):e1000613) and by others, and we do not present that dependence as a novel finding.

      On review, the histograms in Figure 5A were reproduced from that earlier study and should not have been included without attribution. We have removed that panel and cited the original finding in its place. Figures 5B and 5C are both new results from this study, and both are retained. Figure 5B, which shows that NK cells degranulate in response to wild-type HIV-infected targets but not to ΔVpr-infected or uninfected targets, establishing that the degranulation response in this system depends on Vpr, becomes Figure 4—figure supplement 1 in the revised manuscript. Figure 5C, the NKG2D blockade experiment, becomes Figure 4A and 4B.

      We would also distinguish Figure 6 (now Figure 5), which we consider a substantive finding rather than a restatement of the established NKG2D-ligand dependence. That figure shows that NKG2D expression differs at the level of the individual subsets, and that the difference is present before stimulation and is specific to NKG2D. In the no-target condition, NKG2D was 1182 gMFI higher on CD56<sup>dim</sup>CD16<sup>dim</sup> cells (p < 0.0001), approximately 1.6-fold, while NKp46 measured on the same cells in the same wells did not differ (mean difference 45.67 gMFI, p = 0.0668). This provides a candidate mechanism for the superior effector function of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset, in addition to their serial-degranulation capacity, and it bears directly on the central question of whether the two subsets differ intrinsically rather than as a consequence of activation. The revised text presents the established NKG2D-ligand dependence as context while making the subset-level NKG2D difference, and its mechanistic significance, more prominent.

      Figure 7

      The authors extended their functional analyses of NK cells to ADCC function. It is very well established that CD16 is downregulated in the context of ADCC following activation of NK cells. Consistent with this, higher degranulation is observed by CD16dim NK cells.

      We agree that CD16 is downregulated during antibody-dependent responses, and this is precisely why we included the ADAM17-inhibition experiments within Figure 7 (now Figure 6), to determine whether the higher degranulation of CD56<sup>dim</sup>CD16<sup>dim</sup> cells is a consequence of that shedding or a property of a distinct subset. As detailed in the central response, these experiments argue against the shedding interpretation. In Fig. 7B (now Figure 6B), inhibiting ADAM17 affects the two subsets in opposite directions: it increases the degranulation of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while decreasing that of CD56<sup>dim</sup>CD16<sup>dim</sup> cells. If the CD16<sup>dim</sup> cells were simply CD16<sup>bright</sup> cells that had shed CD16, blocking shedding would be expected to move the two in the same direction; the opposite responses instead indicate two distinct populations with opposite functional dependence on ADAM17 activity. The same opposition is seen when serial degranulation is resolved by the number of events per cell (Fig. 9, now Figure 8). Thus, while CD16 downregulation during antibody-dependent responses is well established, these data indicate that the superior response of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset is not explained by it. This interpretation is now explicit in the revised Discussion.

      ADAM17-inhibition data (final figures)

      These data are of interest, but should be presented in a more structured way. First of all, does the addition of ADAM17 inhibitors change the overall proportion of CD16bright and dim NK cells following activation, independent of whether these cells degranulate or not? Overall, the proportion of CD16dim NK cells that degranulate appears to be reduced in the presence of the ADAM inhibitor, which is consistent with reduced CD16 shedding and maintenance of CD16 expression on activated NK cells - and this is supported by the increase in CD107a-positive NK cells that express CD16 (Figure 8a). Overall, the differences between CD16bright and dim NK cells in their level of activation appear to disappear in the presence of an ADAM17 inhibitor, based on the data shown in Figure 8b, suggesting that CD16 downregulation is occurring in response to activation of NK cells as a consequence of CD16 shedding, and can be inhibited by an ADAM17 inhibitor.

      We thank the reviewer for these suggestions, which we have used to present the ADAM17-inhibition data more clearly.

      (1) Effect on subset proportions, independent of degranulation. This is shown in Fig. 8B (now Figure 7B). Because the two subsets differ greatly in baseline frequency, with CD56<sup>dim</sup>CD16<sup>bright</sup> cells constituting the large majority of CD56<sup>dim</sup> NK cells before stimulation, a change in raw bulk proportion is small and difficult to interpret, for example a shift from roughly 95% to 92.5% of the bright population. To place the two subsets on comparable footing, the panel reports, for each subset, the frequency following target exposure minus its frequency in the matched unstimulated condition. Presented this way, ADAM17 inhibition clearly reduces the activation-associated change in subset proportions, consistent with reduced CD16 shedding. This normalization is stated in the legend and is now described in the Results text so that the analysis is not overlooked.

      (2) Interpretation of the ADAM17-inhibition data. We agree that CD16 downregulation occurs as a consequence of activation-induced shedding and is prevented by ADAM17 inhibition; this is not in dispute. We would, however, offer an additional observation that bears on whether the between-subset functional difference is itself a product of that shedding. In Fig. 8A (now Figure 7A), combining ADAM17 inhibition with NKG2D blockade reduces CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation to 2.03%, below both anti-NKG2D alone at 5.24% and ADAM17 inhibition alone at 5.05% (p < 0.0001 and p = 0.0001 respectively). If the CD16<sup>dim</sup> advantage were solely a consequence of CD16 shedding, and if NKG2D blockade acted only by reducing that shedding, the combination could not reduce degranulation further than ADAM17 inhibition alone.

      The same figure also shows that the two perturbations act in opposite directions within the CD56<sup>dim</sup>CD16<sup>bright</sup> subset: anti-NKG2D reduced their degranulation from 2.58% to 0.92% (p = 0.0263), whereas ADAM17 inhibition increased it to 3.99% (p = 0.0679). NKG2D therefore supports the response of CD56<sup>dim</sup>CD16<sup>bright</sup> cells while ADAM17 activity constrains it, the reverse of the pattern in CD56<sup>dim</sup>CD16<sup>dim</sup> cells. Together with the opposite responses of the two subsets to ADAM17 inhibition described in the central response, this indicates that NKG2D and ADAM17 contribute through separable mechanisms and that the two subsets are not one population at different stages of shedding. These data are now presented in a more structured form and the interpretation is explicit in the revised text.

      Summary statement

      Taken together, many of the data presented in the manuscript are consistent with the very well-established downregulation of CD16 expression on activated NK cells, suggesting that the observed association between reduced CD16 expression on CD56dim NK cells and enhanced effector functions is a consequence of higher activation of these NK cells.

      We appreciate the reviewer articulating the central concern so clearly. We agree that CD16 downregulation on activated NK cells is well established and occurs in our assays; where we reach a different conclusion is on whether the enhanced function of the CD56<sup>dim</sup>CD16<sup>dim</sup> subset is a consequence of that downregulation. As set out in the central response, three observations argue that it is not: the advantage is present in cells sorted into subsets before any target contact, where post-assay CD16 changes cannot apply (Fig. 2); the two subsets respond to ADAM17 inhibition in opposite directions, both in magnitude (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8), which is difficult to reconcile with their being one population at different activation levels; and blocking NKG2D together with ADAM17 reduces CD16<sup>dim</sup> degranulation below either alone (Fig. 8, now Figure 7), indicating that the advantage is driven by NKG2D through a mechanism separable from shedding. We therefore interpret the association between low CD16 and enhanced function not as activation-induced downregulation of a single population, but as a property of a distinct, pre-existing subset. This interpretation is stated and defended explicitly in the revised Discussion, and will be strengthened with the expanded pre-sorted experiments.

      Reviewer #2 (Public review):

      (1) The central conclusion is weakened by the use of CD16 as a stable phenotypic marker. CD16 is well established to be rapidly downregulated following NK-cell activation and target cell (K562 or infected cells) engagement through ADAM17-mediated shedding. NK cell shedding regulates NK cell effector functions by promoting target cell detachment, boosting serial killing capacity, and preventing overstimulation. Therefore, NK cells displaying a CD56dimCD16dim phenotype after co-culture cannot be assumed to represent a pre-existing subset with intrinsically superior cytotoxic activity, but may instead correspond to activated CD56dimCD16bright NK cells that have downregulated CD16 during the assay. Because the vast majority of the functional experiments classified NK cell subsets based on post-assay CD16 expression, it is difficult to distinguish intrinsic functional differences between NK cell subsets from activation-induced phenotypic conversion. This limitation affects the interpretation of most of the study's principal findings.

      We thank the reviewer for this careful and well-articulated concern, which we recognize as the central issue of the review, and which we address in full in the central response above. We agree with the reviewer's premises: CD16 is rapidly shed by ADAM17 upon activation, and this shedding is itself functionally important, promoting target detachment, supporting serial engagement, and limiting overstimulation. Indeed, ADAM17-mediated shedding is integral to the serial degranulation mechanism we propose. We also agree that classifying subsets by post-assay CD16 expression alone cannot, on its own, distinguish a pre-existing subset from activation-induced conversion.

      For this reason, our conclusion does not rest on post-assay classification. As detailed in the central response, the CD56<sup>dim</sup>CD16<sup>dim</sup> advantage is demonstrated in cells sorted into subsets before any target contact, where the readout is direct lysis rather than post-assay gating and where activation-induced shedding therefore cannot account for the difference (Fig. 2), an experiment both this reviewer and the peer reviewer identify as the strongest in the manuscript. This is reinforced by evidence that the two subsets are functionally distinct rather than one population caught at different stages of shedding: they respond to ADAM17 inhibition in opposite directions, both in the magnitude of degranulation (Fig. 7B, now Figure 6B) and in the number of serial degranulation events per cell (Fig. 9, now Figure 8); blocking NKG2D together with ADAM17 reduces CD16<sup>dim</sup> degranulation below either alone, indicating a mechanism separable from shedding (Fig. 8, now Figure 7); and the subsets differ before stimulation, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells and no corresponding difference in NKp46 (Fig. 6, now Figure 5). We will strengthen this further by expanding the sorted-subset experiments across additional donors, and the revised text rests the manuscript's conclusions explicitly on the pre-sorted data.

      (2) The "killing frequency" analysis presented in Figure 3 is based on a mathematical estimate rather than a direct experimental measurement. Since total target cell killing is measured in mixed NK cell populations, it cannot be attributed to individual NK cell subsets. This experiment must be repeated using purified NK cell subsets.

      We agree that the killing frequency in Fig. 3 (now Figure 2—figure supplement 1C) is a mathematical estimate rather than a direct measurement. We would add that this is intrinsic to the metric: killing frequency is a derived quantity whether calculated from mixed or purified populations, so repeating it on purified subsets would not convert it into a direct measurement. This analysis has been moved to the supplementary material, and its limitations are stated in the Discussion, namely that killing frequency is an estimate and should be interpreted as such. Direct, subset-resolved killing is instead provided by Figure 2, in which NK cells sorted before target exposure show that purified CD56<sup>dim</sup>CD16<sup>dim</sup> cells lyse HIV-infected targets more efficiently than purified CD56<sup>dim</sup>CD16<sup>bright</sup> cells; this is the measurement on which our conclusion regarding direct killing rests, and it is the experiment we will expand across additional donors.

      (3) The serial degranulation assay presented in Figure 4 does not directly measure serial target cell killing and therefore does not support the conclusion that CD56dimCD16dim NK cells possess superior serial killing capacity. Furthermore, the increased serial degranulation observed in the CD16dim population could simply reflect activation-induced CD16 downregulation rather than an intrinsic property of this subset. This experiment should therefore be repeated using purified NK cell subsets.

      We agree that Fig. 4 (now Figure 3) measures serial degranulation, not serial killing directly. The text has been revised throughout to describe this as serial degranulation rather than serial killing, so that our conclusions match what was measured. Regarding the concern that increased serial degranulation in CD16<sup>dim</sup> cells reflects activation-induced CD16 downregulation, we address this in the central response; the opposite responses of the two subsets to ADAM17 inhibition (Figs. 7B and 9, now Figures 6B and 8) argue against that interpretation. As the reviewer suggests, we will repeat the serial degranulation assay on subsets purified before target exposure if cell yields from the sort permit. We note that this assay requires four sequential labelling and washing steps and that the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56<sup>dim</sup> NK cells, so the number of sorted cells recovered may be limiting; we will report the outcome either way.

      (4) The finding that CD56dimCD16dim NK cells exhibit greater ADCC activity is somewhat counterintuitive given the central role of CD16 in mediating ADCC. Moreover, these experiments are likely confounded by activation-induced CD16 downregulation, which is expected to be even more pronounced during ADCC. Thus, the apparent superiority of the CD56dimCD16dim subset may simply reflect the conversion of activated CD56dimCD16bright NK cells into the CD16dim gate rather than intrinsically greater ADCC activity. To directly compare the intrinsic ADCC capacity of each subset, these experiments should be repeated using purified NK cell populations prior to target-cell stimulation.

      We agree that the greater antibody-dependent response of CD56<sup>dim</sup>CD16<sup>dim</sup> cells is counterintuitive given the central role of CD16, and we regard it as an informative finding rather than an artifact. As set out in our revised Discussion, the surface density of gp120 on HIV-infected primary T-cells is approximately 6.4 × 10<sup>2</sup> molecules per cell (Vasiliver-Shamis et al., 2008), two orders of magnitude below high-density antigens such as CD20 on Raji cells at approximately 5 × 10<sup>4</sup> molecules per cell (Lallemand et al., 2017). Antibody-dependent responses against HIV-infected cells therefore proceed under conditions of limiting antigen, and both CD56<sup>dim</sup> subsets face the same constraint. What differs between them is not the constraint but the NKG2D available to meet it: CD56<sup>dim</sup>CD16<sup>dim</sup> cells carry higher NKG2D before target contact and respond more strongly, despite their lower CD16. Consistent with a requirement for a second signal under these conditions, ADAM17 inhibition reduced CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation even at 0 µg/mL VRC01 (Figs. 7B and 7C, now Figures 6B and 6C), where no antibody is present to engage CD16.

      Regarding the concern that this superiority reflects conversion of CD16<sup>bright</sup> cells into the CD16<sup>dim</sup> gate, we address this in full in the central response; the opposite responses of the two subsets to ADAM17 inhibition, in both degranulation magnitude (Fig. 7B, now Figure 6B) and serial degranulation (Fig. 9, now Figure 8), argue against it. As the reviewer recommends, we will directly compare the intrinsic antibody-dependent capacity of each subset using cells purified before target-cell stimulation, extending the pre-sorted approach of Figure 2 to the antibody-dependent setting across additional donors.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As discussed in the public review, the majority of the functional assays should be repeated using purified NK cell subsets. This approach would eliminate the confounding effect of activation-induced CD16 downregulation and allow the intrinsic functional properties of each subset to be directly compared.

      We agree, and this is the central experimental commitment of our revision. As described in the central response, we will repeat the specific-lysis and antibody-dependent degranulation assays using NK cell subsets purified before target-cell exposure, so that the intrinsic functional properties of each subset are compared directly and are not subject to activation-induced changes in CD16 expression. We will also perform the serial degranulation assay on sorted subsets if cell yields permit; that assay requires four sequential labelling and washing steps, and the CD56<sup>dim</sup>CD16<sup>dim</sup> subset constitutes fewer than 5% of CD56 <sup>dim</sup> NK cells, so we cannot commit to it in advance of the sort. This extends the pre-sorted approach already used in Figure 2, which the reviewer identifies as the strongest evidence in the manuscript, across additional donors and across the functional readouts. These experiments require cell sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      (2) The experiments performed with purified NK cell subsets in Figure 2 provide the strongest evidence supporting the authors' conclusion that CD56dimCD16dim NK cells exhibit greater direct cytotoxicity against HIV-infected target cells. These data are the most convincing in the manuscript because they are not confounded by post-assay changes in CD16 expression. However, unlike the other functional assays, no representative gating strategy or raw flow cytometry plots are provided, and the results appear to be based on a single representative experiment. Given the importance of these data to the manuscript's central conclusion, this experiment should be expanded to include biological replicates from additional donors, representative flow cytometry plots, and validation using additional HIV-1 infectious molecular clones.

      We appreciate the reviewer identifying the sorted-subset experiments in Figure 2 as the strongest evidence for our conclusion, and we agree these data warrant expansion. In the revised manuscript we will:

      (1) Expand the experiment to include biological replicates from additional donors, with all donors shown rather than a single representative experiment.

      (2) Provide the representative gating strategy and flow cytometry plots for the sorted subsets. The reviewer is correct that these should be included, and we will add them, including for the expanded experiments.

      (3) Validate the finding using additional HIV-1 strains. We note, for clarity, that the virus used throughout this study is a primary patient isolate (HIV-1<sup>SHM-1</sup>), as stated in the Materials and Methods, rather than an infectious molecular clone. We have now compared NK cell degranulation against autologous CD4<sup>positive</sup> T-cells productively infected with HIV-1<sup>SHM-1</sup>, with the X4-tropic infectious molecular clone HIV-1<sup>NL4-3</sup>, and with the R5-tropic laboratory-adapted strain HIV-1<sup>BaL</sup>, at three effector cell to target cell ratios. CD56 <sup>dim</sup> CD16 <sup>dim</sup> cells degranulated more than CD56 <sup>dim</sup>CD16<sup>bright</sup> cells against every virus at every ratio, in all nine comparisons at p < 0.0001. Both subsets responded less to HIV-1<sup>NL4-3</sup> and HIV-1<sup>BaL</sup> than to HIV-1<sup>SHM-1</sup>, and did so in proportion: the ratio of CD56 <sup>dim</sup>CD16 <sup>dim</sup> to CD56 <sup>dim</sup> CD16<sup>bright</sup> degranulation ranged from 2.7 to 3.8 across all nine conditions. The magnitude of the response therefore varies with the virus, whereas the relationship between the two subsets does not. These data are included in the revised manuscript as Figure 1—figure supplement 3, with the statistical analysis in a supplemental table.

      The experiments described in points 1 and 2 require cell sorting and primary-cell work and will be completed within approximately six to eight weeks and provided with the revised manuscript.

      (3) The mechanism underlying the enhanced effector function of CD56dimCD16dim NK cells remains unclear. Although the phenotypic characterization presented in Figure 5 (and related supplement figures) is informative, NK cell receptor expression was assessed after target cell stimulation, when it may already have been altered by activation and CD16 downregulation. Receptor expression should therefore be evaluated prior to stimulation. In addition to NKG2D, the authors should also consider assessing additional activating receptors, notably NKp30, which has recently been implicated in the elimination of autologous HIV-1-infected cells (PMID: 41079618).

      We agree that receptor expression should be assessed before stimulation, and in fact it is. In Fig. 6 (now Figure 5), the receptor gMFI data include the no-target condition, showing that CD56<sup>dim</sup>CD16<sup>dim</sup> cells express 1182 gMFI more NKG2D than CD56 <sup>dim</sup>CD16<sup>bright</sup> cells before any target-cell contact (p < 0.0001), approximately 1.6-fold, and therefore before any activation-induced change in receptor expression. NKp46, measured on the same cells in the same wells, did not differ between the subsets under the same condition (mean difference 45.67 gMFI, p = 0.0668), indicating that the difference is specific to NKG2D rather than a general difference in activating receptor density. The baseline condition is now labelled explicitly as 1:0 in the revised figure and described as such in the Results text.

      Regarding NKp30, we have assessed this receptor. Degranulation did not differ between NKp30 positive and NKp30 negative cells within either CD56<sup>dim</sup> subset, whereas both CD56<sup>dim</sup>CD16<sup>dim</sup> groups exceeded both CD56<sup>dim</sup>CD16<sup>bright</sup> groups regardless of NKp30 status, indicating that the enhanced degranulation of the subset is not attributable to NKp30, paralleling our finding for NKp46. These NKp30 data are included in the revised manuscript as Figure 5—figure supplement 1, with the corresponding statistical analysis in a supplemental table. We note that this analysis addresses whether NKp30 accounts for the difference between the subsets; it does not exclude a role for NKp30 in NK-cell recognition of HIV-infected cells more generally, consistent with the study the reviewer cites, which we now discuss.

      (4) In Figure 5, the histograms corresponding to the uninfected and ΔVpr conditions appear to be identical. If this is indeed the case, this represents a serious concern, as these are two distinct experimental conditions and should not be represented by the same flow cytometry plot. This raises the possibility of an inadvertent panel duplication. The authors should carefully verify the figure and replace the duplicated panel if necessary.

      We thank the reviewer for this careful observation. On review, the histograms in Figure 5A were reproduced from our earlier study (Ward et al., PLoS Pathog 2009;5(10):e1000613) and should not have been included without attribution. We have removed that panel and cite the original finding in its place.

      Figures 5B and 5C are both new results from this study, and both are retained. Figure 5B, which shows that NK cells degranulate in response to wild-type HIV-infected targets but not to ΔVpr-infected or uninfected targets, becomes Figure 4—figure supplement 1 in the revised manuscript. Figure 5C, the NKG2D blockade experiment, becomes Figure 4A and 4B.

      (5) The ADCC experiments and calculation require additional methodological clarification, particularly the analyses presented in Figure 7C. Although NK cell degranulation is commonly used as a surrogate marker of ADCC, the data presented in Figure 7 do not appear to isolate the antibody-dependent component of the response. To specifically quantify ADCC-mediated degranulation, the degranulation induced by HIV-infected target cells alone (i.e., in the absence of VRC01) should be subtracted from that measured in the presence of VRC01. Notably, in Figure 7C (DMSO), the CD56dimCD16dim population appears to exhibit similar levels of degranulation in the absence and presence of VRC01, suggesting that antibody-dependent degranulation may be limited in this subset, which does not support the author's conclusions.

      We thank the reviewer for raising this, and we agree that the antibody-dependent and antibody-independent components of the response should be distinguished. We would, however, respectfully argue against the subtraction as a means of doing so, and we believe the experiments already in the manuscript address the underlying question more directly.

      The condition without VRC01 is not a background to be removed. It is the NKG2D-driven response of the same cells to the same infected targets, measured through the same degranulation machinery, and it is one of the principal findings of the study. Subtracting it treats the two components as though they were independent and additive, when both converge on a single immunological synapse and a single degranulation event per cell. The difference between the two conditions is therefore not the antibody-dependent response; it is the increment in total degranulation produced by adding antibody, which is a different quantity and one that carries no clean interpretation at the level of the individual cell.

      The question the reviewer raises, whether the antibody-dependent component differs between the subsets, is answered directly by the two-way ANOVA of these data. Across the VRC01 titration, the effect of NK cell subset accounts for 85.91% of the total variation (F(1, 16) = 424.0, p < 0.0001) and the effect of VRC01 concentration for 10.18% (F(3, 16) = 16.74, p < 0.0001), while the subset × VRC01 interaction is not significant (F(3, 16) = 1.112, p = 0.3732) and accounts for 0.68% (Supplemental Table 20 in the revised manuscript). The absence of an interaction means that adding antibody raises the response of both subsets by a comparable amount, and that the difference between the subsets is the same at every VRC01 concentration tested. This is a statistical statement about the antibody-dependent component, obtained without subtracting one condition from another.

      We agree with the implication the reviewer draws from this, and we state it plainly in the revised Discussion: the antibody-dependent increment is modest in both subsets. We attribute this to the very low surface density of gp120 on HIV-infected primary T-cells, approximately 6.4 × 10<sup>2</sup> molecules per cell (Vasiliver-Shamis et al., 2008), which is two orders of magnitude below high-density antigens such as CD20 on Raji cells, approximately 5 × 10<sup>4</sup> molecules per cell (Lallemand et al., 2017). Under these conditions the antibody-dependent signal available to any NK cell is limited, and this applies equally to both subsets.

      Where we differ from the reviewer is on the conclusion this supports. That the antibody-dependent increment is modest in both subsets does not weaken our central claim, which is comparative: at every VRC01 concentration tested, including in the presence of antibody, CD56<sup>dim</sup>CD16<sup>dim</sup> cells degranulate more than CD56<sup>dim</sup>CD16<sup>bright</sup> cells against antibody-coated HIV-infected targets. That comparison is what the manuscript reports, and it is unaffected by how the response is partitioned between its antibody-dependent and antibody-independent components.

      We also note that the experiments in Figs. 7B and 7C (now Figures 6B and 6C) do isolate a component of the response experimentally rather than arithmetically. Inhibiting ADAM17 removes the contribution that depends on CD16 turnover, and it does so in opposite directions in the two subsets, reducing CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation and increasing that of CD56<sup>dim</sup>CD16<sup>bright</sup> cells at every VRC01 concentration. These are direct experimental manipulations of the antibody-dependent pathway, and they are more informative than the arithmetic difference between two conditions.

      Finally, we take the reviewer's point that the analyses in Fig. 7C require clearer explanation. In the revised manuscript we state explicitly what is plotted, namely the percentage of each CD16 subset among CD107a positive CD56<sup>dim</sup> NK cells, we describe the background subtraction that applies to all CD107a data in this study, and we report the statistical analysis of each panel in full.

      (6) It is also unclear how the authors interpret the effects of ADAM17 inhibition. While ADAM17 inhibition increases the ADCC activity of the CD56dimCD16bright population, it simultaneously decreases that of the CD56dimCD16dim population. An alternative explanation is that inhibition of CD16 shedding prevents activated CD56dimCD16bright NK cells from transitioning into CD56dimCD16dim during the assay. This possibility should be discussed and experimentally addressed, as it provides a plausible alternative interpretation of the observed phenotype.

      We thank the reviewer for articulating this alternative, which we address in full in the central response. We agree that the opposite effects of ADAM17 inhibition on the two subsets are central to interpreting these experiments, and we interpret them as evidence that the two are distinct populations rather than one transitioning into the other.

      The reviewer's alternative, that ADAM17 inhibition prevents CD56<sup>dim</sup>CD16<sup>bright</sup> cells from transitioning into the CD56<sup>dim</sup>CD16<sup>dim</sup> gate, predicts that blocking shedding should reduce the CD56<sup>dim</sup>CD16<sup>dim</sup> population by cutting off its supply from CD16<sup>bright</sup> cells. Three observations argue against this being the explanation for the functional difference. First, the effect is not merely a change in population size but a change in per-cell function in opposite directions: ADAM17 inhibition increases the number of serial degranulation events in CD56<sup>dim</sup>CD16<sup>bright</sup> cells while decreasing them in CD56<sup>dim</sup>CD16<sup>dim</sup> cells (Fig. 9, now Figure 8), which is difficult to explain if the dim cells were simply bright cells prevented from converting. Second, blocking NKG2D together with ADAM17 reduces CD56<sup>dim</sup>CD16<sup>dim</sup> degranulation below either treatment alone (Fig. 8A, now Figure 7A); if NKG2D blockade acted only by reducing the shedding that drives the putative transition, the combination could not exceed the effect of ADAM17 inhibition alone. Third, we have tracked the fate of each subset sorted before target exposure: at one hour, when degranulation is maximal, only approximately 4% of sorted CD16<sup>bright</sup> cells were found in the CD16<sup>dim</sup> gate, while approximately 35% had moved to the CD16<sup>negative</sup> gate (Author response image 1 accompanying this response). Shedding therefore directs CD16<sup>bright</sup> cells past the CD16<sup>dim</sup> gate rather than into it.

      We discuss this alternative explicitly in the revised Discussion and will address it further experimentally by repeating these assays on subsets purified before target exposure, where no transition can occur during the assay.

      Reviewing Editor Comments:

      The conclusion that CD56dimCD16dim NK cells are intrinsically superior effectors against HIV-infected target cells requires additional evidence because CD16 is rapidly downregulated following NK-cell activation. Throughout most of the study, NK-cell subsets are classified after target-cell encounter, making it difficult to distinguish pre-existing CD56dimCD16dim cells from activated CD56dimCD16bright cells that have undergone ADAM17-mediated CD16 shedding. The authors should repeat functional experiments using NK-cell subsets purified before target-cell exposure and determine the extent to which ADAM17 inhibition alters subset frequencies and functional readouts. These experiments are essential to establish whether the observed functional differences reflect intrinsic biology rather than activation-induced phenotypic conversion.

      We thank the editor for this clear synthesis of the central concern, which we address in full in the central response above. In brief, our conclusion does not rest on post-encounter classification: the CD56<sup>dim</sup>CD16<sup>dim</sup> advantage is established in cells sorted into subsets before any target contact, using direct lysis as the readout (Fig. 2), and is reinforced by the opposite responses of the two subsets to ADAM17 inhibition in both degranulation magnitude (Fig. 7B, now Figure 6B) and serial degranulation (Fig. 9, now Figure 8), by the separable contributions of NKG2D and ADAM17 (Fig. 8, now Figure 7), and by pre-stimulation differences between the subsets, with NKG2D 1182 gMFI higher on CD16<sup>dim</sup> cells and no corresponding difference in NKp46 (Fig. 6, now Figure 5). We agree these questions are central and will repeat the functional experiments on subsets purified before target exposure, and the effect of ADAM17 inhibition on subset frequencies is presented explicitly in Figs. 7C and 8B (now Figures 6C and 7B).

      Major conclusions should be supported by more rigorous experimental validation. In particular, the sorted NK-cell experiments should be expanded using multiple independent donors, include killing of uninfected target cells as controls and provide representative gating strategies and flow cytometry plots. Likewise, the current analyses of killing frequency, serial killing, and ADCC should be strengthened by direct measurements using purified NK-cell subsets rather than mathematical estimates or analyses performed in mixed NK-cell populations.

      We agree and will strengthen the validation as follows. The sorted-subset experiments will be expanded across multiple independent donors, with all donors shown. Uninfected-target controls will be included for the sorted-cell lysis experiments (Fig. 2); the corresponding CD107a controls against uninfected targets are already presented in Figure 1—Figure Supplement 4C of the revised manuscript. Representative gating strategies and flow cytometry plots will be provided for the sorted-cell experiments, as for our other assays. Regarding direct measurement: the killing-frequency metric (Fig. 3, now Figure 2—figure supplement 1C) is a mathematical estimate whether derived from mixed or purified populations, and it has been moved to the supplementary material with this limitation noted in the Discussion, while direct, subset-resolved killing is provided by the pre-sorted lysis experiment (Fig. 2), which we will expand; the serial degranulation assay (Fig. 4, now Figure 3) is now described as serial degranulation rather than serial killing, and will be repeated on purified subsets if cell yields from the sort permit; and the antibody-dependent comparison will be performed on subsets purified before stimulation.

      Some aspects of the data analysis and presentation require clarification. The authors should evaluate receptor expression before target-cell stimulation, clarify the ADCC analyses and interpretation of ADAM17 inhibition, verify the apparent duplicated flow-cytometry panel, apply appropriate statistical corrections for multiple comparisons where necessary, and streamline the presentation by emphasizing the principal conclusions rather than extensive descriptive analyses.

      We have addressed each of these. Receptor expression before stimulation is shown in the gMFI data of Fig. 6 (now Figure 5) at the no-target condition, where CD56<sup>dim</sup>CD16<sup>dim</sup> cells carry 1182 gMFI more NKG2D than CD56<sup>dim</sup>CD16<sup>bright</sup> cells (p < 0.0001) with no corresponding difference in NKp46 (p = 0.0668); this condition is now labelled explicitly as 1:0 in the figure and described as such in the Results text. The antibody-dependent analyses and the interpretation of ADAM17 inhibition are clarified in the revised text, as detailed in our responses to the three reviewers and the central response.

      On the duplicated panel: the histograms in Figure 5A were reproduced from Ward et al. (2009) and should not have been included without attribution. That panel has been removed and the original finding is cited in its place. Figures 5B and 5C are both new results and are retained, becoming Figure 4—figure supplement 1 and Figures 4A and 4B respectively.

      We have applied appropriate multiple-comparison corrections, using the post-hoc test matched to each comparison structure and reporting adjusted p-values throughout; the design and post-hoc test used for each figure and panel are given in a new supplemental table. Finally, we have streamlined the presentation by opening each Results section with a statement of the principal finding before the supporting data, and by reducing the inhibitory receptor section from approximately 1,800 words and more than 80 reported p-values to approximately 700 words and 24, with the detailed data retained in the supplemental tables and their significance synthesized in the Discussion.

      References

      Amand M, Iserentant G, Poli A, Sleiman M, Fievez V, Sanchez IP, Sauvageot N, Michel T, Aouali N, Janji B, Trujillo-Vargas CM, Seguin-Devaux C, Zimmer J. 2017. Human CD56<sup>dim</sup>CD16<sup>dim</sup> cells as an individualized natural killer cell subset. Frontiers in Immunology 8:699. doi:10.3389/fimmu.2017.00699.

      Lallemand C, Liang F, Staub F, Simansour M, Vallette B, Huang L, Ferrando-Miguel R, Tovey MG. 2017. A novel system for the quantification of the ADCC activity of therapeutic antibodies. Journal of Immunology Research 2017:3908289. doi:10.1155/2017/3908289.

      Vasiliver-Shamis G, Tuen M, Wu TW, Starr T, Cameron TO, Thomson R, Kaur G, Liu J, Visciano ML, Li H, Kumar R, Ansari R, Han DP, Cho MW, Dustin ML, Hioe CE. 2008. Human immunodeficiency virus type 1 envelope gp120 induces a stop signal and virological synapse formation in noninfected CD4+ T cells. Journal of Virology 82:9445-9457. doi:10.1128/JVI.00835-08.

      Ward J, Davis Z, DeHart J, Zimmerman E, Bosque A, Brunetta E, Mavilio D, Planelles V, Barker E. 2009. HIV-1 Vpr triggers natural killer cell-mediated lysis of infected cells through activation of the ATR-mediated DNA damage response. PLoS Pathogens 5(10):e1000613. doi:10.1371/journal.ppat.1000613.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We have made major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We have changed the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions have greatly improved our work and presentation and we are grateful to the reviewers and our editors. We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this has been clarified in the main text. We have now included an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS strengthens our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results are reported in the new Supplemental Table 6.

      We have now included a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results are now reported in the manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We have addressed other recommendations from this reviewer as follows. We cite and discuss previous SCD rodent model work observing decreased butyrate in disease, and we have updated our results and discussion sections regarding associations between the microbiome and virome and clinical and molecular features.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses:

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself. (REVISION POINT 3)

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this has been clarified in the main text. We have tempered our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We have strengthened our control of clinical confounders, and added clearer statistical correction, as described below, with corresponding updates to the manuscript. We look forward to conducting future studies to test causality and understand mechanism.

      We have now done sensitivity analysis within SCD patients to determine whether our microbiome and virome results associate with treatment and clinical severity. To evaluate whether microbiome and virome features were explained by clinical or demographic heterogeneity within the SCD cohort, we restricted analyses to SCD participants. We fit a separate multivariable regression model for each feature. Each model included age, sex, hydroxyurea use, transfusions in the past year, and acute care utilization in the past year as predictors.

      feature_z <sup>~</sup> age_z + sex_F + HU + log1p(TxPastYr)_z + log1p(AcuteCarePastYr)_z

      Non-negative abundance, ratio, pathway, viral, and count-like variables were log-transformed to reduce skew, using feature-specific pseudocounts for zero-containing microbiome/virome variables and ln (1 + x) transformation for count covariates. Diversity and indicator scores were not log-transformed. Continuous variables were then standardized to Z-scores before modeling. Models were fit using ordinary least squares with HC3 robust standard errors. FDR correction was applied separately for each model term across the tested microbiome and virome features. No microbiome or virome feature showed an FDR-significant association with hydroxyurea use, transfusions in the past year, or acute care utilization in the past year. These results are reported in the manuscript and in the new Supplemental Table 7.

      We have reported multiple-testing correction results for all associations between microbiome and virome features and clinical and molecular features and updated manuscript figures accordingly.

      To evaluate relationships between microbiome/virome features and clinical or immune markers within the SCD cohort, we performed Spearman correlation analyses. Microbiome and virome features were organized into four prespecified feature groups: community metrics, taxa, functional pathways, and viral features. Clinical and immune markers were grouped into marker sets for visualization and multiple-testing correction, including hematologic clinical markers, hemolysis markers, creatinine, acute care burden, and cytokines/chemokines. Spearman correlation coefficients were calculated for each feature–marker pair. Benjamini-Hochberg FDR correction was applied within each prespecified feature group by marker group block. Nominal associations were defined as p < 0.05, FDR-significant associations as q < 0.05, and trends as q < 0.10. In the heatmap figure, boxes now indicate nominal p < 0.05, asterisks indicate q < 0.05, and daggers indicate q < 0.10.

      We have addressed other recommendations from this reviewer as follows. We have modified our language describing prophage/immune associations. We have revised the Methods to clarify how abundance data were processed before MaAsLin2 modeling. Taxonomic profiles were analyzed using MaAsLin2 with total-sum scaling normalization and log transformation, while pathway profiles were analyzed without additional normalization because the input pathway table had already been normalized prior to MaAsLin2 analysis; MaAsLin2 log transformation was then applied. We agree that relative abundance metagenomic data are compositional, and we have revised the text to clarify that these analyses identify covariate-adjusted associations with transformed relative abundance rather than absolute abundance. We also now note this as a limitation of the study. We also note the limitations of F:B as a metric. We have added text describing the need for further analysis of the prophages to understand their patterns of host range and transmission and their associations with features such as shared geography, health care exposure, and diet. Finally, we have updated Figure 5 to reflect our updated analysis with significance indicated.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We have now noted in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single-centre, cross-sectional. We have further made modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review)

      Summary:

      The authors report the results of a tDCS brain stimulation study (verum vs sham stimulation of left DLPFC; between-subjects) in 46 participants, using an intense stimulation protocol over 2 weeks, combined with an experience-sampling approach, plus follow-up measures after 6 months.

      Strengths:

      The authors are studying a relevant and interesting research question using an intriguing design, following participants quite intensely over time and even at a follow-up time point. The use of an experience-sampling approach is another strength of the work.

      Comments on revised version.

      With the last round of revisions, the authors have now addressed my concerns.

      Thank you to re-review this revision, and we all appreciate you kindly contributing to substantially improve the conceptualization, statistics and statements for this manuscript.

      Reviewer #4 (Public review):

      Summary:

      The current study tested the effects of repeated sessions of tDCS targeting the DLPFC on procrastination behavior. The main outcome is that anodal versus sham DLPFC tDCS reduces procrastination behavior on both a short-term and a long-term scale up to six months after the stimulation sessions.

      Strengths:

      The current study tests competing models of procrastination with state-of-the-art high-definition transcranial electric stimulation. The study assesses stimulation effects on procrastination on both a short-term and a long-term scale, suggesting that repeated stimulation of the prefrontal cortex reduces procrastination on a time scale of up to six months.

      Weaknesses:

      The manuscript has already been reviewed and revised before, and it seems that the quality of the manuscript has substantially improved as a result of this revision process. I agree with the other reviewers that one must be cautious with drawing conclusions regarding the cognitive mechanisms underlying this effect, as many different cognitive functions are implemented by the DLPFC.

      We do appreciate you to take valuable time offering those insightful and helpful comments on this revised manuscript. As you kindly raised, this revision has redrawn conclusions and statements on domain-specific mechanistic roles of DLPFC in interpreting why this neuromodulation treatments are effective.

      One aspect of the current results that puzzles me is the strength of the current stimulation effects. Meta-analyses suggest that tDCS shows only small-to-moderate effect sizes (with Cohen's d around 0.5). While the authors report no effect sizes for their statistical models, the small p values, in combination with the unusually small sample size of 18 participants per group, suggests that the effect size must be rather large. Can the authors provide an estimate of the effect size of their stimulation effects? If they are considerably larger than to be expected, could the authors give an explanation for why their stimulation setup is showing much stronger effects than comparable high-definition tDCS studies on cognition or decision making?

      Thank you for raising this very crucial question in effect size determination. We fully understand that this large effect size makes you puzzled, and that the limited sample size indeed attenuates detectable power in the statistics. We completely agree that reporting the actual effect sizes is essential for interpreting the magnitude of our findings, and we appreciate the opportunity to clarify why the observed effects in our study appear substantially larger than the small-to-moderate effect sizes typically reported in meta-analyses of tDCS studies on cognition and decision-making.

      Following your suggestion, we have calculated the effect sizes for our primary outcomes. Based on the simple effect analyses (pre- vs. post-neuromodulation within the active neuromodulation group), we computed Cohen’s d for the within-group changes: for task-execution willingness, d = 2.37 (95% CI [1.49, 3.25]); for the actual procrastination rate, d = 1.52 (95% CI [0.87, 2.16]).

      We acknowledge that these effect sizes are considerably larger than the typical d ≈ 0.5 reported in the tDCS literature. After careful consideration, we attribute this discrepancy to three methodological and conceptual differences between our study and typical cognitive/decision-making tDCS studies. First, the small-to-moderate effect sizes are observed in studies utilizing a single-session tDCS protocol, yet our study employed an intensive 7-session HD-tDCS protocol over 15 days. Therefore, multi-session tDCS that induces cumulative, activity-dependent long-term potentiation (LTP)-like plasticity may substantially amplify and consolidates behavioral effects compared to single-session stimulation (Ke et al., 2023; Zhong et al., 2021). Second, most tDCS studies on cognition recruit healthy young adults who often perform near ceiling on laboratory tasks, leaving little "room for improvement" and thereby constraining the observable effect size. In our study, we strictly screened for severe chronic procrastinators. Because our participants had severe baseline deficits in task execution, the "ceiling space" for behavioral improvement was much larger, naturally inflating the observable effect size of the intervention. Lastly, given all the procrastinators completed tasks in the last session (0% procrastination rate in the active group, without within-group variance), mathematically, this boundary variable (0% vs 100%) artificially inflates the effect size estimate when calculating Cohen’s d with a near-zero post-test standard deviation.

      Nevertheless, as a sensitivity analysis, those findings are confirmed by Beta regression model addressing the risks of boundary variables, indicating that the potential inflation of effect sizes is statistically acceptable:

      In summary, while the observed Cohen's d values are unusually large, they are contextually justified by the cumulative nature of our multi-session protocol, the targeted clinical-like population, and the mathematical properties of bounded behavioral metrics.

      Results Section (Page 9, Line 441-444)

      “... For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001, Cohen d = 2.37, 95% CI: 1.49-3.25; Fig. 3A and Fig. S2a).”

      Results Section (Page 9, Line 454-458)

      “... Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004, Cohen d = 1.52, 95% CI: 0.87-2.16), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group.”

      Regarding the strengths of the stimulation effects, I moreover found remarkable that the post-test procrastination rate was 100% in all (!) participants in the DLPFC group (figure 3F). I admit that it is hard to trust results that have no individual variation at all. This means that all participants are perfect responders to tDCS, which is again at variance what one typically expects for tDCS (where one usually has many non-responders). Do the authors have an explanation for this?

      Thank you for raising this highly important and helpful comment. Indeed, we fully understand that this result (a 100% task complete rate among all participants in the DLPFC group) is somewhat extraordinary. This pattern was equally striking to us when unblinded the data. After carefully scrutinizing the data and statistics, we are thrilled to confirm that this pattern is true. In support of this observation, we were gratified to receive numerous thank-you letters from participants who engaged in active neuromodulation. They expressed gratitude to us, and reported that they have substantially ameliorated procrastination behavior in real-life activities after completing the trial. While this does not constitute formal scientific evidence, we are also glad to see the benefits of this neuromodulation for those procrastinators.

      Two reasons could account for this pattern herein. One interpretation is to attribute this pattern to “floor effect”. In the present study, the procrastination rate was calculated as 1 minus the task-completion rate (e.g., 80%, 60%, 40%) by the deadline. At last stimulation sessions (#6 and #7), all the participants completed their real-life tasks before the deadline, yielding a 0% (1 minus 100% completion rate) procrastination rate, without any between-individual variation. Thus, rather than there being no individual variation in procrastination, this scalar – the procrastination rate - is too insensitive to capture subtle differences per se. For instance, although participants #1 and #2 both showed a 0% procrastination rate - meaning that both completed their tasks before the deadline - Participant #1 might have completed it 3 hours before the deadline, whereas Participant #2 might have completed it only 10 minutes before. In this case, the “scalar inflation” emerges to let us perceive that both participants have equivalent procrastination rates, although participant #2 may have a higher procrastination level than #1. As conceptually defined in the field, procrastination is contextualized as “not completing a task before the deadline”. Thus, if this task is completed before the deadline, regardless of whether it was finished close to or far in advance of the deadline, this case is defined as “no procrastination”. In the present study, the primary outcome is whether a participant procrastinated on a real-life task before the deadline in real-world settings, irrespective of when she/he completed this task. Thus, this scalar - procrastination rate - fits our conceptualization of procrastination.

      Another reason is the potential accumulative effects from sequential multi-session tDCS stimulation, as we explained above. As shown in Mann-Kendall trend tests, the procrastination rates show a significant linear downtrend in the active neuromodulation group across sessions, even after removing sessions #6 and #7. This indicates that the improvements of going against procrastination may be sequentially accumulative along with the increase in sessions, implying a potential “dose-dependent effect”. Despite a speculative interpretation, this “dose-dependent effect” in neuromodulation has been well-documented in previous studies, showing the robustly linear association between the number of sessions and effectiveness (c.f., Cole et al., 2020; Hutton et al., 2023; Sabé et al., 2024; Schulze et al., 2018). Therefore, although this extreme pattern is somewhat extraordinary compared to previous observations, it makes sense.

      We also conducted robustness check by removing sessions #6, #7, and both, to validate whether this results were biased by “scalar inflation”. We do believe that this analysis could support statistical robustness to go against potential biases from extreme cells. By doing so, we found that all the group*treatment_day interaction effects remained significant when removing either session #6 or session #7 (or even both, all p-values < .05), indicating high statistical robustness. Please see Table S3 and Table S4.

      Taken together, in spite of their being extraordinary, we confirm that those findings are statistically robust to extreme outliers. As you kindly suggested, we have added those findings of the robustness check into the revised Supplemental Materials section.

      In any case, I am surprised by the rather small sample size. Due to the small effect sizes for tDCS, it is common to have a minimum of 30 subjects per group in between-subject designs. According to G*Power, a between-subject design with 17 subjects per group could detect only relatively large effect sizes of Cohen's d = 0.99 (alpha = 5%, power = 80%, independent-samples t-test). As explained above, this is far above the effect size that can be expected for tDCS. In addition, small samples bear the risk that results strongly depend on outliers in the data, which might explain the strong effect size observed in the current study. The small sample size should be discussed as a major limitation of the current study and that the results need to be replicated by studies with larger sample sizes. Moreover, to rule out that the results are driven by outlier in the data, the authors should show individual data points in all plots showing empirical data.

      We sincerely thank you for this highly constructive and methodologically sound critique. We completely agree that sample size is a critical consideration in tDCS research, and that visualizing individual data points is essential to rule out the possibility that our findings are driven by outliers.

      We acknowledge that our sample size is smaller than the ~30 per group often recommended for detecting small-to-moderate effects in general cognitive tDCS meta-analyses. We have determined this a priori effect size based on the existing work we published previously (Xu et al., 2023, J Exp Psychol Gen;152(4):1122-1133). In our pilot study (Xu et al., 2023), we identified a significant interaction effect between the single-session tDCS stimulation (active vs sham) and time (pre-test vs post-test) (t = 2.38, p = .02, n = 27; 95% CI [0.14, 1.49]) for changing procrastination willingness in the laboratory settings, indicating a medium effect size. Based on this specific empirical foundation, GPower indicated that a total sample size of 34 (17 per group) was sufficient to achieve 80% power (please see GPower output below). To account for potential attrition, we aimed to recruit 36 participants (18 per group), ultimately retaining 46 participants (23 per group) after exclusions. While we stand by this a priori justification, we fully agree with your overarching point that this remains a constraint.

      We completely agree with your observation regarding the plots. The apparent absence of data points in the previous versions of Figures 3B and 3F was not due to data exclusion, but rather to severe overplotting. Because multiple participants in the active neuromodulation group achieved identical scores (e.g., 0% procrastination rate or 100% task-execution willingness in later sessions), their data points perfectly overlapped, making it appear as though only ~10 points were present. As you helpfully suggested, we now employ jittered scatter plots with adjusted transparency, ensuring that all 23 individual data points per group are clearly visible, even when values are identical. As these revised figures demonstrate, the significant group differences reflect a consistent, cohort-wide shift rather than the influence of isolated outliers. Those

      As you rightly suggested, we have explicitly framed the small sample size as a major limitation and emphasized the necessity for large-scale replication. We have strengthened the wording in the Limitations section to explicitly mention the risk of outlier dependency and the need for larger cohorts.

      Legend Section (Page 28, Line 1161-1163)

      “… To ensure transparency and rule out outlier-driven effects, individual data points for all participants (N=23 per group) are overlaid on the bars using a jittered distribution to prevent overplotting of identical values.”

      Discussion Section (Page 13, Line 691-696)

      “… a major limitation of the current study is the relatively small sample size (total N = 46). While this was determined a priori based on our specific pilot study, small samples inherently bear a higher risk of being influenced by outliers and may overestimate effect sizes compared to large-scale meta-analytic expectations for tDCS. Therefore, these findings warrant caution in generalization and necessitate rigorous replication in larger, adequately powered cohorts.”

      Related to this, in the figure showing individual data points (3B/F), I count only around 10 data points per tDCS group for the 18 participants per group. I ask the authors to modify the plot that the data points from all participants can be seen (for example, by adding some noise on the x-axis for participants with the same value on the y axis).

      Thank you for this kind reminder. As we replied above, those plots have been redrawn by adding the jitters, which favor the readability as you kindly suggested.

      Another surprising aspect of the data is that repeated sessions of tDCS change procrastination behavior up to six months after stimulation. Do the authors think that their tDCS setup leads to such long-lasting neuroplastic changes, and if yes, can they cite prior work where similar dosages of tDCS also showed such long-lasting effects? Or could the results be explained by learning effects, for example because participants in the DLPFC group learned during the repeated tDCS sessions that it feels internally rewarding to finish one's tasks instead of procrastinating them, and they still benefit from this kind of "learned industriousness" 6 months later? In any case, in my view it is important to be more specific about how seven sessions of tDCS can affect behavior half a year later.

      We sincerely thank the reviewer for this highly insightful and thought-provoking comment. The concept of "learned industriousness" is particularly apt and captures a crucial alternative mechanism that we must address. We agree that explaining how seven sessions of tDCS can affect behavior half a year later requires a nuanced discussion of both neurobiological and behavioral learning mechanisms.

      Regarding the first point, we do believe that our multi-session protocol can induce long-lasting neuroplastic changes. While single-session tDCS effects are typically transient, cumulative neurobiological evidence demonstrates that repeated, multi-session protocols (typically ranging from 5 to 10 sessions) can induce activity-dependent, long-term potentiation (LTP)-like plasticity that consolidates over time (Agboada et al., 2020; Au et al., 2017; Jannati et al., 2023). Our 7-session protocol falls squarely within this range of "intensified dosing" designed to promote such consolidation. Meta-analyses and empirical studies on multi-session tDCS have shown that such protocols can produce behavioral and neurophysiological effects lasting weeks to months, particularly when targeting prefrontal regions involved in value-based decision-making and cognitive control (e.g., Brunoni et al., 2013; Sabé et al., 2024; Woodham et al., 2025).

      Furthermore, we completely agree with you for this alternative explanation regarding learning effects. It is highly plausible that participants in the active group, experiencing reduced task aversiveness and increased outcome value during the intervention, learned that completing tasks is internally rewarding. This aligns perfectly with the psychological concept of "learned industriousness" (Eisenberger, 1992), where the reinforcement of effortful behavior makes future engagement more likely. We explicitly acknowledge that repeated exposure to the experience-sampling protocol and the positive feedback of task completion could facilitate this kind of behavioral learning. More importantly, we argue that the learning effects are not bad things in this neuromodulation, and the learning effect and the neuroplasticity may be synergistic. The sham control group underwent the exact same experience-sampling protocol, reported real-life tasks, and had the identical opportunity for "learned industriousness" through feedback. However, as identified in the half-year follow-up, the sham group did not exhibit the same progressive improvement during the intervention, nor did they sustain a significant reduction in procrastination at the 6-month follow-up (their rates returned to near-baseline levels). This divergence suggests that while learning may play a role, the active neuromodulation likely provided the necessary neuroplastic "boost" (e.g., by enhancing prefrontal value-encoding circuits) that facilitated, accelerated, and consolidated this learning, making the behavioral change durable. Without the neuromodulatory enhancement, the mere exposure to the protocol was insufficient to produce long-term change.

      As you kindly suggested, we have explicitly incorporated this nuanced discussion into the revised manuscript, by citing relevant literature on multi-session tDCS plasticity, explicitly acknowledging the "learned industriousness" hypothesis, and reiterating the limitation of having only a single follow-up point.

      Discussion Section (Page 12, Line 627-634)

      “... Despite statistically supporting the TDM, we acknowledge that alternative neurocognitive mechanisms could contribute to the observed reductions in procrastination. For instance, repeated exposure to the experience-sampling protocol may have enhanced participants’ awareness of task progress or facilitated feedback-based learning, thereby increasing the subjective value of goal completion independent of DLPFC neuromodulation. Participants in the active group may have learned during the repeated sessions that completing tasks feels internally rewarding, thereby benefiting from a form of “learned industriousness” (Eisenberger, 1992) that persists months later.”

      Discussion Section (Page 14, Line 719-723)

      “... we explicitly note that a single 6-month follow-up timepoint cannot definitively establish the stability or trajectory of these effects. Future studies incorporating multiple longitudinal assessments (e.g., 1-month, 3-month, 6-month, 12-month) are required to substantiate claims about long-term retention and to disentangle the precise contributions of neuroplasticity versus behavioral learning.”

      Lastly, the link to the data repository works, but I could not inspect the data because I was asked to request access to the data, which I did not do in order to remain anonymous.

      Thank you a lot to take invaluable to review our data and code in this repository. As we reported previously, all the data and code to support those findings have been deposited in the eLife online submission system for your reviews and scrutiny before this manuscript is formally published. As the editorial policy of eLife on VOR (Version of Record) instructed, to prevent from mixture of codes and data across multiple round of revisions, those data and codes in the final version would be released once this paper is formally published. Please do not worry for the anonymity policy. This is a public peer review, and it thus enables those helpful comments that you kindly suggested to be public when this manuscript is formally published. Again, thank you to substantially contribute on this revised manuscript by sharing those helpful suggestions.

      Recommendations for the authors:

      Editors note: We encourage the authors to consider the remaining reviewer concerns and revise the manuscript accordingly.

      Thank you so much for this warm and kind reminder. We have addressed all of those concerns that remained by the new Reviewer #4, point-by-point. All the co-authors do appreciate you for handling our manuscript, and for contributing those fruitful and helpful comments. We do believe that the quality of this manuscript has been substantially improved, benefiting from this editorial process.

      References

      Agboada, D., Mosayebi-Samani, M., Kuo, M. F., & Nitsche, M. A. (2020). Induction of long-term potentiation-like plasticity in the primary motor cortex with repeated anodal transcranial direct current stimulation - Better effects with intensified protocols? Brain Stimulation, 13(4), 987–997. https://doi.org/10.1016/j.brs.2020.04.009

      Au, J., Karsten, C., Buschkuehl, M., & Jaeggi, S. M. (2017). Optimizing transcranial direct current stimulation protocols to promote long-term learning. Journal of Cognitive Enhancement, 1(1), 65–72. https://doi.org/10.1007/s41465-017-0007-6

      Brunoni, A. R., Boggio, P. S., Ferrucci, R., Priori, A., & Fregni, F. (2013). Transcranial direct current stimulation: challenges, opportunities, and impact on psychiatry and neurorehabilitation. Frontiers in Psychiatry, 4, 19. https://doi.org/10.3389/fpsyt.2013.00019

      Cole, E. J., Stimpson, K. H., Bentzley, B. S., Gulser, M., Cherian, K., Tischler, C., Nejad, R., Pankow, H., Choi, E., Aaron, H., Espil, F. M., Pannu, J., Xiao, X., Duvio, D., Solvason, H. B., Hawkins, J., Guerra, A., Jo, B., Raj, K. S., Phillips, A. L., … Williams, N. R. (2020). Stanford accelerated intelligent neuromodulation therapy for treatment-resistant depression. The American Journal of Psychiatry, 177(8), 716–726. https://doi.org/10.1176/appi.ajp.2019.19070720

      Eisenberger, R. (1992). Learned industriousness. Psychological Review, 99(2), 248–267. https://doi.org/10.1037/0033-295X.99.2.248

      Hutton, T. M., Aaronson, S. T., Carpenter, L. L., Pages, K., Krantz, D., Lucas, L., Chen, B., & Sackeim, H. A. (2023). Dosing transcranial magnetic stimulation in major depressive disorder: Relations between number of treatment sessions and effectiveness in a large patient registry. Brain Stimulation, 16(5), 1510–1521. https://doi.org/10.1016/j.brs.2023.10.001

      Jannati, A., Oberman, L. M., Rotenberg, A., & Pascual-Leone, A. (2023). Assessing the mechanisms of brain plasticity by transcranial magnetic stimulation. Neuropsychopharmacology, 48(1), 191–208. https://doi.org/10.1038/s41386-022-01453-8

      Ke, Y., Liu, S., Chen, L., et al. (2023). Lasting enhancements in neural efficiency by multi-session transcranial direct current stimulation during working memory training. npj Science of Learning, 8(1), Article 23. https://doi.org/10.1038/s41539-023-00200-y

      Sabé, M., Hyde, J., Cramer, C., Eberhard, A., Crippa, A., Brunoni, A. R., Aleman, A., Kaiser, S., Baldwin, D. S., Garner, M., Sentissi, O., Fiedorowicz, J. G., Brandt, V., Cortese, S., & Solmi, M. (2024). Transcranial magnetic stimulation and transcranial direct current stimulation across mental disorders: A systematic review and dose-response meta-analysis. JAMA Network Open, 7(5), e2412616. https://doi.org/10.1001/jamanetworkopen.2024.12616

      Schulze, L., Feffer, K., Lozano, C., Giacobbe, P., Daskalakis, Z. J., Blumberger, D. M., & Downar, J. (2018). Number of pulses or number of sessions? An open-label study of trajectories of improvement for once- vs. twice-daily dorsomedial prefrontal rTMS in major depression. Brain Stimulation, 11(2), 327–336. https://doi.org/10.1016/j.brs.2017.11.002

      Woodham, R. D., Selvaraj, S., Lajmi, N., Hobday, H., Sheehan, G., Ghazi-Noori, A.-R., Lagerberg, P. J., Rizvi, M., Kwon, S. S., Orhii, P., Maislin, D., Hernandez, L., Machado-Vieira, R., Soares, J. C., Young, A. H., & Fu, C. H. Y. (2025). Home-based transcranial direct current stimulation treatment for major depressive disorder: a fully remote phase 2 randomized sham-controlled trial. Nature Medicine, 31(1), 87–95. https://doi.org/10.1038/s41591-024-03305-y

      Zhong, M., Cywiak, C., Metto, A. C., Liu, X., Qian, C., et al. (2021). Multi-session delivery of synchronous rTMS and sensory stimulation induces long-term plasticity. Brain Stimulation, 14(4), 884–894. https://doi.org/10.1016/j.brs.2021.05.003

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This interesting paper demonstrates that transgenic over-expression of sphingosine 1-phosphate receptor 1 (S1PR1) on neutrophils alters their phenotype, resulting in (1) accumulation of neutrophils in blood, spleen, lung, and liver; (2) a shift in homing receptor expression with reduced CXCR2 and elevated CXCR4; (3) altered transcriptional profile with an increase in "G5c" neutrophils and reduced "module scores" for apoptosis and inflammatory response; (4) reduced ROS production upon fLMP stimulation; and (5) altered responses to bacterial and viral infections of the lung. It raises many interesting questions about how S1P signaling regulates neutrophil biology, and hence will be the basis of future studies. These include: (1) What is the physiological role of S1PR1 signaling in neutrophils? Although there is no dramatic effect on numbers upon S1PR1 loss, is there an effect on any of the other parameters measured? (2) What is unique about the lung that S1PR1 over-expression is particularly impactful there? (3) What distinguishes the bacterial context in which S1PR1 over-expression is maladaptive from the viral context in which S1PR1 over-expression is protective? and (4) Can treatment with an S1PR1 agonist mimic S1PR1 over-expression? As a possibly related question, when in neutrophil development does S1PR1 signaling function to shift the phenotype?

      Strengths:

      (1) A comprehensive characterization of S1PR1-transgenic neutrophils.

      (2) Opens many interesting areas of investigation.

      We thank the reviewer for a very positive assessment of our work and for raising very interesting questions, which will be useful to extend this work in the future.

      Weaknesses:

      Although some characterization of the neutrophil-specific Mrp8-Cre is done, most of the experiments use the more widely expressed LysM-Cre. The redistribution phenotype is much stronger with LysM-Cre than with Mrp8-Cre, so it is unclear what effects are attributable to a cell-intrinsic role of S1PR1, even in studies of neutrophils analyzed ex vivo.

      We acknowledge that the data from Mrp8-Cre mouse strain are more limited than the LysM-Cre counterparts. This is because we initially characterized the LysM-Cre S1pr1 KO and TG strains and confirmed key findings relevant to neutrophils in the Mrp8-Cre counterparts. Going forward, more studies will be done in the Mrp8-Cre strain as suggested by the reviewer.

      Reviewer #2 (Public review):

      The authors have utilised two main models to assess the function of S1PR1 in neutrophils in mice. The knockout of this receptor shows no conclusive effect on neutrophil numbers or functions; it was only the overexpression that resulted in significant alterations. Therefore, often the conclusions do not describe normal or disease physiology but could be useful in a bioengineering context.

      We agree with the reviewer that some of the key findings described in our manuscript, for example, neutrophil survival, spleen size, and ROS reduction, etc., were not observed in the S1pr1 KO strains. Our interpretation is that other receptors, for example S1PR4, could be involved in compensating for the loss of S1PR1. We will explain this better in the revisions.

      Strengths:

      From a bioengineering standpoint, this seems like an important study - showing enforced expression of S1PR1 in neutrophils has improved outcomes for influenza infection (Figures 6 and 7).

      We agree with the reviewer that overexpression of S1PR1 could be useful from “bioengineering standpoint”.

      Weaknesses:

      Although the strength is the influenza model, genetic modification of human neutrophils cannot be a strategy, and therefore, is there any way to increase this receptor for mouse, or more importantly, human neutrophils? This study only looks at mice with a non-physiological model of overexpression. It does not offer a real therapeutic option, which drastically hinders the importance of the study. I have other concerns with the data analysis and interpretation, which I detail on a figure-by-figure basis (and how it relates to conclusions) below:

      The reviewer's comments are acknowledged. However, at this early stage of discovery, we feel that it would be premature to address the issue of “a real therapeutic option”. This can be addressed in the future.

      Main specific issues:

      (1) Figure 2A+B: This is unconvincing; in the surface staining there seem to be real cells positive for the receptor (high staining in the histogram), but none of the transgenic protein is getting there? This undermines the idea that the effects of the transgene are related to S1P signalling. In the 'Total S1PR1' this is both underwhelming and misleading, as an isotype control (or better S1PR1 knockout) is missing, which would give a better representation of actual expression (flow cytometry autofluorescence famously increases in the red laser channels with fix/perm). The Imagestream chosen images are showing best-case scenarios - and aren't representative. What does the isotype/ KO look like here? All in all, the conclusion on receptor internalization is not well supported, especially when theoretically the TG overexpression should overload S1P availability. This also highlights the lack of another control - does overexpression of another random/non-functional protein have the same effect? To play devil's advocate, perhaps overloading of the ubiquitin-proteasome system is responsible?

      We will conduct S1PR1 antibody staining with knockout neutrophils as requested by the reviewer. The results will be shown in the revision. The comment of the reviewer on “ubiquitin-proteosome” system is not relevant in our opinion and does not impact the validity of our findings. The approach that we used is standard mouse genetics, and the results should be interpreted from that perspective and not from “devil's advocate”.

      (2) Figures 2E-H: In the text, the authors should fix the statement 'Additionally, surface CXCR2 was downregulated and CXCR4 upregulated in LysM-S1pr1 TG neutrophils across bone marrow, spleen, and blood (Fig. 2, E and F)' to better reflect that there is no significant difference in the bone marrow regarding CXCR4. Of note, the total MFI from this data would also be informative, another noticeable absence being the gating strategies for much of the data. Also, alter the statement: 'CD62L expression was largely preserved across compartments, with only a modest reduction in bone marrow neutrophils (Fig. 2G)'. A 50% reduction in CD62L is not modest.

      These statements will be changed in the revised manuscript, which we hope to submit soon.

      (3) Supplemental Figure 3. A common theme: the wrong statistics have been used here, which has led to a false conclusion. Megakaryocyte/erythrocyte progenitors (MEPs) were only elevated in 2/3 TG mice, and the numbers are so small that this is not significant by any measure of the word. This is certainly not statistically significant if the correct test of (log-normalized) two-way ANOVA is performed (with Sidak's post hoc test). Another acceptable test would be Kruskal-Wallis with Dunn's post-test just for MEPs.

      We will reanalyze these data with different statistical tests and discuss these in the revision.

      (4) Starting at Figure 3, the authors refer to 'S1PR1hi neutrophil accumulation'. Crucially, the authors must here and throughout be explicitly clear in which cells they are referring to, as this can be misleading - particularly as there are real S1PR1-high cells identified in Figure 2A surface staining. It is my understanding that the authors here mean the transgenic artificially high mice - a very large distinction.

      We will clarify and edit this terminology throughout the manuscript in the revision. S1PR1<sup>hi</sup> does not strictly imply receptor on the cell surface. It also includes internalized S1PR1 receptors.

      (5) Figure 3A: It is difficult to interpret the figure with the necessary details about the experiment. For instance, there is no mention that this is sterile inflammation or what caused it.

      This will be edited in the revision.

      (6) Figure 3B and C: It should be made clear whether these splenic neutrophils are related to the time course of peritoneal inflammation in 3A. Why are there so many apoptotic neutrophils in the spleen? The low numbers here suggest a processing issue rather than real death in vivo (which usually is absent).

      This will be addressed in the revision.

      (7) Figure 3D: This can also be misleading - the wrong statistics are again used. This should be a log-transformed two-way ANOVA. Regardless of this, the data is not strong enough to be conclusive, a minor effect at best that could also just be related to the type of cell tracker used.

      Same as #3 above.

      (8) Figure 5E: It is stated that 'LysM-S1pr1 TG mice exhibited a higher bacterial burden in the lungs than controls (Fig. 5E).' Again, misleading results, first the wrong statistical test was used (correct = log norm one-way ANOVA with Tukey's or Kruskal Wallis with Dunn's), secondly the only significance is between S1PR1(fsf) and the Mrp8-S1PR1, not with the LysM TG. 5F is also not strong, with only 2/7 values appearing outside the range of the control - P values can be misleading when poor statistics are used.

      Same as #3 above.

      (9) Figure 6G: Some discussion should be given for why Neutrophils are lower in BALF in the IAV model - even though higher in the lung in the non-IAC mice in Figure 1. In general, rather than focusing on the non-physiological differences, the discussion could better reflect the inconsistencies and more fully address the difference between the TG and KO and what this means going forward.

      Same as #5 above.

      Reviewer #3 (Public review):

      Summary:

      Using mice that overexpress S1PR1 in myeloid cells or specifically in neutrophils, the authors show that increased S1PR1 promotes neutrophil release from the bone marrow and accumulation in blood and peripheral tissues without causing baseline tissue injury. These cells acquire a CXCR4-high, CXCR2-low, CD101-low phenotype, survive longer, and display enhanced mitochondrial metabolism and mTOR signaling, together with reduced apoptotic, inflammatory, and ROS-related programs. Although phagocytosis is preserved, ROS production is markedly reduced. This is associated with impaired bacterial clearance in the lung but improved outcomes during influenza infection, including better survival, less weight loss, improved oxygenation, lower viral burden, and reduced lung inflammation. In contrast, myeloid S1PR1 deletion produces little detectable phenotype. The authors therefore propose that S1PR1 separates neutrophil persistence from inflammatory function, improving tolerance to viral lung injury at the expense of antibacterial defense.

      Strengths:

      This is a technically solid paper using novel mouse models to overexpress S1PR1 specifically in myeloid cells as well as neutrophils. The data are striking with respect to neutrophil expansion. The diverse roles of neutrophils and their population heterogeneity are an important scientific area that has led to many recent breakthroughs - PMC11785525; PMC12823425, thus this is a timely study.

      We thank the reviewer for positive comments.

      Weaknesses:

      The study mainly demonstrates what S1PR1 overexpression is sufficient to do, rather than establishing the physiological role of endogenous S1PR1. The conclusions should therefore be narrowed unless the authors provide stronger loss-of-function and physiological validation. As written, the abstract ("S1PR1 promotes mitochondrial fitness, enhances survival, and reduces inflammatory output") and the conclusion ("S1PR1 serves as a key regulatory axis") are sufficiency claims but should not be promoted as necessity claims. The honest sentence is: "Thus, a better conclusion would be that enforced S1PR1 expression is sufficient to reprogram neutrophils".

      We will edit the text to better reflect sufficiency versus necessity terms.

      The authors do not confirm efficient S1pr1 deletion in neutrophils. Furthermore, the knockout is examined only under steady-state conditions and limited in vitro stimulation, but not in the bacterial or influenza models where the transgenic phenotype is observed. Without these experiments, the study cannot establish whether endogenous S1PR1 is necessary for the reported functions.

      We will provide qRT-PCR data (that we have done already but did not include in the original version) confirming efficient deletion of the S1pr1 gene.

      The degree of S1PR1 overexpression is not quantified relative to normal physiological levels. The authors should determine whether naturally occurring S1PR1-high neutrophils display the same survival, metabolic, trafficking, and inflammatory features observed in the transgenic cells.

      We agree that the level of S1PR1 overexpression in our transgenic neutrophils should be quantitatively compared with endogenous S1PR1 expression. We will provide the quantitative result in the revision. Our analyses of independent human (GSE216009) and mouse (GSE243466 and GSE266518) single-cell RNA-seq datasets indicate that endogenous mRNA expression for this receptor is higher in specific neutrophil states associated with inflammatory, immature, and tissue-associated populations. Notably, these endogenous S1PR1-expressing neutrophils show transcriptional features involving altered oxidative/inflammatory programs, chemotaxis, and mitochondrial/metabolic regulation that partially overlap with the phenotype of S1PR1-transgenic neutrophils. We will provide additional quantitative and dataset analyses in the revised manuscript.

      Analysis of relevant human or mouse datasets, including sepsis, ARDS, viral infection, cancer, or aging, would also help establish whether this neutrophil state exists physiologically.

      As described above, we have analyzed independent human and mouse single-cell RNA-seq datasets from sepsis, cancer, and other inflammatory conditions to determine whether endogenous S1PR1-expressing neutrophil states occur naturally across biological and disease contexts. These analyses support the presence of distinct endogenous S1PR1-expressing neutrophil populations across multiple biological contexts. We will provide the expanded analyses and corresponding data in the revised manuscript.

      Surface S1PR1 expression appears similar between control and transgenic neutrophils, whereas total intracellular receptor is increased. This suggests that the phenotype may depend on receptor internalization or endosomal signaling. An internalization-deficient S1PR1 model, such as S1P1-S5A, would help distinguish sustained surface signaling from internalization-dependent signaling. The authors should also determine whether the phenotype requires ligand binding, Gi signaling, and mTOR activity.

      We agree that future experiments will use internalization-defective S1PR1 S5A knock-in neutrophils.

      The reduction in CXCR2 and decreased neutrophil accumulation in the airways could alone explain the protection from influenza-induced lung injury. The current experiments do not clearly distinguish neutrophil reprogramming from defective migration into the alveolar space.

      We agree. We did not examine the role of CXCR2 in the influenza experiments. Our interpretation is that reduced ROS from transgenic neutrophils reduced lung injury.

      Although this may be outside the scope of the current study, the authors should directly test whether CXCR2 inhibition reproduces the phenotype.

      This is a good point that the reviewer brought up. This can be addressed in a future study.

      The reported reduction in viral load should also be confirmed using plaque assay or TCID50, and the possible contribution of NET formation should be examined.

      This is again a good suggestion that can be addressed in a future study.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      This is an interesting paper, with the primary finding being that localizing MreB or PBP2 to the cell poles in E. coli is primarily demonstrated via an aggregate formed by expressing M. xanthus MreB.

      I have 2 main concerns:

      (1) First, the authors should clarify and adjust their interpretation of FDAA incorporation: As written, the authors interpret FDAA incorporation as being caused by incorporation during PG polymerization. However, in E. coli, FDAA incorporation does not result from the elongation of PG strands or their initial 4-3 crosslinking by DD transpeptidases following polymerization, but rather by the remodeling of L,D-transpeptidases.

      Thus, it is not accurate to refer to the FDAA incorporation as PG elongation, but rather the modification of crosslinks from 4-3 to 3-3 crosslinks at that location. To claim a link to PG polymerization, other experiments or substantial explanations are needed.

      (E. coli Cells Incorporate FDAAs by L,D-TPases in a Growth Independent Manner) - https://doi.org/10.1021/acschembio.

      We appreciate the reviewer for bringing up this important question. We will address this concern by further explaining our results in the revised manuscript.

      However, we respectfully disagree with the reviewer for two reasons:

      First, the paper by Kuru et al. (mentioned by the reviewer) revealed that (exact quote): “Our in vitro and in vivo data unequivocally demonstrate that these bacteria incorporate FDAAs using two extra cytoplasmic pathways: through activity of their D, D-transpeptidases, and, if present, by their L, D-transpeptidases…These mechanistic findings enabled development of a new, FDAA-based, in vitro labelling approach that reports on subcellular distribution of muropeptides, an especially important attribute to enable the study of bacteria with poorly defined growth modes” (Kuru et al., 2019).

      Thus, the paper by Kuru et al. identified two FDAA labeling patterns, a D, D-transpeptidase-dependent, concentrated labeling for PG growth and an L, D-transpeptidase-dependent, growth-independent labeling for PG modification along the entire cell envelope (Kuru et al., 2019). The polar FDAA foci we presented do not match the reported pattern of L, D-transpeptidase-dependent incorporation.

      Second, we provided the evidence in Fig. 3b that the polar FDAA labeling is due to the activity of PBP2, a D, D-transpeptidase, because mecillinam that inhibits PBP2 is sufficient to abolish polar FDAA incorporation.

      (2) Second, the evidence provided that this is polar elongation is not sufficient to prove elongation. In all images claiming polar growth, the FDAA focus appears as a single spot, which corresponds to the MreB aggregate visible in bright-field. That might indicate incorporation, but it does not demonstrate polar elongation. To prove this, the authors should do different-length pulses of FDAAs and demonstrate an increasing length of labeled PG along the cell. Cells with one focus should show increasing length; polar foci should elongate from both ends. I find the current 2-color labeling insufficient, as BADA labels the entire cell.

      We appreciate this comment and totally agree with the reviewer. The wide BADA labeling band could indeed come from L, D-transpeptidases. We will follow the reviewer’s recommendation to address this comment with additional staining experiments.

      Small points:

      (3) Lines 350- 302: "In this case, as the nonpolar region is no longer the growth zone, the established cylindrical PG structure is sufficient to maintain cell width, where MreB filaments become nonessential." Does polar elongation give robustness to rod shape? The authors should include an analysis of cell width and its variation within a cell and between cells.

      We appreciate this comment. Judging from the bright-field images we presented, we believe that polar elongation does give robustness to rod shape. We will follow the reviewer’s recommendation and provide the said analysis.

      (4) In the abstract: "This reprogrammed growth mode bypasses the requirement for MreB filaments, highlighting a plasticity of the Rod system that suggests polar elongation may have emerged through the evolutionary loss of MreB." This argument should not be made without evolutionary analysis or reference to such work indicating this is the case.

      We appreciate this comment and will remove this statement from our revised manuscript.

      (5) 365-367 "PG-depleted spheroplasts can spontaneously regenerate rod shape through curvature-dependent localization of MreB filaments [21, 57]". This should be amended. The Billing paper did indeed study spheroplasts, but the Hussain paper used teichoic acid-depleted cells that still had a cell wall.

      We thank the reviewer for pointing out this mistake. We will amend this statement in our revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Based on observations of localisation of MreBEc at the poles within an aggregate-like structure, upon heterologous expression of MreBMx, the authors set out to investigate how this non-canonical localisation of MreBs leads to a reprogramming of peptidoglycan synthesis to the poles. This is analogous to the polar growth observed in phyla which are not dependent on dispersed growth of PG, but only at the poles, and are MreB independent.

      The authors proceed to establish that PG synthesis is MreB-dependent, Rod enzyme-dependent, and requires the prior establishment of a pole.

      Strengths:

      (1) It is a very interesting idea to design experiments to demonstrate reprogramming of non-polar to polar growth based on the observation of localisation of a heterologously expressed MreB.

      (2) The experiments to demonstrate the factors that determine polar growth and the observation of the PG in each of these experimental situations are convincing.

      (3) I find the observation of an extra layer of PG in the heterologously expressed system very intriguing. It will be interesting to see if this layer merges with the other PG layer at some stage or branches from the non-polar growth near the poles.

      Weaknesses:

      (1) It is not clear what exactly the identity of the polar aggregates is and how much of this activity is an artefact of partially functional MreBs.

      We appreciate this comment and will address it by further explaining our results in the revised manuscript. While we can only say that the polar aggregates resemble inclusion bodies, we do believe that they cause polar PG growth because in the cells that express MreB<sub>Mx</sub>, polar PG growth does not occur at the poles that lack MreB aggregates.

      (2) I find it intriguing that the localisation and growth are predominantly at one pole only. It is unclear to me how this can be reconciled with growth and shape maintenance, and an increase in length and width. Is the increase in length and width a consequence of misshapen cells that are bulged in the absence of a normal PG layer?

      We appreciate this comment. Judging from the bright-field images we presented, we believe that polar elongation does not generate bulges and is thus sufficient for maintaining rod shape. We will follow the reviewer’s recommendation to clarify this.

      (3) The authors do not follow up on the observations in the first figure on the length and width changes and the extra peptidoglycan layer (which I feel are the most interesting aspects), and how this can be connected to the polar growth observed in the later sections of the manuscript.

      We appreciate this comment. We believe that the thickened PG patches are integral parts of the polar PG, rather than an extra layer, which is, however, technically challenging to prove. Thus, we will relay on fluorescence microscopy to visualize polar PG growth.

      (4) The claim that this could be a precursor of an MreB-independent polar growth mechanism appears to be a bit far-fetched, because the system is still dependent on having an established pole for PG synthesis to occur in the new place.

      We appreciate this comment and will remove such speculations from our revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Most rod-shaped bacteria grow by one of two mechanisms: growth from the pole or growth from the midcell. It is rare for a single species to utilize both modes of growth, although a few examples do exist. Here, the authors have artificially induced E. coli cells to grow from the poles, either by expressing mreB from Myxoccocus xanthus in E. coli, leading to the mislocalization of MreB to the poles in large aggregates, or by forcing the localization of major cell wall synthesis proteins to the cell pole. The fact that cells switched modes of growth suggests an evolutionary pathway from midcell to polar growing cells as well as suggests that there might be unknown conditions in nature when cells may switch growth modes.

      Strengths:

      (1) The authors use a strain that has replaced mreB with a functional fluorescent version at the native site. This eliminates any effects of having two copies of mreB. Because MreB is fluorescently tagged, they can monitor its localization when mreB from M. xanthus is expressed in E. coli. They notice that MreBec now forms bright polar foci and that there appear to be changes to the cell wall at the pole.

      (2) D-amino acids are specific to the cell wall, and fluorescent versions (FDAA) have been used to mark sites of new cell wall insertion. The authors use these FDAAs to determine how cell wall synthesis correlates to MreB and if that changes when MreBmx is expressed. Again, there is pretty clear evidence that cell wall synthesis follows MreB localization to the pole.

      (3) MreB itself does not synthesize the cell wall, but localizes the proteins, such as PBP2, that do. Using a published method to force proteins to the pole, the authors show that when they target PBP2 to the pole, they can phenocopy the polar growth seen when MreB is polar. Interestingly, these cells become resistant to A22, a drug that targets MreB, suggesting that localized growth at the pole does not require MreB and is sufficient to maintain rod shape.

      Weaknesses:

      (1) The authors do not show what the poles of control cells look like, making it difficult to determine if there is a change when MreBmx is expressed. However, the localization of both MreBec and MreBmx clearly forms bright foci at the pole.

      We appreciate this comment and totally agree with the reviewer. We will provide reference images in the revised manuscript.

      (2) While more quantification is needed, the authors show some evidence that RodZ, an MreB interaction partner, is needed for this polar growth, as cells lacking rodZ still form foci at pole-like regions when MreBmx is in the cell; however, these cells remain spherical and do not elongate from these foci.

      These results strongly support our conclusions. In the cells that express MreB<sub>Mx</sub>, MreB<sub>Mx</sub> causes MreB<sub>Ec</sub> to mislocalize, and mislocalized MreB<sub>Ec</sub> recruits Rod enzymes (RodA and PBP2) to poles through the connector protein RodZ. Here when we delete rodZ, while MreB<sub>Ec</sub> still forms aggregates, it is unable to recruit Rod enzymes to those aggregates.

      In contrast, when we directly relocalize PBP2 to cell poles through the PopZ tag, polar PG growth can bypass the requirement for MreB<sub>Ec</sub> filaments (Fig. 4).

      When MreB is deleted, and cells become spherical, the authors were unable to cause the polar growth mode. They suggest that this is due to the lack of a preexisting pole; however, experimental evidence to test this is missing.

      We appreciate this comment. We plan to localize PBP2 to cell poles in an mreB depletion strain to test if cells still remain rods when mreB is depleted.

      Conclusion:

      Overall, the authors do a good job of showing that E. coli can grow with a polar method rather than a midcell method of cell wall insertion. It is unclear why MreBec forms at poles when MreBmx is present and even if this MreB is functional. The foci look similar to inclusion bodies, which are normally aggregates of misfolded proteins that migrate to the poles. Past work has shown that when MreB is more polarly localized, branches form, which is not seen here. Importantly, the authors also show that there is feedback between the localization of MreB and PBP2 as both appear to regulate the localization of the other.

      Because CFP-labeled MreB<sub>Ec</sub> is still fluorescent in the aggregates, we believe that some MreB<sub>Ec</sub> molecules are still correctly folded there, at least on the surfaces of those aggregates.

      Reference

      Kuru, E., Radkov, A., Meng, X., Egan, A., Alvarez, L., Dowson, A., Booher, G., Breukink, E., Roper, D.I., Cava, F., Vollmer, W., Brun, Y., and VanNieuwenhze, M.S. (2019) Mechanisms of incorporation for D-amino acid probes that target peptidoglycan biosynthesis. ACS Chem Biol.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study leverages a large global dataset of tens of thousands of tuberculosis samples to place recurrent protein-coding mutations into their three-dimensional structural context, offering an expanded view of how antibiotic resistance emerges compared to traditional genetic analyses alone. The strength of evidence is convincing, supported by the scale and breadth of the dataset and the systematic structural analysis, although some of the assumptions made in the the modeling approach are only partially supported. Overall, the work will be of broad interest to researchers studying microbial evolution, antibiotic resistance, and structure-function relationships in pathogens.

      We thank the reviewers and editors for their careful critique of our work. We believe the work has been strengthened by addressing the comments and are delighted to submit a revised version. This version has a detailed discussion of prior literature on structure analysis of antibiotic resistance variants in Mycobacterium tuberculosis, more details on dataset origins and data processing to improve reproducibility, better explanations of the evolutionary assumptions underlying the scoring method we developed, and more discussion of proteins that show homoplasic and clustered mutational signals that are not known to confer antibiotic resistance. We include detailed responses to the comments below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Green et al. attempt to use large-scale protein structure analysis to find signals of selection and clustering related to antibiotic resistance. This was applied to the whole proteome of Mycobacterium tuberculosis, with a specific focus on the smaller set of known antibiotic-resistance-related proteins.

      Strengths:

      The use of geospatial analysis to detect signals of selection and clustering on the structural level is really intriguing. This could have a wider use beyond the AMR focussed work here and could be applied to a more general evolutionary analysis context. Much of the strength of this work lies in breaking ground into this structural evolution space, something rarely seen in such pathogen data. Additional further research can be done to build on this foundation, and the work presented here will be important for the field.

      The size of the dataset and use of protein structure prediction via AlphaFold, giving such a consistent signal within the dataset, is also of great interest and shows the power of these approaches to allow us to integrate protein structure more confidently into evolution and selection analyses.

      Weaknesses:

      There are several issues with the evolutionary analysis and assumptions made in the paper, which perhaps overstate the findings, or require refining to take into account other factors that may be at play.

      (1) The focus on antimicrobial resistance (AMR) throughout the paper contains the findings within that lens. This results in a few different weaknesses:

      (a) While the large size of the analysis is highlighted in the abstract and elsewhere, in reality, only a few proteins are studied in depth. These are proteins already associated with AMR by many other studies, somewhat retreading old ground and reducing the novelty.

      (b) Beyond the AMR-associated proteins, the proteome work is of great interest, but only casually interrogated and only in the context of AMR. There appears to be an assumption that all signals of positive selection detected are related to AMR, whereas something like cas10 is part of the CRISPR machinery, a set of proteins often under positive selection, and thus unlikely to be AMR-related.

      We agree that environmental pressures beyond AMR may impose positive selection. In response to the reviewer’s comment we have now included more results and a supplementary figure of the findings about Cas10. We have expanded our results about the proteins found with significant clustering that are not known AMR-associated proteins, and clarified in the discussion that we don’t believe positive selection is caused only by AMR.

      We note that data from homoplasic substitutions in proteins (Figure 1) does indeed support that AMR is the strongest driver of positive selection in Mtb. A challenge is that knowledge about protein function varies in depth by protein, and in Mtb the AMR proteins are among the most well-studied in the proteome. Thus, explanations from literature are most readily available for AMR proteins. Moreover, proteins may exhibit positive selection for more than one reason – for example, we find significant clustering in the proteins GlmM, GlmS, GlmU, and MurA, all of which are involved in amino sugar metabolism, a key component of the cell wall. These proteins could plausibly have a role in AMR via cell wall permeability mechanisms, and could plausibly have a role in adaptation to host environment via the same (or other) mechanisms.

      (2) The strength of the signal from the structural information and the novelty of the structural incorporation into prediction are perhaps overstated.

      (a) A drop of 13% in F1 for a gain of 2% in PPV is quite the trade-off. This is not as indicative of a strong predictor that could be used as the abstract claims. While the approach is novel and this is a good finding for a first attempt at such complex analysis, this is perhaps not as significant as the authors claim.

      (b) In relation to this, there is a lack of situating these findings within the wider research landscape. For instance, the use of structure for predicting resistance has been done, for example, in PncA (https://academic.oup.com/jacamr/article/6/2/dlae037/7630603, https://www.sciencedirect.com/science/article/pii/S1476927125003664, https://www.nature.com/articles/s41598-020-58635-x) and in RpoB (https://www.nature.com/articles/s41598-020-74648-y). These, and other such works, should be acknowledged as the novelty of this work is perhaps not as stark as the authors present it to be.

      We appreciate this comment and have made efforts to better situate our work in the wider context of structure-based prediction of antibiotic resistance. We have included description of and citation to these works and others in the introduction, results, and discussion. A differentiator between our work and previous is that we have trained a predictor across all proteins in the WHO catalogue of known resistance variants, not limited ourselves to a single protein at a time. We feel this makes the case that structure (and specifically proximity to known resistance-conferring variants) is a universally useful feature for resistance mutation prediction. We have also updated our abstract to reflect the exact performance of our method.

      Introduction: “Protein three-dimensional structure has shown utility as an input feature for identifying resistance-conferring variants in known resistance-conferring proteins such as RpoB,[26,27] PncA [28–30], and AtpE.[31]While past work has sought to reannotate parts of the M. tuberculosis proteome with computationally predicted protein structures using older structure prediction methods,[32] we can now infer a protein structure for nearly every protein in the proteome using AlphaFold,[33] leading to new works examining the 3D location of mutations in known and suspected resistance-conferring proteins.[34,35]”

      (3) The authors postulate that neutral AA substitutions would be randomly distributed in the protein structure and thus use random mutations as a negative control to simulate this neutral evolution. However, I am unsure if this is a true negative control for neutral evolution. The vast majority of residues would be under purifying selection, not neutral selection, especially in core proteins like rpoB and gyrA. Therefore, most of these residues would never be mutated in a real-world dataset. Therefore, you are not testing positive selection against neutral selection; you are testing positive against purifying, which will have a much stronger signal. This is likely to, in turn, overestimate the signal of positive selection. This would be better accounted for using a model of neutral evolution, although this is complex and perhaps outside the scope. Still, it needs to be made clear that these negative controls are not representative of neutral evolution.

      The goal of our negative control was to simulate the random accumulation of amino acid substitutions without the effects of selection, which we had originally referred to as “neutral evolution” but is better described as “randomly accumulating substitutions.” We agree that in the absence of antibiotics, essential proteins like RpoB and GyrA are probably under purifying selection, and thus will have depletion of mutations in their hydrophobic cores. We have revised our wording in the results section “A protein-level statistic to test for mutational clustering” to make it clear that our negative control is that of randomly accumulating substitutions. We have included possible extension to more realistic evolutionary scenarios in the discussion.

      As a side note, if we were to compare the observed mutation 3-D pattern to a purifying selection model, that may overestimate the signal compared with the randomly accumulating substitution model that we currently present in the manuscript. Purifying selection would tend to result in slower evolutionary rates than positive selection or randomly accumulation substitutions. So, a control based on purifying selection would have to have fewer mutations to account for the same evolutionary time, and this could lead to underestimating clustering.

      (4) In a similar vein, the use of 15 Å as a cut-off for stating co-localisation feels quite arbitrary. The average radius of a globular protein is about 20 Å, so this could be quite a

      large patch of a protein. I think it may be good to situate the cut-off for a 'single location' within a size estimator of the entire protein, as 15 Å could be a neighbourhood in a large protein, but be the whole protein for smaller ones.

      We interrogated the use of 15 Å as a cutoff and found that it is indeed not very stringent, and functions more as a filter to remove the most egregious examples of proteins lacking single-location clustering. We include a new supplementary figure showing the number of significant hits as the cutoff is varied from 2 Å to 40 Å, and summarize these results in the main text. We note that we in fact find a weak negative relationship between protein length and the distance between the top two residues with highest G-score (R<sup>2</sup> = 0.008, b = -3.1806, p-value = 0.048), the opposite of what would be expected under a scenario the distance between residues is simply driven by protein size and not a signal for clustering.

      Reviewer #2 (Public review):

      Summary:

      This is an important study that, for the first time, systematically places the homoplastic genetic variation observed in the coding regions in a large collection of >31,000 M. tuberculosis samples into the protein structural context. This should be much more informative when, e.g. predicting antimicrobial resistance. The authors imaginatively apply the Getis-Ord score, which originated in geographical spatial analysis but has also been used in human disease to demonstrate that missense mutations in M. tuberculosis known to be associated with antimicrobial resistance are clustered in space. That they are able to consider almost all of the proteome using a large dataset of 31,000 M. tuberculosis complex clinical samples, which makes the evidence convincing.

      Strengths:

      To my knowledge, this is the first study to place the homoplastic missense mutations from a large clinical dataset into their protein structural context and attempt to look for clustering in space, which could be indicative of a recent evolutionary pressure, such as the use of antibiotics. The field usually only views resistance through the genetic paradigm, so it is delightful to see a structural paradigm being brought to bear, as this should, in theory, be much more informative, as protein structure is much closer to function. In addition, the dataset used is large (>31,000 clinical M. tuberculosis samples), and the authors are able to consider almost all of the ORFs (3,687/3,996) in the M. tuberculosis reference, and hence the analysis is comprehensive.

      Weaknesses:

      It is not apparent at the time of this review if the study could be reproduced by other researchers as e.g. whilst the authors state that the raw sequencing files (FASTQ) underpinning the dataset of 31,428 M. tuberculosis isolates can be downloaded the table in the Supplement containing the sample and accession identifiers contains rows that do not contain NCBI accessions e.g. '01R0685' or 'IDR 1600023875' or '1479144813357T181715lib5022nextseqn0035151bp' instead of the expected form e.g. 'SAMEA1016138'. I have searched the NCBI SRA using these terms and got no results, so they cannot be used to download any FASTQ files. There is also no information in the preprint on how the reads were processed (which is a complex process) and the dataset of SNPs subsequently built. One can trace back through the references, but I cannot find anywhere where one can download the SNP dataset, which would permit researchers to reproduce at least the latter stages of the work -- one obvious option would be to make the SNP dataset available. Likewise, the authors have constructed a "M. tuberculosis structureome", which would be very useful for the community but does not appear to be publicly available. At the time of the review, not all the GitHub repositories were public, so these points may have been rectified when that was corrected.

      We have made a number of changes to improve reproducibility of the manuscript.

      First, we have updated Table 1 with additional information to reflect the dataset of origin. While most of the isolates used are available from NCBI (94.5%), the remainder are from other sources. An additional 4.8% are exclusively from PATRIC (now the BV-BRC) and 0.4% are from ENA. Some of the isolates were originally named by their internal identifiers, not their NCBI BioSamples, which has been rectified. Of the three identifiers the reviewer cites, two were originally from Reseq-TB and are now listed with their NCBI BioSamples, and the third, 01R0685, corresponds to one isolate deposited with others in a single BioProject (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB26000; individual isolate available here: https://www.ncbi.nlm.nih.gov/biosample/10125872). We have clarified in the table metadata that the relevant search identifier.

      Second, we have added a new section to our methods about the origins of the Mtb genomic data used in our manuscript, and describing the process of variant calling and SNP dataset construction.

      Third, we have provided the data in a zenodo repository along with instructions for using the protein distance map files provided: 10.5281/zenodo.20766453

      Lastly, we have ensured that the github is publicly available.

      The authors correctly point out in the Introduction that supervised methods like GWAS or ML need datasets with matching genetic and phenotypic drug susceptibility data, which are much difficult/expensive to obtain, but don't then close the loop by comparing their results back to such supervised methods. They pick out RnJ as having previously been identified by a GWAS, but it would have provided a useful validation of their method to e.g. demonstrating that X% of the genes they identify were also identified by GWAS/ML studies, and therefore their method can achieve similar results but without having to collect pDST data.

      We agree that this is a compelling possible extension of the work but due to time constraints have chosen not to pursue it for this manuscript.

      Whilst the authors acknowledge that assuming all sites are equally likely to mutate in their random shuffling procedure is a shortcoming, a bigger weakness is, I suspect, that one should also only consider which amino acids could arise at each codon due to a SNP. Shuffling assumes any amino acid can arise at any codon which is only possible with multiple nucleotide changes, which is possible but highly unlikely.

      Our approach is based on analyzing, in the wild-type protein structure, the 3D location at which mutations occur. In this calculation we do not consider the identity of the amino acid change per se. We have now clarified in the methods that we are not explicitly simulating biochemical change of the wild-type amino acid to any given mutant, rather we are analyzing the wild type amino acid in its structural context. We have added the following text in the methods section: In Computing inter-residue distances, “The EVcouplings Python package was used to compute the distance between wild-type amino acid residues in all protein structures”; in Computing the Getis-Ord score for clustering of homoplastic mutations, “The two values input to the Getis-Ord statistic computation are a per-residue score x, here the per-amino acid homoplasy score, and a weight matrix W that contains the inverse of the inter-residue distances computed from the wildtype amino acids”; and in Preparing GeO score calibration data, “Note that we do not recompute inter-residue distances when simulating mutations in an amino acid, as the distances used as input to GeO score are the wild-type inter-residue distances.” We hope this addresses the reviewer’s concern.

      Finally, the authors implicitly assume that the mutations do not perturb the structure of the proteins, which is likely to be generally true for essential genes but less likely to be true for non-essential genes. This assumption underpins their entire approach and should be borne in mind when evaluating the results.

      Our approach is based on analyzing, in the wild-type protein structure, the 3D location at which missense mutations occur and are observable in a naturally evolving population. It is true that we have not undertaken an analysis of whether any given mutation does or does not perturb the protein structure in which it occurs. However, location alone is a useful piece of information to analyze, as the location of naturally occurring mutations gives a readout of what types of mutations are allowed to persist under natural selection. Among our findings is that mutations display significant clustering even in non-essential genes, which we address in our discussion, “for proteins where mutations that lead to loss of function are known to cause resistance, such as PncA and RsmG (GidB), it is not necessarily expected to find clustering of mutations. We suspect that the observed clustering is due to mutations in a certain region of the protein being more likely to cause loss of function.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It would be good if a more detailed description of the sequencing dataset origins were presented. The Supplementary Table 1, which is meant to hold these data, points to another paper, which in turn points to another paper which does not detail collection strategies for this data. This needs to be clear so that any bias in collection, which could over-inflate selection signals, can be assessed.

      [From public reviews]

      First, we have updated Table 1 with additional information to reflect the dataset of origin. While most of the isolates used are available from the NCBI (94.5%), the remainder are from other sources. An additional 4.8% are exclusively from PATRIC (now the BV-BRC) and 0.4% are from ENA. Some of the isolates were originally named by their internal identifiers, not their NCBI BioSamples, which has been rectified. Of the three identifiers the reviewer cites, two were originally from Reseq-TB and are now listed with their NCBI BioSamples, and the third, 01R0685, corresponds to one isolate deposited with others in a single BioProject (https://www.ncbi.nlm.nih.gov/bioproject/?term=PRJEB26000; individual isolate available here: https://www.ncbi.nlm.nih.gov/biosample/10125872). We have clarified in the table metadata that the relevant search identifier.

      (2) Line 351: The exact download date is missing here.

      We have added the exact download date

      (3) Line 363: Did you mean lowest e-value? Higher would be worse.

      Thank you for catching this, it is indeed the lowest e-value (confusion stemmed from looking at highest negative log e-value).

      Reviewer #2 (Recommendations for the authors):

      (1) The GitHub repo* was not public at the time of review (nor was it listed under user aggreen), so I could not check how reproducible the results are -- please make it public.

      Absolutely, this has been addressed.

      (2) The authors say "we envision structure being added as an additional feature in future work to predict resistance phenotypes from sequences" - this is not true, as some work has already been published predicting resistance in MBTC going back to 2019**

      We have addressed this with wording changes and citations to the mentioned work (see public responses)

      (3) Throughout the term 'non-synonymous' is used; that would include premature stop codons. Would 'missense' be more appropriate?

      You’re correct in pointing out that the term “non-synonymous” is too general for what we mean in this paper. Our analysis included missense mutations and in-frame indels, but did not include premature stop (nonsense) mutations or frameshift mutations. So, the term missense (alone) is narrower than what we wish to convey. We have made clarifications throughout.

      (4) The aminoglycosides are an important, albeit less used, class of antibiotics, and mutations arise in the ribosomal genes, e.g. rrs (which hence do not encode protein). They have therefore been excluded for obvious reasons, but it would help a reader from the tuberculosis field if this were acknowledged. Likewise, a reader might wonder why Rv0678 isn't in Figure 1 - I suspect it is because most of the samples were sequenced before the introduction of bedaquiline, but again, it would help if this were explained.

      See response to (5)

      (5) On a related note, it is not surprising that rpoC appears in Figure 1 due to its role in compensating for the fitness cost that arises when a rifamipicin-resistance mutation occurs in rpoB: obviously not central but a nice "oh yeah that makes sense" point for the reader if it were briefly mentioned.

      These are both great points about the relevance of our results to the Mtb community. We have added an additional paragraph interpreting the results of Figure 1 that mentions the reason for the appearance of RpoC and Cas10, and non-appearance of non-coding genes and genes relevant to resistance to newly introduced and repurposed drugs.

      (6) How is the "minimum coordinate difference" calculated? I assume all the structures are missing hydrogens as usual, so for two amino acids A and B, is it the smallest distance between any pair of heavy atoms from A and B? That would, I assume, introduce some bias for larger amino acids like Trp, or did you calculate from shared atoms like the backbone C_alpha atoms? That in turn will tend to make the distances a bit larger. It would be useful to know, as you explicitly mention a 1.5 nm threshold.

      We have explained this in a new section of the results, “The EVcouplings Python package was used to compute the distance between amino acid residues in all protein structures [49]. The package calculates the distance between all heavy (non-hydrogen) atoms in residue i and residue j, then returns the minimum of those distances.”

      (7) Given the reference used (H37Rv) is Lineage 4, one wonders about deeprooted/phylogenetic mutations, but then I suspect this sentence is doing a lot of that heavy-lifting: "We performed ancestral sequence reconstruction to determine the number of independent arisals of each mutation (homoplasy)". For the more general reader, it would be useful to touch on exactly what you mean and the importance of only considering homoplastic mutations.

      We have expanded our explanation in this section to better make the case for the use of homoplastic variants in our analysis, “Because analyzing the frequency of alleles in a population can be biased by oversampling of particular lineages, and by evolutionary recency, we chose to analyze the number of independent arrivals of each mutation (homoplasy) rather than their population-level frequency. This ensures that more recent evolutionary events are not underrepresented due to lack of time to spread in the population. To accomplish this, we used a previously compiled a dataset of genomes of 31,428 isolates from the Mycobacterium tuberculosis complex (MTBC), with ancestral sequence reconstruction to determine the number of independent arrivals of each mutation (Supplementary Data 1).”

      (8) Whilst this is true: "Evolution-based approaches are an alternative for finding variants associated with antibiotic resistance without requiring resistance phenotype data", the dataset used to, e.g. build the second edition of the WHO catalogue of resistanceassociated variants has >50,000 samples and therefore is larger than the dataset you have analysed here. The last time I looked, there were >100k M. tuberculosis samples in NCBI, and therefore, to be valid, your approach should really use more samples than are available with WGS and pDST data. I appreciate, however, that this will not be possible for this manuscript, but it is an obvious criticism.

      We appreciate this critique and acknowledge that the number of isolates with both WGS and pDST has increased rapidly in recent years. The first edition of the WHO catalogue (2021) used 38,215 isolates, which increased to over 50k in the second edition. The dataset of homoplastic mutations on which we based this paper was originally published in 2021.

      To address your comment, we have softened our assertions in the introduction about the utility of evolution-based approaches, and emphasize instead the different nature of the underlying signal, “Evolution-based approaches are an alternative for finding variants associated with antibiotic resistance by analyzing their mutational frequency and phylogenetic distribution”

      (9) Minor point, but the second sentence in the Introduction ignores that one can diagnose MDR-TB using phenotypic methods as well as genetic methods.

      We have added an additional citation and mention of laboratory phenotypic methods.

      (10) There are a few typos: "genic" "G-sore"

      Addressed.

    1. Author response:

      We thank all three reviewers for their detailed and constructive reviews, which will help us improve the manuscript. Concerning the simplicity of the model, we will elaborate on the limitations that arise from the modelling choices and better justify these choices. With that in mind, we believe this level of abstraction is well suited to the study’s objective. As a rate-coded model, it is designed to capture interactions between multiple circuit pathways, and lays the foundation to delve into neuronal dynamics in the future for further insight, once the network-level questions we address here - regarding the exploration-exploitation tradeoff and evasion of sub-optimal performance convergence - are well examined. For this purpose, we believe the model’s simplified network-level perspective is not a bug, but a feature. We will also explore in depth why and how the dual pathway model performs better in non-convex sensorimotor landscapes, building on prior analytical work that speaks to the question (Sankar, Leblois & Rougier, ICDL 2022). We will further substantiate and discuss our choice concerning the potential mechanisms underlying the proposed overnight synaptic volatility. We will include comparable approaches (including relevant machine learning analogues) in the introduction. Finally, we will clarify our interpretation of the sign of synaptic weights. We hope to incorporate all valuable feedback by the reviewers in the revised manuscript.

      References

      Remya Sankar, Arthur Leblois, Nicolas P. Rougier (2022). Dual pathway architecture underlying vocal learning in songbirds. In IEEE International Conference on Development and Learning, ICDL 2022, London, United Kingdom, September 12-15, 2022. pages 265-271, IEEE, 2022.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents an interesting behavioral paradigm and reveals interactive effects of social hierarchy and threat type on defensive behaviors. However, addressing the aforementioned points regarding methodological detail, rigor in behavioral classification, depth of result interpretation, and focus of the discussion is essential to strengthen the reliability and impact of the conclusions in a revised manuscript.

      Strengths:

      The paper is logically sound, featuring detailed classification and analysis of behaviors, with a focus on behavioral categories and transitions, thereby establishing a relatively robust research framework.

      Weaknesses:

      Several points require clarification or further revision.

      (1) Methods and Terminology Regarding Social Hierarchy:

      The study uses the tube test to determine subordinate status, but the methodological description is quite brief. Please provide a more detailed account of the experimental procedure and the criteria used for determination.

      We have included more details about how the tube test was performed in the revised manuscript. Social rank within each mouse pair was determined using a standard tube test paradigm. To minimize stress, mice were pair-housed for at least two weeks with a 15-cm tube placed in their home cage to allow voluntary exploration. Prior to rank assessment, mice were trained to traverse a 30-cm tube over two consecutive days (10 trials per day), with alternating entry from either end to prevent side bias. On the test day, each mouse pair underwent up to seven competitive trials in the same 30-cm tube. In each trial, the two mice were simultaneously released from opposite ends of the tube. A “win” was defined as one mouse successfully advancing through the tube while the opponent retreated completely out of the tube (all four paws outside) for at least 5 seconds. The first mouse to achieve four wins was designated as the dominant individual, whereas the opponent was classified as subordinate. Social rank stability was reassessed one day after threat exposure using the same criteria. Only pairs with consistent ranks were included in subsequent analyses.

      The dominance hierarchy is established based on pairs of mice. However, the use of terms like "group cohesion" - typically applied to larger groups - to describe dyadic interactions seems overstated. Please revise the terminology to more accurately reflect the pairwise experimental setup.

      Thanks for the comment. We have replaced the term “group cohesion” with “social engagement”.

      (2) Criteria and Validity of Behavioral Classification:

      The criteria for classifying mouse behaviors (e.g., passive defense, active defense) are not sufficiently clear. Please explicitly state the operational definitions and distinguishing features for each behavioral category.

      Passive defense was defined as an immobility-based defensive strategy characterized by suppression of locomotor activity, including freezing and tail rattling. Active defense was defined as movement- or posture-dependent defensive strategy, including approach, investigation, withdrawal, and stretch-attend. We have clarified these in the revised manuscript.

      How was the meaningfulness and distinctness of these behavioral categories ensured to avoid overlap? For instance, based on Figure 3E, is "active defense" synonymous with "investigative defense," involving movement to the near region followed by return to the far region? This requires clearer delineation.

      Defensive behaviors in the rat exposure paradigm were grouped into two categories: passive and active defense, each comprising distinct behaviors. All the manually annotated behaviors were mutually exclusive; that is, each video frame was assigned a single behavioral label to avoid overlap across behaviors. Active defense includes four behaviors: approach, investigation, withdrawal, and stretch-attend. We have clarified these points in the revised manuscript.

      The current analysis focuses on a few core behaviors, while other recorded behaviors appear less relevant. Please clarify the principles for selecting or categorizing all recorded behaviors.

      Thank you for pointing this out. In the current study, we focused primarily on defensive and social behaviors. We also included several neutral solitary behaviors related to anxiety and defensive state, such as sniffing, grooming, and rearing, which were consistently expressed across animals and closely linked to our main findings. We have clarified these in the revised manuscript.

      (3) Interpretation of Key Findings and Mechanistic Insights:

      Looming exposure increased the proportion of proactive bouts in the dominant zone but decreased it in the subordinate zone (Figure 4G), with a similar trend during rat exposure. Please provide a potential explanation for this consistent pattern. Does this consistency arise from shared neural mechanisms, or do different behavioral strategies converge to produce similar outputs under both threats?

      Thanks for bringing up this important question. The consistent increase in proactive bouts in dominant mice across both paradigms suggests a consistent rank-dependent reorganization of dyadic interaction under threats. We propose that this convergence reflect a shared neural mechanism that links defensive state with social-rank information, potentially involving top-down regulation from the mPFC to threat-specific midbrain and hypothalamic defensive circuits. We have expanded the discussion to incorporate this explanation.

      (4) Support for Claims and Study Limitations:

      The manuscript states that this work addresses a gap by showing defensive responses are jointly shaped by threat type and social rank, emphasizing survival-critical behaviors over fear or stress alone. However, it is possible that the behavioral differences stem from varying degrees of danger perception rather than purely strategic choices. This warrants a clear description and a deeper discussion to address this possibility.

      We thank the reviewer for this insightful comment. We agree that, in principle, behavioral differences could arise from variations in perceived danger rather than strategic choice. In humans, decisions can sometimes reflect value-based strategies that override perceived danger. In contrast, under naturalistic threat conditions, mice likely rely predominantly on danger perception to make behavioral decisions, and such responses are expected to be consistent with value-based strategies shaped by natural selection. In the revised manuscript, we have expanded the Discussion to address the role of threat perception and its relationship to decision-making in our behavioral paradigms.

      The Discussion section proposes numerous brain regions potentially involved in fear and social regulation. As this is a behavioral study, the extensive speculation on specific neural circuitry involvement, without supporting neuroscience data, appears insufficiently grounded and somewhat vague. It is recommended to focus the discussion more on the implications of the behavioral findings themselves or to explicitly frame these neural hypotheses as directions for future research.

      We have revised the Discussion to focus more directly on behavioral findings and added explicit neural hypotheses as potential future directions.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate how dominance hierarchy shapes defensive strategies in mice under two naturalistic threats: a transient visual looming stimulus and a sustained live rat. By comparing single versus paired testing, they report that social presence attenuates fear and that dominant and subordinate mice exhibit different patterns of defensive and social behaviors depending on threat type. The work provides a rich behavioral dataset and a potentially useful framework for studying hierarchical modulation of innate fear.

      Strengths:

      (1) The study uses two ecologically meaningful threat paradigms, allowing comparison across transient and sustained threat contexts.

      (2) Behavioral quantification is detailed, with manual annotation of multiple behavior types and transition-matrix level analysis.

      (3) The comparison of dominant versus subordinate pairs is novel in the context of innate fear.

      (4) The manuscript is well-organized and clearly written.

      (5) Figures are visually informative and support major claims.

      Weaknesses:

      Lack of neural mechanism insights.

      The current study focused on behavior. In the revised manuscript, we have incorporated a discussion of potential neural mechanisms and highlight this as an important direction for future work.

      Reviewer #3 (Public review):

      Summary:

      This study examines how dominance hierarchy influences innate defensive behaviors in pair-housed male mice exposed to two types of naturalistic threats: a transient looming stimulus and a sustained live rat. The authors show that social presence reduces fear-related behaviors and promotes active defense, with dominant mice benefiting more prominently. They also demonstrate that threat exposure reinforces social roles and increases group cohesion. The work highlights the bidirectional interaction between social structure and defensive behavior.

      Strengths:

      This study makes a valuable contribution to behavioral neuroscience through its well-designed examination of socially modulated fear. A key strength is the use of two ethologically relevant threat paradigms - a transient looming stimulus and a sustained live predator, enabling a nuanced comparison of defensive behaviors. The experimental design is robust, systematically comparing animals tested alone versus with their cage mate to cleanly isolate social effects. The behavioral analysis is sophisticated, employing detailed transition maps that reveal how social context reshapes behavioral sequences, going beyond simple duration measurements. The finding that social modulation is rank-dependent adds significant depth, linking social hierarchy to adaptive defense strategies. Furthermore, the demonstration that threat exposure reciprocally enhances social cohesion provides a compelling systems-level perspective. Together, these elements establish a strong behavioral framework for future investigations into the neural circuits underlying socially modulated innate fear.

      Weaknesses:

      The study exhibits several limitations. The neural mechanism proposed is speculative, as the study provides no causal evidence.

      Establishing causal evidence for neural mechanisms is beyond the scope of the current behavioral study. We highlight this as an important direction for future work in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Clarify the definitions of all behavioral categories (escape, assessment, passive defense, active defense, etc.) in detail.

      We have clarified the definitions of all behavioral categories in detail in the revised manuscript.

      (2) Reduce speculative statements about SC-VMHdm-mPFC circuitry, and provide the data about this circuit, if possible.

      We have reduced the speculative statements about this circuitry in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      We commend the authors on a carefully executed and conceptually clear study that makes a valuable contribution to behavioral neuroscience. Below are some issues and suggestions for this manuscript:

      (1) Please add the following to the discussion: the reason why, compared to looming exposure, social behaviors were more often followed by defensive behavior during rat exposure.

      We have added this discussion to the revised manuscript. This difference reflects the distinct temporal characteristics of the two threats. Following the transient looming stimulus, the threat rapidly ceases once the stimulus ends, reducing the need for sustained defensive behavior. Consequently, social interactions primarily occur after threat termination and are less frequently interleaved with subsequent defensive behaviors. Consistent with this interpretation, looming exposure selectively increased the duration of social behavior in subordinate mice. In contrast, rat exposure represents a sustained multisensory threat that maintains a persistently elevated defensive state. Under these conditions, both the frequency and duration of social interactions increased in dominant and subordinate mice. Moreover, huddling emerged as the predominant social behavior, and the frequent transitions between freezing and huddling suggest that social interactions become integrated with ongoing defensive responses, potentially serving as a safety-seeking or cohesive defense during sustained threat.

      (2) Figures 1B and 1C showed that Grooming has significantly decreased. Please verify whether the statistical methods and results are correct.

      The grooming data in the original Figure 1C did not distinguish dominant and subordinate mice and did not show significant decrease. We have removed it in the revised manuscript. In original Figure 2I (current Figure 1P), the social modulation on grooming behavior was observed only in dominant mice (Two-way ANOVA with post hoc Tukey’s range test).

      (3) To investigate how social context modulates the expression and progression of defensive responses, the authors analyzed behaviors in two time windows: the early phase (0-5 seconds after stimulus onset) and the late phase (20-60 seconds after onset) in Figure2. What is the rationale for selecting 5 seconds as the cutoff for the early-phase behavioral analysis? Were the behaviors of mice between 5 and 20 seconds also analyzed?

      The reason to select 5 seconds as the cutoff for early-phase behavioral analysis is that most behavioral decisions are made within this time window. We also analyzed the behaviors of mice between 5 and 20 seconds and have integrated them into the revised manuscript.

      (4) Figure 2 demonstrated that the social context attenuates looming-evoked defensive behavior in a rank-dependent manner. Then, I would like to ask if the influence of the social context on rat-evoked defensive behavior is also present in a rank-dependent manner? Discuss the similarities and differences between the social context's effect on looming-evoked defensive behavior and rat behavior.

      The influence of the social context on rat-evoked defensive behavior is also present in a rank-dependent manner. Specifically, total stretch-attend (SA) time and SA frequency were increased only in dominant mice (revised Figures 2Q and 2R), whereas average SA duration was decreased only in subordinate mice (revised Figure 2S). Approach-investigation-withdraw (AIW) frequency was increased only in dominant mice while approach speed was increased only in subordinate mice (revised Figures 2U and 2V).

      Social context exerts a broad protective influence across threat types, but the form of this modulation differs depending on the nature of the threat. Similarity: social presence consistently alleviates threat-induced stress and reshapes defensive behavior in a rank-dependent manner, with dominant benefiting more strongly under both threats, suggesting higher social rank may associate with greater flexibility to integrate social modulation. Difference: the behavioral outcomes of social regulation are distinct across threat types. During looming, social presence primarily suppresses immediate defensive responses and alleviates post-looming anxiety, suggesting under transient and unpredictable threat, social context dampens acute defense and facilitates behavioral recovery. In contrast, during the sustained rat exposure, social presence promotes a shift in defensive strategy from passive to active defense, rather than mere suppression of defensive output. These results suggest that social context flexibly adjusts defensive behavior according to ecological demands by reducing excessive defense and anxiety in response to transient looming and facilitating active coping when threatened by a sustained live predator. These discussions have been included to the revised manuscript.

      (5) In Figure 4, looming exposure increased these two measures only in subordinate individuals, whereas rat exposure affected both ranks, indicating threat-specific modulation of social behaviors. The visual looming paradigm primarily simulates visual stimuli triggered by aerial predators, while the rat exposure stimulus involves not only visual cues but also other sensory inputs, such as olfactory signals for the experimental animals. This may lead to differences in the defensive behaviors of mice, particularly increasing the proportion of proactive behaviors in dominant individuals. If olfactory information transmission is blocked during rat exposure, can the behavioral phenotypes shown in Figure 4 still be observed?

      The rationale for incorporating both looming and rat exposure paradigms was to mimic the distinct, ethologically relevant predator encounters by rodents in natural environments. As the reviewer pointed out, the multimodality nature of the rat threat may contribute to differences in the defensive behaviors displayed by mice. This point is also acknowledged in our manuscript, where we state that rat exposure “imposes prolonged stress and elicits a broader repertoire of defensive behaviors”.

      We agree that systematically dissecting the contribution of specific sensory modalities (e.g., olfactory, visual, or auditory) to defensive behaviors during rat exposure is an interesting and important question. Based on prior literature, we speculate that olfactory cues likely play a major role. However, the main aim of the present study is to investigate how social regulation of defensive behaviors depends on dominance hierarchy and threat type, rather than isolating modality-specific sensory mechanisms. Within this framework, we preserved the multisensory features of the rat exposure to maintain its ecological validity and to emphasize its distinction from the visual-only looming threat. We thank the reviewer for raising this insightful question, but it is beyond the scope of current study. We will consider it as an important direction for future investigations.

      (6) Discuss the potential mechanisms for the similarities and differences in the impact of the looming threat and the rat threat on cohesive behavior?

      We have expanded the Discussion to address the potential neural mechanisms underlying both the similarities and differences in the effects of looming and rat threats on social cohesion. As discussed in response to issue (1), we propose that the behavioral differences between the two threats arise from their distinct nature. Here, we further discuss a potential circuit mechanism. Specifically, we propose that the mPFC serves as a common hub integrating social context and dominance-related information with threat processing, thereby contributing to the shared enhancement of social engagement under both threats. We further speculate that the distinct patterns of social behavior may arise from partially distinct defensive circuits. SC-centered visual threat circuits may facilitate rapid post-threat social engagement by recruiting vigilance-related networks, whereas VMHdm-centered predator-defense circuits may promote sustained social cohesion by engaging neural circuits that support coordinated coping during persistent defensive states. We emphasize that these are hypotheses, and that future studies combining circuit-level recordings and causal perturbations within the behavioral framework established here will be required to test them.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this work, the authors investigate the mechanisms of low-frequency synaptic depression at cerebellar parallel fiber to interneuron synapses using unitary recordings that allow direct quantification of synaptic vesicle release. They show that sparse stimulation can induce robust synaptic depression even in the absence of substantial vesicle consumption, and that this depressed state is rapidly reversed when stimulation frequency is increased. To account for these observations, the authors propose a model in which low-frequency depression reflects a redistribution of vesicles within the readily releasable pool, in particular, a reduction in docking site occupancy due to vesicle undocking.

      Strengths:

      I found the experimental work to be of high quality throughout. The use of simple synapse recordings to count individual vesicle release events is particularly powerful in this context and allows questions to be addressed that are difficult to approach with more conventional approaches. The demonstration that low-frequency depression can occur independently of prior vesicle release, together with the rapid recovery observed during high-frequency stimulation, places strong constraints on possible underlying mechanisms and represents a clear strength of the study.

      The modelling framework is clearly laid out and helps organize a broad set of observations across stimulation frequencies. Several of the experimental tests appear well-motivated by the model, including the recovery train experiments, the analysis of failures, and the use of doublet stimulation. Taken together, the data provide a coherent phenomenological description of low-frequency depression and its relationship to vesicle availability within the readily releasable pool.

      We thank the Reviewer for his positive assessment of our work.

      Weaknesses:

      While the experimental results are strong, the manuscript would benefit from rebalancing the strength of the mechanistic conclusions drawn from the modelling in light of its limitations. The framework is clearly useful and provides a coherent interpretation of the data, but it is not uniquely constrained by the experimental observations, and alternative models or interpretations could plausibly account for the findings. The use of different model regimes concatenated across time, with substantially different parameter values, highlights the abstract nature of the approach. For these reasons, the model seems best presented as one plausible explanatory framework rather than a definitive biological mechanism. Clarifying the distinction between data-driven observations and model-based inferences would help readers assess which conclusions are strongly supported and which remain more speculative.

      The interpretation of the Ca2+-related experiments would benefit from more cautious wording. The absence of detectable changes in presynaptic Ca2+ signals does not exclude more localized or subtle Ca2+-dependent mechanisms, and conclusions regarding Ca2+ independence should therefore be framed accordingly. In addition, while low-frequency depression is still observed at reduced extracellular Ca2+, these experiments appear less diagnostic of the specific model-derived mechanism emphasized elsewhere in the manuscript - namely, a selective reduction in docking-site occupancy - and should be discussed with appropriate qualification in the text.

      Concerning Ca2+ signals, the Reviewer is right. While we found no change in Ca2+ signalling apart from a slow Ca2+ accumulation during long trains at 1 Hz, the possibility of an undetected change cannot be excluded. We have added a word of caution in this direction on p. 11. Concerning the 1.5 mM Ca2+ experiments, the Reviewer presumably alludes to the first recovery train (yellow) point in Supplementary Fig. 2C. This is also the last point (s11) of the slow train at 0.5 Hz because no delay at all was interposed between the slow train and the recovery train. We have now included one more experiment (with a present total number n = 6), and we have corrected Fig. S2C accordingly. In the new version the depression measured for s4-s10 vs s1 during the 0.5 Hz trains is 0.69 +/- 0.05 (p = 0.00058, paired one-tail t-test). The ratio of the s1 value of the recovery train compared to control s1 is 0.83 +/- 0.08 (p = 0.028, paired one-tail t-test).

      Major points:

      (1) Clarify and qualify mechanistic claims derived from the model.

      Throughout the manuscript, changes in model parameters are at times described as if they directly reflected underlying physiological mechanisms. As a result, the conceptual distinction between experimentally observed phenomena, model-derived variables, and biological interpretation is not always clear. Several conclusions in the Results and Discussion are phrased as mechanistic statements, although they rest on assumptions intrinsic to the modelling framework. The authors should systematically review the text and explicitly distinguish between (i) experimentally observed changes in synaptic responses and (ii) inferences about vesicle docking states or transitions within the model.

      In particular, statements implying that vesicle undocking is the mechanism underlying low-frequency depression should be rephrased to reflect that this is an interpretation within the proposed framework rather than a uniquely demonstrated biological process. For example, statements such as "Low-frequency depression is caused by synaptic vesicle undocking" should be replaced with formulations such as "Within the framework of our model, low-frequency depression is accounted for by a redistribution of synaptic vesicles away from docking sites" or "Our results are consistent with a model in which changes in vesicle docking-state occupancy contribute to low-frequency depression."

      A particularly problematic example is the statement that "these experiments further confirm that LFD only involves a decrease in δ, without accompanying changes in ρ or IP size." Here, an experimentally defined phenomenon (LFD) is directly equated with changes in model-derived variables. Such statements should be revised to make clear that δ, ρ, and IP size are inferred quantities within the model, and that the experimental data are interpreted through this framework rather than directly confirming changes in these parameters. Similarly, overgeneralizing statements such as "Undocking therefore represents the key mechanism controlling short-term depression across stimulation frequencies" should be softened to reflect that this conclusion emerges from the model rather than from direct experimental evidence.

      As suggested, we clarify the distinction in the revised version between experimental data and modelling, and we refrain from making definitive statements on underlying cellular mechanisms.

      (2) Address the biological interpretation of time-dependent model regimes.

      The model relies on distinct parameter regimes applied at different time points, with some transitions effectively suppressed in certain regimes. While this approach captures the data well, its biological interpretation remains unclear. The authors should either (i) expand the discussion to outline plausible biological processes that could give rise to such regime changes (for example, calcium-dependent modulation of transition rates or activity-dependent changes in vesicle state stability), or (ii) more explicitly frame this aspect of the model as a descriptive abstraction rather than a mechanistic proposal. This further underscores the need to clearly separate the descriptive role of the model from claims about underlying biological mechanisms.

      We thank the Reviewer for drawing our attention to this important point. Below 10 ms, rate constants are largely determined by the large-amplitude, fast-decaying Ca2+ signal occurring near voltage-dependent Ca2+ channels (‘Ca2+ nanodomain’). After 10 ms, the rate constants depend on the low-amplitude, slowly decaying Ca2+ signals averaged over the entire varicosity (‘volume-averaged Ca2+’). We explain this better in the revised version (Materials and Methods, p. 21).

      (3) Reframe conclusions drawn from calcium-related experiments.

      The calcium imaging data demonstrate no detectable changes in the measured presynaptic calcium signals under the tested conditions, but they do not rule out that calcium signals contribute in ways undetectable by the assay. Conclusions should therefore be revised to reflect this limitation, avoiding statements that exclude a role for calcium-dependent mechanisms. Wording such as "we did not detect evidence for..." would be more appropriate than conclusions implying the absence of an effect.

      Similarly, while low-frequency depression is still observed at reduced extracellular calcium (1.5 mM Ca<sup>2+</sup>), the specific mechanistic signature emphasized elsewhere in the manuscript - namely a selectively reduced first response during a high-frequency recovery train - is no longer apparent. These experiments should therefore be discussed as consistent with the proposed framework, but not as providing independent support for a selective reduction in docking-site occupancy. Explicitly acknowledging this limitation would improve clarity and avoid overinterpreting these data.

      This has been discussed above (‘weaknesses’).

      (4) Soften interpretations based on non-significant comparisons.

      In several places, comparisons that do not reach statistical significance are used to argue for equivalence between conditions (for example, comparisons involving failure versus non-failure trials or different LFD conditions). These conclusions should be revised to emphasize the limits of statistical power and framed as a lack of evidence for a difference rather than evidence of independence.

      We have amended this point in the revised version.

      Reviewer #2 (Public review):

      Summary:

      Silva and co-workers exploit their previously established methods of analyzing release events at single parallel fiber to molecular layer interneuron synapses. They observed synaptic depression at low transmission frequencies (< 5 Hz), which rapidly recovers during high-frequency transmission. Analysis of the time course of low-frequency depression revealed an initial rapid and a slow linearly increasing time course. Strikingly, the initial depression occurred even in the absence of preceding release, arguing against vesicle depletion as the underlying mechanism.

      Strengths:

      The main strength of the study is the careful demonstration of an interesting synaptic phenomenon challenging the classical vesicle-centered interpretation of synaptic depression.

      We thank the Reviewer for his positive assessment of our work.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      The finding of release-independent synaptic depression is important and would have widespread implications. Therefore, some more analyses to increase the confidence in these findings could be performed.

      My concern is whether rundown could explain the findings. If the rate of failures in s1 increases and at the same time the amplitude decreases during the experiments, an apparent depression in s2 could arise. The Supplementary Figure 5A addresses run-down, but the figure is not easy to understand, and, as far as I understood, it does not address the question of whether the release-independent depression could be caused by a rundown. To address this, the analysis of Figure 5 could be repeated by investigating the failure rate and amplitude separately or by analyzing the 1st and 2nd half of the recordings separately.

      The Reviewer makes a very important point that had escaped our attention. If the responses were declining over the course of an experiment, near the end of the recordings, a high proportion of failures would be associated with a weak response to the second AP. This could distort the relation between initial failures and amount of LFD, perhaps to the point of indicating LFD after failures when there were none. As suggested by the Reviewer, we tested this possibility by examining the stability of the synaptic responses during experiments. We found a mean s<sub>1</sub> value of 0.87 ± 0.13 for the first half of the experiments used in Fig. 5, and of 1.10 ± 0.17 for the second half (p > 0.05, n = 10). This analysis shows that there was no rundown during these experiments. We show in Author response image 1 a plot of s1 as a function of train number in these experiments, for two examples as well as for the average across frequencies and across experiments. These plots do not suggest any artefactual correlation between failures, mean s1, and rundown.

      Author response image 1.

      Plot of s1 as a function of train number for the experiments of Fig. 5

      In response to a request of Reviewer 2, Author response image 1 illustrates the evolution of s1 values as a function of train number for the experiments used to produce Figure 5. In each experiment, about 20 s1 values were obtained at two ISIs (either 10 ms and 500 ms, or 800 ms and 1600 ms). Author response image 1 shows two examples of s1 values as a function of train number (these values fluctuate widely between 0 and 3), and the average across cells and ISI values. There is no indication of a rundown of s1 values as a function of train number.

      Reviewer #3 (Public review):

      Summary:

      The manuscript builds on the observation that, at some synapses, low-frequency stimulation causes synaptic depression, which can be reversed by subsequent high-frequency stimulation. Such low-frequency depression (LFD) cannot be easily explained by the depletion of a single vesicle pool. Here, Silva and colleagues propose a model of activity-dependent vesicle trafficking to explain LFD at synapses between cerebellar granule cells and molecular layer interneurons.

      Strengths:

      Overall, LFD is interesting and worthy of examination, and the authors provide new experimental results that are of the high quality expected from this group.

      Weaknesses:

      The study proposes a novel model of vesicle trafficking that is not explained by known biological mechanisms, and the manuscript does not adequately compare or discuss alternative models.

      I have several concerns about how the authors interpret the data. First, the manuscript's primary conceptual advance is the idea that LFD involves vesicle undocking, rather than depletion. However, most experiments were performed under conditions that promote vesicle depletion (3 mM extracellular Ca2+). When experiments were repeated in physiological Ca2+, there appeared to be little or no LFD (stats are not provided). Second, the RS/DS/DU/undocking model, though not outside the realm of possibility, is not readily explained by known mechanisms and is only loosely supported by experimental findings. Third, when simulating LFD, the authors do not compare alternative models and use inappropriate language to imply that a model fit represents the truth (e.g., "the finding of identical experimental and simulated values confirms that the undocking mechanism accounts for LFD"). Finally, the model is presented in an overly complicated manner. The sheer amount of terms and nomenclature makes the manuscript confusing and difficult to read. Overall, the manuscript would benefit from added experiments and more statistics, a better justification and evaluation of the model, and more nuanced language.

      We respectfully disagree with these sweeping criticisms, as described in more detail below.

      Major concerns:

      (1) Most experiments were performed under conditions that exacerbate depletion

      In order to attribute LFD to vesicle undocking rather than depletion, it is important to show LFD under conditions where depletion is minimal. As mentioned above, the authors only report significant LFD in elevated extracellular Ca2+. In a small number of experiments performed in more physiological Ca2+ (1.5 mM), there is no depression after a single stimulus, and it is not clear that there was statistically significant depression during a low-frequency train. Several studies cited in support of LFD share this problem:

      Abrahamsson et al., (2007) recorded from Schaffer collaterals in 4 mM Ca, 3-4X physiological Ca2+.

      Doussau et al., (2010) recorded from Aplysia synapses in 3X Ca compared to seawater.

      Rudolph et al., (2011) is cited as an example of LFD. However, this study performed experiments at high release probability cerebellar climbing fibers, and reported depression that increased monotonically with stimulation frequency, so it does not resemble the phenomenon studied in this paper. Lin et al., (2022) also largely describe monotonic depression at the calyx.

      The Reviewer suggests that LFD may only occur under non-physiological conditions, if the release probability has been increased by artificially elevating the extracellular Ca2+.

      The implication is that LFD is at best a curiosity with little or no significance for brain signalling. We disagree with this point of view for several reasons.

      Concerning the statement ‘In order to attribute LFD to vesicle undocking rather than depletion, it is important to show LFD under conditions where depletion is minimal’: This is the purpose of the analysis shown in Fig. 5.

      The statement ‘the authors only report significant LFD in elevated extracellular Ca2+’ is inaccurate. Fig. S2C shows a clear LFD in 1.5 mM Ca2+, as acknowledged by Reviewer 1 (‘low-frequency depression is still observed at reduced extracellular Ca2+’). However, we failed to provide a p-value for the depression in the initial version of the paper (p was 0.004, n = 5 with this data set; paired t-test, one-tail). In the revised version, we document the 1.5 mM results more extensively, including the incorporation of the results of an additional experiment, and an explicit statistical analysis of the data (p = 0.00058, n = 6; paired t-test, one-tail).

      Concerning the statement ‘there is no depression after a single stimulus’: We find that the onset kinetics of LFD is slower in 1.5 Ca2+ than in 3 Ca2+ (respectively 1.8 ISI and 0.51 ISI, Fig. 2C and Fig. S2C). This explains that the PPR is not significantly <1 in 1.5 Ca2+ without implying any weakening of the extent of LFD at steady state.

      As explained in the manuscript (p. 5), in a previous work, we developed a method to ascribe changes in SV pools, within the RS/DS model, with specific modifications of s1, s2 and s5-s8 during test 100 Hz trains (Tran et al., 2022). This method was developed in 3 mM Ca2+ conditions, and for this reason, we performed most experiments for the present work in 3 mM Ca2+.

      Chiu and Carter (2024) demonstrated LFD in neocortical synapses; they performed their study in 1.2 mM Ca2+, not in elevated Ca2+.

      Rudolph et al., (2011) showed low-frequency depression not only in elevated external Ca2+, but also in 0.5 mM Ca2+. While Rudolph et al., (2011) did not make an explicit link between their observations and LFD, there is no reason to doubt that these observations are an example of LFD. They showed a biphasic depression when switching the stimulation frequency from 0.05 Hz to 2 Hz. In one of the founding papers of LFD, Doussau et al., (2010) describe a biphasic depression when switching the stimulation frequency from 0.025 Hz to 1 Hz; Fig. 1 of the two papers (Rudolph 2011 and Doussau 2010) are strikingly similar.

      Lin et al., (2022) would probably not agree with the statement that the depression at the calyx is ‘largely monotonic’, as they stress the finding of quasi-constant depression between 5 and 50 Hz.

      The authors note that their results differ from those of Atluri and Regehr, but do not mention that a possible reason for the difference is the increased release probability in their experiments.

      In fact, we clearly listed the difference in external Ca2+ as a likely source of the discrepancy by saying ‘This discrepancy presumably stems from differences in experimental conditions (room temperature, stimulation of multiple presynaptic PFs and 2 mM external Ca<sup>2+</sup> concentration in the previous work, vs. near-physiological temperature, single presynaptic stimulation and 3 mM external Ca<sup>2+</sup> here)’.

      The authors should provide statistics for the data obtained in 1.5 mM Ca, and discuss why LFD is increased in conditions that also elevate vesicle release probability.

      See our comments above: the revised version includes the requested statistics. On p. 6 of the manuscript, we do provide an explanation for the apparent lack of LFD at 1.5 Ca2+ and 2 Hz, namely a superimposition of LFD with facilitation. At 1.5 Ca2+ and 0.5 Hz, our LFD numbers are not weaker than at 3 mM Ca2+ and 0.5 Hz of 1 Hz.

      Altogether, it is correct that many LFD experiments have been carried out in high release probability synapses and/or under conditions of elevated Ca2+. However, the reasons underlying these choices are diverse (in our case, to build on the previous SV pool analysis developed in Tran et al., 2022 in 3 Ca2+ conditions) and do not imply a limitation to the phenomenon. LFD is present in physiological conditions for low-to-moderate release probability synapses (as shown in our work), and altogether, there is no reason to dismiss LFD as non-physiological.

      (2) Lack of biological mechanisms supporting the model

      The model is presented without compelling biological support. The evidence in support of vesicle undocking comes from experiments by the Watanabe lab, which showed fewer-than expected docked vesicles under EM when cultured synapses were stimulated immediately prior to high-pressure freezing. Kusick et al were careful to note that these vesicles may have been lost to fusion.

      The Watanabe lab showed an SV deficit at docking sites at times ranging from about 100 ms to several seconds (Kusick et al., 2020, their Fig. 5E). This corresponds to the ISI values where we see paired-pulse depression. In their Summary, Kusick et al., raise the possibility of SV fusion as an alternative to undocking at the 100 ms time point. But the same issue had previously been considered in Miki et al., 2018 with other techniques (their Fig. 2d), where it was shown that the SV deficit seen in paired-pulse experiments could not be explained by fusion. This leaves undocking as the most likely explanation, at least in our preparation. We have added a new paragraph on p. 14 to clarify this point.

      The putative undocking Kusick describes is immediate (< 5 ms after stimulation), and it was not shown to be Ca2+ sensitive. This manuscript describes "calcium-dependent undocking" that proceeds from 10 ms - 200 ms. Multiple studies from the Watanabe lab show that a single stimulus lowers the number of docked vesicles, and subsequently, there is a transient redocking of vesicles that can be blocked by EGTA or Syt7 knockout.

      This is not an accurate description of the Kusick results or of our results. In the Kusick paper, the SV deficit seen at <5 ms after stimulation is attributed to exocytosis, not to undocking. Clearly, it is Ca2+-dependent. Our manuscript describes potential calcium-dependent undocking not during the time 10 ms- 150 ms, during which our undocking rate is assumed to be calcium-independent, but starting at 150 ms and lasting a few hundred ms thereafter.

      I also question the rationale for the authors' model that 2 vesicles are coupled in series to a single release site. Previous papers from this lab cited EM studies from frog and neuromuscular that showed filamentous connections between vesicles (do these synapses show LFD?). Here, the authors primarily cite their previous models to support their arguments. I encourage them to continue searching for ultrastructural evidence for 2-vesicle-docking-units and to cite such studies.

      It is important to remember that our sequential two-step model was not based on EM data, but on a series of functional data including variance-mean analysis of summed SV release numbers; covariance analysis among subsequent SV release numbers; analysis of release latencies as a function of stimulus number during an AP train; analysis of SV release numbers under conditions of very high release probability. We note that the phenomenon of Ca2+-dependent docking that we proposed based on these observations has been consistent with flash-and-freeze or zap-and-freeze results from several laboratories. Concerning potential filamentous connections between SVs and the AZ plasma membrane at a distance of several 10s of nm, this has been seen not only in frog or mice neuromuscular junctions, but also at brain synapses (ex: Siksou et al., Journal of Neuroscience 2007; Cole et al., Journal of Neuroscience 2016; Fernandez-Busnadiego, Journal of Cell Biology 2010; 2013).

      (3) Comparison to other vesicle models

      The authors use overly assertive language to suggest that the model proves a mechanism. "Altogether, these results indicate that the slow phase of LFD ... reflects a δ decrease without significant changes in pr, in ρ or in IP size". Simulating data does not conclusively "indicate" the underlying mechanism, but the authors could state their data can be "explained by a model where..".

      Please see our response above to a similar point by Reviewer 1.

      However, LFD does not require activity-dependent undocking. Instead, the phenomenon has been explained by high-release probability, paired with an activity-dependent increase in either docking or release probability (Chiu and Carter, 2024; Doussau et al., 2017). Does the new model do a better job of replicating some facet of the data? If multiple models can explain the same data, how can we determine which model is correct? The "Alternative Presynaptic Depression Mechanisms" should be expanded to discuss these issues.

      We could not find statements in the Chiu and Carter paper or in the Doussau et al., paper explaining LFD ‘by high-release probability, paired with an activity-dependent increase in either docking or release probability’. As far as we can see, Chiu and Carter do not propose any specific mechanism for LFD, beyond saying that depression and facilitation must be separate. Doussau et al., (their Fig. 6) clearly frame their interpretation in a sequential two-step model. As in the preceding Miki et al., paper (which they cite extensively), they assume a rapid (a few ms), Ca-dependent transition between their ‘reluctant pool’ and their ‘fully-releasable pool’, respectively homologous to RS and DS. Thus, the Doussau et al., interpretation is close to that presented in our present work, even though significant differences exist. An important difference is that Doussau et al., did not use simple synapses, so that they did not have access to key synaptic parameters such as the number of docking sites or the release probability per docking site. Consequently, the model in Doussau et al., does not have the same level of detail as ours. The revised version explains better the differences and similarity between the models of Doussau et al., and that exposed in our work (new paragraph on p. 14).

      Recommendations for the authors:

      Reviewing Editor Comments:

      Three reviewers have seen the paper. They agree that the evidence for low-frequency depression (LFD) is solid, but they all suggest that the data with 1.5 mM extracellular calcium should be analyzed and discussed in more detail. The finding of release-independent depression is surprising, but further data analysis to investigate whether it could be caused by run-down is required. The reviewers think the study will be more complete if the authors add N's and stats for LFD in 1.5 mM Ca. In addition, the most important change would be adding clarity in the text about the limitations and simplifications of the model. Please edit the text to clearly state that the presented model provides one possible framework to account for the data, and that other models might also be plausible, and clearly describe the limitations and simplifications of the model.

      Thank you for your positive assessment of our work. As explained in more detail below, we have added the requested statistics for 1.5 mM Ca data. We have also thoroughly revised our text to separate more clearly facts and interpretation. Finally, we now explain in more detail the limitations of our model.

      Reviewer #1 (Recommendations for the authors):

      Minor points

      (1) Statistical analysis and reporting.

      Please justify the use of one-tailed statistical tests and consider whether two-tailed tests would be more appropriate in some cases.

      In Materials and Methods (p. 17), we explain that one-tailed statistics were used when the sign of any possible deviation was predicted; otherwise two-tailed statistics were used.

      Exact p-values should be reported consistently rather than using threshold notation.

      We have corrected this as far as we could. Exact p-values that are still missing will be provided in the version of record.

      For analyses using rank-based statistics (e.g., calcium imaging and recovery experiments), please clarify which comparisons were paired and which were unpaired, and ensure that this is clearly indicated in the Methods and figure legends.

      We have now clearly indicated whether these statistics were paired or unpaired.

      (2) Figure clarity and presentation.

      In Figure 1, clarify whether the responses shown in panel C represent measured EPSCs or derived quantities, as this is not immediately clear from the labelling.

      Thank you for spotting this, the figure has been corrected.

      In Figure 2, please show the variability of control responses used for normalization.

      The variability of control responses used for normalization is now clearly explained in Materials and Methods (p. 18).

      In Figure 3, the use of similar color schemes for different model states and for data-model comparisons is confusing and should be revised.

      Thank you for spotting this; has been corrected.

      In Figures 4 and 5, consider reintroducing or clarifying schematic representations of the model states referred to in the text to aid reader comprehension.

      Figure 4 had already such representations. We have now modified these representations to make them simpler and clearer. We considered adding some model to Fig. 5 but could not come to any satisfactory option.

      (3) Controls and interpretation of pharmacological experiments.

      For experiments involving intracellular BAPTA, it would strengthen the interpretation to either include or explicitly discuss positive controls demonstrating that the manipulation was effective under the recording conditions. Alternatively, loosen the interpretation of those data.

      We have modified the text as suggested by the Reviewer.

      (4) Terminology and consistency.

      Ensure consistent use of terminology for vesicle states and model variables throughout the text and figures, and clarify whether different labels (e.g., RS/DS versus alternative notations) refer to equivalent states.

      We have revised our manuscript as suggested.

      (5) Model fit at short inter-stimulus intervals.

      In Figure 3B, the model appears to capture the overall trend of facilitation but shows a noticeable deviation from the data at the shortest inter-stimulus intervals. It would be helpful to clarify whether this reflects limitations of the current parameterization, known simplifications in the model (e.g., assumptions about fast calcium-dependent processes), or variability across recordings. Briefly commenting on this mismatch would help readers interpret the strengths and limits of the model in this regime.

      Because of the statistical nature of data (quantal fluctuations) and simulations (Monte Carlo trials), some differences between data and simulations are unavoidable. The low simulation point at 10 ms ISI in Fig. 3B may in addition result from an imperfect transition between the two time domains (time domains 1 and 2) of the simulation. We have added a sentence in Materials and Methods (p. 22) to explain this.

      (6) Model parameterization and fitting.

      Please clarify how model parameters were estimated, including the fitting procedure, cost function, and any constraints applied. It would be helpful to report confidence intervals or other measures of uncertainty for key fixed parameters, and to indicate how variable these parameters are across cells or recordings. In addition, a brief discussion of how sensitive the main conclusions are to variation in key parameters (such as release probability or transition rates) would help readers assess the robustness and identifiability of the model-derived interpretations.

      In addition, it is not entirely clear how model parameters and transitions differ across the concatenated temporal regimes used in the simulations, nor which parameters are held fixed versus altered or effectively suppressed. Summarizing, for each regime, the parameter values (or ranges) and active transitions-either in a table or schematic-would greatly improve transparency and help readers assess parameter identifiability and the robustness of the model-derived conclusions.

      For each group of experiments, results under control conditions (8-APs at 100 Hz) were averaged together, leading to the determination of one set of parameters (N, ⍴, δ, r, s, p<sup>r</sup>). Next the parameters for the second time domain were obtained by trial and error to optimize the fit for Fig. 1 (PPR) data. The parameters were constrained such the PPR recovered to 1 with a 10-second ISI. Unfortunately, the parameter domain was too complex to make an unsupervised fitting procedure practical. We have now added a paragraph on p. 22 to indicate which model parameters were heavily constrained by data and which ones were less constrained. Individual synapses were not simulated. Regarding the different time domains, the probabilities rb and sb are introduced for the second time domain and rf and sf are changed as shown in Fig. Supp. 3. The kinetic parameters are reset at the beginning of each simulated AP.

      (7) Clarify the scope of the model with respect to the slow component of low-frequency depression.

      The manuscript identifies a slow component of low-frequency depression that persists over long timescales. While this is an interesting observation, it is not fully clear whether this component is intended to be captured by the current model or represents an additional process outside its scope. Making this distinction explicit would help readers better grasp the scope and limitations of the model.

      As stated in our manuscript, modelling the slow component of LFD falls outside the scope of the present work. The set of rate constants in Fig. S3 does not produce any slow component, and we failed to find a combination of rate parameters that would produce a slow component. This is now clearly stated in the Materials and Methods section.

      Reviewer #2 (Recommendations for the authors):

      I have only the following suggestions to further improve the manuscript.

      (1) It is argued that LFD is alleviated when using doublets rather than singlets (Figure 8), but is the steady-state depression in s1 between singlets and doublets at the end of the experiment statistically different?

      It is unclear what should be considered steady state here. Average s<sub>1</sub> values are smaller for singlets than for doublets over the last 25 points (0.58 +/- 0.05 and 0.74 +/- 0.05, p = 0.02, unpaired t-test).

      (2) I do not understand Supplementary Figure 4. Does the value of zero in this plot indicate that the stimulation does not evoke any release anymore? How can this be differentiated from rundown? Can the LFD be reversed by a high-frequency stimulation?

      We have rewritten the description of this figure. We have also modified the figure labelling. Finally, we have changed the size and color of the symbol at zero delay to stress that the linear fit is constrained to include this point. Altogether we trust that these changes have clarified the presentation of these results.

      (3) The findings are related to Kusick et al., (2020). However, I think that in Kusick et al., (2020) as well as in Watanabe et al., (2013; doi:10.1038/nature12809), all tested time points ranging from a few milliseconds to several seconds after a stimulus show vesicle depletion except the time point of 14 ms in Kusick et al., (2020). I think these findings are, therefore, inconsistent with the much slower observed LFD.

      The increase in docked SV# at 14 ms is reported not only in Kusick et al., 2020 but also in Wu et al., 2023 as well as in Ogunmowo et al., 2025. Later, at 100 ms, the number of docked SVs drops again (Kusick et al., 2020, their Fig. 5 d-e, and Ogunmowo et al., 2025), to finally recover on a time scale of seconds (Kusick et al., 2020). These data fit with our observations if LFD is due to undocking, as we propose.

      (4) Page 13, 4th paragraph:

      Eshra et al., (2021) provide experimental evidence for "a Ca2+ -dependent rate of SV entry into the RRP".

      The Reviewer is right: Eshra et al., actually reported a weak Ca2+-dependence of the replenishment. We have therefore eliminated the Eshra reference in this sentence (note that this paper is mentioned elsewhere in the manuscript).

      Reviewer #3 (Recommendations for the authors):

      (1) Complicated terminology: I had a hard time understanding the model due to terms like δ, ρ, IP, DS, RS, RS gate, etc. Anything the authors can do to reduce the number of acronyms and simplify their description/depiction of the model will be helpful for readers.

      This resembles recommendation (4) of Reviewer 1. As stated in our response above, we have entirely revised the manuscript with these issues in mind. In addition, the abbreviation list at the onset of the manuscript should help the readers.

      (2) Figure 5 may represent the strongest evidence in support of an undocking model. This is the smallest figure, and I found it harder to understand. For example, I was initially confused because the blue "fail S1" traces seemed to suggest a very large PPR value. I suggest vastly expanding the size of this figure and the associated results section.

      Thank you for your interest in Fig. 5. We have explained the normalization procedure in more detail in the revised version. Also, we have modified the axis label of Fig. 5B to improve clarity. Finally, we have added a paragraph to explain the reason why the RS/DS model accounts for the results of Fig. 5.

      (3) Given the similar results but differing mechanistic conclusions of Doussau et al., (2017), that study merits more discussion in the manuscript.

      See our response to this point above (main comment (3) by the same Reviewer).

      (4) Several studies of LFD synapses have found a role for kinase activity (Silverman-Gavrila et al., 2005; Doussau et al,. 2010). This mechanism could be discussed or experimentally tested here.

      As mentioned in the Discussion section, the nature of the proteins involved in LFD, or in docking/undocking, remains uncertain. It is not surprising that broad-spectrum kinases and phosphatases affect LFD, as reported in the papers quoted by the Reviewer, but the full implications of these findings will only become clear

    1. Author response:

      The following is the authors’ response to the previous reviews

      Summary of revision for all reviewers:

      We are encouraged that the reviewers recognized the importance of the central question addressed by this study, the value of comparing acoustic- and cochlear-implant-evoked cortical responses within a common framework, and the potential relevance of these analyses for auditory neuroscience and neuroprosthetic design.

      At the same time, the reviewers made clear that the manuscript would be strengthened by:

      (1) Clearer calibration of claims regarding spatial organization 

      (2) Clarified explanation of the Figure 8 cross-modal analyses

      (3) More explicit discussion of the methodological and interpretive limitations of the present dataset, particularly the acute, anesthetized, and partially indirectly validated nature of the experiments

      In response, we substantially revised the manuscript to improve methodological clarity, narrow claims where appropriate, and align the title, abstract, results, discussion, figures, and legends with the precision of the data.

      Summary of major changes to revised manuscript:

      We revised the manuscript throughout to distinguish non-random spatial organization, coarse topographic/cochleotopic structure, and locally graded tonotopy/cochleotopy, and we softened claims accordingly, especially for the cochlear implant and TCA-derived analyses.

      We substantially clarified the Figure 8 cross-modal decoding framework, including:

      - The within-modality normal-hearing control

      - The normal-hearing-trained / cochlear-implant-tested analysis

      - The shuffled baseline used to estimate chance-level information transfer

      We expanded the methods and discussion to make the scope and limitations more explicit, including:

      - The acute and anesthetized nature of recordings

      - The use of monopolar stimulation

      - The lack of direct deafening validation in the main iEEG cohort

      - That ECAP forward-masking measurements were obtained in a separate acute cohort

      - The possibility that the observed acoustic–electrical mismatch is transient rather than fixed

      We also improved the consistency of our terminology, panel labeling, figure legends, and cohort descriptions.

      Public Reviews:

      Reviewer #1 (Public review):

      We thank the reviewer for recognizing the importance of the question addressed by this study and for pushing us to sharpen the distinction between non-random spatial structure, coarse topography, and graded cochleotopy. These comments also helped us better calibrate our claims about TCA-derived maps and cross-modal generalization, and prompted clearer discussion of the limits of the present dataset.

      (1) The main weakness is that the evidence for spatial organization remains difficult to interpret. In Figure 2, the authors argue that both tone-evoked and cochlear implant-evoked responses are spatially organized, but the slope analyses are not significant for the cochlear implant condition. The revised vector-strength analysis supports the presence of non-random spatial structure, but this is not the same as demonstrating a clear graded cochleotopic organization. The manuscript would be strongest if it consistently distinguished between non-random spatial structure, coarse topography, and true graded tonotopy or cochleotopy.

      We agree. We revised the manuscript to distinguish these levels of interpretation more carefully and consistently. Specifically, we now reserve “non-random spatial organization” for results supported by the vector strength/shuffle analyses, and describe cochlear-implant-evoked maps as showing coarse and variable spatial structure rather than uniformly robust graded cochleotopy.

      We now also emphasize more explicitly that implanted animals varied considerably: some showed clear nonrandom organization, whereas others exhibited effectively random global maps. At the same time, the aggregate implant-evoked data remained non-random and showed decreasing spatial correlation with increasing electrode separation, which we interpret as evidence for coarse cochleotopic structure at the population level, rather than strong local graded cochleotopy in every animal.

      Accordingly, we revised the manuscript to distinguish:

      “Non-random spatial organization”

      “Coarse topographic/cochleotopic structure”

      “Locally graded tonotopy/cochleotopy”

      (2) A related issue is that some figure titles and interpretive statements still appear stronger than the data justify. For example, the TCA results in Figure 7 are described as revealing topographically organized latent spatial factors, but the statistical support appears strongest for normal-hearing high-gamma responses, with weaker or non-significant results in other conditions. These data remain interesting, but they would be better framed as evidence for weak or coarse spatial structure rather than robust topographic organization across all modalities.

      Now in our updated manuscript we revised the Figure 7 title, legend, and results to avoid implying robust topographic organization across all conditions. The manuscript now describes the TCA-derived spatial factors as showing coarse, non-random spatial structure, with the strongest support for local tonotopy in the normal hearing high-gamma condition and weaker or non-significant support for robust local topography in the other conditions.

      (3) The decoder analyses are improved, especially with the added tone-to-tone control. This control supports the conclusion that poor acoustic-to-CI transfer is not simply a failure of the TCA/LDA pipeline. However, the analysis remains model-dependent, and the absolute information transfer values are low. It would be helpful either to include an analogous analysis using raw ERP/high-gamma features or to explain more explicitly why the TCA-based approach is the appropriate primary test. The data support poor generalization between acoustic and implant-evoked cortical responses, but claims about perceptual qualities should remain speculative because perception is not directly measured in these experiments.

      Thanks, we now revised the Figure 8 results section to first explain that TCA is especially appropriate here because it constrains the decomposition into separable spatial, temporal, and trial factors. This makes it possible to learn spatial and temporal factors in the normal-hearing condition, hold those fixed, and re-optimize only trial factors in withheld normal-hearing data or cochlear-implant-evoked data, providing an interpretable test of within-modality recovery and cross-modal generalization.

      Thus, our emphasis on TCA is primarily conceptual, not merely computational. Raw-feature decoding could also be informative, but for the specific cross-modal question addressed here, TCA provides the more interpretable framework.

      We also now state more explicitly that the absolute information-transfer values are low, and we correspondingly limit our conclusion: normal-hearing-trained representations generalize poorly to acute cochlear-implant-evoked responses in this framework. We further and thoroughly revised the abstract, results, and discussion so that any perceptual implications remain explicitly speculative, since perception was not measured directly in these experiments.

      (4) Finally, although methodological reporting is much improved, some verification remains indirect. The authors provide useful implantation criteria and cite prior validation of their deafening approach, but the manuscript would be clearer if it explicitly distinguished between validation performed in the present animals and validation based on previous cohorts. This distinction is important because surgical variability, implantation efficacy, and deafening completeness can influence the interpretation of cochlear implant experiments.

      We revised the manuscript to make this distinction explicit. In the methods, we now state clearly that we did not obtain ABR measurements, hair-cell counts, or behavioral confirmation of deafening in the animals used for the present iEEG dataset, and that support for deafening efficacy in this cohort therefore relies on prior validation of the same procedure in separate cohorts.

      We also now clarify that the ECAP forward-masking measurements were obtained in a separate acute cohort (N=3) and were not recorded from the animals used in the main iEEG dataset.

      Our goal in these revisions was to make the provenance of each validation step fully explicit and to avoid implying that all validation measures were obtained in the same animals used for cortical recording.

      Reviewer #2 (Public review):

      We thank the reviewer for highlighting the value of the decoder-based analyses while also challenging us to frame the study’s novelty more precisely and to make Figure 8 substantially clearer. In response, we narrowed our novelty claims, improved the explanation of the cross-modal decoding framework, and clarified the logic of the shuffled baseline.

      (1) The observation that responses to cochlear implant stimulation (stimulation) is spatially organized is not new (e.g. Adenis et al. 2024)

      Thanks, good point. We revised the Introduction to make clear that prior studies have already demonstrated spatial organization of cochlear-implant-evoked cortical responses, including prior animal work cited in the manuscript. Our intended contribution is therefore not the demonstration of spatial organization per se, but rather the use of a shared analytical framework to test whether acoustic and electrical stimulation produce overlapping or transferable spatiotemporal cortical population representations, including single-trial decoding and cross-modal generalization analyses.

      (2) The claim that spatial and temporal dimensions contribute information about the sound is also not new there is a large literature on this topic.

      We revised the manuscript to avoid implying novelty for this general principle. Our intended point is narrower: in the present work, spatial and temporal dimensions were recorded simultaneously across a 60-channel cortical surface array and integrated into trial-by-trial population analyses that could be directly compared across normal-hearing and cochlear-implant conditions within a shared framework.

      (3) The analyses supporting the claim that there is a mismatch between cochlear implant and sound representation are still unclear, particularly in Fig. 8.

      We revised both the Figure 8 legend and the results section to explain the analysis step-by-step. In particular, we now distinguish clearly among:

      The normal-hearing to normal-hearing re-optimization control, which tests whether the TCA/LDA framework can recover stimulus-related structure when modality is unchanged;

      The normal-hearing-trained / cochlear-implant-tested analysis, which tests cross-modal generalization; the shuffled baseline, which estimates chance-level information transfer by destroying any structured relationship between predicted labels and actual implant channels while preserving matrix dimensions.

      We now also explain explicitly why the shuffled baseline can equal or slightly exceed the measured normal hearing-to-cochlear-implant transfer. Because observed cross-modal transfer was extremely small, shuffling does not restore meaningful structure; rather, it shows that the measured transfer lies at or below chance level. We believe that this substantially improves the clarity and interpretability of Figure 8.

      Reviewer #3 (Public review):

      We thank the reviewer for recognizing the strengths of the study design, the single-trial decoding analyses, and the potential clinical relevance of this approach, while also pressing us to better address the limitations of monopolar stimulation, possible place-frequency mismatch, and the acute nature of the recordings. These comments prompted important revisions to both interpretation and discussion.

      (1a) The conclusion of the paper, especially the concept of distinct cortical encoding for each modality, is unfortunately partially supported by the results as the authors ignored fundamental limitations of CI related stimulation. First, the authors stimulated in a Monopolar mode which, albeit being clinically relevant, notoriously generates a high current spread in rodent models.

      We agree that monopolar stimulation is an important limitation, especially in the rodent cochlea, where current spread may be broader than in human subjects. We now acknowledge this more explicitly in the discussion.

      At the same time, our ECAP forward-masking measurements in a separate acute cohort provided supportive evidence for spatially and temporally tuned peripheral activation under the stimulus intensities used here. We therefore interpret the data as indicating that some peripheral selectivity was present, while also acknowledging that broader current spread under monopolar stimulation may have contributed to the coarse cortical organization observed in the implant condition.

      We also now state explicitly that the observed coarse cortical organization likely reflects a combination of peripheral current spread, downstream cortical pooling, and the spatial resolution limits of mesoscale surface iEEG, rather than any single factor alone.

      (1b) Comparing the averaged BF maps for iEEG (Fig-2A, C), BFs ranged from 4 to 16kHz with a predominance of 4kHz BFs. The lack of BFs at higher frequencies might reveal a potential location mismatch between the frequency range sampled at the level of the cortex (low to medium frequencies) and the frequency range covered by the CI inserted mostly in the first turn-and-a-half of the cochlea (high to medium frequencies). Looking at Fig2F (and to some extend 2A) most of CI electrodes elicited responses around the 4kHz regions and averaged maps show a predominance of CI-3-4 across cortex (Fig-2C, H and Sup Fig. 3) from areas with 4kHz BF to areas with 16kHz BF. It is doubtful that CI-3-4 are located near the 4kHz region based on Müller's work (1991) on the frequency representation in the rat cochlea.

      We appreciate this point and agree that a precise one-to-one alignment between cortical best-frequency maps and intracochlear electrode position cannot be established from the present data. We now clarify this limitation in the manuscript, noting that the iEEG-based maps are relatively coarse and appear to overrepresent mid-frequency regions, limiting direct inference from implant electrode number to a precise acoustic-frequency equivalent.

      We also emphasize that our central conclusion does not depend on assigning each implant electrode a specific best frequency, but rather on comparing the spatial organization and cross-modal generalization of cortical population responses.

      (1c) Moreover, Supplemental figure 3 shows that only a couple of CI electrodes are predominately represented at the level of the cortex. Thus, it seems possible that current spread ended stimulating indistinctly higher turns of the cochlea or even the modiolus in a non-specific manner, greatly reducing (or smearing) the placecoding/frequency resolution of each electrode, which in turn could explain the coarse topographic (or coarsely tonotopic according to the manuscript) organization of the cortical responses.

      We agree that the predominance of only a subset of implant electrodes in the cortical maps is consistent with relatively coarse peripheral and/or cortical place coding. We revised the discussion to acknowledge more explicitly that possible current spread, broad peripheral excitation, and the limited spatial resolution of surface iEEG could all contribute to the coarse spatial organization observed in the implant condition.

      We also emphasize in the revised text that this possibility does not negate the evidence for non-random organization, but it does limit the strength of any claim about finely graded cochleotopy.

      (2) Second, although the authors acknowledge that post-lingual CI users always have an adaptation period, their conclusion is based on measurements that are relatively "early" in the CI-use timeline so to speak since iEEG were collected a) acutely right after mono-aural implantation and stimulation, b) under anesthesia, c) using unmodulated pulse train fixed at 900pps regardless of the electrode used and thus lacking any temporal information shifts in relationship to electrode cochleotopic placement. Basically, all CI electrodes had the same rate whereas you would expect basal CI electrodes to be amplitude modulated at higher frequencies than apical electrodes.

      We agree. We now emphasize more clearly that the present experiments probe only an early and simplified stage of cochlear implant use. The recordings were acute, performed under anesthesia, and used short constant-rate pulse trains chosen to isolate fundamental organizing features of primary cortical responses rather than to reproduce the full complexity of clinical stimulation.

      We therefore frame our conclusions specifically in terms of acute cortical encoding under these conditions, and now state explicitly that these data do not address how chronic use, wakefulness, adaptation, or more complex modulation-rich stimuli may alter more global cortical representations.

      (3) As much as the reviewer likes the overall approach with the use of PCA-LDA and TCA, and agrees that information transfer seems inexistant at time of measurement, authors should be more careful in their strong conclusion that two distinct encoding exist. The non-overlapping between sound and electric stimulation representations might exist only transiently and this should be acknowledged a bit more in the discussion. Without repetition of iEEG measurement at later period with chronic use of the CI, it is not possible to definitively claim that two distinct, non-overlapping coding co-exist at all times.

      We agree and appreciate this point. We revised the Discussion to make this much more explicit. Our data support poor overlap or poor cross-modal generalization at the time of acute implant activation, but they do not establish that this relationship is fixed over chronic implant use.

      We now state directly that the observed mismatch may be transient, and that future longitudinal studies will be required to determine whether cortical representations become more aligned with experience, whether downstream readout adapts to a novel code, or whether both processes contribute.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The authors have provided new analyses to support the claim that CI and NH representations do not match. However, the analyses performed are presented in an unclear manner. Most particularly, it is very hard to understand what has been done in Fig. 8D and Fig. 8G. The legend in D is incomplete. The legend for G does not explain why there are lines with different color and what are the different lines. The reasoning behind the shuffled CI dataset is not explained. Why does shuffling restore information transfer, this is very counter intuitive and raises again the question of the value of this analysis.

      Agreed, we have revised the Figure 8 legend and rewritten the Figure 8 results section to initially explain the analysis workflow more clearly.

      We now explain that the shuffled condition is used to estimate a chance-level baseline for information transfer in the normal-hearing-trained / cochlear-implant-tested confusion matrix. Specifically, we compute mutual information after randomizing the predicted labels in the confusion matrix, thereby destroying any structured relationship between predicted tone labels and actual implant channels while preserving matrix dimensions and marginal structure (results, Fig. 8 section).

      We also now state explicitly that while this shuffled condition can yield slightly higher mutual information than the non-shuffled NH→CI condition, the point is not that shuffling restores meaningful decoding. Rather, the point is that the observed NH→CI transfer is so low that it does not exceed this chance-level baseline (results, Fig. 8 section).

      (2) Note that the legends of fig. 8 are mislabel (goes up to H although the last panel is G, a legend for E is missing).

      Thank you for catching this error. We have corrected the Figure 8 legend so that all panels are accurately labeled and described.

      Reviewer #3 (Recommendations for the authors):

      (1) Fig. 2C and 2H are from 2 to 16kHz (as announced in the previous responses) but then Sup. Fig. 3 goes from 2 to 32kHz. Still on Sup. Fig. 3, color code for CI electrodes is inverted compared to the rest of the MS.

      Thanks, good catches. We have updated Figures 2C and 2H to represent the full tested frequency range, and we have corrected the inverted color code for cochlear-implant electrodes in Supplemental Figure 3.

      (2) Fig. 2A. says n=1, Fig. 2F should say the same. Fig. 2C and 2H should also have n=1 on bottom map. Fig. 2D and 2I should mention n=1. Same thing with Fig. 3A/3C, Fig. 3B/3D (n=7), Fig. 7A/7B/7C.

      We have revised the relevant figure legends to make clear that data are from a single animal unless otherwise noted, while minimizing visual crowding in the figure panels.

      (3) The reviewer once again thinks it would be better to not truncate the axis of Figs. 4C, 6C, 7D as it creates confusion with Figs. 5D and 8G, where suddenly, there are 8 electrodes. The legend justification isn't enough.

      We appreciate this concern and have clarified the issue in the revised manuscript. In some animals (N=3), some electrodes in the 8-channel array were non-functional prior to implantation, so those animals contributed six rather than eight usable implant channels. We now make this explicit in the manuscript and supplemental figures, including the number of channels used in the respective animals in the methods under “Cochlear implant programming” and in Supplemental Figure 3. 

      (4) Although the reviewer appreciated the justification of 15 PCA for their analysis, this justification should be provided as is in the methods, especially as the Sup. Fig. 4 does not even explain the significance of the linear regression of components 16 to 30. Some other reviewers would say that 10 components were more than enough without this justification.

      We agree and have provided further justification of the 15 PCA components to provide justification as is in the Methods “Principal component analysis” and in the results describing Figure 4. 

      (5) Sup. Fig. 1 and the eCAP measurements are a nice and needed addition to the MS but the authors should make it clearer that it came from 3 animals that are not part of the cohort presented in the rest of the paper / in Sup. Fig. 2. The method section certainly not make that distinction, giving the impression that all animals have been tested for channel interactions and temporal recovery.

      We agree and have revised the methods section to state clearly that the ECAP forward-masking measurements were obtained in a separate group of acutely implanted rats (N=3), rather than the main iEEG cohort in Supplemental Figure 1.

      (6) As said in the public review, the authors should acknowledge more the potential transiency of the nonoverlapping representation in the discussion, especially since most of the recordings are acute.

      We agree, and we have revised the discussion accordingly. As described in our public response above, we now state directly that the poor overlap between acoustic and electrical representations was observed under acute recording conditions and may not persist unchanged with chronic implant use or behavioral adaptation.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study aimed at replicating two previous findings that showed (1) a link between prediction tendencies and neural speech tracking, and (2) that eye movements track speech. The main findings were replicated which supports the robustness of these results. The authors also investigated interactions between prediction tendencies and ocular speech tracking, but the data did not reveal clear relationships. The authors propose a framework that integrates the findings of the study and proposes how eye movements and prediction tendencies shape perception.

      Strengths:

      This is a well-written paper that addresses interesting research questions, bringing together two subfields that are usually studied in separation: auditory speech and eye movements. The authors aimed at replicating findings from two of their previous studies, which was overall successful and speaks for the robustness of the findings. The overall approach is convincing, methods and analyses appear to be thorough, and results are compelling.

      Weaknesses:

      Eye movement behavior could have presented in more detail and the authors could have attempted to understand whether there is a particular component in eye movement behavior (e.g., blinks, microsaccades) that drives the observed effects.

      Reviewer #2 (Public review):

      Summary

      Schubert et al. recorded MEG and eye tracking activity while participants were listening to stories in single-speaker or multi-speaker speech. In a separate task, MEG was recorded while the same participants were listening to four types of pure tones in either structured (75% predictable) or random (25%) sequences. The MEG data from this task was used to quantify individual 'prediction tendency': the amount by which the neural signal is modulated by whether or not a repeated tone was (un)predictable, given the context. In a replication of earlier work, this prediction tendency was found to correlate with 'neural speech tracking' during the main task. Neural speech tracking is quantified as the multivariate relationship between MEG activity and speech amplitude envelope. Prediction tendency did not correlate with 'ocular speech tracking' during the main task. Neural speech tracking was further modulated by local semantic violations in the speech material and by whether or not a distracting speaker was present. The authors suggest that part of the neural speech tracking is mediated by ocular speech tracking. Story comprehension was negatively related with ocular speech tracking.

      Strengths

      This is an ambitious study, and the authors' attempt to integrate the many reported findings related to prediction and attention in one framework is laudable. The data acquisition and analyses appear to be done with great attention to methodological detail. Furthermore, the experimental paradigm used is more naturalistic than was previously done in similar setups (i.e.: stories instead of sentences).

      Weaknesses

      While the analysis pipeline is outlined in much detail, some analysis choices appear ad-hoc and could have been more uniform and/or better motivated (other than: this is what was done before).

      Reviewer #3 (Public review):

      I thank the authors for their extensive revision of this paper, and I found some elements greatly improved.

      In particular, the authors do embrace a somewhat more speculative tone in the current version, which I think is fitting for this work, as the data seem (to me) to be not fully conclusive. The data set collected here is clearly valuable and unique (and I would encourage the authors to make it publicly available!), however, my overall impression is that the specific analyses reported here might not fully

      Despite the revised description of methods, results and figures, I still have trouble understanding many of the results and the authors conclusive interpretation of them. These are my main reservations:

      (1) Regarding "individual prediction tendency" - thank you for adding clarifying methodological details and showing the data in a new Figure (#2). Honestly, however, I still can't say that I fully understand the result. For example, why is there also a significant response in the random condition as well? And how do you interpret the interesting time-course (with a peak ~200ms prior to the stimulus, and a reduction overtime from there? Also (I may have missed this, but..) what neural data was used to train the classifier and derive the "prediction tendency" index? Was it just the broadband neural response? Is there a way to know which sensors contributed to this metric (e.g., are they predominantly auditory? Frontal?)? And is there a way to establish the statistical significance of this metric (e.g., how good the decoder actually was in predicting behavioral sensitivity?). I don't see any statistics in the results section describing the individual prediction tendency.

      (2) Regarding the TRF analysis - Thanks for clarifying the approach used to obtain 2-second long "segments" of speech tracking. This is an interesting approach, however I think quite new(?), and for me it raises a whole new set of questions, as well as additional controls and data that I would have liked to see, to be convinced that results are significant. I will elaborate:

      - Do I understand correctly that you segment the real and predicted neural response into 2-second long segments and then calculate the Pearsons' correlation between them to assess the goodness of the model? This is very unclear, since in the methods section you state only that "the same" analysis was performed as for the full data - but what exactly? Clearly, values will be very different when using such short segments. I feel that additional details are still required (and perhaps data shown) to fully understand the "semantic violation" analysis of TRFs.

      - I would like to reiterate my previous comment regarding the use of permutation tests to verify the validity of TRF-based measures derived. This would be especially important when using new approaches (such as the segmentation used here). The authors argue that this is not needed since this was not done in their previously published study. However, this sounds a bit like "two wrongs make a right" argument... why not just do it, and let us know that this 2-second segmentation approach allows estimating reliable speech tracking?

      - Following up on my previous comment that defining "clusters" as at least two neighboring channels (Figure 3) - the fact that this is a default in Fieldtrip is by no means sufficient justification! This seems quite liberal to me, especially given the many comparisons performed. Here too, permutations can help to determine the necessary data-driven threshold for corrections. This is of course critical for interpreting the result shown in Figures 3E&G that are critical "take home messages" of the paper - i.e., that the prediction-index from the first part of the experiment is related to speech tracking in the second part of the experiment. To my eyes, this does not look extremely convincing, but perhaps the authors can show more conclusive data to support this (e.g., scatter plots of the betas across participant?). - A similar point can be made for the effect of semantic violations (though here the scalp-level result is somewhat more clustered). The authors point out that the semantic effect is a "replication" of their result reported in Schubert et al. 2023, but if I am not mistaken the results there were somewhat different (as was the manipulation). It would be nice to explicitly discuss the similarity/difference between these effects.

      (3) Regarding the ocular-TRFs -

      - Maybe this is just me, but I believe that effects that are robust should be clearly visible in the data, without the need for fancy "black-box" statistical models. In the case of the ocular TRFs, it is hard for me to see how these time-courses are not just noise (and, again, a permutation test would have helped to convince me.). The inconsistent results for horizontal and vertical eye-movements vis a vis the experimental conditions (single vs. multi-speaker conditions) don't help either, despite the authors argument that these are "independent" - but why should this be the case, especially if there is nothing really to look at in this task? - I remain with this scepticism for the mediation-portion of the analysis as well... But perhaps replications from other groups or making the data public will help shed further light on this in the future.

      Minor

      - Thanks for adding information about the creation of semantic-violation stimuli. Since the violations and lexical-controls were taken from different audio recordings, it would have been nice to verify that differences between neural responses cannot be attributed to differences in articulations (e.g., by comparing their spectro-temporal properties)

      We are grateful to the Reviewing Editor and the three reviewers for the care and time they have devoted across the review rounds; their input has strengthened the paper. We are happy to confirm that we have carried out the two additional analyses the Assessment identifies as the path to a "solid" rating for methodological rigor. In brief:

      (1) Multiple-comparison correction for the prediction-tendency effect (Figure 3). We would like to be transparent about why our implementation differs from the specific procedure that was suggested. Testing the relationship between prediction tendency and neural speech tracking properly requires a mixed-effects model: the nested, repeated-measures structure of the data demands a per-subject random intercept (1|subject) to absorb the between-subject variability that the fixed effects do not capture. This is a standard and widely recommended approach for nested data, not an exotic one. Crucially, to our knowledge no implementation of a cluster-based permutation test exists for mixed-effects models — so the permutation/cluster procedure that was recommended cannot be applied to the very model the data structure requires. This constraint, rather than preference, is what led us to Bayesian estimation in the first place. We would add that the effect replicates Schubert et al. (2023), with a consistent left-frontal localisation across studies, which further speaks to its robustness.

      We fully share the concern underlying the recommendation, however — that a Bayesian model must equally guard against multiple comparisons across channels — and we have addressed it directly within the framework the data require. We tightened the HDI inclusion criterion to be analogous to a frequentist multiple-comparison correction (with boundaries comparable to a Bonferroni-corrected p-value) and re-estimated the model (encoding ~ prediction tendency * condition + (1|subject)) with 200,000 draws per chain (previously 2,000) to obtain stable posterior tails. The spatial pattern reproduces Figure 3E. At a Bonferroni-comparable per-channel boundary — itself widely regarded as overly conservative, just as the conventional 5% threshold is often criticised as arbitrary — no single channel survives in isolation; but considering the three-channel cluster jointly, the accumulated posterior places only ~0.7% of its mass at or below zero, i.e. a ~99.3% posterior probability of a positive effect (see Author response image 1). Rather than commit to a single arbitrary boundary, we report this accumulated posterior and invite the reviewers and readers to apply whatever criterion they consider appropriate.

      Author response image 1.

      (2) Circular-shift procedure for the mediation analysis (eye movements). Following Reviewer 1's suggestion, we replaced the previous shuffling control with a circularly shifted eye-movement predictor (shifted by half its total duration) as the control model. This isolates the genuine envelope contribution from any reduction in envelope weights that arises simply from adding a second predictor, and provides a stricter test of the specific temporal relationship between eye movements and the speech envelope. The mediation results are robust under this procedure (now Figure 5): vertical eye movements significantly mediate neural tracking of clear speech across all three principal components, and horizontal eye movements mediate tracking of target speech in the multi-speaker condition, with an early, sustained component peaking at ~0.18 s. Importantly, this stricter control did not materially change the interpretations we draw from these results: the pattern of mediation, and the conclusions about ocular speech tracking, remain consistent with those obtained under the previous approach.

      We have chosen to focus this revision specifically on these two points, which the Assessment identifies as what is needed to move the work from "incomplete" to "solid." We are sincerely grateful for the many further suggestions the reviewers raised, several of which we agree are thoughtful and point to worthwhile directions for future work. Since the previous round, however, the circumstances of the two authors who led this work have changed substantially: the corresponding author has since left academia altogether, and the shared first author's professional circumstances have likewise changed considerably. We are therefore not in a position to take on further analyses, or to prepare a point-by-point reply to the remaining comments, beyond the two above; we hope the editors and reviewers will understand that this necessarily defines the final scope of our revision. We remain sincerely grateful for the time and care they have invested in the manuscript.

    1. Author response:

      We are grateful to the reviewers and eLife editors for providing thoughtful and constructive commentary on our manuscript. The major concerns raised during review center on (A) the broader impact of methionine depletion on cellular metabolism beyond protein synthesis, including potential effects on cellular bioenergetics, and B) the extent to which assay conditions may induce cellular stress that influences downstream functional measurements. We will address both points during the revision period through experiments we have outlined below, most of which are already underway. 

      Related to (A), we will define the metabolic consequences of methionine deprivation on cellular metabolism with greater precision and are currently optimizing a third non-canonical alkyne amino acid, β-ethynylserine (β-ES), for labeling surface-exposed proteins in living cells. β-ES is a clickable threonine analogue that is efficiently incorporated into the proteome in the presence of physiological threonine and therefore does not require metabolic deprivation (1). We are in the process of optimizing β-ES for the MARBL workflow as a substitute for homopropargylglycine (HPG), which reduces the duration of methionine deprivation from six hours to two hours. Thus, β-ES eliminates the SAM-depletion concern for the baseline-translation half of the MARBL ratiometric measurement and constitutes a genuine methodological advance beyond the original submission. 

      Related to (B), we plan to generate additional data benchmarking MARBL against Seahorse and other established assays for measuring cellular bioenergetics, while also performing phenotypic and functional cellular characterization at key stages of the workflow. Experiments that are planned during the revision period are described in our point-by-point response below.

      Public Reviews:

      Reviewer #1 (Public review):

      The idea behind this paper is to have an alternate, reliable and quantitative approach to assess cell-cell metabolic heterogeneity. This study tries to achieve that using translationally-coupled energetic responses to metabolic stress. This is interesting because, in general, most quantitative measurements of metabolic outputs are 'bulk' and average for many cells. To overcome this, many recent studies use some read-outs of translation (presuming that translation is the single major energy sink in cells - however, this is objectively correct only in rapidly proliferating cells). That said, the authors take an interesting approach - to use two clickable methionine analogs, and assess baseline vs metabolically coupled translation within the same cell.

      The highlight is the methodology development where two distinct, clickable CMAs are used (to replace methionine in proteins). The labelling and approach are clever, and can be useful if carefully used. But there is going to be a challenge in using this, since the depletion of methionine itself (required for labelling), and a bias towards incorporation in some proteins (because these reagents are not highly permeable) will make it challenging to obtain precise metabolic state information - which otherwise can be obtained directly and far more precisely using a combination of other methods (ATP/flux measurement, respiratory capacity, translation rates etc). I therefore only make broader comments in my review below - to help structure this study better, clearly identify key limitations (and there are several that are clearly seen), and better clarify what MARBL may be useful for.

      (1) These reagents used for MARBL are not highly cell permeable/transported into cells, and largely work on the surface proteome. Which means the measurements related to changes in translation are indirect - quantified based on changes on the surface proteome (seen with labelling), and not the entire proteome.

      We appreciate the reviewer’s careful consideration of the MARBL labeling strategy and the opportunity to clarify an important feature of the method. We interpret this comment to mean that the detection reagents used in MARBL are not cell-permeable, rather than that the methionine analogues AHA and HPG are inefficiently transported into cells. HPG and AHA are transported through the ubiquitously expressed sodium-dependent neutral amino acid transporter SLC1A5 (2), and are incorporated into the proteome by intracellular translational machinery. The bias identified by the reviewer arises from the click detection step, not from metabolic labeling. We agree with the reviewer that MARBL measures the appearance of a subset of newly synthesized proteins where the amino acid analogue is accessible at the cell surface rather than the nascent production of the entire proteome. However, this is an intentional feature of MARBL and represents a key technical advance. It enables MARBL to generate a translation-dependent signal in live cells without requiring destructive interventions like fixation and permeabilization to access the entire proteome, as is required for other approaches such as BONCAT, THRONCAT, CENCAT, and SCENITH (1,3–5). Importantly, we empirically demonstrate that the surfaceaccessible fraction of the nascent proteome provides sufficient signal for quantitative measurements of translation that respond predictably to inhibition of protein synthesis as well as to metabolic perturbations that alter cellular energetics (Main Figure 1D-G, 2C). MARBL is also readily compatible with fixation and permeabilization, allowing the same labeling strategy to quantify analogue incorporation when a measure of total protein synthesis is desired and there is no need for the recovery of live cells. In the revised manuscript, we will include data directly comparing MARBL surface labeling with total nascent protein synthesis measured following fixation and permeabilization to show that these signals track with one another.

      (2) A major limitation - which will confound any interpretation- is the need to use methionine-free media; this can be a problem beyond protein synthesis/met incorporation since this will almost instantly deplete SAM pools in cells. There is little data provided on the impact of using this approach on SAM pools (time kinetics, how quickly SAM pools are affected, how much of the impact on metabolism comes from purely that, etc).

      This is important to establish because (i) of the continuous, very high flux of SAM -> SAH (eg. in the folate pathway, other methylations), and a constant need for SAM synthesis from methionine. This information will set the limits of capabilities of this method, as well as help delineate how much you can interpret results related to metabolic states between compared cells/states, etc. The labelling process is ~2 hours, while effects on SAM can be seen within minutes of methionine starvation in media, in metabolically active cells.

      The reviewer is correct in pointing out that cellular SAM pools are dynamic and adapt rapidly to methionine availability within hours of its deprivation or add-back (6,7). In the current version of the manuscript, we observe an ~100-fold decrease in intracellular SAM with 2 hours of methionine-free incubation in Jurkat cells, corresponding to the same time window over which we label with AHA in MARBL (currently presented as a min-max scaled heat map in Supplementary Figure 1A-B). We agree this needs to be presented more clearly as a potential limitation and quantified more explicitly. To address this, we plan to (i) reformat the existing SAM/ SAH/ methionine kinetics as absolute peak areas (Author response image 1A-C), (ii) add a novel dual-isotope simultaneous measurement of ATP turnover and SAM turnover across a panel of human cell lines to define the extent to which changes in SAM metabolism during the MARBL labeling workflow are likely to influence cellular energy demand (using <sup>13</sup>C<sub>5</sub>-methionine and H<sub>2</sub><sup>18</sup>O), and (iii) add β-ES as a threonine-based alternative to HPG for measuring baseline translation that substantially reduces the duration of methionine deprivation required for MARBL. β-ES is a clickable, alkyne-modified threonine analogue that is incorporated into newly synthesized proteins by mammalian cells in complete medium (1,4). Compared to methionine, which is a major part of the methionine cycle and 1-carbon metabolism that are required for cell proliferation and survival, threonine plays a less extensive role in mammalian intermediary metabolism. We plan to optimize and validate β-ESàAHA as well as AHAàβ-ES dual labeling and determine whether this modified MARBL workflow preserves the dynamic range and metabolic responsiveness of the HPGàAHA approach as in the original submitted manuscript. These experiments are currently underway, and we have already found that β-ES produces robust surface signal above background within 2-4 hours of labeling using our established MARBL protocol in the presence of normal threonine levels (Author response image 2A-D). We are currently optimizing the surface labeling protocol to further reduce the duration of baseline β-ES incubation in the revised manuscript. We are unable to eliminate methionine depletion entirely from the MARBL workflow. This is because endogenous methionyl-tRNA synthetases possess drastically higher affinities for canonical methionine (8–10), which prevents the incorporation of methionine analogues into proteins. However, these planned experiments will better define the impact of methionine limitation while also providing an alternative MARBL implementation that restricts methionine withdrawal to the shorter AHA-labeling window.

      Author response image 1.

      Dynamics of intracellular methionine and its derived metabolites. (A-C) LC-MS/MS raw peak areas in Jurkat cells of methionine (A), SAM (B), and SAH (C) after 2, 4, 6, or 8 hours of incubation in Met-free RPMI media supplemented with Met, AHA, or HPG. Abbreviations: AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, SAM = Sadenosylmethionine, SAH = S-adenosyl-L-homocysteine. Statistics: Graphs display mean ± SD (A-C).

      Author response image 2.

      Extension of live surface labeling protocol using clickable threonine analogue to monitor surface translation. (A) Chemical structures for threonine and its clickable analogue β-ES. (B-C) Representative flow cytometric histograms (B) and corresponding gMFIs of extracellular Azd-647 signal after 1, 2, or 4 hours of β-ES incorporation. (D) gMFIs of Alk-647 (corresponding to AHA) or Azd-647 (corresponding to HPG or β-ES) after 1, 2, or 4 hours of incorporation. Abbreviations: β-ES = β-Ethynylserine, AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, CHX = Cycloheximide, gMFI = Geometric mean fluorescence intensity. Statistics: Graphs display mean ± SD (C-D).

      (3) What is the effect on overall adenylate charge/ATP due to shifting to methionine-free media + label addition? How does it vary between the cells tested (suspension vs adherent)? This should be established before data related to Fig. 2. Does it correlate with extent of AHA incorporation?

      The primary conclusion that this method suitably reflects overall changes in energetics comes from the titration of 2DG/glycolytic inhibition.

      We thank the reviewer for raising these important points, which are related to point #2 above, as both ask how methionine-free labeling conditions alter cellular metabolism. We agree that a more comprehensive set of experiments exploring methionine deprivation will better delimit interpretations that can be drawn from a MARBL assay related to cellular energetics. As described in our response to point #2, we will measure baseline adenylate energy charge across a panel of adherent and suspension cell lines under methionine-replete, methionine-free, and methionine-free +AHA/HPG labeling conditions and determine its relationship to AHA incorporation. We will also perform analogous experiments during β-ES supplementation to determine how this alternative labeling condition affects cellular energetic state and β-ES incorporation. 

      We also wish to clarify that the adenylate energy charge measurements in Main Figure 2D and the measurements of newly synthesized ATP by H<sub>2</sub><sup>18</sup>O turnover in Main Figure 2E were performed in methionine-free media supplemented with either AHA or methionine. We apologize that this was not made clear in the figure key and legend. In addition, we inadvertently covered the methionine-replete control in the original figure with the key. This has been corrected in Author response image 3A-B and will be updated in the revised manuscript. In Jurkat cells, both adenylate energy charge and ATP synthesis are similar in the presence and absence of methionine.

      We note, however, that cells with intact energy-generating systems maintain adenylate energy charge within a relatively narrow range (11). Given this buffering capacity, we do not expect methionine deprivation to produce a large change in adenylate energy charge in most cell types. 

      Author response image 3.

      Validation of surface translation as a readout of cellular energetics in methionine-free and AHA-supplemented media. (A) Correlation between changes in adenylate energy charge versus normalized Alk-647 Alkyne flow cytometric signal in methionine-free RPMI supplemented with either AHA or Met (Pearson’s R = 0.9044, Pearson’s R<sup>2</sup> = 0.8179, p = 0.0052). (B) Correlation between changes in newly synthesized ATP versus normalized Alk-647 Alkyne flow cytometric signal in methionine-free RPMI supplemented with either AHA or Met (Pearson’s R = 0.8854, Pearson’s R<sup>2</sup> = 0.7840, p = 0.008). Abbreviations: AHA = Azidohomoalanine, HPG = Homopropargylglycine, Met = Methionine, CHX = Cycloheximide, gMFI = geometric mean fluorescence intensity, 2DG = 2-Deoxy-D-Glucose, Omy = Oligomycin A. Statistics: Pearson’s correlation coefficient was used for correlation and significance (A-B). The Met + DMSO condition was excluded from the Pearson correlation analysis. Graphs display mean ± SD (A-B).

      (4) Relatedly, if this label incorporation experiment is carried out (for ~2 hrs), and subsequently there is a washout/replacement with fresh, methionine-supplemented medium, (how quickly) do the cells recover and restore their energetic allocations?

      If we observe substantial changes in adenylate energy charge associated with methionine restriction, as assessed in the experiments planned in Points #2 and #3 above, we will perform methionine add-back experiments to define the kinetics of recovery. In addition, our plan to provide an alternative workflow that replaces HPG with β-ES will reduce the total duration of methionine depletion and further address this concern.

      (5) One possible advantage of a system like this can be to address questions in single cells/study cell metabolic heterogeneity. However, these are best done if the attaching moiety has a (selective) fluorescence increase and/or other read-out that can be quantitatively obtained at a single cell level. Largely, using AHA or HPG effectively only leads to bulk estimates (which can be sub-sorted towards single-cell estimates indirectly). This means that this method cannot really be used to study cell-cell metabolic heterogeneity effectively - compared to far simpler approaches, for example using a mitochondrial potentiometric dye with high fluorescence, or reporters for glycolytic activity, etc. This would also be related to Figure 5 - at best, this approach may complement existing approaches towards identifying heterogeneous sub-populations of cells. 

      However, I do agree that MARBL is flexible, stable, and can be internally normalised and used through flowbased platforms. It can supplement existing approaches to perturb bioenergetics, and also supplement existing approaches to understand metabolic state in live cells, particularly in suspension cells.

      We thank the reviewer for raising this thoughtful point, which we will address with textual changes in the discussion. We note that MARBL provides a quantitative single-cell readout when flow cytometry is used as the analytical endpoint, because the dual-color labeling scheme provides an internal control for baseline translation that accounts for variation in protein synthesis rates independent from cellular energetics. This permits a single metabolic resilience index to be calculated per cell and interpreted either at single-cell resolution or after grouping cells into sub-populations as a bulk estimate. Both approaches have utility depending on the biological question and intended downstream application. We agree with the reviewer that MARBL can be paired with other reporters for metabolism to obtain a more comprehensive understanding of bioenergetic heterogeneity and appreciate the reviewer’s recognition of its value as a complementary approach for studying metabolic state in live cells.

      Reviewer #2 (Public review):

      Summary: 

      Delacruz et al. describe a new method, called “MARBL” (Methionine Analogues for Ratiometric Bioenergetics in Live cells) to measure metabolic activity in single cells. The concept is similar to the SCENITH (anti-puromycin flow cytometry) assay to measure energy metabolism by measuring protein translation activity, yet offers, in theory, two advantages: 1) it keeps cells alive for downstream biological assays and 2) it is a ratiometric measurement, measuring both baseline translation and translation in the presence of metabolic inhibitors to correct for inherent cell-to-cell translation differences.

      Specifically, this method takes advantage of two click-chemistry-active methionine analogs, and then clicks fluorophores onto newly-synthesized surface proteins that have incorporated these analogs to measure translational activity. One methionine analog is given to cells for 2-4 hours to measure baseline translational activity, then metabolism is blocked using 2-deoxyglucose and oligomycin and the second methionine analog given to measure “metabolically-linked” translation activity. The authors establish this technique and show that mouse T cells polarized as pathogenic Th17 cells are more translationally active (“resilient”) compared to nonpathogenic Th17 cells, and when sorted, the resilient cells produce more interferon-gamma. This latter finding requires live cells after the metabolic measurement assay, showcasing findings that are inaccessible to the SCENITH assay.

      Strengths:

      The approach used is conceptually clever. It is appealing to measure metabolic/translational activity and to then be able to carry out further assays on sorted cell populations with different degrees of metabolic activity. This would indeed represent a useful advance.

      We thank the reviewer for recognizing the conceptual strengths of MARBL and its potential to enable downstream analysis of live cell populations with distinct metabolic states.

      Weaknesses:

      In principle, one key benefit of this technique is that cells can be used for biological assays after the metabolic measurement. Indeed, this would represent a valuable tool in the field.

      However, in this technique, cells are subjected to methionine deprivation, addition of non-natural methionine analogs, click chemistry, and high doses of toxic metabolic inhibitors 2-deoxyglucose and oligomycin. Indeed, the authors show in Figure S5F that 1/3 more of the post-MARBL cells die relative to cells not subject to this technique (60% viability in unclicked control, 40% in MARBL-measured cells). This data suggests that cells after this technique may be stressed and not reflective of the biological function of unmanipulated cells. More controls on viability and cell function (e.g. cytokine production) at more time points after the MARBL assay would have been valuable to address this issue.

      The reviewer makes important points regarding the conditions needed for the workflow of a MARBL assay that may impact cellular fitness. Perturbational methods, particularly techniques that interrogate bioenergetics, inherently require media-based or pharmacologically induced stress to evaluate cellular responses. Still, we agree that comprehensively profiling the fitness of cells following a MARBL assay and sorting is important since our technique aims to link cellular bioenergetics to functional outcomes. First, we would like to highlight that the MARBL-processed pathogenic and non-pathogenic TH17 cells depicted in Main Figure 5 were rested overnight after Fluorescence-Activated Cell Sorting (FACS) in complete RPMI medium prior to the restimulation assay. This is in line with standard practice to allow cells to recover from the shear stress of sorting prior to subsequent experiments (12,13). In addition, both pathogenic and non-pathogenic TH17 cells maintained the expression of lineage-defining transcription factors throughout the MARBL workflow, as shown by analyzing rested cells stained with antibodies for T-bet (expressed by pathogenic TH17) as well as RORγt (expressed in both cell types) by flow cytometry (Author response image 4A-C). To address this in the revised manuscript, we plan to repeat our non-pathogenic and pathogenic TH17 dual-MARBL and sorting experiment (Main Figure 5A-B) and rest sorted cells in RPMI with IL-2 for longer periods of time (24 or 48 hours) before restimulation for viability and cytokine analysis. These controls will provide more information on cellular fitness throughout the MARBL workflow, and we appreciate the reviewer’s suggestion.

      Author response image 4.

      Expression of lineage-defining transcription factors is maintained post-MARBL processing and fluorescence-activated cell sorting (FACS). (A) Experimental schematic. Ex vivo differentiated pTH17 and npTH17 cells were stained with CD45.2 antibodies conjugated to different color fluorophores, processed via MARBL, mixed at a 1:1 ratio, sorted, and then re-cultured in IL-2-supplemented RPMI media. After resting overnight, the expression of RORγt and T-bet were evaluated by intracellular staining and flow cytometry, distinguishing pTH17 from npTH17 cells based on prior CD45.2 staining. (B-C) Frequency of RORγt (B) and T-bet (C) positivity in murine Th17 cells post-MARBL processing, sorting, and overnight rest in IL-2-supplemented media. Abbreviations: npTH17 = non-pathogenic TH17, pTH17 = pathogenic TH17, HPG = Homopropargylglycine, AHA = Azidohomoalanine, Met = Methionine, 2DG = 2-Deoxy-D-Glucose, Omy = Oligomycin A. Statistics: Graphs display mean ± SD (B-C).

      Another weakness of the paper is limited benchmarking against established metabolic assays in the field. The main assays used currently in the field are SCENITH and Seahorse. The authors do not compare their findings to SCENITH. They do compare their results to Seahorse, but the data shown don't address the key question: how does energy production measured by Seahorse, say in unmanipulated vs 2dg+oligomycin-treated cells, compare to the MARBL measurement? (Instead, they show a calculated "glucose dependence" metric in cells subjected to low vs high inhibitor dose, not showing the underlying data or cells that didn't receive an inhibitor).

      We appreciate the opportunity to perform additional benchmarking to define how the MARBL signal compares to existing methods for measuring cellular energetics. For our initial validation, we chose to benchmark the single-color MARBL signal against direct LC-MS/MS quantification of adenylate energy charge as well as newly synthesized ATP (Main Figure 2A-E). Although Seahorse and SCENITH are widely used standards in the field, these assays still provide indirect measures of cellular energetics, whereas LC-MS/MS quantifies the high-energy nucleotide pools that dictate cellular energy status. Both the adenylate energy charge and newly synthesized ATP decreased in response to increasing concentrations of 2-deoxy-D-glucose (2DG) and oligomycin A (Omy) treatments that impair ATP regeneration, and this energetic response displayed a linear relationship with the AHA click signal (Main Figure 2D-E). Nevertheless, we agree that additional benchmarking suggested by the reviewer will strengthen the methodological foundation of MARBL and help users understand how MARBL measurements relate to those obtained using more established metabolic assays. For this purpose, it is important to account for differences in what each assay measures. Seahorse resolves oxidative and glycolytic activity through simultaneous measurements of OCR and ECAR, whereas MARBL (as well as SCENITH and CENCAT) integrate the energetic contributions of these pathways into a single translation-dependent readout. This is the reason why we compared MARBL with Seahorse using the glucose-dependence calculation employed by SCENITH in the current version of the manuscript (Main Figure 2FH) (5). As suggested by the reviewer, we will assess how OCR and ECAR measured by Seahorse vary relative to the MARBL signal across different oligomycin and 2-deoxyglucose treatment conditions, providing a more comprehensive view of the bioenergetic responses captured by MARBL. 

      References

      (1) Ignacio BJ, Dijkstra J, Mora N, Slot EFJ, van Weijsten MJ, Storkebaum E, et al. THRONCAT: metabolic labeling of newly synthesized proteins using a bioorthogonal threonine analog. Nat Commun. 2023 Jun 8;14(1):3367. doi:10.1038/s41467-023-39063-7 PubMed PMID: 37291115; PubMed Central PMCID: PMC10250548.

      (2) Pelgrom LR, Davis GM, O’Shaughnessy S, Wezenberg EJM, Van Kasteren SI, Finlay DK, et al. QUAS-R: An SLC1A5-mediated glutamine uptake assay with single-cell resolution reveals metabolic heterogeneity with immune populations. Cell Reports. 2023 Aug 29;42(8):112828. doi:10.1016/j.celrep.2023.112828

      (3) Dieterich DC, Link AJ, Graumann J, Tirrell DA, Schuman EM. Selective identification of newly synthesized proteins in mammalian cells using bioorthogonal noncanonical amino acid tagging (BONCAT). Proceedings of the National Academy of Sciences. 2006 Jun 20;103(25):9482–7. doi:10.1073/pnas.0601637103 PubMed PMID: 16769897.

      (4) Vrieling F, van der Zande HJP, Naus B, Smeehuijzen L, van Heck JIP, Ignacio BJ, et al. CENCAT enables immunometabolic profiling by measuring protein synthesis via bioorthogonal noncanonical amino acid tagging. Cell Rep Methods. 2024 Oct 21;4(10):100883. doi:10.1016/j.crmeth.2024.100883 PubMed PMID: 39437716; PubMed Central PMCID: PMC11573747.

      (5) Argüello RJ, Combes AJ, Char R, Gigan JP, Baaziz AI, Bousiquot E, et al. SCENITH: A flow cytometry based method to functionally profile energy metabolism with single cell resolution. Cell Metab. 2020 Dec 1;32(6):1063-1075.e7. doi:10.1016/j.cmet.2020.11.007 PubMed PMID: 33264598; PubMed Central PMCID: PMC8407169.

      (6) Mentch SJ, Mehrmohamadi M, Huang L, Liu X, Gupta D, Mattocks D, et al. Histone Methylation Dynamics and Gene Regulation Occur through the Sensing of One-Carbon Metabolism. Cell Metab. 2015 Nov 3;22(5):861–73. doi:10.1016/j.cmet.2015.08.024 PubMed PMID: 26411344; PubMed Central PMCID: PMC4635069.

      (7) Chen Z, Chen W, Reheman Z, Jiang H, Wu J, Li X. Genetically encoded RNA-based sensors with Pepper fluorogenic aptamer. Nucleic Acids Res. 2023 Sep 8;51(16):8322–36. doi:10.1093/nar/gkad620 PubMed PMID: 37486780; PubMed Central PMCID: PMC10484673.

      (8) Kiick KL, Saxon E, Tirrell DA, Bertozzi CR. Incorporation of azides into recombinant proteins for chemoselective modification by the Staudinger ligation. Proc Natl Acad Sci U S A. 2002 Jan 8;99(1):19–24. doi:10.1073/pnas.012583299 PubMed PMID: 11752401; PubMed Central PMCID: PMC117506.

      (9) Beatty KE, Liu JC, Xie F, Dieterich DC, Schuman EM, Wang Q, et al. Fluorescence visualization of newly synthesized proteins in mammalian cells. Angew Chem Int Ed Engl. 2006 Nov 13;45(44):7364–7. doi:10.1002/anie.200602114 PubMed PMID: 17036290.

      (10) Kiick KL, Weberskirch R, Tirrell DA. Identification of an expanded set of translationally active methionine analogues in Escherichia coli. FEBS Lett. 2001 Jul 27;502(1–2):25–30. doi:10.1016/s0014-5793(01)02657-6 PubMed PMID: 11478942.

      (11) De la Fuente IM, Cortés JM, Valero E, Desroches M, Rodrigues S, Malaina I, et al. On the dynamics of the adenylate energy system: homeorhesis vs homeostasis. PLoS One. 2014;9(10):e108676. doi:10.1371/journal.pone.0108676 PubMed PMID: 25303477; PubMed Central PMCID: PMC4193753.

      (12) Pollizzi KN, Patel CH, Sun IH, Oh MH, Waickman AT, Wen J, et al. mTORC1 and mTORC2 selectively regulate CD8<sup>+</sup> T cell differentiation. J Clin Invest. 2015 May 1;125(5):2090–108. doi:10.1172/JCI77746 PubMed PMID: 0.

      (13) Roth TL, Puig-Saus C, Yu R, Shifrut E, Carnevale J, Li PJ, et al. Reprogramming human T cell function and specificity with non-viral genome targeting. Nature. 2018 Jul;559(7714):405–9. doi:10.1038/s41586-018-03265

  3. Aug 2026
    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors are trying to characterize the sources of H. pylori in an island system. They find that, like the humans, the bacteria are admixed, but there is no correlation within the island between human ancestry and bacterial ancestry.

      We thank the reviewer for their consideration of our work.

      Strengths:

      The study has taken particular care to characterize the humans from which isolates were obtained. Thus, it is a particularly convincing demonstration of a bacterial "melting pot".

      Weaknesses:

      The GWAS is highly confounded by population structure. In particular, there is a large group of strains that lack the cag pathogenicity island, and also differ in frequencies of other genes. So it's not clear that these differences are other than in cag status.

      We ran the mentioned GWAS using the Cabo Verdean European lineage as controls group and the European gastric cancer lineage as cases to investigate the genomic differentiation between the Cabo Verdean European lineages and gastric cancer-associated lineages. Our objective was not to replicate previously reported associations or to identify novel H. pylori loci associated with gastric cancer. Therefore, we will revise the manuscript to make our objective clearer and remove statements that suggest direct associations with gastric cancer. In addition, because we employed a linear mixed model that accounts for population structure through the estimation of a genome-wide kinship matrix, we will include the genomic inflation factor. These measures will provide readers with a clearer assessment of the extent of population structure influencing the association analysis.

      Figure 1 seems very inconclusive.

      We agree that Figure 1 shows overlap between H. pylori-seropositive and seronegative distributions; in addition, its visual interpretation might be also influenced by a number of outlying observations. Nevertheless, the figure illustrates differences in pepsinogen I and II concentrations and pepsinogen I/II ratios between H. pylori-seropositive and seronegative individuals. To improve clarity,  we will include statistics in the plot, and we will revise the text to provide a more complete description of the key findings shown in Figure 1. If the reviewers and editor consider that these descriptive data are not central to the main conclusions of the study, we would be happy to move Figure 1 to the Supplementary Material, where it can provide supporting context without detracting from the presentation of the principal results.

      Reviewer #2 (Public review):

      Summary:

      This study investigates the population structure and ancestry of Helicobacter pylori in Cabo Verde, where the human population has mixed West African and European ancestry. The authors combine a population survey, serum markers, bacterial genome analysis, and paired human-bacterial ancestry data. They report high H. pylori seropositivity, several distinct bacterial groups, and limited correlation between human ancestry and bacterial ancestry. They also identify one European-derived bacterial group that appears to have undergone a recent expansion and carries fewer well-known virulence-related genes.

      The study is interesting, and the dataset is valuable, especially because population-based H. pylori genomic data from Cabo Verde and West Africa are limited. The results provide useful information on bacterial diversity, historical migration, and host-bacterial ancestry. However, some of the main conclusions are stronger than the evidence currently supports, particularly the claims of host adaptation, increased transmission, and reduced virulence.

      We thank the reviewer for their comments and careful consideration of our work.

      Strengths:

      (1) A major strength is the study population. Participants were recruited from the general population and not only from patients with gastrointestinal disease. This gives a broader view of H. pylori diversity in Cabo Verde than studies based only on hospital patients.

      (2) The number of participants tested for H. pylori antibodies is substantial, and the authors also obtained a relatively large number of bacterial genomes. The combination of human and bacterial genomic information is another important strength. This allows the authors to directly examine whether human ancestry is related to the ancestry of the colonising bacteria.

      (3) The population genetic analyses are extensive. The authors use several different approaches, and these generally support the existence of African-derived and European-derived bacterial groups in Cabo Verde. The identification of two low-diversity European-derived groups is also interesting and suggests a relatively recent expansion.

      (4) The addition of new strains from Ghana and Portugal improves the reference dataset. The results may help future studies of H. pylori population structure in Africa, Europe, Cabo Verde, and populations affected by historical Atlantic migration.

      (5) The finding that human ancestry and bacterial ancestry are only weakly related in this population is potentially important. It suggests that the long-term relationship between human and bacterial ancestry may be less stable in recently admixed populations.

      Weaknesses:

      The main weakness is that several biological conclusions are based on indirect evidence. The genomic results support recent expansion of one bacterial group, but they do not directly show that this expansion was caused by adaptation to local hosts or by increased transmission. Founder effects, population history, geographic clustering, household transmission, or random expansion may also explain the pattern. The wording should therefore be more cautious.

      We thank the reviewer for underlining the limitations of inferring the causes of lineage expansion from indirect genomic patterns alone. We will review the manuscript to more carefully reflect the uncertainty surrounding the drivers of this lineage expansion. We will also expand the Discussion to outline analytical approaches that may help distinguish between demographic and selective scenarios in a highly recombining species (where high levels of recombination may have obscured local signatures of selection). Where feasible, we will explore additional analyses of the genomic data to assess the extent to which the observed patterns are consistent with these alternative hypotheses.

      The conclusion of reduced virulence is also not fully supported. The expanded lineage often lacks the cag pathogenicity island and carries less virulent forms of vacA, which suggests lower virulence potential. However, this does not prove that the strains cause less gastric damage or lower disease risk. There are no endoscopic or histological data, and serum pepsinogen values are only indirect markers.

      We will review our manuscript to adopt a more cautious description of these strains. However, we emphasize that, as show in  in Figure 4 – source data3, we could not find strains belonging to this expanded lineage carrying an active cagPAI,  or a virulent vacA allele. Our GWAs analyses show high divergence between these strains and gastric cancer strains both at a genome-wide level, and at previously identified virulence genes (as sabA and BabA, in addition to the ones already mentioned above). Finally, serum pepsinogen serum analysis, although indirect, have been shown to compare with  to the “gold standard” method, histopathological biopsy microscopy (ex: Telaranta-Keerie et al 2010, 10.3109/00365521.2010.487918; Kitamura et al 2015, 10.1111/jgh.12987; Miftahussurur et al, 2020, 10.1371/journal.pone.0230064; please see more about this in our following comment). Taken together these results provide minimal support for the hypothesis that the expansion of this lineage is driven by increased virulence. We will review the manuscript in order to reflect this.

      The description of the study population as having limited gastric inflammation is too strong. Serum pepsinogen measurements are useful for estimating gastric atrophy, but they do not directly measure the degree of histological gastritis. In addition, participants were recruited independently of symptoms, but this does not mean that they were all asymptomatic.

      The epidemiological estimate is based on antibody testing. This measures seropositivity and cannot clearly distinguish current from previous infection. Therefore, terms such as active infection or colonisation should be used carefully.

      As mentioned above, serum pepsinogen serum analysis has been widely shown to reliably detect both chronic and atrophic gastritis (e.g. Miftahussurur et al, 2020, 10.1371/journal.pone.0230064). We adopted conservative cut-offs for serum pepsinogen values after a careful review of the literature in our analysis. However, we acknowledge that these cut-offs may vary between populations. Therefore, we agree with the reviewer that a more precise assessment would have included a sensitivity and specificity analysis of  serum pepsinogen measurements against histopathological biopsy microscopy in a subset of the individuals. Taking this is into consideration, we will review the manuscript to include this discussion and to adopt a more cautious description of these results. Although we acknowledge that seropositive does not distinguish active from past infection in the Discussion, we will bring this discussion into the Results section as well.

      The proposed new West-Central African bacterial group is based on a small number of reference strains from Ghana and Nigeria. The result is interesting, but broader sampling from African countries is needed before this group can be considered firmly established.

      We thank the reviewer for pointing the need to assess higher African diversity, which is an issue that we also mention in the Discussion. Although the new West-Central African group comprised fewer isolates than the other comparison populations, the clustering pattern is unlikely to be explained solely by sample size because the analysis was based on a normalised chromosome-painting coancestry matrix, which reduces the influence of unequal donor numbers. In our review, we will provide additional analyses to confirm the observed clustering pattern: 1) we will provide fineSTRUCTURE MCMC tree assignments; 2) we will repeat the Chromopainter/ fineSTRUCTURE analyses under random downsampling of the larger groups; 3) as suggested in further reviewer recommendations, we will repeat the Chromopainter/ fineSTRUCTURE analyses after removing the new Ghanaian sequenced strains.

      The interpretation related to the trans-Atlantic slave trade is plausible, but the data mainly show patterns consistent with known historical migration. They do not directly demonstrate when or how the bacterial lineages moved.

      We thank the reviewer for highlighting this point. The chromosome-painting analysis identifies shared ancestry and gene flow between populations but does not directly estimate the timing of those events. However, several lines of evidence suggest that the majority of the strains have been present in Cabo Verde for an extended period rather than representing recent introductions. First, the isolates were obtained from individuals who, as well as both of their parents, were born in Cabo Verde. Second, Cabo Verdean African strains show an excess of European ancestry relative to their putative parental African populations. And vice-versa: Cabo Verdean European strains exhibit an excess of African ancestry compared with their putative parental European populations. To further investigate this question, as also suggested in reviewer recommendations, we will explore the feasibility of identifying clonal or near-clonal relationships between Cabo Verdean and assess whether dating approaches such as BactDating can provide additional insights into the timescale over which these lineages have diversified and admixed within Cabo Verde.

      The gastric cancer comparison may also be affected by bacterial population structure. Differences between the Cabo Verdean lineage and gastric cancer strains may reflect ancestry or lineage differences rather than disease association alone.

      We thank the reviewer for highlighting this point. As outlined in our response to Reviewer 1, the primary objective of this analysis was to investigate genomic differentiation between the Cabo Verdean and gastric cancer-associated lineages. Therefore, the strong contribution of ancestry and lineage effects is central to the interpretation of this comparison. We will make this clearer in our final manuscript.

      Overall, the study achieves its main aim of describing H. pylori diversity and ancestry in Cabo Verde. The evidence is strong for the population structure and ancestry findings, but less strong for the proposed mechanisms of adaptation, transmission, and reduced disease-causing potential. The work will be useful to the field, but the main conclusions should be stated more carefully.

      We thank  the reviewers and the editor for their time and effort in assessing the manuscript. We will submit a revised version that carefully addresses these points and incorporates the suggested changes.

    1. Author response:

      We are grateful to the editor and reviewers for providing their time and expertise in the assessment of this article. We are glad that the overall evaluation is supportive.

      Reviewer #1 (Public review):

      This interesting manuscript challenges the current interpretation of the well-established stop signal reaction time task (SSRT), commonly used in many research areas. SSRT has been traditionally thought of as primarily a measure of inhibitory control. In this work, the authors argue that this is influenced significantly by sensory and motor transmission times, and that these low-level processes may systematically confound SSRT estimates both in individual groups and also in many clinical populations.

      Conceptually, this raises an important and significant question regarding the construct validity of one of the more widely used behavioral measures of response inhibition. The authors provide a clear theoretical framing and importantly address the overlooked assumptions in SSRT modelling, the underexplored source of variability.

      Further, direct evidence separating peripheral sensory and motor contributions from central inhibitory processes is certainly needed, as well as a more balanced interpretation of prior methodological refinements in the field. Is a correlation between T0 and SSRT sufficient to conclude that SSSRT may be predominantly driven by peripheral delays? What proportion of SSRT variance remains unexplained after one accounts for T0? If T0 and inhibitory processes share neural processing speed, does that mean they covary? If one corrects T0, do the group differences become smaller or perhaps disappear?

      We thank the reviewer for their positive assessment of our work. We agree all these questions are important. Our revision will provide repeatability estimates for all our indices, and use the repeatability of SSRT and T0 to estimate the proportion of SSRT variance that remains unexplained. We will mention that it is theoretically possible that shared neural processing speed between T0 and inhibitory processes contributes to the reported effect but, since the neural pathways are anatomically largely distinct, the available cognitive neuroscience literature suggests such contribution is unlikely to be strong enough to explain the observed relationship. The impact of correcting group differences in SSRT for T0 will entirely depend on where differences in SSRT between these groups come from. If it comes exclusively from peripheral delays, then the effect may indeed disappear. If it is central, then stronger differences may be revealed after correction. We appreciate these are key questions we would like answers to already, but identifying and reanalysing suitable datasets to answer it will be the purpose of future papers.

      Furthermore, how robust is the T0 estimate in noisy environments of all sorts and especially across modalities (visual and motor)? The authors argue that this relationship is clearer in "good quality data"; how do you define that, and what happens with noisy data? For instance, the authors refer to a range of clinical disorders such as ADHD and PD where noise is abundantly present, partly due to the disease itself or due to treatment.

      By good quality data, we mean enough trials at the optimal RT and SOA combinations, and an adequate preprocessing pipeline that excludes non-standard trials (poor fixation, pre-emptive responses, large undershoot …). Participants with increased intra-individual variance or low compliance will need more trials overall to get enough trials around divergence times for these to be accurately estimated. Supplementary figure 5 illustrates how low trial numbers lead to an overestimation of T0, which can be mitigated by pooling across participants. We expect noisy data to have a similar effect, although its impact may differ across clinical conditions based on where the additional variability comes from. We shall have more clarity on the reliability of T0 in clinical populations once relevant datasets have been reanalysed, and use this to define constraints for future data collection. Again, this will need to wait for future papers.

      In conclusion, this is certainly a thought-provoking and potentially influential contribution in the literature that raises important questions about the interpretation of the stop signal reaction time task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions

      Reviewer #2 (Public review):

      Continuing their work distinguishing sensory latencies of "cognitive" processes, the authors turn their attention to "response inhibition". The "square quotes" are being used to highlight how this manuscript aims to challenge previous descriptions of performance data and inferred computational processes. The authors assert that previous descriptions of the measure known as "stop signal reaction time" (SSRT) are flawed because they did not account for sensory latencies empirically or theoretically.

      Enthusiasm for the manuscript cannot be high in light of many weaknesses countering the possible strengths. Strengths include offering an opportunity to more carefully characterize the quantity SSRT and a specific empirical approach offered to the research community. However, these strengths are countered by the following structural, theoretical, and empirical weaknesses:

      As announced by the elephant in the title, the writing could be described as excessively polemical. However, the characterization and interpretation of previous empirical and theoretical work is disputable.

      The major theoretical claim regarding sensory delays inherent in SSRT is not novel. The authors assert, "...this corpus of work may have been misinterpreted because the SSRT is systematically influenced by low level sensory and motor transmission times, arguably more so than by inhibition or cognitive processes." This was certainly recognized by Logan and Cowan in their original work. They wrote, "An act of control, like any other act, must take time. The theory provides methods for measuring the latency of control even when the act of control is not directly observable." (page 298) Also, "... the estimate of stop-signal reaction time includes the latency of the internal response to the stop signal and the duration of the ballistic process." (page 316-317). Moreover, subsequent computational and empirical work, some noted by the authors, has distinguished the sensory encoding interval from the interval during which the STOP process interrupts the GO process.

      We agree that our manuscript should acknowledge that Logan and Cowan (1984) explicitly stated that internal and ballistic delays contribute to SSRT and thank the reviewer for highlighting the need to clarify the relationship between our work and the original Logan and Cowan framework. We will clarify that the novelty of our claim is not that peripheral delays contribute to SSRT, which is a logical necessity, but that differences in peripheral delays (across conditions or people) contribute to differences in SSRT, sometimes to a large extent. Such differences in SSRT are very widely assumed to reflect inhibitory control in the large corpus of work that followed this initial literature. This corpus has essentially ignored the message about stimulus processing and ballistic delays and their implications for individual differences or changes across conditions. Therefore, we maintain that SSRT differences may have been widely misinterpreted. We reference the articles where these implications were clearly spelled out: these empirical and modelling studies were based on a few monkeys or human participants, and therefore unable to provide the large-scale demonstration we provide here.

      The theoretical suggestion that an accounting for sensory delays undermines the functional interpretation of SSRT mischaracterizes the original literature. For example, in the Abstract the authors write "Sensory and motor contributions must be ruled out before linking SSRT results to inhibition or cognition". The original Logan and Cowan theory was about what happens at the end of SSRT, and that was described only as an "act of control", in perfectly positivist fashion. For example, Logan and Cowan wrote, "Estimates of stop-signal reaction time provide a measure of the latency of control." (page 315). Thus, the authors are misstating what was meant originally by SSRT. In addition, the authors offer no specific or formal definition to specify what they mean by "inhibition or cognitive processes".

      We thank the reviewer for pointing out that their original approach was mechanistically agnostic, which we will explicitly clarify in revision. We will remove the words “top-down” from our 4th sentence and reword the quoted sentence into "Changes in peripheral delays must be ruled out before linking changes in SSRT to inhibition or cognition". It remains the case that many hundreds of studies have since interpreted “ability to inhibit” and “latency of control” as specific to inhibition and control, and therefore have assumed that changes in SSRT directly reflect an inhibitory cognitive process. We agree that we do not currently offer a definition of what we mean by cognitive processes, except that they do not include incompressible sensory and motor delays. Based on the reviewer’s clarification, this common shortcut in the literature appears inconsistent with the spirit of the initial work, which we seek to rectify.

      Confidence in the new empirical conclusions of the manuscript must be low because the new performance data are of questionable quality. The first issue is that the stopping accuracy (or inhibition functions in original terminology) shown in Figure S3 is very problematic for the interpretation of the authors' empirical work in this manuscript. There are two problems. First, these plots should span from nearly 0% to nearly 100%. It is not possible to resolve the span of each individual in the figure, but it is clear that many, if not most, in both the Manual and Saccadic data span just 20-30%. Second, the plots should span the 50% success value. It is clear that the maximum or minimum values for many participants do not reach the 50% value. These two problems indicate that many (most?) participants were not really sensitive to the stop signal.

      We acknowledge our inhibition functions are narrow, but we do not believe this undermines the main conclusions, for several reasons. On the question of sensitivity to the stop signal, average spans for inhibition functions after participant exclusions were 37% for manual and 29% for saccades. This limited span is mainly attributable to our fixed SOA design, which was a necessary feature of a direct comparison between manual and saccadic behaviours. Figure S3 covers only 80 ms spread of SOA. The slopes are commensurate with most other studies, which cover much wider differences in SOA. If one selects the central 100 ms (where the slopes are steepest) from the figures in most previous papers, one will find stopping accuracy changes of around 30%. Therefore, sensitivity is similar.

      On the question of some functions not crossing 50%, the correlations between SSRT and T0 remain the same if we only keep those participants who crossed the 50% point (R(26)=0.64 for manual, R(11)=0.4 for saccades, same statistical significance levels). We will additionally rerun our SSRT and T0 correlation with stopping accuracy as a covariate. As stopping accuracy affects SSRT but not T0, it is unlikely to drive our results.

      The second issue concerns the pattern of response times (RTs) on "ignore" trials. The authors portray performance as exemplifying a "pause-then-go" strategy. This is not uncommon, but it is not the only way participants perform. Many participants across multiple studies of selective stimulus stopping produce RTs on "Ignore" trials essentially indistinguishable from RTs on no-stop trials. The authors must acknowledge and account for such individual variability. In fact, the "T_s" value is measured by the difference in distributions of RT on no-signal and ignore trials. If these distributions are not different, then the measurement and interpretation of this quantity is questionable.

      We agree that the issue raised by the reviewer would be important if no difference were present between the distributions. In our dataset, however, all participants showed a measurable distributional difference. We believe the difference in perspective that ‘many participants…produce RTs on ignore trials essentially indistinguishable from RTs on no-stop trials’ may be attributed to differences in the way we analyse data (RT distributions versus mean RT).

      Nearly all participants in our final sample had mean RTignore – RTgo > 10 ms (significant at the individual level), except for 2 in the manual condition (and none for saccades). Following Bisset & Logan (2014), a lack of clear mean RTignore - RTgo difference in these 2 participants might have been interpreted as reflecting a different strategy. However, all our participants showed clear dips between go and ignore RT distributions when locked on signal onset, and these two manual participants were no exception. Therefore, accounting for response probability and RT at each SOA in our distributional analysis revealed clear ignore versus go differences, masked when relying on mean RT. The lack of mean RT difference for these two participants in the manual modality had no impact on our main hypothesis testing because T0 was extracted by comparing signal-absent and signal-present trials (pooling ignore and stop), while TS was extracted by comparing ignore and stop (not ignore and go).

      In terms of interpretation and whether participants employ a pause-then-go strategy, we understand performance in the selective stopping task as reflecting a combination of automatic activation and interference, and endogenous activation and inhibition. Our interpretation is that the pause component primarily reflects automatic interference triggered by stimulus onset, although strategic factors may also contribute in some circumstances. Individual differences in mean RTignore - RTgo could reflect both automatic interference and endogenous pausing (both of which would increase the difference between ignore and go trials), as well as subsequent failing to go on ignore trials (omissions, which would decrease differences by removing longer latency responses from ignore distributions just as in stop signal distributions). Although some of this can be described as strategic, some won’t be, and we therefore refrain from inferring strategy based on mean RTignore – RTgo.

      Related, the distributions of RT on stop trials, particularly for saccade responses, are portrayed with a second mode in the schematic illustrations and clearly peaking at SSRT in Figure S1. This second mode is not observed in other saccade stop signal studies. This indicates that the participants in this study were in a peculiar mode of performance.

      The second mode indicates that participants occasionally ignore the stop signal, which is why it peaks at the same latency as the rebound for the ignore distribution. These are not unusual, in particular in selective stopping designs, but are not as easily seen on cumulative functions, which are the standard way of plotting the results in this field.

      Finally, given the pivotal role of measures of differences of RT distributions and the pronounced variation of stopping accuracy (Figure S3), the authors must show the distributions for all of their new participants. The authors' claim to higher resolution obliges them to reveal every step of analysis.

      We will save figures showing the individual distributions in the OSF folder. Note that these figures can be produced by running the code we shared, so each step is already fully transparent, but we will create tidy versions that also highlight manually corrected indices.

      In its current form, this manuscript is unlikely to change the thinking of modelers or practitioners of the stop signal task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions, so that the paper can be revised to address the concerns.

      Reviewer #3 (Public review):

      Summary:

      Statham and colleagues test an assumption underpinning a very large literature: that the stop-signal reaction time (SSRT) indexes the speed or efficacy of top-down inhibitory control. They argue instead, and support their claims with a total of eight datasets, that SSRT is substantially occupied by visuomotor deadtime (i.e., incompressible sensory and motor delays common to all visually guided responses), which varies across individuals, conditions and populations in ways that mimic effects usually attributed to inhibitory control. They propose two remedies: subtracting an independent estimate of visuomotor deadtime (T₀) from SSRT, and a new index, the selective stopping delay (ΔT), from the stimulus-selective stopping task.

      Strengths:

      The paper's principal strength is the combination of these components. That SSRT must contain peripheral delays is not itself new, as the authors point out (Boucher et al., 2007; Salinas and Stanford, 2013; Bompas et al., 2020). What is new is the quantification of the problem at scale, across seven archival datasets and a preregistered replication, together with the demonstration that T₀ can be recovered from existing stop-task data. That is important, as it provides a diagnostic that can be applied to data already collected. The authors' offer to assist others in doing so is exemplary. The supplementary analyses of trial numbers and participant pooling are very useful, and the paper provides important sanity checks, notably confirming that stop and ignore signals produce indistinguishable initial interference before pooling them.

      Weaknesses:

      The evidence for the central claim is strong but presented in a way that overstates it. Figure 2 reports 85% and 80% shared variance between SSRT and T₀, but these pool across datasets and, more critically, across response modality: manual and saccadic estimates from the same participants are plotted together with a single regression line through both. Because manual and saccadic deadtimes differ by roughly 130 ms, the resulting correlation largely reflects a between-condition difference rather than covariation among individuals. The numbers that speak to individual differences are more modest (40% for manual responses; 7% for saccades). The manual result is convincing and consequential; the saccadic result is not, and the explanation in terms of restricted range, while plausible, is offered after the fact and is directly testable by reporting the reliability of saccadic T₀ or correcting the correlation for attenuation. This limitation is arguably good news for the paper's practical message, since it implies saccadic measures are relatively protected, but the manuscript should make clear (including in the abstract) that the strong individual-differences case rests on the manual data, where motor execution delay is the main driver.

      We will add separate regression lines and R-values for manual and saccadic modalities on Fig.2A and an inset showing the variance only driven by individual differences across all archival data (i.e. z-scored per condition and datasets, R(215)=0.43, p<0.001). We will also state more explicitly that the overall correlation may not be the relevant one for researchers specifically interested in individual differences.

      While a large portion of SSRT literature is about individual differences, there are also many studies about differences between conditions, including comparing different response modalities. Therefore, it is a general question whether differences of any kind in SSRT reflect differences in inhibitory control or differences in sensory-motor delays. At a conceptual level, most users of the SSRT are intending to measure control, and have a conceptual model in which control is separate from the modality of response or the exact characteristics of stimulus delivery. Thus, it is important to point out that their measure of control is very much dependent on these things, and in fact to a much larger degree than the more subtle individual differences, group differences or conditions of interest.

      We agree that repeatability is critical for interpreting null results and will provide split-half repeatability for all our indices, including T0 and SSRT, and use these to correct their correlations. We thank the reviewer for this suggestion.

      A related point concerns interpretation rather than analysis. Since SSRT is, on the authors' own account, approximately the sum of T<sub>0</sub> and a decision-related component, covariation between the two is expected on structural grounds; the preregistered correlation with reaction times from separate speeded blocks mitigates this, but the finding is less surprising than its current framing implies. What would determine whether past conclusions must be revised is not whether SSRT correlates with T<sub>0</sub> across individuals, but whether the decision-related component tracks the independent variable in any given study. The alcohol reanalysis could be a test case for this: the authors show that alcohol raises T<sub>0</sub> commensurately with SSRT and conclude the effects are "consistent with these effects being fully driven by visuomotor delays," yet (unless I missed something) they do not report the corrected measure for these data, while they do so for signal contrast and response modality. Running that analysis, and stating plainly what Campbell et al. (2017) would have concluded under the proposed treatment, would be an important demonstration.

      We fully agree that the presence of a correlation is indeed entirely expected and obvious in our own account, as conveyed early on in the manuscript (“From Fig. 1D, it seems clear that SSRT and T<sub>0</sub> are inevitably connected”). We agree the main question is what is left for SSRT to explain. We will update our analysis of the Campbell et al. (2017) alcohol study as suggested. Future work can then focus on other “independent variables”.

      The case for ΔT is the least developed part of the paper. ΔT is a difference between two independently estimated, individually noisy quantities, extracted by a non-trivial procedure (see also below), and no reliability estimates are reported for T<sub>0</sub>, TS or ΔT. This would be possible based on the two-session design (and the group has prior work on the reliability of cognitive control measures). This matters because the argument that ΔT is superior rests, to some extent, on null findings: ΔT does not correlate with SSRT, with stopping accuracy, or with the differential response to stop and ignore trials. These null correlations are interpreted as freedom from confounds, but an unreliable measure would produce the same pattern, and the seven participants with implausible negative ΔT values indicate that noise is not negligible.

      We fully agree with all this and will update the wording surrounding the lack of correlation between ΔT and the other measures in light of its repeatability

      In addition, the subjective correction of dip onsets ("Departure points were visually inspected and adjusted if it was deemed that the algorithm had placed them in inappropriate places"), which is critical to the paper's central measurement, should be blinded to condition or show inter-rater agreement. Since T<sub>0</sub> and TS are compared across conditions and ΔT is their difference, this introduces researcher degrees of freedom.

      All divergence times (T0, T0,stop, T0,ignore and TS) were confirmed by two of the authors. Each index for each modality is plotted on a separate figure (showing all individuals). It is technically easy to compare, say, T0 and TS, for one individual, but we refrained from doing this (and indeed ended up with many T0 > TS). We agree that blinding and inter-rater reliability are important safeguards and that our current analysis fell short of this. We will explore whether this can be done retrospectively, and report on this exercise alongside guidance on criteria used for manual corrections.

      This is particularly critical when a dip is not easy to extract. Figure 3 depicts an idealized ignore-trial distribution with a clean, deep dip. Real distributions are unlikely to look like this, and dip depth should depend on the behavioral relevance and salience of the ignored event; published work on rapid manual inhibition indicates that dips to behaviorally irrelevant events can be very shallow. Since ΔT is extractable only where the dip is resolvable, the generality of the method can be questioned. Ideally, the empirical distributions underlying every dataset analyzed should be shown to alleviate this concern.

      We will make figures available in the OSF folder with individual distributions that supported the extraction of each index, flagging those that got manually corrected. This will make apparent that the vast majority of dips were very clear, for both T0 and TS. Unclear dips led to missing indices, and were therefore excluded from our hypothesis testing.

      One uncontrolled procedural difference also deserves comment. Corrective feedback about stopping too often or stopping too rarely was given after manual blocks only; saccadic blocks received none, and fixed rather than staircased delays were used throughout. Since the manual-saccadic contrast carries much of the argument, and saccadic blocks yielded both lower stopping accuracy (36% vs 52%) and far more exclusions (8 vs 1 of 37), this asymmetry offers an alternative to the interpretation that saccades are simply harder to inhibit.

      Indeed, blockwise feedback would have been hard to implement reliably for saccades. As manual and saccadic blocks were interleaved, our hope was that participants could use the feedback received for manual to adjust their strategy for both modalities. We will check how often the feedback was triggered for manual blocks and, if more than negligible, we will note the reviewer’s suggestion as a possible driver for modality differences.

      A final point concerns the comparison between response modalities. Raw SSRT suggests that saccadic inhibition is faster than manual (174 vs 266 ms), while both proposed corrections reverse this, with SSRT−T<sub>0</sub> and ΔT each indicating that saccadic inhibition is slower (the latter consistently across nearly every participant). This is one of the clearest illustrations of the paper's thesis, but it is not taken up in the discussion, which returns to modality only to note that saccadic T<sub>0</sub> varies little (the one reference to variation across action modalities appears in the modeling section, without stating its direction). It would also benefit from a caveat. Both corrected measures subtract the same T<sub>0</sub>, and manual and saccadic T<sub>0</sub> differ by roughly 130 ms, so the two do not corroborate one another independently (TS is itself longer for manual responses, and yields a shorter ΔT only once the larger manual T<sub>0</sub> is removed). The accuracy of the subtraction therefore matters here: if the manual regression slope of 0.75 reflects sub-additivity rather than attenuation, subtracting the full T<sub>0</sub> would overcorrect manual responses more than saccadic ones, and could produce the reversal on its own.

      We agree with all this. We will use the repeatability of SSRT and T0 to disattenuate the slopes and consider alternatives to subtraction to correct for visuo-motor deadtime. Before we can elaborate on the modality effect on the speed of inhibition, we need to simulate the effect of motor variability on T0, TS and SSRT. If motor noise affects TS or SSRT more or less than it affects T0, this will affect our conclusions. We will explore this issue in the revision and report any analyses that bear on the robustness of the modality effect.

      These concerns qualify rather than undermine the contribution. The core observation is robust, the diagnostic is practical and immediately applicable, and the case that a large body of work requires re-examination is well made. If the corrected analyses are carried through on the datasets already in hand, this will be an important paper for anyone who uses the stop-signal task.

      We thank the reviewer for their in-depth and thoughtful comments and suggestions

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this paper, Chen et al. identified a role for the circadian photoreceptor CRYPTOCHROME (CRY) in promoting wakefulness under short photoperiods. This research is potentially important as hypersomnolence is often seen in patients suffering from SAD during winter times. The mechanisms underlying these sleep effects are poorly known.

      Strengths:

      The authors clearly demonstrated that mutations in cry lead to elevated sleep under 4:20 Light-Dark (LD) cycles. Furthermore, using RNAi, they identified GABAergic neurons as a primary site of CRY action to promote wakefulness under short photoperiods. They then provide genetic and pharmacological evidence demonstrating that CRY acts on GABAergic transmission to modulate sleep under such conditions.

      Weaknesses:

      The authors then went on to identify the neuronal location of this CRY action on sleep. This is where this reviewer is much more circumspect about the data provided. The authors hypothesize that the l-LNvs which are known to be arousal promoting may be involved in the phenotypes they are observing. To investigate this, they undertook several imaging and genetic experiments.

      While the authors have made improvements in this resubmitted manuscript, there are still multiple concerns about the paper. I think the authors provide enough evidence suggesting that CRY plays a role in sleep under short photoperiod. The data also supports that CRY acts in GABAergic neurons. However, there are still major issues with the quality of the confocal images presented throughout the paper. In many cases it appears that the images are oversaturated with poor resolution, making it hard to understand what is going on. In addition, none of the drivers used in this study are specific to the neurons the authors aim to manipulate. Therefore, the identity of the GABAergic neurons involved in this CRY dependent sleep mechanism remains unclear. Similarly, whether l-LNvs are the target of this GABA mediated sleep regulation under short photoperiod is not fully demonstrated. The data presented suggests that but does not prove it.

      Major concerns:

      (1) While the authors provided sleep parameters like consolidation or waking activity for some experiments. These measurements are still not shown for several experiments (for example Figures 2E, 3, 4, 5, and 6). These data are essential, these metrics must be reported for all sleep experiments.

      These metrics have now been added to Fig.2 S4 and 5, Fig.3 S1-3, Fig.4 S2 and 3, Fig.5 S2 and 4, as well as Fig.6 S1.

      (2) Line 144 "We fed flies with agonists of GABA-A (THIP) and GABA-B receptor (SKF-97541) (Ki and Lim, 2019; Matsuda et al., 1996; Mezler et al., 2001). Both drugs enhance sleep in WT," The proper citation is needed here, Dissel et al., 2015 PMID:25913403. Both THIP and SKF-97541 were used in that paper.

      Thank you for pointing this out. We have modified our manuscript accordingly.

      (3) Figure 2C and 2F: it appears that the control data is the same in both panels. That is not acceptable.

      Thank you for pointing this out. We are now using data from control flies that were monitored in the same experiments as the experimental groups.

      (4) Figure 4A: With the quality of the images, it is impossible to assess whether GABA levels are increased at the l-LNvs soma.

      We apologize for the poor quality. Unfortunately, the GABA immunostaining does not work very well in our hands and thus the background is high. We have now commented on this issue in the fourth paragraph of discussion and have toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript.

      (5) Fig 4 S1A shows colabeling of l-LNvs and Gad1-Gal4 expressing neurons. They are almost 100% overlapping signals. This would indicate that the l-LNvs are GABAergic themselves, or that there is a problem with this experiment.

      Fig 4 S1A demonstrates the expression pattern of SYT-GFP driven by Gad1GAL4, which should label the synaptic terminals of GABAergic neurons. Therefore, the labeling observed at l-LNvs suggest that GABAergic neurons project to l-LNvs. This is further validated by the GRASP and trans-Tango experiments.

      (6) Fig 4 S1B: Again, I can see colabelling of the GFP and PDF staining, suggesting that Gad1-Gal4 expresses in l-LNvs.

      Fig 4 S1B demonstrates anatomical sites where GABAergic neurons project to and form synaptic connections with PDF neurons. Therefore, GFP signals at the l-LNvs suggest that these cells receive synaptic inputs from GABAergic neurons, echoing the results shown in Fig 4 S1A.

      (7) Line 184: "Consistently, knocking down Rdl in the l-LNvs rescues the long sleep phenotype of cry mutants (Figure 4-figure supplement 1D)." This statement is incorrect as the driver used for this experiment, 78G01-GAL4 is not specific to the l-LNvs, so it is possible that the phenotypes observed are not coming from these neurons.

      Thank you for pointing this out. We have modified our manuscript to note this.

      (8) Figure 4G-K: None of these manipulations are specific to the l-LNvs. The authors describe 10H10-GAL4 and 78G01-GAL4 as l-LNvs specific tools, but this is not the case. Why not use the SS00681 Split-GAL4 line described in Liang et al., 2017 PMID: 28552314? It is possible that some of the effects reported in this manuscript are not caused by manipulating the l-LNvs.

      Thank you for pointing this out. We have now modified our manuscript to avoid misleading remarks. We have used SS00681 Split-GAL4 to express TrpA1 but did not observe any substantial effect on sleep duration under short photoperiod. Therefore, we did not use this line for further experiments.

      (9) Similarly for the manipulation of s-LNvs, the authors cannot rule out effect that are coming from other cells as R6-GAL4 is not specific to s-LNvs.

      We have now modified our manuscript to avoid misleading remarks.

      (10) The staining presented in Fig 5 S1 is not very convincing. Difficult to see whether Gad1-GAL4 only expresses in the s-LNvs.

      We have now quantified the GFP signal in the l-LNvs and s-LNVs in Fig.5 S1B and D. As can be seen, the s-LNvs show prominent signal above the background while the l-LNvs do not.

      Reviewer #3 (Public review):

      Summary:

      In humans, short photoperiods are associated with hypersomnolence. The mechanisms underlying these effects is however, unknown. Chen et al. use the fly Drosophila to determine the mechanisms regulating sleep under short photoperiods. They find that mutations in the circadian photoreceptor cryptochrome (cry) increase sleep specifically under short photoperiods (e.g. 4h light: 20 h dark). They go on to show that cry is required in GABAergic neurons and that the effects of the cry mutation on sleep are mediated by alterations in GABA signalling. Further, they suggest that the relevant subset of GABAergic neurons are the well-studied small ventral lateral neurons that they suggest inhibit the arousal promoting large ventral neurons via GABA signaling

      Strengths:

      Genetic analysis to show that cryptochrome (but not other core clock genes) mediates the increase in sleep in short photoperiods, and circuit analysis to localise cry function to GABAergic neurons.

      Weaknesses:

      The authors' have substantially revised their manuscript, and the manuscript is better for the revisions. However, the conclusion that the sLNvs are GABAergic is unfortunately still not well supported by the data. A key sticking point remains the anti GABA immunostaining, and specific driver lines for sLNvs and lLNvs.

      The authors should tone down their conclusions to reflect the fact that their data, as presented, does not support the model that cry acts in sLNvs to modulate GABA signalling onto lLNvs and thus modulate sleep.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript in the Introduction, Results and Discussion.

      Reviewer #4 (Public review):

      Summary:

      Short photoperiod is an important experimental manipulation in neurobiology, endocrinology, and metabolism studies. However, the molecular mechanisms by which short photoperiod gives rise to behavioral phenotypes that are seen in seasonal affective disorders remain unknown. Using the classic circadian model organism Drosophila, this study examines short photoperiod-induced hypersomnolence and identifies the circadian photoreceptor cryptochrome as a regulator of GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions. The discovery has broad implications for understanding how short photoperiod modulates neural inhibition in circadian circuits in regulating sleep.

      Strengths:

      The Drosophila model provided a powerful platform to dissect the molecular mechanisms underlying short photoperiod-induced hypersomnolence. A battery of behavioral, imaging, circuit-manipulation approaches was employed to test the novel hypothesis that the circadian photoreceptor cryptochrome modulates GABAergic tone within the clock neural circuit to promote wakefulness under short photoperiod conditions.

      Weaknesses:

      The current model proposed by the authors suggests that the small ventral lateral neurons of the Drosophila clock circuit are GABAergic; however, this remains unclear. At present, the field lacks sufficient data and validated reagents to definitively establish the GABAergic identity of these neuropeptidergic neurons.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript.

      Recommendations for the authors:

      The manuscript has improved after revisions. However, the evidence in support of the claim that the sLNVs secrete GABA onto the lLNvs remains unconvincing. The evidence that loss of cry in GABAergic neurons modulates sleep is solid. However, the authors' claim that the sLNVs are the relevant GABAergic neurons is not sufficiently backed up by the evidence presented. We suggest that the authors tone down their conclusions to reflect this.

      Thank you for the comments. We have now toned down our conclusions regarding the GABAergic s-LNv—l-LNv circuitry in this revised version of the manuscript in the Introduction, Results and Discussion.

      Reviewer #3 (Recommendations for the authors):

      Minor points:

      (1) The authors suggest that the effects of cry on sleep and mediated by the Rdl receptor, and use the GABA agonist THIP as support of this argument. However THIP acts on the Lcch3 and Grd receptors, not Rdl

      Thank you for pointing this out. We have modified relevant discussion accordingly.

      (2) In several instances (e.g. line 66, line 148), the authors use 'consistently' in the sense of 'consistent with previous data'. It would be better if they rephrase this.

      This has been fixed.

      Reviewer #4 (Recommendations for the authors):

      It is my pleasure to serve as a reviewer for this revised manuscript. The authors have carefully revised the manuscript in response to the critiques raised by all previous reviewers and have used all the available reagents to conduct additional experiments to assess the GABAergic properties of the small ventral lateral neurons (sLNv). Although it remains unclear in the field whether sLNvs co-transmit GABA, this study raises this possibility within an interesting biological relevant context. I recommend this manuscript for final publication.

      Thank you for your comments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      (1) The study emphasizes H3K4me2, which often serves as a precursor to H3K4me3, a well-studied modification during early development. Analyzing the new H3K4me2 dataset alongside published H3K4me3 data is crucial for comprehensively understanding epigenetic reprogramming post-fertilization and the interplay between histone modifications. However, the current analysis is preliminary and lacks depth.

      We fully agree with this valuable suggestion. Our research group has previously systematically profiled H3K4me3 dynamics in human and mouse early embryos, and the relevant results have been published in Science (2019). The core objective of the current study is to explore the erasure, re-establishment and biological functions of H3K4me2 during mammalian parental-to-zygote transition. To enrich our analysis, we have now integrated our H3K4me2 data with publicly available H3K4me3 datasets for joint analysis. The results clearly demonstrate that H3K4me2 is not merely a precursor of H3K4me3. These two histone marks present distinct genome-wide distribution patterns and perform independent regulatory roles in embryonic epigenetic reprogramming. We have added the joint analysis results and relevant discussions in the revised manuscript to elaborate the crosstalk between H3K4me2 and H3K4me3.

      Manuscript Revisions

      All supplementary analyses and discussions are located in the Results and Discussion sections, Results section (Page 6, Lines 156–165) (Page 7, Lines 197–201) (Page 10, Lines 274–278) and Discussion section (Page 13- 14, Lines 374–382) (Page 15, Lines 424–429.

      (2) Tranylcypromine (TCP) is known as an irreversible inhibitor of monoamine oxidase and LSD1. While the authors suggest TCP inhibits the expression of LSD2, this assertion is questionable. Given TCP's potential non-specific effects in cells, conclusions related to the experiments using TCP should be made with caution.

      We highly appreciate this important reminder about the off-target effects of TCP. We have supplemented two classic literatures (J Am Chem Soc, 2010; Mol Cell, 2010) which have proven that TCP acts as an irreversible inhibitor targeting both LSD1 (KDM1A) and LSD2 (KDM1B). According to our transcriptome and protein detection data, the endogenous expression level of LSD1 is extremely low in mouse early embryos. Therefore, LSD2 is the primary functional target of TCP in our embryonic experimental system. All conclusions derived from TCP treatment experiments are described prudently in the full text to avoid over-interpretation.

      Manuscript Revisions

      Relevant supplements are made in the Results section (Page 7, Lines 197–201).

      (3) Some batches of H3K4me2 antibody are known to cross-react with H3K4me3. Has the H3K4me2 antibody used in CUT&RUN been tested for such cross-reactivity? Heatmaps in the figures indeed show similar distribution for H3K4me2 and H3K4me3, further raising concerns about antibody specificity.

      Thank you for raising this critical question regarding antibody specificity, which is essential for the reliability of CUT&RUN experiments. The H3K4me2 antibody used in this study was purchased from Millipore (Cat. No. 07030). Based on the manufacturer’s product specification and our internal verification, this antibody has very low cross-reactivity with H3K4me3.The similar distribution shown in heatmaps is not caused by antibody contamination. Instead, it reflects the inherent spatial correlation between H3K4me2 and H3K4me3 on chromatin in early embryos.

      (4) Certain statements lack supporting references or figures (examples on page 9 can be found on line 245, line 254, and line 258).

      We apologize for the inadequate citation in the original manuscript. We have comprehensively checked the full text and added standard peer-reviewed references to all statements without literature support. Specifically, we have supplemented corresponding references for the content on Page 9, Line 259 and Line 266 as suggested. We also completed a full-text inspection to fix similar problems in other positions.

      Manuscript Revisions

      References are supplemented on Page 9, Line 259 and Line 266.

      (5) Extensive language editing is recommended to clarify ambiguous sentences. Additionally, caution should be taken to avoid overstatement - most analyses in this study only suggest correlation rather than causality.

      We fully accept this suggestion. We have thoroughly revised all ambiguous, redundant and grammatically problematic sentences throughout the manuscript to improve readability and academic rigour. Furthermore, we have carefully modified all overstated expressions. For all experimental results and bioinformatics analyses, we only use words such as correlate with, suggest, indicate to describe correlative relationships. All inappropriate causal inferences have been completely removed to ensure objective presentation of our data.

      Manuscript Revisions

      Full manuscript is polished and revised.

      Reviewer #2 (Public Review):

      (1) The authors claim that the Cut & Run worked for MII oocytes, zygotes, and the 2-cell embryos. However, it is unclear if H3K4me2 is erased during the stage or if the Cut & Run did not work for these samples. To support the hypothesis of the erasure of H3K4me2, the authors conducted immunofluorescence staining, and H3k4me2 was undetected in the MII oocyte, PN5, and 2-cell stage. However, the published papers showed strong staining of H3K4me2 at the zygote stage and 2-cell stage ((Ancelin et al., 2016; Shao et al., 2014)). The authors need to cite these papers and discuss the contradictory findings.

      The authors used 165 MII oocytes and 190 GV oocytes for the Cut & Run. The amount of DNA in MII oocytes is halved because of the emission of the first polar body. Would it be a reason that H3K4me2 has fewer H3K4me2 peaks in MII oocytes?

      Thank you for putting forward these thoughtful questions. Firstly, we have cited two published literatures (Ancelin et al., 2016; Shao et al., 2014) in the revised manuscript and discussed the inconsistent immunofluorescence results. The main reason for the discrepancy lies in different confocal microscope parameters including laser power, gain and exposure time adopted by different laboratories. In our study, we used unified imaging parameters to continuously observe samples from GV oocytes to blastocysts, so weak H3K4me2 signals at zygote and two-cell stages could not be detected. When we adjust parameters specifically for these stages, weak fluorescence signals can be observed. We have elaborated this point in the Discussion section.

      Secondly, we clarify that the reduction of H3K4me2 peaks in MII oocytes is not caused by decreased DNA content. Although MII oocytes extrude the first polar body during maturation, we collected the polar body together with oocytes in all CUT&RUN experiments, so the total DNA content of MII samples is not reduced. Combined with previous studies on human oocytes, we confirm that the loss of H3K4me2 peaks from GV to MII stage is a real physiological epigenetic change accompanying oocyte meiotic maturation and chromatin remodeling.

      (2) The authors claim that Kdm1a is rarely expressed during mouse embryonic development (Figure 4A). However, the published paper showed that KDM1a is present in the zygote and 2-cell stage using immunostaining and western blotting ((Ancelin et al., 2016)). Additionally, this paper showed that depletion of maternal KDM1A protein results in developmental arrest at the two-cell stage, and therefore, KDM1a is functionally important in early development. The authors should have cited the paper and described the role of KDM1a in early embryos.

      We apologize for the ambiguous expression in the original manuscript. What we described is a relative expression level: in mouse early embryos, the expression of KDM1A is lower than KDM1B, rather than the absolute absence of KDM1A.

      (3) The authors used the published RNA data set and interpreted that KDM1B (LSD2) was highly expressed at the MII stage (Figure S3A). However, the heat map shows that KDM1B expression is high in growing oocytes but not at 8w_oocytes and MII oocytes. The authors need to interpret the data accurately.

      We sincerely apologize for the data misinterpretation caused by improper data normalization in the original heatmap. We have completely re-normalized the RNA-seq data and redrawn Supplementary Figure S3A.

      The updated heatmap clearly shows that KDM1B is highly expressed in growing oocytes, while its expression decreases in 8-week oocytes and MII oocytes. Combined with Figure 4A, we have rewritten the description of KDM1B expression trends across different oocyte stages, and all textual descriptions are now consistent with the corrected data.

      Manuscript Revisions

      Supplementary Figure S3A is remade; data interpretation is revised in the Results section. Supplementary Figure S3A (remade); Results section (Page 42).

      (4) All embryos in the TCP group were arrested at the four-cell stage. Embryos generated from KDM1b KO females can survive until E10.5 (Ciccone et al., 2009); therefore, TCP-treated embryos show a more severe phenotype than oocyte-derived KDM1b deleted embryos. Depletion of maternal KDM1A protein results in developmental arrest at the two-cell stage ((Ancelin et al., 2016)). The authors need to examine whether TCP treatment affects KDM1a expression. Western blotting would be recommended to quantify the expression of KDM1A and KDM1B in the TCP-treated embryos.

      We dig the transcriptome data to confirm the specificity of TCP to KDM1b. In addition, the intervention of TCP on the whole fertilized egg in this study increased the H3K4me2 content, and the embryo development retarding effect was more significant than that obtained by crossing with normal paternal lines after knocking down KDM1B from the mother.

      (5) H3K4me2 is increased dramatically in the TCP-treated embryos in Figure 4 (the intensity is 1,000 times more than the control). However, the Cut & Run H3K4me2 shows that the H3K4me2 signal is increased in 251 genes and decreased in 194 genes in the TCP-treated embryos. The authors need to explain why the gain of H3K4me2 is less evident in the Cut & Run data set than in the immunofluorescence result.

      Thank you for this valuable question. The inconsistent data performance between immunofluorescence (IF) and CUT&RUN is determined by the essential differences between the two technical principles.

      Immunofluorescence is a global semi-quantitative method, which reflects the total content of H3K4me2 in the whole nucleus. The 1000-fold increase refers to the overall fluorescence intensity of the nucleus. In contrast, CUT&RUN combined with high-throughput sequencing is a locus-specific quantitative method, which detects H3K4me2 enrichment changes at individual gene loci. Different analytical models and threshold settings also lead to differences in final data presentation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper asks how the NK cell receptor KIR2DL4 binds HLA-G and undergoes endocytosis. The authors propose that an allosteric disulfide-bond switch controls whether the receptor is in a ligand-binding or non-binding state, and they support this model using mutagenesis, imaging, mass spectrometry, and structural prediction.

      Strengths:

      A major strength is the use of diverse, complementary approaches to validate the central claim. The authors combined unbiased random mutagenesis to identify key residues, confocal microscopy to track cellular localization, and mass spectrometry to quantify the redox states of specific disulfide bonds. These methods consistently support a single model: an allosteric disulfide switch. The transition between a Cys10-Cys28 bond and a Cys28-Cys74 bond serves as a functional switch that controls whether the receptor resides at the plasma membrane to bind ligand or remains inactive in endosomes.

      Weaknesses:

      (1) The core model is interesting, but some of the strongest mechanistic claims still rely heavily on structure prediction rather than direct structural evidence, especially the proposed HLA-G contact surface in Figure 6 (now in Figure 7).

      The crystal structure of KIR2DL4 has a D0 domain in the C10-C28 disulfide configuration [1]. The AlphaFold prediction is different, having a C28-C74 disulfide bond in the D0 domain. It is understood that any prediction could be wrong. Nevertheless, the AlphaFold structure did point to the possibility that KIR2DL4 exists in two different disulfide-bonded forms. We went on to demonstrate experimentally that these two forms coexist in human cells. This conclusion is independent of the structure predicted by AlphaFold.

      The second difference predicted by AlphaFold is an allosteric change in a loop distant from the disulfide bond, suggesting the possibility that it could control binding of HLA-G. Again, this prediction could be wrong. Nevertheless, considering that the KIR2DL4 used to obtain a crystal structure was in a C10-C28 bond configuration and did not bind HLA-G [1], we wondered if HLA-G would bind to KIR2DL4 in a C28-C74 configuration.

      New experiments included in the revision have shown that a purified Cys10Leu KIR2DL4 mutant binds HLA-G (new Figure 6). Solving the structure of a KIR2DL4–HLA-G complex would be ideal, but this has not been possible thus far. The difficulty in crystallizing KIR2DL4 may be due, in part, to its propensity to form oligomers [1], as shown in the new Figure S6.

      The addition of both direct binding of HLA-G to KIR2DL4 and functional data showing that KIR2DL4 induces an ISG response when in the C28-C74 but not in the C10-C28 configuration strengthens the conclusion that disulfide switching controls ligand binding and downstream signaling relevant to NK cell interactions with HLA-G in early pregnancy.

      (2) The paper supports an effect of the disulfide state on trafficking and uptake, but the case for direct KIR2DL4-HLA-G binding still feels somewhat indirect. The manuscript itself notes that direct binding had not been previously shown, and the current explanation partly depends on inference about which disulfide state is present.

      Direct binding and affinity measurements of HLA-G bound to the Cys10Leu KIR2DL4 mutant (in a C28-C74 disulfide form) are in the new Figure 6. This crucial result is also consistent with functional data. New experiments (new Figure 5D) have shown that the ability of HLA-G to stimulate a transcriptional interferon-stimulated gene (ISG) response occurred with the C28-C74 form, but not the C10-C28 form of KIR2DL4.

      Surface plasmon resonance data showed for the first time direct binding between the C28-C74 form of KIR2DL4 and soluble HLA-G, with a K<sub>D</sub> of 1.6 mM (new Figure 6). Binding to WT KIR2DL4, which is in both configurations, C10-C28 and C28-C74, was also detected but with a lower affinity (K<sub>D</sub> = 19.4 mM). The purified WT KIR2DL4 formed oligomers (new Figure S6). In addition, binding of HLA-G to KIR2DL4 depended on the sequence of the peptide presented by HLA-G, as only one out of three peptides tested was compatible with KIR2DL4 binding.

      This new data was obtained in the laboratory of Jamie Rossjohn at Monash University, Victoria, Australia. He, along with Jan Peterson and Priyanka Chaurasia are new co-author on our revised manuscript.

      (3) Most of the main experiments are done in transfected 293T cells, so it is still not fully clear how strongly this mechanism carries over to the more relevant NK-cell setting discussed in the paper.

      Primary resting NK cells are not amenable to transfection. Despite this technical hurdle, we have included two key findings with primary NK cells in the revision.

      (1) As in the 293T transfected cell system, we have shown that inhibition of PDI caused reduced uptake of HLA-G in primary resting NK cells (New Figure 4E, F, G). This is consistent with uptake of HLA-G by the C28-C74 form of KIR2DL4 and with a switch from C10-C28 to C28-C74 catalyzed by PDI.

      (2) We have shown that cell-surface C28-C74 KIR2DL4 on primary NK cells, as detected by mAb 2388, decreased upon inhibition of PDI, again consistent with the role of PDI in maintaining a pool of C28-C74-bonded KIR2DL4 at the cell surface (New Figure 5E, F). As shown in the original Figure 3, PDI could reduce the C10-C28 bond in purified WT KIR2DL4 in vitro.

      (4) The cellular evidence for the PDI story is not specific, since it depends a lot on inhibitor and blocking experiments that could affect the broader extracellular redox environment.

      Using inhibitors that target PDIA1 selectively, namely Rutin (PDI-specific up to 30 microM), and a PDI-specific monoclonal antibody, we found that HLA-G uptake by primary NK cells was inhibited (Figure 4C, D). We admit that pCMPS and thiol blockade by DTNB (Figure 4A, B) affect the extracellular redox environment. Data that were obtained without PDI inhibitors include the reduction of the C10-C28 bond by PDI in WT KIR2DL4 in vitro (Figure 3F), direct binding of HLA-G to KIR2DL4 in a C28-C74 disulfide conformation (Figure 6), and a functional transcriptional response to HLA-G by C28-C74 KIR2DL4 and not with the C10-C28 KIR2DL4.

      Reviewer #2 (Public review):

      Summary:

      Rajagopalan et al show how extracellular domain features regulate KIR2DL4 internalization. The trafficking phenotypes of cysteine mutants are logically organized, and well-summarized in a Table. The disulfide mapping and differential alkylation strategy are appropriate and provide strong support for alternative disulfide configurations in D0. The higher accessibility or more selective reduction of Cys10-Cys28 as compared to Cys28-Cys74 by PDI is a key mechanistic anchor.

      Strengths:

      The identification of a conformational switch in KIR2DL4 is conceptually novel. Experimental elegance, detailed and well-written.

      Weaknesses:

      Most of the mechanistic work was shown in HEK293. The authors should exhibit relevance using primary NK cells (using primary NK)

      As primary NK cells are not amenable to transfection, it is difficult to dissect the role of each disulfide form of the receptor KIR2DL4.

      Instead, we have now included PDI inhibition experiments using primary NK cells and shown that PDI inhibition reduces HLA-G uptake by primary NK cells (New Figure 4E, F, G). This is consistent with uptake of HLA-G by the C28-C74 form of KIR2DL4 and with a switch from C10-C28 to C28-C74 catalyzed by PDI.

      Furthermore, inhibition of PDI caused a decrease of C28-C74 KIR2DL4 at the cell surface of primary NK cells (New Figure 5E, F). This data is consistent with a requirement for a switch from C10-C28 to C28-C74 catalyzed by PDI, which maintains a pool of C28-C74 KIR2DL4 at the cell surface for HLA-G binding and internalization. As shown in the original Figure 3, PDI can reduce the C10-C28 bond in purified WT KIR2DL4 in vitro.

      Recommendations for the authors:

      Reviewing Editor Comments:

      To improve the strength of the evidence and the overall impact of the paper, please address the following major points:

      (1) Validation in Primary Cells:

      The central biological framing of the paper involves decidual NK cell responses to soluble HLA-G. We strongly recommend performing a critical experiment using primary NK cells to test whether PDI inhibition or thiol blockade alters KIR2DL4 surface retention and HLA-G uptake in a manner consistent with your observations in 293T cells.

      We have added new experiments with primary, resting NK cells, as described in our response to the major point 3 of reviewer #1, and to the weakness raised by reviewer #2.

      Briefly, we have included experiments in the revised manuscript on the effect of PDI inhibition on HLA-G uptake in primary NK cells (New Figure 4E, F, G) and on transient accumulation of KIR2DL4 at the cell surface (in a C28-C74 bonded form) of primary NK cells (New Figure 5E, F). The data showed that HLA-G endocytosis by primary NK cells and the presence of KIR2DL4 at the plasma membrane of primary NK cells were reduced after inhibition of PDI.

      (2) Clarification of the "Switching" Mechanism:

      The current data points toward the coexistence of the Cys10-Cys28 and Cys28-Cys74 states. Please clarify or provide evidence regarding whether a dynamic conversion occurs (e.g., prior to binding, upon ligand engagement, or during trafficking) versus a model of stable coexistence of two distinct receptor pools.

      Stable coexistence of two distinct KIR2DL4 receptor pools was a plausible hypothesis but one that is not supported by some of our data. In such a scenario, the C10-C28 form would not bind HLA-G and would reside in endosomes. It could have a role that is not related to HLA-G nor to the transcriptional response induced by HLA-G. However, our recent paper [2] showed that the transcriptional response of primary NK cells to soluble mAb #33 (bound to C10-C28) is very similar (R<sup>2</sup>=0.89) to that of resting NK cells incubated with soluble HLA-G (bound to C28-C74). These two ligands were tested at the same time, at the same molarity, and with the same primary NK cells [2].

      We don’t have answers yet to some obvious questions: is there switching after internalization of KIR2DL4 bound to mAb #33? What is the fate of C28-C74 that internalizes with HLA-G? We are not aware of technology that would answer these questions.

      A C28-C74 form, as a separate pool with residency at the cell surface, could be functional and respond to HLA-G by internalization and signaling from endosomes. However, there is no stable pool of C28-C74 KIR2DL4 at the cell surface and C28-C74 is depleted from the cell surface in the presence of PDI inhibitor (new Figure 5E, F), suggesting that C28-C74 KIR2DL4 is generated by the activity of PDI (new Figure S5). The sum of our experiments points to a tightly regulated control of KIR2DL4 biology, rather than the coexistence of two separate pools. A separate pool of C10-C28 KIR2DL4 would remain in an inactive state as far as the response to HLA-G is concerned. We favor the model whereby functional C28-C74 is generated from C10-C28 by the activity of PDI.

      Why could the response to HLA-G not be simpler? We address this point in the Discussion. One reason is that C28-C74 KIR2DL4 signaling at the plasma membrane of NK cells could be subject to inhibition by LILRB1 and NKG2A-CD94, co-expressed on NK cells, which bind to HLA-G and HLA-E, respectively, on fetal trophoblasts that encounter maternal NK cells in the decidua. These inhibitory receptors are known to be dominant against activation signals [3]. Trophoblast cells that invade the maternal decidua and encounter decidual NK cells selectively express HLA-C, HLA-E, and HLA-G. Strong inhibition signals by LILRB1 and NKG2A-CD94 could prevent activation through KIR2DL4. However, KIR2DL4 signaling, which occurs in endosomes [4] where signaling is sustained [5], can bypass these inhibitory signals at the plasma membrane.

      (3) Specificity of the PDI Model:

      Please elaborate on the relevance of extracellular PDI. Specifically, how does PDI perturbation affect the relative abundance of the two disulfide forms in a cellular context?

      We show in Figure 5E that two mAb for KIR2DL4 recognize different forms of the receptor. While mAb #33 recognizes only the C10-C28 form of the receptor, which is not at the cell surface, mAb 2238 recognizes both forms of the receptor. This allowed us to examine the effect of PDI on surface expression of the C28-C74 form of KIR2DL4 as detected by mAb 2238. We show that PDI inhibition reduces surface staining of C28-C74 (new Figure 5F), consistent with a model whereby a switch from C10-C28 to C28-C74 is catalyzed by PDI.

      A quantitative assessment of the relative abundance of the two forms of KIR2DL4 upon inhibition by PDI in a cellular context would have to be carried out by mass spec analysis of the two forms before and after treatment. That would be a very challenging experiment to perform with intact cells rather than purified proteins.

      (4) Agonist Antibody Mechanism:

      The manuscript mentions mAb #33 as a KIR2DL4 agonist. It would be highly informative for the reader if you could elaborate on whether this antibody activates the receptor by stabilizing a specific disulfide state or by driving internalization independently of HLA-G.

      We have shown that the agonist mAb #33 recognizes only the C10-C28 form (Figure 5E). We do not yet understand how it activates KIR2DL4. We do know that mAb #33 is not driving internalization considering that the receptor internalizes constitutively and is predominantly located in endosomes in the absence of HLA-G. Instead, it is the C10-C28 form of the receptor that carries mAb #33 into endosomes. Understanding how mAb #33 may function as a receptor agonist will require crystallization of the antibody bound to the receptor and is beyond the scope of this study. Structural studies of KIR2DL4 have been very difficult, due in part to its isoforms and tendency to form oligomers. It is not possible to answer your interesting question at this time.

      Minor Revisions:

      (1) Imaging Quantification:

      Ensure all figure legends include the number of independent experiments (n), specific statistical tests used, and precise alignment with the Methods section.

      This information is now included in the Methods section.

      (2) Textual Flow:

      To enhance engagement, please integrate the logic of Table 1 more explicitly into the main text of the Results section.

      This has been done.

      (3) Structural Discussion:

      Acknowledge the limitations of using structure prediction for the binding interface and discuss how these models align with existing literature on KIR-ligand interactions.

      We have described the use of AlphaFold solely as a tool to make predictions. Predictions can be wrong. Even so, they can generate new and useful hypotheses, as they did here. Existing, traditional KIR-ligand interactions are not informative in the context of the D0 domain in KIR2DL4 for the following reasons:

      The KIR2DL1/2/3 receptors with 2 Ig domains (hence 2D) have a D1 and a D2 domain. A comparison with KIR2DL4, which has a D0 and a D2 domain, may not be informative.

      The KIR3D receptors have the three domains, D0, D1 and D2. A structure of KIR3DL1 bound to HLA-B has been solved [6] by our collaborator for the revision, Dr. Jamie Rossjohn. As shown and mentioned in our manuscript (Fig. S7C and Legend), “predicted” contacts of the KIR2DL4 D2 domain with HLA-G involve residues conserved in the heavy chains of HLA-B and HLA-G and residues conserved in the KIR3DL1 and KIR2DL4 D2 domains. It is therefore likely that the KIR2DL4 D2 domain contacts HLA-G in a similar way.

      As for the KIR2DL4 D0 domain, it is very different. Due to the similarity between D2 domains of KIR3DL1 and KIR2DL4, and to the lack of a D1 domain in KIR2DL4, the KIR2DL4 D0 domain is in a completely different space than the D0 domain of KIR3DL1. “Predictions” by AlphaFold show that there could be interactions between the KIR2DL4 D0 domain and HLA-G (Figures 7 and S7). These predictions could be wrong. Nevertheless, the disulfide switch in the KIR2DL4 D0 domain correlates with a predicted change elsewhere on D0 at a position compatible with proximity to HLA-G. Furthermore, the KIR2DL4 isoform with a Cys28-Cys74 bond is “predicted” to be more aligned with a potential binding site than the Cys10-Cys28 isoform. Having no structural guide as a reference on how KIR2DL4 D0 domain may interact with HLA-G, such predictions may generate testable hypotheses.

      As we clearly state in the manuscript: “Structures of KIR2DL4–HLA-G complexes obtained experimentally are required to determine how HLA-G distinguishes the D0 domain in the alternative disulfide-bonded configurations.” (Results), and “Rules that dictate HLA-G binding to KIR2DL4 await further studies and structures of KIR2DL4–HLA-G complexes.” (Discussion).

      In the revised manuscript, we have now included SPR binding data for KIR2DL4 with HLA-G. We also show a higher affinity of HLA-G for the C28-C74 form of KIR2DL4. This has strengthened the study as it validates our model whereby switching to the functional form of the receptor allows binding of HLA-G. In this regard, we also include data showing that only the C28-C74 form of KIR2DL4 can respond to HLA-G to induce transcription of an ISG response. This provides a functional correlate to the role of the different disulfide forms of the receptor.

      Reviewer #2 (Recommendations for the authors):

      Major points to address:

      (1) Exhibit relevance using primary NK cells (using primary NK). The central biological framing is decidual NK responses to soluble HLA-G during early pregnancy, yet most mechanistic work is in 293T transfectants. The authors can perform one of the critical experiments using primary NK cells with soluble HLA-G stimulation. They should test whether PDI inhibition/thiol blockade similarly alters KIR2DL4 surface retention and HLA-G uptake in primary NK cells

      These experiments have been performed in primary NK cells and are described in the new Figure 4E, F, G and Figure 5F.

      (2) The authors should detail more about the relevance of extracellular PDI and the effect of PDI perturbation on the abundance of the two disulfide forms in cells. They should also provide evidence or discuss whether switching occurs prior to ligand binding, upon ligand engagement, or during trafficking.

      Such experiments would be very challenging. The predicted structural change is minor and may not be detectable by changes in proximity of labeled reporters. Ligand is not required for switching. We do know that PDI can convert C10-C28 into C28-C74, presumably by accessibility to the KIR2DL4 Cys28 when bonded in a C10-C28 configuration (Figure 2). How ligands (mAb #33 or HLA-G) impact KIR2DL4 structure is unknown. Data are compatible with the possibility of a stabilization of C10-C28 by mAb #33 and of C28-C74 by HLA-G.

      (3) The authors should elaborate on whether mAb #33 activates by stabilizing or by driving internalization independent of HLA-G. This is very interesting to the reader, given mAb #33 as a KIR2DL4 agonist.

      The question is undeniably interesting. mAb #33 is not required for internalization but is required for signaling. The C10-C28 KIR2DL4 configuration to which it binds internalizes constitutively and resides mainly in endosomes. How mAb #33 internalization by KIR2DL4 (not the reverse) results in signaling is not known. Nor is it known for the alternative form, C28-C74, which binds HLA-G, internalizes it, and signals for a transcriptional response very similar to that of C10-C28 bound to mAb #33 [2]. The C28-C74 KIR2DL4 configuration is retained, probably transiently, at the cell surface, to be available for HLA-G binding and internalization.

      Minor points to address:

      (1) The authors should ensure that all imaging quantifications include n, the number of experiments, and statistical treatment. Some are described in the Methods section. Please align figure legends with the method in detail.

      Details of the imaging experiments are provided in the Methods section.

      (2) Please summarize the Table 1 logic in the main text for enhanced reader engagement.

      This has been done.

      (3) The authors identify both Cys10-Cys28 and Cys28-Cys74 states in human cells. The data points towards coexistence rather than towards dynamic conversion. Please provide clarity on switching versus stable coexistence of two forms.

      Stable coexistence of two distinct KIR2DL4 receptor pools was a plausible hypothesis but one that is not supported by some of our data. In such a scenario, the C10-C28 form would not bind HLA-G and would reside in endosomes. It could have a role that is not related to HLA-G nor to the transcriptional response induced by HLA-G. However, our recent paper [2] showed that the transcriptional response of primary NK cells to soluble mAb #33 (bound to C10-C28) is very similar (R<sup>2</sup>=0.89) to that of resting NK cells incubated with soluble HLA-G (bound to C28-C74). These two ligands were tested at the same time, at the same molarity, and with the same primary NK cells [2].

      We don’t have answers yet to some obvious questions: is there switching after internalization of KIR2DL4 bound to mAb #33? What is the fate of C28-C74 that internalizes with HLA-G? We are not aware of technology that would answer these questions.

      A C28-C74 form, as a separate pool with residency at the cell surface, could be functional and respond to HLA-G by internalization and signaling from endosomes. However, there is no stable pool of C28-C74 KIR2DL4 at the cell surface and C28-C74 is depleted from the cell surface in the presence of PDI inhibitor (new Figure 5E, F), suggesting that C28-C74 KIR2DL4 is generated by the activity of PDI (new Figure S5). The sum of our experiments points to a tightly regulated control of KIR2DL4 biology, rather than the coexistence of two separate pools. A separate pool of C10-C28 KIR2DL4 would remain in an inactive state as far as the response to HLA-G is concerned. We favor the model whereby functional C28-C74 is generated from C10-C28 by the activity of PDI.

      Why could the response to HLA-G not be simpler? We address this point in the Discussion. One reason is that C28-C74 KIR2DL4 signaling at the plasma membrane of NK cells could be subject to inhibition by LILRB1 and NKG2A-CD94, co-expressed on NK cells, which bind to HLA-G and HLA-E, respectively. These inhibitory receptors are known to be dominant against activation signals [3]. Trophoblast cells that invade the maternal decidua express HLA-E and HLA-G and encounter decidual NK cells that express LILRB1 and NKG2A-CD94. Strong inhibition signals induced by these two receptors could prevent activation through KIR2DL4. KIR2DL4 signaling in endosomes protects it from these inhibitory signals and benefits from the sustained signaling property of endosomal signaling platforms [5].

      (1) S. Moradi et al., The structure of the atypical killer cell immunoglobulin-like receptor, KIR2DL4. J Biol Chem 290, 10460-10471 (2015).

      (2) S. Rajagopalan et al., The fetal trophoblast cell marker HLA-G activates a type I interferon response in primary NK cells through the receptor KIR2DL4. Sci Signal 19, eadv2400 (2026).

      (3) E. O. Long, H. S. Kim, D. Liu, M. E. Peterson, S. Rajagopalan, Controlling natural killer cell responses: integration of signals for activation and inhibition. Annu Rev Immunol 31, 227-258 (2013).

      (4) S. Rajagopalan et al., Activation of NK cells by an endocytosed receptor for soluble HLA-G. PLoS Biol 4, e9 (2006).

      (5) M. Miaczynska, L. Pelkmans, M. Zerial, Not just a sink: endosomes in control of signal transduction. Curr Opin Cell Biol 16, 400-406 (2004).

      (6) J. P. Vivian et al., Killer cell immunoglobulin-like receptor 3DL1-mediated recognition of human leukocyte antigen B. Nature 479, 401-405 (2011).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors characterize the phospholipid scramblase Xkr in Drosophila. They generate null mutants in both S2 cells and flies and find that phosphatidylserine (PS) exposure is reduced during apoptosis; they show reduced engulfment of apoptotic cells, and that the protein is localized partially within the cytoplasm, overlapping with the ER. They go on to identify Xkr binding partners and show that they overlap with plasma membrane-ER contact sites, suggesting that Xkr facilitates PS transfer from the ER to PM. Overall, this reveals a new role for Xkr and identifies new binding partners, which are valuable contributions to the field.

      Strengths:

      (1) The generation of new Xkr reagents in both S2 cells and flies to analyze its function. Tools are used to quantify both PS exposure and efferocytosis, and the effects of Xkr knockout are significant.

      (2) The discovery of new binding partners of Xkr which also affect PS exposure and efferocytosis.

      (3) The authors demonstrate that the binding partners are conserved in mammalian cells.

      Weaknesses:

      (1) Throughout the manuscript (e.g, lines 105, 165, 274 and discussion), the authors describe Xkr as being activated in a caspase-independent manner, and use this as the rationale for identifying binding partners. However, this is never shown in the manuscript or clearly referenced. Interestingly, there is a TEVDA sequence in the fly ortholog at the same location as the caspase cleavage site in C. elegans Ced-8 (Figure S1), suggesting the caspase cleavage site is conserved. This should be further investigated, or the statements regarding caspase independence should be modified. I don't think the N- and C-terminal GFP fusions indicate caspase independence, especially since apoptosis was not induced in Figure 1A, B. If cleavage occurred at the TEVDA site in Figure S1A, it would not lead to a noticeable change on the Western blot, although the size does look a bit smaller in Figure S2B at the 8 h time point.

      We thank the reviewer for pointing out that, as E/DXXD has been considered a conserved caspase-3 cleavage site, TEVDA has also been validated as a caspase-6 cleavage site, which we have missed. We will further confirm this using site-mutated expression vectors in S2 cells.

      (2) The authors examine overlap between tagged Xkr and cellular compartment markers and find substantial overlap with Lamp (and other vesicle markers to a lesser extent) (Figure S2). This is not addressed in the paper and could indicate engulfment of other cells since S2 cells are macrophages. To test this, the staining could be tested on the mixed cells (vesicle-GFP tagged S2 + apoptotic xkr-mcherry). Similarly, calreticulin is an eatme signal that gets translocated to the PM of apoptotic cells. This could affect interpretation of colocalization (Figure 2J), and ideally another ER marker should be used.

      We thank the reviewer for the suggestion. We will attempt to label Xkr-mCherry under apoptosis with other vesicles and change the ER marker to Cnx99A (Calnexin ortholog in Drosophila).

      (3) There are some places where there is over- or incorrect interpretation, and these instances should be corrected.

      We thank the reviewer for their careful reading, and we will correct the mistakes in the revised manuscript.

      Specific examples:

      a) Line 342 "Relative expression analysis by RT-qPCR showed that all three mutants were likely null alleles." This does not make sense since there is still mRNA present. In Figure S7A, the tm9sf4 allele is expressed at 75% of the control. The others show a greater reduction, but this is not proof of a null allele.

      We agree with the reviewer’s opinion. These mutants from the BDSC are not completely deleted but partially deleted; therefore, the RT-qPCR assay may not be very accurate. We will detect the mRNA levels of tm9sf4, dorp9, and sac1 using RT primers from different cDNA regions to make the results more convincing.

      b) Figure S3I - It looks like mCherry-Lact:C2 does get localized to the PM with AcD treatment in the xkr[ko], although the authors conclude "this disrupted PS localization to the PM could not be restored by apoptosis induction". However, the PM localization does look disrupted in the tm9sf4 and sac1 knockdowns.

      We thank you for raising this intriguing hypothesis. Indeed, PM localization of Lact:C2 was reduced in xkr<sup>ko</sup> cells, and the distribution could not be rescued after apoptosis. Unlike xkr<sup>ko</sup>, tm9sf4, and sac1 RNAi-treated cells displayed weak PS disorder, which may be due to the efficiency of knockdown. However, the statistical results indicated that the ratio of PM/Cyto was reduced in tm9sf4 and sac1 RNAi-treated cells.

      c) Figure 3I. The control Lact:C2 staining looks very different from the staining in Figure 2J, with abundant Lact:C2 outside the cell. Given the variability in the staining, were the contact sites quantified? On lines 287-288, it is stated that "fewer ER-PM MCSs were detected in xkrko cells than in WT", but no quantification is provided.

      We thank for the reviewer’s suggestion. We will add the statistical results of Fig. 3I in the revised version.

      d) Line 299-300 - "the interaction between Xkr and dORP9 was enhanced after apoptosis induction". The interaction does not look enhanced in Figure S5F, so this statement should be removed or data supporting the statement should be provided. The interaction between Xkr and dORP2 looks enhanced upon apoptosis induction, but also paradoxically looks even more enhanced when apoptosis is blocked.

      We thank you for raising this intriguing hypothesis. We will delete the relevant statement to eliminate unnecessary misunderstandings.

      e) The data in Figure S6 are highlighted in the abstract. If this is a major conclusion, it would be best to move it to the main text and provide quantification.

      We thank for the reviewer’s suggestion. We will move this to the main text and provide quantification in the revised version.

      f) Lines 392-4. The concluding statement seems overstated given that there was only a modest inhibition of PS exposure in the osbpl5 knockdown (Figure 6A) and no defects in efferocytosis (Figure 6C). The osbpl8 showed a stronger effect on PS exposure but still a very modest effect on efferocytosis.

      We thank for the reviewer’s suggestion. We will weaken the statement in the Results section of Figure 6 and perform osbpl9 knockdown to observe efferocytosis in Raw264.7 cells, as OSBPL9 interacts with Xkr8 strongly.

      Reviewer #2 (Public review):

      In this study, the authors investigate the mechanisms underlying phosphatidylserine (PS) exposure during efferocytosis in Drosophila. They first show that Xkr promotes PS exposure and apoptotic cell clearance in both S2 cells and Drosophila embryos. As Drosophila Xkr lacks the canonical caspase cleavage site found in mammalian XKR proteins, the authors further explore the underlying mechanism by which Xkr regulates PS externalization. Through protein interaction studies, they identify TM9SF4 as an interacting partner of Xkr that regulates PS distribution and show that non-vesicular PS transport contributes to apoptotic PS exposure and efferocytosis. Using protein interaction studies, they further demonstrate that Xkr interacts with the lipid transfer protein dORP9 at ER-PM contact sites to facilitate non-vesicular PS transport to the plasma membrane. Loss of these proteins affects PS externalization and efferocytosis in Drosophila. Finally, using human cells, they demonstrate that human OSBPL8 interacts with XKR8 to regulate apoptotic PS exposure. Overall, the study supports a model in which Xkr promotes efferocytosis by facilitating lipid transport in addition to its role as a phospholipid scramblase.

      Thank you for your comprehensive and generous assessment of our work and for the time and expertise you have devoted to reviewing our manuscript. We will revise the manuscript accordingly and provide a point-by-point response in the revised version.

      Reviewer #3 (Public review):

      Summary:

      The manuscript investigates the function of the Drosophila Xkr protein, a homolog of mammalian Xkr8 that lacks the canonical caspase-cleavage motif. The authors show that apoptotic stimuli increase Xkr protein abundance through a post-transcriptional mechanism and that Xkr promotes phosphatidylserine (PS) exposure during apoptosis. Using immunoprecipitation coupled with mass spectrometry, they identify TM9SF4 as an Xkr-interacting protein and further implicate TM9SF4, Sac1, dORP2, dORP9, and Vap33 in regulating apoptotic PS exposure and efferocytosis. Based on these findings, the authors propose that Xkr regulates PS transport at ER-PM contact sites. Similar observations are also presented in human cells.

      Strengths:

      Overall, this is an interesting study. The authors provide convincing evidence that Drosophila Xkr participates in apoptotic PS exposure and employ multiple complementary approaches to support the involvement of several proteins in this pathway. The identification of TM9SF4 as a potential regulator of Xkr-mediated PS exposure is likely to be of broad interest.

      Weaknesses:

      I am less convinced by the evidence supporting the proposed role of ER-PM contact sites, and several mechanistic conclusions appear to extend beyond the data presented. Addressing the following points would substantially strengthen the manuscript.

      We sincerely thank you for your careful reading and accurate summary of our manuscript. We appreciate the time, effort, and expertise you have dedicated to evaluating our work, and we will try our best to improve our manuscript according to your suggestions.

      Major concerns:

      (1) In Figure 2A and related text, it is unclear whether the mass spectrometry analysis was performed using untreated cells or AcD-treated cells. If the objective was to identify apoptosis-associated Xkr interactors, it would be helpful to clarify the experimental condition and explain whether apoptosis-specific interactors were analyzed separately.

      We thank for the reviewer’s suggestion. We used AcD-treated S2 cells and untreated S2 cells to perform mass spectrometry. To clarify this, we will add a detailed method description in the method section.

      (2) In Figure 2B, 2E, and several other co-IP results, a negative control of Flag tag only is required to exclude experimental errors like insufficient washing, etc.

      We thank for the reviewer’s suggestion. We used anti-HA magnetic beads to perform immunoprecipitation, and single HA-TM9SF4 was used as a negative control.

      (3) In Figure S3B, S3F, and several other BiFC results, an mVC-only negative control would be important to exclude nonspecific fluorescence complementation.

      We thank for the reviewer’s suggestion, we will add the negative control for BiFC results in the revised version.

      (4) In Figure 2G, the quantitative values appear inconsistent with the flow cytometry histograms. The peak shift following Sac1 knockdown appears smaller than that of TM9SF4 knockdown, whereas the quantified values suggest the opposite. Please clarify this apparent discrepancy.

      We sincerely thank you for the careful consideration of our statistical results, which were obtained from 3 repeats. We will choose another flow cytometry histogram of tm9sf4 and sac1 to make the data and images more consistent.

      (5) I find the interpretation in Lines 223-227 difficult to reconcile with the data. Knockdown of both tm9sf4 and sac1 impaired apoptotic PS exposure to a similar extent as xkr knockout. However, while xkr deficiency significantly reduced efferocytosis, sac1 knockdown produced only a modest, statistically insignificant effect. These observations suggest that impaired PS exposure alone may not fully account for the efferocytosis phenotype observed in xkr-deficient cells. These results appear difficult to reconcile with the proposed model, which needs careful discussion.

      We sincerely thank the reviewer for their careful and thoughtful observations. Given the results we have observed, we will add this to the discussion section in the revised version.

      (6) In Lines 274-275, the authors state that 'increased Xkr may accelerate non-vesicular PS transport for efficient apoptotic PS exposure'. However, Xkr protein levels increase only ~8 h after AcD treatment, whereas PS exposure occurs much earlier. Thus, alternative explanations like Xkr relocalization (Figure S5C), rather than increased abundance, may also explain how Xkr mediates PS transport. An Xkr overexpression experiment could be helpful to support this statement.

      We thank for the reviewer’s suggestion. We will overexpress Xkr with or without AcD treatment to observe whether the localization or amount of Lact:C2 changes and to re-evaluate the role of Xkr in PS exposure.

      (7) The interpretation of the MAPPER experiments requires further clarification. In Line 283, the authors refer to "the intracellular proportion of the signal for each protein overlapping with MAPPER." Since MAPPER is designed to label ER-PM contact sites, which are located on the plasma membrane, intracellular MAPPER fluorescence likely represents the ER network rather than bona fide ER-PM contacts. Throughout the manuscript (including Figure S6, etc.), intracellular MAPPER puncta appear to be interpreted as ER-PM contacts, which may not be appropriate. In contrast, the peripheral MAPPER puncta observed along the cell cortex (e.g., Figure S5C after AcD treatment) are more consistent with authentic ER-PM contact sites. It is also not obvious that these cortical MAPPER signals colocalize with Xkr(Figure S5C). Thus, while the data support a role for the ER, they do not yet convincingly demonstrate Xkr clustering at ER-PM contact sites.

      We thank the reviewer for the suggestion, and we believe that the TIRF technique can help us demonstrate the ER-PM signal. Since our college has no TIRF microscope, we will try our best to seek cooperation from other colleges to achieve this experiment.

      (8) In the Xkr knockout cells, all fluorescence signals appear substantially low in intensity. Differences in protein distribution are difficult to interpret when overall probe expression also appears altered. It would be helpful to demonstrate that probe expression levels are comparable between conditions. Furthermore, as noted above, intracellular MAPPER signal may primarily represent ER rather than ER-PM contacts. Finally, despite the reduced signal intensity, the remaining MAPPER and PS signals still appear well colocalized in the knockout cells, similar to the observations in Figure 2J. The interpretation in Lines 285-288 should therefore be reconsidered.

      We sincerely thank the reviewer for this careful and thoughtful observation, and we agree that the interpretation in Line 285-288 is overstated. To explain this, we plan to detect the Lact:C2 and MAPPER signals in S2 and xkr<sup>ko</sup> cells with or without AcD to confirm how Xkr regulates PS via ER-PM under apoptotic conditions.

    1. Author response:

      We are pleased that the reviewers found the study conceptually novel and the analytical framework rigorous. In response we have substantially revised the manuscript to clarify methodological details, temper several interpretations, expand discussion of alternative explanations, and include additional analyses using the existing dataset. We have deliberately revised the manuscript so that our conclusions are limited to those directly supported by the data, namely that physiologically identified RVM pain-modulatory neurons exhibit structured dynamics spanning multiple temporal scales. We do not interpret the slow fluctuations as evidence for a specific intrinsic oscillator or for a causal role in physiological state regulation. We have also expanded the rationale for the lightly anaesthetized preparation, emphasizing that it provides both the recording stability required for prolonged single-unit recordings from sparse neurons in the deep RVM and a controlled physiological setting in which the baseline temporal organization of the circuit can be characterized while minimizing ongoing sensory, motor, and behavioral influences.

      Regarding the rationale for the lightly anaesthetized preparation, these experiments take advantage of the well-validated lightly anaesthetized Sprague-Dawley rat in which much of the foundational data concerning physiology and function of RVM neurons was obtained. This “middle-out” strategy [1] has allowed direct connections between the activity and pharmacology of identified RVM neurons and altered nociceptive behavior. This protocol demonstrably spares the essential links between brainstem pain-modulating neurons and nociceptive transmission pathways. Although the focus here was on ongoing activity, precluding the repeated nociceptive testing needed to link neuronal activity to nociceptive threshold, previous work has demonstrated that ongoing activity of OFF and ON-cells is correlated with nociceptive sensitivity [2] and that alterations in OFF- and ON cell firing in response to pharmacological manipulation and in models of persistent pain states, stress, and sickness have behavioral relevance [3–6,6–24]. Further, conclusions from work in lightly anaesthetized rats have repeatedly been found to be congruent with behavioral observations by other groups in awake rats and mice [19,25–36] and with functional imaging evidence in humans [37–40]. The lightly anaesthetized model has thus established a circuit-level explanatory framework for behavioral findings obtained in several species in multiple laboratories.

      A further consideration for the present study is that the lightly anaesthetized preparation allows us to examine the underlying temporal organization of the RVM under controlled conditions, without the additional factors that would necessarily come into play in an awake animal. Dynamics would inevitably be influenced by ongoing sensory input, behavioral priorities, arousal and other internal state changes. These factors would make it difficult to distinguish the intrinsic dynamics of the descending pain-modulatory system from the effects of the animal’s constantly changing experience.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors hypothesized that “RVM neurons operate across multiple temporal scales, integrating fast responses associated with reflex-linked control with slower fluctuations reflecting ongoing network or state-dependent modulation”. The hypothesis was tested with the established ON/OFF-cell model and probabilistic modeling. The study is conceptually interesting and methodologically sophisticated. The findings build toward the conclusion that pain-control circuits operate across multiple timescales.

      Strengths:

      The use of Bayesian regression and Gaussian process modeling to quantify and characterize recovery dynamics and ongoing oscillatory activity.

      The authors show that slow rhythmic activity appears preferentially in ON- and OFF-cells but not in NEUTRAL-cells, suggesting that the oscillations are related to pain-modulatory circuitry rather than being a generic feature of all recorded neurons.

      The observation that some oscillatory activity is coherent with autonomic measures aligns with broader views of the RVM as a hub integrating nociceptive and homeostatic regulation.

      Some pitfalls are appreciated and discussed by the authors, including the functional significance of slow fluctuations, the influence of anesthetics on global brain-state dynamics, the molecular profiles of the studied ON- and OFF-cells, and the heart rate as a covarying signal of RVM neuronal activity.

      Weaknesses:

      A general weakness is that the work is mostly descriptive and relies on anesthetized preparations. Whether the observed rhythms occur in awake animals and are linked to fluctuations in pain behavior needs to be confirmed in future studies.

      The study measures limited autonomic variables. The causal relationship between “slow fluctuations” and “ongoing physiological state” is unclear and overstated, since the data presented appear correlational.

      ON- and OFF-cells in the RVM are identified by their responses correlated with reflexive activity. The significance of the observed oscillations in spontaneous pain conditions is unclear.

      It is uncertain whether the observed rhythms truly reflect intrinsic RVM organization rather than anesthesia-dependent phenomena; the authors appreciated this pitfall, though.

      The Gaussian process analysis suggests predictability and quasi-periodicity, but predictability alone does not necessarily imply a true biological oscillator.

      Conclusion:

      The results support the authors’ hypothesis. The findings provide a compelling conceptual message about the multiscale organization and dynamics of descending pain-control circuits and encourage further studies on the topic.

      We thank Reviewer R1 for their thoughtful and balanced assessment of our work. We are grateful for the reviewer’s positive evaluation of the conceptual framework, the analytical methodology, and the conclusion that the results support the hypothesis that RVM pain-modulatory neurons operate across multiple temporal scales. We agree with the reviewer’s central assessment that the present study is primarily descriptive and that several important questions regarding the origin and functional significance of the slow dynamics remain unresolved. We also appreciate the reviewer’s emphasis on clearly distinguishing observations directly supported by the data from their mechanistic interpretation. Although the manuscript already acknowledged that the present findings do not establish the mechanistic origin of the slow dynamics, we agree that this distinction could be made more explicit. We have therefore revised the manuscript to clarify that the approximately 5-minute fluctuations represent structured, quasi-periodic activity whose underlying origin cannot be determined from the present experiments. Throughout the manuscript we now explicitly acknowledge that these dynamics may arise from interactions between RVM circuitry and broader physiological or network processes, including anaesthesia-related state modulation, autonomic regulation, or other slow network influences. We also emphasise that the relationship between RVM activity and heart rate is correlational and does not establish a causal interaction. Finally, we now discuss more explicitly the complementary roles of controlled lightly anaesthetized and awake preparations. The present preparation was chosen to characterize baseline RVM dynamics under controlled sensory and behavioral conditions, whereas future awake studies will be important for determining how these dynamics are expressed and modulated during ongoing behavior, sensory experience, and chronic pain.

      (1) In the abstract, the ”timescales” are vaguely stated as ”rapid activation”, ”fast recovery dynamics,” and ”slow dynamics”. Quantifying these expressions with approximate ranges (milliseconds, seconds, tens of seconds, minutes, etc.) whenever possible would benefit readers.

      We thank the reviewer for this helpful suggestion. We have revised the Abstract to provide approximate timescales for the different phases of neuronal activity, distinguishing the rapid stimulus-evoked response (sub-second), recovery dynamics (seconds to hundreds of seconds), and ongoing quasi-periodic fluctuations (approximately 5 minutes). We believe these revisions improve the clarity of the Abstract and better convey the central findings of the study.

      (2) The abstract states that ”Effective pain therapies increasingly target neural circuits...” The connection to therapy is not clarified in the manuscript. A brief statement about how temporal dynamics might influence neuromodulation, analgesic interventions, or chronic pain could strengthen translational impact.

      We appreciate this suggestion. We have revised both the Abstract and Discussion to better explain the potential translational relevance of our findings. Rather than making a broad statement regarding pain therapies, we now briefly discuss how understanding the temporal organisation of descending pain-modulatory circuits may ultimately inform the design and timing of neuromodulatory interventions. We also emphasise that these implications remain speculative and require future investigation.

      (3) To address inter-animal and inter-neuron variability. Are the effects consistent across animals? Are all neurons oscillatory? Are the reported timescales driven by a subset of cells?

      We thank the reviewer for raising this important point. We have expanded the Results and Discussion to clarify the degree of variability observed across neurons and animals. In particular, we now emphasise that the slow quasi-periodic dynamics are not uniformly expressed across all neurons, but rather represent a structured population-level phenomenon with variability in predictability and modulation strength between cells, particularly within the OFF-cell population. We also clarify the consistency of the observed timescales across animals and discuss this variability as an important feature of the underlying circuitry rather than evidence for a single homogeneous oscillatory process.

      (4) Discuss what circuit mechanisms generate the oscillations. Are they driven by inputs from the PAG or intrinsic to RVM?

      We agree that the mechanisms underlying the slow temporal dynamics are an important question. We have expanded the Discussion to consider several possible sources of these dynamics, including intrinsic RVM circuitry, descending inputs from higher-order structures, and broader physiological or brain-state fluctuations. We emphasise that the present experiments cannot distinguish between these possibilities and have revised the manuscript to make this limitation more explicit while highlighting it as an important direction for future work.

      Reviewer #2 (Public review):

      Using electrophysiological recordings in a well-characterized animal model of acute pain, and analytical and modeling methods, the authors show that descending pain-modulatory neurons in the rostral ventromedial medulla (RVM) operate across various timescales. They have both rapid multi-phase responses to noxious stimuli that unfold over tens of seconds, with distinct fast and slow recovery dynamics. Additionally, they generate slow quasi-periodic oscillations with approximately 5-minute periods during ongoing activity. These oscillations are statistically predictable and cell-type specific, demonstrating that descending pain control is organized through structured temporal dynamics that encompass immediate stimulus-evoked responses and slower fluctuations associated with physiological state.

      A novel discovery is a 5-minute quasi-periodic oscillation in ongoing ON- and OFF-cell activity. This oscillation, along with its coherence with heart rate, forms the basis for the claim that descending pain circuits exhibit intrinsic multi-timescale organization. However, it’s crucial to demonstrate that this periodicity is independent of external experimental cycles such as methohexital infusion pharmacokinetics, servo-controlled temperature regulation, or slow autonomic feedback loops, all of which operate on similar timescales. For instance, the 300-second period closely matches typical drug infusion cycling and thermoregulatory feedback intervals. Therefore, heart-rate coherence peaks at multiples of this period could equally reflect a shared external driver rather than intrinsic RVM organization. Although the absence of this cyclic structure in Neutral cells argues against this possibility, the authors might want to explicitly discuss this potential confound.

      The findings are important and novel in that they characterize an intriguing structure in the activity of ON and OFF neurons in the RVM. However, in the absence of a causal manipulation causality can only be inferred. That there is no phase-dependence of withdrawal latency argues against a causal role. The author are encouraged to qualify their conclusions (and their title) accordingly. Because anesthesia can affect global dynamics, this might affect the oscillations reported. Without awake validation, it remains uncertain whether these rhythms reflect an intrinsic property or an anesthesia-induced regime. Again, the absence of oscillations in Neutral cells argues against this possibility, but it is still possible that ON/OFF cells are embedded in different circuits that are affected differently by anesthesia.

      We thank Reviewer R2 for their careful and constructive assessment of our work and for recognising the novelty of identifying structured multi-timescale dynamics in physiologically characterised RVM neurons. We particularly appreciate the reviewer’s thoughtful consideration of alternative explanations for the observed low-frequency temporal structure.

      The reviewer raises an important question regarding the extent to which the approximately 5-minute quasi-periodic dynamics reflect processes generated within descending pain-modulatory circuitry versus broader physiological or experimental influences. As discussed in the original manuscript, the present experiments cannot determine the precise mechanistic origin of these dynamics, and we have revised the Discussion to make this distinction more explicit. We now consider possible contributions from autonomic regulation, thermoregulatory processes, anaesthesia-related state modulation, and other slow physiological influences. We also clarify that the NEUTRAL-cell population argues against a uniform global effect acting similarly across all RVM neurons, but cannot exclude systemic influences that preferentially engage ON- and OFF-cell circuitry.

      We have additionally expanded the rationale for the lightly anaesthetized preparation. This preparation was not used solely for technical convenience. Stable single-unit recordings from physiologically identified ON- and OFF-cells are technically challenging because the RVM is a deep brainstem structure and these functional cell classes are relatively sparse; suppression of spontaneous movement therefore permits substantially greater recording stability over the prolonged epochs required here. Importantly, the preparation also provides a controlled physiological setting in which the underlying temporal organization of RVM activity can be examined while reducing the continuously changing sensory, motor, arousal, and behavioral influences that would necessarily contribute to RVM activity in an awake animal. Awake preparations are essential for determining how RVM neurons respond during ongoing behavior and natural sensory experience, but that is a complementary question to the one addressed here: whether physiologically identified RVM neurons exhibit structured temporal dynamics under controlled conditions.

      We have therefore revised the manuscript to present the lightly anaesthetized preparation as both a methodological choice and an important boundary condition on interpretation. We continue to acknowledge that anaesthesia may influence slow network dynamics, and that future awake recordings will be required to determine how the temporal structure identified here is expressed in the behaving animal. We have also revised the title and several sections of the manuscript to ensure that our conclusions consistently reflect the correlational nature of the data and do not imply mechanistic or causal interpretations beyond those directly supported by the experiments.

      (1) Analyses of many of the ON-cells had longer training windows (> 1 sec) compared to those for the NEUTRAL cells. Could this have reduced the ability to fit and validate periodicity for the latter cell type?

      We thank the reviewer for raising this important point. The difference between ON/OFFand NEUTRAL-cell analyses reflects the available recording durations rather than differences in the Gaussian process fitting procedure. All cell classes were fitted using the same GP model and training strategy; however, some NEUTRAL-cell recordings were shorter (960 s versus 1500 s for ON- and OFF-cells), resulting in correspondingly shorter training segments. Whilst the minimum frequency recoverable from a 960 s training segment is 0.00104 Hz, meaning that the 0.0033 Hz frequency observed in ON- and OFF-cells would be recoverable if present in a 960 s recording. We agree that shorter recordings could, in principle, reduce the ability to estimate slow periodic structure. We have therefore clarified this point in the Methods and Discussion. Importantly, the absence of predictable low-frequency dynamics in NEUTRAL-cells is supported not only by GP prediction performance but also by the independent power spectral analysis, the low latent GP variance, and the near-flat phase-normalised reconstructions, suggesting that the difference between cell classes is not solely attributable to recording duration. To address this concern, we will revise the manuscript to repeat the GP analysis after truncating the ON- and OFF-cell recordings to match the duration of the NEUTRAL-cell recordings. We will also include a 960-second-long simulated NEUTRAL-cell recording with periodic structure, to demonstrate that this would be located by our method if present.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript by Ashworth and colleagues, the authors investigate the temporal dynamics of the rostral ventromedial medulla (RVM), a key output node in a major descending pain-modulation circuit. Using data from extrasellar single-unit recordings of RVM ON, OFF, and NEUTRAL cells in lightly anesthetized rats, the authors’ computational modeling yielded two major findings: (1) heat-evoked ON burst and OFF pause, followed by exponential recovery components in10s of seconds; and (2) ON and OFF cells exhibit periodic fluctuations in 5-minute cycles that are statistically predictable.

      Strengths:

      The manuscript’s concept is innovative, offering the first quantitative analysis of multitimescale dynamics in physiologically characterized RVM pain-modulating neurons. This advances a field that has mostly depended on qualitative or single-timescale descriptions. The authors use contemporary Gaussian process and probabilistic models to capture statistically predictable slow dynamics. The study is further strengthened by identifying ON-, OFF-, and NEUTRAL-type cells using well-established criteria grounded in decades of RVM research. The combination of rapid reflex-related responses and slower ongoing rhythms supports a dual-timescale framework, providing a more integrated understanding of how these neurons may regulate reflex activity and state-dependent processes.

      Weaknesses:

      Several limitations are noted. Incomplete characterization of light anesthesia during recording sessions, such as methohexital stability and clear criteria for identifying “lightly anesthetized” states. While the NEUTRAL cell control is helpful, it does not fully address concerns about circuit specificity or systemic confounds. The findings are male-dominant, which may limit their generalizability. The synchrony between ON and OFF cells was suggested but not directly tested. The heart rate coherence with ON, OFF, and NEUTRAL cell activity results is intriguing but does not fully clarify how these neurons influence heart rate, particularly within the “lightly anesthetized” model.

      We thank Reviewer R3 for their thoughtful and constructive assessment of our work. We appreciate the reviewer’s emphasis on providing additional methodological detail and placing the findings within the context and limitations of the experimental preparation. In response, we have substantially expanded the Methods to provide a more complete description of the lightly anaesthetized preparation, the methohexital infusion protocol, physiological monitoring, and the rationale for the ongoing recording paradigm.

      We have also clarified why this preparation was appropriate for the question addressed here. In addition to enabling stable long-duration single-unit recordings from sparse, physiologically identified neurons in the deep RVM, the lightly anesthetized preparation provides a controlled physiological setting in which baseline temporal dynamics can be characterized while minimizing ongoing sensory, motor, and behavioral influences. We nevertheless acknowledge that anaesthesia may alter slow brain-state dynamics, and we now make this limitation more explicit throughout the manuscript. We have also revised the Discussion to more clearly acknowledge the predominantly male sample, the interpretation of the NEUTRAL-cell population as a comparison group rather than a definitive control for systemic effects, the limitations of inferring synchrony from pseudo-population data, and the correlational nature of the heart-rate coherence analysis.

      (1) How does the lightly anesthetized preparation affect evoked and oscillation activity modeling? Given that cell activities can be highly influenced by the state of sedation and the pharmacology of methohexital, detailing how light anesthesia was achieved and determined can help interpret the limitations of the current model. For example, did the methohexital rate adjustments occur during the ongoing activity period used for GP modeling? What specific criteria defined “lightly anesthetized” beyond stable paw withdrawal latency, such as stable respiratory rate, EMG (reflex vigor?), and core temperature? Given RVM activity coupled to autonomic/thermoregulatory circuits, data on these variables should be reported, or their absence should be acknowledged.

      We thank the reviewer for this important comment. We have substantially expanded the Methods and Discussion to describe both the rationale for the lightly anaesthetized preparation and the criteria used to maintain it.

      The preparation offers both technical and conceptual advantages for the present question. Technically, the RVM is a deep brainstem structure and physiologically identified ON- and OFF-cells are relatively sparse. Prolonged extracellular recordings therefore depend on maintaining stable electrode–neuron contact, which is readily disrupted by spontaneous movement. Light methohexital anaesthesia suppresses spontaneous movement while preserving nocifensive withdrawal responses and the canonical physiological response patterns used to identify ON-, OFF-, and NEUTRAL-cells. Conceptually, the aim of the present study was to characterize the baseline temporal organization of identified RVM neurons rather than to determine which sensory, cognitive, or behavioral events drive their activity in an awake animal. An awake preparation would necessarily introduce continuously changing sensory input, motor activity, arousal, behavioral priorities, and other internal-state variables, all of which are known to influence RVM activity. These are important influences in their own right, but for the present question they would make it more difficult to distinguish underlying temporal structure from activity driven by ongoing experience. We therefore view controlled lightly anaesthetized and awake preparations as complementary: the former is useful for identifying foundational circuit dynamics under controlled conditions, whereas the latter will be essential for determining how those dynamics are modified and expressed during natural behavior.

      This preparation has also been extensively used to establish the canonical relationship between ON-/OFF-cell activity and nocifensive responses, pharmacological modulation of the RVM, and top-down control from structures including the hypothalamus and amygdala, with many of these functional relationships subsequently confirmed in awake behavioral experiments. We have added this context to the revised manuscript.

      With respect to physiological monitoring, core temperature was continuously monitored and maintained at 36–37 °C, heart rate was monitored by EKG, and EMG was recorded to monitor withdrawal responses. Light anesthesia was defined functionally by preservation of a stable nocifensive withdrawal response in the absence of spontaneous movement. Respiratory variables were not recorded, and we now acknowledge this explicitly as a limitation. We have also clarified in Methods that the Methohexital rates were not adjusted during the recording windows used for the gaussian process analysis.

      We therefore agree that the findings must be interpreted within the context of the lightly anaesthetized preparation, but we do not view awake recordings as a direct substitute for the present experiment. Rather, awake studies provide the important next step of determining how the structured dynamics identified under controlled conditions are modulated by sensory experience, behavioral state, and ongoing cognition.

      (2) It is unclear how the absence of slow oscillations in NEUTRAL cells can be used as an internal control for anesthesia and systemic drift. It is unlikely that NEUTRAL cells are identified in every single-cell recording session for them to be used as a consistent internal control. Also, as the authors suggested that the shared modulatory inputs to ON/OFF cells explain the coordinating mechanism for ON/OFF rhythmicity, the lack of rhythmicity or coherence in majority of the NEUTRAL cells may indicate that they do not receive the same modulatory inputs as ON/OFF cells. Would this make NEUTRAL cells insensitive to systemic changes throughout the recording sessions? Do rhythmic vs. non-rhythmic cells differ in location within the RVM?

      We appreciate the reviewer’s important distinction. We agree that NEUTRAL-cells should not be considered a definitive internal control for anaesthesia or systemic physiological drift. NEUTRAL-, ON-, and OFF-cells were not necessarily recorded simultaneously within the same session, and the functional classes may differ in the systemic or modulatory inputs they receive. Our intended inference is therefore narrower: the absence of comparable low-frequency temporal structure in most NEUTRAL-cells argues against a uniform global process that imposes the same temporal pattern on all RVM neurons. It does not exclude anaesthesia-related, autonomic, thermoregulatory, or other systemic processes that preferentially influence ON- and OFF-cell circuitry. We have revised the manuscript throughout to make this distinction explicit and now refer to NEUTRAL-cells as an informative comparison population rather than as a definitive control for systemic influences.

      Indeed, as the reviewer suggests, differential sensitivity to common modulatory inputs could itself contribute to the distinction between ON/OFF- and NEUTRAL-cell dynamics. This interpretation is also compatible with the observation that a subset of NEUTRAL-cells shows low-frequency coherence with heart rate despite lacking the structured approximately 5-minute temporal dynamics observed in the ON/OFF populations.

      We additionally examined the reconstructed recording locations and found no obvious anatomical segregation between neurons showing stronger versus weaker low-frequency structure within the sampled RVM region. We now state this in the revised manuscript. We appreciate the reviewer’s point that there may be locational differences between RVM rhythmic and non-rhythmic cells, which should be addressed in future work; however, determining this would require a substantially larger sample size, for example with multichannel probe recording, for a valid analysis.

      (3) It is important to acknowledge that findings are effectively male-only (77 M and 6 F). Although a recent publication demonstrated that RVM ON and OFF cell activities do not differ substantially on an individual level between male and female rats, sex differences in RVM population dynamics remain unexplored. The current finding may not be generalizable to females.

      We thank the reviewer for highlighting this important limitation. We now explicitly acknowledge in the Discussion that the present dataset is predominantly male (77 males, 6 females) and therefore does not permit meaningful assessment of sex differences in population dynamics. Although previous studies suggest that individual ON- and OFF-cell responses are broadly comparable between sexes, the generalisability of the present findings to female animals remains unknown and should be addressed in future work.

      (4) It was suggested that strong synchrony exists within each functional population (e.g., ON and OFF cells). However, phase-relationship or coherence analyses were lacking. Since there were recoding sessions with > 2 cells/animals, were there enough recordings that contain simultaneous ON/OFF pairs to allow for these analyses?

      We thank the reviewer for this helpful suggestion. We agree that direct analyses of synchrony between simultaneously recorded neurons would provide valuable additional information. However, the number of simultaneous recordings containing identifiable ON/OFF-cell pairs was insufficient to support a robust phase or coherence analysis. We have therefore revised the Discussion to avoid implying that synchrony has been directly demonstrated and instead describe the results as evidence for consistent low-frequency temporal structure across recordings. We also identify direct analysis of synchrony in larger simultaneously recorded neuronal populations as an important direction for future work.

      (5) It was intriguing that the ON-cell population’s ongoing activity shows a predictive structure, while the OFF-cell population does not (Figure 5). However, this interesting asymmetry in ongoing activity between two cell classes was not adequately explained in the discussion. For example, since shared modulatory inputs were proposed as the coordinating mechanism for ON/OFF rhythmicity, how may this difference in ON and OFF rhythm predictivity occur?

      We appreciate the reviewer drawing attention to this interesting observation. We have expanded the Discussion to consider possible explanations for the greater predictability observed in ON-cells relative to OFF-cells. Although both populations exhibited similar dominant timescales, ON-cells displayed larger latent GP variance and more consistent predictive performance, whereas OFF-cells exhibited greater heterogeneity across recordings. We now discuss several possible explanations for this asymmetry, including differences in intrinsic cellular properties, network coupling, or modulation amplitude, while emphasising that the present data do not allow these possibilities to be distinguished.

      (6) The relationship between RVM activity oscillations and cardiac rhythms appears to be covariate but may not support the ”physiologically meaningful” claim with the current analysis. Additional discussion could help clarify the findings of a) how the RVM oscillation period of 300s relates to the heart rate peak/oscillation period of 600s (Figure 6d) and b) how NEUTRAL cells show heart rate coherence but lack rhythmicity.

      We thank the reviewer for this thoughtful comment. We have revised the Discussion to more carefully interpret the heart-rate coherence analysis. In particular, we now emphasise that the observed coherence demonstrates shared low-frequency temporal structure but does not establish a causal relationship between RVM activity and cardiac dynamics. We also discuss the relationship between the approximately 300-s RVM timescale and the broader low-frequency components observed in the heart-rate spectrum, noting the limited frequency resolution available at these timescales. Finally, we expand our discussion of the NEUTRAL-cell results, clarifying that significant coherence in NEUTRAL-cells despite the absence of comparable structured low-frequency firing dynamics is consistent with shared physiological influences acting on multiple cell classes without implying that the slow temporal structure originates within NEUTRAL-cells.

      References

      (1) Noble, D. The Music of Life: Biology beyond the Genome (Oxford University Press, 2006).

      (2) Heinricher, M. M., Barbaro, N. M. & Fields, H. L. Putative Nociceptive Modulating Neurons in the Rostral Ventromedial Medulla of the Rat: Firing of On- and Off-Cells Is Related to Nociceptive Responsiveness. Somatosensory & motor research 6, 427–39 (1989).

      (3) Barbaro, N. M., Heinricher, M. M. & Fields, H. L. Putative Nociceptive Modulatory Neurons in the Rostral Ventromedial Medulla of the Rat Display Highly Correlated Firing Patterns. Somatosensory & Motor Research 6, 413–425 (1989).

      (4) Heinricher, M. M., Haws, C. M. & Fields, H. L. Evidence for GABA-mediated control of putative nociceptive modulating neurons in the rostral ventromedial medulla: Iontophoresis of bicuculline eliminates the off-cell pause. Somatosensory & Motor Research 8, 215–225 (1991).

      (5) Heinricher, M. M. & Kaplan, H. J. GABA-mediated inhibition in rostral ventromedial medulla: Role in nociceptive modulation in the lightly anesthetized rat. Pain 47, 105–113 (1991).

      (6) Heinricher, M. M. & Tortorici, V. Interference with GABA transmission in the rostral ventromedial medulla: Disinhibition of off-cells as a central mechanism in nociceptive modulation. Neuroscience 63, 533–546 (1994).

      (7) Heinricher, M. M., McGaraughty, S. & Grandy, D. K. Circuitry Underlying AntiOpioid Actions of Orphanin FQ in the Rostral Ventromedial Medulla. Journal of Neurophysiology 78, 3351– 3358 (1997).

      (8) Heinricher, M. M., McGaraughty, S. & Farr, D. A. The role of excitatory amino acid transmission within the rostral ventromedial medulla in the antinociceptive actions of systemically administered morphine. Pain 81, 57–65 (1999).

      (9) Heinricher, M. M., McGaraughty, S. & Tortorici, V. Circuitry Underlying Antiopioid Actions of Cholecystokinin Within the Rostral Ventromedial Medulla. Journal of Neurophysiology 85, 280–286 (2001).

      (10) Heinricher, M. M., Schouten, J. C. & Jobst, E. E. Activation of brainstem N-methyl-daspartate receptors is required for the analgesic actions of morphine given systemically. Pain 92, 129–138 (2001).

      (11) McGaraughty, S. & Heinricher, M. M. Microinjection of morphine into various amygdaloid nuclei differentially affects nociceptive responsiveness and RVM neuronal activity. Pain 96, 153–162 (2002).

      (12) Heinricher, M. M. & Neubert, M. J. Neural Basis for the Hyperalgesic Action of Cholecystokinin in the Rostral Ventromedial Medulla. Journal of Neurophysiology 92, 1982–1989 (2004).

      (13) Heinricher, M. M., Martenson, M. E. & Neubert, M. J. Prostaglandin E2 in the midbrain periaqueductal gray produces hyperalgesia and activates pain-modulating circuitry in the rostral ventromedial medulla. Pain 110, 419–426 (2004).

      (14) Heinricher, M. M., Neubert, M. J., Martenson, M. E. & Gonc¸alves, L. Prostaglandin E2 in the medial preoptic area produces hyperalgesia and activates pain-modulating circuitry in the rostral ventromedial medulla. Neuroscience 128, 389–398 (2004).

      (15) Kincaid, W., Neubert, M. J., Xu, M., Kim, C. J. & Heinricher, M. M. Role for Medullary Pain Facilitating Neurons in Secondary Thermal Hyperalgesia. Journal of Neurophysiology 95, 33–41 (2006).

      (16) Ortiz, J., Heinricher, M. & Selden, N. Noradrenergic agonist administration into the central nucleus of the amygdala increases the tail-flick latency in lightly anesthetized rats. Neuroscience 148, 737–743 (2007).

      (17) Xu, M., Kim, C. J., Neubert, M. J. & Heinricher, M. M. NMDA receptor-mediated activation of medullary pro-nociceptive neurons is required for secondary thermal hyperalgesia. PAIN 127, 253 (2007).

      (18) Ortiz, J. P., Close, L. N., Heinricher, M. M. & Selden, N. R. α2-Noradrenergic antagonist administration into the central nucleus of the amygdala blocks stress-induced hypoalgesia in awake behaving rats. Neuroscience 157, 223–228 (2008).

      (19) Edelmayer, R. M. et al. Medullary pain facilitating neurons mediate allodynia in headache-related pain. Annals of Neurology 65, 184–193 (2009).

      (20) Martenson, M. E., Cetas, J. S. & Heinricher, M. M. A possible neural basis for stress-induced hyperalgesia. Pain 142, 236–244 (2009).

      (21) Heinricher, M. M., Maire, J. J., Lee, D., Nalwalk, J. W. & Hough, L. B. Physiological Basis for Inhibition of Morphine and Improgan Antinociception by CC12, a P450 Epoxygenase Inhibitor. Journal of Neurophysiology 104, 3222–3230 (2010).

      (22) Heinricher, M. M., Martenson, M. E., Nalwalk, J. W. & Hough, L. B. Neural basis for improgan antinociception. Neuroscience 169, 1414–1420 (2010).

      (23) McGaraughty, S., Farr, D. A. & Heinricher, M. M. Lesions of the periaqueductal gray disrupt input to the rostral ventromedial medulla following microinjections of morphine into the medial or basolateral nuclei of the amygdala. Brain Research 1009, 223–227 (2004).

      (24) Rogness, V. M. et al. Descending Facilitation of Nociceptive Transmission From the Rostral Ventromedial Medulla Contributes to Hyperalgesia in Mice with Sickle Cell Disease. Neuroscience 526, 1–12 (2023).

      (25) Smith, D. J. et al. Dose-Dependent Pain-Facilitatory and -Inhibitory Actions of Neurotensin Are Revealed by SR 48692, a Nonpeptide Neurotensin Antagonist: Influence on the Antinociceptive Effect of Morphine1,2. The Journal of Pharmacology and Experimental Therapeutics 282, 899–908 (1997).

      (26) Hurley, R. W. & Hammond, D. L. The Analgesic Effects of Supraspinal µ and δ Opioid Receptor Agonists Are Potentiated during Persistent Inflammation. Journal of Neuroscience 20, 1249–1259 (2000).

      (27) Kovelowski, C. J. et al. Supraspinal cholecystokinin may drive tonic descending facilitation mechanisms to maintain neuropathic pain in the rat. Pain 87, 265–273 (2000).

      (28) Porreca, F. et al. Inhibition of Neuropathic Pain by Selective Ablation of Brainstem Medullary Cells Expressing the µ-Opioid Receptor. Journal of Neuroscience 21, 5281–5288 (2001).

      (29) Zhang, Y. et al. Identifying local and descending inputs for primary sensory neurons. Journal of Clinical Investigation 125, 3782–3795 (2015).

      (30) Franc¸ois, A. et al. A Brainstem-Spinal Cord Inhibitory Circuit for Mechanical Pain Modulation by GABA and Enkephalins. Neuron 93, 822–839.e6 (2017). URL https://www.ncbi.nlm. nih.gov/pmc/articles/PMC7354674/.

      (31) Kim, J.-H. et al. Yin-and-yang bifurcation of opioidergic circuits for descending analgesia at the midbrain of the mouse. Proceedings of the National Academy of Sciences 115, 11078–11083 (2018).

      (32) Nguyen, E. et al. Medullary kappa-opioid receptor neurons inhibit pain and itch through a descending circuit. Brain 145, 2586–2601 (2022).

      (33) Jiao, Y. et al. Molecular identification of bulbospinal ON neurons by GPER, which drives pain and morphine tolerance. The Journal of Clinical Investigation 133, e154588 (2023).

      (34) Nguyen, E., Grajales-Reyes, J. G., Gereau, R. W. & Ross, S. E. Cell type-specific dissection of sensory pathways involved in descending modulation. Trends in Neurosciences 46, 539–550 (2023).

      (35) Fatt, M. P. et al. Morphine-responsive neurons that regulate mechanical antinociception. Science 385, eado6593 (2024).

      (36) Wang, Q. et al. Deconstruction of a spino-brain–spinal cord circuit that drives chronic pain. Nature 1–10 (2026).

      (37) Brooks, J. C., Davies, W.-E. & Pickering, A. E. Resolving the Brainstem Contributions to Attentional Analgesia. The Journal of Neuroscience 37, 2279–2291 (2017).

      (38) Mills, E. P. et al. Brainstem Pain-Control Circuitry Connectivity in Chronic Neuropathic Pain. Journal of Neuroscience 38, 465–473 (2018).

      (39) Mills, E. P., Keay, K. A. & Henderson, L. A. Brainstem Pain-Modulation Circuitry and Its Plasticity in Neuropathic Pain: Insights From Human Brain Imaging Investigations. Frontiers in Pain Research (Lausanne, Switzerland) 2, 705345 (2021).

      (40) Oliva, V., Hartley-Davies, R., Moran, R., Pickering, A. E. & Brooks, J. C. Simultaneous brain, brainstem, and spinal cord pharmacological-fMRI reveals involvement of an endogenous opioid network in attentional analgesia. eLife 11, e71877 (2022).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Liao et al. present SCOPE (Spatial reConstruction via Oligonucleotide Proximity Encoding), a method for reconstructing spatial organization from diffusion-defined DNA barcode interactions without the use of optical imaging. In SCOPE, hydrogel beads bearing unique DNA barcodes contain both "sender" and "receiver" oligonucleotides. Upon enzymatic release, sender oligos diffuse locally and hybridize to receiver oligos on neighboring beads, forming chimeric molecules that encode spatial proximity. Sequencing these products yields an interaction matrix, which is then used to reconstruct a spatial coordinate map.

      The authors demonstrate reconstruction of synthetic two-dimensional shapes, a large multicolor Snellen eye chart, and the interior surface of three-dimensional molds. The work expands the conceptual and experimental landscape of optics-free spatial sequencing.

      Thank you for this accurate summary of the work.

      Strengths:

      SCOPE employs bidirectional sender and receiver oligonucleotides on every bead, rather than using asymmetric transmitter-receiver architectures found in other diffusion-based methods. The symmetric design may improve detection sensitivity and reconstruction strategies, and represents a meaningful variation on optics-free spatial encoding.

      A notable strength of this study is the physical scale achieved. The authors reconstruct a Snellen chart spanning approximately 704 mm² and demonstrate molded 3D structures on the order of 75-100 mm³. Although some larger-scale warping is evident, and is discussed as potentially due to non-uniform diffusion, the relative local positioning across these large areas appears impressively accurate.

      The authors extend reconstruction beyond two-dimensional arrays to three-dimensional molded surfaces. This demonstrates that the assay and the computational methods for interpreting proximity graphs can support nonplanar spatial relationships, expanding the scope of optics-free spatial inference.

      Thank you for highlighting these strengths of SCOPE.

      Weaknesses:

      Although the method is discussed in the context of spatial genomics and potential tissue applications, it is currently demonstrated only on engineered two-dimensional bead arrays and three-dimensional shapes fabricated in molds. It remains unclear how SCOPE would perform in heterogeneous biological environments, where diffusion may exhibit additional non-uniformities. A biological proof-of-concept, even limited in scope, would help define the method's strengths and limitations more clearly.

      We concur with the reviewer that a biological proof-of-concept is a key next step, and that diffusion will be more heterogeneous in this more complex environment. To this end, we are actively working to further develop SCOPE for use in tissue sections, with the goal of capturing transcriptomes, accessible chromatin, and genomes. As part of this work, we also hope to systematically explore a range of tissue permeabilization and tissue clearing approaches to mitigate the impact of heterogeneity on performance.

      The reconstruction of three-dimensional structures lacks strong sampling from volume interiors. This is speculated to be due to several possible factors; however, this limitation constrains the method to reconstruction of volume surfaces rather than comprehensive three-dimensional profiling.

      Thank you for highlighting this important limitation. The 3D reconstructions are indeed constrained by undersampling of volume interiors. We anticipate that this might be addressed via relatively minor adjustments to the protocol, e.g. using light- or base-labile linkers to trigger oligo release, with the expectation that this will improve reaction consistency throughout the volume. However, even if we are unable to resolve this issue, we note that surface-resolved reconstructions may be useful for some goals, e.g. embedding a bead-packed gel within a tissue lumen, such as the gut. This could enable surface beads to capture RNA transcripts from adjacent cells, while bead–bead associations serve to define the surface topology.

      The reconstruction workflow involves multiple preprocessing steps and embedding choices. While these appear to work well for synthetic shapes with known geometry, it is less clear how parameter choices would be made in contexts where ground truth is unknown. Clarifying how reconstruction robustness is assessed without prior knowledge of spatial structure would help readers understand how the method could be practically deployed, particularly in more heterogeneous tissue contexts.

      Thank you for the opportunity to clarify. The computational pipeline used for 2D SCOPE reconstruction is designed to operate on a standardized input format and can be applied to arbitrary datasets without prior knowledge of spatial structure. For example, as shown in Figure 3, both the circle and “swoosh” geometries were reconstructed using the same algorithm and identical initial parameters. While certain hyperparameters are pre-specified (e.g. the number of k-nearest neighbours used to compute the pairwise distance matrix for UMAP), these are fixed across datasets. Other parameters, such as UMAP’s “min_dist,” are selected via an automated heuristic grid search that proceeds without user intervention. The agreement with ground truth in these controlled settings, together with the reproducibility of stochastic reconstructions (see Figure 3E-F), supports the robustness of the approach.

      Importantly, there was one exception. Reconstruction of the Snellen eye chart dataset required a manual step, involving an initial 3D UMAP embedding followed by a 2D projection to “flatten” the result. We suspect this reflects radial non-uniformities in sender/receiver oligo diffusion at larger spatial scales. Addressing such confounders algorithmically by explicitly modelling diffusion heterogeneity represents an important area for future work, with the goal of entirely eliminating the need for manual intervention.

      Finally, we note that these benchmark shapes represent somewhat contrived examples, and the geometries encountered in practice may often be much less complex. For example, in conventional spatial genomics, the geometry consists of a bead monolayer forming a flat, regular surface on a rectangular slide of known dimensions. Regardless of the tissue architecture overlaid on this surface, the reconstruction problem is defined by the bead monolayer itself, inferred through sender-receiver interactions.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It would be helpful to further clarify the limitations in interior sampling of three-dimensional structures by providing a more explicit comparison with Qian and Weinstein's volumetric DNA microscopy (UMI-UEI) approach. In particular, do the authors anticipate that the current limitation in SCOPE's interior sampling can be mitigated through experimental optimization, or might this represent an inherent challenge associated with hydrogel-based bead scaffolds relative to substrate-free approaches? A more detailed discussion of this point would help readers understand whether the observed volumetric constraint is technical and potentially solvable, or structural to the platform design.

      We thank the reviewer for this suggestion and agree that a more explicit comparison to volumetric DNA microscopy helps clarify the origin of the current limitation in interior sampling. Based on our experiments to date, we view this constraint as primarily technical and, in principle, addressable.

      A key distinction between SCOPE and volumetric DNA microscopy(Qian and Weinstein, 2025), using the UMI– UEI framework, lies in the recovery of recorded molecules from the interior of the sample. In volumetric DNA microscopy, the hydrogel-embedded specimen can be fully digested and treated with Proteinase K, enabling efficient liberation and recovery of molecules throughout the volume for downstream sequencing(Qian et al., 2026). In contrast, in the current implementation of SCOPE, we do not dissolve the polyacrylamide hydrogel, and recovery therefore relies largely on diffusion of chimeric molecules out of the scaffold. This likely biases against molecules generated in the interior and leads to reduced sampling of internal regions. This interpretation is supported by a control experiment in which barcoded beads were allowed to settle in solution at the bottom of a tube in the absence of a polymerized hydrogel scaffold. In this setting, 3D UMAP reconstruction yielded a solid, non-hollow structure consistent with the expected conical geometry of the tube bottom, indicating that SCOPE is capable of recovering volumetric structure when recovery is not diffusion-limited. Taken together, these observations suggest that the apparent “hollowing” in current 3D reconstructions reflects a limitation in molecule recovery from hydrogel scaffolds, rather than an inherent constraint of the SCOPE framework itself.

      We are currently exploring potential solutions, including the use of reducible crosslinkers to enable hydrogel dissolution and/or mechanical shearing of the gel. If these experiments are successful, we would plan to include the results in revisions to the manuscript, together with an appropriately edited version of the paragraph above. If they are unsuccessful and major experimental effort is going to be required to address this issue, we would likely move forward with textual changes only, incorporating the points made in the paragraph above into the discussion.

      References

      Qian N, Li J, Yasser R, Yu M, Weinstein JA. 2026. Volumetric DNA microscopy for mapping spatial transcriptomes in three dimensions. Nat Protoc. doi:10.1038/s41596-025-01329-3

      Qian N, Weinstein JA. 2025. Spatial transcriptomic imaging of an intact organism using volumetric DNA microscopy. Nat Biotechnol 1–11.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study by Akhtar et al. aims to investigate the link between systemic metabolism and respiratory demands, and how sleep and the circadian clock regulate metabolic states and respiratory dynamics. The authors leverage genetic mutants that are defective in sleep and circadian behavior in combination with indirect respirometry and steady-state LC-MS-based metabolomics to address this question in the Drosophila model.

      First, the authors performed respirometry (on groups of 25 flies) to measure oxygen consumption (VO2) and carbon dioxide production (VCO2) to calculate the respiratory quotient (RQ) across the 24-hour day (12h:12h light-dark cycle) and assess metabolic fuel utilization. They observed that among all the genotypes tested, wild type (WT) flies and per0 flies in LD and WT flies in DD exhibit RQ >1. They concluded the >1 RQ is consistent with active lipogenesis. In contrast, the short-sleep mutants fumin (fmn) and sleepless (sss) showed significantly different RQ; the fmn exhibits a slight reduction in RQ values, suggesting increased reliance on carbohydrate metabolism, while sss exhibits even lower RQ (0.94), consistent with a shift toward lipid and protein catabolism.

      The authors then proceeded to bin these measurements in 12-hour partitions, ZT0-12 and ZT12-24, to assess diurnal differences in average values of VO2, VCO2, and RQ. They observed significant day-night differences in metabolic rates in WT-LD flies, with higher rates during the day. The diurnal differences remain in the short-sleep mutants, but the overall metabolic rates are higher. WT-DD flies exhibit the lowest respiratory activity, although the day-night differences remain in free-running conditions. Finally, per01 mutants exhibit no significant change in day-night respiratory rates, suggesting that a functional circadian clock is necessary for diurnal differences in metabolic rates.

      They then performed finer-resolution 24-hour rhythmic analysis (RAIN and JTK) to determine if VO2, VCO2, and RQ exhibit 24-hour rhythmic and if there are genotypespecific differences. Based on their criteria, VCO2 is rhythmic in all conditions tested, while VO2 is rhythmic in all conditions except in fmn-LD. Finally, RQ is rhythmic in all 3 mutants but not in WT-LD and WT-DD. Peak phases for the rhythms were deduced using JTK lag values.

      The authors proceeded to leverage a previously published steady-state metabolite dataset to investigate the potential association of RQ with metabolite profiles. Spearman correlation was performed to identify metabolites that exhibit coupling to respiratory output. Positive and negative lag analysis were subsequently performed to further characterize these associations based on the timing of the metabolite peak changes relative to RQ fluctuations. The authors suggest that a positive lag indicates that metabolite changes occur after shifts in RQ, and a negative lag signifies that metabolite changes precede RQ changes. To visualize metabolic pathways that exhibit these temporal relationships, a clustered heatmap and enrichment analysis were performed. Through these analyses, they concluded that both sleep and circadian systems are essential for aligning metabolic substrate selection with energy demands, and different metabolic pathways are mis regulated in the different mutants with sleep and circadian defects.

      We thank the reviewer for summarizing the contributions made by this manuscript.

      Strength:

      The research questions this study explores are significant, given that metabolism and respiratory demand are central to animal biology. The experimental methods used, including the well-characterized fly genetic mutants, the newly developed method for indirect calorimetry measurements, and LC-MS-based metabolomics, are all appropriate. This study provides insights into the impact of sleep and circadian rhythm disruption on metabolism and respiratory demand and serves as a foundation for future mechanistic investigations.

      We thank the reviewer for the positive comments.

      Weaknesses:

      There are some conceptual flaws that the authors need to address regarding circadian biology, and some of the conclusions can be better supported by additional analysis to provide a stronger foundation for future functional investigation.

      At times, the methods, especially the statistical analysis, are not well articulated; they need to be better explained.

      Thank you for this suggestion and have revised and we expanded the Methods and figure legends to improve transparency and reproducibility.

      Specifically, we have:

      (i) Strengthened the rhythmicity description by specifying that rhythmicity was assessed in Nitecap using the RAIN algorithm with FDR-adjusted p-values (significant p ≤ 0.05; trending 0.05–0.1), and that period and peak phase were estimated with JTK_CYCLE (JTK lag = peak phase), with per-genotype period, phase, and p-values reported in Table 1;

      (ii) Clarified sample size and replication by adding these lines to the methods section and indicating sample sizes in figure legends. Figure legends now report n (chambers) and SEM.

      “Each genotype measurement represents an average of ~300 flies (25 flies per chamber × 4 chambers per experiment × 3 experimental days). The chamber was treated as the experimental unit for all analyses.”

      In addition, we have expanded the description of metabolomics-respirometry correlation analyses to include dataset structure, time-matching across ZT, normalization steps, the use of Spearman correlations, and interpretation of lagged associations.

      (iii) We have added these lines to the methods section:

      “To integrate respirometry with metabolomics, RQ was recorded continuously at 1-second resolution and averaged into 5-minute bins. Because steady-state metabolite measurements were acquired at 2-hour intervals, we extracted the RQ values corresponding to each 2-hour Zeitgeber Time (ZT) sampling point from the 5-minutebinned dataset to generate time-matched RQ-metabolite pairs. We additionally evaluated temporal relationships using a lag analysis by systematically shifting the RQ time series relative to the metabolite time points (−120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes). Metabolite abundances were normalized as described above, and associations between RQ and individual metabolites were quantified using Spearman rank correlations (ρ) at each lag. Metabolites showing strong associations (e.g., |ρ| > 0.7 with nominal p < 0.05) were carried forward for visualization and summary, and lag direction was interpreted as metabolites preceding (negative lag) or following (positive lag) changes in RQ.”

      Reviewer #2 (Public review):

      This is an innovative and technically strong study that integrates dual-gas respirometry with LC-MS metabolomics to examine how sleep and circadian disruption shape metabolism in Drosophila. The combination of continuous O<sub>2</sub>/CO<sub>2</sub> measurements with high-temporal-resolution metabolite profiling is novel and provides fresh insight into how wild-type flies maintain anticipatory fuel alignment, while mutants shift to reactive or misaligned metabolism. The use of lag-shift correlation analysis is particularly clever, as it highlights temporal coordination rather than static associations. Together, the findings advance our understanding of how circadian clocks and sleep contribute to metabolic efficiency and redox balance.

      We thank the reviewer for the positive comments.

      However, there are several areas where the manuscript could be strengthened.

      The authors should acknowledge that their findings may be gene specific. Because sleep deprivation was not performed, it remains uncertain whether the observed metabolic shifts generalize to sleep loss broadly or are restricted to the fmn and sss mutants. This concern also connects to the finding of metabolic misalignment under constant darkness despite an intact clock.

      We agree that our findings should be framed as genotype- and condition-specific. The phenotypes we report arise from chronic, genetically encoded sleep loss (fmn, sss) and clock loss (per01); because acute sleep deprivation was not performed, we do not claim these effects generalize to sleep loss broadly. This also bears on the reviewer's point about constant darkness: the metabolic misalignment we observe in WT-DD occurs despite an intact clock, and we therefore interpret it as a consequence of removing external light-dark cues under our conditions. We have scoped the claims accordingly in the Abstract, Results, and Discussion (subsection “Metabolic Desynchrony and Redox Imbalance in Wild-Type Flies Under Constant Darkness, DD”).

      The text now reads as follows:

      “We restrict our conclusions to the genotypes and conditions tested (fmn, sss, and per<sup>01</sup>), and we do not generalize these effects to acute sleep deprivation because sleep deprivation was not performed in this study. Accordingly, the ‘metabolic misalignment’ observed in constant darkness (DD) likely results from a decrease in synchrony due to the removal of external light:dark cues.”

      The conclusion that external entrainment is essential for maintaining energy homeostasis in flies may not translate to mammals. It would help to reference supporting data for the finding and discuss differences across species. Ideally, complementary circadian (lightdark cycle disruption) or sleep deprivation (for several hours) experiments, or citation of comparable studies, would strengthen the generality of the findings.

      Thank you. We have tempered the interpretation and expanded both the discussion and its citations. We now (i) avoid stating that external entrainment is universally “essential” for energy homeostasis, (ii) explicitly discuss fly-mammal differences (sleep architecture, thermoregulation, feeding control, and entrainment mechanisms), and (iii) anchor the translational comparison to the mammalian circadian-misalignment and sleep-loss literature already integrated in our Discussion (refs [3, 43-46]), noting that establishing cross-species generality will require additional paradigms (constant light, acute sleep deprivation).

      The text now reads as follows:

      “These phenotypes parallel mammalian systems, where sleep loss and circadian misalignment are linked to elevated basal metabolic rate, a shift toward carbohydrate oxidation and lipid/protein catabolism, and blunted, phase-shifted respiratory oscillations [3, 43-46]; physiological differences between flies and mammals nonetheless caution against direct mechanistic extrapolation. In constant darkness, our DD data show that endogenous free-running regulation persists but that removing external light-dark cues degrades temporal coordination between respiration and metabolism; however, we acknowledge that the coupling may be different in mammals.”

      Figures 1-4 are straightforward and clear, but when the manuscript transitions to the metabolite-respiration correlations, there is little description of the metabolomics methods or datasets, which should be clarified.

      Thank you for noting this. We agree that the transition to the metabolite–respiration correlation analyses required clearer description of the metabolomics datasets and processing. We have revised the Methods and the corresponding Results text to briefly summarize the metabolomics dataset parameters and workflow, including how metabolomics and respirometry measurements were time-matched across ZT, the normalization procedures applied prior to analysis, the use of Spearman rank correlations, and how we interpret lagged relationships between metabolite abundance and respiratory outputs.

      The text now reads as follows:

      Methods:

      “Metabolomics-respirometry integration and lag analysis

      To integrate respirometry with metabolomics, RQ was recorded continuously at 1-second resolution and averaged into 5-minute bins. Because steady-state metabolite measurements were acquired at 2-hour intervals, we extracted the RQ values corresponding to each 2-hour Zeitgeber Time (ZT) sampling point from the 5-minutebinned dataset to generate time-matched RQ-metabolite pairs. We additionally evaluated temporal relationships using a lag analysis by systematically shifting the RQ time series relative to the metabolite timepoints (−120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes). Metabolite abundances were normalized as described above, and associations between RQ and individual metabolites were quantified using Spearman rank correlations (ρ) at each lag. Metabolites showing strong associations (e.g., |ρ| > 0.7 with nominal p < 0.05) were carried forward for visualization and summary, and lag direction was interpreted as metabolites preceding (negative lag) or following (positive lag) changes in RQ.”

      Results:

      “Temporal Profiling of Respiratory Quotient in Wild-Type Flies Under Light-Dark Conditions

      RQ values corresponding to each 2-hour Zeitgeber Time (ZT) point were extracted from the 5-minute-binned dataset. Building on this alignment, we explored temporal relationships by systematically shifting the RQ time series by −120, −60, −30, −15, −5, +5, +15, +30, +60, and +120 minutes relative to the metabolite dataset. The continuous respirometry time series showed an oscillatory day-night pattern in RQ; metabolomics was then used to relate time-matched and lagged metabolite dynamics to RQ patterns (Figure 4).”

      The Discussion is at times repetitive and could be tightened, with the main message (i.e., wild-type flies align metabolism in advance, while mutants do not) kept front and center.

      Thank you for this helpful suggestion. We have revised the Discussion to reduce repetition and improve focus by keeping the central takeaway explicit throughout, and by consolidating overlapping paragraphs into a more streamlined narrative.

      We added this revision at the start of the Discussion, in the opening subsection “Temporal Misalignment Alters Fuel Utilization and Respiratory Rhythms.”

      The Discussion now reads as follows:

      “Across the manuscript, the central takeaway is that wild-type flies under LD exhibit anticipatory alignment of fuel selection with time of day, whereas short-sleep mutants (fmn, sss) and clock-disrupted flies (per01) show reactive or misaligned metabolism under our conditions. We therefore focus the Discussion on loss of temporal coordination between respiratory output and pathway-level metabolism, rather than reiterating rate changes alone.”

      Terms such as "anticipatory" and "reactive" should be defined early and used consistently throughout.

      Thank you for this suggestion. We agree and have revised the manuscript to define these terms early (at first use) and apply them consistently throughout. We added this definition in two places:

      (i) In the Results, at the start of the metabolomics-respirometry integration section where we first introduce the lag analysis, and

      (ii) In the Methods, within the paragraph describing the lag analysis workflow, using identical wording.

      The text now reads as follows:

      “We define ‘anticipatory’ as metabolite changes that precede the associated respiratory shift (negative lag) and ‘reactive’ as changes that follow or coincide with the respiratory shift (positive lag), and we use these terms consistently throughout.”

      Overall, this is a strong and novel contribution. With clarification of scope, refinement of presentation, and a more focused Discussion, the paper will make a significant impact.

      We again thank the reviewer for the positive comments.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate how sleep loss and circadian disruption affect whole-organism metabolism in Drosophila melanogaster. They used chamber-based flow-through respirometry to measure oxygen consumption and carbon dioxide production in wild-type flies and in mutants with impaired sleep or circadian function. These measurements were then integrated with a previously published metabolomics dataset to explore how respiratory dynamics align with metabolic pathways. The central claim is that wild-type flies display anticipatory coordination of metabolic processes with circadian time, while mutants exhibit reactive shifts in substrate use, redox imbalance, and signs of mitochondrial stress.

      We thank the reviewer for summarizing the contributions made by this manuscript.

      Strengths:

      The study has several strengths. Continuous high-resolution respirometry in flies is challenging, and its application across multiple genotypes provides good comparative insight. The conceptual framework distinguishing anticipatory from reactive metabolic regulation is interesting. The translational framing helps place the work in a broader context of sleep, circadian biology, and metabolic health.

      We thank the reviewer for the positive comments.

      Weaknesses:

      At the same time, the evidence supporting the conclusions is somewhat limited. The metabolomics data were not newly generated but repurposed from prior work, reducing novelty.

      Thank you for raising this point. We now make the provenance of the metabolomics dataset explicit in the manuscript. Importantly, the current study uses this dataset in a new analytical context by integrating it with continuous VCO<sub>2</sub>/VO<sub>2</sub> respirometry through timematched and lag-aware analyses. This approach allows us to evaluate dynamic relationships between respiratory output and metabolite profiles across circadian time, which was not addressed in the original metabolomics study. We have clarified this point in the Introduction and Methods.

      The text now reads as follows:

      In the Introduction:

      “To provide a more comprehensive perspective on metabolic regulation, we complemented newly generated respiratory measurements with steady-state metabolomic profiling using liquid chromatography-mass spectrometry (LC-MS) data previously published from our group [27]. This integrative framework enabled timematched and lag-aware analysis of respiratory output and metabolite profiles across Zeitgeber time in the LD cycle….”

      In the Methods:

      “The metabolomics dataset analyzed in this study was previously published and is publicly available, as described in detail in [27, 31]. In the present study, these data were integrated with respirometry measurements to assess temporal relationships between metabolite abundance and respiratory output.”

      The biological replication in the respirometry assays is low, with only a small number of chambers per genotype.

      Thank you for highlighting this concern. We suggest that this is a lack of clarity in our initial description of the design and that the replication structure should be stated more explicitly. We have revised the Methods and all relevant figure legends to clearly report biological replication using the chamber as the experimental unit, including n (number of chambers) per genotype and the associated error structure. We also clarify sampling depth by stating that each genotype measurement reflects an average of ~300 flies (25 flies per chamber × 4 chambers per experiment × 3 experimental days). This information is now reported consistently to make the unit of analysis transparent.

      We added this clarification in the Methods under “Respirometry Setup” (where chamber loading and experimental design are described) and ensured that each relevant figure legend explicitly reports n (chambers) and SEM.

      The text now reads as follows:

      “Each genotype measurement represents an average of ~300 flies (25 flies/chamber × 4 chambers/experiment × 3 experimental days), with the chamber as the experimental unit; n (chambers) and SEM are reported in each figure legend.”

      Importantly, respiratory parameters in flies are strongly influenced by locomotor activity, yet no direct measurements of activity were included, making it difficult to separate intrinsic metabolic changes from behavioral differences in mutants.

      A detailed timing comparison between behavior (feeding and locomotion) compared to respirometry is given in our master response to Reviewer 1, Major comment 4.

      In addition, repeated claims of "mitochondrial stress" are not directly substantiated by assays of mitochondrial function.

      Thank you. We agree that “mitochondrial stress” requires direct functional evidence. We therefore directly assayed mitochondrial respiration (baseline gut-tissue OCR in fmn and per01 versus iso31 controls), added a new Methods subsection and Figure 9, and reframed our wording from “mitochondrial stress/impairment” to altered (elevated) baseline mitochondrial respiration. The experimental details are now in the Methods and the result in the Results (both quoted below). sss was not assayed, so we removed the functional mitochondrial-stress claim for sss; we retain Vaccaro et al. (2020) as prior support for fmn.

      Methods — new subsection “Gut Tissue Respirometry” now reads: “Oxygen consumption rate (OCR) was measured in dissected gut tissue from iso31, fmn, and per01 flies using the Resipher System (Lucid Scientific, GA, USA). Baseline OCR (fmol/mm<sup>2</sup>/s) was averaged over a 24-hour window following a 12-hour acclimation; values from 3 independent runs were pooled, median-normalized to iso31 within each run, and log2(x+10)-transformed. A single iso31 outlier (third per01 run) was excluded; no other values were removed. Each mutant was compared with iso31 using the MannWhitney test (GraphPad Prism 10), with significance at p < 0.05 (fmn vs iso31, n = 12 vs 12; per01 vs iso31, n = 12 vs 11 after excluding one iso31 outlier).”

      The text has been added to Results:

      “Because pathway-level metabolomics implicated mitochondrial pathways in the sleep and circadian mutants, we directly assayed mitochondrial respiration by measuring baseline oxygen consumption rate (OCR) in dissected gut tissue from fmn and per01 relative to iso31 controls. Both fmn and per01 guts showed significantly elevated baseline OCR (fmn vs iso31, n = 12 vs 12; per01 vs iso31, n = 12 vs 11 across 3 runs; Mann-Whitney test, p<0.05; Figure 9A,B). Because only baseline OCR was measured, we interpret this as altered (elevated) baseline mitochondrial respiration rather than reduced capacity or a specific coupling defect (Figure 9).”

      Discussion- fmn:

      “These interpretations are further supported by gut-tissue respirometry showing elevated baseline mitochondrial respiration in fmn relative to iso31 controls. Together with prior evidence of ROS accumulation (oxidative stress) in fmn (Vaccaro et al., Cell, 2020), these functional data indicate that chronic sleep loss in fmn is associated with altered mitochondrial respiration.”

      Discussion- per01:

      “Gut-tissue respirometry in per01 likewise showed elevated baseline mitochondrial respiration, providing functional evidence consistent with the metabolomic signatures of disrupted redox balance and mitochondrial metabolism.”

      The study also excluded female flies entirely, despite well-documented sex differences in metabolism, which narrows the generality of the findings.

      Thank you for raising this point. We agree that sex is an important biological variable in metabolic regulation. While our Methods state that male flies were collected, we have now made this explicit and unambiguous by stating that only males were used for respirometry and metabolomics integration, and we have added this as a limitation in the Discussion, noting that sex-specific physiology could influence the magnitude and/or timing of the effects we report. We also highlight inclusion of females as an important future direction.

      The text now reads as follows (Methods):

      “Only male flies were used for all respirometry experiments and for integration with the metabolomics dataset. This was done to reduce variability introduced by female reproductive status (e.g., mating/egg production) and associated metabolic differences, enabling a clearer comparison across genotypes and lighting conditions.”

      The text now reads as follows (Discussion):

      “Because only males were analyzed, our conclusions may not generalize to females, which can show sex-specific metabolic physiology. Inclusion of female flies and direct sex comparisons across LD and DD conditions will be an important future direction. More specifically, females carry a higher reproductive and biosynthetic load (egg production) that typically raises metabolic rate and can shift RQ toward lipogenesis and alter the amplitude and phase of diurnal respiratory rhythms; females might therefore show larger or differently-timed effects than the males studied here.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments:

      (1) The authors appear to be using WT-DD as a condition to disrupt circadian rhythm (line 216). Although rhythmicity is often dampened in DD compared to in LD, circadian rhythm is defined as rhythm in a constant condition (e.g., DD) after entrainment. WT-DD is not a condition that the authors should use if they want to disrupt circadian rhythms in flies. WTLL would be a better condition to use since flies become arrhythmic in LL, not DD.

      Thank you for this clarification. We agree that DD is not a circadian-disrupting condition; circadian rhythmicity is defined by rhythms that persist under constant conditions after entrainment. Our intent was not to treat WT-DD as arrhythmic, but to use DD to assess free-running circadian regulation in the absence of external light–dark cues. We have revised the manuscript to clearly distinguish diurnal rhythms under LD from free-running circadian rhythms under DD, and to avoid implying that DD abolishes rhythmicity.

      The Results text now reads as follows:

      “We used constant darkness (DD) to assess free-running circadian regulation in the absence of external light–dark cues.”

      In the Discussion, in the subsection “Metabolic Desynchrony and Redox Imbalance in Wild-Type Flies Under Constant Darkness,” we revised the interpretation of WT-DD to clarify that the observed metabolic effects reflect removal of external light–dark cues under our experimental conditions, while avoiding overgeneralization.

      The Discussion text now reads as follows:

      Our DD data show that endogenous free-running circadian regulation persists but that removing external light–dark cues degrades the temporal coordination between respiration and metabolism, indicating that external entrainment normally strengthens this coordination under our conditions; however, we acknowledge that the coupling may be different in mammals.

      (2) The authors need to be more precise in the use of the term "circadian" throughout the manuscript. When describing 24h rhythmicity in the LD condition in flies, they can use the diurnal rhythm. Circadian rhythm is an endogenous rhythm without external time cues (e.g., DD rhythm).

      Thank you for this point. We agree and have revised the manuscript to use terminology consistently: rhythms measured under LD are now referred to as diurnal (LD) rhythms/patterns, and the term circadian is reserved for endogenous free-running rhythms in DD. We updated the Methods, Results, Discussion and figure legends throughout to correct instances where LD rhythmicity was previously labeled as “circadian.”

      We added this clarification in the Methods (“Drosophila Strains, Entrainment and Collection”) and in the Results/figure legends where LD time courses are described.

      The text now reads as follows (Methods):

      “Male flies were collected shortly after eclosion and entrained in light-dark (LD) incubators for a minimum of three days before diurnal (LD) time-course collection across Zeitgeber time (ZT).”

      We revised the Results text where WT-DD and LD time courses are described, replacing imprecise references to ‘circadian disruption’ or ‘circadian cycle’ with ‘free-running conditions in DD,’ ‘24-hour cycle,’ or ‘diurnal pattern under LD,’ as appropriate. We also revised the Discussion to avoid describing LD patterns as circadian and to avoid implying that DD disrupts circadian rhythmicity.

      (3) Lines 253-256: The authors' interpretation of this data is not accurate. The authors observed significant day-night differences in VO2 and VCO2 in WT-DD. This suggests there is circadian control over metabolism rhythms. The authors noted there is "limited circadian control".

      Thank you for pointing this out. We agree that the significant day-night differences in VO<sub>2</sub> and VCO<sub>2</sub> in WT-DD support persistent endogenous (circadian) control of respiratory rhythms under constant darkness. We have revised the Results text (lines 253-256) to remove the statement implying “limited circadian control” and instead describe the WTDD effect as maintained rhythmicity with altered amplitude and/or phase relative to LD, rather than loss of rhythmic regulation.

      We added this revision in the Results section under “Diurnal Variation in CO<sub>2</sub> Production, O<sub>2</sub> Consumption, and Respiratory Quotient Across Genotypes” (WT-DD description; lines 253-256).

      The text now reads as follows:

      WT-DD flies, maintained in constant darkness, exhibited the lowest overall respiratory activity. Despite the absence of environmental light cues, VCO<sub>2</sub> and VO<sub>2</sub> retained significant day-night differences (Figure S4), consistent with persistent free-running circadian control in constant darkness (with altered amplitude and/or phase relative to LD).

      (4) Besides sleep, metabolic rates are known to be affected by food consumption. Measuring food consumption of the sleep and circadian mutants might provide insights into whether the metabolic rates are more affected by changes in sleep profile or food consumption. This might also be important given fumin displays impaired dopamine transport function and defective dopamine reuptake, and dopamine is known to affect eating behavior. This was not considered and/or discussed.

      It is technically challenging to measure feeding during respirometry, so we acknowledge it as a limitation. To test whether behavioral timing could explain our results, we compared our respiratory rhythms with the feeding and activity rhythms reported in Malik et al. (2026); the comparison and its interpretation are now in the Discussion (quoted below). This is our master response to the activity/feeding concern and is crossreferenced from Reviewer 3’s public review and Recommendation 1; feeding and activity were not measured for sss.

      We added this clarification in the Discussion (Limitations/confounds) where we address potential behavioral contributors (activity/feeding) to respirometry outcomes.

      The text now reads as follows:

      “Locomotor activity and feeding could not be measured during the respirometry recordings, so genotype differences in respiratory parameters should be interpreted with caution. To assess whether behavioral timing could account for these differences, we compared our respiratory rhythms with the feeding and activity rhythms reported for these genotypes in Malik et al. (2026): wild-type feeding peaked at ZT ~3.25 and fmn feeding was phase-delayed to ZT ~4.5, while fmn also showed elevated locomotor activity, particularly during the dark period. In our data the fmn VCO2 rhythm peaks at ZT ~3.25 (RQ at ZT ~4.25), so the respiratory peak slightly precedes the feeding peak; the phase of the fmn respiratory rhythm is therefore not driven by feeding, although the elevated activity of fmn may contribute to its higher overall metabolic rate. Feeding and activity were not measured for sss.”

      (5) It is not clear whether the authors are simply analyzing the SAME dataset in Figure 1, Figure S4, and Figure 2-3, but just with different resolutions. They need to better articulate this point.

      We thank the reviewer for pointing this out. While the source data for Figures 1, 2-3, and Figure S4 use the same underlying respirometry dataset, they present different analyses to address specific questions. Figure 1 shows the full time-course traces (fine time bins), Figures 2-3 extract rhythmicity metrics (e.g., period/phase) from those same time-series, and Figure S4 collapses the same data into simple day (ZT0-12) vs night (ZT12-24) averages.

      We added this clarification in the Figure legends for Figure 1, Figures 2-3, and Figure S4, and also noted it in the Methods where the respirometry analysis outputs (time-series binning, rhythmicity analysis, and day/night averaging) are described.

      (6) Figure 3 and page 14: What are their criteria for differentiating "conserved" vs "distinct" phase alignment? It is not clear whether this conclusion is supported by any statistical analysis.

      We now specify that the “conserved” vs “distinct” phase descriptions refer to early- vs late-peaking rhythms, and we state the criterion explicitly: a phase was called “conserved” when it fell within ±3 h of the WT-LD peak. The revised text reads: “VO<sub>2</sub> peaked … suggesting conserved phase alignment, with all phases within ±3 h of WT-LD (Figure 3, Table 1)”; “RQ peaked … indicating distinct phase alignment, with the sleep-mutant phases ~8 h apart (Figure 3, Table 1).”

      The text now reads as follows:

      “VO<sub>2</sub> peaked …suggesting conserved phase alignment, with all phases within ±3 h of WT-LD (Figure 3, Table 1)”

      “RQ peaked … indicating distinct phase alignment compared to respiratory output, with phases of the sleep mutants 8 h apart (Figure 3, Table 1).”

      (7) Figure 4: The authors need to provide more details as to how the metabolite dataset was utilized to generate this figure and how they made the conclusion that their analysis "revealed a distinct circadian rhythmicity in RQ, characterized by oscillatory patterns indicative of coordinated substrate utilization across the day-night cycle".

      We have clarified how the respirometry and metabolomics data are used for Figure 4 and corrected the overstated rhythmicity claim:

      (1) The RQ patterning in Figure 4 is derived from the continuous respirometry time series, not from the metabolomics dataset, which is used only for the time-matched and lagged correlation analyses that relate metabolite dynamics to RQ. (2) We removed the statement that the analysis “revealed a distinct circadian rhythmicity in RQ”: RQ was not statistically rhythmic in WT-LD or WT-DD (Table 1), and Figure 4 instead shows the day-night RQ pattern that serves as the reference for the lag-based metabolite correlations. We revised the Methods, the Results paragraph introducing Figure 4, and the Figure 4 legend accordingly.

      The text now reads as follows:

      “Respiratory quotient (RQ) was recorded continuously and averaged into 5-minute bins. To integrate with metabolomics collected every 2 hours, we extracted the corresponding 2-hour ZT RQ values and performed a lag analysis (−120 to +120 min) to relate metabolite dynamics to RQ patterns; rhythmicity of RQ itself was assessed from the respirometry time series.”

      (8) It is unclear why the examples of hydroxyhexadecenoylcarnitine and quinolinate were chosen to be presented in Figure 5a. The authors should clarify their choice of these two examples. Also, the authors should generate a supplemental table with the "several metabolites demonstrating strong correlations (line 294).

      Thank you for this suggestion. Hydroxyhexadecenoylcarnitine and quinolinate were selected as representative examples, and we agree that the metabolites supporting the strong RQ-associated correlations should be provided more explicitly. We have now added a Supplementary Table listing the metabolites demonstrating strong correlations with RQ across WT-LD, fmn, sss, per<sup>01</sup>, and WT-DD conditions, using the same selection criterion applied in the heatmap analyses (|ρ| ≥ 0.7, p < 0.05).

      We also revised the Results text near the statement describing strong metabolite-RQ correlations to direct readers to this new table.

      The text now reads as follows:

      Hydroxyhexadecenoylcarnitine and quinolinate are among the strongest positively- and negatively-lagged RQ-correlated metabolites in WT-LD (ρ = +0.78 at +120 min and ρ = −0.77 at −120 min; Supplementary Table 1), illustrating the two opposite lag directions of the workflow. The full set of metabolites showing strong RQ-associated correlations across WT-LD, fmn, sss, per<sup>01</sup>, and WT-DD conditions is provided in Supplementary Table 1.

      (9) Although Spearman correlation analysis suggests some correlation between RQ and the two metabolites shown in Figure 5, the correlation shown in Figure 5b does not appear to be compelling. Results shown in Figure 5b do not provide confidence that conclusions based on clustered heatmap analysis shown in Figures 6 to 8 are meaningful. In addition to Spearman correlation, the authors might consider performing additional statistical methods to provide further support.

      Thank you for this comment. We suggest that Fig. 5b was not explained clearly and have clarified. Figure 5b is a lag analysis, not a separate correlation result: it shows how the Spearman correlation changes when the RQ time series is shifted forward or backward in time relative to the metabolite timepoints. The goal is to illustrate lead-lag timing (which shift gives the strongest association), rather than to present a single “strong” correlation as standalone proof.

      To address the concern about confidence in the heatmap-based results (Figs. 6-8), we have strengthened the reporting by providing effect sizes (Spearman ρ) and lag for the metabolite-respirometry associations (now included as Supplementary Table 1). This allows readers to evaluate the statistical support underlying the clustering, beyond the visual patterns in the heatmaps.

      We additionally report multiple-testing–corrected significance for the metabolite–RQ correlations (Benjamini–Hochberg FDR) alongside nominal p in Supplementary Table 1, using the same correction already applied to the pathway enrichment in Supplementary Table 2.

      Minor comments:

      (1) Line 82: The authors should clarify what they mean by "circadian collection". Except for WT-DD, my interpretation is that they collected their samples in LD, so that would not be "circadian collection".

      Thank you for catching this. We agree that “circadian collection” was imprecise. We have revised line 82 to clarify that samples collected under LD were collected across diurnal (LD) time (ZT), and we now reserve “circadian” specifically for collections under constant conditions (DD). We also updated the wording throughout the manuscript to maintain this distinction consistently. We added this clarification in the Methods section “Drosophila Strains, Entrainment and Collection” (line 82).

      The text now reads as follows:

      “Male flies were collected shortly after eclosion and entrained in light-dark (LD) incubators for a minimum of three days before diurnal (LD) time-course collection across Zeitgeber time (ZT).

      (2) The authors cited Frayn 1983 to indicate how the RQ value can be used to reflect metabolic fuel utilization. Is this interpretation accepted for all animals, including flies?

      RQ is widely used in indirect calorimetry as an index of relative substrate utilization, including in small model organisms, but we agree that it should be interpreted with appropriate caveats. We have revised the manuscript to clarify that we interpret RQ conservatively as reflecting relative shifts in substrate utilization over time and between genotypes, rather than as a precise quantitative measure of absolute carbohydrate versus lipid oxidation.

      We made this change in two places: in the Introduction, where RQ is introduced and Frayn is cited, and in the Methods, under “Carbon Dioxide and Oxygen Analysis and Calculations,” where RQ is defined.

      The text now reads as follows:

      Introduction: “These measurements allow for the estimation of energy expenditure and respiratory quotient (RQ), which can provide an index of relative substrate utilization, with appropriate caveats[13, 14]. In this study, we interpret RQ conservatively as reflecting relative shifts in substrate utilization over time and between genotypes, rather than as a precise quantitative measure of absolute carbohydrate versus lipid oxidation.”

      Methods: “RQ was used as an index of relative shifts in substrate utilization over time and between genotypes, interpreted conservatively rather than as a precise measure of absolute carbohydrate versus lipid oxidation as this has not been directly characterized in flies.

      (3) Figure S4: Since the authors are comparing day-night differences, they should plot them in the same graph to make it easier to compare.

      We have revised Figure S4 to plot day and night within the same graph/panel for each metric (VCO<sub>2</sub>, VO<sub>2</sub>, RQ), using side-by-side day vs night groupings per genotype to facilitate direct visual comparison, and we updated the legend accordingly.

      (4) Line 252: When comparing diurnal differences, the authors should not use the word "rhythm". Pattern or profile might be a better word to use.

      We appreciate the suggestion. Where we use “rhythm” we refer specifically to 24-hour oscillations established statistically by JTK_CYCLE and RAIN (Figures 2-3, Table 1); for the coarser day-versus-night comparisons we agree “pattern” or “profile” is preferable and have adopted it there. We have gone through the manuscript to apply this distinction consistently.

      (5) Figure 6a: larger font labels are necessary for the metabolites.

      Thank you. We have revised Figure 6a to increase the metabolite label font size for improved readability.

      Reviewer #3 (Recommendations for the authors):

      (1) Activity controls: To strengthen the paper, include or reference direct measures of locomotor activity (e.g., DAM system). This would allow the separation of metabolic changes from behavioral differences and would enable better analysis of circadian patterns.

      This activity/feeding confound is addressed in full in our response to Reviewer 1, Major comment 4; we cross-reference it here to avoid repetition.

      (2) Mitochondrial function: "mitochondrial stress" should be supported by additional assays such as mitochondrial enzyme activities, high-resolution respirometry, or reactive oxygen species measurements.

      See our full response to the mitochondrial point in Reviewer #3's public review above, including new Figure 9 and the Gut Tissue Respirometry Methods.

      (3) Sex differences: Provide a clear justification for excluding female flies. If feasible, incorporate female data or explicitly discuss how sex differences could alter metabolic outcomes.

      This is addressed in the public review for reviewer 3.

      (4) Statistical presentation: Increase n. Clarify in figure legends whether error bars represent SD or SEM and ensure consistency across all figures (Figure legends).

      We have revised all figure legends to clearly state that error bars represent SEM and ensured this is applied consistently across all figures. We also explicitly report the n for each genotype (with chambers as the experimental unit) in the relevant legends.

      The text now reads as follows:

      “Error bars represent SEM, and n denotes the number of chambers (experimental units) per genotype.”

      (5) Sample size and replication: Indicate more clearly that the chamber, not the individual fly, is the experimental unit. Discuss limitations of replication and statistical power in the text.

      Thank you for this comment. We have clarified throughout the Methods and figure legends that the chamber (not the individual fly) is the experimental unit. We now state that each genotype measurement reflects an average of 300 flies (25 flies/chamber × 4 chambers/experiment × 3 experimental days), and we report n (number of chambers) and the error structure (SEM) for each genotype.

      Added to Discussion:

      “We also acknowledge the limits of this replication: with n = 12 chambers per genotype, statistical power to detect small-magnitude differences and subtle phase shifts is limited, and negative calls (e.g., arrhythmicity) should be interpreted with this caveat.”

      (6) Writing and clarity: (a) Streamline the Discussion to focus on mechanistic themes (anticipatory vs reactive alignment, substrate shifts, redox imbalance).

      (b) The opening phrase ("Precise temporal regulation of metabolism by sleep and circadian rhythms is essential for dynamic energy homeostasis") is dense, vague, and difficult to interpret. Consider rephrasing to something more concrete, for example: "Sleep and circadian rhythms tightly control when and how the body uses energy, but we do not yet know exactly how this timing connects to oxygen use and breathing needs."

      Thank you for these helpful suggestions. We have revised the Discussion to reduce repetition and improve focus by organizing it around the key mechanistic themes raised by our data, including anticipatory vs reactive metabolic alignment, substrate-use shifts, and redox/mitochondrial imbalance. We revised the opening sentence of the Abstract to make the biological question more concrete and accessible, following the reviewer’s suggestion.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Some of the authors proposed in a PNAS paper in 2016 the occurrence of the EntnerDoudoroff (ED) pathway in cyanobacteria and plants, on the basis of several lines of biochemical and genetic evidence. However, more recent results indicated that one of the two specific enzymes of the ED pathway (EDD) is missing in Synechocystis PCC 6803. The authors carried out additional experiments, which demonstrated that EDD is missing, and one of the enzymes (ED aldolase) is a promiscuous enzyme which seems to be involved in proline metabolism and is not actually participating in the ED pathway as initially believed. The results described in this paper are strong evidence that this new interpretation is appropriate, and therefore, it corrects the previous proposal, providing an honest description of the reasons why the authors had reached the wrong conclusion about the existence of the ED pathway in cyanobacteria and plants.

      We thank Reviewer 1 for the summary and comments. Based on our finding that EDA is a promiscuous aldolase that, in addition to the cleavage of KDPG to GAP and pyruvate (a reaction of the ED pathway) catalyzes other reactions in vitro, we proposed potential in vivo functions of EDA, including its involvement in proline metabolism. However, these assumptions require further experimental testing. We do not yet have definitive findings regarding the function of the promiscuous aldolase EDA in Synechocystis in vivo, but respective studies are currently underway.

      Strengths:

      Thorough reanalysis of the experimental results obtained in previous studies, which led to the publication of the PNAS paper in 2016.

      New experimental evidence to confirm that enzymes previously considered as participating in the ED actually are not catalyzing the ED biochemical reactions, but are involved in other metabolic pathways. Also, the authors completely discarded the occurrence of the GDH/GK shunt in Synechocystis PCC 6803. Generally speaking, the manuscript is very clearly written, with a precise description of the previous findings, the mistakes which took place in the 2016 paper, and the strategies they have used to address those issues, in order to reach a thoroughly revised vision of the glucose metabolic pathways in Synechocystis PCC 6803. In this regard, the drawings shown in Figures 1 and 7 are very helpful for the reader to follow the story and understand the possible metabolic transformations depending on the working hypothesis.

      Also, I commend the authors for openly describing previous mistakes. In this paper, they reassess past observations in light of more recent findings and to integrate the information in this manuscript. The scientific conclusions are solid and very interesting, and besides, they use the opportunity to offer valuable advice to researchers. This is especially focused on the importance of careful biochemical characterization of enzymes, which should always be carried out when studying proteins which have been identified as a specific enzyme on the basis of sequence homology. In a similar way, they found that an insertional mutant was the cause of the absence of specific metabolites, which had been attributed to particularities of a metabolic pathway in that mutant, when it was actually due to a nucleotide insertion; this could have been easily prevented by confirming the correct generation of the mutant by DNA sequencing.

      We thank the reviewer for this kind comment. We agree that biochemical characterization of enzymes as well as DNA sequencing to check deletion mutants, are important and valuable tools. As outlined in the manuscript and additionally in more detail in a recently submitted article, which is available at bioRxiv (Theune et al. 2026, doi: https://doi.org/10.64898/2026.04.08.717167) and is currently under review at PLOS One, we suggest that genome sequencing of deletion mutants in combination with complemented strains as controls are required to minimize the risk of misinterpretation based on secondary mutations (1). During the early stages of our research on the ED pathway, and later as well when we were already trying to resolve the conflicting results that had accumulated concerning the ED pathway, genome sequencing for Synechocystis mutants was not affordable as a routine procedure (2-4). Therefore, we could not have easily prevented this misconception based on this technique at that time. However, we strongly encourage genome sequencing of deletion mutants (in combination with complemented strains) as routine procedures these days (1).

      Weaknesses:

      The authors propose that EDA might be involved in the PEP-pyruvate-OAA node, or in the proline metabolism, but this requires further experimental work for clarification; what their results indicate clearly is that this enzyme is not actually catalyzing the transformation of KDPG to GAP, which is the second specific enzyme of the ED pathway. But the real physiological function in this cyanobacterium is still unconfirmed.

      As stated above and in the manuscript, we agree that the in vivo role of EDA requires further experimental work which is in progress. However, our results demonstrate that EDA splits KDPG into GAP and pyruvate in vitro, but we assume that this reaction does not play a role in vivo due to the absence of its substrate.

      Another aspect which could be improved is that the recombinant expression of some genes was carried out in E. coli; even if this is a useful and valid research strategy, in studies like this (where there is a strong focus on the physiological function of enzymes in the original organism, Synechocystis PCC 6803), I think it would have been more appropriate to express the 6803 genes in another cyanobacterium easily amenable for genetic transformation and gene expression, which would produce the protein in a physiological environment more similar to another cyanobacterium (compared to E. coli, which is an heterotrophic bacterium). I am not sure this would change any of the obtained results, but it certainly would confer additional robustness to the enzymatic results.

      Synechocystis is easily amenable to genetic manipulation, and we agree that expression and purification of all enzymes from this host would have been ideal. However, the first characterization of Synechocystis EDA was performed with proteins that were purified from Synechocystis and showed activity on KDPG at comparable rates as proteins that were purified from E. coli in this study (2). Moreover, most biochemical characterizations of EDAs from archaea, bacteria and plants were performed after recombinant expression in E. coli and yielded highly active enzyme as in the case of Synechocystis is this study (5-7). Therefore, we currently have no reason to worry that the expression in E. coli might affect the enzymatic activity of EDA. The main reason for utilizing E. coli as an expression strain in this study was to gain higher yields of protein for in-depth analyses.

      Bibliography:

      I think the list of papers used in this manuscript is complete and up to date. However, I do miss recent papers which addressed one aspect that was proposed in the original 2016 PNAS paper: the authors wrote, "We therefore suggest that Prochlorococcus might oxidize glucose via the ED pathway under mixotrophic conditions, as shown for Synechocystis." Recent studies checked this hypothesis and have shown that the ED pathway seems to be also missing in Prochlorococcus and marine Synechococcus, and I think this manuscript is a good place to cite them, since these results are consistent with the findings of this paper.

      We will include a references from Moreno-Cabezuelo et a. 2023 (DOI: 10.1128/spectrum.03275-22) in which the proteomes of three marine Prochlorococcus and three marine Synechococcus strains were investigated upon exposure to glucose (8). Protein levels of EDA were either downregulated or not affected while proteins involved in OPP pathway and CBB cycle were upregulated. The authors of this study conclude that this indicates that the latter processes rather than the ED pathway are involved in photomixotrophy in these strains. However, flux analyses are still missing.

      Reviewer #2 (Public review):

      Summary:

      The study presents novel results on the presence of the Entner-Doudoroff pathway in Synechocystis sp. PCC 6803. In contrast to an earlier study, compelling evidence is given that this strain lacks both an ED pathway and a glucose dehydrogenase/glucokinase bypass but contains a promiscuous aldolase, which also decarboxylates oxaloacetate and cleaves 2-keto-4-hydroxyglutarate (as it occurs in proline degradation). The study concludes with successfully reconciling data from different studies and with lessons learned from the previous misconception.

      Strengths:

      Solid biochemical data are presented to reconcile contradicting data of earlier studies and to serve as a basis for disclosing possible functions of a promiscuous aldolase. Earlier misconceptions and lessons to be learned are well discussed.

      Weaknesses:

      The materials and methods section is rather lengthy, suffering from a lack of conciseness and repetition, and nevertheless misses some specifications.

      We thank Reviewer 2 for the kind summary and comments and will improve the materials and methods part accordingly in a revised version.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Some additional aspects that could be improved:

      (1) L182-184: Are there any known in vitro attempts to determine whether some DHADs can accept 6PG as substrate? If not, did the authors check this possibility in the lab?

      To the best of our knowledge, as mentioned in the manuscript, some DHADs are tested for their activity towards gluconate and some other substrates but not 6PG. It was suggested that since gluconate is smaller it might fit in the catalytic site of DHADs normally occupied by DHIV (7). We discussed this point in the manuscript (Line 177-185).

      In this study we tested the DHAD Slr0452 from Synechocystis and DHAD from Synechococcus with 6PG as substrate but the enzyme did not catalyze 6PG dehydration. The result for Slr0452 is shown in Figure S3.

      (2) L234-241: This paragraph shows how important it is to avoid relying only on sequence alignments to assign functions to proteins (and enzymes in particular). The authors stress this idea elsewhere in the paper, but I think it should receive even more attention in a manuscript like this. Physiological characterization of the protein function is paramount, especially in the case of enzymes. Also, the GDH1 overexpression mutant was done in E. coli; this adds an additional layer of uncertainty, since the protein processing in E. coli might not be entirely identical to that carried out in cyanobacteria, as I mentioned above. This, in turn, could be one of the reasons for not finding the expected enzymatic activity. This comment is also relevant for the results shown in Table 1 (page 12).

      As outlined following lines L234-241 we tested crude cell extracts from Synechocystis WT, a Synechocystis strain overexpressing putative GDH1 (Sll1709) and a Dzwf deletion mutant which can be assumed to upregulate a GDH/GK bypass if present in Synechocystis for GDH activity but did not find any. This strongly indicates that sll1709 does not code for an active GDH in Synechocystis. We thereafter also overexpressed GDH1 in E. coli and again could not detect any GDH activity. In case of GDH1, enzyme activity was therefore tested both in Synechocystis and E. coli and yielded similar results.

      (3) L309-310: "we currently have no explanation for the gluconate that was detected in previous IC-ESI-MSMS measurements in Synechocystis". This is a serious issue, given how important this evidence was for the conclusions of the 2016 PNAS paper and the hypothesis of the ED pathway in cyanobacteria. I would suggest that the authors provide some possible explanations, even if it is based on studies from other teams.

      It would be very speculative and therefore in our eyes not helpful to search for explanations in this case as indeed the measurements were done in another lab. One explanation might be that 6P gluconate got dephosphorylated and yielded gluconate as an artifact. However, as we have no experimental validation for this idea and furthermore cannot test it. We therefore prefer to not further comment on this aspect.

      (4) L347: The presence of an insertion in the sequence of zwf in the ∆gnd mutant is a welcome explanation for the observed results: abolition of 6PG production in this mutant. This is another important message to stress in this manuscript: construction of mutants should always be double checked by DNA sequencing of the relevant genomic regions, to ensure that these kinds of side problems do not appear, leading to confusing results.

      We agree with the reviewer and would get even one step further and rather suggest combining whole-genome sequencing with complemented mutants as it is difficult to know all relevant genomic regions that should be sequenced. We discuss this issue in even more detail in another manuscript that is currently available as online preprint in bioRxiv and is under review (9). This work was now added as a citation in line 566 and the end of the following statement (line 564-566): Routine complementation of deletion mutants and sequencing of selected genes or the entire genome are effective means of identifying secondary mutations that can lead to misleading phenotypes (9).

      (5) L418-419 and Table 2: The results observed for the authors (i.e., that EDA could also catalyze reactions with OAA and KHG, albeit with substantially lower catalytic efficiency than with KDPG), is a matter for concern: if their hypothesis is correct (meaning that this EDA is fundamentally involved in the PEP-pyruvate-OAA node and/or proline metabolism, and not with the ED pathway), then why should it keep in the evolution of these organisms such a strong preference for KPDG, when it is not being used physiologically for the ED pathway)? Furthermore, the Km for KDPG is lower than for OAA or KHG, leading to a very big difference in the Kcat/Km values.

      To solve these questions further respective studies are underway. As EDD is absent from Synechocystis no KDPG should be available in the cells so that catalytic activity on KDPG should be irrelevant in vivo. The in vivo role of Eda requires further clarification.

      Hereafter, I will mention some aspects, following the instructions of eLife, which are related to suggestions for improved experiments/data/analyses, improvements of writing and presentation, and minor corrections to text/figures.

      (1) L81: I would modify the text to "Accordingly, this raises further questions...".

      Thanks for this hint. We modified the text accordingly.

      (2) L189: Add "pages" after "following".

      Thanks for pointing this out. We replaced “following” by “below”.

      (3) Page 7: The whole beginning of the Results section is actually more discussion than description of results, and the first mention of figures appears in L207 of page 8. Given the content of the paper, I think the authors might reconsider using "Results and Discussion" rather than different, specific sections for Results and Discussion. This is one of the papers where I think the combined use of both makes sense and will allow an easier understanding of the message.

      We thank the reviewer for this valuable suggestion and changed the heading to Results and Discussion.

      (4) L181: Add reference regarding the llvD/EDD superfamily.

      We added the following references in lines 172-176 and 185-189:

      (1) Melse, O., Sutiono, S., Haslbeck, M., Schenk, G., Antes, I., & Sieber, V. (2022). Structure Guided Modulation of the Catalytic Properties of [2Fe− 2S]-Dependent Dehydratases. ChemBioChem, 23(10), e202200088.

      (2) Ren, Y., Vettenranta, E., Penttinen, L., Jänis, J., Rouvinen, J., & Hakulinen, N. (2025). The engineered dimer of L-arabinonate dehydratase from Rhizobium leguminosarum bv. trifolii: The role of intersubunit interactions in IlvD/EDD family. Biochemical and Biophysical Research Communications, 757, 151610.

      (3) Ahmed, H., Ettema, T. J., Tjaden, B., Geerling, A. C., Van Der Oost, J., & Siebers, B. (2005). The semi-phosphorylative Entner–Doudoroff pathway in hyperthermophilic archaea: a reevaluation. Biochemical Journal, 390(2), 529-540. ff

      (4) Bräsen, C., Esser, D., Rauch, B., & Siebers, B. (2014). Carbohydrate metabolism in Archaea: current insights into unusual enzymes and pathways and their regulation. Microbiology and Molecular Biology Reviews, 78(1), 89-175.

      (5) Figure 2: The data shown in column plots in Fig 2A, B and C, and 4B, could be presented in tables, which would save space while providing the same information: basically, very little/no activity in some cases vs high levels of activity in others.

      We would like to keep the column plots as we still think that they visualize our data well.

      (6) L255 "Unfortunately, we were not able to overexpress putative GDH2". It would be interesting to give more details about the possible reasons for this fact.

      We tested different growth conditions for recombinant GDH2 expression. The expression culture was incubated at 37 °C for 3 hours as well as overnight at 18 °C for overnight after induction. Both experiments did not resolve the expression problem.

      (7) L409-410: I think this sentence should include a brief part explaining the kind of essay used to test this activity.

      We added now that the LDH-coupled continuous assay was used (see line 403).

      (8) Figure 6E: Please give the specific activity in U/mg, as in Fig 6F, instead of percents.

      100% is given in U/mg units in the figure legend as “control without effector (100 %; specific activity of 4.3 U/mg)”. For easy comparison of effectors, the relative activity (%) is often used. We would therefore prefer to keep the current data presentation.

      (9) L576: Provide the origin of the utilized PCC 6803 strain, given there is a certain level of variability in this strain (glucose tolerance, etc). Also, even if there are some methods which are very widely used, I think the Materials and Methods section should either properly describe them or else cite the source. For instance, BG11 medium is mentioned, but no further information is given.

      We included the information that the glucose-tolerant Synechocystis strain was utilized and added the receipt of and a citation for BG11 medium (10).

      (10) L582 Generation of mutants: This section mentions the Gibson Assembly cloning method, but I miss further information to allow the reader to reproduce the methodology with as many details as possible, or at least cite papers which do so.

      We added a reference in which Gibson Assembly is described (11). Together with the primers listed in Table S3 the generation of mutants is now reproducible.

      (11) L609 Please give information in g, not rpm, for centrifugation. Also, mention the model and brand of the centrifuge and rotors used. Also, immunoblotting is very loosely described. This is also valid for other sections, for instance, L618, L636.

      We now added the following information: Cells were harvested by centrifugation at an RCF (relative centrifugal force) of 3,992 x g in a Beckman Coulter with a JLA-8.1000 Rotor for 20 minutes at 4°C. We now added a reference (12) in which immunoblotting is described in more detail.

      (12) L613: Describe the "small scale purification".

      We now added the information that the small-scale purification was performed using a 50-ml aliquot of the large culture which was treated as described below for the remaining sample.

      (13) L619-620: Describe the composition of the lysis buffer.

      The composition of the lysis buffer is already described as follows: lysis buffer (50 mM NaPO<sub>4</sub> pH=7.0; 250 mM NaCl; 1 tablet complete protease inhibitor EDTA-free

      (Roche) per 50 mL)

      (14) L691: Specify which amounts of auxiliary enzymes in coupled enzymatic assays were used.

      Thanks for pointing this out. We have now integrated the information that 1U of each of the auxiliary enzymes was utilized in the coupled enzymatic assays.

      (15) L716: The authors mention several times using a "double beam spectrophotometer". Please provide the model and brand.

      Model and brand were now added for the double-beam spectrophotometer (Uvikon 810, Kontron, Augsburg, Germany).

      (16) L717 and 718: define "∆absorption".

      In line 715, the information is given that absorption was measured at 340 nm, "∆absorption" is accordingly the ∆absorption at 340 nm. This information was added.

      (17) L723: "Synechocystis" should be in italics.

      Synechocystis is now written italics.

      (18) L749-759: This section should be described in more detail: preparation of protein extracts, SDS, immunoblotting, etc, or cite references of the same team where these methods were properly described.

      In this section the listed methods are already described in detail.

      (19) 798-799: "frozen cell pellets". Please provide numbers to specify the amount of material used.

      Thanks for pointing this out. We now added the information that frozen cell pellets with a wet weight of 2.4 g wet weight were resuspended.

      (20) L871-872: "It was ensured that auxiliary enzymes were not rate-limiting. One unit (1 U) of enzyme activity is defined as 1 µmol substrate consumed or product formed per minute" is repeated several times in the manuscript (L 907-909, L936-938). I would advise using it the first time, and on other occasions, refer to the same conditions as described above.

      We have accordingly circumvented the repetition of 1 U definition from the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Interpunctuation, especially comma placement, should be improved.

      We improved interpunctuation, especially comma placement to the best of our knowledge.

      (2) Line 63: delete the first "which".

      “Which” was deleted.

      (3) Lines 81/82: revise sentence.

      We revised the sentence to: Accordingly, this raises further questions about the presence of the ED pathway in cyanobacteria and plants.

      (4) Line 164: "presumed" instead of "presumes".

      The word was changed accordingly.

      (5) Line 228: by others? especially in reference 1?

      We deleted by others as the reference is given.

      (6) Line 240: "or" instead of "no".

      “no” was replaced by “or”

      (7) Figure 2: The axes are not well visible, and the explanation for the positive control in panel B is missing in the legend.

      Axes from figures 2, 3 and 4 were enlarged. For Figure 2B the following information was added in the legend: As a positive control, 0.05 U glucose dehydrogenase from Pseudomonas sp. was added to Δzwf cultures and to purified putative GDH1 (Sll1709). Axes from figures 2, 3 and 4 were enlarged.

      (8) The investigation on the general absence/presence of the GDH/GK bypass in cyanobacteria may not be necessary for this study.

      We included this data in this manuscript as the mistaken assumption that the GDH/GK bypass exist in Synechocystis lead among other observations to the misinterpretation of an existing ED pathway in Synechocystis. We would therefore prefer to keep these data in the manuscript.

      (9) Line 316: delete "on".

      “on” was deleted.

      (10) Line 352: delete "or".

      “or” was deleted.

      (11) Lines 352/353: ZWF expression level appears to be reduced accordingly. This should be stated.

      We agree that Zwf expression might be lower, however, we are not entirely sure if this is truly valid and would rather test this assumption further as described in the following lines.

      (12) Figure 4: The axes are not well visible.

      Axes from figures 2, 3 and 4 were enlarged.

      (13) Line 363: values are not only normalized to protein content, but also to activity found for the WT.

      We now added: The values are normalized to Zwf enzyme activity found in the WT based on protein content.

      (14) Line 408: no separate subsection required.

      The title for a new subsection was deleted.

      (15) Lines 437-439: These are results descriptions, which should not be placed in the legend, but in the main text, as is partially done.

      We deleted these result descriptions in the legend.

      (16) Lines 441/442: formatting: one or no bracket pair.

      The brackets were corrected.

      (17) Lines 443/444: refer to Table 2 instead of giving the values in the legend to avoid duplication.

      We deleted the values and now refer to Table 2.

      (18) Line 503: delete "identified".

      We deleted the second “identified” in the sentence and changed the wording to: Apart from four identified cyanobacteria that possess potential EDDs. In addition, we also added the names of the four cyanobacteria that were found including the sequence IDs of the putative EDDs.

      (19) Lines 529/530: revise sentence and format.

      We added one sentence and revised the following sentence: In contrast to Synechocystis EDA, EDA from Synechococcus prefers OAA over KDPG. The catalytic efficiency of Synechococcus EDA on oxaloacetate is rather low (OAA 0.437 s<sup>-1</sup> mM<sup>-1</sup>), however, its activity can be enhanced by NADP<sup>+</sup>(13).

      (20) Line 542: revise sentence.

      We revised the sentence to: It remains to be investigated whether this reaction could play a role in vivo, with KDPG potentially acting as a regulatory metabolite at low concentrations.

      (21) Line 577: The glass tubes used for cultivation should be specified.

      We now added the following information: Custom-made glass tubes were placed in a photobioreactor (manufactured by Willi Hilke, Uslar, Germany).

      (22) Lines 584-585: unclear, was the resistance cassette not placed in the gene to be deleted?

      Yes, the resistance cassette was placed in the gene to be deleted and was fused for homologous recombination to two DNA fragments approximately 200 bp directly upstream and downstream of the gene. This information was now added.

      (23) Line 607: cultivation equipment to be specified.

      The following information was now added: For the purification of GDH1 from Synechocystis, a 6 L photoautotrophic culture of the P3-His-GDH1 overexpression strain was grown in a 10 L glass flask at 28°C, illuminated with constant light (50 µmol m<sup>-2</sup> s<sup>-1</sup>) and gassed with filter sterilized ambient air to an OD<sub>750</sub> of about 1.

      (24) Line 670: GTS should be specified.

      Thank you for this hint. This was a typo. GST was meant not GTS. This was now corrected and GST was specified as Glutathione S-Transferase.

      (25) Line 679: delete "gluconate kinase (GK) and".

      The second gluconate kinase (GK) was deleted and sentence was revised to:

      For gluconate kinase (GK) activity measurements in Synechocystis crude cell extracts the GK reaction was enzymatically coupled to 6-phosphogluconate dehydrogenase (GND) reaction, the latter providing NADP<sup>+</sup> reduction, which was monitored photometrically at 340 nm.

      (26) Line 686: again GK activity? Difference unclear. Was the previously described procedure for GND activity determination?

      GK activity measurements in Synechocystis crude cell extracts and GK activity measurements with recombinant enzyme that was expressed in E. coli were done in two different labs with different protocols. Therefore, the first description refers to measurements with Synechocystis while the second measurement refers to measurements with E.coli. This is now specified more clearly.

      (27) Type/supplier of spectrophotometers and centrifuges used should be given.

      Model and brand were now added for the double-beam spectrophotometer (Uvikon 810, Kontron, Augsburg, Germany). As this study was performed in two different labs over a period of 8 years including one lab moving to a new location, it is now impossible to specify all centrifuges that were utilized. Even though we agree that it would be good to provide this information, we now would have difficulties to be specific.

      (28) Consider the referencing of published methods to streamline the materials and methods section.

      We now streamlined the materials and methods section by deleting repetitions as outlined below. However, as protein expression, protein purification and enzymatic tests were performed in different labs, in some cases several protocols are given.

      (29) Line 757: give specifics of anti-rabbit antibody and define PBS-T and PBS-T Cytiva.

      Specifics were added to the text.

      (30) Lines 761ff: It is not given for all genes used how they were derived. All synthesized?

      We now added detailed information for all genes.

      (31) Line 762: codon-optimized for? E. coli?

      The information was added that genes that were expressed in E. coli were codon-optimized for E. coli.

      (32) Lines 782-786: Rationals for experimental strategies do not belong to materials and methods sections, but to results sections.

      The part was deleted in the materials and methods section and transferred to the results section.

      (33) Lines 818/819: repetitive.

      We removed the repetition and refer to the purification method as stated above in the materials and methods section.

      (34) Lines 847-851: True for all EDA-type assays? Kinetic parameters are shown in Table 2 rather than Table 1.

      Yes, true for all EDA-type assays. We changed the Table number to 2.

      (35) Line 863: delete "in".

      “in” was deleted

      (36) Lines 888-893: sounds repetitive.

      The lines were modified accordingly.

      (37) Lines 908/909: repetitive.

      The repetitive comment on the definition of 1U was deleted.

      (38) Lines 928-938: repetition

      The repetition was deleted.

      References

      (1) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in <em> Synechocystis </em> PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (2) X. Chen et al., The Entner–Doudoroff pathway is an overlooked glycolytic route in cyanobacteria and plants. Proceedings of the National Academy of Sciences 113, 5441-5446 (2016).

      (3) D. Schulze et al., GC/MS-based 13C metabolic flux analysis resolves the parallel and cyclic photomixotrophic metabolism of Synechocystis sp. PCC 6803 and selected deletion mutants including the Entner-Doudoroff and phosphoketolase pathways. Microbial Cell Factories 21, 69 (2022).

      (4) A. Makowka et al., Glycolytic Shunts Replenish the Calvin–Benson–Bassham Cycle as Anaplerotic Reactions in Cyanobacteria. Molecular Plant 13, 471-482 (2020).

      (5) V. Zaitsev et al., Insights into the Substrate Specificity of Archaeal Entner–Doudoroff Aldolases: The Structures of Picrophilus torridus 2-Keto-3-deoxygluconate Aldolase and Sulfolobus solfataricus 2-Keto-3-deoxy-6-phosphogluconate Aldolase in Complex with 2-Keto-3-deoxy-6-phosphogluconate. Biochemistry 57, 3797-3806 (2018).

      (6) J. S. Griffiths et al., Cloning, isolation and characterization of the Thermotoga maritima KDPG aldolase. Bioorg Med Chem 10, 545-550 (2002).

      (7) S. E. Evans et al., Plastid ancestors lacked a complete Entner-Doudoroff pathway, limiting plants to glycolysis and the pentose phosphate pathway. Nature Communications 15, 1102 (2024).

      (8) J. Moreno-Cabezuelo, G. Gómez-Baena, J. Díez, J. M. García-Fernández, Integrated Proteomic and Metabolomic Analyses Show Differential Effects of Glucose Availability in Marine Synechococcus and Prochlorococcus. Microbiol Spectr 11, e0327522 (2023).

      (9) M. Theune et al., Easy-to-use whole-genome sequencing workflows and standardized practices to uncover hidden genetic variation in Synechocystis sp. PCC 6803 wild-type and knock-out strains. bioRxiv 10.64898/2026.04.08.717167, 2026.2004.2008.717167 (2026).

      (10) R. Y. Stanier, R. Kunisawa, M. Mandel, G. Cohen-Bazire, Purification and properties of unicellular blue-green algae (order Chroococcales). Bacteriol Rev 35, 171-205 (1971).

      (11) D. G. Gibson et al., Enzymatic assembly of DNA molecules up to several hundred kilobases. Nature Methods 6, 343-345 (2009).

      (12) M. Boehm et al., Comprehensive study on ferredoxin isoforms in the cyanobacterium Synechocystis sp. PCC 6803. bioRxiv 10.64898/2026.04.08.717189, 2026.2004.2008.717189 (2026).

      (13) N. Xie, C. Sharma, K. Rusche, X. Wang, Phosphoketolase and KDPG aldolase metabolisms modulate photosynthetic carbon yield in cyanobacteria. The Plant cell 10.1093/plcell/koae291 (2024).

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      The authors test specific but related hypotheses regarding anti-predator responses of wild marmoset groups to predator and human playback sounds triggered to play via a motion sensor on camera-trap devices. The differential responses they observe to human noises and natural predator sounds are interesting, but greater inferences are limited due to a lack of clarity in the methods and analyses.

      Strengths:

      The authors create an excellent experimental design using a customised ABR system for testing the behavioural responses of wild, social-living, tiny, arboreal primates: pygmy marmosets. Much of the work is described with great transparency, and figures and tables are helpful in facilitating this.

      Weaknesses:

      The current study requires improvement in three areas, in my opinion, to permit readers to better evaluate the validity and importance of these results.

      (1) Improve the framing of the paper:

      The current title and justification for this study appear to point to a lack of previous studies testing specific hypotheses (line 51/52: "ABRs have not been applied to hypothesis testing". I find this a rather strange argument to make, given that a quick read through of other ABR papers, cited by the authors (e.g., Kasper et al., 2025, Epperly et al., 2021), are testing predictions set by ecological theory in the cascading effects of predator-prey dynamics. To me, even if these are not explicitly stating "X hypothesis" in their paper, they still appear to be studies guided by implicit hypotheses. To say that previous work with ABR did not test hypotheses is presumptuous, in my opinion. The entire paper would be much better appreciated if the authors could reframe the study for its importance to arboreal mammal/ tropical ecology, anthropogenic effects, and so on. Similarly, the authors should avoid use of phrasing such as "this study is the first direct test of .... " (lines 59/60) and should emphasize the true significance of their work, beyond it being the 'first' of something.

      Similarly, on line 87, "demonstrating that the ABR system can be used to generate data for hypothesis testing" should be removed, as firstly sufficient sample size for any study depends on a number of study-specific parameters, and the authors do not actually demonstrate this in my opinion, given that many of their models end up suffering from singular fit. This is due to a lack of sample size, and also because they do not actually do any type of power analysis or something similar to demonstrate that they actually assessed sample size. So again, my suggestion is to reframe the paper to focus on the behavioural ecology and conservation-related impacts rather than this emphasis on methodology.

      We agree that the reviewer makes a valid point that while we were focusing on explicit hypothesis testing in our statements that the other papers mentioned are making predictions based off ecological theory. We will remove line 87 and make sure to limit these comments in the manuscript. We will reframe the introduction and discussion to better reflect this and change the verbiage throughout to focus on behavioural ecology, the impacts of anthropogenic noise and the conservation implications of the study as well as the novel arboreal aspect of the work. We will also update the title to better fit this framing of the study.

      (2) Methods:

      The authors generally do a great job providing sufficient detail on the ABR system and how each experiment was designed. Still, there is room for improvement, as I was confused a number of times. I also would recommend that the authors include a limitations section somewhere which considers the drawbacks of their study, in particular the lack of individual identity for behavioural responses of marmosets (especially given that they used focals, it seems), the groups being in close vicinity of one another/potentially related (?), the specific stimuli used, etc.

      We do touch on the limitation of not identifying individuals in the discussion (lines 252-258), but we will draw this into a specific section which will address this and the other limitations mentioned here.

      Points of confusion for me included what the control was for Experiment 1. Line 93 - 70 videos without playbacks are used (Table 1), but it is not clear how these videos were selected, and it is not a suitable control comparison for assessing the difference in behaviour associated with playbacks (playbacks with control sounds are). At most, these videos will give basal rates of behaviour (like vocalizations, etc.), but then it is not clear why these '70 videos' and how they were chosen to avoid bias. So, for example, in line 105 the authors write that focals were more likely to flee when hearing playback stimuli than in these "control" videos, but this is not convincing. If there was no fleeing after playback of control sounds (cicadas, macaws) - i.e., the true control in this experiment - then this should be the comparison that is emphasized.

      For one group, there were only 13 videos without a playback where a marmoset was present. So, we used these 13 videos for this group, and selected 13 videos at random from the other groups to match sample size. For these groups, we assigned each video without a playback but with a pygmy marmoset a sequential number, and then used a random number generator to select 13 videos for analysis. We refer to these videos as controls as they are negative controls, and the reviewer is correct – they do measure basal levels of behaviour. In contrast, the playback of control sounds is a procedural control (Bui et al., 2022). We will change how we refer to these controls in the manuscript to reflect the types of control they are. We use the comparison between negative controls and videos with playbacks in experiment 1 to assess the impact of the playback procedure itself, though the reviewer is correct that our conclusions would be better supported if we explicitly compared the intervention playbacks with the procedural control. We will add post-hoc tests to make this comparison explicit.

      Can the authors also clarify how they considered/assessed the sound playback level (normally done in Decibels) and if they did not normalize the sound level across the playback stimuli, why not, and what potential effect this could have on the results (i.e., something else to consider for the limitations section)?

      We edited the audios so that they were at similar volume levels. We will update the methods with this information and touch on this in the updated limitations section discussed above.

      Something else not discussed is the rate of exposure to predator and human noise for these wild monkeys. Are these rates within normal range/expectation for these monkeys?

      Although we do not have information about exposure rates to predators, we do mention levels of exposure of human noise in our methods section on lines 318-320 and we also touch on this in our discussion lines 206-212.

      We will update the methods section to be clearer that all groups are exposed to high levels of anthropogenic noise due to their proximity to the ecotourism lodge and community. We will also expand on this in the discussion as a limitation as having a broader array of groups with varying levels of exposure to humans would allow us to see the broader behavioural reactions to these playback stimuli.

      For the predators we mention in the methods section line 382 “All four species have been found in the study area (Barker and Papworth, 2024)” but we will further expand on this to provide information on the density of raptors in the area based on our previous study.

      Thinking here of the number of videos captured for each group presumably means exposure to a playback unless 'control' videos were videos where no playback sound was emitted (see question above re: control videos). There were a lot more unsuccessful videos than successful ones that the authors could use, so trying to understand the potential impacts of this (see question re: trial/video # below as well).

      The number of videos was the total number of times the camera trap triggered. These included videos triggered by another animal or foliage movement where there was no marmoset present, and includes both videos with and without a playback.

      We will make this clearer in our description of the results and will report the number of unsuccessful playbacks.

      (3) Analysis:

      A few things are unclear and need more explanation in the way the authors conducted their analyses, although they do well to detail all steps of their statistical methods, which was great.

      For assessing model fit, it is not clear what exactly was assessed with the 'performance package' line 467, as the authors do not go on to provide us with any results of the performance/fit. Instead, they tell us that the models did not fit well, with no parameter provided (e.g. lines 481-488). I'm familiar with overdispersion as a parameter that is reported for Poisson models (that does not seem to be provided here). Or by looking at changes in model estimates if one datapoint (and/or one group) is removed after another (with replacement, so keeps sample size static). On line 468/469, it says that model fit was assessed via conditional R2; could the authors provide a citation for this practice, and then provide the R2 parameter for the other models that were used/included in the end?

      We used various tests from the performance package (e.g. check_overdispersion) to test the fit of different distributional models (e.g. Poisson, negative binomial) to the same data. Although some models were not overdispersed and did not show evidence of zero-inflation, they did have singular fits, and/or a conditional R<sup>2</sup> of 1.0 suggesting overfitting. We therefore did not choose these models. We will clarify this and provide further details of our approach in the manuscript.

      The models for the behaviours not reported did not fit well using any distributional model. These behaviours had very low occurrence (0 seconds in most videos), and so there was very sparse data for generating estimates, and very low variation within / between groups and conditions. Therefore, these models generated the errors ‘Model nearly unidentifiable’ and warnings about singular boundaries. These suggest we would not be able to reliably generate model estimates, so we do not report the results. We will change the manuscript to make this reasoning more explicit.

      Also, please standardize how the GLMM results are presented. There should be the estimate, SE, Z or t, then p value. (line 99, 137).

      We will update the results reporting to be standardised as the reviewer has suggested.

      Given the high rates of exposure to playbacks, I think the authors should test trial # (or video #) for a potential habituation effect, with earlier captures more likely to draw stronger responses than later video captures for each group.

      We will include video number in a reanalysis of the data.

      Reviewer #2 (Public review):

      Summary:

      The article describes an interesting methodology to test hypotheses about the impact of anthropogenic noise on a small arboreal primate, the pygmy marmoset. The authors used a motion-triggered combination of camera traps and speakers to play back control sounds, avian predator calls, and anthropogenic noise to test the risk-disturbance hypothesis and the distracted prey hypothesis. In addition, the authors implemented a technique that is usually used for larger mammals and has not been used before for smaller arboreal animals. The authors are careful in their interpretation of the results and do not favor one hypothesis over the other. The authors also elaborate extensively in their discussion on how to improve this kind of data collection in the future.

      Strengths:

      This study provides a method for rapid data collection while minimizing observer impact. The sample size is comparatively large for a wild animal in a reserve, given the overall observation time. The article also benefits from a solid analysis of the data.

      Weaknesses:

      Though the authors tested two contrasting hypotheses, the discussion would benefit from more detail on the ecological relevance of the observed behaviors.

      We thank the reviewer for their comments and we will update the discussion to add more detail on the ecological relevance of the behaviours that we observed.

      References

      Bui, S., Madaro, A., Nilsson, J., Fjelldal, P.G., Iversen, M.H., Brinchmann, M.F., Venås, B., Schrøder, M.B. and Stien, L.H. 2022. Warm water treatment increased mortality risk in salmon. Veterinary and Animal Science, 17, 100265. https://doi.org/10.1016/j.vas.2022.100265

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Li et al. used genetically engineered murine intestinal organoids to investigate how the temporal order of oncogenic mutations influences cell state and tumourigenicity of colorectal epithelial cells. By sequentially introducing Apc and Trp53 loss-of-function mutations in alternate orders within a Kras^G12D background, the authors generated isogenic organoid lines for both in vitro and in vivo characterisation. Bulk RNA-seq reveals expected transcriptional changes with relatively modest differences between the two triple-mutant configurations (KAT vs KTA). The key finding emerges from transplantation assays: while KAT and KTA organoids show equivalent tumourigenic potential in immunodeficient mice, only KAT organoids form tumours in immunocompetent hosts (5/10 vs 0/10), suggesting that mutation order shapes susceptibility to immune-mediated clearance. The experiments are well-executed, and the conclusions are generally supported by the data.

      Strengths:

      The experimental system is well-designed for the question. By combining a Kras^G12D transgenic background with sequential CRISPR-mediated knockout of Apc and Trp53 in alternate orders, the authors generated truly isogenic organoid lines that differ only in mutational sequence. This is technically non-trivial and provides a clean platform for dissecting order effects, a question otherwise difficult to address experimentally.

      The authors performed comprehensive baseline characterisation of these organoids, including morphological and histological assessment, quantification of organoid-forming efficiency and proliferation, and bulk RNA-seq profiling. While these analyses revealed no major differences between KAT and KTA organoids, and the observed enhancement of epithelial stemness upon Apc loss and proliferative advantage conferred by Trp53 loss are largely expected, the systematic nature of this characterisation establishes a useful methodological template for future organoid-based studies.

      The authors further investigated the functional impact of mutational order using subcutaneous transplantation assays. By comparing tumour formation in immunodeficient versus immunocompetent hosts, the authors uncover a genuinely unexpected finding: KAT and KTA organoids behave equivalently in the absence of adaptive immunity, but diverge dramatically when immune pressure is applied (KAT: 5/10; KTA: 0/10). This observation is arguably the most compelling aspect of the study and opens an interesting line of inquiry.

      We greatly appreciate your comments on this study.

      Weaknesses:

      The authors acknowledge that initiating with Kras^G12D does not reflect the typical human sporadic CRC trajectory, where APC loss is usually the first event. While this design choice was pragmatic, it means the observed order effects are contextualised within an artificial starting point. It remains unclear whether the Apc/Trp53 order would matter in a Kras-wild-type background, or whether the Kras-driven cellular state is a prerequisite for these phenotypes to emerge.

      We agree with the reviewer that initiating tumorigenesis with Kras<sup>G12D</sup> does not fully recapitulate the most common trajectory of sporadic human CRC, where APC loss typically occurs first. We had noted this point in the original Discussion and further clarified it more explicitly in the Introduction part of the revised manuscript as shown in Line 97–103.

      Our experimental design was intended to establish a controlled and genetically tractable system to interrogate the principle of mutation order effects. In this context, Kras<sup>G12D</sup> activation provides a stable oncogenic baseline that facilitates sequential genome engineering and comparison of isogenic lines.

      Although APC loss is frequently the initiation event, a recent study has suggested that Kras<sup>G12D</sup> priming can reshape the selective landscape for subsequent driver events, including Apc alterations (PMID: 41339549). Consistent with this notion, our data indicate that Kras<sup>G12D</sup> activation induces a permissive oncogenic cellular state that may influence the phenotypic consequences of later mutations. We therefore speculate that the Kras<sup>G12D</sup>-primed context may contribute to the observed order-dependent effects.

      We agree that testing Apc Trp53 order in a Kras-wild-type background would be an important future direction, and we have pointed this out explicitly in the revised Discussion as shown in Line 549–554.

      Subcutaneous implantation provides a tractable readout of tumourigenicity, but the cutaneous immune microenvironment differs substantially from that of the intestinal mucosa. Given that the central claim concerns immune-mediated selection, orthotopic transplantation would more directly test whether the observed order effects hold in a physiologically relevant context.

      In the present study, we employed subcutaneous transplantation as a widely used platform to assess tumorigenic potential under controlled immune conditions. This approach offers high reproducibility, straightforward tumour monitoring, and has been broadly applied in organoid-based cancer studies in both immunodeficient (PMID: 23273993, 23776211, 32209571, 33055221) and immunocompetent (PMID: 32209571, 33055221, 41672595) settings.

      Importantly, our primary goal was to determine whether mutation order influences susceptibility to immune-mediated clearance, rather than to model the full complexity of the intestinal niche. The clear divergence between KAT and KTA specifically in immunocompetent hosts supports the existence of intrinsic mutation order-dependent immune vulnerability.

      Nevertheless, we fully agree with the reviewer that orthotopic transplantation would provide a more physiologically relevant immune microenvironment and represents also an important direction for future investigation. We have explicitly discussed this limitation and highlight orthotopic validation as an important future direction in the revised Discussion as shown in Line 563–571.

      The ssGSEA comparison involves only 14 ATK tumours, and the key comparisons (Figure 6E) yield borderline significance (p=0.052). More fundamentally, since mutation order cannot be inferred from the clinical samples, the authors are correlating organoid-derived IFN signatures with tumour immunophenotypes without direct evidence that these patients' tumours followed a KAT-like trajectory. The reasoning becomes circular: KAT organoids define the signature used to identify KAT-like clinical tumours.

      We thank the reviewer for raising this important point. We would like to clarify that our intention was not to infer the actual mutation order in clinical samples, which indeed cannot be reliably reconstructed from bulk tumour RNA-seq data.

      Instead, our goal was to determine whether the transcriptional programs distinguishing KAT and KTA organoids could be observed in human CRC cohorts. In this context, the organoid-derived IFN-related signature was used as a molecular reference to assess potential clinical correlation, rather than to classify tumours by evolutionary trajectory.

      We agree that the statistical significance in Figure 6E is modest (p = 0.052), and we have revised the text (Line 478–480) to present this analysis more cautiously as a suggestive trend rather than definitive evidence. We also clarified this limitation explicitly in the revised manuscript (Line 537–542) to avoid overinterpretation.

      Furthermore, the most striking finding of the study, that KTA organoids fail to form tumours in immunocompetent hosts while KAT organoids can, lacks a mechanistic follow-up. The transcriptomic differences between KAT and KTA are modest when cultured as monocultures, yet their in vivo fates diverge dramatically. The authors do not address why these subtle intrinsic differences translate into such divergent immune susceptibility, nor do they characterise the immune response adequately (beyond limited CD4/CD8 IHC at tumour peripheries).

      We thank the reviewer for this important point. We agree that the mechanistic basis underlying the differential immune susceptibility between KAT and KTA remains incompletely resolved.

      A practical limitation of the current study is that KTA grafts failed to establish tumours in immunocompetent hosts, which precluded downstream histological and immune profiling of established lesions. As a result, our in vivo immune characterization of KTA grafts is nearly impossible.

      Nevertheless, our transcriptomic analyses indicate that KAT and KTA organoids differ in interferon-response and immune-related programs prior to transplantation, and those differentially expressed genes were consistently preserved in tumour cells derived from immunodeficient hosts. These results suggest the presence of intrinsic tumour-cell-autonomous differences may influence immune recognition or clearance.

      We have expanded the Discussion to outline several non-mutually exclusive mechanisms that could account for this phenotype, including altered interferon responsiveness, differential antigen presentation capacity, and changes in tumour cell-intrinsic immune escape programs (Line 527–533). These hypotheses are consistent with the transcriptional differences observed prior to transplantation and provide a framework for future mechanistic investigation. We agree that deeper immune profiling (e.g., immune infiltrate composition, antigen presentation status, and functional immune assays) will be important to fully elucidate the mechanism and represents a key direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This study addresses an important and timely question in colorectal cancer biology by systematically examining the effects of the common driver mutations APC, KRAS G12D, and TP53 in murine colorectal organoids, with particular emphasis on how the order of APC and TP53 acquisition influences tumor phenotype. These mutations are well known to be frequent, truncal, and often co-occurring in colorectal cancer. While it is increasingly appreciated that mutational order can shape tumor behavior, studies directly comparing the phenotypic consequences of alternative APC-TP53 mutation orders remain rare. This work, therefore, addresses a relevant and timely question.

      Strengths:

      A major strength of the study is its focus on previously unexplored biology, combined with the generation of multiple isogenic murine organoid models with controlled mutational sequences. The authors employ careful and robust quality control of the CRISPR-mediated alterations, and the inclusion of both in vitro and in vivo experiments strengthens the relevance of the work.

      We greatly appreciate your comments on this study.

      Weaknesses:

      There are, however, several limitations that should be considered when interpreting the findings. First, KRAS G12D activation is used as the initiating alteration, whereas APC loss is generally believed to be the initiating event in most human colorectal cancers.

      We sincerely thank the reviewer for their insightful comments regarding the initiation of tumorigenesis with a Kras mutation rather than the more canonical Apc loss, which was also raised by the reviewer #1. We fully agree that the Apc-first represents the most prevalent sequence in human colorectal cancer (CRC), We have more clearly explained the rationale for our experimental design in the revised Introduction part as outlined in our response to reviewer #1.

      Second, the analysis is restricted to comparing only two mutation orders (KAT versus KTA), which limits the breadth of conclusions that can be drawn about mutation ordering more generally.

      We thank the reviewer for this critical concern, which we agree is essential for strengthening the robustness and generality of our findings. However, as a proof-of-concept study of Apc and Trp53 loss, two major oncogenic events in CRC, serves as a biologically meaningful starting point for dissecting order-dependent effects. Although it is of great significance to compare all six possible mutation orders of these three driver genes, generating and thoroughly characterizing all genotypes (with identical replicates) represents a substantial undertaking beyond the scope of this initial study.

      Finally, key RNA-sequencing and in vivo experiments rely on a single isogenic line, which substantially constrains interpretability.

      The aim of the study was to systematically investigate how mutation accumulation and order influence colorectal cancer initiation. While the data suggest that the relative timing of APC and TP53 loss may be particularly important for tumor initiation, the absence of biological replication makes it difficult to draw robust conclusions. Engraftment efficiency and tumor behavior can be influenced by many factors for a single clone, including additional passenger mutations acquired during culturing, as well as epigenetic differences that are independent of the engineered mutations.

      We thank the reviewer for this concern. We apologize that we have not made a clear presentation of our data source. Indeed, for all major in vitro and in vivo assays of double and triple mutants, we analyzed at least two independently derived clones per genotype. These independent clones harbour distinct mutations in target genes and were treated as biological replicates throughout the study.

      To improve clarity and transparency, we have revised the relevant figure legends and further provided a Table S5 to explicitly indicate the clonal origin of each data point throughout the study.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      (1) It would be useful to have a timeline of the study, much like the one provided for the water maze protocol. On that timeline, please include the sample sizes examined, ages at exposure, and other pertinent procedures, indicating which animals remained alive for testing, etc.

      We agree with the reviewer that a timeline would be helpful for clarifying the PAE exposure paradigm and the subsequent sample collection and processing. Sample sizes vary across the different analyses; therefore, we have indicated the sample size for each experiment in the corresponding figure. To further improve clarity, we will add Author response image 1, which will include a schematic of the overall experimental timeline, including the PAE regimen, collection time points, and sample processing, as shown below.

      Author response image 1.

      (2) What were the attrition rates for each study group?

      We thank the reviewer for raising this important point. There was no attrition of animals within the experimental groups; the number of animals included at the beginning and end of the study remained the same. However, we did observe a reduction in litter size following prenatal alcohol exposure (PAE). In our 3xTg-AD colony, litters typically consisted of approximately 8 pups under control conditions, whereas PAE litters occasionally contained 4–6 pups. Thus, the reduction in animal numbers reflects decreased litter size associated with PAE rather than attrition during the study. We would also like to clarify that the primary scope of this study was not to provide a terminal/end-point analysis of disease progression, but rather to investigate the emergence of Alzheimer’s disease (AD)-related phenotypes during early adulthood following PAE. Accordingly, our longitudinal experimental design focused on identifying the earliest molecular, synaptic, behavioral, and neuropathological alterations that emerge during this period. This approach allowed us to examine whether PAE accelerates or precipitates the onset of AD-related symptomatology in the 3xTg-AD model, rather than following the animals through advanced disease stages. We will clarify this rationale in the revised manuscript.

      (3) What are the human age equivalents of the maternal mice?

      We appreciate the reviewer’s question regarding the age of the maternal mice. The dams used in our study were young adult females (2 to 3 months of age) at the time of breeding. Because chronological age does not translate linearly between mice and humans, particularly during development and reproductive maturation, we have avoided assigning a precise human-age equivalent. Based on established comparative developmental frameworks and calculations, these animals represent a 20 years old young-adult in the reproductive stage, rather than an advanced maternal-age condition [1]. We will clarify this point in the revised manuscript.

      (4) It would be helpful to see a graph of the BECs of each animal relative to the doses given. That would clarify how the alcohol exposure amount and timing are the same and where they are different for all exposed mice, given that the alcohol levels were somewhat different by group, as noted in the Methods. Were these BEC differences at all related to group differences in outcome measures or memory performance?

      We appreciate the reviewer’s suggestion to provide a more detailed representation of the BEC data. We agree that displaying the BECs for individual animals would provide additional clarity regarding the consistency of alcohol exposure across groups. We have therefore included the individual BEC values in the new Figure 1. We observed some variability in BECs between the 3xTg-AD and B6129 groups, as noted in the Methods. Importantly, however, the average alcohol consumption was comparable between the two genotypes, indicating that the difference in BECs was not due to differences in the amount of alcohol consumed. All dams in the PAE groups reached BECs above 0.08 g/dL, the commonly used legal blood alcohol concentration limit in the United States, supporting the use of our paradigm as a binge-like alcohol exposure model. We further examined whether the variability in BECs was associated with the differences observed in outcome measures, including memory performance. We did not find evidence that the differences in BECs accounted for the group differences in behavioral or molecular outcomes. Thus, although some intergroup variability in BECs was present, the overall alcohol exposure was comparable, and the observed phenotypic differences were not attributable to differences in alcohol consumption.

      Reviewer #2 (Public review):

      (1) Some figures lack prenatal alcohol treatment in the 3xTg-AD mice.

      We appreciate the reviewer’s observation and agree that the rationale for the different experimental groups across the figures should be clarified. The primary focus of this study is to characterize the effects of prenatal alcohol exposure (PAE) in wild-type B6129 mice, with the 3xTg-AD mice serving primarily as a disease-model reference to determine whether the effects observed following PAE in wild-type animals overlap with or resemble features of AD pathology. Accordingly, the initial figures focus on the effects of PAE in B6129 mice and include the non-exposed 3xTg-AD group as a reference for the AD phenotype. The last two figures specifically address the effects of PAE in the 3xTgAD model, with the 3xTg-AD mice becoming the experimental subject of interest rather than serving solely as a disease reference. For this reason, the PAE-3xTg-AD group is not included in the earlier figures, whereas it is included in the final two figures where the effect of PAE on the AD model is directly evaluated. We will clarify this experimental rationale in the revised manuscript and figure legends.

      (2) Some overstatements should be tempered. For instance, one cannot conclude that the changes in CTFs are driving the changes in learning and memory (as suggested in the last line of the abstract) without a direct intervention testing this. For instance, though PAE caused a more robust learning deficit at 6 mo in WT, the impact on CTFs was less than it was at 3 mo. PAE did not significantly change CTFs or learning/memory in 3xTg-AD mice at 4 months, suggesting the genotype effect takes over at this point. The text should be adjusted to reflect this.

      We appreciate the reviewer’s careful consideration of this point. We agree that the relationship between APP CTF accumulation and learning and memory deficits should not be interpreted as causal in the absence of a direct intervention experiment. We were careful in choosing the wording throughout the manuscript to describe these findings as associated changes rather than evidence of a causal interaction. Our data demonstrate the presence of APP CTF accumulation and learning and memory deficits following PAE, but they do not establish that CTF accumulation directly drives the behavioral phenotype. We therefore will temper the language in the Abstract and throughout the manuscript to avoid overstatement. We also acknowledge that the relationship between these phenotypes is not necessarily linear across age: although PAE produced a more pronounced learning deficit at 6 months in B6129 mice, the magnitude of APP CTF accumulation was greater at the earlier time point. Importantly, we consider the possibility that the greater APP CTF accumulation observed at earlier ages may represent an early molecular insult whose functional consequences become evident later in life. In this context, the temporal dissociation between the molecular and behavioral phenotypes could be consistent with a “two-hit” model, in which an early-life insult induced by PAE creates or primes a pathological vulnerability that subsequently manifests as cognitive dysfunction with ageing [2]. We recognize, however, that this interpretation remains a hypothesis and would require longitudinal mechanistic studies to establish. This temporal relationship may also contribute to the broader concept of early-life origins of AD/ADRD, suggesting that prenatal environmental exposures may initiate molecular alterations during neurodevelopment that remain detectable or predispose the brain to later-life dysfunction. Similarly, the absence of significant changes in APP CTFs or learning and memory in 4-month-old 3xTg-AD mice suggests that the effects of the AD genotype may become dominant at this stage. Consistent with this interpretation, we state in the Discussion that future studies are needed to identify and experimentally test the direct molecular pathways affected by PAE that ultimately contribute to learning and memory impairment. We will revise the text accordingly to make this distinction clear while highlighting the potential significance of an early molecular insult preceding the later emergence of behavioral phenotypes.

      Reviewer #3 (Public review):

      (1) It is unclear as to whether there are sex differences, particularly in the adult cohort.

      We appreciate the reviewer’s comment regarding potential sex differences. We agree that considering sex as a biological variable is important, particularly for the adult cohorts. To address this point, we will include identifying marks for male and female animals separately in our plots where sample size permits. This will allow the reader to better evaluate potential sex-dependent effects of PAE and to determine whether the observed phenotypes are consistent across sexes. We will also clarify this approach in the revised manuscript.

      (2) More clarity is needed on sample size per cohort and whether mice that were used for anatomy and biochemical analyses were previously used for behavior. Including a table and noting any overlap would be useful.

      We appreciate the reviewer’s suggestion and agree that greater clarity regarding the sample sizes and use of animals across analyses is important. The sample size for each cohort and experimental group is indicated in the corresponding figures, and we will make this information more explicit in each figure legend to facilitate interpretation. Animals that underwent behavioral testing were subsequently used for biochemical analyses, allowing us to examine molecular changes in the same animals in which behavioral phenotypes were characterized. In contrast, for the neonatal cohort, we performed both anatomical and biochemical analyses, to assess the distribution and extent of APP CTF accumulation across the brain during this early developmental period. This approach was selected to provide a broader assessment of the spatial distribution of APP CTF accumulation at birth. We will clarify these experimental details in the revised Methods and figure legends.

      (3) In many instances, two-way ANOVAs with treatment (PAE vs vehicle) and genotype as factors will be useful to report (e.g., Figure 1).

      We appreciate the reviewer’s suggestion regarding the use of two-way ANOVA with treatment and genotype as factors. However, we respectfully disagree that this approach is appropriate for all of the analyses presented in this manuscript. Our experimental design and the specific biological questions addressed in each experiment were not uniform across cohorts. In particular, the primary objective of the study was to characterize the effects of PAE in B6129 mice, with the 3xTg-AD mice serving primarily as a disease-model reference, while the effects of PAE in the 3xTg-AD model were specifically examined in the later experiments. Therefore, combining genotype and treatment as factors across all datasets would not always reflect the experimental questions or the structure of the cohorts. In addition, some experiments did not include all four groups, making a two-way ANOVA inappropriate for those analyses. Nevertheless, we agree that a two-way ANOVA may be informative for experiments in which both genotype and treatment are fully represented and the experimental design supports this analysis. We will therefore consider and apply two-way ANOVA, where appropriate, to those datasets, including evaluation of the main effects of genotype and treatment and their interaction. We will clarify the statistical approach and its rationale in the revised Methods and figure legends.

      (4) In Figure 5 and line 253, it is stated that older mice have more severe deficits, but there are no direct statistical comparisons with younger AD mice.

      We appreciate the reviewer’s observation. We agree that, in the absence of a direct statistical comparison between age groups, the statement that older mice have “more severe deficits” may be too strong. Our intention was to describe the apparent progression of the phenotype across age rather than to imply that we had statistically demonstrated an age-dependent increase in severity. We have therefore revised the text in Figure 5 and at line 253 to use more cautious language, describing the greater magnitude of the observed deficits in older mice without implying a direct statistical comparison between age groups. We agree that a formal conclusion regarding age-dependent progression would require a statistical analysis, which we will include in the revised manuscript.

      (5) Lines 270-271 refer to mice as "presymptomatic", but these mice do have behavioral symptoms. Do the authors mean no neuropathology yet? Any data showing lack of robust neuropathology would be useful.

      We appreciate the reviewer’s careful observation. We agree that the term “presymptomatic” was not sufficiently precise, particularly because the mice already exhibit measurable behavioral alterations at this age. Our intention was not to suggest that these animals were free of phenotypic abnormalities, but rather that they were at an early stage of disease progression, before the emergence of robust neuropathological features. We have therefore revised the terminology to avoid referring to these mice as “presymptomatic.” Instead, we describe them as being in an early stage of disease development, characterized by emerging behavioral and molecular alterations but without the extensive cardinal neuropathology typically associated with later stages of the 3xTg-AD phenotype. We agree that the distinction between behavioral symptoms and neuropathological progression is important. In this study, our focus was on the emergence of early AD-related phenotypes during young adulthood, rather than on establishing the absence of neuropathology. We have therefore avoided making a definitive claim regarding the lack of neuropathology and have revised the text to more accurately reflect the scope of our data.

      Additional References

      (1) Dutta, S. & Sengupta, P. Men and mice: Relating their ages. Life Sci. 152, 244–248 (2016).

      (2) Gunn, J. S. et al. Exploring the ‘Multiple-Hit Hypothesis’ of Neurodegenerative Disease: Bacterial Infection Comes Up to Bat’. Frontiers in Cellular and Infection Microbiology | www.frontiersin.org 1, 138 (2019).

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank you for the time you took to review our work and for your feedback! We have performed the following additional analyses:

      (1) We analyzed the optomotor mismatch response as a function of time spent in the optogenetic closed-loop session.

      (2) We show the correlation between optomotor mismatch response and visuomotor mismatch response for functionally identified PE neurons.

      The main changes to the manuscript are:

      (3) Clarification of our terminology (moved and refined definition of the teaching signals).

      (4) A new figure panel summarizing the functional influence patterns we identified.

      (5) Clarification of our argumentation for the JEPA-inspired proposal.

      All comments are addressed individually in the following.

      Public Reviews:

      Reviewer #1 (Public review):

      Vasilevskaya and Keller test different models of cortical function through the lens of predictive processing, a powerful framework for the brain to learn and predict the statistics of the world via generative internal models. The authors use a clever combination of behavioral perturbations in closedloop and open-loop visuomotor virtual reality assays, a paradigm the Keller lab pioneered and used effectively in the past decade, in conjunction with two-photon imaging of neuronal calcium responses and targeted optogenetic perturbations of activity. They specifically put to test proposed hierarchical vs. non-hierarchical circuit implementations of predictive processing by analyzing the logic of inter-lamina interactions (superficial vs. deep; L2/3 vs. L5/6).

      The authors conclude that both versions of predictive processing architectures they analyze are likely invalid, and instead formulate an alternative novel model of cortical function based on a recently developed machine learning algorithm for self-supervised learning (joint embeddings of predictive architectures, JEPA) and its further refinements. JEPA borrows elements from predictive processing, engaging two encoder networks and training the output of one network to predict the output of the other. In their new model of cortical computations, prediction error neurons in L2/3 compare the deep layers (L5/6) activity, which is taken as a teaching signal, to a local, L2/3 prediction of this latent representation.

      Specifically, the authors build on their previous work and reports from other groups that different sets of L2/3 neurons compute positive prediction errors (fire when sensory stimuli appear unexpectedly with respect to the movements of the animal; e.g., grating onsets in the absence of locomotion) and respectively negative prediction errors (fire when sensory stimuli are absent, while the brain expected them to be present; e.g. mice locomote but visual flow is suddenly halted - visuomotor mismatches). These L2/3 positive and negative prediction error neurons exchange messages with neurons in the deeper cortical layers that, the authors propose, build an internal representation (R) of the sensory stimuli given the animals' movements.

      In the hierarchical model, internal representation neurons (R) are supposed to act as a teaching signal for both types of prediction error neurons; the output of the positive prediction error neurons is assumed to suppress activity of R such that the error between the teaching signal and the prediction is minimized; similarly, in the non-hierarchical version, R serves as a prediction for the prediction error neurons, and in turn it receives excitatory drive from the positive prediction error neurons and negative input from the negative prediction error neurons.

      The authors find that the functional impact of L5 neurons on L2/3 neurons is not compatible with the non-hierarchical architecture they and other groups proposed, but rather in accordance with the hierarchical model. At the same time, the functional impact of L2/3 neurons (positive vs. negative prediction error neurons) on L5 neurons (internal representation) appears not compatible with the hierarchical model, but rather in accordance with the non-hierarchical implementation.

      They further hypothesize that L2/3 prediction error neurons don't use sensory input, but rather the L5 activity as a teaching signal, and test it using perturbations (halts) of optogenetic stimulation of L5 neurons coupled with locomotion (Figure 7).

      All in all, the question is topical, and the new model addresses a decades-long quest to develop a unifying model of cortical function. The findings reported here transform our understanding of cortical computations, opening new, exciting avenues for future investigation. The experimental design and execution are rigorous; the arguments are clearly laid out (in spite of ample potential for confusion given the numerous loops and sign flips). These include a discussion of why the non-hierarchical model proposed by the same group does not hold, as well as potential caveats in interpreting the results and novel testable proposed experiments emerging from the JEPA-like model.

      I have several questions about the interpretations of some of the claims and suggestions for potential additional experiments and analyses.

      We thank the reviewer for their comments. We address them below.

      (1) Some of the pieces of the puzzle remain to be identified and demonstrated: the existence of internal representation neurons in L2/3 and ascertaining that the L5/6 neurons analyzed function indeed as internal representation neurons. The authors find that stimulation of L2/3 positive prediction error neurons enhances activity of L5 neurons...If L5 neurons hold a latent representation that serves as a teaching signal for L2/3 neurons (as the authors posit), wouldn't one expect that the input they receive from the positive prediction neurons be suppressive, such that the error is further minimized?

      Not necessarily - this depends on the model one has for how cortex works. In the hierarchical predictive processing model, PE+ neurons are expected to suppress the local internal representation neurons in L5. Our data, however, are not consistent with this model. This is one of the key arguments we build the idea on that JEPA is a better model for cortex than hierarchical predictive processing. In JEPA we would not expect the L2/3 prediction error to update L5 directly, but instead update the prediction of L5 activity (the source of this signal remains to be identified; see our speculations on the origin of prediction signals in the comments to question 3), and drive plasticity in the local L5 encoder network. That something acts like a teaching signal for the L2/3 comparator does not, by itself, determine how the outcome of the comparison influences the source of the teaching signal.

      (2) Do the authors envision any specific differences between the representations of the two encoder networks posited to exist in L2/3 and L5 in the JEPA-like implementation? Are they synchronous/offset in their temporal representations, or any other features?

      Given theoretical work (Mohammadi et al., 2025), one would expect to find differences in learning rates between the predictor and the encoder networks. Assuming the predictor network is also implemented in L2/3, we would expect to see faster learning rates in L2/3 compared to L5. Implementations inspired by related theoretical works also predict differences in learning rate between the two encoder networks (Grill et al., 2020), again with higher learning rates expected in L2/3 compared to L5. Beyond that, however, we are not aware of any experimentally observable differences one might expect to find. We are hoping the computational community will remedy this soon.

      (3) Where is the prediction coming from onto L2/3 neurons? Is it emerging locally in L2/3 from the putative internal representation neurons, or is it long-range - as work from the authors previously proposed? Or a mix of both?

      We expect the predictions to come from long-range inputs. In classical (hierarchical) JEPA one would expect these to be the lateral communication within L2/3. In a non-hierarchical implementation that is capable of operating on arbitrary graphs (as would be necessary for it to work in cortex), we suspect that the L2/3 network will use both long-range L2/3 and long-range L5 input. In JEPA terminology: the encoder A network uses information from non-local sources of both networks to predict local activity of encoder B network (assuming that information is statistically useful in predicting that activity).

      (4) What is the role of the indiscriminate L4 input that appears to enhance activity of both positive and negative prediction error neurons in L2/3?

      The short answer is, we don’t know. We would have indeed expected to find some asymmetry of influence. There are a few options: A) We might be failing to activate a specific interneuron that mediates feedforward inhibition on this pathway by the artificial stimulation of Scnn1a neurons. B) We know that Scnn1a neurons are only a subset of L4 neurons – there are other populations of L4 neurons that might exhibit the opposing influence. C) The Scnn1a population is more than 1 layer of a JEPA away from the comparator and not part of the predictor that forms the actual representation compared against L5 (e.g. L4 could provide an input to L2/3 internal representation neurons, and thus only indirectly influence L2/3 PE neurons). D) The JEPA analogy is wrong.

      (5) Does Figure 7D change in a meaningful manner if the authors plot the correlation between optomotor mismatch response and visuomotor mismatch response specifically for the negative prediction error neurons in L2/3 (Adamts-2) rather than for all L2/3 cells sampled?

      We might be misunderstanding. If the reviewer means genetically identified negative prediction error neurons (Adamts2), we do not have the data to address this question, as we did not perform any recordings of molecularly defined Adamts2 population in L2/3. If the reviewer means functionally identified negative prediction neurons, this would be the two rightmost data points in Figure 7D (the x-axis in this panel is the visuomotor mismatch response strength we use to functionally identify PE- neurons). These two data points on the right-hand side of the plot correspond exactly to what we classify as negative prediction error neurons throughout the rest of the manuscript (15% of the most responsive neurons to visuomotor mismatch).

      Does the reviewer mean, is there also a positive correlation between optomotor mismatch response and visuomotor mismatch response when looking only at neurons that we identify as PE-? If so, the answer is yes (Author response image 1). Interestingly, while there is a positive correlation, there is also nonuniformity in response patterns. We think that this is expected, given that our visuomotor coupling paradigm captures only a very small subspace of stimuli that prediction error neurons are tuned to. Bulk stimulation of L5, in contrast, might work better to separate all PE neurons. In other words, we speculate that L5 stimulation is a better predictor of the functional role of an L2/3 neuron.

      Author response image 1.

      Optomotor mismatch response as a function of visuomotor mismatch response for neurons that are functionally classified as PE. Red line shows a linear fit estimated with a bootstrap approach.

      (6) Do the optomotor mismatch responses in L2/3 neurons depend on how long the closed-loop coupling of optogenetic stimulation of Tlx3 L5 neurons and locomotion speed has been in place for?

      No, not that we can measure. We performed an analysis in which we split the optogenetic closed-loop session into two equal parts, early and late. We then quantified the average optomotor mismatch response independently for early and late parts of the session (Author response image 2A). Based on this quantification, we find no evidence of a change as a function of time in the closed-loop session. We also performed a sliding window analysis on a shorter timescale. While it did look like the responses may be smaller in the first few minutes, we do not have sufficient data to address this. None of the differences in response size were significant (Author response image 2B).

      Author response image 2.

      Optomotor mismatch response as a function of experience with artificial closed-loop coupling. (A) Mean L2/3 population response to optomotor mismatch in the first half (Early MM) and in the second half (Late MM) of the optogenetic closed-loop session. (B) Mean L2/3 population response to optomotor mismatch as a function of time in the optogenetic closed-loop session. Error bars indicate SEM. Differences between the mean values were estimated by hierarchical bootstrap and are not significant.

      Reviewer #2 (Public review):

      This manuscript reveals the functional connectivity of two different classes of cortical neurons that respond in opposite ways to mismatches between sensory and top-down inputs. These data are very valuable because different theories of information processing in the cortex make different predictions on the patterns of connectivity of these neurons. Therefore, these data strongly constrain possible theories of cortical processing.

      We thank the reviewer for their comments. We address them below.

      General comments:

      (1) The methods of statistical testing are insufficiently described. I did not understand the description in lines 1105-1119. The authors should provide sufficient details so the reader can reproduce their analyses. For example, it may be helpful to provide specific details of the testing procedure for one of the comparisons (e.g. the first comparison in Table S1).

      We assume the reviewer is not familiar with hierarchical bootstrapping in general. If so, the explanation below would summarize the procedure. This is the procedure with particular emphasis on its application to neuroscience data is described in the paper we reference in that part of the methods (Saravanan et al., 2020). Given that the analysis has become relatively standard (and is described in the reference provided), we think it might be an overkill to add the full explanation below to the manuscript. In addition to the general procedure, there are only 2 pieces of information relevant to fully reconstructing the analysis:

      (1) What are the “levels” (mice, recording sites, neurons, trials)?

      (2) What is the number of bootstrap samples used.

      Thus, we think all information is already provided in the methods. We now also explicitly added the levels when describing the nested structure of the data in the manuscript to increase clarity.

      Hierarchical bootstrap analysis:

      Our data are naturally nested: multiple neurons are recorded within a single mouse, and multiple mice are tested within an experimental group. Standard bootstrapping (sampling with replacement from the entire pool of neurons) fails because it assumes all observations are independent. In reality, neurons from the same mouse are more similar to each other than to neurons from a different mouse. The hierarchical bootstrap (or multi-level bootstrap) preserves this nested structure, ensuring your confidence intervals are not artificially narrow due to pseudoreplication.

      To illustrate the problem, assume you have 100 neurons from Mouse A and 10 neurons from Mouse B, a simple bootstrap will be heavily biased toward Mouse A. Furthermore, the simple bootstrap ignores the fact that the true variance in your population comes from two sources:

      (1) Between-mouse variance (differences in surgery, genetics, or behavior).

      (2) Within-mouse variance (differences in tuning or activity between individual cells).

      Hierarchical bootstrap addresses this problem, and is implemented as follows: To estimate the mean response while accounting for different sample sizes per mouse, a two-level resampling scheme is used

      (1) Resample the higher level (Mice)

      First, you account for the variability between animals.

      Suppose you have N mice in total.

      Randomly draw N mice with replacement from your original pool.

      Note: Because this is with replacement, a single mouse’s data might be included multiple times in one bootstrap iteration, while another mouse might be left out entirely.

      (2) Resample the lower level (Neurons)

      For each mouse selected in Step 1, you must now account for the variability within that specific animal.

      Look at the number of neurons actually recorded from that mouse (let’s call it k<sub>i</sub>).

      Randomly draw k<sub>i</sub> neurons with replacement from that mouse’s specific pool of recorded cells.

      This step is crucial: you always resample the same number of neurons that were originally recorded for that specific mouse. This maintains the "weight" or "influence" that animal had in the original dataset.

      (3) Calculate the resampled statistic

      Calculate the mean of all neurons collected in this “bootstrap sample”.

      (4) Iterate

      Repeat Steps 1–3 many times (typically B = 1,000 or 10,000 iterations).

      The distribution of these B bootstrap means represents your sampling distribution and is used to calculate confidence intervals and p-values. 

      (2) The authors should clarify how the problem of multiple comparisons was addressed for comparisons performed in multiple moments of time, where significance is indicated by a black bar (e.g. in Figure 2F).

      There is no family-wise error correction in cases of comparing response time courses implemented in our analysis. If the reviewer has a good suggestion for how to implement family wise error correction, we would be happy to implement it. We are not aware of anything that is better than what we currently do (the time bin-wise comparison). To briefly explain the problem: In most neuroscience papers, response curves are compared by choosing a time window (e.g. 0.5 to 1 s following a trigger) and calculating mean values of the curves in these windows. This hides the problem of multiple comparisons that arises from the fact that the experimenter is free to choose a response window used for analysis. Note, this also creates a strong incentive to “optimize” choice of an analysis window – a part of the analysis that a reader is typically completely blind to. We could of course also choose an analysis window that sounds reasonable and yields significant differences for our analyses. However, to provide a more unbiased image of the data, we have come to do bin-wise comparisons with a fixed p-value (typically 0.05). This is used in all of our papers at the moment. Given that samples from neighboring timepoints are correlated via a combination of actual responses and a subset of noise sources, the samples are not independent. We now also implemented a correction for spurious positive values by requiring at least 2 neighboring bins to have a p-value below 0.05 to be shown. If the samples were independent, this would mean a false positive rate of 0.0025. Given that they are not, this is a lower bound only. Additionally, any family-wise error correction would be a function of the number of time bins we show in the plot. This would mean that our choice of the time window shown in a plot (-1s to +4s, or +5s, etc.) would change the p-value we consider significant. Thus, there is no explicit family-wise error correction, and we have come to the conclusion that the bin-wise comparison with a fixed p-value is the most unbiased representation of the data we can provide. The alternative would be to additionally plot z-scores or the p-values as a function of time, but in our experience these types of plots are even harder to read for readers not used to it.

      (3) It would be helpful to add a figure in the Discussion summarising the functional connectivity suggested by all experiments.

      We now added a panel that summarizes the functional connectivity observed in our experiments to Figure S10. 

      (4) Throughout the manuscript, the authors use the term "teaching signals", but I am unclear what they mean by it: after reading the definition in lines 45-46, I thought that they corresponded to values (as they are compared to sensory signals). Later (428-430), the text suggests that they correspond to error neurons. But then lines 605-607 say it is not an error signal. The authors should define teaching signals very precisely or remove this term.

      The formal definition of the teaching signal was in footnote 1 of the manuscript. We assume the reviewer may have missed this. We now moved this definition into the main text to increase clarity.

      We use the term teaching signal to mean exactly this definition throughout the manuscript. We suspect, a second source of confusion may come from ambiguity in regards to anatomical vs. functional definitions. We have attempted to emphasize that the definition of teaching signal is a functional one, not an anatomical one (as is the case for ‘prediction’ – predictions are functionally defined, not anatomically – hence it makes sense to ask questions of the form “what are the potential sources of predictions” etc.). A teaching signal is a signal that is compared against a prediction (the ‘ground truth’ the prediction is compared and trained against). In different circuit implementations of predictive processing, different inputs function as predictions and teaching signals. In the hierarchical implementation, activity of the internal representation neuron at the lower level serves as a teaching signal for a prediction signal that is formed by the internal representation neuron from the higher level. In the non-hierarchical implementation, external inputs from the thalamus or other cortical areas serve as teaching signals for the respective prediction signals that are formed by internal representation neurons. This terminology is most intuitive when thinking from the perspective of a prediction error neuron, since a prediction error neuron computes the difference between two inputs signals – one of which functions as a teaching input, and the other one as a respective prediction. Hence, lines 428-430 specify that layer 5 input onto prediction error neurons of layer 2/3 serves as a teaching input.

      Reviewer #2 (Public review):

      Vasilevskaya and Keller set out to experimentally distinguish between two variants of predictive processing: a hierarchical and a non-hierarchical variant. The hierarchical variant assumes a hierarchical organization in which internal representation neurons (believed to be a subset of layer 5 excitatory neurons) serve as a source of a teaching signal for local prediction error neurons as well as for the next higher level of the hierarchy, while simultaneously providing prediction signals to the preceding lower level. In contrast, the non-hierarchical variant posits that these layer 5 internal representation neurons provide local predictions to layer 2/3 prediction error neurons.

      The interaction between internal representation neurons and prediction error neurons differs fundamentally between the two variants. In the hierarchical variant, internal representation neurons excite positive prediction error neurons and inhibit negative prediction error neurons, while at the same time being inhibited by positive prediction error neurons and excited by negative prediction error neurons. In the non-hierarchical variant, this pattern of connectivity is reversed.

      This work is very exciting, timely, and carefully executed. The authors functionally, and later molecularly, identify layer 2/3 prediction error neurons in V1 and probe their interactions with genetically defined neuron types in cortical layers 5 and 6 using optogenetics. They demonstrate that the functional influence of putative prediction error neurons in layer 2/3 onto layer 5 is incompatible with the hierarchical variant, whereas the influence of layer 5 onto putative prediction error neurons in layer 2/3 is incompatible with the non-hierarchical variant. They then test an alternative hypothesis, in which layer 2/3 responses resemble prediction errors with respect to perturbations of artificial layer 5 activity patterns. To investigate this, they designed an experiment in which optogenetic activation of L5 IT neurons was closed-loop coupled to the mouse's locomotion speed in the absence of visual feedback, allowing them to probe the causal influence of L5 activity on layer 2/3 responses.

      Finally, the authors hypothesize that their data are more consistent with a joint embedding predictive architecture (JEPA) and outline experimentally testable predictions arising from this framework.

      We thank the reviewer for their comments. We address them below.

      While the work is overall convincing and significantly advances our understanding of the circuit-level implementation of predictive processing, there are a few weaknesses that should be addressed or discussed:

      (1) The authors define putative positive prediction error neurons as the 15% of neurons most responsive to grating onset and putative negative prediction error neurons as the 15% most responsive to visuomotor mismatch. While this selection would be expected to overlap with negative and positive prediction error neurons, the criterion is not sufficiently stringent (independent of the exact percentage chosen). In particular, classification of a neuron as a prediction error neuron should ideally be accompanied by evidence that it does not exhibit a significant increase in activity when the prediction matches the sensory input or teaching signal.

      We understand the reviewer’s intuition. We can indeed use other stimuli to identify prediction error neurons, like the relative suppression of responses in closed-loop running onset vs open-loop running onsets. This was the reason behind including Figure S1, to show that our selection results in expected pattern of running onset responses. We don’t typically use the running onset responses as they have an additional confound we have not fully understood. This is that running onset always tends to result in an increase of calcium activity in all neurons. This could have a variety of reasons: A) Contamination of hemodynamic occlusion signals (blood vessels tend to constrict at running onset, making it appear like an increase in calcium activity) – see Yogesh et al., 2025. B) The virtual coupling in our VR is not good enough to provide a true “closed loop” experience. Humans typically notice lags of larger than 30ms – in our VR it is approximately 100 ms. C). Running onset in head-fixed animals is not accompanied by a vestibular input. D) Predictive processing is wrong. We tend to think it is a combination of the three.

      We can also use combinations of the two criteria to select neurons - if the reviewer has a specific selection criteria in mind (top XXX% MM responsive AND top XXX% closed-loop suppressed, etc.) we are happy to repeat the analysis for that specific set of criteria, but the fundamental problem that we are using a functional response to select these neurons does not go away. We know that our functional selection criteria mean we select a population of neurons that is enriched for prediction error neurons. If the enrichment is too weak, we would expect to find no effects in terms of functional influence. It is hard to explain, however, how a weak enrichment could result in a strong effect on functional influence. Our arguments in more lengthy form, for why the visuomotor mismatch is a good stimulus to identify negative prediction error neurons can be found here: Attinger et al., 2017; Jordan and Keller, 2020; Leinweber et al., 2017; O’Toole et al., 2023; Vasilevskaya et al., 2022; Zmarz and Keller, 2016.

      (2) The authors "speculate that the prediction error responses in layer 2/3 may not be computed with respect to sensory input, but with respect to layer 5 activity as a teaching signal." However, it is unclear how this perspective differs from earlier statements in the manuscript. In the Introduction, the authors note that "these signals, typically referred to as sensory signals, we will refer to as teaching signals," and later describe the hierarchical variant as one "in which internal representation neurons act as a source of the teaching signal." Given this framing, it is difficult to identify what is conceptually novel in the updated view. Is the key distinction that layer 2/3 neurons are now proposed to generate predictions in an internal representation space rather than in sensory input space, as briefly suggested in the Discussion? Or are the authors introducing a distinction between an external (sensory) and an internal (cortical) teaching signal? If so, this distinction should be made explicit. Clarifying this point would considerably strengthen the manuscript.

      There might be a misunderstanding regarding our usage of the term teaching signal. In hierarchical predictive processing the teaching signal is typically referred to as a sensory signal, as e.g. in: “prediction error neurons compare predictions to sensory input”. In non-hierarchical predictive processing, or far away from the sensory input (think prefrontal cortex), or for cross-modal interactions “sensory” input is misleading. Also, in non-predictive-processing type models (like JEPA), sensory input has a different functional role. Thus, we operationally define teaching signal as the signal that is compared against the prediction by prediction error neurons.

      The two primary options we are comparing are:

      (1) Is the teaching signal to the L2/3 comparator a bottom-up input to V1 (as one would expect in predictive processing). 

      (2) Is the teaching signal to the L2/3 comparator L5 input (as one would expect in JEPA).

      Our data argue in favor of option 2. We have rephrased parts of the manuscript to try to make this clearer. 

      (3) The authors propose that "L2/3 neurons predict L5 activity, hence making predictions in the internal representation space rather than the input space," and further suggest that, since both deep and superficial cortical layers receive thalamic input, the cortex may function like a JEPA. This idea appears closely related to the model introduced by Nejad et al. (2025), which effectively implements a JEPA-like architecture: L5 activity serves as a target against which L2/3 predictions are compared in a selfsupervised manner, with both L5 and L2/3 (via L4) receiving thalamic input. It would be helpful for the authors to clarify how their framework differs from that model, and to specify the key conceptual or mechanistic distinctions between the present proposal and the approach described by Nejad et al.

      The two proposals indeed share similarities in assuming that bottom-up input for both L2/3 and L5 arrives from thalamus, and that representations formed in L2/3 are used for predicting the activity of L5. However, there are a few important differences between the JEPA implementation proposal formulated here and the Nejad et al. model.

      (1) There is no proposed mapping of computations in the Nejad et al. model onto different JEPA networks. We assume that the suggested mapping would be L4 and L5 as encoder networks, and L2/3 as a predictor network? In that case, it is different to our proposal, in which L2/3 is part of the encoder network.

      (2) Our proposal contains explicit prediction error neuron cell types within L2/3, while prediction errors in Nejad et al. are encoded in the gradients, and the layer origin of these signals is hypothesized to be L5 (‘the learning-driving error signal originates in L5’). Hence, also the role of L5-L2/3 connection is distinct between the two proposals. In Nejad et al. this connection serves the role of error propagation and update for predictions in L2/3, while in our proposal this connection contains teaching signal (target representations) that are compared to predictions within L2/3. Similarly, the functional role of L2/3-L5 connection is also different, since in Nejad et al, it is supposed to carry predictions of L5 activity, whereas in our proposal we expect it to drive plasticity in L5 encoder.

      (3) The difference outlined above also makes it evident that the two proposals should differ in how deep and superficial layers are expected to influence the activity of one another. Indeed, the proposal in Nejad et al. is based on the cortical column idea, and according to eq. 2 and 3 in the Methods, activity in L5 is a function of activity in L2/3, while activity in L2/3 is not a function of activity in L5. Our proposal is based on idea of layers forming parallel networks, where horizontal communication is the dominant mode of cortico-cortical interactions, and activity in deep layers serve as a teaching signal for L2/3. In our case, we expect the opposite - that activity in L2/3 depends on activity of L5, while activity of L5 is not immediately dependent on activity of L2/3 (only via plasticity route). This led us to propose one of direct tests for our framework – silencing L2/3 in a familiar setting should result in no immediate changes to L5 activity and behavior of the animal.

      (4) The proposal in Nejad et al. relies on input reconstruction or variance maximization within the L5 autoencoder network to avoid collapse. Instead, our proposal has no reconstruction objective.

      (5) Lastly, there is a time delay between inputs to L5 and L2/3 that is proposed in Nejad et al., while this is not something inherent to our proposal.

      We expect that the most useful future models should move beyond JEPA, with the emphasis on nonhierarchical models capable of operating on arbitrary graphs. We think cortex functions according to principles of a JEPA (predictions in latent space), and that the role of cell types and the exact computational organization remain to be constrained.

      REFERENCES

      Attinger, A., Wang, B., Keller, G.B., 2017. Visuomotor Coupling Shapes the Functional Development of Mouse Visual Cortex. Cell 169, 1291-1302.e14. https://doi.org/10.1016/j.cell.2017.05.023

      Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P.H., Buchatskaya, E., Doersch, C., Pires, B.A., Guo, Z.D., Azar, M.G., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M., 2020. Bootstrap your own latent: A new approach to self-supervised Learning. https://doi.org/10.48550/arXiv.2006.07733

      Jordan, R., Keller, G.B., 2020. Opposing Influence of Top-down and Bottom-up Input on Excitatory Layer 2/3 Neurons in Mouse Primary Visual Cortex. Neuron 108, 1194-1206.e5. https://doi.org/10.1016/j.neuron.2020.09.024

      Leinweber, M., Ward, D.R., Sobczak, J.M., Attinger, A., Keller, G.B., 2017. A Sensorimotor Circuit in Mouse Cortex for Visual Flow Predictions. Neuron 95, 1420-1432.e5. https://doi.org/10.1016/j.neuron.2017.08.036

      Mohammadi, A.G., Halvagal, M.S., Zenke, F., 2025. Understanding cortical computation through the lens of joint-embedding predictive architectures. https://doi.org/10.1101/2025.11.25.690220

      O’Toole, S.M., Oyibo, H.K., Keller, G.B., 2023. Molecularly targetable cell types in mouse visual cortex have distinguishable prediction error responses. Neuron 111, 2918-2928.e8. https://doi.org/10.1016/j.neuron.2023.08.015

      Saravanan, V., Berman, G.J., Sober, S.J., 2020. Application of the hierarchical bootstrap to multi-level data in neuroscience. Neurons Behav. Data Anal. Theory 3, https://nbdt.scholasticahq.com/article/13927-application-of-the-hierarchical-bootstrap-tomulti-level-data-in-neuroscience.

      Vasilevskaya, A., Widmer, F.C., Keller, G.B., Jordan, R., 2022. Locomotion-induced gain of visual responses cannot explain visuomotor mismatch responses in layer 2/3 of primary visual cortex. https://doi.org/10.1101/2022.02.11.479795

      Yogesh, B., Heindorf, M., Jordan, R., Keller, G.B., 2025. Quantification of the effect of hemodynamic occlusion in two-photon imaging of mouse cortex. eLife 14, RP104914. https://doi.org/10.7554/eLife.104914

      Zmarz, P., Keller, G.B., 2016. Mismatch Receptive Fields in Mouse Visual Cortex. Neuron 92, 766–772. https://doi.org/10.1016/j.neuron.2016.09.057

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the Reviewing Editor for their careful evaluation of our manuscript and for their constructive and insightful comments. Their suggestions have helped us to improve the clarity, rigor, and presentation of our work. In response to these comments, we have substantially revised the manuscript and performed several additional analyses and experiments, as summarized below.

      Major additions and modifications made during revision

      In response to the reviewers' comments, we have substantially revised the manuscript and performed several additional analyses and experiments:

      New analyses

      - Quantification of MyoF recovery following auxin washout using MyoF-mAID-HA immunofluorescence (Figure 8B, Figure S10D).

      - Quantification of maternal MIC2 fluorescence intensity following 24 h MyoF depletion and subsequent redistribution after auxin washout (150 micronemes per condition; Figure S10A,B).

      - Pearson correlation analysis of ANKER1-Halo and HDEL-GFP localization (Pearson's R = 0.92 ± 0.04; n = 30 parasites).

      - Additional probability-based analysis supporting regulated microneme inheritance.

      - Expanded analysis of microneme redistribution across larger replication stages (Figure S5D).

      New figures

      - Figure S6: Dual-labelling analysis of additional Group 2 organelles (ER, apicoplast, glideosome).

      - Figure S10A, B: MIC2 fluorescence intensity analysis following RB retention and redistribution.

      - Figure S10D: Correlation between MyoF recovery and phenotype rescue.

      - Figure S11: Schematic overview of quantification and analysis workflow.

      Additional experimental efforts

      - Generation of a MIC2-Halo / IMC1-mKATE / Cb-Emerald parasite line to improve visualization of RB-associated trafficking.

      - Multiple attempts to perform higher-temporal-resolution live-cell imaging. However, prolonged acquisition resulted in severe phototoxicity, replication arrest, and parasite death, preventing reliable long-term recordings.

      Textual and methodological revisions

      - Expanded Materials and Methods section with detailed descriptions of fluorescence quantification, colocalization analyses, and statistical procedures.

      - Re-evaluation of statistical analyses using two-tailed tests throughout.

      - Revision of manuscript text to clarify the evidence supporting RB-associated trafficking and to better acknowledge current limitations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work asks the question of how different organelles and structures in the apicomplexan parasite Toxoplasma gondii are recycled and/or segregated to the daughter cells during cell replication. In particular, they consider an unusual cell structure called the residual body that links replicating cells during the intracellular infection stage of this parasite. The residual body has historically been considered a 'dumping ground' for unnecessary relics of the mother cell during division, but this notion is increasingly being revised. Indeed, cell replication in Toxoplasma is often misinterpreted as cell division (cytokinesis), but in fact, the cell replicates its organelles and structures to multiple 10s of copies in seemingly distinctly formed daughter cells, but cytokinesis is delayed for many such cycles and typically only occurs simultaneously with parasite egress from its host cell. The residual body is, in fact, the connection between these pre-cytokinetic replicated daughters, and effectively, this is still a single cell at this stage. The authors have previously shown that an actin network extends through the residual body between these daughter cells, and ER and mitochondria common to all cells are also linked through this structure. This study examining the fates of organelles during cell replication is timely for continuing our understanding of how this fascinating component of the cell participates in these processes. The authors use Halo-tags as their principal tool to track discrete populations of proteins, labelling their organelle locations, and this provides beautiful insight into these processes.

      Strengths:

      Using dyes conjugated to Halo tags, this work elegantly tracks the fates of proteins synthesised by an original 'mother' cell over several replication cycles of pre-cytokinetic 'daughters'. Using this tool, they show that some organelles are made intact just once and that some of these can be subsequently sorted to the daughters (micronemes and rhoptries) while others are dismantled (IMC) and the daughters must make their own. A third set of organelles (largely synthesis, sorting, and metabolic compartments) is divided and inherited, and new daughter-synthesised proteins are added to the preexisting maternal proteins in these structures. A role for actin and myosin is clearly demonstrated for micronemes and rhoptries, and this correlates with their relatively late inheritance into the developing daughters. Overall, this work gives clarity to the behaviours of several cell structures during replication and paves the way to a better understanding of the mechanisms that drive the differences between structures and the universality of these processes in other apicomplexan parasites.

      Weaknesses:

      In addressing the question of residual body participation in sorting of organelles, it would be useful to clearly define this structure and when and where it is delineated from the posterior of a mother cell during the formation of daughter structures. This might seem like a moot point, but it would give clarity to notions of recycling and 'reservoirs'. Mother cells retain their active invasion apparatus until very late in daughter formation, and the need for micronemes and rhoptries to be released from this service late in the process might explain why they are only then trafficked to the cell posterior and then into the daughters. So, is this a distinct 'residual body' body function/reservoir or just a spatial constraint of this sequence of daughter formation? In subsequent cell replications (4, 8, 16... stages), is there a separation between the residual body that links them all and the posterior of each new 'mother cell', and if so, when is this distinction lost? This is important because without a definition, we might be confusing different processes.

      We thank the reviewer for this excellent and thoughtful question. The residual body (RB) emerges at the end of the first replication cycle, where it is delineated by the basal complex and persists as an IVN-associated compartment connecting all daughter parasites through both plasma membrane and cytoplasm. Previous EM and live-cell studies, including ours, have shown that the RB is not a passive remnant but a dynamic structure dependent on F-actin and unconventional myosins, supporting recycling, inter-parasite connectivity, and synchronous growth (Delbac et al., 2001; Muñiz-Hernández et al., 2011; Frénal et al., 2017; Periz et al., 2017).

      In the present study, the MyoF reversibility experiment provides strong support for a model in which RB functions as an active recycling hub. Upon MyoF depletion, maternal microneme and rhoptry proteins accumulate within the RB. Following restoration of MyoF expression, this material is redistributed to daughter organelles. We interpret this reversible phenotype as evidence that the RB represents a distinct and regulated trafficking intermediate rather than simply a by-product of late daughter cell formation.

      We agree with the reviewer that mother cells retain a functional invasion apparatus until very late during daughter formation, and that the delayed release of micronemes and rhoptries likely contributes to their late trafficking toward the cell posterior. However, our data indicate that once released, these organelles transit through a defined RB compartment that actively participates in their recycling rather than merely reflecting positional constraints. This has been previously well illustrated for micronemes, which are trafficked along F-actin filaments within the residual body (Periz et al., 2019).

      At later rounds of replication (4, 8, 16 parasites), previous studies have demonstrated the presence of multiple residual body centres within the same vacuole. However, the precise temporal and structural distinction between the RB linking parasites within the vacuole and the posterior of newly formed mother cells remains insufficiently resolved and is beyond the scope of the present study. Importantly, available ultrastructural and live-cell imaging supports the persistence of shared RB compartments connecting parasites within a vacuole, arguing against a simple conflation of posterior membranes and residual body material.

      While the primary aim of the current work was to investigate the RB's role in organelle recycling, we fully agree that a more precise definition of when and how the RB is formed, remodelled, and ultimately resolved during successive replication cycles will be essential to distinguish recycling from spatial constraints. We have revised the Discussion to better acknowledge this limitation and to avoid overinterpreting the role of the RB in organelle inheritance.

      Are rhoptries/micronemes that originate in one 'mother' able to be sorted to the 'daughters' from a distinct mother in this syncytium? If so, this would make it a sorting centre, but otherwise we could be just capturing the activities at the posterior of any given cell during replication. The authors' further thoughts on this would be very interesting.

      We agree with the reviewer that our current data do not definitively demonstrate whether rhoptries or micronemes originating from one “mother” parasite can be redistributed to daughters derived from another mother within the same syncytial vacuole. Nevertheless, our MyoF chase experiments are consistent with a model in which the RB/IVN functions as an active recycling and sorting hub rather than simply representing posterior trafficking events associated with individual parasites.

      Upon MyoF depletion, maternal micronemes accumulated within the RB. Following restoration of MyoF expression, these accumulated micronemes were subsequently redistributed to daughter parasites. This reversible redistribution is more consistent with an active recycling process than with passive accumulation alone.

      To further support this interpretation, we expanded the analysis presented in Figure S5 by including additional vacuoles and larger replication stages (new panel D). These analyses show that maternal micronemes are redistributed broadly and relatively evenly among daughter parasites. We additionally performed a probability-based analysis demonstrating that the recurrent and homogeneous redistribution patterns observed are highly unlikely to arise from stochastic capture events occurring independently at the posterior end of each parasite during replication. Together, these analyses support the interpretation that microneme redistribution is a regulated process.

      Direct demonstration of recycling between all parasites within a vacuole would require a system allowing simultaneous differential labeling of (i) daughter parasites derived from a specific mother cell and (ii) the maternal organelles originating from that same mother during a subsequent replication cycle. To our knowledge, such an approach is not currently technically feasible. Nevertheless, our live-cell imaging experiments provide additional support for communal redistribution, as microneme material accumulated within the RB was subsequently observed redistributing, albeit unevenly, across multiple tachyzoites within the same vacuole.

      The Group 2 structures are described as those that are divided between daughters and receive newly synthesised proteins that add to the maternal protein of these compartments. While this is a logical conclusion for several that are mentioned, where the maternal protein signal is seen to be depleted with replication (including for the apicoplast, ER, glideosome, and Golgi). Data for the addition of new proteins to these existing structures is actually only presented in direct support of this for the Golgi.

      We thank the reviewer for this important clarification. We initially selected the Golgi as a representative example because its morphology and restricted localization provide the clearest visualization of the dual-labeling dynamics. However, the same experimental approach was applied to all Group 2 organelles analyzed in this study. To address the reviewer's concern more directly, we have now included a new supplementary figure (Figure S6) showing that the same pattern is also observed for the apicoplast, ER, and glideosome.

      We would also like to clarify that the maternal protein signal is not lost during replication but instead becomes progressively diluted as these organelles expand, are partitioned into daughter parasites, and incorporate newly synthesized proteins. The Golgi was originally highlighted because these dynamics are most readily visualized in this compartment; however, the same principle applies to all Group 2 organelles analyzed in this study, as now illustrated in Figure S6.

      Reviewer #2 (Public review):

      Summary:

      Toxoplasma gondii is an obligate intracellular parasite and the causative agent of Toxoplasmosis. Parasite invasion into host cells, intracellular replication, and then egress, which results in the destruction of the infected cell, is central to pathogenicity. This manuscript focuses on understanding how maternal resources (in this case, cellular organelles) are shared between daughter parasites during cell division. Many organelles are single copy, meaning that division and inheritance by the daughters is crucial for successful replication. The major strength of this study was the use of a Halobased pulse chase assay to characterize patterns of organelle inheritance. The results show that both microneme and rhoptries (secretory vesicles) previously thought to be synthesized de novo are inherited by daughter parasites. Thus, this paper adds new insight to our understanding of cell division in this important parasite.

      Strengths:

      This study demonstrated that pulse labeling of proteins can be used to monitor protein synthesis, turnover, and movement. This approach will be of great interest to the field. Using this method, the authors demonstrate three main modes of organelle inheritance.

      (1) Organelles, where there are multiple copies (such as secretory vesicles, micronemes, and rhoptries), are divided between the daughter parasites, with additional contribution of newly formed vesicles. New and old material remain as separate entities in the cell.

      (2) Single-copy organelles, which are expanded to include newly synthesized material prior to division, such as the Golgi and apicoplast.

      (3) Cytoskeletal structures that are synthesized anew during each round of division. These studies provide more refined insight into patterns or organelle inheritance and demonstrate that secretory organelles are not made de novo during each round of division as previously thought. The paper has a logical flow, and overall, the data is presented in a clear and organized fashion.

      Weaknesses:

      (1) Descriptions of methodology and statistical analysis were incomplete.

      We agree with the reviewer that the description of the methodology and statistical analyses required further clarification. To address this, we have added a new supplementary figure (Figure S11) illustrating the experimental workflow, quantification strategy, and analysis pipeline. We have also expanded the Materials and Methods section to provide detailed descriptions of the experimental design, fluorescence quantification procedures, statistical analyses, and the number of biological replicates. These revisions provide a clearer and more comprehensive description of the methodology and data analysis.

      (2) There are inconsistencies between the data in Figures 1 and 5. In Figure 1, a small amount of maternal IMC is visible in stage 2 parasites. Although this is a ~90% reduction, these parasites should be quantified as parasites with material IMC. However, the graph in Figure 5C indicates that no material parasites have GAPM1a, given that graph 5C is a binary measure (present vs. absent), one would expect a non-zero percent of parasites to have maternal material.

      We agree with Reviewer 2 that, based on the raw fluorescence signal, one might expect a non-zero percentage of parasites to retain maternal IMC material after the first replication. The apparent discrepancy between Figures 1 and 5 reflects our thresholding strategy rather than inconsistent data.

      Figure 5C presents a binary analysis (presence versus absence) using a threshold calibrated from stage 1 parasites and applied uniformly across all markers. Under this criterion, the residual GAPM1a signal after the first replication falls below the detection threshold, resulting in 0% positive vacuoles. Although normalization to stage 2 parasites would detect this weak residual signal, such a protein-specific threshold would compromise direct comparison across the dataset.

      To clarify this point, we have updated the Figure 5C legend to explain the analytical approach and the asterisk associated with GAPM1a. The residual maternal IMC signal visible in Figure 1 represents a rare example selected to illustrate the remaining ~10% signal and is consistent with the absence of detectable maternal IMC1 after replication in Figures 2C and 5E.

      (3) The conclusion from Figure 6 was not justified based on the data. I agree with the author's conclusion that the accumulation of micronemes and rhoptries in the residual body was timedependent. In Figure 6A, the signal observed in the residual body at times 6:30, 13, and 14 hours is not observed in subsequent time points. However, the fate of these micronemes and rhoptries is unclear. It cannot be concluded that these vesicles are recycled back to the mother. They could also have been degraded. In fact, the graphs of microneme inheritance in Figure 2B show a decrease in maternal signal from 100% to 80% between stages 1 and 2, indicating that some microneme degradation is taking place.

      We agree with the reviewer that Figure 6 alone does not definitively establish the fate of micronemes and rhoptries accumulating within the residual body (RB), and that both recycling and degradation remain possible interpretations. Our conclusion that maternal micronemes are predominantly recycled is therefore based on the integration of Figure 6 with our MyoF depletion and recovery experiments, additional quantitative analyses, and previous work demonstrating F-actin-dependent microneme trafficking through the RB (Periz et al., 2019).

      Consistent with this model, MyoF depletion results in the accumulation of maternal micronemes within the RB, whereas restoration of MyoF expression following auxin washout leads to their redistribution across multiple tachyzoites within the same vacuole (Figure 8). Furthermore, maternal microneme signal remains detectable even after prolonged MyoF depletion (up to 48 h) and multiple rounds of replication (Figures 7 and 8), arguing against extensive degradation.

      To further address this possibility, we quantified the fluorescence intensity of individual maternal MIC2-positive micronemes retained within the RB after 24 h of MyoF depletion and following redistribution after auxin washout (150 micronemes per condition). No significant difference in fluorescence intensity was observed compared with control maternal micronemes (Figure S10A,B), indicating that maternal microneme signal is preserved during RB retention and redistribution.

      We therefore interpret the decrease in maternal microneme signal observed between stages 1 and 2 in Figure 2B primarily as a consequence of redistribution and dilution rather than degradation, consistent with the stable fluorescence intensity of individual micronemes (Figure 3). Regarding rhoptries, we note that the majority (~90%) are incorporated into daughter parasites before budding is complete, limiting their accumulation within the RB and suggesting that RB-associated trafficking primarily reflects redistribution rather than bulk degradation.

      (4) To convincingly demonstrate that the redistribution of micronemes and rhoptries was due to recovery of MyoF protein levels after auxin washout, a Western blot should be performed to show MyoF protein levels over time. In addition, the decrease in mMIC2 protein levels in the residual body in Figure 8F should be measured and normalized for photobleaching. Both apical and basal signals appear to be reduced over the time course of imaging.

      We agree with the reviewer that demonstrating MyoF recovery following auxin washout is important. Rather than performing a Western blot, we monitored MyoF recovery by immunofluorescence using the HA tag in the MyoF-mAID-HA strain, allowing direct correlation between MyoF reappearance and microneme redistribution at the single-vacuole level. These data are now included in Figure 8B, with the corresponding MyoF presence–phenotype association analysis presented in Figure S10D.

      Regarding photobleaching, we agree that fluorescence loss during time-lapse imaging is an important consideration. However, in this experiment, changes in fluorescence intensity reflect not only photobleaching but also biological redistribution of micronemes and movement of parasites in and out of the imaging plane. In the absence of a stable internal reference fluorophore, applying a standard photobleaching correction could therefore introduce additional inaccuracies. For this reason, we did not quantify fluorescence intensity during the redistribution phase.

      Instead, to assess whether maternal microneme signal is lost during RB retention and redistribution, we quantified the fluorescence intensity of individual maternal MIC2-positive micronemes following 24 h of MyoF depletion and subsequent auxin washout (Figure S10A,B). No significant difference was observed compared with control maternal micronemes, supporting the conclusion that redistribution occurs without substantial loss of the maternal microneme pool.

      Reviewer #3 (Public review):

      Summary:

      Knoerzer-Suckow et al. explore the mechanisms of organelle inheritance during endodyogeny in Toxoplasma gondii using an innovative dual-labeling approach to track the distribution of maternal organelles into daughter parasites. They can clearly distinguish between maternal and daughterderived organelles using their dual-labeling Halo Tag approach. They reveal that different organelles are trafficked to daughter parasites in three broad patterns, which they have binned into groups. Their findings reveal a role for MyoF in the inheritance of micronemes and rhoptries, and notably, they observe that the inner membrane complex (IMC) is not recycled. Instead, the IMC undergoes a pronounced relocalization to the posterior of the maternal cell, where it is likely targeted for degradation.

      Strengths:

      The data surrounding their MyoF knockdown experiments, IMC degradation, and trafficking of MIC2 after auxin washout are compelling. These data add to the knowledge of how organelle inheritance occurs in T. gondii, increasing the field's understanding of endodyogeny.

      Weaknesses:

      (1) The evidence provided to support the claim that microneme and rhoptry inheritance specifically traffics through the residual body does not sufficiently substantiate the claim. The temporal resolution of the imaging is inadequate to precisely trace the path of microneme and rhoptry inheritance. From the data shown in the manuscript, it can be concluded that at least some of the micronemes and rhoptries might be recycled through the residual body, but it is unclear whether many or most of these organelles do so.

      We thank the reviewer for this important comment and refer also to our response to Reviewer 1 above.

      Previous work has demonstrated F-actin-dependent trafficking of micronemes within the residual body (RB) (Periz et al., 2019). Consistent with these findings, our data support a model in which RB-mediated trafficking contributes to maternal microneme recycling. We acknowledge, however, that the temporal resolution of our imaging does not allow continuous tracking of every individual organelle throughout the entire replication process.

      In contrast, our observations indicate that the majority of maternal rhoptry material is incorporated into daughter cells before replication is complete and therefore does not necessarily transit through the RB under normal conditions (Figure 6). Nevertheless, rhoptry inheritance remains dependent on the actin–MyoF trafficking machinery, as MyoF depletion results in the accumulation of rhoptry material within the RB (Figures 6 and 7).

      Taken together, our data support a model in which the RB serves as an important recycling hub for maternal micronemes and can, under conditions of impaired trafficking, also transiently accommodate rhoptry material. However, our imaging resolution does not allow us to conclude that all microneme or rhoptry inheritance obligatorily transits through the RB, and we have revised the manuscript to reflect this limitation more explicitly.

      (2) The absence of specific markers for the residual body brings into question whether microneme inheritance occurs through a discrete residual body or simply via the basal end of the maternal parasite. The authors need a robust way to visualize and define the residual body to claim that micronemes and rhoptries are specifically transported through this structure.

      We agree with the reviewer that the absence of a dedicated residual body (RB) marker remains a limitation and that such a tool would improve the precision of our analyses. To date, no specific RB marker has been identified (see also our response to Reviewer 1). The most reliable proxy currently available is the F-actin chromobody, which labels the dense F-actin network associated with the RB. Using this approach, previous work demonstrated F-actin-dependent trafficking of micronemes within the RB (Periz et al., 2019).

      Building on these findings, our data support a model in which RB-associated trafficking contributes to maternal microneme recycling, whereas rhoptries are more frequently incorporated directly into daughter cells without obvious RB transit. In addition, functional perturbation of the actin–MyoF transport machinery, through MyoF depletion and subsequent recovery, supports the interpretation that the RB represents a discrete actin-associated compartment involved in organelle redistribution. Nevertheless, we acknowledge that our current imaging resolution does not allow us to determine the extent to which all microneme or rhoptry inheritance occurs through the RB, and we have revised the manuscript accordingly.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Comments for revision where either the clarity or accuracy could be improved:

      (1) The methods could do with some further detail with respect to the fluorescence intensity measurement. For example, where the Z-series was taken, and for measurements, were maximum projects taken or single Z-planes? Were all measurements made on unprocessed images, or was any deconvolution, etc, undertaken?

      We updated the Materials and Methods for better clarity and now read as: “Parasites were labeled as described above and allowed to replicate for 24 h on HFF-coated Ibidi live-cell dishes. Approximately 15 fields of view were imaged using Z-stacks spanning 3 μm centred on the vacuoles. For each replication stage (1, 2, 4, and 8 parasites per vacuole), individual tachyzoites were sampled across multiple vacuoles.

      Maximum-intensity projections were generated from non-deconvolved images. Fluorescence intensity (FI) was quantified as the maximum grey value measured within regions of interest (ROIs) drawn on individual tachyzoites from each vacuole stage. ROIs excluded overlapping parasites, neighbouring vacuoles, and regions with atypical signal intensity. For each biological replicate, up to 25 tachyzoites per replication stage were analyzed, and the mean FI value was calculated for each stage. The highest mean FI observed among the stages within a replicate was defined as 100%, and FI values for the other stages were expressed relative to this maximum. Relative FI values were then averaged across three independent biological replicates. Data are presented as mean ± SD.”

      (2) Line 93: Is more JF549 added at each of stages 2, 4, and 8? I assume so to get the progressive increase, but it would help to clarify this here.

      As illustrated in Figure 1a, the second dye is added only once, after 24 h of replication, and is then washed off before imaging. It is not maintained throughout the replication steps. The observed increase in signal is related to the presence of newly synthesized (de novo) proteins generated during replication. The Halo tags of these new proteins are initially free of ligand, as no ligand is present during replication. During the second labelling step, all free Halo tags can bind the dye. The signal intensity at each replication step therefore reflects the amount of de novo material present, which is higher at step 4 than at step 2, because the proteins generated de novo during stage 2 are also present in stage 4. The signal will be determined by both the number of newly produced molecules and their concentration at the same localization.

      (3) Line 103: Check that the Carruthers and Sibley, 1997, ref for Tic20 in the apicoplast is correct. I don't think this could be correct given the date.

      The reference it will be corrected to G van Dooren et al. 2008.

      (4) Figure legends: It would be useful to state what form of microscopy was used in each figure.

      Following the reviewer's advice the legends has been updated.

      (5) Figure 2, S2: How are single organelles tracked, such as rhoptries? I'd assume that a cell will gain new de novo organelles as well, and that this would reduce the signal per cell. Stage 2 has a rhoptry signal in both daughters, so I'd expect the signal to be roughly half for the whole cell, unless the authors can resolve individual rhoptries (which would surprise me with this microscopy). If individual rhoptries were resolved, how was this done, and what was the confidence in this (were there controls?)

      We thank the reviewer for this important point. We do not resolve individual rhoptries with the imaging conditions used in this study. Instead, fluorescence intensity was measured as the maximum grey value within a representative region of interest (ROI) encompassing the apical rhoptry signal while excluding overlapping parasites and regions with atypical fluorescence intensity. The same ROI selection strategy was applied consistently across all replication stages, allowing direct comparison with stage 1 parasites.

      The analyses presented in Figures 3 and S2 show an increasing proportion of tachyzoites lacking detectable maternal rhoptry signal as replication progresses, while the fluorescence intensity of the remaining maternal signal remains relatively stable. Together, these observations are consistent with the redistribution of intact maternal rhoptries rather than a progressive loss of rhoptry fluorescence. 

      (6) Line 150: it is unclear what is meant by 'regulated partitioning'. The more equal inheritance of micronemes versus rhoptries might not indicate a 'regulated partitioning' but just a more uniform distribution, given the larger number of micronemes versus rhoptries.

      We agree with the reviewer that this statement required clarification, and we have revised the text accordingly. Our intention was to emphasize that microneme inheritance appears to be a regulated process, rather than to directly compare it with rhoptry inheritance. The observed differences between these organelles are likely influenced, at least in part, by their different abundances.

      Across successive rounds of replication, daughter parasites consistently inherit comparable amounts of maternal micronemes, even at later replication stages (Figure S4). Given that a single mother parasite contains approximately 30–40 micronemes, whereas successive rounds of endodyogeny can generate up to 32 daughter parasites, a purely stochastic segregation would be unlikely to produce the relatively uniform distribution observed (~1–2 maternal micronemes per tachyzoite). To support this interpretation, we performed an additional probability-based analysis, which indicates that the observed redistribution patterns are unlikely to arise by chance alone. We therefore interpret these findings as supporting the existence of mechanisms that promote balanced microneme inheritance during parasite replication.

      (7) Line 184: How is the maternal signal measured without detecting the internal daughter signal? Is this an average signal for the full parasite, or just for a cross-section of the IMC? And if the latter, how are the different profiles of mother and daughter accounted for? Also, the abbreviation in the brackets doesn't make sense here.

      We thank the reviewer for this important point. We have revised the Materials and Methods section to provide a clearer description of the fluorescence intensity (FI) measurements and added a new supplementary figure (Figure S11) illustrating the analysis workflow.

      Briefly, FI measurements were performed on maximum-intensity projections generated from Z-stack images without deconvolution. A representative region of interest (ROI) was selected, and the maximum grey value was used for quantification. This approach minimizes variability arising from differences in ROI size and provides a robust metric for comparison across replication stages.

      Daughter cell fluorescence was measured using the same approach while excluding overlapping signals from neighboring daughter cells and the maternal IMC. Maternal and daughter signals were distinguished based on their spatial localization and fluorescence labeling. Finally, the abbreviation in brackets has been corrected for clarity.

      (8) Line 191: Subheading a bit unclear. Distinct from other organelles, or are miconeme and rhoptry pathways distinct from each other?

      We agree with the reviewer and have updated the subheading to “Whole-organelle inheritance of micronemes and rhoptries occurs via distinct recycling pathways”

      (9) Line 193: The site of disassembly of the IMC (suggested RB here) might not be the same as the site of degradation. I suggest using 'disassembly' instead here.

      We agree that “disassembly” is an appropriate term to describe the breakdown of the IMC at the residual body (RB). However, we also believe that the RB represents the primary site of IMC degradation, for two reasons. First, if IMC material were not degraded at this site, we would expect to detect Halo-positive signal elsewhere following IMC collapse, which we do not observe. Second, transport of IMC material to an alternative degradation site would be required, but no IMC-positive vesicles are observed, arguing against significant redistribution. Together, these observations support the conclusion that the RB is both the site of disassembly and degradation of maternal IMC.

      (10) Line 215: The conclusion for a difference in timing of microneme and rhoptry segregation is not clearly supported by the data presented. Also, if there are more micronemes than rhoptries, then the frequency of observing a microneme being trafficked through the RB would need to be higher than for rhoptries if the mechanisms were the same. So, a difference in frequency here cannot be used to argue for a different mechanism.

      We agree with the reviewer and the text have been edited to soften our conclusion. 

      (11) Line 234: 'segregation' might be a better term than 'recycling' here because it is actually the sorting into daughter cells that is the important process.

      The text have been edited

      (12) Line 235: I don't think this can be what the authors intend to say. If the maternally-inherited rhoptries are not trafficked through the RB (every time), then how do they get into the daughters? Perhaps this is a case where a clear definition of the RB is required.

      Our observations indicate that maternally inherited rhoptries are frequently incorporated into daughter cells before collapse of the mother cell and establishment of the residual body (RB). Although F-actin is enriched within the RB, an actin network is also present throughout the parasite cytoplasm, where MyoF is likewise localized. We therefore propose that, unlike micronemes, rhoptries do not necessarily transit through the RB during every replication cycle but can be incorporated directly into developing daughter cells while still relying on the same actin–MyoF-dependent trafficking machinery.

      (13) The MyoF Rescue, the experimental plan is not fully described in order to be clear. If the endomembrane architecture was disrupted by MyoF depletion, and this secondary effect caused the segregation phenotype, restoration of MyoF might also simply restore the endomembrane system. So a direct role for MyoF doesn't seem to have been tested in this case.

      We appreciate the reviewer's concern that the segregation phenotype could, in principle, arise indirectly from disruption of endomembrane architecture following MyoF depletion. However, although Golgi morphology is altered in MyoF-depleted parasites, its core functions appear largely preserved. This is supported by the normal biogenesis of de novo micronemes, their correct targeting to the apical pole, their efficient secretion, and the previously reported preservation of parasite invasion. In addition, Golgi markers are not detected in the residual body, where maternally inherited micronemes accumulate, arguing against Golgi-mediated trafficking as the primary cause of the segregation phenotype.

      Taken together, these observations support the interpretation that the segregation defects are more likely to reflect a direct role of MyoF in organelle trafficking and inheritance than a secondary consequence of generalized disruption of endomembrane organization.

      (14) Line 279: Why call it a checkpoint? What is the evidence for its presence here being sensed before a further process is activated, which is what a checkpoint does?

      We agree the reviewer that checkpoint is a misleading term and have been updated to trafficking hub. 

      (15) Line 287 confuses replication of the daughters from cytokinesis, which only happens when each cell loses cytoplasmic connectivity with the other.

      We will clarify this point. In Toxoplasma gondii, cytokinesis represents the final step of daughter cell formation, during which the two fully assembled daughter parasites separate from the mother cell following collapse of the maternal cytoplasm. Historically, the residual body was proposed to arise simply as leftover material from this process. However, multiple studies have now shown that residual body formation is an active and regulated process, dependent on specific cytoskeletal and trafficking factors. Importantly, although cytokinesis marks the physical separation of daughter cells from the mother, parasites within a vacuole remain connected via the residual body and continue to share cytoplasmic and plasma membrane components until egress. 

      (16) Line 298: I don't think there is direct evidence of degradation in the RB. There might be disassembly, but degradation implies proteolysis, which hasn't been tested for.

      We agree with the reviewer that our data do not provide direct biochemical evidence of proteolysis within the residual body (RB) and primarily demonstrate disassembly of the maternal IMC at this site. However, several observations are consistent with local degradation. Following IMC collapse, we do not detect Halo-positive signal elsewhere in the parasite, nor do we observe IMC-positive vesicles or other structures that would suggest transport to a distinct degradation compartment.

      In addition, previous work identified the E3 ubiquitin ligase CSAR1 as a mediator of protein turnover within the RB, supporting the idea that this compartment is associated with degradation-related processes (O'Shaughnessy et al., 2023). While we cannot formally demonstrate proteolysis, these observations support a model in which IMC disassembly is closely coupled to local degradation within the RB.

      (17) The paragraph structure gets a bit confusing at times. See single sentence paragraph, Line 224. Does this sentence justify its own paragraph?

      The text has been edited.

      (18) Make sure Toxoplasma gondii is in italics throughout.

      The text has been edited

      (19) Line 279 cites Figure 10. But there is none.

      The text has been edited

      (20) I advocate introducing a few new acronyms, like DCs. I find that this ultimately reduces the ease with which readers read the work if they don't learn them all quickly.

      We agree that excessive use of acronyms can negatively impact readability. In the present manuscript, all abbreviations used in the text are introduced at their first occurrence in the Introduction, including DCs (line 32), IMC (line 41), PV (line 35), ER (lines 38–39), RB (line 50), and IVN (line 49). We have carefully limited the use of abbreviations to commonly used terms in the field and to those that recur frequently throughout the manuscript, with the aim of balancing clarity and readability. Nevertheless, we are happy to reduce or remove specific abbreviations if the reviewer feels this would further improve clarity.

      Reviewer #2 (Recommendations for the authors):

      (1) Descriptions of methodology and statistical analysis were incomplete as follows:

      (1a) It was unclear how the fluorescence intensity measurements (used to evaluate inheritance vs. new synthesis) were carried out. The y-axis on the graph is labeled average fluorescence intensity (% of max intensity). However, it does not state what was averaged (average fluorescence per vacuole?) and what was max intensity (max pixel intensity in each image or time point with the highest average intensity, relative to the other time points?)

      We agree with the reviewer that the original description of the fluorescence intensity (FI) measurements lacked clarity. We have therefore revised the Materials and Methods section and added a new supplementary figure (Figure S11) illustrating the analysis workflow.

      Briefly, vacuoles were imaged as Z-stacks, and maximum-intensity projections were used for analysis. FI was quantified as the maximum grey value measured within representative regions of interest (ROIs) drawn on individual tachyzoites, rather than as an integrated fluorescence signal across the vacuole. This approach minimizes variability arising from differences in ROI size and allows direct comparison between replication stages.

      For each biological replicate, up to 25 tachyzoites per replication stage were analyzed. The mean FI for each stage was normalized to the highest mean value within that replicate, and data from three independent biological replicates were subsequently averaged.

      (1b) Given the uncertainties with how these measurements were performed, it is difficult to interpret the data. For example, one would expect that the fluorescence intensity of newly synthesized IMC1 in 8-parasite vacuoles would be 4 times higher than that of a 2-parasite vacuole; however, based on the graph in Figure 1B, the measured increase was only 30%.

      We agree that the original description of the fluorescence intensity (FI) measurements required further clarification and have revised the Materials and Methods accordingly. As the reviewer correctly notes, a fourfold increase in FI between 2- and 8-parasite vacuoles would be expected if total IMC fluorescence across the entire vacuole had been measured. However, this was not the parameter quantified.

      Instead, FI was measured as the maximum grey value within representative regions of the daughter IMC, providing a per-cell rather than a whole-vacuole measurement. Using this approach, FI increases between the 2- and 4-parasite stages and then reaches a plateau.

      This behavior is consistent with the biology of IMC biogenesis. Although the total amount of IMC per vacuole increases with parasite number, the amount of IMC protein incorporated into each daughter parasite remains relatively constant. Consequently, once daughter IMCs are fully assembled from de novo-synthesised material, additional rounds of replication increase the total IMC content per vacuole but not the fluorescence intensity measured for individual parasites.

      (1c) T. gondii replicates in an asynchronous manner, so that at the 24-hour time point, a single dish can contain vacuoles containing 2, 4, and 8 parasites. This should be stated explicitly so readers unfamiliar with T. gondii's growth patterns can understand how the experiment was performed.

      We agree with the reviewer and have updated the text line 93. “As Toxoplasma gondii replicates in an asynchronous manner, after 24 of replication, vacuoles containing 1, 2, 4, and 8 parasites can be observed in a single dish.”

      (1d) Colocalization package in Fiji used for ANKER1-Halo with HDEL-GFP and MIC2/RON2 with CbEmeraldFP should be specified.

      We thank the reviewer for this suggestion. Following this recommendation, we performed Pearson correlation analysis for the ANKER1–HDEL-GFP experiment using the Coloc 2 plugins of FiJi. ANKER1Halo and HDEL-GFP showed a strong spatial correlation (Pearson's R = 0.92 ± 0.04, n=30 parasites from three independent biological replicates), supporting localization of ANKER1 to the ER.

      We note, however, that this analysis should be interpreted as evidence for co-distribution within the same organelle rather than direct molecular colocalization, as ANKER1 is a transmembrane protein whereas HDEL-GFP labels the ER lumen.

      For all the rest of our analysis, no automated colocalization package or plugin in Fiji was used for the analyses involving ANKER1-Halo with HDEL-GFP or MIC2/RON2 with Cb-EmeraldFP. Colocalization was assessed manually across all experiments by inspecting both full Z-stacks and maximum-intensity projections to ensure robust spatial overlap.

      For MIC2 and RON2, the presence of signal within the Cb-Emerald–positive filament was scored as either cytoplasmic, on the residual body or absence of colocalisation.

      In total, more than 300 and 500 vacuoles were analyzed for MIC2 and RON2–Cb-Emerald colocalization respectively (stable expression of both markers), and more than 150 vacuoles were analyzed for ANKER1-Halo and HDEL-GFP colocalization (transient expression of HDEL-GFP).

      This information has now been added to the Methods section.

      (1e) Statistical methods should be described on an experiment-by-experiment basis. The authors should justify why a one-tailed t-test was conducted. A two-tailed t-test seems more appropriate.

      We agree that statistical methods should be clearly justified on an experiment-by-experiment basis. All the statistical analysis have been performed using two tails and updated in the figures.

      (1f) In Figures 5E and 5F, using boxes to indicate the exact areas of the cell that were used in the fluorescence intensity measurements, rather than arrows, would make this data easier to interpret.

      The figure has been updated.

      Reviewer #3 (Recommendations for the authors):

      The current time-lapse images and videos do not clearly demonstrate microneme movement from the maternal parasite apical end to the residual body and back to the apical end of daughter parasites. As such, the route by which micronemes enter daughter parasites remains inconclusive. To strengthen their claims, the authors should employ higher temporal resolution imaging to definitively capture the movement of micronemes from the maternal apical region into the daughters. From the current data, it also seems plausible that the micronemes may be trafficked into the daughters through the conoid as well, as there is no evidence provided showing a movement of micronemes away from the apical end of the maternal parasite before being present in the daughter parasites.

      We agree with the reviewer that higher temporal resolution imaging would provide a more definitive view of microneme trafficking. However, long-term live imaging of replicating Toxoplasma gondii requires a compromise between temporal resolution and parasite viability. In our experiments, images were acquired every 15–30 min over periods of up to 16 h, as more frequent acquisition consistently induced phototoxicity and prevented completion of parasite replication.

      Despite this limitation, our imaging reliably tracked maternal micronemes over successive rounds of endodyogeny and consistently showed microneme signal associated with the residual body. These observations are in agreement with previous high-temporal-resolution studies, which demonstrated F-actin-dependent microneme trafficking within the residual body over shorter imaging periods (Periz et al., 2019).

      We cannot formally exclude the possibility that some micronemes are transferred directly to daughter parasites through the apical end. However, together with previous studies showing enrichment of F-actin at the basal region of developing daughter cells rather than at the apical tip (Periz et al., 2017), our observations support a model in which RB-mediated trafficking contributes to maternal microneme inheritance.

      The lack of a clear residual body marker needs to be addressed, as the distinction between the basal end of the maternal cell and a bona fide residual body must be explicitly defined to substantiate the major claim of the study. As it stands, it remains unclear whether micronemes and rhoptries as a whole travel through the residual body to be transported into the daughter parasites.

      We agree with the reviewer that a marker specific to the residual body would strengthen this study. Unfortunately, no such marker has been identified to date. The F-actin chromobody is currently the best available proxy, as previous studies have shown that the F-actin network is enriched within the residual body (Periz et al., 2017; Kellermeier et al., 2024). Moreover, high-resolution live-cell imaging has previously demonstrated F-actin-dependent microneme trafficking within this compartment (Periz et al., 2019). We have revised the manuscript to more clearly acknowledge this limitation.

      The conclusions drawn from the actin colocalization data in Figure 6C are based entirely on fixed samples, despite all experimental tools being compatible with live-cell imaging. Supplementing the fixed imaging with live cell data would increase its biological relevance. Published studies have shown that fixation of the actin chromobody results in the loss of resolution of an appreciable amount of the cytosolic F-actin network, and while the localizations analyzed here are primarily along the periphery, since the quantification and text make claims about the colocalization within the cytosol, this potential loss of cytosolic F-actin becomes an issue as there may be more actin available for analysis that is lost due to fixation within the cytosol of the parasites.

      We thank the reviewer for raising this important point and agree that conventional fixation can compromise preservation of the F-actin network. However, we used the same fixation protocol described by Periz et al. (2019), which allows reliable visualization of the RB-associated F-actin network. We have also corrected the description of the fixation protocol in the Materials and Methods.

      Fixation was necessary to image entire vacuoles with sufficient spatial resolution and signal-to-noise ratio for the volumetric analyses presented in Figure 6C. Although some loss of cytosolic F-actin cannot be excluded, this would be expected to reduce, rather than artificially increase, the detection of organelle–actin associations.

      Importantly, previous live-cell imaging studies demonstrated F-actin-dependent microneme trafficking (Periz et al., 2019), and our observations are consistent with these findings. Moreover, the defects observed following MyoF depletion provide independent functional evidence that the trafficking events described here rely on the actin–MyoF transport machinery.

      The statement of colocalization should be backed up by quantitative coefficients like Pearson's coefficient.

      We thank the reviewer for this helpful suggestion. Following this recommendation, we performed a Pearson correlation analysis of ANKER1-Halo and HDEL-GFP using the Coloc 2 plugin in Fiji. ANKER1-Halo showed a strong spatial correlation with HDEL-GFP (Pearson's R = 0.92 ± 0.04, n = 30 parasites from three independent biological replicates), supporting localization of ANKER1 to the ER. As ANKER1 is a transmembrane protein and HDEL-GFP labels the ER lumen, this analysis should be interpreted as evidence of co-distribution within the same organelle rather than direct molecular colocalization.

      In contrast, we do not consider Pearson's coefficient appropriate for evaluating the association of micronemes or rhoptries with F-actin. These organelles are predominantly concentrated at the apical pole and, when associated with F-actin, are typically positioned along rather than directly overlapping the filaments. Consequently, Pearson's coefficient would underestimate these biologically relevant associations. We therefore relied on morphological and spatial criteria, which we consider more appropriate for assessing organelle–cytoskeleton interactions.

      In addition, from the methods and presented figure images, specifically in Figure 6C for RON2, how the cytosolic and residual body actin is separated is difficult to discern, as there is a clear residual body actin signal overlapping a parasite. The methods for how this was separated and analyzed should be clearer to remove doubts about how this area was measured, as the current description raises concerns about the counting of the residual body actin within the cytosol.

      We agree with the reviewer that the distinction between cytosolic and residual body (RB)-associated F-actin required further clarification. All analyses were performed manually, as described in our response to Reviewer 2 (comment 1d). The RB was identified by the presence of thick, bundled F-actin filaments at the basal pole that formed a continuous structure connecting parasites within the vacuole, whereas cytosolic F-actin was defined as the thinner filamentous network within the parasite body.

      No automated or threshold-based segmentation was used because the marked differences in filament morphology and fluorescence intensity make reliable thresholding difficult and prone to misclassification. Manual annotation based on spatial localization and filament morphology was therefore considered the most appropriate approach. We have clarified these criteria in the Materials and Methods section.

      Line comments:

      (1) 113: round to rounds - "did not obtain maternal organelles after successive rounds of replication...".

      The text has been updated

      (2) 146: grammatical, "As consequence a progressive" -> "As a consequence", or "Consequently".

      The text has been updated

      (3) 164-166: "Autonomous duplication" implies the separation and duplication of the Golgi occurs on its own, i.e., without any outside intervention, when we know from Carmeille et al. 2021 and Figure 7C here that the Golgi becomes fragmented over rounds of division in the absence of MyoF. I think this is primarily a word choice error with "autonomous".

      We agree with the reviewer and the word autonomous has been removed

      (4) 184: The wording suggests that DC's refers to daughter IMC's, when DC has already been given as an abbreviation for daughter cells previously.

      The text has been updated to correct this error

      (5) 189: de novo is not italicized.

      The text has been updated

      (6) 192: The data shown so far do not show that the RB plays a selective role in organelle recycling.

      The text has been edited to fit better our results “Our data suggest that the organelles trafficking through the residual body (RB) have different fate”

      (7) 208: State that it depends on F-actin, but never show that it is dependent on F-actin through actin disruption, such as cytochalasin D treatment or a specific conditional disruption of F-actin.

      We agree with the reviewer that we did not repeat F-actin disruption experiments (e.g., cytochalasin D or jasplakinolide treatments) in this study. These experiments were performed in our previous work, where pharmacological disruption of F-actin was shown to impair microneme trafficking (Periz et al., 2019). We therefore chose not to repeat these assays.

      Instead, the present study provides complementary evidence by demonstrating that depletion of Myosin F (MyoF), a motor that uses F-actin as a transport track (Kellermeier et al., 2024), disrupts the trafficking of both maternal micronemes and rhoptries. Together, our previous F-actin perturbation experiments and the MyoF depletion data presented here support the interpretation that these trafficking events depend on the actin–MyoF transport machinery. 

      (8) 209: This suggests that the chromobody was transiently expressed in the RON2-Halo line, but the methods suggest MIC2-Halo and RON2-Halo were integrated into a parasite line stably expressing Cb-Emerald.

      The text has been edited.

      (9) 229: The section is confusing with the mention of (now maternal). If I understand correctly, the point being made is that the de novo synthesized MIC2 at stage 2 is now the maternal MIC2 for stage 4, but coloring-wise within the figure, the now maternal MIC2 at stage 4 from stage 2 would still be green. The methods suggest these images were all taken simultaneously, and not at specific timepoints of the same vacuole, so the now maternal line remains confusing.

      The text has been revised for clarity and now reads: “Because Toxoplasma gondii replicates asynchronously, vacuoles at different replication stages coexist within the same culture after 24 h. The second labeling step marks all proteins synthesized since the beginning of the experiment, allowing discrimination between proteins present in the original mother parasite and those synthesized during subsequent replication cycles. As daughter parasites form, they inherit material from their mother, such that proteins synthesized during one replication cycle become maternal proteins in the next. Under MyoF depletion, these newly synthesized protein pools accumulate within the residual body instead of being redistributed to daughter parasites during subsequent rounds of replication (Figure 7, stage 4).”

      (10) 238: The sentences here indicate that Golgi inheritance occurs without issue: "golgi inheritance remained unaffected by MyoF depletion". But, it is evident from the images shown that the Golgi is extremely fragmented, with many more Golgi fragments by stage 8 than there are parasites. This could be solved by rewording and including a line along the lines of "in accordance with the results found in Carmeille et al. 2021".

      The reference to this article is already stated later in the text now line 299-303 “This active role is further supported by the dependence of RB-mediated recycling on F-actin and the class XXII myosin MyoF, which we show to be essential for retrieval of maternal MIC2 and RON2 but dispensable for Golgi inheritance although we noticed a fragmentation of the Golgi, which has been described to depend on MyoF (Carmeille et al., 2021).”

      (11) 319: missing a comma, "Many of these, particularly...".

      The text has been edited.

      (12) 437: Images -> imaged.

      The text has been edited.

      (13) 439: a fresh media -> and fresh media.

      The text has been edited.

      (14) 439: Replication -> replicate.

      The text has been edited.

      (15) 440: images -> imaged.

      The text has been edited.

      Other grammatical issues within the methods:

      (1) Figure 5C: Within the y-axis label of Figure 5C, there is an asterisk with no asterisk explanation within the legend.

      We thank the reviewer for pointing out this oversight. The figure legend has been updated to include an explanation of the asterisk, providing a clearer understanding of the results.

      (2) Figure 5E: The addition of an IMC1 label to match the magenta color would be helpful to readers.

      The figure has been updated

      (3) Figure 6D: No statistics showing significance of colocalization.

      We thank the reviewer for this comment. In Figure 6D, we report the frequency of observed colocalization between organelles (MIC2 and RON2) and F-actin filaments, based on manual analysis across >300 vacuoles for MIC2 and >500 vacuoles for RON2. Specifically, we observed cytoplasmic F-actin colocalization in 94% of vacuoles for MIC2 and 79% for RON2, and residual body F-actin colocalization in 85% of vacuoles for MIC2 and 11% for RON2. This analysis is descriptive and is not intended to compare MIC2 versus RON2 quantitatively; rather, it illustrates the general association of these organelles with F-actin filaments. Standard statistical measures such as Pearson correlation are not meaningful in this context because the organelles are punctate and primarily localized along filaments rather than overlapping continuously. Importantly, the percentages reported represent the fraction of vacuoles in which colocalization can be observed, not the percentage of colocalization between Cb-Emerald and MIC2/RON2 within individual vacuoles.

      (4) Figure 7: The legend title of Figure 7 is at the end of the legend of Figure 6.

      The text has been edited.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Here, Pinto and colleagues set out to investigate whether the cow udder is a potential mixing site for the influenza virus. The authors have demonstrated that bovine mammary epithelial cells can be infected with both avian and human influenza A viruses, supporting the idea that the cow udder may be a potential site for reassortment. Furthermore, they demonstrate that the bovine-adapted IAV replicates to similar titers in avian epithelial cells when compared to an AIV precursor virus. Thus, suggesting there is no fitness trade-off, and confirms the potential for spill-back of the cattle B3.13 into poultry, which has already been observed. Overall, I believe the authors achieved their aims. However, there are instances in which the results do not entirely support the conclusions (noted in weaknesses). Given the ongoing questions surrounding highly pathogenic avian influenza A virus in dairy cows, this work provides valuable evidence for the potential of the cow udder as a site of reassortment. These findings highlight the need for surveillance of influenza A virus incursions into livestock species, particularly cows. Some specific strengths and questions regarding weaknesses have been outlined below.

      Strengths:

      (1) The authors use a diverse range of cell types and influenza A virus strains, as well as a wide range of techniques to address the questions at hand.

      (2) The use of cells from multiple bovine breeds for the MAC-T, bMEC and explants suggests the phenomenon is not unique to a single breed.

      (3) The results suggesting there is no fitness trade-off for Cattle Texas in an avian host are interesting, and confirm the potential for spill-back of the cattle B3.13 into poultry, which has been observed.

      Weaknesses:

      I have listed my complete questions/concerns below. However, there are two main weaknesses of the article in its current state. Firstly, there is no apples-to-apples comparison in terms of determining a preference for IAV to infect the cow udder over other organs (Q4). The mammary gland and respiratory tract are represented by epithelial cells, but for other organs, fibroblasts were chosen. I think the fairer comparison would be to compare epithelial cells from different organs to demonstrate a preference for the mammary gland. Secondly, the main premise of the article relies on bMEC and MAC-T (primary and immortalised mammary epithelial cells), facilitating higher viral growth than the cells from other organs. Yet throughout the article, a 10x higher dose of IAV is used in the bMEC cells compared to everything else (Q6). This raises the question of how much of the results are due to a preference for the mammary epithelial cells, and how much is simply due to the increased dose.

      (Q4) When we set out to test if cow mammary gland cells were particularly susceptible to IAV infection compared to other bovine cell types, we used what was available in the Roslin Institute – a mix of primary and continuous cells from various anatomical sites: three epithelial cell types (two mammary, one respiratory tract) two immune cell types and four sets of fibroblasts from various organs. Given the representation of different anatomical sites, cell types and differentiation statuses, we considered this a suitably diverse panel with which to characterise infection dynamics of a broad range of IAVs, before more focussed investigations using the bMEC and explant tissues. Both mammary epithelial cell types grew our library of influenza challenge strains significantly better than the BAT-II respiratory epithelial cells, as well as the two immune cell types and all four fibroblast populations. Of the fibroblast cells, those derived from the brain grew IAV significantly better than the skin and turbinate fibroblasts, while blood-derived macrophages grew virus significantly better than the lymphocytes and non-brain fibroblasts. So there are “apple-to-apple” comparisons as well as apple-to-pear comparisons that give significant differences. We therefore think that our conclusions (in the abstract) that mammary cells are particularly replication competent for IAV, (at the end of the introduction) that “a wide range of cow-derived cells are susceptible” and that (in the results section) that “mammary cells showed the highest susceptibility” are justifiable. However, we agree that testing a wider variety of epithelial cells would be useful and have added text to the Discussion (lines 224-228) to acknowledge this.

      (Q6) We used a higher MOI for bMECs because test experiments with WT PR8 and the Cattle Texas 6:2 reassortant virus showed that MOI 0.01 infections gave more variable results than those run at MOI 0.1, perhaps because of the intrinsic variability of mixed primary cell populations. However, the end-point titres between the two conditions were not significantly different, so we therefore chose to go with the higher MOI. Accordingly, we do not think this choice is a confounding issue. This explanation (line numbers 340-345) and a new Supplementary Figure 11 showing the results of the two MOI tests have been added to the manuscript.

      Reviewer #2 (Public review):

      The authors use a library of influenza A viruses from different strains, classified in lab-adapted, human, avian, and swine according to the animal from which they were isolated. They propose that the cow mammary gland serves as a mixing vessel for influenza A viruses. As a first approach, the authors assess susceptibility to infection across different cell types, including continuous and primary cell lines, bovine mammary cells, and mammary explants. All these cells support polymerase activity. Then, they analyzed changes in the bovine virus's viral fitness relative to an avian precursor. The authors use single-gene replacement to study whether and which RNP segments improve viral transcription. As part of this section, they also test IFN-specific antagonism by NS1 to assess the input of segment 8. Quantitative glycomic analysis was performed on the continuous bovine mammary cell line to demonstrate the presence of both a2,3 and a2,6, which is consistent with their observation that these cells can be co-infected with human and avian IAVs simultaneously. The main question, however, is: what is the glycome in the explants, or directly from tissues?

      We report quantitative glycomics for the primary bovine mammary epithelial cells as well as the continuous line the referee highlights. However, we agree with R2 that a detailed glycomic analysis of primary bovine mammary tissue would allow a better understanding of the actual glycosylation status in vivo. This has been undertaken by the authors and is available as a bioRxiv preprint. This is now cited (ref 25) in the relevant part of the results (line 184-185)

      Overall, the manuscript is clearly written and provides new insights into the behaviour of the cattle isolate, now compared with a representative group of model or precursor HAs of different origins.

      It would be great if a consistent nomenclature for the IAV strains could be used in the study. There is a mix of origin (Texas), animal from which the virus was isolated (mallard), or abbreviations that do not follow guidelines (IAV07). Are the USSR and Udorn not lab-adapted?

      We chose the abbreviated names for a variety of reasons. Partly from common usage (e.g. PR8, Udorn), partly for consistency with other already published papers from the FluTrailMap consortia (e.g. Cattle Texas; Dholakia et al 2026), partly to make diversity obvious in certain figures (e.g. H3N1, H5N2 etc) and partly to avoid confusion between viruses that originate from the same geographic area (e.g. AIV07, AIV09, H5N8-20 etc which are all A/Ck/England/isolate numbers). Overall, we found it more confusing to use the expanded nomenclature. Re AIV07 which the referee criticises for not following naming guidelines – if this is a reference to the EURL nomenclature, AIV07 is the abbreviation for the specific virus A/Chicken/England/053052/2021, our representative virus for EURL genotype EA-2020-C, as we say in the text. This nomenclature has now been added to Table 1, to provide a fuller cross-reference for all the names.

      As to whether USSR and Udorn are lab-adapted – that depends on definitions. There is a continuum of adaptive changes and/or sequence drift starting from the very first growth cycle of an isolate in the laboratory. The viruses we define here as lab adapted are ones that have been deliberately adapted to other host species or which have very long passage histories in multiple laboratory systems resulting in known functionally significant changes; for example, one lineage of PR8 was passaged 77 times in mice, 717 times in cell culture, 30 times in chick embryos, 5 times in ferrets and a further 50 times in chick embryos (https://www.medscape.com/viewarticle/812621_3?form=fpf), rendering it unarguably lab-adapted. We admit that A/USSR/77 and A/Udorn/307/1972 are probably further along this adaptive pathway than more recent isolates such as A/Norway/3433/2018, but are unaware of any specific reason that would put them into our lab-adapted category.

      The experimental setup includes bovine mammary primary and continuous cells, as well as mammary explants. Some of the most significant differences, for example, in viral fitness studies and co-infection experiments, are observed in these explants. Perhaps there could be some additional focus on this observation. The implications in comparison to the results obtained in cultured cells could be described. How will the human and other HA subtype viruses fare in the explants?

      We agree that this is an important and interesting question, and had already tested the strains we used for co-infections: human seasonal pdm09 H1N1 “Norway” and low pathogenic avian influenza “H3N1”, in the mammary explants. Both replicate the avian virus to 20-fold higher titres. We have added this information to the revised manuscript as new Figure panels S6E-H, called out on line 200-201 of the results.

      Reviewer #3 (Public review):

      Summary:

      This excellent manuscript by Pinto, Sharp, and colleagues examines bovine tissue tropism for influenza viruses. They find that bovine flu, as well as other strains, has strong replication in mammary tissue. They also map the genetic changes to influenza that improve replication in bovine cells. Overall, the study is well designed and executed, and the results are very timely.

      Strengths:

      (1) The experiments are well-controlled.

      (2) The figures are well-constructed and easy to follow.

      (3) The Methods and legends are detailed, with sufficient information.

      Weaknesses:

      (1) A comparison to human cells would strengthen the overall impact of the results. Are human mammary cells also uniquely susceptible to influenza? Are bovine mammary cells special in some way?

      This is an interesting question, but we have not tested mammary gland cells from humans (or any other species of mammal). We have however reported elsewhere (Dholakia et al., Nat Commun. 2026 Jan 16;17(1):1603. doi: 10.1038/s41467-026-68306-6.) that Cattle Texas grows well in a variety of human respiratory cells. Here, we are considering the bovine mammary organ as a potential reassortment site for IAVs because of the ongoing viral mastitis epidemic in US dairy cattle; human mammary organs seem unlikely to create a similar opportunity.

      (2) For the virus infection studies with segment 8 swaps, it should at least be noted that some of the phenotypes could be driven by NEP.

      We agree; as Table S1 indicates, NEP has two changes (one shared with NS1) between AIV07 and our B3.13 isolate, so we should not have conflated segment and NS1. We have changed the text to acknowledge this throughout the results (lines 127, 137 and 149) and in the discussion (line 243-244).

      (3) The data demonstrating that bMEC can support co-infection are compelling and important, but would be strengthened with a comparison from a different cell type or species. Do mammary cells uniquely support higher co-infection?

      We have data showing that co-infection also occurs in the continuous MAC-T udder cell line and have now included these data in a revised Figure 4D (described/called out on lines 198-206). We have not tested bovine cells from other organs for co-infection potential as they do not seem to be significant sites of infection in vivo.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) How nasal turbinate and cardiac fibroblasts are acquired/cultures is missing from the methods.

      Apologies for the omissions and thank you, because rectifying this brought to light an error in cell naming. The cells originally called bovine cardiac fibroblasts were in fact a second independent preparation of skin fibroblasts. We have corrected the labelling in Figs 1, 2 and S2. The nasal turbinate cells were bought in from ATCC (code CRL-1390). This information has now been added to the methods (line 310).

      (2) Please specify what cell types make up the 2D enteroids, the mammary explants and the nasal turbinates.

      The composition of the 2D enteroids is described in detail in reference [60]. In precis, they are comprised predominantly of epithelial cells, including Paneth, goblet and enteroendocrine cells, as well as stem cells. This information has been added to the methods (lines 374-376) The mammary explants include duct epithelium, connective and muscle tissue, defined by H&E staining of cut sections (see new Figure S12 and text added to the methods on lines 384-386). We have also added the person who did the histology (Rebecca Ross) as an author and to the credit taxonomy (line 877). The nasal turbinate cells appear to be predominantly fibroblast morphology (Methods line 310-311).

      (3) Epithelial cells are the main target for IAV, so why were fibroblasts chosen as a comparison to mammary epithelial cells? This needs justification in text. Without justification, how much can be attributed to the results being mammary-specific, rather than epithelial-specific? The brain (choroid plexus epithelium), heart (epicardium) and skin all contain epithelial cells.

      We think this query is what the referee calls “Q4’ in the public part of their review. Please see our answer above.

      (4) Figure 1, S1, S2 and S3 seem to suggest that mammary cells are more susceptible to IAV infection than cells from other organs. But Figure 2 demonstrates that when it comes to Cattle Texas and AIV07, most of the cell types show high viral titers. If the question is about whether the cow udder is the primary mixing site, would it not be more relevant to investigate which cell type facilitates the best growth of the potential precursor viruses, similar to Figure 4A?

      Figs 1, S1, 2 and 3 all use “full” viruses whereas Fig 2 uses 6:2 reassortants between PR8 and the HPAIVs for biosafety reasons. WT PR8 replicates well in most of the bovine cells tested (Figs S1-3) so we do not see any contradiction. Figure 2 examines the contributions the internal genes make to replication in bovine cells. Re the question over the udder being a potential mixing site – this is where the virus is replicating in the real world; at least in part because of the transmission mechanism, but also because the mammary gland epithelium is highly susceptible to infection, as we show.

      (5) It is misleading to compare the bMEC infection of MOI 0.1, to all other infections of MOI 0.01. This is a consistent problem throughout the article - Figures 1, S1, 2, 4. Why was the bMEC infection at a 10x greater dose? The main premise of the article relies on bMEC and MAC-T (primary and immortalised mammary epithelial cells), facilitating higher viral growth than the cells from other organs. If we compare the MOI 0.01 experiments alone, then the evidence relies on the immortalised MAC-T cells, compared to primary cell types. In this case, how much can be said about it being mammary specific, rather than immortalised vs primary? I do wonder, for example, how the epithelial nasal turbinates or type II pneumocytes would compare to the primary mammary epithelial cells if they were at the same MOI.

      Please see our answer to this query earlier in the rebuttal.

      (6) Why are the 2D enteroids excluded from Figure 1?

      We had limited supplies of a difficult-to-grow cell model, so we only used them to test the 6:2 viruses (Fig 2).

      (7) The colour scheme for Figure 3 is confusing. In Figure 3A, blue indicates European ancestry, and yellow represents North American ancestry. However, in Figure 3B, these colours now mean something different. To a reader, when there is a colour-coded schematic, it is instinctual to think that this then corresponds to the following panel(s). Since consistently throughout the article, yellow has been used for Cattle Texas, and blue has been used for AIV07, I would suggest choosing different colours to represent European and North American ancestry in Figure 3A

      We’ve changed the figure as the referee suggests and modified the Fig 3 legend accordingly (line 895)

      (8) I am unsure about the conclusions drawn from the results of Figure 3B. In the results, it is framed as trying to determine which segments contributed to the improved activity of Cattle Texas compared to AIV07. In lines 110-111, "... PB2 or PA from AIV07 significantly decreased Cattle Texas minireplicon activity". If PB2 is indeed significant, there is a missing yellow asterisk in Figure 3B.

      Apologies, there was indeed a missing asterisk on the figure; now added.

      Given the significance of PA, why was it not investigated in terms of growth kinetics similar to Figure 3C? Was it overlooked because it doesn't have a North American ancestry? The results of Figure 3B suggest that the 4 amino acid mutation in PA has significantly contributed to changes in polymerase activity.

      The PA changes do indeed matter for minireplicon activity – the key change is K497R, as detailed in our related publication in Nat Comms (citation 17). However, it is less important than changes in PB2, and the PA segment swap by itself has little effect on overall virus replication.

      (9) Similarly, in Figure 3C and lines 116-117, the error bars on the graph are overlapping at 48 hours, suggesting no difference in overall replication. The kinetics are slowed for AIV07 seg1-3, but not for AIV07 seg 1, indicating PB1 does not have an effect. This would then suggest that something in segment 2 or 3 contributes to the slowed kinetics in Figure 3C, which, from the Figure 3B results, is unlikely to be due to PB2. While it was not reassorted, the PA segment is potentially the driver, with its 4 amino acid mutations. I think it is worth performing growth kinetics with and without these 4 amino acid changes in PA.

      We agree that visually on a log<sub>10</sub> scale, the titres of the “WT” 6:2 Cattle Texas and 5:2:1 segment 1 reassortment appear close, but the average titres are 5 and 7-fold different at 24 and 48h respectively, while a 2-way ANOVA with Dunnet’s multiple comparison post-test gives statistical significance at 48h. We have added this information to the figure and its legend (lines 905-906).

      (10) In Figure 3C and E, why did you choose to perform the growth kinetics in the immortalised cell line, when you have access to primary cells? The primary cells would be a more accurate representation of what happens in situ.

      The primary cells were difficult to work with and only available intermittently, so we used what was available at the time.

      (11) In lines 145-147, "thus overall, the reassortment event that replaced segments 1, 2 and 8 alongside drift adaptations in segment 3 may have contributed to the ability of the B3.13 genotype virus to infect cattle". This is not clearly supported by the evidence presented. In terms of segment 1/PB2, the growth kinetics of Figure 3C have overlapping error bars at 48 hours. Where is any evidence presented for the role of segment 2/PB1? There is no change in Figure 3B.

      The referee is correct, calling out seg2 here was an error; we have revised the text (line 144).

      Segment 3 is overlooked in Figure 3 (as highlighted in Q9 and 10), and shows no difference in Figure S4.

      Please see response to Q8; we think segment 3 contributes via PA adaptation, not via PA-X.

      (12) In line 146 "... drift adaptation in segment 3". Make it clear here that you are talking about genetic drift. However, is this likely to be genetic drift? The 4 amino acid mutations are shown to have a significant impact on polymerase activity in Figure 3B, and in Figure 3C, PA potentially contributes to the reduced kinetics. When there are amino acid mutations that correspond to a beneficial phenotypic change, attributing this to drift alone rather than host adaptation is strange.

      Yes, wording clarified (line 145). “Drift” was used to distinguish it from reassortment but we agree this was an incorrect term in the context.

      (13) Figure 4A MAT-C cells: this is ostensibly the same experiment as Figure 3C in terms of the Cattle Texas and AIV07 viruses. If this is the case, how can you explain the difference in kinetics and overall titer? In Figure 3C, Cattle Texas reaches 10^6, and in Figure 4A it reaches almost 10^9. That's almost 3 log difference. Similarly, in Figure 3C Cattle Texas reaches 10^3, but in Figure 4A it reaches 10^6, a 3-log difference. At 24 hrs, they have roughly a 3-log difference between them in Figure 3C, but in Figure 4A this difference is much smaller. As far as I can tell, these are the same viruses, same dose and same cell model. The t0 titer is also vastly different between the two experiments.

      The experiments were done at different times (several months apart, so different cell passage numbers and/or serum batches) and by different people. We have no explanation other than biological variability. However, both groups of experiments include genuine biological replicates done over the course of 2-3 weeks, so in our view represent coherent tests within themselves.

      (14) In Figure 4, why weren't the a-2,6 a-2,3 proportions analysed for the explants and/or used for the co-infection experiment? It showed the greatest difference between the Cattle Texas and precursor viruses in Figure 4A.

      Our data on the proportion of 2,6 and 2,3 SA in bovine udder tissue are now available in a separate preprint (now cited as [25] in our MS). We did not use the explants for co-infection experiments because it would have been technically difficult to read out the outcome by flow cytometry.

      (15) Figure 4D requires a supplementary figure demonstrating the gating strategy, including one of the samples as an example.

      We have compiled a figure of this and added it as new Figure S13 (called out line 556).

      (16) In Figure 4D, why was an MOI of 5 chosen instead of the MOI of 0.01 used throughout the article for MAC-T cell infection? An MOI of 5 (so in a co-infection, a total of 10 virus particles per cell) is completely overloading the cells. At this dose, 10% of the cells were able to be co-infected, but how representative is this of a real co-infection scenario? While it demonstrates it is possible, it potentially remains highly unlikely, similar to the discussion around the swine respiratory tract in lines 235-240.

      We had also performed the co-infections at lower MOI (1) with very similar results – this is now included in Figure 4D. Furthermore, we redid the experiments at a lower MOI of 0.05 and still see co-infection; this now replaces the MOI 5 data in Figure 4D. The text has been revised accordingly (lines 202-205)

      (17) In Figure 4D, why were immortalised cells used when primary bMEC and mammary explants are available? Primary cells would provide more convincing evidence for the potential of the cow udder to be a mixing vessel. Considering that throughout the paper, a 10x higher viral dose is used in the bMEC culture, I wonder if you would need a significantly higher MOI than 5 to produce similar results in a co-infection experiment. The bMEC also has a more even a-2,3 to a-2,6 ratio compared to MAC-T in Figure 4C.

      The bMECs in the original figure are primary cells. In response to other queries, we now include data from the immortalised MAC-T cell line as well.

      Reviewer #2 (Recommendations for the authors):

      Figure 1A, the coloured underline to discriminate continuous and primary cells is lost upon printing... perhaps another way is better?

      We have changed the primary cells to italic text to make the distinction clearer.

      We have made some other minor changes to wording to correct grammatical errors or improve clarity as we went through.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary:

      This paper proposes a non-decision time (NDT)-informed approach to estimating timevarying decision thresholds in diffusion models of decision making. The manuscript motivates the method well, outlines the identifiability issues it is intended to address, and evaluates it using simulations and two empirical datasets. The aim is clear, the scope is deliberately focused, and the manuscript is well written. The core idea is interesting, technically grounded, and a meaningful contribution to ongoing work on collapsing thresholds.

      Strengths:

      The manuscript is logically structured and easy to follow. The emphasis on parameter recovery is appropriate and appreciated. The finding that the exponential NDT-informed function produces substantially better recovery than the hyperbolic form is useful, given the importance placed on identifiability earlier in the paper. The threshold visualisations are also helpful for interpreting what the models are doing. Overall, the work offers a well-defined, methodologically oriented contribution that will interest researchers working on time-varying thresholds.

      We appreciate the positive and constructive feedback. We have addressed your comments in the revised manuscript, as detailed in our responses to the specific comments below.

      Weaknesses / Areas for Clarification:

      A few points would benefit from clarification, additional analysis, or revised presentation:

      (1) It would help readers to see a concrete demonstration of the trade-off between NDT and collapsing thresholds, to give a sense of the scale of the identifiability problem motivating the work.

      Thank you for this constructive suggestion. We conducted a new simulation study in which we considered the non-decision time as a fixed parameter and estimated the remaining parameters of the collapsing threshold model. In this simulation study, we contaminated the non-decision time with noise at six levels (i.e., 0%, 2%, 4%, 6%, 8%, and 10%). The simulation results showed that increasing the noise level in non-decision time worsens the estimation of the starting threshold and decay rate parameters, indicating a trade-off between non-decision time and collapsing threshold parameters. The results of this simulation are presented in Appendix 1 in the new version of the manuscript.

      (2) Before moving to the empirical datasets, the manuscript really needs a simulation-based model recovery comparison, since all major conclusions of the empirical applications rely on model comparison. One approach might be to simulate from (a) an FT model with across-trial drift variability and (b) one of the CT models, then fit both models to each of the simulated data sets. This would address a longstanding issue: sometimes CT models are preferred even when the estimated collapse in the thresholds is close to zero. A recovery study would confirm that model selection behaves sensibly in the new framework.

      We are grateful for this constructive comment. The revised manuscript includes a model recovery study (see Simulation study 2, in particular Table 1 and the accompanying text). In this simulation, as the reviewer suggested, we generated data from FT-DDM with across-trial variability in drift rate and CT-DDM with hyperbolic and exponential collapsing thresholds. Then, for each generated dataset, we fitted two models and classified the datasets based on goodness-of-fit and estimated decay rate. The results show that incorporating the decay rate, which can be reliably estimated by the NDT-informed modeling framework, significantly improves the precision of model recovery.

      (3) An additional subtle point is that BIC is defined in terms of the maximised log-likelihood of the model for the data being modelled. In the joint model, the parameter estimates maximise the combined likelihood of behavioural and non-decision-time data. This means the behavioural log-likelihood evaluated at the joint MLEs is not the behavioural MLE. If BIC is being computed for the behavioural data only, this breaks the assumptions underlying BIC. The only valid BIC here would be one defined for the joint model using the joint likelihood.

      We thank the reviewer for raising this important methodological point. We agree that, strictly speaking, the behavioral log-likelihood evaluated at the joint maximum likelihood estimates is not guaranteed to be equal to the behavioral maximum likelihood estimate. Therefore, if one were to interpret BIC<sub>Behavior</sub> for joint (NDT-informed) models as a conventional BIC derived from a purely behavioral maximum likelihood fit, this would indeed violate the standard assumptions underlying BIC. Our intention, however, was not to claim that BIC<sub>Behavior</sub> represents a formally valid BIC in the strict information-theoretic sense. Rather, we used it as a diagnostic measure to assess how well the jointly estimated parameters account for the behavioral data relative to the uninformed models. Importantly, in our datasets, the behavioral log-likelihood evaluated at the joint estimates is higher than that obtained from the uninformed models. This suggests that the additional constraint introduced by the non-decision time information helps guide the optimization procedure toward parameter regions that provide a better account of the behavioral data. In other words, uninformed models appear to converge to suboptimal parameter estimates, but the joint modeling framework helps regularize the estimation process and yields better estimates of the optimal parameters. We fully acknowledge that this use of BIC<sub>Behavior</sub> for the joint models departs from standard practice in the cognitive modeling literature. For this reason, we rely primarily on the joint likelihood–based model comparison as the formally valid criterion. The behavioral BIC is reported only to provide additional intuition regarding the goodness of fit on behavioral data and is used exclusively to compare NDT-informed models with their uninformed counterparts under identical evaluation criteria.

      Also, to warn readers about this limitation, we included the following statement in the section where we defined BIC measures:

      “It is worth noting that comparing joint and behavioral models using BIC<sub>Behavior</sub> is uncommon and constitutes a limitation of the model comparison study.”

      (4) Table 1 sets up the Study 1 comparisons, but there’s no row for the FT model. Similarly, Figures 10 and 13 would be more informative if they included FT predictions. This matters because, in Study 1, the FT model appears to fit aggregate accuracy better than the BIC-preferred collapsing model, currently shown only in Appendix 5. Some discussion of why would strengthen the argument.

      In the revised manuscript, we have included the results for NDT-informed FT-DDM in the main text. However, we kept the FT-DDM with drift variability in the appendix, since none of the models in the main text include drift rate variability.

      (5) In Figure 7, the degree of decay underestimation is obscured by using a density plot rather than a scatterplot, consistent with the other panels of the same figure. Presenting it the same way would make the mis-recovery more transparent. The accompanying text may also need clarification: when data are generated from an FT model with across-trial drift variability, the NDT-informed model seems to infer FT boundaries essentially. If that’s correct, the model must be misfitting the simulated data. This is actually a useful result as it suggests across-trial drift variability in FT models is discriminable from collapsing-threshold models. It would be good to make this explicit.

      Regarding the visualization of decay rate estimation in Figure 7, we would like to clarify that, in the data-generating process for the FT model with across-trial drift variability, the decay rate is fixed at zero. That is, unlike the other parameters, there is only a single true value on the x-axis. If we were to present the decay rate recovery using a standard scatter plot (true vs. estimated values), all points would lie vertically above the single true value (zero), resulting in a vertical strip of points. While such a plot would technically be consistent with the other panels, it would not clearly convey the distributional properties of the estimated decay rates, specifically, which values are more likely under model misidentification. For this reason, we chose a density plot to more transparently illustrate the distribution of inferred decay rates when the true generating process is an FT model. We believe this representation more effectively communicates the extent and structure of mis-recovery. To avoid confusion, we have added a clarifying footnote in the revised manuscript explaining why a density plot was used in this specific panel.

      Regarding distinguishing fixed-threshold models from collapsing-threshold models, as suggested in your comment 2, we conducted an additional simulation, and the results showed that when using the NDT-informed diffusion model, we can distinguish between both models with high accuracy. Specifically, we showed that incorporating the estimated decay rate value provided by the NDT-informed modeling approach can significantly improve the model recovery accuracy.

      (6) Given the large recovery advantage of the exponential NDT-informed function over the hyperbolic one, the authors may want to consider whether the results favour adopting the former more generally. Given these findings, I would consider recommending the exponential NDT-informed model for future use.

      Consistent with the reviewer’s argument, we included the following text in the general discussion:

      “Importantly, the exponential collapsing threshold exhibited substantially better parameter recovery and superior model recovery performance, suggesting that this specification may be preferable in future cognitive modeling applications.”

      (7) In Study 2 (Figure 13), all models qualitatively miss an interesting empirical pattern: under speed emphasis, errors are faster than corrects, while under accuracy emphasis, errors become slower. The error RT distribution in the speed condition is especially poorly captured. It would be helpful for the authors to comment, as it suggests that something theoretically relevant is missing from all models tested.

      Thank you for mentioning this point. Because the models considered in the main text do not include across-trial variability in the starting point, the model cannot predict the fast error pattern observed in the speed condition of the second study. In the revised manuscript, we included a note on this point:

      “The NDT-informed models’ predictions are depicted in Figure 13. This figure shows that the FT-DDM overestimates the last RT quantiles for both correct and incorrect responses in both speed and accuracy conditions. However, the qualitative predictions of CT-DDMs align more closely with the empirical data. It is also worth noting that all the considered computational models misfit the incorrect responses in the speed condition. This misfit is to be linked to the presence of fast errors, specifically in the speed condition. Including the starting-point variability parameter in the model enables the model to predict fast errors. However, as the aim here was not merely to fit the data with the best possible model, but to test the NDT-informed modeling framework, we did not include starting-point variability in the model.”

      (8) The threshold visualisations extend to 3 seconds, yet both datasets show decisions mostly finishing by 1.5 seconds. Shortening the x-axis would better reflect the empirical RT distributions and avoid unintentionally overstating the timescale of the empirical decision processes.

      In the new version of the manuscript, we shortened the x-axis in these plots. See Figures 9 and 12 in the new version of the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript should explicitly state how critical the log-normal assumption for NDT is, and whether there are caveats if it doesn’t hold.

      We agree with this suggestion. Therefore, we conducted an additional simulation study in which we assumed a normal distribution for non-decision time measurements and replicated the main results reported in simulation study 1 (see Appendix 8 in the new version of the manuscript). These simulation results reveal that independent of distributional assumptions on non-decision time measurements, constraining non-decision time can improve the estimation of the collapsing threshold.

      (2) On page 14, one paragraph refers to five models and another to six. It is unclear which is correct.

      We are sorry for the confusion. In the new version of the manuscript, we addressed this issue.

      (3) In the Discussion: Hawkins & Heathcote (2021) found that NDT estimates in the TRDM recover well, but the timer-offset parameter does not (and hence is set to a fixed value). It would be interesting to test whether the NDT-informed approach could extend to that parameter. The TRDM can also generate error RT distributions that are faster than correct RT distributions, which is the qualitative pattern that none of the present models capture in the speed condition of Study 2.

      We appreciate the reviewer’s suggestion, as it would be another important demonstration of the framework we built in the paper. Nevertheless, given the number of models already considered in the manuscript and the focus on collapsing threshold diffusion models, we believe that this addition would blur the focus of the present manuscript. Therefore, we have decided not to include TRDM in the manuscript. Nevertheless, we have addressed the reviewer’s concern regarding the non-decision time estimation in the TRDM in the revised manuscript:

      “Notably, this model can also predict faster error responses than the correct response, the pattern that is observed in the speed condition of Study 2. However, the authors reported poor parameter recovery for the onset of the timing process (Hawkins and Heathcote, 2021). Thus, the reliability issue here specifically concerns the estimation of the shift parameter of the timing accumulator. Informing the model with external estimates of non-decision time might, therefore, improve parameter recovery in this model as well.”

      Reviewer #2 (Public review):

      Summary:

      The authors use simulations and empirical data fitting in order to demonstrate that informing a decision model on estimates of single-trial non-decision time can guide the model to more reliable parameter estimates, especially when the model has collapsing bounds.

      Strengths:

      The paper is well written and motivated, with clear depth of knowledge in the areas of neurophysiology of decision-making, sequential sampling models, and, in particular, the phenomenon of collapsing decision bounds.

      Two large-scale simulations are run to test parameter recovery, and two empirical datasets are fit and assessed; the fitting procedures themselves are state-of-the-art, and the study makes use of a very new and well-designed ERP decomposition algorithm that provides single-trial estimates of the duration of diffusion; the results provide inferences about the operation of decision bound collapse - all of this is impressive.

      We appreciate your feedback and comments. Below, we provided a response for each comment.

      Weaknesses:

      (1) This is an interesting and promising idea, but a very important issue is not clear: it is an intuitive principle that information from an external empirical source can enhance the reliability of parameter estimates for a given model, but how can the overall BIC improve, unless it is in fact a different model? Unfortunately, it is not clear whether and how the model structure itself differs between the NDTinformed and non-NDT-informed cases. Ideally, they are the same actual model, but with one getting extra guidance on where to place the tau and/or sigma parameters from external measurements. The absence of sigma (non-decision time variance) estimates for the non-NDT-informed model, however, suggests it is different in structure, not just in its lack of constraints. If they were the same model, whether they do or do not possess non-decision time variability (which is not currently clear), the only possible reason that the NDT-informed model could achieve better BIC is because the non-NDT-informed model gets lost in the fitting procedure and fails to find the global optimum. If they are in fact different models - for example, if the NDT-informed model is endowed with NDT variability, while the non-NDT-informed model is not - then the fit superiority doesn’t necessarily say anything about an NDT-informed reliability boost, but rather just that a model with NDT variability fits better than one without.

      To respond to this comment, we would like to note that the structural difference between NDT-informed and uninformed models is the assumption about non-decision time. In principle, the behavioral parts (i.e., parameters related to the diffusion part) of both NDT-informed and uninformed models are identical. However, during the estimation, the non-decision time in the NDT-informed model is subject to an additional constraint imposed by the neural data. In other words, in the uninformed model, non-decision time is estimated using the behavioral data by maximizing the likelihood of a CT-DDM. However, the NDT-informed model incorporates an additional data type, resulting in a different likelihood function (see Equation (3)). Specifically, in the NDT-informed model, we make an additional assumption regarding the non-decision time: it is set to the mean of the trial-level non-decision time measurements distribution derived from neural data (Equation (3) specifies the joint model structure). This additional assumption constrains the search space of the non-decision time parameter in the NDT-informed model. As in the main text, we assumed that non-decision time measurements obtained from neural data are log-normally distributed, where sigma is the shape parameter of the log-normal distribution. Therefore, sigma can represent the variability of the non-decision time measurements, and it is not the trial-to-trial non-decision time variability parameter. Indeed, the non-decision time parameter is fixed across trials in both models. Also, it should be clear from Equation (3) that sigma only appears in the second term of the joint likelihood, which corresponds to non-decision time measurements and not the CT-DDM term. Therefore, the behavioral parts of the NDT-informed models and Uninformed models are identical, and the NDT-informed models include one additional parameter (i.e., sigma) corresponding to the variability in non-decision time measurements.

      Also, to explain how constraining non-decision time improves the BIC, we would like to clarify that constraining non-decision time using an additional data source constrains the search space for non-decision time, thereby leading to better parameter identification and, consequently, a better fit to behavioral data. We included the following text in the discussion section to make this explicit.

      “This improvement likely reflects more accurate parameter estimation enabled by the additional information. In other words, constraining the non-decision time using neural measurements led the optimizer to estimate the CT-DDM parameter more accurately and, as a result, improve the fit to empirical data.”

      (2) One reason this is unclear is that Footnote 4 says that this study did not allow trial-to-trial variability in nondecision time, but the entire premise of using variable external single-trial estimates of nondecision times (illustrated in Figure 2) assumes there is nondecision time variability and that we have access to its distribution.

      We are sorry for the confusion. To respond to this comment, we would like to highlight that this modelling approach does not include any mechanism for across-trial variability in the behavioral part, as the likelihood of the choice behavior (see Equation (3)) does not include any variability parameter. As mentioned before, sigma represents the variability in non-decision measurements. However, the model does not propagate the across-trial variability of the neural data on the behavior side (unlike the models in Ghaderi-Kangavari et al. 2023). In other words, in this approach, none of the diffusion model’s parameters include across-trial variability, and we have only considered and estimated neural data variability in the NDT-informed model. Also, to improve the manuscript’s coherence, we removed this footnote.

      (3) It is good that there is an Intro section to explain how the tradeoff between NDT and collapsing bound parameters renders them difficult to simultaneously identify, but I think it needs more work to make it clear. First of all, it is not impossible to identify both, in the same way as, say, pre- and postdecisional nondecision time components cannot be resolved from behaviour alone - the intro had already talked about how collapsing bounds impact RT distribution shapes in specific ways, and obviously mean (or invariant) NDT can’t do that - it can only translate the whole distribution earlier/later on the time axis. This is at odds with the phrasing “one CANNOT estimate these three parameters simultaneously.” So it should be first clarified that this tradeoff is not absolute. Second, many readers will wonder if it is simply a matter of characterising the bound collapse time course as beginning at accumulation onset, instead of stimulus offset - does that not sidestep the issue? Third, assuming the above can be explained, and there is a reason to keep the collapse function aligned to stimulus onset, could the tradeoff be illustrated by picking two distinct sets of parameter values for non-decision time, starting threshold, and decay rate, which produce almost identical bound dynamics as a function of RT? It is not going to work for most readers to simply give the formula on line 211 and say ”There is a tradeoff.” Most readers will need more hand-holding.

      We are grateful for this comment. In response to this comment, which was also partly mentioned by the first reviewer, we first highlight that, in the presence of a nonlinear collapsing threshold, the effect of non-decision time is no longer linear, as it forces the threshold to take a specific value at the final stopping point. To illustrate the tradeoff between imprecise non-decision time estimation and collapsing threshold estimation, we conducted a simulation study. The results for the simulation study are presented in Appendix 1. In this study, we contaminated the non-decision time with noise at six levels (i.e., 0%,2%,4%,6%,8%, and 10%). The simulation results showed that increasing the noise level in non-decision time worsens the estimation of the starting threshold and decay rate parameters, indicating a trade-off between non-decision time and collapsing threshold parameters.

      Regarding the second point, we would like to clarify that in this paper, consistent with other works on collapsing threshold, we assumed that the collapsing starts with evidence accumulation and not with stimulus onset, as the decision makers need a short amount of time for perceiving and encoding the stimulus (i.e., perceptual encoding time). Fixing the collapsing onset to the stimulus onset introduces an ad hoc assumption into the model (that the encoding time is zero), which is cognitively implausible. Therefore, although this assumption can mathematically resolve the issue, it is not cognitively plausible. However, one important point we did not consider in the previous version is the two-stage accumulation process models in which the collapsing onset occurs later than the evidence-accumulation onset. We have discussed the estimation of such models as a limitation in the general discussion:

      “The collapsing threshold dynamics considered in this work (e.g., exponential and hyperbolic) impose a monotonically decreasing threshold over time. Although these dynamics are theoretically well motivated (Fudenberg et al., 2018; Frazier and Yu, 2007) and have been employed in several previous studies (e.g., Olschewski et al., 2025; Milosavljevic et al., 2010; Voskuilen et al., 2016), some research has proposed delayed-collapsing threshold models (e.g., Diederich and Oswald, 2016), in which the onset of threshold collapse does not coincide with the onset of evidence accumulation. Estimating such delayed-collapsing models may require more than simply constraining non-decision time, as the onset of threshold collapse must also be identified. A promising approach for addressing this challenge is the HMP method, which may provide additional temporal information about distinct cognitive processing stages. In particular, HMP may allow the onset of threshold collapse to be estimated as a separate cognitive stage. Future research should therefore investigate the estimation of multi-stage evidence accumulation models (e.g., Diederich and Oswald, 2016; Diederich and Colonius, 2021) within the HMP framework.”

      (4) A lognormal distribution is used as line 231 says it “must” produce a right-skew. Why? It is unusual for non-decision time distribution to be asymmetric in diffusion modeling, so this “must” statement must be fully explained and justified. Would I be right in saying that if either fixed or symmetrically distributed nondecision times were assumed, as in the majority of diffusion models, then the non-identifiability problem goes away? If the issue is one faced only by a special class of DDMs with lognormal NDT, this should be stated upfront.

      We would like to clarify that this assumption is about the non-decision time measurements and not about the across-trial variability parameter of non-decision time. Although for computational convenience, non-decision time is often assumed to follow a normal distribution in diffusion modelling literature, some empirical studies have shown that the perceptual encoding time and total approximated non-decision time follow a right-skewed distribution. Moreover, as HMP estimations of non-decision time are usually right-skewed, we formalized the model with a log-normal distribution. However, we would like to note that the identifiability issue in collapsing-threshold diffusion models is not related to the distributional assumption over the non-decision time measurements, as poor parameter recovery of the collapsing threshold was also reported in Evans et al. (2020). To show that the distributional assumption about the non-decision time measurements does not affect the results, we conducted an additional simulation study in which we assumed a normal distribution for non-decision time measurements and showed that constraining non-decision time improves parameter recovery for collapsing threshold parameters. The results for this simulation are presented in Appendix 8 in the new version of the manuscript. We also revised line 231 as follows (see line 236 in the new version):

      “Empirical studies on non-decision time measurement usually have reported a right-skewed distribution for their measurements (e.g., Weindel, 2021; Weindel et al., 2025). For instance, the measured perceptual encoding time and motor execution time reported by Weindel et al. (2025) are right-skewed. Therefore, we assume that the observed non-decision time measurements Z<sub>n</sub> follow an approximate log-normal distribution, which is right-skewed (we will discuss how this distributional assumption can affect the results later). Thus, we model the non-decision time measurements Z<sub>n</sub>, using a log-normal distribution with parameters µ and σ<sub>z</sub>. The available measurements (i.e., observed data) at trial n can be represented as

      follows:”

      (5) In the simulation study methods, is the only difference between NDT-informed and non-informed models that the non-NDT-informed must also estimate tau and sigma, whereas the NDT-informed model “knows” these two parameters and so only has the other three to estimate? And is it the exact same data that the two models are fit to, in each of the simulation runs? Why is sigma missing from the uninformed part of Figure 4? If it is nondecision time variability, shouldn’t the model at least be aware of the existence of sigma and try to estimate it, in order for this to be a meaningful comparison?

      As mentioned in the response to your first comment, the difference between NDT-informed and uninformed models is the access to an additional source of data related to non-decision time, and sigma represents the shape parameter of the non-decision time measurements distribution. Therefore, sigma belongs only to the NDT-informed model, and, as in the uninformed model, there is no additional data source, so the model does not include sigma.

      (6) I am curious to know whether a linear bound collapse suffers from the same identifiability issues with NDT, or was it not considered here because it is so suboptimal next to the hyperbolic/exponential?

      Thank you very much for this comment. The main reason we did not include linear collapsing threshold models was the assessment by Evans et al. (2020), which indicated that, with sufficient trials, these models can be estimated reasonably well. However, to investigate whether constraining non-decision time can also improve the estimation of linear collapsing threshold models, we conducted an additional simulation study, which is reported in Appendix 3 in the new version of the manuscript. The simulation results confirm those reported by Evans et al. (2020) and show that the parameters of the uninformed linear models can be identified using more than 500 trials. Constraining the non-decision time using the NDT-informed diffusion modeling framework still improves parameter estimation in the linear collapsing threshold model and reduces the required number of trials for reliable estimation to 250. Especially, the estimation of the starting threshold improves significantly.

      (7) The approach using HMP rests on the assumption that accumulation onset is marked by the peak of a certain neural event, but even if it is highly predictive of accumulation onset, depending on what it reflects, it could come systematically earlier or later than the actual accumulation onset. Could the authors comment on what implications this might have for the approach?

      Thank you for mentioning this point. We first would like to point out that Weindel et al. (2024) showed that HMP can predict the underlying generative distribution of cognitive states with high precision and without systematic bias. Second, it is worth highlighting the results reported in Appendix 5 (i.e., “Bias in non-decision time measurements”). In this appendix, we discussed how bias in non-decision time measurements (i.e., systematic underestimation or overestimation) affects the estimation of the collapsing threshold. Particularly, see Figure 3 in Appendix 5. To make these results clearer in the main text, we included the following paragraph at the end of the results section in simulation study 1:

      “Additionally, we examined the effect of systematic bias in non-decision time measurement on parameter recovery. Appendix 5 presents the parameter recovery simulation results in the presence of biased non-decision time measurement (i.e., systematically underestimated or overestimated). The results revealed that, even in the presence of biased non-decision time measurement, the actual generating parameters show a high correlation with the estimated parameters. Underestimation in non-decision time leads to overestimation in the starting threshold and decay rate. Conversely, overestimation in non-decision time leads to underestimation in both the starting threshold and non-decision time.”

      (8) Figure 7: for this simulation, it would be helpful to know the degree to which you can get away with not equipping the model to capture drift rate variability, when the degree of that d.r. variability actually produces appreciable slow error rates. The approach here is to sample uniformly from ranges of the parameters, but how many of these produce data that can be reasonably recognised as similar to human behaviour on typical perceptual decision tasks? The authors point out that only 5% of fits estimate an appreciable bound collapse but if there are only 10% of the parameter vectors that produce data in a typical RT range with typical error rates etc, and half of these produce an appreciable downturn in accuracy for slower RT, and all of the latter represent that 5%, then that’s quite a different story. An easy fix would be to plot estimated decay as a scatter plot against the rate of decline of accuracy from the median RT to the slowest RT, to visualise the degree to which slow errors can be absorbed by the no-dr-var model without falsely estimating steep bound collapse. In general, I’m not so sure of the value of this section, since, in principle, there is no getting around the fact that if what is in truth a drift-variability source of slow errors is fit with a model that can only capture it with a collapsing bound, it will estimate a collapsing bound, or just fail to capture those slow errors.

      Thank you for this comment. We would like to first note that the aim of this section is to illustrate that NDT-informed modeling enables us to distinguish between CT-DDM and FT-DDM with drift variability (i.e., the two competing models that are relatively hard to distinguish). Therefore, we changed the name of this section to “Simulation study 2: Model recovery”. In the new version of the manuscript, we included a cross-model fitting simulation and a model recovery simulation in this section. The estimated decay rates in the cross-model fitting study indicate that, when the underlying generating model is FT-DDM with drift variability, parameter estimation using the NDT-informed approach yields precise inference. This is due to the estimated decay rate, which is very close to zero. The model recovery results also confirmed that incorporating the decay rate value into the model inference improves precision.

      Moreover, to address the concern regarding the slow-error pattern in the simulated data, we examined the relationship between the slow–fast accuracy difference and the estimated decay rate. Specifically, we computed the difference between the accuracy of responses with response times below the median (ACC1) and those above the median (ACC2), and plotted this difference against the estimated decay rate. Author response image 1 presents the resulting scatter plot, with color indicating the drift rate. As shown in Author response image 1, incorrect inferences about the decay rate primarily occur at high drift rates. This pattern emerges because high drift rates produce very fast responses, leaving little time for the threshold to meaningfully collapse. Consequently, the behavioral signatures of collapsing-threshold and fixed-threshold models become increasingly similar under high drift conditions.

      Author response image 1.

      Illustration of the difference between the accuracies of responses with response time below the median and those with response time above the median against the estimated decay rate. Colour shows the drift rate value.

      Reviewer #2 (Recommendations for the authors):

      (1) Abstract: improves fit to behaviour in what way? Reliability or absolute quant fit to behavior? I.e., is it just helping constrain it so it finds the global opt?

      (2) Line 72 - Are these “neuroimaging” studies? Perhaps use the broader term “neuroscience.”

      (3) Line 95 - It’s important because it may confuse readers how something dynamic like a collapsing bound could be resolved with, say, fMRI.

      (4) Line 85 – “greater variability”... in what?

      (5) Line 87 -Revise to avoid misconstruing a time on task effect - e.g., “higher error rate for trials with longer RT” is more explicit.

      (6) Line 89-94 - It’s not clear what findings are being referenced here.

      (7) Line 103 - External biases such as priors and relative value?

      (8) Line 175 - Clarify this applies to any ddm, not just ctddm.

      (9) Line 268 - Explain what portion N200 accounts for.

      (10) Line 327 - Is this equation supposed to be for Delta-X, as opposed to X(t+delta-t)? If you want X(t+delta-t) on the LHS, then X(t) must be added to the RHS.

      We are very thankful for such a precise evaluation and constructive comments. We addressed all the comments raised by the reviewer in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The current paper addresses an important issue in evidence accumulation models: many modelers implement flat decision boundaries because the collapsing alternatives are hard to reliably estimate. Here, using simulations, the authors demonstrate that parameter recovery can be drastically improved by providing the model with additional data (specifically, an EEG-informed estimate of nondecision time). Moreover, in two empirical datasets, it is shown that those EEG-informed models provide a better fit to the data. The method seems sound and promising and might inform future work on the debate regarding flat vs collapsing choice boundaries. As an evidence-accumulation enthusiast, I am quite excited about this work, although for a broader audience, the immediate applicability of this approach seems limited because it does require EEG data (i.e., limiting widespread use of the method or e.g., answering questions about individual differences that require a very large N).

      We are very grateful for your positive evaluation and your comments.

      Reviewer #3 (Recommendations for the authors):

      This is a very decent study, very well written and properly executed. Most of my comments below are suggestions for the authors as to how to make the story more compelling.

      (1) I think the authors can do more to explain why the NDT-informed models fit better to empirical data. If the NDT-estimates are equal to the ground truth, then isn’t it possible that such models win simply because they have one parameter less? However, this is not what the authors are claiming, though, on l.567 it says that the fit is better because parameters are better estimated. I think it should be possible to dissociate these two accounts.

      Thank you for this important comment. For clarification, both NDT-informed and uninformed models estimate the non-decision time parameter, and neither treats it as equal to the ground truth. However, the difference between these two models is that, in the NDT-informed model, the non-decision time is subject to an additional constraint; therefore, the number of behavioral parameters (i.e., parameters of the diffusion part) is identical for both NDT-informed and uninformed collapsing threshold models. Although the number of behavioral parameters is identical in the NDT-informed model, one parameter is constrained, thereby limiting the model’s complexity and flexibility compared to uninformed models. Consequently, the improvement in the fit can only be attributed to better parameter estimation in the NDT-informed models, resulting from constraining the non-decision time using neural measurements. To clarify that in the manuscript, we included the following text in the revised version:

      “This improvement likely reflects more accurate parameter estimation enabled by the additional information. In other words, constraining the non-decision time using neural measurements led the optimizer to estimate the CT-DDM parameter more accurately and, as a result, improve the fit to empirical data.”

      (2) Given that so much of the writing focuses on the conclusion that collapsing boundary models ¿ flat models, it is very odd that there is no comparison to a flat boundary model reported in the text. Why is the model in Supplement 5 not just included in the main text (and in Table 1)? This would make it so much easier for the reader.

      To address this comment and also the similar point raised by Reviewer 1, we included the NDT-informed FT-DDM in the results section and compared the FT-DDM with CT-DDMs with respect to BIC<sub>Joint</sub>. However, we retained the other FT-DDM in the appendix because this model includes across-trial variability in drift rate, whereas the models in the main text do not.

      (3) I would have appreciated a bit more background about the importance of the number of trials per participant. Given that the proposed method requires collecting EEG data (which is time and labour-intensive) I wonder to what extent you get similar improvements in parameter recovery by collecting more data per participant (which is usually cheap). Put differently, I would appreciate an additional matrix in Figure 5 for vanilla models.

      The new version of Figure 5 in the revised manuscript now contains the goodness of parameter recovery for the uninformed CT-DDMs for different numbers of trials. As the simulation results suggest, even with 1000 trials, the parameters of the uninformed CT-DDMs are still not reliably identifiable. Therefore, these results suggest that although increasing the number of trials can slightly improve the parameter recovery of the CT-DDMs, it cannot fully resolve the reliability issue in their parameter recovery.

      In addition, it is worth clarifying that, although the paper focuses on extracting non-decision time from the EEG signal using the HMP method, as discussed in the general discussion, the NDT-informed approach can also be employed with purely behavioral methods.

      (4) L116-117: minor detail: I don’t think it’s fair to write that it’s an open question whether or not thresholds collapse. I think it’s fair to say that this is hard to show, and that the conditions under which it appears are unclear; but saying that it’s still unclear whether this occurs at all seems unfair with regard to previous work.

      Thank you for mentioning this point. We revised the mentioned sentence as follows in the new version of the manuscript:

      “This issue is particularly critical given that the conditions under which individuals adjust their decision thresholds during a single trial remain an open question.”

      (5) Figure 5: Minor detail: It would be useful to mention in the figure or caption how many trials underlie these simulations, which is now somewhat buried in the text.

      As Figure 5 shows sensitivity to the number of trials, and the number of trials is explicitly mentioned in the figure, we assume that the reviewer intended Figure 6. In the revised manuscript, we explicitly specify the number of trials for the results reported in Figure 6. We simulated 500 trials for each noise level in this graph.

      (6) Figure 10 and similar figures: I tried to figure out how corrects and incorrects differ, but couldn’t see the difference. Can the authors use something more colorblind friendly (hope I didn’t give up on being anonymous here)?

      We are so sorry for the inconvenience. In the revised manuscript, we used different symbols to distinguish correct and incorrect data.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study investigates the peptide-binding principles of promiscuous chicken MHC molecules. The data from crystallography, mass spectrometry, and modeling are convincing. However, the presentation would benefit from streamlining and clear links between data and conclusions. This paper will be of broad interest to immunologists and those interested in vaccine development.

      Overall, we are delighted and grateful to the eLIFE editors and the two reviewers for the careful and thoughtful assessments and reviews of our paper. We are glad that the strengths of the paper were apparent and appreciated. And of course, every paper has weaknesses, especially for a story as complex as this one.

      We made only minor changes to accommodate the reviewer comments, along with additions for which we only became aware upon this submission of a revised manuscript. In particular, we shortened the title and abstract to fit what is usual for an eLIFE paper, added Key Resources table with accompanying references, changed the numbering of the figures throughout the manuscript to ensure that each page represented a figure (rather than panels of a figure), moved the figure legends from the embedded figures to a list near the end of the manuscript, and split the supplemental spreadsheet into two renamed Data Source files.

      Before answering the comments and questions directly, perhaps a few points would help clarify why the paper is as it is.

      First, the experiments cover over three decades of work, with the first gas phase sequencing results done in 1992. Unlike some of the chicken class I alleles which immediately gave completely clear stringent motifs (B4, B12 and B15 in Wallny et al 2006 PNAS, B19 in Han et al 2023 J Immunol), we harvested nothing but confusion from the B21 class I results (Fig. 1). Initially, we thought that the lack of a clear motif for B21 was due to multiple well-expressed class I molecules but only one dominantly-expressed class I molecule was found (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol) and, to our surprise, bacterially-expressed BF2*21:01 heavy chain and b<sub>2</sub>-microglobulin refolded with two synthetic peptides without sequence in common, and the crystal structures showed that this molecule remodeled the binding site to accommodate two such disparate peptides (Koch et al 2008 Immunity). This was the beginning of our understanding of the spectrum of class I alleles from promiscuous generalists to fastidious specialists, which we have explored in a series of further papers (in particular, Chappell et al 2015 eLIFE, Tresgaskes et al 2016 PNAS, Kaufman 2018 Trends Immunol, Tregaskes and Kaufman 2022 Mol Immunol).

      Second, over these many years, we continued to explore the binding properties of BF2*21:01 in ever more detail, resulting in the current manuscript. We learned only slowly how to probe this unexpected promiscuity, unprecedented in the MHC literature, so that the experiments proceeded with our best understanding at the time, including taking advantage of new approaches as they become available. Each experiment built on the previous set of experiments and each brought us closer to an understanding.

      Third, having amassed a collection of data, we chose eLIFE exactly because it allows us to present the entire story from beginning to end without compromise, not just the highlights with the major points illustrated by a few main figures and with the supporting data in many supplementary figures. We include all the data, because it is all part of the story, and so interested researchers to look at the data from their own perspective. Although mostly we provide bar graphs, we include spreadsheets for the raw data (or close to them) for the final experiments (illustrated by Figs. 10 and 14-22) in the two source data files, so these can be assessed easily by others in the field, perhaps using approaches that we may not feel competent to perform.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Combining in vitro refolding, SEC-based assembly assays, peptide-library screening, MALDI-TOF, LC-MS/MS, structural analysis and immunopeptidomics, this manuscript investigates the peptide-binding principles of the promiscuous chicken MHC-I molecule BF2*21:01.

      Strengths:

      Although the peptide motif of BF2*21:01 is highly complex, this manuscript identified several principles, including a preference for 10-mer peptides, co-variation between P2 and Pc-2, effects of P3 and Pc-3, and a strong cellular preference for Leu at Pc. The results are important for avian MHC biology and poultry vaccine epitope prediction.

      Weaknesses:

      The manuscript is sometimes difficult to follow because the authors present a large amount of peptide-library, structural and immunopeptidomics data. without always clearly explaining how these datasets support the proposed simplifying principles.

      We are delighted and grateful to the reviewer 1 for the careful and thoughtful comments and questions concerning our manuscript. We are glad that the strengths of the paper were apparent and appreciated, and acknowledge the weaknesses that come with such a complex story with experiments performed over decades.

      Major Issues - Points Requiring Clarification or Additional Support:

      (1) Line 282-301, 537-545)

      The immunopeptidomics conclusions are mainly based on one B21 cell line with one biological replicate and at least two technical replicates. Given the complexity of the BF2*21:01 peptide repertoire, this is a major limitation. The authors should either provide additional biological replicates or clearly state this limitation in the Abstract, Results and Discussion.

      This limitation is clearly stated in lines 537-545, as part of a paragraph covering the various ways in which the data presented in this manuscript could be improved. In fact, we have performed immunopeptidomics of several different B21 cell types, with many replicates and found similar data as presented, giving us confidence in our interpretations. However, these other experiments belong in different stories, so it is not appropriate that the data be reported in this manuscript.

      (2) (Lines 290-313)

      The B21 cell preparations contain both BF2 and the lowly expressed BF1 molecule. Some peptides, especially 8-mers or peptides with atypical motifs, may derive from BF1*21:01. The authors should clarify how BF2*21:01-bound peptides were distinguished from possible BF1-derived peptides, or interpret the immunopeptidomics motif more cautiously. The authors should also provide or cite evidence confirming the B21 haplotype identity of the cell line and chicken materials used for immunopeptidomics.

      The concern about the contribution of BF1*21:01 to the immunopeptidomics is clearly stated in the manuscript, both lines 290-313 and as part of the paragraph describing the limitations of the experiments (lines 542-543). In fact, the expression of BF1 molecules has long been known to be less than 10% of BF2 molecules at the RNA level, and much less at the protein level (Wallny et al 2006 PNAS, Shaw et al 2007 J Immunol). The proportion of 8mers identified by immunopeptidomics is also low (Fig. 14), and it is not impossible that most 8mers are due to BF1*21:01. We have used assembly assays with peptide libraries, immunopeptidomics and a crystal structure to determine the peptide motif for typical BF1 molecules, of which BF1*21:01 is one and found it may contribute to 8mer peptides but very seldom to longer peptides. This work is unpublished but gives us confidence that the characteristics of BF2*21:01 are not misrepresented by the data in this manuscript.

      The sources of the chicken samples and the cell lines are described in detail under Materials and Methods (lines 577-590), citing relevant publications.

      (3) (Lines 217-221, 243-253)

      The authors acknowledge that MALDI-TOF cannot reliably distinguish peptide combinations with identical or similar masses, nor determine residue positions in some cases. Therefore, MALDI-TOF results should not be overinterpreted as precise evidence for residue preference. The authors should clearly indicate which conclusions are supported by LC-MS/MS.

      As described, the experiments follow each other in temporal sequence, so that we started with single peptides, then peptide libraries that varied in one position, then peptide libraries that varied in two positions first analysed by MALDI-TOF and later by LC-MS/MS. The final experiment (Fig. 10, with the original data in the supplementary spreadsheet) directly compares MALDI-TOF and LC-MS/MS results for six peptide libraries, so that the strength of the evidence for residue preference is clear. Throughout the manuscript, we do our best to not to overstate conclusions based on the data of any particular experiment.

      (4) (Lines 297-301, 316-330)

      The authors suggest that longer peptides may bulge in the middle or extend out of the groove at the C-terminal end. The rationale for the C-terminal extension is not clearly explained. Why is the C-terminal extension considered rather than the N-terminal extension? If the binding register is uncertain, long peptides should be analyzed separately from canonical-length peptides.

      When the first sequence of a chicken class I cDNA was determined, an immediate mystery was why one of the so-called invariant residues that coordinate the N- and C-termini of the bound peptide is not conserved (Kaufman et al 1992 J Immunol). In fact, this residue Tyr at position 86 in HLA-A2 and the equivalent position in all mammalian classical class I molecules is an Arg in the classical class I molecules of all non-mammalian vertebrates and is common with class II molecules (Kaufman et al 1995 Semin Immunol). Similar to class II molecules, this Arg in chicken class I molecules allows the peptide to extend out of the C-terminus, as shown by a crystal structure (Xiao et al 2018 J Immunol). The concern that we might be misidentifying the C-terminal amino acid was the basis for the analysis in Figs. 23 and 24, but in the absence of crystal structures, we are not able to provide a final answer this question. Perhaps relevant is the fact that a chicken class II molecule can bind exactly the same peptide in two conformations, one with a canonical 9mer core and the other with an unexpected 10mer core (Goryanin et al 2026 J Virol).

      By contrast, N-terminal extensions are only found for some class I alleles and thus far depend on the substitution of small amino acid sidechains for W166 (Li et al 2011 J Virol for bovine, Ma et al 2020 J Immunol for Xenopus, Wei et al 2022 J Immunol for ovine). Thus far, no chicken BF2 sequences have this substitution, consonant with the many crystal structures, including those for BF2*21:01 (Koch et al 2008 Immunity, Chappell et al 2015 eLIFE, this manuscript). However, in unpublished data, we find that most BF1 sequences have sequence differences that could allow N-terminal extensions, although we have no crystal structures to support this possibility.

      (5) (Lines 406-439)

      In vitro assembly assays show that several hydrophobic residues can be tolerated at Pc, whereas immunopeptidomics shows a strong Leu preference at this position. The authors should clarify whether this Leu preference reflects intrinsic BF2*21:01 binding specificity, TAP-mediated peptide transport, antigen processing, peptide loading, or a cell-line-specific effect. Additional experimental support, such as TAP transport analysis, would strengthen this conclusion.

      The preference for Leu at the final position of the peptide by immunopeptidomics of the B21 cell line is strong but not absolute and is certainly affected at the least by the length of the peptide (Figs. 23 and 24). Unpublished immunopeptidomics results (mentioned above) show that this is not a cell line-specific result. The evidence from assembly assays of various peptides is that several hydrophobic amino acids are tolerated with sufficient stability of BF2*21:01 that they are detected in the assay (Figs. 3, 5, 9 and 10). Thermostability assays (Fig. 6) show that peptides with these same hydrophobic amino acids are stable to at least body temperature of chickens. These experiments show that such stability is peptide-dependent (that is, whether a particular amino acid is tolerated depends on the stability conferred by the rest of the peptide). Finally, peptide translocation assays using B21 cells have been done (Tregaskes et al 2016 PNAS) and show that peptides with several hydrophobic amino acids can be pumped into the lumen of the endoplasmic reticulum. However, the assays are with single synthetic peptides, so the data are not extensive enough to separate the effects of the final amino acid from the rest of the peptide. Certainly, peptides with amino acids other than Leu at the C-terminus can be translocated. So, it is not yet clear at which point the preference for Leu at the C-terminus of the peptide arises.

      (6) (Lines 172-178, 243-279, 442-457)

      The structural analysis explains some residue combinations, such as Arg at P2 with Glu at Pc-2 or Trp at Pc. However, the structural interpretation is not fully integrated with the large-scale peptide library and immunopeptidomics results. Representative high- and low-frequency combinations should be discussed structurally.

      Six crystal structures show that BF2*21:02 remodels the binding to accommodate a variety of anchor residues (Koch et al 2008 Immunity, Chappel et al 2015 eLIFE). These crystal structures are representative of sequences found by the immunopeptidomics from very frequent (H-E at roughly 15% 8-12mers) to moderately frequent (E-L at roughly 6% 8-12mers) to infrequent (N-F, A-D and E-D at roughly 1.5%, 1.6% and 0.7% 8-12mers) based on Fig. 18. All but one of the structures has Leu at the C-terminus, with the last one having Val which is found but not frequently by immunopeptidomics.

      Similar numbers are found by LC-MS/MS of double-substitution libraries of the two original peptide sequences in Fig. 10 with H-E found frequently (8.1% in P390, 3.8% in P498) and the others infrequently (0.1, 0.9, 1.0, 0.3% in P390, 0, 1.4, 1.0, 0.3% in P498), as calculated from the numbers in the Supplementary data spreadsheet. As discussed in the manuscript, for single-substitution peptide libraries of the two original peptides, Ile/Leu at the C-terminus was very frequent but at the same or slightly less level as Phe, with Met less frequent and Val even less so (Fig. 7).

      In addition, there are two more structures along with models explicitly testing some substitutions (Fig. 5). Attempting more current modelling approaches, we found AlphaFold 3 was unable to correctly predict most of the conformations that are found in the crystal structures of BF2*21:01, so we don’t feel confident in using them to predict unknown structures of this kind.

      (7) The inference of co-variation between P2 and Pc-2, as well as the modulatory effects of P3 and Pc-3, should be better explained. At present, some conclusions appear to be based mainly on residue-frequency patterns, and the logical connection between these observations and the proposed binding principles is not always clear. Statistical analyses, such as mutual information, chi-square tests or permutation tests, and representative structural explanations would strengthen this conclusion.

      We endeavored to do our best to explain the data, our interpretations and our reasoning, so we apologise if we have not managed to be as clear as might be desired. We have included as close to raw data as possible for the LC-MS/MS and MALDI-TOF (Fig. 10) and for the immunopeptidomics (Fig. 14 and 18) in the Supplementary Data spreadsheet, exactly so that competent practitioners can carry out further analyses (including the sophisticated statistical tests mentioned).

      Reviewer #2 (Public review):

      Summary:

      The study presents an in-depth analysis of the peptide repertoire bound by a promiscuous chicken MHC molecule using mass spectrometry, x-ray crystallography and modelling. While the MHC can bind a very diverse set of peptides, the authors have found some new rules that govern peptide binding to this MHC that could help to build a predictive model to study the repertoire of pathogen-derived peptides.

      Strengths:

      The study uses a range of well performed experiment across multiple techniques and provides an in-depth analysis of the peptide repertoire, including peptide sequences, length, preferred residues, stability and MHC presentation.

      Weaknesses:

      The data overall support the analysis and conclusion well. The only caveat is linked to Figure 4, which does not describe the stability of the peptide-MHC complex, but instead shows refold yield, and the two are not always linked.

      We are grateful for the clear understanding of the strengths of the work. With regards to Fig. 4, we agree with the reviewer that there are differences in refold yield but that measure may not be correlated with stability of the peptide-MHC complex. However, we were basing our interpretation of stability on the position and quality of the monomer peak, as illustrated by the trace in Fig. 2, in which a sharp peak at the monomer position represents a stable complex (as seen for the 10 and 11mer peptides) and later peaks represent unstable complexes falling apart during the chromatography (as seen for the 7, 8 and 9mer peptides).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Issues: Editorial and Data Presentation Modifications

      (1) Lines 53-62, 155-170, 303-314: The terms Pc, Pc-2 and Pc-3 should be clearly defined early in the manuscript and figure legends.

      The Abstract introduces the abbreviations as “peptide positions P<sub>2</sub> and P<sub>c-2</sub>” followed by P<sub>C</sub>, which are standard usage and seem clear. The first usage in the text is “anchor residues at three positions, with co-variation between the anchor residues at P<sub>2</sub> and P<sub>c-2</sub>” which again seems clear, particularly in the context of the text and the figure. However, a parenthetical description has been added to read “…anchor residues at three positions, with co-variation between the anchor residues at P<sub>2</sub> and P<sub>c-2</sub> (position 2 and the position two before the C-terminus) …”. Given these usages, it seems unlikely that the reader will fail to understand P<sub>C</sub> and P<sub>C-3</sub>.

      (2) Lines 255-279: The term "peptide backbone" should be defined as the fixed sequence context outside the randomized positions, if this is what the authors mean. Suggest clarifying the meaning of "peptide backbone".

      The phrase reads “double-substitution libraries based on four peptide backbones”, which in context of the figures seems clear. However, a parenthetical description has been added: “The examination of double-substitution libraries analysed by MALDI-TOF was expanded to other backbones (that is, other sequences in which two positions were randomised): …”.

      (3) Several figures are complex. The authors should add brief take-home messages to figure legends.

      Every figure legend starts with a (sometimes quite long) take-home message. It is not clear what more should be added.

      (4) A concise summary table of the proposed binding rules, including preferred peptide length, P2, Pc-2 and Pc preferences, and the effects of P3/Pc-3, would be useful for readers.

      A concise summary table of binding rules would be very helpful, but the rules are complex, both qualitative and quantitative. For example, immunopeptidomics shows that 10mers are preferred, but that fails to capture the quantitation. The co-variation of P<sub>2</sub> and P<sub>c-2</sub>, which is the strongest and best characterized correlation, nevertheless is complex since the occupancy of P<sub>2</sub> by different amino acids (presumably independently of P<sub>c-2</sub>) varies considerably. At our current level of understanding, it is hard to imagine a table that is both concise and precise. With time and more data, perhaps code for quantitative prediction might be constructed (dare one suggest machine learning…), which is the hope of presenting all the available data in one place.

      (5) The Results section contains a large amount of peptide-library, structural and immunopeptidomics data. The authors should improve the logical flow and add clearer transition sentences to explain how each dataset supports the proposed simplifying principles.

      The reviewer is of course correct that any written text can be improved (although each critic may have a different idea about which part should be fixed), but we have done our best with the material and time available. We could respond productively to a more detailed critique.

      (6) Some statements in the Abstract and Discussion should be softened, especially those related to in vivo peptide preferences, BF2*21:01 promiscuity and peptide prediction, given the limited biological replication and uncertainty in peptide assignment.

      The statements in the both the Abstract and Discussion are very general, reflecting what we believe to be careful interpretations based on the data. We could respond productively to concerns about specific claims.

      Reviewer #2 (Recommendations for the authors):

      Overall, the data presented in this study are interesting; however, it is complex, and some results could be merged and simplified, as well as the figures. The data provide an in-depth analysis, using mass spectrometry, of the interplay between the different positions of the peptide and the residues favourable to bind within the antigen-binding cleft.

      (1) From the abstract, the concept of "promiscuous generalists and fastidious specialists" is not explored after or defined within the results.

      The Abstract introduces the concept of promiscuous generalists and fastidious specialists to provide the basis for exploring the peptide-binding specificity of the most promiscuous class I known, BF2*21:01. This overall concept is described in enormous detail in several publications cited in the current manuscript, but it is not particularly germane to the analyses.

      (2) From the abstract "These simplifying principles may eventually allow predictions of pathogen peptides", I'm not sure how "simplifying" the principles are with the data, if anything, it does show a rather complex interplay between the different residues of the peptide that enable the MHC to bind a large and diverse number of peptides.

      Compared to any combination of anchor residues being permitted at equal frequency, there are clear preferences which the experiments identify. Instead of an enormous range of possible peptide lengths, roughly 50% of peptides are 10mers. Of course, the structural reasons behind these results are certainly complex and likely must be understood in detail in order to attempt peptide predictions. 

      (3) Line 48. "less well-expressed". Less than what? Do you mean the level of expression was lower? And if yes, are there values of fold change for comparison?

      The sentence in the Abstract reads “Chicken BF2 alleles … are less well-expressed on the cell surface … while certain human HLA-B alleles … are well-expressed …”. Read as a complete sentence, the meaning is clear. This difference for chicken BF2 alleles has been quantified as reported in several publications cited in the current manuscript, ranging from 3-5 fold for peripheral blood lymphocytes to ten-fold for erythrocytes (Kaufman et al 1995 Immunol Rev, Chappel et al 2015 eLIFE), with similar numbers for a few HLA-B alleles on human peripheral blood lymphocytes and monocytes (Chappell et al 2015 eLIFE).

      (4) Line 149. "with individual peptides" which peptides are we referring to here?

      This introductory sentence to a paragraph outlines the general method used for the experiments in this section of the Results, “refolding in vitro … with individual peptides.” Which “individual peptides” are described in the following paragraphs, with the next section of the Results using “refolding in vitro … with peptide libraries”.

      (5) Line 170. If 9 mers and below are not stable, which is not really quantified or shown with the data on Figure 4, why is refolding material observed for peptides with different lengths of 9 aa and below on Figure 4? A lower yield of refolded material can have a different origin, and there is no association between stability and refold yield. The notion of stability here probably needs to be changed, as it does not apply to the data.

      As described in our response to a similar concern above, we are not basing our interpretation of stability on the quantity of refolded material, but on the position and quality of the monomer peak, as described clearly in the legend to Fig. 4: “The original 10mer (REVDEQLLSV) and 11mer (GHAEEYGAETL) peptides refold with BF2*21:01 to give stable monomers as do 11mer and 10mer derivative peptides, but 9mer, 8mer or 7mer peptides give heavy chain only peaks” and “The peptides 11mer GHAEEYAETL (top panel), 10mer REVDEQLLSV (middle panel), 11mer GHAEAAAAETL and 10mer GHAEAAAETL (bottom panel) gave sharp monomer peaks, while the 9mer GHAEAAETL, 8mer GHAEAETL and 7mer GHAEETL gave a delayed broad peak indicative of unstable binding or heavy chain.” This concept is illustrated by the trace in Fig. 2 (discussed in the text at the beginning of this section of the Results), in which a sharp peak at the monomer position represents a stable complex, and later peaks represent unstable complexes falling apart during the chromatography or free heavy chains.

      (6) Line 175. As the 3BEV structure had a P2-His, is a comparable structure expected?

      This sentence reads “Structures with amino acid substitutions in the 11mer peptide GHAEEYGAETL bind with Asp at P<sub>c-2</sub> and either His or Arg at P<sub>2</sub> (5AD0 and 5ADZ), comparable to the original structure (3BEV) (Fig. 5A).” Minor changes in positions and orientations of individual amino acid sidechains are expected and are clear from the crystal structures presented in Fig. 5A, but are overall comparable to the original peptide in 3BEV, which has a His at P<sub>2</sub> and a Glu at P<sub>c-2</sub>.

      (7) Line 175 "modelling the substituted". How was the modelling done?

      The sentence reads “Modelling the substituted peptide with Arg at P<sub>2</sub> and Glu at P<sub>c-2</sub> shows a steric clash that can explain why this peptide did not refold with BF2*21:01 (Fig. 5A).” The legend to Fig. 5A states that “modelling done as detailed in Materials and Methods”, but apparently that section was omitted. A section has been added now to the Materials and Methods which states “Beginning with known crystal structures, modelling was carried out using PyMol with the protein mutagenesis Wizard, the rotamer toggle, show bumps and show surfaces.”

      (8) Line 176 "shows a steric clash" with what? Figure 5 is too small to see, and there is no label on the residue to follow where the steric clash is coming from.

      As stated in the legend to Fig. 5, “Structures were determined for GHAEEYGAETL (3BEV), GHAEEYGADTL (5AD0) and GRAEEYGADTL (5ACZ), which all refolded successful to give stable monomers, while GRAEEYGAETL did not (see Fig. 3), all models of which showed steric clashes (one depicted, red arrow).” In the model shown, the clash is between R9 of the BF2*21:01 with the Glu at peptide position 9. Parenthetically, this depiction of the key MHC residues for the co-variation has been used repeatedly, starting with Chappell et al 2015 eLIFE. The picture can be zoomed to make it large enough to see.

      (9) Line 247 "many combinations". It is not clear here if combinations are referring to a set of double-substitutions or different peptides?

      Each peptide in the library has a different double-substitution, so the meaning is the same either way. The point of Fig. 10 is to compare the identification of peptides from LC-MS/MS (in which each peptide is identified exactly) with the identification of sets of peptides from MALDI-TOF (in which the order of the double substitution is not clear, as well as the exact amino acid in the cases of I/L and Q/K). As the figure shows, the numbers are generally very similar, but this sentence describes the percentage of cases for which only one method or the other identified a peptide.

      (10) Lines 252-253. The conclusion of the MS data that the LC-MS/NS is more accurate and sensitive than MALDI-TOF is not very surprising. What was the rationale for using both?

      Our examination of the binding properties of BF2*21:01 for peptides and peptide libraries developed over a long time-span, so in the beginning we looked at single peptides with size exclusion chromatography peaks as the measure, then single- and later double-substitution libraries, first by MALDI-TOF and later by LC-MS/MS. Each set of experiments built on the previous work, so that together they tell the story. For example, after we optimized the use of double-substitution libraries, we were worried about the effects of temperature, so we tested that by MALDI-TOF. While repeating the temperature experiment once we optimized the LC-MS/MS approach might have yielded some additional data, there were other questions to answer.

      (11) Line 272-273 "support the idea that P3 is an important position within the peptide (Figures 8-10) despite not contacting the MHC molecule" The structure of 3BEV clearly shows interaction between the P3-Ala and the Tyr156 of the MHC. Residues that are fully or partially buried in the MHC cleft almost all contact the MHC molecule. Maybe I've missed something, but I think this statement is inaccurate.

      We agree with the reviewer that nearly every peptide residue contacts the MHC molecule (but of course some much more than others). The statement is now changed in the text to read “support the idea that P<sub>3</sub> is an important position within the peptide (Figures 8-10) despite not being an anchor residue."

      (12) Line 294. Figure 14 clearly shows that the number of 8 and 9-mer peptides eluted is at the same level as the 11mer and above, so how does the data fit with the statement that 9mer and shorter peptides are not stable with the MHC?

      Fig. 4 shows that 7, 8 and 9mer derivatives of the original 11mer failed to refold to give a single peak of stable monomers, while Fig. 6 shows that the 10mer derivative of the original 11mer was more thermostable than both the 9mer derivative and the original 11mer. The reason why the 9mer yielded so little monomer in Fig. 4 while giving enough to test by thermostability in Fig. 6 is no longer remembered. A key point is that these results are peptide sequence-specific, so it is not impossible to imagine stable binding of an appropriate 8mer (or perhaps even a 7mer). Another uncertainty, mentioned in the Results and Discussion, is that BF1*21:01 molecules bind primarily 8mers, and the contribution of peptides from BF1*21:01 is not known with certainty.

      (13) Line 316. What was the rationale for choosing peptides > 12aa to see if there is some overhang or bulge? As even 11-12mer can exhibit such features.

      We were just looking for any obvious patterns, but we didn’t find any.

      (14) The section starting at line 366 would have benefited from some structure prediction or modelling to illustrate the findings.

      We would have been delighted to model peptides, but we have used AlphaFold3 to compare the models to our crystal structures for seven chicken class I alleles (including BF2*21:01) with one or more peptides. The models sometimes fit the experimental data but they often didn’t, often by a wide margin, and with BF2*21:01 the worst (presumably because the system is not trained on MHC molecules which utilise charge transfer). Therefore, we do not feel confident in using any such modelling approaches except in the most defined situations (such as illustrated in Fig. 5).

      (15) Typo - Alleles should be italic, and in vitro as well.

      Alleles of genes are in italics, but alleles of proteins are not. To write in vitro in italics is customary, which we have corrected.

      (16) Figures

      (a) Some of the figures could be merged together.

      Of course, any presentation can be improved, but which figures we should merge (some already being three pages in length) is not clear. We could productively respond to more detailed suggestions.

      (b) Figure 1. I can't see the different colours mentioned in the figure legend

      Our apologies if the colours are not clear enough, but they are present only as an aid (as elsewhere in the manuscript), with red D and E, blue H, K and R, green N, Q, S and T, and all other amino acids black.

      (c) Figure 12. Is this figure only with 11mer peptides?

      Figure 12 shows the percentage of peptides with particular amino acids at P<sub>2</sub> and P<sub>c-2</sub> for three 11-mer double-substitution peptide libraries, the sequences of which are written on the graphs and described in the figure legends.

      (d) Table 1. The name of the protein should be added to the table.

      OK.

    1. Author response:  

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version.

      The authors have revised their manuscript to clarify that molecular analyses have not yet occurred and have resolved the technical/publication issues with the figures. I look forward to seeing these tools used in future publications to answer important questions in Blastocystis research.

      We are grateful to Reviewer 1 for their careful and constructive assessment of our revised manuscript. We are pleased that the revisions have satisfactorily addressed their previous concerns, and we sincerely appreciate their recognition of the value of our work.

      Reviewer #3 (Public review):

      Comments on revised version.

      The revised version provides sufficient clarity and appropriate visual presentation. Some confusion evidently arose due to my misunderstanding, so I thank the authors for their comprehensive clarifications and patience.

      We are grateful to Reviewer 3 for their careful reading of the revised manuscript and for the additional constructive suggestions. We appreciate that these recommendations are aimed at improving clarity, accessibility, and presentation, and we have addressed them as detailed below.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      (1) Table 1

      (1a) The description of the column "Promoter Lengths Tested" is a bit confusing because it actually shows acronyms of the used promoters. Suggestion for clarity: either keep the name and mention just the lengths, or rename to "Promoter abbreviation" or "Promoter acronym".

      We thank the reviewer for pointing this out. To avoid confusion, we have renamed this column “Promoters Tested”, which better captures the information presented. The column contains both promoter identifiers and the corresponding promoter lengths, which are further clarified in the table notes.

      (1b) It is not entirely clear based on the formatting of the table, which columns are relevant for "Blastocystis ST4-WR1" and which for "Blastocystis ST7-B".

      We thank the reviewer for this helpful suggestion. We have revised the formatting of Table 1 to more clearly distinguish the data corresponding to Blastocystis ST7-B and Blastocystis ST4-WR1. Specifically, we have added a vertical divider between the relevant column groups to improve readability and make the species-specific information easier to follow.

      (2) Figure 2:

      If the authors wish to preserve the decorative colouring in the parts B and C, it would be better to at least keep the datapoint and boxplot hues consistent for individual conditions (or plot the datapoints in the same, neutral hue throughout, e.g., black or grey). Currently, this is only the case in the part C, but not in the part B.

      We thank the reviewer for this useful suggestion. We have revised Figure 2B to improve visual consistency between the data points and boxplots while preserving the overall colour scheme of the figure. This adjustment improves readability without changing the data or its interpretation.

      (3) Figure 3: With the modifications everything is clear. However, two issues are now apparent.

      (3a) After the part A was modified, the arrow colours changed (had been yellow and red, now are cyan and magenta), but the description in the legend remained (yellow and red), so the text should be corrected. [By the way, the original colours worked well.]

      We apologise for overlooking this inconsistency. The Figure 3 legend has now been corrected to match the revised arrow colours in panel A.

      (3b) The care taken by the authors to make the images more accessible is really greatly appreciated, but being a person on the colourblind spectrum, I can attest that the issue is not always just about red and green discrimination: the greyscale and blueish green used in the part A photo in particular are discernible only at huge magnification (to some people). If the signal in the part A were of the same hue as in the part B (yellowish green instead of blueish green), it would readily pop out. For that matter, what is the reason for the use of a different colour profile of the UnaG fluorescence in these two images?

      We sincerely thank the reviewer for this important accessibility-related comment. The difference in colour profiles between panels A and B was intentional because the images were acquired using different imaging modalities, as indicated in the figure legend. We wanted to avoid implying that the two panels were generated under identical imaging conditions. However, we appreciate the reviewer’s point that this distinction should not come at the expense of readability. We have therefore adjusted the display of the UnaG signal in panel A to improve contrast and visibility. 

      (4) Conclusion:

      The expression "bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements" feels a bit clunky. What the part "endogenous regulatory part discovery" alludes to is now clearer, but consider reformulating it to "discovery of endogenous regulatory elements". This would make it clear without the need to add the explanation "namely the identification of native promoter and terminator elements". [It is now clear that the accumulation of noun adjectives and the significance of the word "part" was what originally blurred the overall meaning.

      We thank the reviewer for this valuable suggestion. We agree that the original phrasing was unnecessarily clunky and have revised the sentence for clarity and flow. Lines 682–684 now read:

      “By integrating endogenous promoter and terminator discovery, DNA delivery, selection, clonal recovery, and reporter validation into a single pipeline, we provide a flexible foundation for routine transgene-based studies.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Vitar et al. probes the molecular identity and functional specialization of pH-sensing channels in cerebrospinal fluid-contacting neurons (CSFcNs). Combining patch-clamp electrophysiology, laser-based local acidification, immunohistochemistry, and confocal imaging, the authors propose that PKD2L1 channels localized to the apical protrusion (ApPr) function as the predominant dual-mode pH sensor in these cells.

      The work establishes a compelling spatial-physiological link between channel localization and chemosensory behavior. The integration of optical and electrical approaches is technically strong, and the separation of phasic and sustained response modes offers a useful conceptual advance for understanding how CSF composition is monitored.

      Several aspects of data interpretation, however, require clarification or reanalysis-most notably the single-channel analyses (event counts, Po metrics, and mixed parameters), the statistical treatment, and the interpretation of purported "OFF currents." Additional issues include PKD2L1-TRPP3 nomenclature consistency, kinetic comparison with ASICs, and the physiological relevance of the extreme acidification paradigm. Addressing these points will substantially improve reproducibility and mechanistic depth.

      Overall, this is a scientifically important and technically sophisticated study that advances our understanding of CSF sensing, provided that the analytical and interpretative weaknesses are satisfactorily corrected.

      (1) The authors should re-analyze electrophysiological data, focusing on macroscopic currents rather than statistically unreliable Po calculations. Remove or revise the Po analysis, which currently conflates current amplitude and open probability.

      We agree with the reviewer that the Po analysis has strong limitations, particularly in experiments where the recording times are short, like when extracellular pH is changed either by photolysis (Figure 4D) or puff applications (Figure 3Aa). In order to circumvent that problem and not to rely only on Po estimations, we used alternative methods as well, including the analysis of the current membrane charge that we have used extensively during the manuscript (Figures 3A and 4D, for example) or the analysis of the event latencies (Figure 4G). Nevertheless, single-channel recordings clearly contain information that is not included in the macroscopic current analysis. We intend to stress in the revised version that the elementary current amplitude is conserved by manipulations such as pH changes, leaving the total number of channels (N) and the channel open probability (Po) as possible culprits for the current changes. Since these changes are rapid and reversible, it is likely that N stays constant and that Po changes. In order to address the reviewer’s concern, we propose the following changes/reanalysis: i) to report in each condition the minimum N (maximum observed openings; for example, in Figure 3Aa the minimum N goes from 4 in control conditions to 1 during the puff of the pH 6.4 solution). This method (to estimate N by counting the maximum number of open channels), while imperfect, provides a tentative estimate of Po; ii) following the previous point, we propose to reword the text (and images) and use the expression “apparent Po” instead of “Po”; iii) to report the fraction of time that the channels remain open. Also, we acknowledge that some traces are confusing (Figure 3Aa, top) as they seem to show macroscopic currents. We will modify those figures by plotting the amplitude histograms (as in Figure 1Bb) in order to show unambiguously that current recordings from CSFcNs only show single-channel activities.

      (2) PKD2L1-TRPP3 nomenclature should be clarified and all figure labels, legends, and text should use consistent terminology throughout.

      We agree with the reviewer that the nomenclature concerning polycystin members is confusing. In this manuscript we have followed the nomenclature that has been proposed in a recent, comprehensive review on polycystin channels by Palomero, Larmore and DeCaen (Palomero et al. 2023), where the authors refer to the channels by their gene name. In this review, the authors indicate that PKD2L1 channels correspond to TRPP2 (formerly TRPP3, their table 1). In another, recent review on TRP channels, however, the authors refer to the PKD2L1 channel as TRPP3 (Zhang et al. 2023). In order to avoid any confusion we will remove from the text any reference to the TRPP nomenclature and stick to the PKD2L1 name.

      (3) The authors should reinterpret the so-called OFF currents as pH-dependent recovery or relaxation phenomena, not as distinct current species. Remove the term "OFF response" from the manuscript.

      We concur with the reviewer that the term “OFF response” is not very helpful from the biophysical perspective and conveys the idea that it is another current. We will remove the term “OFF response” or “OFF current” in the revised manuscript and replace it by the term “photolysis-evoked PKD2L1 current”. Also, we will condense two sections (“The proton-induced current is an off-current” and “The off-current is mediated by the activation of PKD2L1 channels”) into a single new section (“The photolysis-induced current is mediated by PKD2L1 channels”), as separating the description of the photolysis-evoked PKD2L1 current compromises its description. Finally, we will rewrite the discussion to better describe this current.

      (4) Evidence for physiological relevance should be provided, including data from milder acidification (pH 6.5-6.8) and, where appropriate, comparisons with ASIC-mediated currents to place PKD2L1 activity in context.

      This is partly addressed in Figure 3. The data there indicates that PKD2L1 channels are very sensitive to pH variations around physiological pH. In order to make this conclusion stronger, we will add to the figure the EC50 values drawn from the fittings. In terms of the ASIC-mediated currents, one of our main conclusions is that ASIC channels are not present in the ApPr, as the effects of proton photolysis in the ApPr and are not blocked by ASIC channels blockers. Our results indicate that PKD2L1 channels are the exclusive pH sensitive channels in the ApPr, while ASIC channels are probably the acid-sensitive channels in the soma, although we have not studied the latter in detail. Following this and the editor’s comments, the subsection “the involvement of ASICs” in the Discussion has been modified.

      (5) Terminology and data presentation should be unified, adopting consistent use of "predominant" (instead of "exclusive") and "sustained" (instead of "tonic"), and all statistical formats and units should be standardized.

      Following the suggestions of the reviewer, an exhaustive work will be performed to unify terminology, data presentation and correct the text following the reviewer’s suggestions.

      (6) The Discussion should be expanded to address potential Ca<sup>2+</sup> -dependent signaling mechanisms downstream of PKD2L1 activation and their possible roles in CSF flow regulation and central chemoreception.

      This is indeed a very interesting and currently unresolved point in the physiology of CSFcNs. Published data indicate that calcium flowing into the cell through PKD2L1 channels is a key regulator of apical process physiology: on the one hand, PKD2L1 channels are calcium permeable and at the same time, they are inhibited by intracellular calcium (DeCaen et al. 2016). Also, ultrastructural data indicate that the ApPr is rich in mitochondria and tubulo-vesicular structures resembling the Golgi apparatus (Bjugn et al. 1988; Bruni et Reddy 1987), which are intracellular organelles that contribute to calcium homeostasis. Altogether, this evidence suggests that intraApPr calcium concentration needs to be finely regulated, both in space and in time, in order for the ApPr to fulfil its physiological roles. Based on what has been published in the literature, we can speculate that these calcium signals can be decoded by several systems: i) calcium can act as the second messenger linking the activation of the multimodal PKD2L1 channels to changes in CSFcNs excitability, which in turn regulate spinal neuronal networks controlling locomotor activity; ii) calcium could initiate neurosecretion of different molecules from the ApPr to the central canal (as has been proposed by the Wyart group in the zebrafish in the context of bacterial infections(Prendergast et al. 2023)); iii) calcium could activate the Hedgehog signaling pathways (as has been shown by (Delling et al. 2013)); iv) calcium could modulate CSF flow directly (by modulating ciliary activity) or indirectly (by modulating ependymal cells ciliary activity through paracrine interactions). Resolving these downstream pathways is essential to fully define the role of CSFcNs as integrators of CSF homeostasis. We will expand this issue in the Discussion of the revised ms.

      Reviewer #2 (Public review):

      Summary:

      Cerebrospinal fluid contacting neurons (CSF-cNs) are GABAergic cells surrounding the spinal cord central canal (CC). In mammals, their soma lies sub-ependymally, with a dendritic-like apical extension (AP) terminating as a bulb inside the CC.

      How this anatomy-soma and AP in distinct extracellular environments relate to their multimodal CSF-sensing function remains unclear.

      The authors confirm that in GATA3:GFP mice, where these cells are labeled, that CSFcNs exhibit prominent spontaneous electrical activity mediated by PKD2L1 (TRPP2) channels, non-selective cation channels with ~200 pS conductance modulated by protons and mechanical forces.

      They investigated PKD2L1 pH sensitivity and its effects on CSFcN excitability. They uncovered that PKD2L1 generates both phasic and tonic currents, bidirectionally modulated by pH with high sensitivity near physiological values.

      Combining electrophysiology (intact and isolated AP recordings) with elegant laser-photolysis, they show that functional PKD2L1 channels localize specifically to the apical extension (AP).

      This spatial segregation, coupled with PKD2L1's biophysical properties (high conductance, pH sensitivity) and the AP's unique features (very high input resistance), renders CSFcN excitability highly sensitive to PKD2L1 modulation. Their findings reveal how the AP's properties are optimised for its sensory role.

      Strengths:

      This is a very convincing demonstration using elegant and challenging approaches (uncaging, outside out patch of the AP) together to form a complete understanding of how these sensory cells can detect the changes of pH in the CSF so finely.

      Weaknesses:

      The following do not constitute weaknesses; rather, they are minor requests that this reviewer considers would complete this beautiful study.

      (1) It would be nice to quantify further the relation in spontaneous as well as in acidic or basic pH between the effects observed on channel opening and holding current: do they always vary together and in a linear way?

      Following the reviewer’s suggestion, we have performed a Spearman’s rank correlation test, which shows a significant correlation between the changes in the apparent open probability and holding current (paired experiments; ctrl vs pH 6.4 pressure applications; p < 0.05, Spearman r = 0.72 and critical value = 0.67). The Pearson correlation coefficient calculated on the same data set = 0.63 and the critical value is 0.632, which indicates that the correlation is not linear. We will add this analysis to the manuscript.

      (2) Since CSF-cNs also respond to changes in osmolarity (Orts Dell Immagine 2013) & mechanosensory stimulations in a PKD2L1 dependent manner (Sternberg NC 2018), it would be nice to test the same results whether the same results hold true on the role of PKD2L1 in AP for pressure application of changes in osmolarity.

      This is a very important point. As the reviewer mentions, previously published experimental evidence indicates that CSFcNs are also sensitive to osmolarity changes and mechanical stimulation in a PKD2L1-dependent manner. It is therefore reasonable to assume that, as for the pH sensitivity, osmotic and mechanical sensitivity depends on channels segregated to the ApPr. For the mechanosensitivity, the spatial segregation could be tested by “touching” either the ApPr or the soma with a piezo-controlled blunted pipette (see, for example, Hao et al. 2013). However, the sensitivity to osmotic changes is much more difficult to assess, as pressure application does not have enough spatial resolution to discriminate among compartments in such a compact cell such as the CSFcNs. In theory, a highly spatially localized osmotic jump could be reached with photolysis, but a caged compound releasing many osmotic particles simultaneously should be used. In typical photolysis experiments, a localized osmotic jump is produced, but it is very low (in the order of 1 to 2 mOsm).

      In mice, like in fish (Sternberg et al, NC 2018), we can observe throughout the figures that a large fraction of the channel activity occurs with partial and very fast openings of the PKD2L1 channel. I recommend the authors analyse the points below:

      (a) To what extent do these partial openings of the channel contribute to the changes in holding current and resting potential?

      As the reviewer indicates, these partial and very fast openings are a characteristic of PKD2L1 single-channel activity that seems to be present in different species. However, estimating what is the exact contribution of these events to the sustained current would require a detailed model of the channel that it is still lacking. Indeed, the exact mechanism by which CSFcNs show this prominent sustained current is unknown and should definitely being addressed in future works.

      (b) In the trace from the outside out AP, it looks like the partial transient openings are gone. Can the authors verify whether these partial openings are only present in somatic recordings?

      The outside-out recordings from the ApPr also show some partial openings (please look at the upper trace in Figure 4Db). We will specifically mention this important point in the revised version of the ms.

      (3) Previous studies have observed expression of metabotropic Glutamate receptors in CSF-cNs (transcriptome from Prendergast et al CB 2023). The authors only used blockers for ionotropic glutamate receptors in their recordings: could it be that these metabotropic receptors influence the response to uncaging of MNI-Glu when glutamate is co-released with a proton?

      We thank the reviewer for pointing out the presence of metabotropic glutamate receptors in CSFcNs. However, our evidence indicates that there is no contribution of metabotropic receptors when uncaging MNIglutamate because: i) the response obtained when uncaging MNI-gLGG (where there is no glutamate release; Figure 5Ab) and ii) the response obtained when uncaging protons from DPNIGABA (a GABA cage that has similar photochemistry than MNI cages which also release a proton when photolysed; data not shown), are the same. Indeed, in both experiments (MNI-gLGG or DPNI-GABA uncaging) a clear photolysis-evoked PKD2L1 current can be observed.

      (4) In the outside out patch of the AP, PKD2L1 unitary currents appear rare. Could it be that the disruption in the cilium or underlying actin/myosin cytoskeleton drastically alter the open probability of the channel?

      Although we have not quantified it, the reviewer is right that the opening frequency of PKD2L1 channels in the outside-out patches is lower than in the whole-ApPr recordings. We interpreted this difference as a difference in channel number. However, another plausible interpretation is that, as the reviewer suggests, the biophysics of the channels are affected because the protein is taken out from its normal ionic environment and/or loses important interactions with regulatory proteins.

      (5) Could the authors use drugs against ASIC to specify which ASIC channels contribute to the pH response in the soma?

      As described in the manuscript, we did perform experiments with ASIC channel blockers, although we did not attempt to characterize the specific ASIC channel involved in the somatic response. Based on what has been published in the literature, we used both psalmotoxin-1 (which blocks ASIC1 channels) and APETx2 (which blocks ASIC3 channels). The presence of ASIC1 channels in mice CSFcNs has been shown by (Orts-Del’Immagine et al. 2012; Orts-Del’Immagine et al. 2016), while the presence of ASIC3 in the lamprey CSFcNs has been shown by (Jalalvand et al. 2016). When we puff an acidic solution aiming at the soma, we can record an inward current that is blocked by psalmotoxin-1, although there is always a small component remaining (as originally shown by Orts-Del’Immagine in the aforementioned articles); however, we have not attempted to block this small component that remains after psalmotoxin-1 bath application.

      (6) This is out of the scope of this study, but we did observe in fish a very rarely-opening channel in the PKD2L1KO mutant. I wonder if the authors have similar observations in the conditions where PKD2L1 is mainly in the closed state.

      We have never seen such kind of openings in our recordings (when the channel is closed or in the presence of dibucaine).

      Bjugn, R, H K Haugland, et P R Flood. 1988. “Ultrastructure of the mouse spinal cord ependyma.” Journal of Anatomy 160 (octobre): 117‑25.

      Bruni, J. E., et K. Reddy. 1987. “Ependyma of the Central Canal of the Rat Spinal Cord: A Light and Transmission Electron Microscopic Study”. Journal of Anatomy 152 (juin): 55‑70.

      DeCaen, Paul G., Xiaowen Liu, Sunday Abiria, et David E. Clapham. 2016. “Atypical Calcium Regulation of the PKD2-L1 Polycystin Ion Channel”. eLife 5 (juin): e13413. https://doi.org/10.7554/eLife.13413.

      Delling, Markus, Paul G. DeCaen, Julia F. Doerner, Sebastien Febvay, et David E. Clapham. 2013. “Primary cilia are specialized calcium signalling organelles”. Nature 504 (7479): 311‑14. https://doi.org/10.1038/nature12833.

      Hao, Jizhe, Jérôme Ruel, Bertrand Coste, Yann Roudaut, Marcel Crest, et Patrick Delmas. 2013. “Piezo-Electrically Driven Mechanical Stimulation of Sensory Neurons”. In Ion Channels, édité par Nikita Gamper, vol. 998. Methods in Molecular Biology. Humana Press. https://doi.org/10.1007/978-1-62703-351-0_12.

      Jalalvand, Elham, Brita Robertson, Peter Wallén, et Sten Grillner. 2016. “Ciliated Neurons Lining the Central Canal Sense Both Fluid Movement and pH through ASIC3”. Nature Communications 7 (janvier): 10002. https://doi.org/10.1038/ncomms10002.

      Orts-Del’Immagine, Adeline, Riad Seddik, Fabien Tell, et al. 2016. “A Single Polycystic Kidney Disease 2-like 1 Channel Opening Acts as a Spike Generator in Cerebrospinal Fluid Contacting Neurons of Adult Mouse Brainstem”. Neuropharmacology 101 (février): 549‑65. https://doi.org/10.1016/j.neuropharm.2015.07.030.

      Orts-Del’immagine, Adeline, Nicolas Wanaverbecq, Catherine Tardivel, Vanessa Tillement, Michel Dallaporta, et Jérôme Trouslard. 2012. “Properties of Subependymal Cerebrospinal Fluid Contacting Neurones in the Dorsal Vagal Complex of the Mouse Brainstem”. The Journal of Physiology 590 (16): 3719‑41. https://doi.org/10.1113/jphysiol.2012.227959.

      Palomero, Orhi Esarte, Megan Larmore, et Paul G. DeCaen. 2023. “Polycystin Channel Complexes”. Annual Review of Physiology 85 (Volume 85, 2023): 425‑48. https://doi.org/10.1146/annurev-physiol-031522-084334.

      Prendergast, Andrew E., Kin Ki Jim, Hugo Marnas, et al. 2023. “CSF-Contacting Neurons Respond to Streptococcus Pneumoniae and Promote Host Survival during Central Nervous System Infection”. Current Biology 33 (5): 940-956.e10. https://doi.org/10.1016/j.cub.2023.01.039.

      Zhang, Miao, Yueming Ma, Xianglu Ye, Ning Zhang, Lei Pan, et Bing Wang. 2023. “TRP (Transient Receptor Potential) Ion Channel Family: Structures, Biological Functions and Therapeutic Interventions for Diseases”. Signal Transduction and Targeted Therapy 8 (1): 261. https://doi.org/10.1038/s41392-023-01464-x.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers were very impressed with your work and definitely feel it is a scientifically important and technically sophisticated study that advances our understanding of CSF sensing. They, however, request some re-analysis of the data and discussion with minimum new experiments, if any. I think, if feasible, this will improve the quality of the study and would look forward to receiving a revised version.

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 1 - Molecular identity and localization of PKD2L1

      Major

      Nomenclature clarity.

      Please clarify the distinction between PKD2L1 and TRPP3. Several parts of the text and figure labels appear to conflate these names. PKD2L1 corresponds to the TRPP3 subfamily member and should not be interchanged with PKD2/TRPP2. Please confirm and update the nomenclature consistently throughout the manuscript (text, figure labels, and captions).

      We have addressed this issue in response to reviewer 1. There seems to be some confusion in the literature concerning the nomenclature of PKD2L1 channels as in some recent publications the PKD2L1 channels are still named as TRPP3. However, the nomenclature of PKD2L1 channels or TRPP2, was updated in 2016 (Wu, Sweet and Clapham, Pharmacological Reviews, 2010). As indicated in response to reviewer 1, we have removed from the text any reference to the TRPP nomenclature and stuck to the PKD2L1 name.

      Physiological meaning of apical restriction.

      Expand the discussion of why apical-restricted localization matters. Specifically, address how segregation to the ApPr could support directional sensing of CSF flow and/or detection of localized pH gradients.

      Following the reviewers and editor’s comments, we have revised the discussion in order to take this and other comments into account.

      Open-state annotation (O3).

      Please include the O3 state in panels Ba and Bb; the figure clearly shows an additional open level consistent with O3.

      The editor is right in that there is another state that presumably corresponds to O3. Following the editor’s recommendation, we now indicate this 3rd level and add a short sentence explaining this in figure 1 legend.

      Minor

      Indicate the ROI definition and background-subtraction method used for fluorescence quantification (ApPr vs soma).

      Not applicable for this figure.

      Ensure the same intensity scale (lookup table and range) is used across panels to enable direct comparison.

      Done.

      (2) Figure 2 - Electrophysiological characterization of ApPr and somatic recordings Major

      Definition of "PKD2L1-dependent current."

      Define this term precisely at its first appearance. Specify whether it denotes currents inhibited by dibucaine, abolished in PKD2L1-knockout preparations, or both.

      Done

      Statistical power of single-channel analysis.

      The number of observed openings (< 1000 events) is too low to estimate open probability (Po) reliably. Please re-analyze the data using macroscopic current traces rather than Po-based kinetics.

      Confounded Po analysis.

      The current Po analysis mixes current amplitude and Po in the same calculation, conflating independent variables. Re-evaluate or remove this analysis.

      Unknown channel count.

      Because the number of channels in each patch is unknown, Po and "closed probability" values cannot be interpreted meaningfully. Focus instead on the averaged macroscopic current density.

      General analytical validity.

      The single-channel analyses in Figure 2 are not interpretable under these experimental conditions. Closed-time distributions and Po-based metrics (e.g., "Po1," "P2") depend critically on channel number and event sampling. Moreover, the manuscript applies essentially the same Po methodology across conditions (Po1 vs P2), which adds no mechanistic resolution and risks circular interpretation.

      Actionable recommendation:

      Remove Po- and closed-time-based analyses from Figure 2 and from the manuscript as a whole. Reanalyze the data using metrics that remain valid when the channel number is unknown:

      Macroscopic current analysis (leak-subtracted current density, I-V relationships, activation time constants).

      Single-channel conductance only (amplitude histograms and unitary slope conductance), without attempting Po or dwell-time inference.

      Report filtering bandwidth and sampling rate, and restrict statistical treatment to these robust parameters.

      Following the reviewers (see above) and editors’ recommendations, we have reanalyzed the data in order to avoid the analysis based on Po and Pc. We have instead calculated from the recordings other 2 parameters, n<sub>max</sub> (the maximum number of channels that open simultaneously during a 500 ms time window) and the total open time of a single channel during the same 500 ms time window. The main text, figures and corresponding figure legends, and the Materials and Methods section have been changed accordingly. Notably, the Po and Pc analysis were removed from Figures 3, 4 and Supplementary Figure 3, and replaced by the above-mentioned parameters. Also, the fact that the recordings are not long enough to calculate Po is now specifically mentioned in the Materials and Methods section, lines 785 to 790. In addition, the analysis in Figure 3Ce has been redone so that the activity of the channel as a function of pH is now plotted as the normalized apparent Po (relative to the apparent Po value at pH 7.4).

      Minor

      State whether input-resistance values (1.8-4.4 GΩ) were leak-subtracted and series-resistance-compensated.

      As already mentioned in the methodology section (line 693), series resistance was not compensated for during the experiments. We have now added a sentence in the methodology section indicating that in the voltage range that was chosen for the analysis of the input resistance, no voltage-dependent conductance was activated (lines 809 to 812).

      Ensure unit consistency: use Po or normalized Po rather than frequency (Hz) throughout.

      (3) Figure 3 - pH-evoked currents and kinetics

      Major

      Invalid Po analysis (Fig. 3Ca-Ce).

      The Po- and Pc-based single-channel analyses in panels 3Ca-3Ce should be deleted. As noted earlier, the event count is insufficient, and the number of active channels in each patch is unknown. Under these conditions, Po and Pc values have no quantitative meaning and could mislead readers. These panels do not contribute additional mechanistic insight beyond the macroscopic current data and therefore, should be removed. If retained for illustrative purposes, they must be explicitly labeled as representative traces without any statistical quantification.

      As we mentioned above, we removed the Po and Pc analysis from the manuscript.

      Minor

      Present regression equations and r<sup>2</sup> values for the linear fits shown in Fig. 3D.

      Done (lines 947951).

      Confirm that all axes include units and identical scaling between conditions for direct comparison.

      Done.

      (4) Figure 4 - Laser photolysis and local stimulation experiments

      Major

      Laser timing annotation.

      Clearly mark laser-pulse timing (e.g., arrow or shaded region) on all current traces to facilitate interpretation.

      We thank the editor for pointing out the inconsistencies in terms of the laser pulse timing. To indicate the laser pulses, we have now added an arrowhead in cases where a single sweep is shown (for example, Figure 3D), and an arrowhead and a dotted magenta vertical line in cases where multiple sweeps are shown (for example, Figure 3C).

      pH calibration within the laser spot.

      Provide quantitative calibration of pH changes induced by laser photolysis, including information on spot size, local diffusion, and estimated pH recovery kinetics.

      This is an important point and we thank the editor for mentioning it. We have now completed the subsection untitled “photolysis” where we provide information on the lateral and axial dimensions of the photolysis laser spot used in this work (lines 734 to 737). We have also rewritten part of Figure 5A legend to highlight the fact that the experiments presented there (photolysis on top of the ApPr and next to it) are compatible with a high spatial resolution of proton release (lines 1023 to 1026).

      On the other hand, we have attempted to perform pH calibrations in the setup using the pH-sensitive dye pyranine (or HPTS: 8-Hydroxypyrene-1,3,6-trisulfonic acid). HPTS is a very useful tool for pH calibrations in the physiological range: its pKa value is close to 7.2, and it can be used as a ratiometric dye (its fluorescence is pH-independent at 405–410 nm and pH-dependent at 450 nm). Unfortunately, when trying to perform a calibration under the conditions of a real experiment,

      where photolysis occurs in a tiny volume (approximately 1 µm³ in a total bath volume of more than 1 ml), we encountered the following problem, which made it impossible to obtain any useful data: the 405 nm uncaging pulse bleaches the dye, and any useful information is lost. Also, our imaging system is not fast enough to follow the pH change. As it is discussed in the Materials and Methods section, subsection “Estimation of the pH drop induced by photolysis” (line 814), the fast protonation of bicarbonate indicates that the pH change induced by the photolysis recovers in the submillisecond time range.

      Repeated stimulation effects.

      Discuss whether repeated photolysis induces adaptation or desensitization of PKD2L1 currents, and indicate whether current amplitude decreases across successive trials.

      This issue is now specifically mentioned in the Materials and Methods section, lines 739 to 741.

      Invalid interpretation of the "OFF response."

      The interpretation of the so-called "OFF response" in Figure 4C is not supported by the presented data. There is no evidence for a bona fide OFF current, and the literature cited does not demonstrate such a phenomenon for PKD2L1 alone. Rather, previous studies implicate PKD1L3-dependent mechanisms in similar biphasic responses. Please reconsider the cited references and remove claims of an OFF current attributed to PKD2L1.

      Done.

      Actionable recommendations:

      Do not use the term "OFF response" throughout the manuscript. Recast these transients as pH dependent recovery or relaxation of current following cessation of acidification.

      Done. We have performed extensive rewriting and reorganization of the Results and Discussion in order to take into account both the reviewer’s and editor’s comments. Please also take a look at comment #3 of Reviewer 1 and point 11 below.

      Include continuous-illumination controls (sustained local acidification) to test whether a steady state current is maintained. This will clarify whether the post-stimulus transient reflects recovery kinetics rather than a distinct current species.

      We thank the editor for suggesting this experiment. However, continuous laser illumination is a difficult manipulation and does not necessarily lead to an acidification of the illuminated volume. Indeed, with continuous illumination the cage is lost from the illumination spot and needs to be replaced by diffusion from the non-illuminated volume, leading to non-homogeneous concentrations. Also, the chances of inducing photo damage are higher. We thus designed a similar experiment where instead of performing continuous illumination we photolysed with short and high frequency trains in order to produce a long-lasting acidification. The results of these experiments have been added to the manuscript as part of the results section and in Figure 5H. Similarly to what is seen with single illuminations, the photolysis trains induce a current that appears almost exclusively at the end of the train, implying that the current is indeed a PKD2L1-dependent recovery current.

      Align the current time course with measured or estimated local pH (or calibrated proxy) to demonstrate causal coupling and avoid implying a separate conductance.

      We have added the calculated pH change to the inset of Figure 4C as an example.

      Revise the schematic/model figure and textual description accordingly, restricting the framework to phasic vs sustained activation modes without invoking a separate OFF current for PKD2L1.

      Done.

      Minor

      Include scale bars, sample numbers (n), and laser parameters (duration, power) in all panels.

      In order not to make the figure and the panels very heavy in the original version, we tried to limit the number of scale bars. We have now performed some modifications, added the missing scale bars, and changed the figure legend in order to take into account the editor’s comments. We have also corrected a few values that were wrongly reported.

      Standardize p-value formatting (e.g., p = 6 × 10 ⁶) throughout the figure and legend.

      Done.

      (5) Figure 5 - Single-channel recordings

      Major

      Mixed parameters (current amplitude and Po).

      The current analysis improperly mixes single-channel current amplitude and Po within the same figure, conflating distinct parameters. These quantities must be analyzed and presented separately, or the Po data should be removed entirely if not independently supported.

      Insufficient event count.

      Given the very limited number of observed openings, Po-based statistics are not meaningful. Please report only representative single-channel traces and corresponding amplitude histograms without attempting quantitative Po estimation.

      Minor

      Convert frequency (Hz) values to Po for consistency with earlier analyses, or remove frequency metrics altogether if Po analysis is omitted.

      Figure 5 does not include Po or event frequency analysis, so we think there must be a misquotation of the figure. However, the Po issue has already been addressed before and alternative analysis have been proposed.

      (6) Introduction

      The introductory paragraph mentions the "five senses" as a framing concept. However, this statement lacks scientific grounding in the context of CSF-contacting neurons and chemosensory physiology. The traditional "five senses" classification is not an evidence-based neurophysiological framework and may be misleading to readers. I recommend removing or rephrasing this part, focusing instead on molecular and cellular mechanisms of sensory transduction (e.g., chemical, mechanical, and pH sensing) rather than on classical sensory categories.

      Following the editor’s recommendation, we have removed this part.

      The manuscript refers to PKD2L1 using the term TRPP2 in some parts of the introduction. This is incorrect, as PKD2L1 corresponds to TRPP3, not TRPP2. Please correct this nomenclature and ensure consistent use of "PKD2L1 (TRPP3)" throughout the entire manuscript to avoid confusion with the distinct PKD2/TRPP2 protein, which belongs to a different subfamily with separate physiological roles.

      We thank the editor for pointing this out. As we mentioned in the responses to the “public reviews”, the literature is confusing, so we decided to remove from the manuscript any mention to TRPP channels.

      (7) Discussion

      The current Discussion reads largely as a descriptive summary of results and lacks conceptual depth. It does not effectively integrate the biophysical properties of PKD2L1 with its physiological role as a neuronal pH sensor, nor does it develop a broader interpretation relevant to CSF homeostasis or chemoreception.

      Following the reviewers and editor’s recommendations, we have now added a new section in the Discussion untitled “PKD2L1 downstream signaling mechanisms”.

      (8) Insufficient biophysical analysis

      The discussion of channel gating and pH dependence is superficial and does not explore the energetic or structural mechanisms underlying proton sensitivity. The authors should analyze their data in the context of known PKD/TRPP family biophysics-for example, protonation sites, subunit composition, or gating kinetics-and explain how these confer bidirectional (acidic vs alkaline) sensitivity within physiological ranges.

      In this work, we studied the pH sensitivity of PKD2L1 channels in the context of CSFcN sensory physiology. From a pure biophysical perspective, the pH sensitivity of PKD2L1 channels has been studied by multiple groups; however, it is still unknown how the gating of the channel responds to pH changes, although it can be proposed that some polar residues in the protein regulate the state of the pore. Likewise, the mechanism of the “off-response” is also unknown. To the best of our knowledge, there is only one article in which the authors have attempted to relate pH, PKD2L1 channel structure, and function. In this work (Su et al., Nature Communications 2018), the authors compare PKD2L1 channels with another pH-sensitive member of the TRP family, TRPML3, whose structures at pH 7.4 and 4.8 are known (Zhou et al., Nature Structural and Molecular Biology, 2017). We have rewritten some sentences of the Discussion in order to be more specific about the pH dependence of PKD2L1 channels and its proposed mechanisms.

      (9) Weak physiological context

      The manuscript does not adequately address how PKD2L1 functions as a true physiological pH sensor. The discussion should connect channel activity to realistic CSF pH fluctuations (6.8-7.6) and to relevant physiological or pathophysiological conditions (e.g., respiratory acidosis, neurogenic regulation of CSF composition). Without this, the relevance of large, artificial acidification (pH 3-3.5) remains unclear.

      We have added a new section in the Discussion where we speculate on how PKD2L1 channels may be activated in physiological and pathophysiological conditions. However, we would like to insist here that the main goal of the photolysis experiments (which induce short and large acidifications) was to assess the spatial segregation of PKD2L1 channels. We now mention this point specifically and also speculate on the conditions that could eventually give rise to the “recovery” current.

      (10) Over-interpretation of unsupported points

      The paragraph describing voltage propagation from the ApPr to the soma/axon is speculative and unsupported by any data in the manuscript. Please delete this section entirely, including the citation to Orts-Del'immagine et al., unless new electrophysiological evidence is added.

      We think this point (the propagation of signals originating from the ApPr to the soma) is important in the context of our work, so we have decided to make new experiments in order measure directly the degree of coupling between the 2 compartments. To do that we made simultaneous, current-clamp and voltage-clamp recordings from the ApPr and the soma. In these conditions we were able to measure experimentally and for the first time both the coupling coefficient and coupling conductance, which confirm that the propagation of voltage signals from the ApPr to the soma is extremely efficient. These new results are now described in a new subsection and in a new Figure 6.

      (11) Clarify the role of "OFF currents."

      The Discussion repeatedly refers to an "OFF response," but this phenomenon is not experimentally demonstrated for PKD2L1 alone. It likely represents pH-dependent recovery rather than an independent current. All discussion of "OFF currents" should be removed or reformulated accordingly.

      Following the editor and reviewer’s comments, we have deleted the term “off response” and “off currents” from the ms and have replaced them with the term “recovery current”. We have also changed the discussion accordingly.

      (12) Integration with ASICs and compartmental sensing

      While the manuscript briefly mentions ASIC involvement, it does not articulate how PKD2L1- and ASIC-mediated signals might complement each other in different compartments (ApPr vs soma). The authors should discuss the potential division of labor between these sensors and how such compartmentalization enhances pH detection in CSFcNs.

      Following the editor’s comments, we have rewritten the part of the subsection ‘the involvement of ASICs’ in the Discussion.

      (12) Broadened physiological perspective

      The Discussion should close by considering Ca<sup>2+</sup> -dependent downstream pathways activated by ⁺ PKD2L1 and their implications for CSF flow regulation, neurosecretion, and central chemoreception. These translational aspects would substantially improve the impact and readability of the manuscript.

      Done

      Overall, the Discussion must evolve from a descriptive narrative to a mechanistically and physiologically integrative synthesis, highlighting why PKD2L1 is not merely present in the ApPr but is a key molecular transducer linking ionic microenvironment to neuronal excitability.

      As it has been detailed above, we have performed several changes in the Discussion that follow the reviewer’s and editor’s recommendations.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We believe that the manuscript has been substantially strengthened through the revision process. The main changes are summarized below:

      We substantially revised the Introduction and Discussion sections to better position our work relative to previous studies on starvation-dependent thermotaxis plasticity, neuropeptidergic modulation, and AWC function.

      We clarified throughout the manuscript the distinction between negative thermotaxis in innocuous thermal ranges and thermonociceptive responses to noxious heat. We now discuss more explicitly that these behaviors involve at least partly distinct molecular, cellular, and circuit-level mechanisms.

      We performed new experiments in ins-1 mutants. Unlike what was previously reported for thermotaxis plasticity, ins-1 does not appear required for starvation-dependent thermonociceptive plasticity in our paradigm (new Figure 6—figure supplement 1).

      We revised the analysis and terminology used for AWC calcium imaging data. We no longer use the “deterministic/stochastic” terminology and instead describe a starvation-induced shift from predominantly excitatory responses to a mixed distribution of excitatory and inhibitory responses. We also added new quantitative analyses and histogram representations of response distributions, as directly suggested by reviewers, to better illustrate this point.

      We performed new genetic interaction experiments using eat-4; flp-6 double mutants. These analyses revealed that glutamatergic and FLP-6 signaling act largely in parallel to mediate heat-evoked reversals after early food deprivation, while prolonged starvation reveals a hierarchical interaction between these pathways.

      We revised and clarified the mechanistic model figures accordingly, particularly regarding the proposed ASI → AWC signaling pathway and the role of ASI-derived neuropeptides.

      We improved the presentation and statistical rigor throughout the manuscript, including:

      - Replacement of heating power values by corresponding temperature increases,

      - Clarification of the rationale for using the 1-hour off-food condition as reference,

      - Expanded statistical reporting and multiple-comparison procedures,

      - Additional methodological details for calcium imaging, rescue validation, and cell ablation approaches,

      - Clarification of replotted datasets in figure legends.

      We also simplified the manuscript by removing experiments whose interpretation remained ambiguous (notably the nsy-1 and nsy-7 analyses).

      Below, we provide a detailed point-by-point response to all reviewer comments.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Thapliyal and Glauser investigates the neural mechanisms that contribute to the progressive suppression of thermonociceptive behavior that is induced under conditions of starvation. Several previous studies have demonstrated that when starved, C. elegans alters its preferences for a variety of sensory cues, including CO2, temperature, and odors, in order to prioritize food seeking over other behavioral drives. The varied mechanisms that underlie the ability of internal states to alter behavioral responses are not fully understood; however, there is growing evidence for a role of neuropeptidergic signaling as well as the capacity for functionally distinct microcircuits, formed by distinct internal states, to trigger similar behavior outcomes.

      Within the physiological range of C. elegans (~15-25{degree sign}C), starvation triggers a profound reduction in temperature-driven thermotaxis behaviors. This reduction involves the recruitment of the amphid sensory neuron pair AWC. The AWC neurons primarily act to sense appetitive chemosensory cues; however, under starvation conditions begin to display temperature responses that previous studies have linked to the reduction in thermotaxis navigation. Here, Thapliyal and Glauser investigate the impact of starvation on thermonociceptive responses, innate escape behaviors that are triggered by exposure to noxious temperatures above 26{degree sign}C or rapid thermal stimuli below 26{degree sign}C. They compare the strength of thermonociceptive behaviors, specifically heat-triggered reversals, in worms experiencing either early food deprivation (1 hour off food) or prolonged starvation (6 hours off food). Their experiments demonstrate a progressive loss of heattriggered reversals that is mediated by AWC and ASI neurons, as well as both glutamatergic and neuropeptidergic signaling.

      At the level of neural activity, this study reports that the transition from early food deprivation to prolonged starvation reconfigures the temperature-driven activity of AWC neurons from largely deterministic to stochastic. This finding is interesting in light of previous work that reported the opposite transition (from stochastic to deterministic) in temperature-driven AWC responses when comparing well-fed worms to those kept from food for 3 hours. This study also identifies neural and genetic mechanisms that contribute to differences in thermonociceptive responses at +1 versus +6 hours of starvation; confusingly, these mechanisms are partially distinct from those that contribute to differences in negative thermotaxis behaviors in well-fed and +3 hours of starvation worms (Takeishi et al, 2020). A limitation of this manuscript is that these differences are not particularly acknowledged or addressed, other than the hypothesis that independent mechanisms underlie negative thermotaxis versus thermonociceptive stimuli. However, this suggestion is not experimentally verified.

      We thank this reviewer for pointing to the interest of our work. The difference between previous work focusing on negative thermotaxis in the range of innocuous temperatures and our work focusing on thermo-nociceptive response is important and indeed deserves further clarification and a deeper discussion in the manuscript.

      Two major empirical evidence for a distinction between negative thermotaxis (as assessed in previous studies) and thermonociceptive plasticity (as assessed in our paradigm) were already included in the initial article version. First, we reported a decrease in average response in AWC neurons due to a shift in the distribution of response polarities from mostly up-response to a mix of ‘up-response’ and ‘downresponse’ after starvation, while previous results showed an increase in response probability of AWCs after starvation (Takeishi et al, 2020). Second, contrary to starvation-evoked thermotaxis adaptation, ASI neurons are required to orchestrate starvation-evoked plasticity in thermonociception. These observations already indicate differences at the circuit and cellular level. For the revision, we conducted further experiments to address the molecular level. We tested ins-1 mutants (see also specific point 3 by reviewer 2, below) and deepened this aspect in the discussion section of the revised manuscript. Previous study found that INS-1 signaling from the intestine is a major mediator of negative thermotaxis plasticity. In contrast, our new data show that INS-1 peptide does not seem critical in regulating starvation-dependent thermonociceptive plasticity (see new Figure 6-supplement 1). Taken together these three lines of empirical evidence support the notion that negative thermotaxis and thermonociceptive starvation-evoked plasticity involves at least partially distinct mechanisms and it seems therefore inappropriate to qualify this notion as purely hypothetical.

      Modification in the revised manuscript include extended introduction about the known thermotaxis regulation mechanisms (Introduction section), new Figure 6-supplement 1 about ins-1 and accompanying text in the result section as well as extended discussion about these differences (Discussion section)

      Multiple additional aspects of this study make the results difficult to synthesize with existing knowledge, including

      (1) Differences in - and insufficient discussion of - the magnitude and kinetics of thermal stimuli;

      We have included a better description of the stimuli characteristics in the revised methods section. The discussion section was deepened to better emphasize that different types of thermal stimuli have been used in different studies.

      (2) This study's use of "heating power" rather than temperature values when presenting behavioral results;

      Thanks for noting this point, which was indeed an unnecessary complication in the result display of the initial manuscript. We have changed ‘power values’ to corresponding ‘temperature increase’ in the revised figures.

      (3) The use of +1 hours starvation as a baseline instead of well-fed worms. Indeed, this last point reflects a noticeable experimental result that differs from previous studies, namely that at room temperature, the basal movements of well-fed and starved worms are not different. Such a surprising result warrants further quantification of worm mobility in general and could have prompted a set of experiments directly testing previously published thermal conditions to demonstrate that the new effects reported arise specifically from the use of thermonociceptive stimuli, as hypothesized.

      The consideration of on-food and off-food behavioral state is an important point indeed. We found that, at room temp, worms shift from dwelling on food (a state with high spontaneous reversal rate) to global search off-food (a state with low spontaneous reversal, after 1hr starvation) (see Figure 1B). Therefore, unlike the reviewer’s statement, the reported data highlighted key differences in the basal locomotion of worms in fed and 1hr starved conditions. Furthermore, these behavioral states have been characterized very deeply using high-content worm behavioural tracking in our recent publication: Thapliyal et al. 2023 (PMID: 37236963). Our choice of using 1hr as a baseline is primarily driven by the fact that an elevated baseline of spontaneous reversals on food decreased the dynamic range to monitor changes in heat-evoked reversals. 1-hour early food deprivation reduced spontaneous reversals and led to a mild attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. Additionally, we observed a clear progressive decrease in heat-evoked reversals with increased duration of food-deprivation, which we further used to dissect the mechanism underlying this plasticity. The choice of the 1hr food deprivation timepoint as a reference is further justified below.

      Finally, a previous report (Yeon et al, 2021) demonstrated differences in the impact of chronic versus acute neural silencing on starvation-dependent plasticity in the context of negative thermotaxis. We therefore wonder whether similar developmental compensation impacts the neural circuits that contribute to starvation-dependent plasticity in the thermonociceptive responses.

      Indeed, this is an interesting question. Our conclusions are so far based on ablation (with chronic effects). In order to gain insight on this question, future studies could address the impact of chronic vs acute silencing approach in starvation-dependent thermonociceptive responses. We have added an opening on this question in the discussion, as follows:

      “Another open question is whether ASI action takes place during development (prior to starvation), or more acutely with active signaling after starvation.”

      A weakness of this manuscript is that the introduction is insufficiently scholarly in terms of citations and the description of current knowledge surrounding the impact of internal state on sensory behavior, particularly given previous work on the impact of feeding state on thermosensory behavioral plasticity (Takeshi et al 2020, Yeon et al 2021) and chemosensory valence (Banerjee et al 2023, Rengarajan et al 2019, etc).

      To address this weakness, we have revised the introduction section of the manuscript and cited previous relevant research on the impact of internal states on animal behavior, including the papers suggested by the reviewer. We note that 2 out of 4 suggested citations were already present in the initial manuscript (though in the discussion section).

      Similarly, the authors' commanding knowledge of the distinction between thermotaxis navigation (especially negative thermotaxis) and thermonociceptive behaviors could be communicated in more depth and clarity to the readers, in order to contextualize this study's new findings within the previous literature.

      As mentioned above, we have deepened this aspect in the discussion section of the manuscript (with a dedicated paragraph). It is quite clear that starvationinduced plasticity in negative thermotaxis and thermonociceptive behaviors engage distinct mechanisms (at least in part). These differences include distinct alterations in AWC calcium activity, role of ASI neurons and INS-1 neuropeptide.

      Nevertheless, this study represents a solid addition to the growing evidence that C. elegans sensory behaviors are strongly impacted by internal states, and that neuropeptidergic signaling plays a key role in mediating behavioral plasticity. To that end, the authors have provided solid evidence of their claims.

      We thank this reviewer for the efforts in evaluating our manuscript, for the positive assessment of our work, and for highlighting some weaknesses which, we believe, have been addressed through the revision.

      Reviewer #2 (Public review):

      In this work, Thapliyal and Glauser tried to provide a mechanistic understanding by which animals modulate their neural circuit responses to control nociceptive behavior on the basis of the dynamic internal feeding state. It is an important study that adds to the growing body of evidence coming from multiple model systems. They have used elegant genetics, behavioral, and Ca-imaging experiments to demonstrate how the auxiliary thermosensory neuron pair, AWC, and one of the internal state-sensing interneuron pairs, ASI, respond to dynamic internal starvation state to modulate behavioral response to noxious heat. Interestingly, these neuron pairs use distinct molecular mechanisms along with some other unidentified neurons to suppress heat-induced reversal response under short-term and prolonged starvation. The experiments are well performed, supporting most of the claims and providing an important framework for future studies.

      I have some queries that, if answered, will certainly enhance the study.

      (1) The results suggest that ASI is one of the primary drivers for the starvation-evoked behavioral plasticity, which regulates AWC activity under prolonged starvation. It raises many important questions, including: (a) how starvation modulates ASI response to heat?, and (b) under prolonged starvation, whether ASI also promotes other, non-AWC, glutamatergic inhibitory neurons to suppress heat-induced reversal, and how?

      We agree with this reviewer that the mechanisms by which ASI detects and mediates starvation-evoked changes in our model is a very interesting (unsolved) question. However, addressing these questions empirically represents a substantial body of work that would go beyond the scope of the present report. E.g., is temperature-dependent activity in ASI even relevant? At present, we envision that ASI could either work acutely (during heat stimuli) or be modulated over much longer time frames (hours of starvation) as an internal state sensor. Therefore, there will be quite some exploration needed before we figure out the ASI-level regulation more fully (including the critical temporal aspect regarding cell activity, as well as quantitative and qualitative transmission aspects). It will be very interesting in future work to address these questions.

      (2) How does ASI regulate AWC activity? In the proposed model (Figure 8) authors suggested an independent, unknown signal, other than INS-32 and NLP-18, from ASI to regulate AWC activity. However, from the results, the existence of another signal is not very clear.

      Thanks for raising this point, which reveals a weakness in our graphical representation (in Fig. 8) that was not properly conveying our point. Our current work shows INS-32 and NLP-18 to be important in modulating heat-evoked reversals upon starvation. However, at the moment, we don't know if INS-32, NLP-18, both, and/or other neuropeptides from ASI modulate AWC activity patterns. The calcium imaging experiments in single, double and potentially triple mutants would answer these questions but are not within our current reach, given the time needed to carry out these experiments. However, we acknowledge this point and have changed the figure and its legend to state that the arrow connecting ASI to AWC activity pattern could potentially reflect the action of these neuropeptides.

      (3) Previously, Takeishi et. al. showed that ins-1 dynamically modulates AWC-AIAmediated thermotaxis behavior based on the feeding state of the animal. It raises questions whether ins-1 also contributes to noxious heat-induced reversal behavior.

      We thank the reviewer for this question. We have now quantified the phenotype of ins-1 mutant in our paradigm. Our data shows that INS-1 neuropeptide is not critical in mediating starvation-evoked thermonociceptive plasticity, unlike plasticity in thermotaxis behavior (See Figure 6- Supplement 1). Together with the differential activity patterns in AWC and the differential need for ASI neurons, these new data further consolidate the notion that starvation-evoked thermotaxis adaptation and noxious-heat avoidance engage separable molecular, cellular and circuit-level modulatory mechanisms. A specific discussion paragraph was added too.

      (4) Experiments with AWC fate conversion mutants (nsy-1 and nsy-7) were very good ideas; however, the results obtained were confusing. flp-6 mutant data suggest AWCoff would be essential for heat-induced reversal, especially at the low intensity stimulus level. However, the nsy-1 mutant-forming two AWCon neurons showed complete rescue at the low heat level, which is quite opposite. Similarly, although less prominent, eat-4 rescue experiments suggested both nsy-1 and nsy-7 should behave normally at high heat conditions, which was not the result observed.

      We appreciate this comment and the legit attempt to infer what we should expect from a worm with two AWCon or two AWCoff, respectively. From previous studies so far, it's not quite clear if cellular properties of newly formed AWCs in nsy-1 and nsy-7 mutants, including response to sensory cues, formed synapses and their partners, expression of neuromodulator and gap junctions, synaptic output are similar or different. We think further studies are required to first establish if FLP-6 and glutamate signaling (expression, release and action) from altered AWCs in nsy-1 and nsy-7 mutants are the same or different. Therefore, direct comparison between cell fate conversion mutants with flp-6 and glutamate would rely on too many assumptions at this stage. Considering this comment, the limited additional value of the data with nsy1 and nsy-7 mutants (in the absence of additional analyses) and the confusion it could trigger, we have decided to remove these non-essential data of the manuscript.

      Reviewer #3 (Public review):

      Summary:

      Thapliyal and Glauser show that hunger alters how C. elegans responds to noxious thermal stimuli. Using targeted neural ablation, mutant analysis, and live-cell functional imaging, the authors demonstrate that hunger changes the properties of AWC sensory neurons, which sense noxious heat. The authors further show that the effects of hunger on nociception require ASI neurons, which are known to respond to hunger and mediate the effects of food deprivation on behavior. Finally, the study uses mutant analysis to implicate glutamate and specific neuropeptides in thermal nociception and in the modulation of nociceptors by hungerresponsive neurons.

      Strengths:

      The study clearly shows a strong effect of hunger on nociception and documents a striking effect of hunger on the intrinsic properties of AWC sensory neurons, which respond to noxious heat. The study also clearly and compellingly demonstrates that ablation of hunger-responsive ASI neurons blocks the effects of hunger on nociceptive AWCs. These data, which constitute the kernel of the manuscript, are striking and exciting.

      Weaknesses:

      The study has some weaknesses that the authors should address.

      (1) Ablation of AWC neurons alters the basal sensitivity to noxious heat stimuli. This should be clearly noted in the description of the result and warrants some discussion.

      We thank this reviewer for raising this legitimate point. We have clarified this aspect in the results section of the revised manuscript, reading as follows:

      “Removal of AWC nearly abolished heat-evoked reversal behavior across all stimulus intensities and timepoints (Figure 2B and E). While one should keep in mind that potential indirect developmental effects might take place in neuro-ablation lines, this observation suggests that AWC plays an essential role in mediating the thermonociceptive response under both early food deprivation and prolonged starvation. Notably, in AWC-ablated animals, the residual response level was unaffected by starvation, suggesting that AWC might also be required for the expression of starvation-dependent plasticity.”

      The contrast with known function of the best-characterized sensory neurons mediating thermal nociception (AFD and FLP) is discussed as follows:

      “...Therefore, noxious heat-evoked activity in AWC varies widely according to context, which is in line with previous literature [18, 20, 41]. Interestingly, the role of AWC is distinct from that of AFD and FLP neurons, which are canonically linked to thermosensation and nociception [5, 14, 42, 43], but contribute only modestly to heat-evoked behavior in our assay conditions with between 1 and 6 hrs of food deprivation.”

      (2) Throughout the study, it seems that data are replotted in multiple figure panels. The authors should clearly indicate in the figure legends when this occurs. Also, the authors should ensure that statistical tests requiring multiple comparisons are correctly implemented and reflect the number of times experimental data are compared to a single set of control data.

      Thanks for raising this important point. We have clarified this aspect in the revised figure legends of the manuscript, and in the method section. In some instances, we reconducted some analyses to be perfectly rigorous in multiple comparison accounting. This did not lead to significantly different conclusions. The one exception was that the small effect of eat-4 mutation on spontaneous reversal went below significance threshold. We therefore removed this aspect of the result reporting and of the corresponding interpretation scheme, which became slightly simpler (Figure 3). Globally, this makes the story more focused.

      (3) How ASIs modulate AWCs remains unclear. The authors find that loss of INS-6, an insulin-like peptide provided by ASIs, partially recapitulates the effect of ASI ablation. This observation is not further developed, and instead, the authors characterize other secreted factors that seem to mediate sensitization of animals to noxious heat stimuli. While it is interesting that there are multiple opposing inputs into the nociceptor circuit, the essential connection between ASIs and AWCs that underlies the foundational observations in Figures 1 and 2 is not sufficiently characterized.

      Whereas we agree that how ASI modulates AWCs is only partially solved by our study, we should emphasize that our work identified two ASI-expressed neuropeptides that function to decrease reversal response after starvation: INS-32 and NLP-18. We initially set a lower priority on INS-6 because the reversal response level in starved mutants appeared lower than that in nlp-18 and ins-32. It is important to note that ins-32 and nlp-18 are not ‘generally potentiated’ mutants, but display reversal upregulation selectively following starvation, which placed them as strong candidates to selectively mediate ASI regulation. This said, it is also true that these two mutants (and ins-6 too) display reduced responsiveness at the early food deprivation time point. Therefore, none of the neuropeptide mutants was strictly identical to ASI ablated line, suggesting that the peptides might also work via non-ASI cells at the early food deprivation timepoint.

      Following this reviewer’s comment, we have attempted to complement our story with the idea of using a similar approach and rescue ins-6 with its endogenous promoter or ASI-specific promoter. Unfortunately, we failed to obtain rescue effects, and therefore these data (with a negative result) remain inconclusive (as we cannot guarantee that the rescue constructs were functional). We decided to keep these data aside in the revised manuscript. Globally, our point made graphically in Figure 7F remains valid. We have complemented the figure legend to mention that INS-6 could also potentially work from ASI, but it is not depicted as no ASI-specific data are available. In summary, our data suggests that the connection between ASI and AWC(s) might be established by the integrated action of multiple peptides and their receptors. Further calcium imaging experiments in single, double and potentially triple mutant(s) of peptides and receptors would be required go deeper in this question, which could be performed in future work.

      “...Additional neuropeptides (such as INS-6) may also be involved, but in the absence of direct evidence for their origin from ASI, they were not included in this scheme.”

      (4) The assertion that 'starvation reshapes AWC responses from deterministic to stochastic' is not clearly supported by the data. AWC neurons seem capable of showing different responses to thermal stimuli, and the probabilities associated with these responses change after fasting. The different kinds of responses are seen under basal and fasted conditions.

      We thank this reviewer for the comment. There is an activity response shift that is quite solidly described, including with new quantitative analyses of distributions (histograms in new Fig. 4CD and new Fig. 5C-D, accompanied by Kruskal-Wallis tests). Yet, we totally agree that the wording choice was inappropriate. We have furthermore changed our terminology to avoid using the terms “stochastic” or “deterministic” that were indeed a cause of confusion. We now use the terms “stimulus-locked responses” and describe the shift as “shift from mostly excitatory responses to a mix of both excitatory and inhibitory responses”. We have also included detailed methodology for characterization of traces and statistical analysis in the revised method section of the manuscript, together with the new analyses on peak polarity distribution.

      Recommendations for the authors:

      Reviewing Editor Comments:

      The reviewers agree that the study is clearly presented and makes good use of behavioral, genetic, and imaging approaches to link starvation state with changes in AWC and ASI function. To strengthen the manuscript and ensure clarity for readers, we ask you to address the following points in revision:

      (1) Positioning and citations.

      Clarify how your findings relate to Takeishi 2020, where the opposite trend in AWC activity was reported, and make a clear distinction between thermonociception and thermotaxis. The introduction should also include additional citations in two specific areas: prior work on AWC and noxious thermal stimuli, and studies demonstrating starvation-dependent behavioral changes via altered neuropeptide release (e.g., Banerjee 2023; Rengarajan 2019).

      We have clarified this aspect with extension of the work cited in the introduction and extensive rewriting of the discussion sections.

      Our data shows that mechanisms underlying starvation dependent changes in thermonociception and thermotaxis show differences at the molecular, cellular and circuit levels. First, we see a decrease in average response in AWC neurons due to shift from mostly excitatory to a mix of excitatory and inhibitory responses in response to noxious heat after starvation, while previous study found an increase in response probability of AWCs after starvation (Takeishi et al, 2020). Second, contrary to thermotaxis behavior ASI neurons are required to orchestrate starvation evoked plasticity in thermonociception. And, finally, previous study found that INS-1 signaling from the intestine regulates thermotaxis behavioral plasticity while INS-1 peptide does not seem critical in regulating starvation-dependent thermonociceptive plasticity (new data in Figure 6 supplement 1).

      We have revised the introduction section of the manuscript and cited previous relevant research on AWC and noxious thermal stimuli and studies demonstrating starvationdependent behavioral changes via altered neuropeptide release including the papers suggested by reviewers. The extended discussion section reads as follows:

      “Starvation regulates thermonociceptive and negative thermotaxis plasticity via at least partly different mechanisms

      Previous studies showed that AWC plays an important role in starvationdependent plasticity in the negative thermotaxis behavior in an innocuous thermal range between 15 and 25°C [26, 33]. Negative thermotaxis involves the detection of thermal changes created by animal movement in spatial thermogradient (0.5°C/cm), the magnitude of the expected thermal changes approximating 0.01°C/s [26]. The starvation impact on negative thermotaxis was shown to (i) involve an up-regulation of AWC cell activity, (ii) rely on INS1 neuropeptide produced in the intestine and (iii) to occur independently of ASI neurons. In contrast, our study used thermo-nociceptive stimuli, with faster raising thermal slopes (~0.5-2°C/s, hence 50-200 times faster than those occurring for thermotaxis) and covering noxious temperatures (up to 28°C). Our results indicate that the regulation of thermo-nociceptive response by starvation (i) is linked to a shift in the distribution of AWC activity response polarities from mostly excitatory to a mix of excitatory and inhibitory response, (ii) relies on ASI and specific neuropeptide produced in ASI, and (iii) works independently of INS-1 neuropeptide. Therefore, our study complements our understanding of the modulation of temperature-dependent behavior in C. elegans with previously undocumented mechanisms at the circuit, cellular and molecular levels.”

      (2) ASI → AWC mechanism.

      Because ASI is central to your conclusions, please expand on how ASI is thought to act on AWC and/or other neurons. If an additional ASI signal is proposed beyond INS-32/NLP-18, mark this as speculative unless further rationale can be provided, and adjust the model figure accordingly.

      We have revised Figure 8 and its legend to clarify what is still hypothetical in the way ASI could affect AWC activity patterns and reversals. Note that the figure was also modified to integrate the conclusions made from epistasis analysis of eat-4 and flp-6.

      (3) AWC response description.

      The data support a shift in response distributions rather than a categorical switch from "deterministic to stochastic." Please adjust the language accordingly and provide a clear description of how traces were classified as "up, variable, or down," ideally with a quantification of the distributional shift.

      We agree that the term “stochastic” can convey different things, and because it was used for something different for AWC in the past, we should have avoided it. We have revised the nomenclature. What we observe can indeed be better described as a shift in the response polarity distribution. The article was revised accordingly. We also included the quantitative analysis and histogram representation, suggested in one of the specific comments, and added detailed methodology on the categorization of traces.

      When ASI is intact, we see a shift from mostly excitatory responses to an ~equal mix of excitatory and inhibitory responses (new Figure 4C-D). This effect is lost when ASI is ablated (new Figure 5C-D).

      (4) Presentation and statistics.

      In figure legends, indicate where datasets are replotted across panels and confirm that multiple-comparison corrections take account of repeated comparisons to the same controls.

      We have included these details in the revised figure legends, and a statement in the method section.

      (5) Methods clarity.

      Provide justification for using 1-h off-food as the baseline, with quantification of baseline mobility/reversal rates. Expand the calcium-imaging methods to describe the processing pipeline (ΔR calculation, baseline period, drift correction), and add a brief rationale if the approach deviates from common normalization procedures. Clarify how cell ablations were performed and verified for specificity, and how cell-specific rescues were confirmed. Please also acknowledge the potential for developmental compensation with chronic ablation.

      The justification of using 1hr off-food as baseline was made more prominent in the revised manuscript.

      Revised result section:

      “Starvation downregulates thermonociceptive responses in C. elegans

      To assess how the feeding state modulates thermonociceptive behavior in C. elegans, we compared responses across different durations of food deprivation (Figure 1A). Synchronized first-day adult animals were stimulated with a series of 4-s infrared pulses of increasing heating power (100, 200, 300, 400 W), causing temperature increase of +2°C, +4°C +6°C and 8°C at the surface of the plate (Figure 1A). Fed animals on food produced robust heat-evoked reversal response to heat, but they also displayed a very elevated baseline of spontaneous reversals (~38%). A 1-hour off-food condition reduced spontaneous reversals (from ~38% to ~10%) and led to an attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. More prolonged food deprivation led to a striking progressive reduction in thermonociceptive responses at every heating level, with responses after 6 hours of starvation approaching baseline spontaneous reversal rates (Figure 1B and C). This suggests a robust inhibition of nociceptive behavior caused by prolonged starvation. To determine whether this attenuation was due to the absence of nutrients or chemosensory cues, we conducted similar starvation experiments in the presence of food odor, with OP50 bacteria present on the petri dish lid (Figure 1D). The reduction in thermonociceptive response persisted, indicating that the effect is driven by the internal starvation state rather than external olfactory input.

      Although fed animals showed high sensitivity to noxious heat, they also displayed an elevated baseline of spontaneous reversals, which limited their utility as a control group by strongly reducing the dynamic range of heat-evoked reversal quantification and by complicating the quantitative comparison with food-deprivation conditions with much-reduced reversal baseline (Figure 1B). In addition, technical limitations in our calcium imaging setup would have prevented the intended follow-up analyses in fed animals. Based on these observations and technical considerations, we focused subsequent analyses, aiming at dissecting the circuit and molecular underpinnings of starvation-dependent plasticity, to the comparison of two off-food conditions with similar spontaneous reversal baseline: the early food deprivation condition (1-hour off-food, with high responsiveness to noxious heat) and the prolonged starvation (6-hour off-food with almost abolished noxious heat responsiveness).”

      In addition, the method section was modified as follows:

      - Calcium imaging details were added regarding ΔR calculation, baseline period, drift correction.

      - We now explicitly refer to the original respective articles describing the neuroablation lines.

      - We clarify that cell-specific transgene expression for rescue was confirmed using SL2::mCherry co-marker

      In the result section, we now explicitly address potential developmental compensation in genetic ablation backgrounds in the result section as follows: “…one should keep in mind that potential indirect developmental effects might take place in neuro-ablation lines”.

      Reviewer #1 (Recommendations for the authors):

      (1) The data availability statement is missing from the reviewed manuscript and should be included.

      Thanks, we have included the data availability statement in the revised manuscript.

      (2) We request additional information on how n's were determined for individual experiments, as well as the inclusion of post-hoc power measurements for all quantification.

      n were determined in agreement with previous studies using similar measures. No a priori power analyses were performed. A posteriori power analyses are not informative beyond the reported effect sizes and p-values (now reported in File S2). We clarified this in the statistical subsection of the method section.

      (3) In many cases, the specific statistical tests used are not clear or justified; more details should be provided, including the non-post-hoc test used. Are all tests one-way ANOVAs? For comparisons across genotype and starvation duration, two-way ANOVAs would likely be more appropriate. Also, the authors switch between Bonferroni post-hoc tests and Holm-Bonferroni post-hoc tests. What determined the use of one versus another?

      We have now clarified the statistical analyses used and provided full details in File S2. We have now more systematically applied two-way ANOVAs across all relevant analyses (with detailed parameters reported in File S2). When particularly relevant (e.g epistasis analysis between eat-4 and flp-6 mutations) the results of the two-way ANOVAs, is also explicitly stated in the result section.

      We also note that all multiple-comparison corrections were performed using the Bonferroni method. Previous mentions of Holm-Bonferroni correction were inaccuracies, and we apologize for this confusion; these mentions have now been corrected throughout the manuscript.

      (4) The use of heating power instead of the temperature experienced by the worms is an unwelcome abstraction. We strongly recommend revisiting that choice.

      We do agree. We have revised the figures to label the axis with temperature increase.

      (5) For calcium imaging, how are the traces categorized into "calcium up", "calcium down", or "no change"? Were those determined blindly - i.e., by individuals unaware of the experimental condition? Did the response direction need to be consistent across different temperatures? Did the change from baseline need to hit a specific threshold, consistent with previous studies in the field (i.e., +/- 3xSD for a minimum amount of time)? We encourage the authors to include these details in their methods section.

      We have complemented the method section to clarify the criteria for the qualitative classification of traces. More importantly, new quantitative peak polarities comparisons were added (see specific points below and above, about histograms).

      (6) For the experiments showing that exposure to food odor does not prevent response reduction, we suggest that feeding worms heat-killed bacteria would be a helpful control for the importance of bacterial nutritional status. In addition, showing that the impact of starvation was reversible with re-feeding would have been a useful experiment in line with standard experimental design in the starvation field.

      Thanks, indeed with our current work we cannot pinpoint the role of additional sensory cues (except food odor) to be mediating starvation-evoked plasticity. Together with refeeding, these are all extremely interesting questions that we aim to answer and potentially link with ASI and AWC activity in our future work.

      (7) For the various AWC rescue experiments, we found it curious that there wasn't an AWCon+off rescue, only each neuron individually.

      Previous studies have identified similar or opposite responses of both AWCs for distinct sensory cues. Though our calcium imaging experiments point to both AWC on and off having similar response patterns to heat, we cannot rule out the possibility that their output (ability of control reversals) is distinct possibly due to recruited neuromodulators. Therefore, in the present work, we examined where these cell types act via the same or distinct combinations of neuromodulators to control reversals.

      Reviewer #2 (Recommendations for the authors):

      Experiments suggested:

      (1) The authors should look into the Ca-dynamics in ASI.

      How does the spontaneous and heat-evoked activity of ASI differ in fed, early food-deprivation and prolong starvation and its link to releases of neuromodulators, modified AWC activity to alter output of thermal nociception are very interesting questions. However, these questions are extremely exploratory (see argumentation above in response to the public review) and addressing them goes beyond the scope of our current manuscript.

      (2) The authors should check AWC activity in ins-32 and nlp-18 mutant animals.

      In this study, we focused on the roles of INS-32 and NLP-18 released from ASI in modulating heat-evoked reversals, as these mutants exhibit relatively strong behavioral phenotypes. However, these effects remain less pronounced than those observed following ASI ablation. In addition, we cannot exclude the contribution of additional signaling molecules, including INS-6 and other neuropeptides.

      A comprehensive analysis of AWC activity in this context would require calcium imaging across multiple genetic backgrounds, including single, double, and potentially higher order peptide and receptor mutants, combined with cell-specific rescue experiments. While we appreciate the suggestion, such an approach would represent a substantial extension of the present work and will be important to pursue in future studies to further elucidate the underlying mechanisms.

      (3) Short-term food deprivation completely eliminated heat heat-induced reversal response to 100W stimulus, while the response to 400W stimulus remained unaffected. This suggests fed, short-term starvation, and prolonged starvation are three distinct states, and authors should also test the response of AWC and ASI ablated animals in the fed conditions.

      We agree that analyzing thermal nociception in fed states, in addition to short-term and prolonged food deprivation states is an important and interesting question, as these 3 conditions likely represent 3 distinct internal states that may recruit different neural pathways.

      Several reasons led us to set the fed condition aside for this study, and we realize we insufficiently explain them in the initial manuscript. There are 2 main reasons.

      (1) It is of paramount importance to consider the ‘baseline’ reversal rate (spontaneous reversals not triggered by heat, but visible in our dataset as the first point in the ‘dose-response’ curve). In Fed animals spontaneous reversal rate is very high (~38%) compared to the 1hr and 6hr food-deprivation conditions (>10%). This has two consequences: first a decreased dynamic range for quantify heat-evoked reversal, and, second, the difficulty in judging quantitative differences in heat-evoked reversals with such major differences in baseline reversals.

      (2) Experimentally, assessing calcium responses in truly fed animals presents technical challenges. With our current setup, animals must be removed from food for at least ~5 minutes prior to recording (followed by ~5 minutes of imaging), which effectively corresponds to a “freshly starved” condition rather than a fully fed state. While previous studies have used serotonin to mimic aspects of the fed state, such manipulations can be difficult to interpret in this context.

      A systematic comparison including fully fed animals, as well as AWC- and ASI-ablated conditions across these states, would be a valuable direction for future work, in particular once the methodological barriers associated with point 2, have been overcome.

      We have clarified these choices in the result section as follows:

      “Starvation downregulates thermonociceptive responses in C. elegans

      To assess how the feeding state modulates thermonociceptive behavior in C. elegans, we compared responses across different durations of food deprivation (Figure 1A). Synchronized first-day adult animals were stimulated with a series of 4-s infrared pulses of increasing heating power (100, 200, 300, 400 W), causing temperature increase of +2°C, +4°C +6°C and 8°C at the surface of the plate (Figure 1A). Fed animals on food produced robust heat-evoked reversal response to heat, but they also displayed a very elevated baseline of spontaneous reversals (~38%). A 1-hour off-food condition reduced spontaneous reversals (from ~38% to ~10%) and led to an attenuation of heat-evoked responses at low stimulus intensities, while responses to stronger stimuli remained comparable to those of fed animals. More prolonged food deprivation led to a striking progressive reduction in thermonociceptive responses at every heating level, with responses after 6 hours of starvation approaching baseline spontaneous reversal rates (Figure 1B and C). This suggests a robust inhibition of nociceptive behavior caused by prolonged starvation. To determine whether this attenuation was due to the absence of nutrients or chemosensory cues, we conducted similar starvation experiments in the presence of food odor, with OP50 bacteria present on the petri dish lid (Figure 1D). The reduction in thermonociceptive response persisted, indicating that the effect is driven by the internal starvation state rather than external olfactory input.

      Although fed animals showed high sensitivity to noxious heat, they also displayed an elevated baseline of spontaneous reversals, which limited their utility as a control group by strongly reducing the dynamic range of heat-evoked reversal quantification and by complicating the quantitative comparison with food-deprivation conditions with muchreduced reversal baseline (Figure 1B). In addition, technical limitations in our calcium imaging setup would have prevented the intended follow-up analyses in fed animals. Based on these observations and technical considerations, we focused subsequent analyses, aiming at dissecting the circuit and molecular underpinnings of starvationdependent plasticity, to the comparison of two off-food conditions with similar spontaneous reversal baseline: the early food deprivation condition (1-hour off-food, with high responsiveness to noxious heat) and the prolonged starvation (6-hour off-food with almost abolished noxious heat responsiveness).”

      (4) The authors should test the effect of ins-1 in noxious heat-mediated dynamic reversal behavior.

      We thank the reviewer for this valuable suggestion. We have now tested the phenotype of ins-1 mutants in our paradigm. Our data shows that INS-1 neuropeptide is not critical in mediating starvation-evoked thermonociceptive plasticity unlike plasticity in thermotaxis behavior (New Figure 6 Sup1). This molecular aspect adds to our initially presented evidence at the cell activity and circuit levels, that noxiousevoked reversal and thermotaxis behaviors are regulated in a clearly separable manner.

      (5) Whether Glutamate and flp-6 work in parallel or in the same pathway to regulate reversals?

      We thank the reviewer for this question. We tested eat-4; flp-6 double mutants and found:

      (1) After short-term food deprivation (1hr), flp-6 and eat-4 separately contribute to heat-evoked reversal at high & low heat and they act in parallel pathways to explain ~90% of animal responsiveness (new version of Fig 3)

      Corresponding new text:

      “Next, we focused on eat-4 and flp-6 mutants, showing the strongest phenotype. We addressed whether glutamate and FLP-6 signaling act dependently of each other in controlling heat-evoked reversal, by testing eat-4; flp-6 double mutants. The residual response seen in each single mutant (Figure 3 A and B) was almost entirely abolished in the double mutant (Figure 3C). A two-way ANOVA for the highest heat stimuli with eat-4 and flp-6 genotypes as factors (two levels each: mutant or wild type) showed significant main effects of eat-4 (F<sub>(1,67)</sub> =48.70, p<.001, η<sup>2</sup>p=0.421) and flp-6 (F<sub>(1,67)</sub> =58.50, p<.001, η<sup></sup>p=0.466), respectively, but no interaction effects (F<sub>(1,67)</sub> =0.093, p=.761, η<sup>2</sup>p=0.001). The significant cumulative effect of the two mutations indicates that the two signaling pathways act mostly independently of each other to mediate heat-evoked reversals.”

      (2) After prolong starvation (6hr), flp-6 mutation has a dominant impact on plasticity and eat-4 mutation cannot cause loss of plasticity, pointing to a hierarchy in this context (new version of Fig. 6, including revised hierarchy in the model in panel G, and also revised model in Fig.

      8).

      Corresponding revised text:

      “Second, we tested whether starvation-dependent plasticity was preserved in eat-4 and flp6 mutant backgrounds, which we had suggested to represent the main AWC transmitters controlling heat-evoked reversals under the early food deprivation condition (Figure 3). Even if the heat-evoked response upon early food-deprivation was reduced relative to wild type in flp-6 mutants, a significant further decline was seen after prolonged starvation (Figure 6B). These results indicate that starvation-dependent plasticity can operate independently of FLP-6. In contrast, eat-4 mutants displayed markedly elevated heat-evoked responses after prolonged starvation, even exceeding the response level seen in the early food deprivation condition for low heat stimuli (Figure 6C). This potentiated response in eat-4 mutants was entirely dependent of an intact FLP-6 signaling, since reversal responses in eat-4; flp-6 mutants were entirely abolished, like in flp-6 single mutant (Figure 6C-E, a two-way ANOVA indicating a significant interaction effect between the two mutations: F<sub>(1,70)</sub> =0.093, p<.001, η<sup>2</sup>p=0.247). Moreover, the potentiated response in eat-4 single mutant could not be rescued by expressing eat-4 rescue transgene selectively in either AWC<sup>OFF</sup> or AWC<sup>ON</sup> neurons (Figure 6F). Interestingly, AWC<sup>OFF</sup>-specific rescue produced a further potentiation of heat-evoked reversal response to high heat stimuli (Figure 6F, 6 and 8°C thermal increases), aggravating the phenotype of eat-4 mutants. These results are consistent with a model in which glutamatergic signaling regulates heat-evoked reversals in starved animals via two bidirectional drives (Figure 6G). On the one hand, glutamatergic signaling—originating from AWC<sup>OFF</sup>—up-regulates reversals in response to high heat stimuli, thus contributing to prevent starvation-induced thermonociceptive plasticity. On the other hand, glutamatergic signaling—originating from neurons other than AWC— down-regulates reversals over a broad range of heat intensities, thus promoting starvation-induced thermonociceptive plasticity. The latter glutamatergic signaling inhibitory effect seems to be more dominant and to depend on intact FLP-6 signaling.”

      (6) The authors should perform flp-6 and eat-4 mutant/rescue experiments in the nsy-1 and nsy-7 background to clarify the results.

      Our results indicate that glutamate release via EAT-4 from both AWC<sup>ON</sup> and AWC<sup>OFF</sup>, as well as FLP-6 from AWC<sup>OFF</sup>, contributes to heat-evoked reversals regulation. The experiments suggested by the reviewer would, in principle, provide further insight into the interaction between AWC subtype identity and the respective roles of glutamatergic and peptidergic signaling.

      However, as discussed in more details above in the public review, the extent and nature of AWC<sup>ON/OFF</sup> remodeling in nsy-1 and nsy-7 mutant backgrounds remain incompletely understood. This introduces significant uncertainty in interpreting results obtained from combining these mutations with eat-4 and flp-6 manipulations. As a result, such experiments would be difficult to interpret in a definitive manner at this stage.

      We therefore consider this an important direction for future work, once the roles of nsy1 and nsy-7 in AWC subtype specification and function are more clearly established. As our preliminary results with nsy-1 and nsy-7 mutants added more confusion than clarity, we have chosen to set them aside (former Fig. 3-figure supplement 2 has been removed).

      Minor comments:

      (1) What is food odor? The experiment should be clearly mentioned.

      Thanks for spotting this unintended omission. Food odor experiments were performed by adding OP50 bacteria on the inward side of the petri dish lid instead of the NGM surface. We have added this description in the method section of the revised manuscript.

      (2) Panel 3C is coming before 3B. This should be rearranged.

      Thanks for pointing this out. Panel arrangement was entirely reorganized in revise Fig. 3, with the addition of eat-4 x flp-6 genetic interaction analysis.

      (3) In Figure 6, if the panels are arranged horizontally, it would be easier to follow.

      Thanks for pointing this out. Panel arrangement was entirely reorganized in revise Fig. 6, with the addition of eat-4 flp-6 genetic interaction analysis.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors should consider moving measurements of AFD-ablated animals into the main Figure 1. AFD is a well-known thermosensor, and it is worth showing that responses to noxious thermal stimuli persist in animals lacking AFD.

      Thank you for this suggestion. We have moved measurements of AFDablated animals to the main figure (revised Fig. 2).

      (2) AFD ablation does affect responses to noxious heat. The authors could consider ablating/silencing AFDs and AWCs simultaneously to determine whether these two neurontypes account for the behavior.

      We agree that investigating the combinatorial contributions of thermosensory neurons, including AFD and AWC, to thermal nociception is an important and interesting question. In principle, simultaneous ablation or silencing of these neuron types could indeed reveal unexpected interactions.

      In our experimental paradigm, however, we observe only a minimal contribution of AFD neurons to heat-evoked reversals, whereas ablation of AWC nearly abolishes the response. Based on these observations, we chose to focus the present study on AWC, which appears to play a more prominent role in this behavior.

      A more detailed dissection of the potential interactions between AFD and AWC, including combinatorial manipulations, would be a valuable direction for future work. In particular, the possibility that AFD exerts a modulatory influence remains an interesting hypothesis to explore.

      (3) Given that EAT-4/VGLUT and FLP-6 neuropeptides each contribute to nociception, the authors should consider testing an eat-4; flp-6 double mutant to determine whether this combination of neurochemical signals accounts for AWC signaling to downstream circuits.

      We thank the reviewer for this suggestion, which is similar to point 5 of Reviewer 2 (above).

      We tested eat-4; flp-6 double mutants and found:

      (1) After short-term food deprivation (1hr), flp-6 and eat-4 separately contribute to heat evoked reversal at high & low heat and they act in parallel pathways to explain ~90% of animal responsiveness (new version of Fig 3)

      Corresponding new text:

      “Next, we focused on eat-4 and flp-6 mutants, showing the strongest phenotype. We addressed whether glutamate and FLP-6 signaling act dependently of each other in controlling heat-evoked reversal, by testing eat-4;flp-6 double mutants. The residual response seen in each single mutant (Figure 3 A and B) was almost entirely abolished in the double mutant (Figure 3C). A two-way ANOVA for the highest heat stimuli with eat-4 and flp-6 genotypes as factors (two levels each: mutant or wild type) showed significant main effects of eat-4 (F<sub>(1,67)</sub> =48.70, p<.001, η<sup>2</sup>p=0.421) and flp-6 (F<sub>(1,67)</sub> =58.50, p<.001, η<sup>2</sup>p=0.466), respectively, but no interaction effects (F<sub>(1,67)</sub> =0.093, p=.761, η<sup>2</sup>p=0.001). The significant cumulative effect of the two mutations indicates that the two signaling pathways act mostly independently of each other to mediate heat-evoked reversals.”

      (2) After prolong starvation (6hr), flp-6 mutation has a dominant impact on plasticity and eat-4 mutation cannot cause loss of plasticity, pointing to a hierarchy in this context (new version of Fig. 6, including revised hierarchy in the model in panel G, and also revised model in Fig. 8).

      Corresponding revised text:

      “Second, we tested whether starvation-dependent plasticity was preserved in eat-4 and flp6 mutant backgrounds, which we had suggested to represent the main AWC transmitters controlling heat-evoked reversals under the early food deprivation condition (Figure 3). Even if the heat-evoked response upon early food deprivation was reduced relative to wild type in flp-6 mutants, a significant further decline was seen after prolonged starvation (Figure 6B). These results indicate that starvation-dependent plasticity can operate independently of FLP-6. In contrast, eat-4 mutants displayed markedly elevated heat-evoked responses after prolonged starvation, even exceeding the response level seen in the early food deprivation condition for low heat stimuli (Figure 6C). This potentiated response in eat-4 mutants was entirely dependent of an intact FLP-6 signaling, since reversal responses in eat-4; flp-6 mutants were entirely abolished, like in flp-6 single mutant (Figure 6C-E, a two-way ANOVA indicating a significant interaction effect between the two mutations: F<sub>(1,70)</sub> =0.093, p<.001, η<sup>2</sup>p=0.247). Moreover, the potentiated response in eat-4 single mutant could not be rescued by expressing eat-4 rescue transgene selectively in either AWC<sup>OFF</sup> or AWC<sup>ON</sup> neurons (Figure 6F). Interestingly, AWC<sup>OFF</sup>-specific rescue produced a further potentiation of heat-evoked reversal response to high heat stimuli (Figure 6F, 6 and 8°C thermal increases), aggravating the phenotype of eat-4 mutants. These results are consistent with a model in which glutamatergic signaling regulates heat-evoked reversals in starved animals via two bidirectional drives (Figure 6G). On the one hand, glutamatergic signaling—originating from AWC<sup>OFF</sup>—up-regulates reversals in response to high heat stimuli, thus contributing to prevent starvation-induced thermonociceptive plasticity. On the other hand, glutamatergic signaling—originating from neurons other than AWC— down-regulates reversals over a broad range of heat intensities, thus promoting starvation-induced thermonociceptive plasticity. The latter glutamatergic signaling inhibitory effect seems to be more dominant and to depend on intact FLP-6 signaling.”

      (4) The authors should consider representing AWC responses to thermal stimuli as histograms to illustrate how fasting increases the probability of some responses and decreases the probability of others.

      We thank this reviewer for the suggestion. The proposed histograms nicely convey the concept of “shift in response polarity distribution” that we observed (using the new terminology we now use, instead of using the term “stochastic”). To create such histograms, we computed the magnitude of the peaks on a trial-by-trial basis. When ASI is intact, we see a shift from mostly excitatory responses to an ~equal mix of excitatory and inhibitory responses (new Figure 4C-D). This effect is lost when ASI is ablated (new Figure 5C-D).

      (5) It seems important to better understand the ins-6 mutant phenotype and determine whether ASI-to-AWC signaling involves this insulin-like peptide (ILP). The authors should consider using some of the tools available for disrupting ILP signaling to more clearly demonstrate that a specific neurochemical signal mediates modulation of AWCs by ASIs.

      Our work identified two ASI-expressed neuropeptides that function to decrease reversal response after starvation: INS-32 and NLP-18. We initially set a lower priority on INS-6 because the reversal response level in starved mutants appeared lower than that in nlp-18 and ins-32. It is important to note that ins-32 and nlp-18 are not ‘generally potentiated’ mutants, but display reversal up-regulation selectively following starvation, which placed them as strong candidates to selectively mediate ASI regulation. This said, it is also true that these two mutants (and ins-6 too) display reduced responsiveness at the early food deprivation time point. Therefore, none of the neuropeptide mutants was strictly identical to ASI, suggesting that the peptides might also work via non-ASI cells at the early food deprivation timepoint (as follow up data indicated at least for nlp-18).

      Following this reviewer’s comment (and the similar one in the public review), we have attempted to complement our story with the idea of using a similar approach and rescue ins-6 with its endogenous promoter or ASI-specific promoter. Unfortunately, we failed to obtain rescue effects, and therefore these data remain inconclusive (as we cannot guarantee that the rescue constructs were functional). We decided to keep these data aside in the revised manuscript. Globally, our point made graphically in Figure 7F remains valid. We have complemented the figure legend to mention that INS6 could also potentially work from ASI, but it is not depicted as no ASI-specific data are available. In summary, our data suggests that the connection between ASI and AWC(s) might be established by the integrated action of multiple peptides and their receptors. Further calcium imaging experiments in single, double and potentially triple mutant(s) of peptides and receptors would be required to go deeper in this question, which could be performed in future work.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      The authors investigated the response of worms to the odorant 1-octanol (1-oct) using a combination of microfluidics-based behavioral analysis and whole-network calcium imaging. They hypothesized that 1-oct may be encoded through two simultaneous, opposing afferent pathways: a repulsive pathway driven by ASH, and an attractive pathway driven by AWC. And the ultimate chemotactic outcome is likely determined by the balance between these two pathways.

      It is not surprising that 1-octanol is encoded as attractive at low concentrations and repulsive at higher concentrations. However, the novel aspect of this study is the discovery of the combinatorial coding of 1-oct in the periphery, where it serves as both an attractant and a repellent. Furthermore, the study uses this dual encoding as a model to explore the neural basis of sensory-driven behaviors at a whole-network scale in this organism. The basic conclusions of this study are well supported by the behavioral and imaging experiments, though there are certain aspects of the manuscript that would benefit from further clarification.

      A key issue is that several previous studies have demonstrated a combinatorial and concentration-dependent coding of odorant sensing in the nematode peripheral nervous system. Specifically, ASH and AWC are the primary receptors for repellent and attractive responses, respectively. However, other neurons such as AWB, AWA, and ADL are also involved in the coding process. These neurons likely communicate with different interneurons to contribute to 1-oct-induced outputs. The authors' conclusion that loss of tax-4 reduces attractive responses and that osm-9 mutants reduce repulsive responses is not entirely convincing. TAX-4 is required for both AWC (an attractive neuron) and AWB (a repulsive neuron), and osm-9 is essential for ASH, ADL, and AWA (attraction-associated). Therefore, the observed effects on the attractive and repulsive responses could be more complex. Additionally, the interpretation of results involving the use of IAA to reduce the contribution of AWC at lower concentrations lacks clarity.

      The authors did not observe any increased correlation between motor command interneurons and sensory neurons, which is consistent with the absence of a consistent relationship between state transitions and 1-oct application. Furthermore, they did not observe significant entrainment of AIB activity with the 2.2 mM 1-oct application. This might be due to the animals being anesthetized with 1 mM tetramisole hydrochloride, which could affect neural activity and/or feedback from locomotion.

      Comments on revisions:

      The authors have addressed all my previously raised concerns.

      Reviewer #2 (Public review):

      Summary:

      The authors used whole-network imaging to identify sensory neurons that responded to the repellant 1-octanol. While several olfactory neurons responded to the initial onset of odor pulses, two neurons consistently responded to all the pulses, ASH and AWC. ASH typically activates in response to repellants, and AWC typically activates in response to the removal of attractants. However in this case, AWC activated in response to the removal of 1-octanol, which was unexpected because 1-octanol is a harmful repellant to the worm. The authors further investigated this phenomenon by testing different concentrations of 1-octanol in a chemotaxis assay, and found that at lower (less harmful) concentrations the odor is actually an attractant, but becomes repulsive at higher concentrations. The amplitude of the ASH response appeared to be modulated by concentration, but this was not true for AWC. The authors propose a model where the behavioral response of the worm is the result of integrating these two opposing drives, where repulsion is a result of the increased ASH activity over-riding the positive drive from AWC. The authors further tested this theory by testing mutants that ablated the AWC response (tax-4 or AWC::HisCl) or ASH response (osm-9 or ASH::HisCl). The chemo-silencing (HisCl) and tax-4 experiments were consistent with their hypothesis, while the osm-9 mutation had a limited impact on chemotaxis behavior, highlighting the potential role of osm-9-independent signaling in ASH in response to 1-octanol. While the interneuron(s) that integrate these signals to influence behavior were not identified, the authors did find that increasing concentrations of 1-octanol did increase the likelihood of AVA activity, a neuron which drives reversals (and hence, behavioral repulsion).

      Strengths:

      This was simple and elegant work that identified specific neurons of interest which generated a hypothesis, which was further tested with mutants that altered neuronal activity. The authors performed both neuronal imaging and behavioral experiments to verify their claims.

      Weaknesses:

      The authors note that other sensory neurons likely contribute to 1-octanol chemotaxis. Given the NeuroPAL data, it would have been nice to identify these other neurons as well. However, the reviewer is aware that this is tangential to the primary focus of this study.

      Reviewer #3 (Public review):

      Summary:

      This work describes how two chemosensory neurons in C. elegans drive opposite behaviors in response to a volatile cue. Because they have different concentration dependencies, this leads to different behavioral responses (attraction at low concentration and repulsion at high concentration). It has been known that many odorants that are attractive at low concentrations are aversive at high concentrations, and the implicated neurons (at least AWC for attraction and ASH for repulsion) have been well established. None the less, by studying behavior and neural responses in a common context (odor pulses, as opposed to gradients) this provides a clear picture of how these sensory neurons may guide the dose dependent response by separately modulating odor entry and odor exit behaviors.

      Strengths:

      (1) This work provides good evidence that worms are attracted to low concentrations and repelled by high concentrations of 1-oct. Calcium imaging also makes it clear that dose-dependence of this response is stronger for ASH than AWC.

      (2) This work presents calcium imaging and behavior with the same stimulus (sudden pulses in volatile odor concentration), while previous studies often focus on using neuronal responses to pulses to understand navigation of gentle gradients.

      Weaknesses:

      (1) As a whole it is not clear precisely how important AWC is (compared to other cells) for the attractive response (as the authors correctly acknowledge).

      (2) The evidence that AIB minus AVA contains relevant information is weak. It appears the entrainment index in Fig. 6H for AIB-AVA could easily be explained by the negative entrainment between AVA and the stimulus (along with no effect or role for AIB). This is suggested by the similar p-values and similar distribution of random EIs (stretched and mirrored) between the first and last rows of this figure.

      (3) The model in Figure 7 would be strengthened if it was demonstrated that IAA is attractive when worms are saturated in a 1/10^4 concentration. Panel 7G (and ref. 39) indicate that 10^-4 IAA activates ASH, which would suggest a different explanation for the change from attraction to repulsion in 7C.

      It was previously published that 1x10^-4 IAA is attractive in a similar microfluidics context (Albrecht and Bargmann), and we have confirmed its attractiveness in our experimental setup (not shown). Specifically, worms accumulate in zones of the arena where 1x10^-4 IAA is present: they are attracted upon first encounter and maintain attraction after prolonged exposure. A sentence to this effect has been added to Results (line 402).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      This looks great! My only small recommendation is to include the 1-octanol concentration used (2.2 mM) in Fig 4 F,G so that the reader can easily compare these data to Fig 4 D,E. You can do this in the figure or legend, though in the figure would be preferable (as you did for Fig 5 and Fig 6).

      [1-oct ] added to Fig. 4D, E

      Reviewer #3 (Recommendations for the authors):

      (1) The AWC traces in Fig. 3C do not match, despite them being from the same animal. There appear to be too many ASH traces. NeuroPAL data/traces from other animals were not available on Zenodo when this review was prepared. This should be fixed.

      Single-animal traces added to figure so reader can see the relationship in the single animal shown and across the full dataset. The averaged ASH/AWC traces are retained in the updated figure because this observation is central to the main hypothesis, so it is important to demonstrate it is a consistent phenomenon. Legend modified (lines 1016-17), ‘consistent’ added to Results (line 147). Data will be uploaded to Zenodo upon finalization of the Version of Record

      (2) The normalization of traces in Fig. 4 is not clear (panel A is not min-max, as panel B appears to be). This normalization may be critical to the interpretation that ASH responds dose-dependently. The code linked on Zenodo was not available when this review was prepared. This should be fixed.

      The reviewer correctly points out a slight error in normalization of the individual worm traces for the ascending concentration series (4A, left panel; the maximal value shown was slightly under 1.0). This is now corrected, interpretation unchanged.

      (3) Chemogenetic data (4F,G and 5D) should specify the promoters used, as they are not ASH and AWC-specific.

      Figures 4F, G and 5D altered to specify the promoters used, as requested. As before, the full range of neurons in which these promoters are active is stated in the Discussion

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the three reviewers for their encouraging and constructive comments. We have addressed them by increasing clarity of the writing and adding details to the methods section that were previously missing or not stated clearly enough. We have included additional experimental data. Specifically, we tested how the osmotic effect of lactulose depends on lactulose dosage and colonization state, added data on cecum sizes of mice with different microbiomes, and tested how lactulose treatment in the active phase affects feeding behavior.

      Public Reviews:

      Reviewer #1 (Public Review):

      Greter et al. provide an interesting and creative use of lactulose as a "microbial metabolism" inducer, combined with tracking of H2 and other fermentation end products. The topic is timely and will likely be of broad interest to researchers studying nutrition, circadian rhythm, and gut microbiota. However, a couple of moderate to major concerns were noted that may impact the interpretation of the current data:

      (1) Much of the data relies on housing gnotobiotic mice in metabolic cages, but I couldn't find any details of methods to assess contamination during multiple days of housing outside of gnotobiotic isolators/cages. Given the complexity of the metabolic cage system used, sterility would likely be incredibly challenging to achieve. More details needed to be included about how potential contamination of the mice was assessed, ideally with 16S rRNA gene sequencing data of the endpoint samples and/or qPCR for total colonization levels relative to the more targeted data shown.

      We thank the reviewer for pointing out that we have not made the experimental setup clear in the text. One of the unique features of our metabolic cage setup is that the mice do not need to be housed outside gnotobiotic isolators, but that the whole system is placed inside an isolator. We have developed and published this system recently (Hoces et al, PLOS Biol 2022), including extensive testing for sterility/gnotobiosis. We have now adapted the main text to increase clarity on this issue (lines 76ff).

      Given that 16S sequencing of germ-free mice will typically produce false-positive reads, we used Blautia pseudococcoides as an indicator strain for contamination. This strain is present in our SPF mouse colony, forms spores that are highly resilient to decontamination measures, and has been the most likely contaminant in our gnotobiotic system. We have checked for presence of this strain in the cecum content of all our animals at the end of each experiment, and only included experiments which had a B. pseudococcoides signal below threshold level. We have now added this information to the methods section (lines 389ff).

      (2) The language could be softened to provide a more nuanced discussion of the results. While lactulose does seem to induce microbial metabolism it also could have direct effects on the host due to its osmotic activity or other off-target effects. Thus, it seems more precise to just refer to lactulose specifically in the figure titles and relevant text.

      We have adapted all figure legends to not contain interpretations, but rather state what was done in the experiments shown in the figures. We have also adapted the text in multiple places to soften the language and avoid overinterpretation of our experimental results.

      Additionally, the degree to which lactulose "disrupts the diurnal rhythm" isn't clear from the data shown, especially given that the markers of circadian rhythm rapidly recover from the perturbation. It is probably more precise to instead state that lactulose transiently induces fermentation during the light phase or something to that effect.

      We tried to make the argument that what we call disruption of the diurnal rhythm is acute, meaning that it is not disrupting the rhythm "chronically" (i.e., for longer), but that it recovers rapidly from this transient disruption. Given the confusion this wording is causing we are introducing this conceptually in the new version of the manuscript (lines 56ff).

      The discussion could also be expanded to address what methods are available or could be developed to build upon the concepts here; for example, the use of genetic inducers of metabolism which may avoid the more complex responses to lactulose.

      We also appreciate the mention of concepts from our study that can be built on in future studies, and we added a paragraph on potential further research. (lines 301ff).

      Despite these concerns, this was still an intriguing and valuable addition to the growing literature on the interface of the microbiome and circadian fields.

      We thank the reviewer for all their encouraging and constructive remarks!

      Reviewer #2 (Public Review):

      Summary:

      The authors aimed to investigate how microbial metabolites, such as hydrogen and short-chain fatty acids (SCFAs), influence feeding behavior and circadian gene expression in mice. Specifically, they sought to understand these effects in different microbial environments, including a reduced community model (EAM), germ-free mice, and SPF mice. The study was designed to explore the broader relationship between the gut microbiome and host circadian rhythms, an area that is not well understood. Through their experiments, the authors hoped to elucidate how microbial metabolism could impact circadian clock genes and feeding patterns, potentially revealing new mechanisms of gut microbiome-host interactions.

      Strengths: 

      The manuscript presents a well-executed investigation into the complex relationship between microbial metabolites and circadian rhythms, with a particular focus on feeding behavior and gene expression in different mouse models. One of the major strengths of the work lies in its innovative use of a reduced community model (EAM) to isolate and examine the effects of specific microbial metabolites, which provides valuable insights into how these metabolites might influence host behavior and circadian regulation. The study also contributes to the broader understanding of the gut microbiome's role in circadian biology, an area that remains poorly understood. The experiments are thoughtfully designed, with a clear rationale that ties together the gut microbiome, metabolic products, and host physiological responses. The authors successfully highlight an intriguing paradox: the significant influence of microbial metabolites in the EAM model versus the lack of effect in germ-free and SPF mice, which adds depth to the ongoing exploration of microbial-host interactions. Despite some methodological concerns, the manuscript offers compelling data and opens up new avenues for research in the field of microbiome and circadian biology.

      We thank the reviewer for their encouraging remarks, specifically on the surprising findings that microbial metabolism seems to affect circadian clock gene expression and behavior differently in EAM and SPF mice.

      Weaknesses:

      The manuscript, while providing valuable insights, has several methodological weaknesses that impact the overall strength of the findings. First, the process for stool collection lacks clarity, raising concerns about potential biases, such as the risk of coprophagia, which could affect the dry-to-wet weight ratio analysis and compromise the validity of these measurements.

      We thank the reviewer for pointing out that our description of the specific methods used for collecting feces were presented in a somewhat confusing manner. In short, dry and wet faecal weights were determined based on fecal pellets that were freshly produced and directly collected from restrained mice. To determine total fecal output over time, we collected all fecal pellets produced in a 5-hour window in a cage, determined their dry weight, and then used the water content determined for fresh faeces to calculate wet weight. Using this method, we cannot account for potential differences in coprophagia between the groups. However, this is not likely to affect the dry-to-wet ratio of faecal output in our results. We have now adapted the section in the methods to increase clarity (lines 440ff), and changed the quantity shown in Figure S2C and E to "water content", which is a more intuitive measure for the same thing.

      Additionally, the use of the term "circadian" in some contexts appears inaccurate, as "diurnal" might be more appropriate, especially given the uncertainty regarding whether the observed microbiome fluctuations are truly circadian.

      Similarly to our answer to reviewer 1 above, we appreciate this remark about imprecise language and have addressed this issue in the text and the figure legends. Indeed, we do not think the fluctuations in microbiota activity are truly circadian, but likely a result of the entrainment through the host's food intake.

      Another significant issue is the unexpected absence of an osmotic effect of lactulose in EAM mice, which contradicts the known properties of lactulose as an osmotic laxative. This finding requires further verification, including the use of a positive control, to ensure it is not artifactual.

      This is a good point. We have used this lactulose dosage specifically to induce microbial metabolism without causing osmotic diarrhoea and went to some lengths do demonstrate this (FigureS2C-E). In response to this comment (and one by reviewer 3 below about transit time), we have now performed additional experiments using higher lactulose dosage (new FigureS3). Our results indicate that the effect of lactulose on water content and transit time depends on microbiota complexity, with no change in fecal water content in 3MM mice even when treated with higher lactulose doses, and a stronger change in SPF mice. We now address this in the main text (lines 127ff).

      The presentation of qRT-PCR data as log2-fold changes, with a mean denominator, could introduce bias by artificially reducing variability, potentially leading to spurious findings or increased risk of Type I error. This approach may explain the unexpected activation of both the positive and negative limbs of the circadian clock.

      While we agree that our description of the qPCR method used for measuring circadian clock gene expression was lacking detail, we do not see how our analysis would lead to an increased risk of Type 1 error.

      Briefly, we use the standard ΔΔCt method to analyze our RT-PCR results. We first normalize gene expression values for each gene of interest to an internal housekeeping gene. Then, we use these normalized values to compare gene expression values in treatment vs control groups (or to the control group at time point 0 in the case of Figure 3C). We then use a log2 transformation to convert the logarithmic RT-PCR values to fold changes. We apologize for the confusing labeling in the figures, where we called the values "log2(fold changes)", which we have now changed to "log2(ΔΔCt) of expression" in Figures 3B, C and S4.

      The simultaneous activation of both limbs of the circadian clock is indeed a surprising result and somewhat complicates interpreting the effect of lactulose treatment on clock gene expression. We take it as evidence that it generally interferes with clock gene expression, while a clearer understanding of the effects would require further research.

      Moreover, the lack of detailed information on the primers and housekeeping genes used in the experiments is concerning, particularly given the importance of using non-circadian housekeeping genes for accurate normalization.

      It seems like the resource table describing these important experimental details was omitted in the original submission. We have now included it in the revised version (Table S1).

      The methods for measuring metabolic hormones, such as GLP-1 and GIP, are also not adequately described. If DPP-IV/protease inhibitor tubes were not used, the data could be unreliable due to the rapid degradation of these hormones by circulating proteases.

      We thank the reviewer for pointing out this omission. We have now added details of how we measured the metabolic hormones to the methods section, including the fact that we have added a DPP-IV inhibitor to the tubes at sampling (lines 469ff).

      Finally, the manuscript does not address the collection of hormone levels during both fasting and fed phases, a critical aspect for interpreting the metabolic impact of microbial metabolites.

      While we agree that it would be interesting to measure hormone levels also in the fed phase, a more thorough examination of hormone levels over the diurnal cycle, as suggested by reviewer 3, would be relevant for a full-scale follow-up. Given our data, we of course cannot exclude that there may be time-point-specific differences and therefore have softened the language around this conclusion to state that hormone levels are not acutely changed after a lactulose intervention at the time-points examined. (lines 318ff).

      These methodological concerns collectively weaken the robustness of the study's results and warrant careful reconsideration and clarification by the authors.

      Because of these weaknesses, the authors have partially achieved their aims by providing novel insights into the relationship between microbial metabolites and host circadian rhythms. The data do suggest that microbial metabolites can significantly influence feeding behavior and circadian gene expression in specific contexts. However, the unexpected absence of an osmotic effect of lactulose, the potential biases introduced by the log2-fold change normalization in qRT-PCR data, and the lack of clarity in critical methodological details weaken the overall conclusions. While the study provides valuable contributions to understanding the gut microbiome's role in circadian biology, the methodological weaknesses prevent a full endorsement of the authors' conclusions. Addressing these issues would be necessary to strengthen the support for their findings and fully achieve the study's aims.

      We thank the reviewer again for their careful and critical reading of our work, and for their constructive input. In the revised version of our manuscript, we address the reviewer's concerns by providing more methodological detail and additional experimental data.

      Despite the methodological concerns raised, this work has the potential to make a significant impact on the field of circadian biology and microbiome research. The study's exploration of the interaction between microbial metabolites and host circadian rhythms in different microbial environments opens new avenues for understanding the complex interplay between the gut microbiome and host physiology. This research contributes to the growing body of evidence that microbial metabolites play a crucial role in regulating host behaviors and physiological processes, including feeding and circadian gene expression.

      We thank the reviewer for their encouraging remarks!

      Reviewer #3 (Public Review):

      Summary:

      In the manuscript by Greter, et al., entitled "Acute targeted induction of gut-microbial metabolism affects host clock genes and nocturnal feeding" the authors are attempting to demonstrate that an acute exposure to a non-nutritive disaccharide (lactulose) promotes microbial metabolism that feeds back onto the host to impact circadian networks. The premise of the study is interesting and the authors have performed several thoughtful experiments to dissect these relationships, providing valuable insights for the field. However, the work presented does not necessarily support some of the conclusions that are drawn. For instance, lactulose is administered during the fasting period to mimic the impact of a feeding bout on the gut microbiota, but it would be important to perform this treatment during the fed state as well to show that the effects on food intake, etc. do not occur.

      We thank the reviewer for this important point. In the revised version, we include an experiment where we administer lactulose during the fed state and do not observe a significant change in food intake. We describe this in the text (lines 189ff) and in the new Figure S5C and D.

      To truly draw the conclusion that the current outcomes are directly connected to and mediated via an impact on the host circadian clock, it would be ideal to perform these studies in a circadian gene knock-out animal (i.e., Cry1 or Cry2 KO mice, or perhaps Bmal-VilCre tissue-specific KO mice). If the effects are lost in these animals, this would more concretely connect the current findings to the circadian clock gene network.

      We agree that these would be interesting experiments to follow up on the question how the observed effects are actuated by host functions. However, they would require a large amount of preparatory work (including rederiving the KO mice to get them germ-free in our gnotobiotic facility), we argue that they are beyond the scope of this study.

      Despite these reservations, the work is promising.

      We thank the reviewer for their encouraging assessment.

      Strengths:

      Attempting to disentangle nutrient acquisition from microbial fermentation and its impact on diurnal dynamics of gut microbes on host circadian rhythms is an important step for providing insights into these host-microbe interactions.

      The authors utilize a novel approach in leveraging lactulose coupled with germ-free animals and metabolic cages fitted with detectors that can measure microbial byproducts of fermentation, particularly hydrogen, in real-time.

      The authors consider several interesting aspects of lactulose delivery, including how it shifts osmotic balance as well as provides calculations that attempt to explain the caloric contribution of fermentation to the animal in the context of reduced food intake. This provides interesting fundamental insights into the role of microbial outputs on host metabolism.

      Thank you!

      Weaknesses:

      While the authors have done a large amount of work to examine the osmotic vs. metabolic influence of lactulose delivery, the authors have not accounted for the enlarged cecum and increased cecal surface area in germ-free mice. The authors could consider an additional control of cecectomy in germ-free mice.

      We thank the reviewer for pointing out the potential effect of the anatomical differences of germ-free and conventionally colonized mice. We agree that when comparing germ-free mice to SPF mice, the enlarged cecum area in germ-free animals could lead to differences in water release or uptake. However, this difference is smaller in gnotobiotic mice colonized with our minimal microbiota, even though their ceca are still slightly smaller than those of germ-free mice (new Figure S2F). While we agree that including control of cecectomy in germ-free mice could be interesting, we do not have the option of doing surgery on germ-free mice given our current experimental setup. We have now added information on cecum weight, a good proxy for cecum size, in the new Figure S2F done.

      The authors have examined GI hormones as one possible mechanism for how food intake is altered by microbial fermentation of lactulose. However, the authors measure PYY and GLP-1 only at a single time point, stating that there are no differences between groups. Given the goal of the studies is to tie these findings back into circadian rhythms, it would be important to show if the diurnal patterns of these GI hormones are altered.

      We fully agree that a deeper investigation of the diurnal fluctuations of hormone levels would be an interesting next step in studying whether perturbations in food intake can disturb these rhythms. Doing this for the whole rhythm would really require a full second study.

      In response to the reviewer's comments, we have changed the statements made around these data to point out just that hormone level fluctuations could not be detected during specific time points after lactulose treatment and therefore do not seem to explain the imminent behavioral changes (lines 318ff).

      Considerations of other factors, such as conjugated vs. deconjugated bile acids, microbial bile salt hydrolase activity, and bile acid resorption, might be an important consideration for how lactulose elicits more influence on ileal circadian clock genes relative to cecum and colon.

      We absolutely agree that investigation of microbial bile acid modification and their metabolism by the host would be an interesting topic for a follow-up study.

      Measurements of GI transit time (both whole gut and regional) would be an important for consideration for how lactulose might be impacting the ileum vs. cecum vs. colon.

      This is also an interesting point. While we did not add an experiment in which we specifically measure transit time to the revised version, we now measure total faecal output in a 5 h time period after PBS or lactulose treatment (Figure S3C). Faecal output is known to be a good proxy for transit time, and we see no significant difference between lactulose treatment (even with a two-fold higher dose than used before, new Figure S3) and PBS treatment.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      (1) Line 126 - see the point in the public review, this data argues against disrupting the rhythm.

      See our response above

      (2) Line 156 - the metric used for water content is confusing. Why not just subtract dry weight from wet weight to get water content? The ratio is much harder to think about for me. Perhaps more importantly, this data is very confusing given that colonization seems to impact the activity of lactulose, which complicates the interpretation. Could be an interesting area for future study that you might highlight more in the discussion.

      We thank the reviewer for pointing out that our presentation of water content could be clearer. We have changed the dry/wet ratio we have used in the previous version to the "water fraction" (new Figure S2C, E; new Figure S3A,B), i.e., the per cent of weight of the wet sample that is made up by water. We would argue that this is a measurement that is easier to interpret than the difference suggested by the reviewer, because it is independent of the absolute sample weight.

      We also agree that it is surprising that colonization state changes the osmotic state of the gut environment and have now added text discussing that (lines 127ff).

      (3) Line 188 - the lack of expression changes in the distal gut (cecum/colon) potentially conflicts with the model, warranting additional discussion/qualifications. Is lactulose getting metabolized in the small intestine? Alternatively, does lactulose have a direct effect on the host? The current text seems to imply that lactulose is fermented in the colon, the fermentation products are absorbed, and then they only impact the ileum through circulation, which doesn't seem physiologically possible.

      It is possible that lactulose affects small intestinal tissue directly, but we show that the effect depends on microbial activity. Microbial activity is much larger in the large intestine than in the small intestine, which is why we hypothesize that it acts through systemic signals rather than locally. These systemic signals might be fermentation products impacting the ileum through circulation. We would argue that this is not implausible, given that there are well-documented systemic effects of fermentation products in circulation (e.g., den Besten et al, 2013). It is, however, also possible that, e.g., metabolism of fermentation products in the liver triggers a secondary signal that acts systemically.

      (4) Line 241 - need to weaken this sub-heading. The experimental design shows that fermentation products impact feeding behavior, but this does not necessarily imply that fermentation products are responsible for the lactulose effect.

      Done.

      (5) Line 256 - The lack of an effect in SPF mice is surprising and potentially conflicts with the model proposed. Given this and other confusing results (gene expression site specificity, osmolarity effects, etc) it seems prudent to be more cautious as to the potential mechanisms through which lactulose supplementation impacts host gene expression and feeding behavior, which would likely require a lot more experiments to provide a definitive answer.

      While the lack of an effect in SPF mice is surprising, we would argue that this is rather points towards the need for a better understanding of host-microbiota-diet interactions than a conflict with our interpretation. We do agree with the reviewer that more work is necessary for providing definitive proof of the underlying mechanisms of the observed effects and have adapted the language throughout the text.

      (6) Line 260 - A lot of text is devoted to the potential caloric effects of the fermentation products and the lactulose itself. Was this in response to a prior reviewer? Either way, it seems too speculative to me and I would recommend trimming it down and moving it to the discussion.

      We have rewritten this section to make it more concise and increase readability (lines 224ff).

      (7) Line 305 - Not fasting, just lower caloric intake, as shown by Figure 1c.

      We have changed the text to reflect that.

      (8) Line 371 - Too definitive given the current data. Need to qualify the interpretation here.

      We have adapted the text and qualified the interpretation.

      (9) Figure 2 - Need to revise the title, no data showing that the rhythm is disrupted.

      Done.

      (10) Figure 3b - Clarify what timepoint is shown in the legend.

      Done.

      (11) Figure 3c - label when the treatment started. Consider changing to 2-way ANOVA which is probably more appropriate than t-tests.

      We have added information on treatment start to the figure legend and have changed the statistical analysis to a two-way ANOVA (time, treatment).

      (12) Figure 4c - move to supplement as this negative data isn't sufficient to rule out an impact on the hypothalamus or liver. Even when only considering transcript levels it's possible that the timepoint is just not ideal.

      We agree that this data only represents a snapshot of gene expression at one timepoint after treatment and does not rule out involvement of hypothalamus or liver in this process. We have now moved the previous Figure 4C to the supplementary information (new Figure S6A).

      (13) Figure 5 - defined the "fermentation products" and their concentrations in the legend. Consider moving panels d, and e to the supplement - negative results with unclear interpretation. Clarify in the legend how many calories/g were assumed for the fermentation products and provide a scientific rationale for this decision. Modify the title to be more cautious - as discussed above.

      We thank the reviewer for pointing out that this was not clear. We have now added the formulation of the fermentation products to the figure legend and moved panels D and E to the supplement. We have also combined the former Figures 4AB and 5ABC into the new Figure 4.

      (14) Figure S1 - The patterns in panel a are really intriguing and could be discussed more. For example, what do you think is driving the rapid oscillations in E. rectale?

      We agree that the patterns are potentially interesting, but the fact that the patterns we observe do not replicate well between the two light-dark cycles we monitor suggest a large contribution of experimental noise. We therefore refrain from interpreting too much into this dataset.

      (15) Figure S2 - Could changing the metric for water content be more easily interpretable? Modify the title to better match the data shown.

      We thank the reviewer for pointing that out. We have now changed the metric in Figure S2 (and the new Figure S3) to % water in feces/cecum content, which is easier to interpret.

      (16) Figure S3 - need to specify the multiple testing correction used.

      Done.

      (17) Figure S4 title - replace "inducing microbial metabolism" with "lactulose".

      Done.

      (18) Figure S5 title - modify to weaken the causal claim.

      We have adapted all figure titles to conform to this comment.

      Reviewer #2 (Recommendations For The Authors):

      Greter and colleagues present an insightful manuscript investigating the effects of microbial metabolites, such as hydrogen and SCFAs, on feeding behavior and circadian gene expression at a single time point. Notably, they observed that these metabolites exert a significant influence on behavior in a reduced community model (EAM). However, this effect was not evident in germ-free or SPF mice, highlighting an intriguing paradox. The manuscript is well-written, with experiments that are thoughtfully designed and clearly rationalized. The study contributes valuable data to the poorly understood relationship between the gut microbiome and host circadian rhythms. While I find the manuscript compelling, I believe that providing additional experimental and analytical details would enhance clarity and rigor.

      Major Comments:

      (1) Additional clarity is needed regarding the stool collection process. Were the samples collected as fresh specimens, or was there a possibility of coprophagia before collection? Clarification on this point is important, as it could impact the results, particularly the dry-to-wet weight ratio analysis. Ensuring the collection process did not introduce this bias is crucial for the validity of these measurements.

      Thank you for pointing out that this was not stated clearly. Wherever we assessed water content in faeces, we used fresh samples that were directly collected from a live animal and immediately frozen in a closed container or analyzed. When we assessed total faecal output, we collected the total bedding from a cage, sorted out the faecal pellets, and only measured dry weight. Total wet weight output was then assessed by using the water content of fresh faeces of animals with the same microbiota and the same treatment as a correction factor. We have now made this clear in the methods section (lines 435ff).

      In cases where we performed total output measurements by collecting bedding, there was the possibility for coprophagia. We did not control for this but assumed that coprophagia will have a small effect that is likely similar between groups and should thus not affect the comparison.

      (2) Caution is advised in the use of the term 'circadian,' which is sometimes used when 'diurnal' might be more appropriate. For example, the title of the first results section could be revised to 'Host Feeding [or Diurnal] Rhythms Influence Microbial Metabolic Fluctuations.' Additionally, line 357 should likely use 'diurnal' instead of 'circadian.' It's important to note that 'circadian' implies that cyclical fluctuations would persist without environmental cues (e.g., feeding). Since it's not clear whether most microbiome compositional or functional fluctuations are truly circadian, 'diurnal' is likely the more accurate term.

      We thank the reviewer for pointing out our imprecise use of the term circadian. We have now adapted this throughout the manuscript.

      (3) The lack of an osmotic effect of lactulose in EAM mice is quite surprising, given that lactulose is known to be an osmotic laxative. Was this finding specific to EAM mice, or was a similar lack of osmotic effect observed in SPF mice? A positive control is necessary to verify that this unexpected result is not artifactual. If there is no osmotic laxative effect in SPF mice, an explanation is needed as to why this medication is not functioning as expected in these mice.

      This is a good point. While we initially thought that the lack of an osmotic effect was due to the specific lactulose dose we were using, we also did not observe an osmotic effect (measured by the water content of fresh faecal pellets produced after treatment) when we used twice the amount of lactulose in EAM mice (new Figure S3). However, in SPF mice, we did see an increase in faecal water content after treatment. This intriguing result suggests that the effect of lactulose as an osmotic laxative depends on the presence of a complex microbiota. We now show this data in the new Figure S3A and discuss it in the text (lines 127ff).

      (4) The authors state that 'To account for faulty measurements due to disruptive events and for measurement noise, some datapoints were excluded from the raw datasets.' It would be important for the authors to confirm that this data exclusion was unbiased, meaning it did not disproportionately affect one group over another, and that any exclusions affected groups randomly.

      This is a good point, and we analyzed this for the experiments we show in Figs 4A/S5A, 4B/S5E, 4C, and S5F, and discuss this in the methods part (lines 493ff). The resulting statistics is not fully conclusive: using a Chi-square test to check whether the probability of excluding values differs between experiments, we get significant differences (p=3.9 x 10<sup>-18</sup>). We would, however, argue that this is not surprising, as different experiments sometimes different in the number of times we needed to do maintenance work on the isolators, which could lead to actuation of the scales measuring feed values, and thus faulty measurements.

      We face the same problem within experiments: we found a significant difference in the probability to exclude values between treatment and control in the experiment shown in Figs 4A/S5A (p=1.7 x 10<sup>-6</sup>), but no significant differences in the experiments in Figs 4B/S5E and S5F (p=0.71, and p=0.10, respectively). While it is hard to strictly exclude an influence of our data exclusion strategy, these findings speak against a systematic effect of treatment.

      (5) The authors should provide details on the primers and housekeeping genes used in their experiments. It's crucial that the housekeeping gene is not circadian and is stable at all time points (PMID 17878933).

      In the previous submission, the main resources table was omitted by accident. We have now added it (Table S1), including details on the primers and reagents used for all experimental work.

      (6) The presentation of qRT-PCR data as log2-fold change is confusing. It's unclear what the numerator and denominator represent for this ratio or why such normalization was deemed necessary. Ideally, transcripts should be normalized to a housekeeping gene (as noted in a previous comment), not to a baseline measure of other same genes acquired from other mice. Log2-fold change is typically appropriate when comparing two measures from the same mice; however, in this study, the mice were euthanized, and the denominator is a mean of genes measured from other samples. This approach could introduce bias and might explain why both the positive and negative limbs appear to be activated by the microbial metabolites. It would be more rigorous to present these values as absolute gene expression levels.

      We thank the reviewer for pointing out that this was not described clearly in the previous version of our manuscript. We have now adapted the methods part to explain that all data showing RT-PCR data is normalized to a housekeeping gene. Only after that, we compare the gene expression levels of the treatment group to the control group (or to gene expression of the control group at timepoint 0 in the case of Figure 3C).

      (7) The use of log-ratios, with a mean as the denominator, could artificially reduce the variability in the data, potentially leading to spurious findings or an increased risk of Type I error.

      As pointed out above, we have used internal normalization to a housekeeping gene before comparing the resulting values of the treatment and control groups. This method (commonly known as ΔΔC<sub>t</sub> method) is a standard way of comparing gene expression values obtained by RT-PCR. The use of a log2 transformation is commonly used to convert the logarithmic data resulting from the RT-PCR measurement to a linear fold-change measurement. We do not see that this data analysis strategy should lead to an increased risk for producing false positives. We now explain this better in the methods section of the manuscript (lines 541ff), and have adapted the labels in Figures 3B, C and S4 to avoid confusion.

      (8) It is unusual that both the positive and negative limbs of the circadian clock are overexpressed following lactulose administration. The authors should provide data confirming that these genes are in counter phase to each other at baseline. This clarification would help readers better understand the effects of the experimental interventions. As it stands, this critical part of the results is quite confusing.

      We agree that this result does not allow a clear interpretation of the effect of lactulose on the diurnal rhythm of the host. While we agree that this would be interesting to understand in detail, we are merely taking this as a first indication that actuation of microbial metabolism during the inactive phase of the diurnal rhythm can lead to changes in clock genes. This claim is supported by our data.

      (9) Could the effects of the microbial metabolites be mediated by AMPK, a known nutrient sensor that can influence the post-translational modification of Cry proteins? It would be beneficial for the authors to explore whether these metabolites have a more direct, previously unknown mechanism of affecting the circadian clock, or if their effects are mediated through known signaling pathways such as AMPK.

      We agree that this would be a valuable path to continue investigating the effect of microbial metabolism on clock gene activity.

      (10) The methods section does not specify how the metabolic hormones (e.g., GLP-1, GIP, leptin, ghrelin) were measured in the experiments. It is important for the authors to confirm that DPP-IV/protease inhibitor tubes were used for hormone measurement, as these proteins can be rapidly degraded by circulating proteases. Without the use of appropriate tubes, this data cannot be reliably interpreted. Additionally, it would have been ideal to collect these hormone levels during both the fasting and fed phases, but it appears this was not done. This represents a significant limitation of the study and should be addressed in the discussion.

      We thank the reviewer for pointing out this omission, we have now added a description of our protocol to measure metabolic hormones to the methods section (lines 469ff).

      (11) If the samples were appropriately collected in DPP-IV/protease inhibitor tubes, the authors should consider measuring active GLP-1, as this would likely provide a more accurate assessment of GLP-1 activity.

      We agree that this would be a valuable next step, in addition to testing the effect of changes in microbial metabolism on the time traces of hormone levels.

      Minor Comments:

      (12) Line 350 appears to have an incomplete sentence, as it seems part of the first sentence in the paragraph has been inadvertently deleted. This should be reviewed and corrected for clarity.

      Done.

      Reviewer #3 (Recommendations For The Authors):

      Major comments:

      (1) Could the authors provide a deeper description about what they are referring to in the following statement? "...higher order interactions and microbial metabolism are variable..." it is difficult to interpret as written. Do the authors mean cross-feeding interactions?

      We have changed this sentence to clarify the meaning.

      (2) Could the authors explicitly state their hypothesis in the introduction and provide a brief, but deeper explanation of the intervention prior to the results section?

      This is a good point, we have adapted the text accordingly (lines 53ff).

      (3) Could the authors include a bit more information regarding the diet provided to the mice? If grain-based chow, please provide insights into the fiber source, etc.

      While we agree that it would be interesting to know what part of the mouse diet is available to the microbes, this is hard for the standard mouse chow that we (and most others doing experiments with mice) feed the experimental animals. We have now added more detailed information on the specific type of chow the mice were fed (lines 347f). We would argue that, because the control and treatment groups were always fed the same chow, the effect of the fiber source and other specifics are controlled for, even if we do not know them.

      (4) In figure 1A - cells/g does not seem to be the correct unit - # of copies/g feces perhaps?

      Cells/g is the appropriate unit, but it seems like we have not explained the way we arrive at this unit in sufficient detail. In short, we use a qPCR run on a known standard curve of bacterial counts (known cells/g values) to estimate these numbers from qPCR results from faeces. We have now explained this better in the methods section (lines 385f).

      (5) Figure 1B/C and Figure S1B/C are confusing - the legend states these measurements were taken over two days, however, the plot shows a single 12:12 LD period. Was the data averaged? It might be best to show each day separately (i.e., over a 48-hour period) rather than in one 24-hour plot. Then, the authors could also show the averages in the light period vs. the dark period in a separate, complementary graph.

      We thank the reviewer for pointing this out. The previous figures were indeed averaged over the two days of measurement and projected onto one 24h period for plotting. We have now changed Figures 1B,C and S1B,C to show the full 48h time windows.

      (6) Line 109 - The reviewer concurs that lactulose is a non-nutritive, synthetic disaccharide, however, in theory, lactulose may have a high heat increment, which could cause the animal to undergo metabolic responses to defend core body temperature (which also exhibits diurnal rhythmicity). Have the authors considered core body temperature rhythms, their connection to microbial metabolism, and the core circadian clock gene network in their model?

      This is an interesting thought. We have not measured body temperature in our experiments. As the heat increment from food is typically associated with metabolic activity, and lactulose is not metabolized, we do not expect its heat increment to be high, at least in GF mice. In mice with a microbiota, we agree that metabolic heat will be produced upon lactulose metabolism by the microbes, which could be a contributor the observed effect.

      (7) Figure 2A and corresponding text in lines 121 - 123 - indicate at what time the bar and whisker plots were taken (assuming at ZT8, but please be explicit in the figure and corresponding text). Could the authors also include statistics for these waveforms?

      Done.

      (8) In Figure 2B, the authors state that microbial metabolism had been restored to normal levels by ZT15, however, did this persist into the next light phase? It would be ideal if the authors could present these data in a similar manner to that shown in Figure S2A for SPF mice. Further, what are the statistical considerations here to describe changes in phase, amplitude, periodicity, etc.

      This is a good point. Unfortunately, we do not have H2 measurements for EAM and GF mice over comparable time periods as shown for SPF mice in Figure S2A. However, the food intake data shown in Figure S5BCD is a indicates that the food intake normalized in the second dark phase after treatment.

      (9) In lines 147 - 149 and in Figure S2B figure legend - assuming these measurements are from individual bacteria? Could this be stated clearly in the text or legend? Also, what are the statistical considerations? Were there significant differences in SCFA production between bacteria?

      We thank the reviewer for pointing out that this was unclear. We have now adapted the legend to explain the way these data were collected and added a statistical analysis.

      (10) The authors have done a large amount of work to examine the osmotic vs. metabolic influence of lactulose delivery - however, have the authors accounted for the enlarged cecum and increased cecal surface area in germ-free mice? Would an additional control be cecectomy in germ-free mice to be more in-line w/ SPF animals? Further, could the authors tie in these findings more explicitly and state how they pertain to the overall goal of the study? Is this simply to draw the conclusion that microbial biomass is increased w/ lactulose?

      We wanted to make sure that the effect we see with lactulose is due to microbial metabolism and not due to the induction of osmotic diarrhoea or other host-dependent effects, and we have added a statement to that effect (lines 127ff). We have not corrected for the change in cecum size between GF and EAM mice, as EAM size (and many other gnotobiotic mouse models) also have enlarged ceca relative to conventionally colonized mice, but have added a dataset showing how cecum size of EAM and GF mice compare and discuss this in the text (new Figure S2F, lines 135ff).

      (11) Is Figure 3A necessary?

      It might not be strictly necessary, but it can help with understanding the relation of the genes tested in B and C, and we would therefore like to keep it in.

      (12) Line 155 - 157 - the authors make the statement that dry/wet feces weight ratio is decreased in GF mice, but this does not appear to be statistically significant. Please adjust to state numerical differences were observed or provide statistics.

      We have changed the statement in the text.

      (13) Could Figure S3A be moved to the main Figure 3 as this may provide a more logical flow? qRT-PCR data is expressed as - log2 (fold expression), but relative to what? Could the authors provide further info about the control?

      The former Figure S3A (new S4A) and Figure 3B show the same data in slightly different ways. We therefore opted to keep only one of those illustrations in a main figure. The fold changes are always relative to the average of the PBS control group, which we now state explicitly in the figure legend.

      (14) The authors state that the qRT-PCR data shows that microbial metabolism of lactulose impacts peripheral circadian gene expression, but this conclusion seems simplified. Lactulose treatment only impacted the ileum circadian gene expression. Additional peripheral tissues (liver, adipose tissue, etc.) could be moved to this figure, i.e., move Figure 4C data. Why do the authors think the ileum was most impacted beyond GI hormones as discussed later in the manuscript? Could changes in bile acid deconjugation (i.e., BSH activity?) and/or bile acid resorption by the host in the distal ileum due to lactulose delivery be involved? Or is it simply due to differences in GI transit time (which was not measured in the current study)? Further, lactulose had minimal impact in SPF ileum, and in fact, shifted Cry1 in the opposite direction relative to EAM mice. Could the authors provide more insight into these disparate observations (line 196 - 200)?

      We agree that the statement "lactulose impacts peripheral gene expression" is oversimplified, and we have now adapted the text to avoid the impression that this is our conclusion. No tested tissues other than the ileum showed significant differences in gene expression at the time point tested, which does not rule out that other tissues would react to the treatment at that time point, or the tested tissues would do so at the tested time point. As we don't have a good enough understanding of what mechanism causes the gene expression changes in the ileum, we refrain from speculating in the text, even though we agree that this is an intriguing question.

      (15) Could the authors provide more insight into the statistical approaches used to assess amplitude, peak, nadir, etc. in Figure 3C? Was the co-sinor waveform tested?

      This is a good question. Even though this was a highly work-intensive experiment using many animals, we would argue that the noise level is too high and the coverage of the time analyzed too sparse to infer meaningful statistics on the fluctuations of gene expression over time. In the new version, we have changed our statistical analysis of this dataset to a two-way ANOVA (treatment, time) to better analyze this dataset.

      (16) The food intake decrease and interpretation following treatment (Figure 4A and S4A) is curious - all animals were gavaged and in EAM mice, many animals, regardless of PBS or lactulose are trending down in food intake rate/total intake. It seems to be more of an impact of gavage and not of treatment, which the authors somewhat acknowledge in Lines 228 - 230.

      We agree that gavage is a possible factor in future food intake of experimental animals, which is precisely why we used the PBS gavage as a control. Even though the difference is not large, we see a significant change in food intake when lactulose is given, but not when PBS is given (Figure 4A, S4A).

      (17) Could the authors provide a deeper rationale for their line of thinking for lines 234 - 240? What is the evidence that systemic effects are likely to occur 3 hours after lactulose delivery? Further, as stated in comment 13, could brain and liver data be moved to Figure 3/Figure S3 as an additional example of peripheral tissue clocks?

      We have added an explanation for the rationale we use to justify the 8h time point (it is 3h after the peak of H2 production, as shown in Fig2AB, which happens 5h after lactulose delivery). While we agree that the brain and liver data would also fit into the Figure 3, we have now moved all negative data to the supplementary information (in response to a comment by reviewer 1, and in a general effort to clean up the data in the manuscript).

      (18) The authors measure PYY and GLP-1 at a single time point and state there are no differences, yet, the goal of the studies is to tie this back to circadian networks. Would it be possible to measure these GI hormones over a 24-hour period to show that the diurnal patterns are altered?

      We fully agree that measuring the metabolic hormones over time would be very interesting. It is possible but would represent a major effort using many animals and a large amount of work. We would therefore argue that it is beyond the scope of this revision, but a good starting point for a follow-up study.

      (19) The authors state that the administration of fermentation products acutely altered circadian food intake, but the studies do not support that this change is connected to the circadian network. Suggest softening the interpretation of the findings.

      We have changed the language there to soften the interpretation.

      Minor comments:

      (1) The authors should consider when it is appropriate to refer to rhythms as diurnal vs. circadian, as each has a distinct meaning. Diurnal follows entrainment cues while circadian is endogenously driven (i.e., line 39, line 58).

      We thank the reviewer for pointing out this important difference, we have adapted this in the whole text accordingly.

      (2) Circadian rhythm should be plural throughout the manuscript (circadian rhythms).

      Thank you, we have changed that where we refer to host circadian rhythms generally.

      (3) Lines 54 - 63. Fermentation should be capitalized when used at the beginning of a sentence.

      Done.

      (4) Line 289 - This should be Figure 5D and 5E.

      Done.

      (5) Line 290 - heart should be cardiac.

      Done.

    1. Author response:

      eLife Assessment

      This paper introduces a valuable optical method for simultaneous in vivo multiphoton imaging of the mouse brain combined with DMD-based one-photon patterned photostimulation in different axial planes. The evidence for effective optical separation of excitation and imaging is convincing, although the in vivo data in the olfactory bulb suggest potential confounding factors arising if stimulation not only affects cell bodies but also neuronal processes. The work will be of broad interest to neurobiologists working in circuit and systems neuroscience, as well as to specialists in optical microscopy.

      We thank the reviewers for their constructive feedback. We are planning to address their concerns as detailed below. Specifically, in the revised manuscript, we will streamline and consolidate the text to include:

      (1) An in-depth discussion of the awake recording results in the main text.

      (2) Further discussion of light scattering, photo-stimulation specificity of targeting individual glomeruli and resolution.

      (3) A summary (including also a table) of the operating regime and comparisons with alternative techniques for patterned photo-stimulation and imaging of the ensuing responses. We will highlight the advantages and constraints of the current implementation of ADePT. In particular, here we explored a small set of spatiotemporal parameters to understand the limits of our technique and provide a proof of principle of the strategy. These parameters can be varied further depending on the exact research question.

      (4) Implementation considerations and technical guidelines for calibration and long-term stability of the rig for ADePT (i.e. ‘a how-to guide’).

      (5) A bill of materials and estimated hardware costs.

      Furthermore, we will provide additional controls and rephrase some of the statements in the text as suggested (e.g. replace ‘accessible’ with ‘simple’, etc.). We will correct the unfortunate grammatical errors, improve clarity of text, and update the references accordingly.

      Public Reviews:

      Reviewer #1 (Public review):

      In this methods paper, the authors introduce a novel and innovative imaging approach for simultaneous in vivo multiphoton imaging of the mouse brain combined with DMD-based one-photon patterned photo-stimulation in different axial planes. This is a highly exciting technique that enables the axial decoupling of optical imaging of deep neural circuits from surface photo-stimulation of spatially precise (tens of micrometres) brain spots. This method builds on previous developments from the same laboratory, combining DMD-based patterned photo-stimulation with in vivo electrophysiological recordings. To my knowledge, this is the first instance in which patterned photo-stimulation has been combined and axially decoupled from two-photon (2P) imaging.

      Beginning with a thorough characterisation of the optical resolution of the photo-stimulation system, the authors applied this method to the olfactory bulb (OB) network, in which sensory inputs are topographically organised at the surface of the OB and thus ideally suited to demonstrate the relevance of this approach. They first showed that this technique can be used to rapidly reveal connectivity patterns of OB output neurons and to identify sister mitral cells. In addition, they manipulated a specific glomerular inhibitory population and demonstrated that these neurons provide spatially heterogeneous long-range inhibition of OB output neurons, with differential effects on mitral and tufted cells (a result previously observed in a paper from the same lab: Banerjee et al., 2015, Neuron). Altogether, the data demonstrate that this technique is well-suited for high-throughput functional mapping of neural circuit properties. The results are compelling and illustrate both the significant advance represented by this method and its feasibility.

      We thank the Reviewer for their constructive input.

      Despite my initial enthusiasm, there are several concerns in the present study that must be addressed in order to rule out confounding observations and to resolve remaining uncertainties regarding photo-stimulation resolution. These include the following:

      (1) Spatial resolution: Although the authors provide convincing data on spatial resolution in vitro, several observations throughout the paper suggest that the effective photo-stimulation precision may be lower than initially reported. For instance, in Figure 2, the authors observe repeated responses in neighbouring glomeruli (e.g., glomeruli #3 & #5, #4 & #6). To what extent could light scattering along the X/Y/Z-axis above the targeted glomerulus recruit en passage axons, resulting in the inadvertent activation of multiple glomeruli?

      Indeed, we cannot rule out this possibility. Fibers of passage are a potential concern. This is why we systematically sample different light intensities and assess their impact on specificity of dendritic mitral and tufted cell responses within the glomerular layer (same axial-plane optical stimulation and imaging experiments, Fig. 2). We identify a range of intensities that on average result mostly in activation of the targeted glomeruli. Within the range of intensities used for identifying sister cells, >90% of responses were on the diagonal (targeted glomeruli) and ~5% pixels that cleared the signal significance criterion used were in off-target glomeruli, as stated in the text and quantified in Fig. 2f. In the revised manuscript, we will further clarify and expand on these points.

      A further observation concerns the presence of "inhibited" sister mitral cells (Figure 3). The authors claim this is reminiscent of the differential spike-timing reported between sister cells (Dwawale et al., 2010, Nat Neuro). However, observing both excitatory and inhibitory responses following stimulation of glutamatergic inputs is an altogether different matter, particularly given that sister mitral cells are reciprocally connected via gap junctions. This observation requires further clarification and raises serious questions about the effective resolution of the stimulation. Could the inhibited cell simply correspond to a non-sister mitral cell receiving disynaptic feed-forward inhibition?

      This is indeed what we think it is happening (i.e. disynaptic feed-forward inhibition as the Reviewer points out). We observe inhibition in some of the mitral cells in the field of imaging when we stimulate not their parent glomerulus, but other glomeruli in the neighbourhood. As the Reviewer points out, sister cells are connected via gap junctions, but they also receive inhibitory chemical synaptic inputs via their secondary (and primary dendrites) from other (not-their-parent) glomeruli mediated by numerous types of interneurons including the DAT+/GABAergic (a.k.a. superficial short axon cells) and granule cells. Our data is consistent with differential inhibitory input from other glomeruli on sister cells getting input from the same parent glomerulus. As it appears that we failed to present this point clearly in the initial submission, we will further expand along these lines in the revised manuscript. Briefly:

      First, we identify a photo-stimulation regime that results mostly in the activation of a given targeted glomerulus (and not of other glomeruli in the field of stimulation). To this end, we photo-stimulate and image ensuing neuronal responses in the same axial optical plane. We strobe (alternate) between monitoring dendritic mitral and tufted cell (enhanced) GCaMP responses within the targeted glomerulus and other glomeruli in the field of imaging, while varying systematically the light intensity (Figs. 2,3a; Suppl. Fig. 4b, Suppl. Fig, 5b-d;h-j). We use as criterion for specificity a condition when >95% of significantly responding pixels (above a statistically defined signal response threshold) lie within the anatomical boundaries of the targeted glomerulus. For each glomerulus (or pixel within a glomerulus) we compared the average light response across trials with the baseline reference distribution in the absence of light stimulation. If this value crossed the 99th percentile of the baseline distribution, the glomerulus/pixel within glomerulus was classified as responsive to the photo-stimulation.

      Second, using the minimal light intensity regime experimentally identified as ‘specific’ for targeting individual glomeruli in the field of photo-stimulation (< 5% significant activation of off-target pixels), we decouple photo-stimulation in the glomerular layer from monitoring responses of mitral and tufted cells in the deeper layers of the olfactory bulb (100-250 µm axial displacement). This approach enables us to map cohorts of sister (daughter) cells associated with any specific target glomerulus in the field of view (1,2,3…n) by monitoring excitatory (enhanced) responses of mitral and tufted cell bodies. A cohort of sister cells associated with glomerulus x<sub>i</sub> (daughters of glomerulus x<sub>i</sub>) is defined by those cells which show statistically significant excitatory responses (4 SD - standard deviations - above their baseline fluctuations) specifically in response to photo-stimulation of glomerulus x<sub>i</sub>.

      Third, in the process, as we photo-stimulate different glomeruli in the field of stimulation, we also observe at times suppressed (inhibitory) responses in a subset of the mitral and tufted cells (exceeding 3 SD in the negative direction their baseline fluctuations). These suppressed responses occur in response to photo-stimulating not the parent glomerulus of a given cell, but other glomeruli in the field. These experiments revealed that within a cohort of sister cells (daughters of glomerulus x<sub>i</sub>), only a subset of cells are suppressed by activation of glomerulus x<sub>j</sub>, and, in a few example cases, different cells are suppressed by activation of different glomeruli (e.g. x<sub>j</sub> vs. x<sub>k</sub>), presumably through disynaptic feed-forward inhibition (Fig. 3b iii; 3d; Suppl. Figs. 5f,g; l,m). These preliminary observations suggest that sister cells receive differential inhibitory inputs from glomeruli in the neighborhood. In the revised manuscript, we will expand to further clarify these points.

      To verify sister cell identity, the authors could confirm that the predicted sister cells share a similar odour receptive field compared to randomly selected mitral cell pairs. In their previous study employing analogous DMD-based photo-stimulation (Dhawale et al., 2010, Nat. Neurosci.), sister mitral cells did not exhibit such opposite response profiles (firing rate correlation of ∼0.7 between sister cells). Could the authors verify that a comparable activity correlation is also observed among the sister cells identified using ADePT in the present study? In Figure S6, the authors show recordings and stimulation of the same neurons co-expressing GCaMP and ChR2. Applying this experimental design to the mitral/tufted cell population (using a Tbet-Cre mouse transduced in the OB with both GCaMP and Chrimson virus) would constitute a valuable control to clarify the nature of these "inhibited" sister cells.

      We thank the Reviewer for the suggestion. We consider that the experiments shown here are proof-of-principle in nature, highlighting the potential of ADePT for mapping functional neural circuit connectivity. In our opinion, further investigating the logic of similarities and differences in the odor responses of sister mitral cells and the nature of inhibitory glomerular interactions forms the focus of future studies. We also note that firing rate correlations can be notoriously difficult to compare and interpret across experimental regimes (spikes vs. calcium imaging).

      An additional concern relates to the 21 out of 162 mitral cells that were activated by two distinct glomeruli - a finding that is incompatible with the established OB wiring diagram and that further challenges the claimed stimulation resolution.

      Indeed, this reflects some degree of non-specific activation of the targeted glomeruli as discussed above. In the revised manuscript, we will further highlight this issue.

      A critical control experiment is also absent: in a Thy1-GCaMP6 mouse lacking any light-sensitive opsin, do the authors observe any unintended side effects of photo-stimulation?

      In the revised manuscript, we will include an additional control as suggested by the reviewer. Within the range of intensities used, we did not observe significant modulation of GCaMP6s activity in mitral and tufted cells in mice lacking light-sensitive opsins.

      Regarding sister cells (Figure 3), tufted cells are not analysed alongside mitral cells in this dataset, whereas this is elegantly performed in Figure 5 using the DAT+ model. Could the authors also demonstrate how the technique can reveal the complete family portrait of sister mitral and tufted cells?

      We thank the Reviewer for the suggestion. We think that the differences between mitral and tufted cells are indeed very interesting to investigate, but in our opinion form the subject of future studies.

      (2) The authors have explored only a limited set of photo-stimulation parameters, primarily varying light intensity. They should present additional tests, such as varying the spot size (which appears to be arbitrarily fixed at 30-50 µm) and the z plane of stimulation. The level of activation can vary considerably: for example, in Figure 3a(iii), identical stimulations elicit responses of markedly different amplitudes (see glom#3 and #4). In Figure 2, 5 out of 15 glomeruli failed to respond - could the choice of z-plane account for this variability? The stimulation duration (50-150 ms) also appears somewhat arbitrary: can the authors demonstrate that the technique is compatible with finer temporal patterns (e.g., 10 Hz stimulation for 500 ms using 20 ms light pulses)? What are the spatiotemporal and axial scanning limits of this approach, and can two or three glomeruli be targeted simultaneously with temporally patterned stimulation?

      Indeed, here we explored a limited set of spatiotemporal parameters to understand the limits of our technique and provide proof of principle. These can be varied depending on the exact question. In the revised manuscript, we will clearly state what the constraints of the current implementation are and provide context for further optimizations. Briefly, in the current version, individual as well as multiple glomeruli can be photo-stimulated together and 20 ms per pulse regime in trains of pulses is feasible.

      (3) One particularly relevant application of this method would be to guide photo-stimulation based on prior functional measurements - for instance, by generating a photo-stimulation mask specifically targeting odour-responsive glomeruli. In the DAT-Cre × Thy1-GCaMP6 experiment shown in Figure 5e, which glomeruli are activated by a given odour, and how does this odor responsiveness influence the efficiency of DAT+ cell-mediated inhibition?

      We thank the Reviewer for the suggestion. We are thinking along exactly the same lines. In particular, we would like to investigate the relationship between the degree of overlap in odor responses of individual glomeruli and the strength and specificity of their inhibitory interactions mediated by DAT+ interneurons. In the revised manuscript, we will further expand on discussing this venue of study. We feel however that this investigation is beyond the scope of this technical report.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Koh and colleagues describe ADePT (Axially Decoupled Photo-stimulation and Two-photon Readout), a modular approach for combining patterned one-photon optogenetic stimulation with two-photon calcium imaging in independently controlled axial planes. The method relies on a digital micromirror device together with a motorized holographic diffuser to generate spatially confined stimulation patterns while imaging deeper neuronal populations. As proof-of-principle applications, the authors use the system to map excitatory and inhibitory functional connectivity in the mouse olfactory bulb by stimulating superficial glomerular circuits and recording responses from mitral and tufted cells in deeper layers.

      This is a well-executed Tools and Resources manuscript. The technical implementation is described in considerable detail, the optical performance is systematically characterized, and the biological experiments provide convincing demonstrations of the types of circuit questions that can be addressed using the method.

      Strengths:

      The greatest strength of the manuscript is the comprehensive technical characterization of the optical system. The authors carefully benchmark the spatial resolution, axial confinement, registration accuracy, calibration procedure, and practical operating limits of the setup. I found the extensive optical benchmarking particularly helpful, as it gives readers a realistic sense of the operating regime and practical limitations of the approach.

      Another strength is the high level of methodological transparency. The optical design, calibration procedures, stimulation strategies, and analysis pipeline are described in sufficient detail that an experienced laboratory could realistically evaluate whether the system is suitable for its own applications. This level of documentation is particularly appropriate for a Tools and Resources article.

      A further strength is the clear positioning of ADePT relative to existing approaches. The authors are transparent about the trade-off between spatial resolution and implementation complexity: ADePT does not provide single-cell photostimulation, but offers flexible axial separation, a large stimulation field, and cellular-resolution two-photon readout in deeper planes without requiring a full holographic stimulation system. This defines a credible and potentially useful experimental niche.

      The biological applications convincingly demonstrate the utility of ADePT. The experiments identifying sister mitral/tufted cells through selective glomerular stimulation and the mapping of heterogeneous inhibitory influences from DAT-positive interneurons illustrate the types of functional connectivity questions that become experimentally accessible with this approach. Importantly, the authors generally avoid overstating these biological findings and appropriately present them as proof-of-principle demonstrations of the technology.

      We thank the Reviewer for their constructive input.

      Weaknesses:

      The primary limitation is inherent to the method itself rather than the execution of the study. Because ADePT relies on one-photon patterned illumination, photo-stimulation remains restricted to relatively superficial structures and does not achieve single-cell spatial resolution. The authors appropriately acknowledge these constraints and clearly position the method within this operating regime. Consequently, ADePT occupies a useful niche for interrogating spatially organized functional units such as olfactory glomeruli or cortical barrels, rather than applications requiring single-cell precision or deeper tissue penetration.

      We agree. In the revised manuscript, we will further expand on these points, highlighting the limitations and advantages of ADePT compared to other techniques in a table discussing various operating regimes.

      Although the manuscript describes the approach as relatively simple and cost-effective, implementation still requires careful optical alignment, registration, calibration, and optimization. This does not diminish the value of the approach, but terms such as modular or accessible may better reflect the practical implementation than simple. Likewise, a brief bill of materials, approximate add-on cost, and indication of which components are essential versus substitutable would help prospective users assess the accessibility of the system.

      We agree. We will change the text accordingly and provide the additional information as suggested.

      Finally, the manuscript provides an impressive level of technical characterization, but much of the practical guidance for adopting the system is distributed across the Results and Discussion. Bringing together the principal limitations, recommended operating regime, expected calibration workflow, evidence for long-term alignment stability, and the circumstances in which ADePT is preferable to alternative approaches would further strengthen the manuscript as a community resource.

      We agree. We will proceed accordingly in the revised manuscript.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors addressed how viral-mediated expression of amyloid in medial septum (MS) cholinergic neurons, or broadband amyloid expression, affects the integrity of MS cholinergic neurons in aging mice, as well as cognition, sleep, and hyperexcitability. Using fiber photometry and viral tracing, they show that MS cholinergic neurons are active during wakefulness and REM sleep and that they also project to many different areas. Next, they show that when they express a viral vector carrying APP to encode amyloid beta in MS cholinergic neurons, these neurons express amyloid as they do in a globally expressing APP model (APP-NLGF). They find that amyloid may spread largely following MS projections and that MS die over time presumably due to amyloid expression. They also describe the emergence of memory deficits and reduced REM sleep attributable to loss of MS cholinergic neurons. Lastly, they report a higher burden of epileptiform activity in mice with broadband amyloid expression and the emergence of neuroinflammation in MS, which may be contributing to cell loss and network dysfunction.

      Strengths:

      (1) New insights on a potential role of MS cholinergic neurons in spreading amyloid.

      (2) Use of several different methods to address effects of MS dysfunction in aging mice (AAV, global, lesioning).

      (3) Combination of activity-related readouts including fiber photometry, EEG coupled to histological, behavioral, tracing, and neuropathology measures.

      (4) Consideration of potential confounds to behavioral measures using proxies of anxiety-related behavior.

      Thank you for the positive assessment and for recognizing the novelty of our findings, the complementarity of our experimental approaches, and the breadth of our multi-modal readouts. We will address the weaknesses raised below point by point.

      Weaknesses:

      (1) The authors aim to model the prodromal phase of Alzheimer's disease (AD) neuropathology, which is a very promising area to target therapeutic intervention. While reduction in basal forebrain volume has been reported early in AD, presumably functional changes may be happening much earlier, i.e., even before MS start to degenerate or before REM sleep is reduced. This view has been proposed by human studies showing increased ChAT reactivity in MCI (PMID: 11835370) and evidence in mouse models showing that MS cholinergic neurons may be hyperactive early and degenerate late with distinct implications for memory (PMID: 41717904). Thus, functional changes could be considered before structural changes could be discussed, as earlier ages in this model could reveal such early changes.

      This is an insightful comment. We fully agree that early functional changes preceding structural degeneration represent an important and exciting avenue, and we will explicitly acknowledge this in the revised manuscript, including the relevant literature on early cholinergic hyperactivity. Examining earlier time points in our model to capture such changes is a compelling perspective that we will discuss as a key direction for future work.

      (2) One limitation of the tracing methodology (Figure 1) that could be improved is sample size, as only 2 mice have been used. Moreover, it would be interesting to conduct the same tracing experiments in APP mice to see how these projections are affected by amyloid pathology.

      We acknowledge this limitation and will increase the sample size for the tracing experiments in the revised manuscript. However, we wish to clarify that the tracing was performed in a separate cohort of young animals specifically to characterize baseline MS cholinergic projections independently of any amyloid-related disturbances. While we agree that replicating these experiments in APP mice would be of great interest, this falls outside the scope of the current study and will be highlighted as an important direction for future work.

      (3) Figure 3 measurements included the whole hippocampal formation, but a region-specific analysis would be warranted as the authors discuss specific accumulation areas.

      We agree with this pertinent suggestion. We will perform and report region-specific analyses of the hippocampal formation in the revised manuscript, in line with our discussion of specific amyloid accumulation areas.

      (4) Figure 5 novel object recognition comparisons use a group of 10 sec exploration, which is unclear why. Novel vs familiar comparisons and reporting of discrimination indexes are considered more robust measurements to report.

      We respectfully maintain our analytical approach. As the test phase was terminated upon reaching a predefined cumulative exploration time of 20 seconds rather than using a fixed trial duration, computing a discrimination index is not appropriate in this context, as total exploration time is constrained by design. We instead followed the validated protocol described by Leger et al. (2013, Nature Protocols; PMID: 24263092), which controls for inter-individual differences in exploratory motivation by fixing cumulative exploration time, ensuring equivalent sampling conditions across animals. We will clarify this methodological choice in the revised manuscript.

      (5) Interictal spike detection would benefit from more methodological detail and examples of spikes detected. Reference 72 does not seem to detail interictal spike detection. Moreover, when during sleep do these spikes happen? It has been shown that they occur primarily during REM sleep when mice show cholinergic hyperactivity (PMID: 37714307). From panel 7B, it seems they occur during NREM, which may be explained by a diminished drive of cholinergic circuits to drive spikes in these mice (vs REM in younger mice). Thus, a NREM vs REM vs Wake analysis will be insightful.

      We appreciate this constructive suggestion. We will provide additional methodological detail on interictal spike detection, include representative examples, and perform a vigilance state-specific analysis in the revised manuscript. We agree this will provide valuable mechanistic insight and will update the reference accordingly.

      Reviewer #2 (Public review):

      Summary:

      In this study, Nollet and colleagues sought to determine whether selective amyloid pathology confined to medial septal (MS) cholinergic neurons is sufficient to recapitulate the prodromal Alzheimer's disease-like phenotypes observed in global AppNL-G-F knock-in mice. To this end, the authors employed a cell-type-specific AAV-mediated approach to selectively express the familial AppNL-G-F allele in MS-ChAT neurons, and subsequently characterized sleep-wake architecture, EEG spectral features, cognitive function, emotional behavior, and histological changes over 13-14 months. By comparing these mice with global AppNL-G-F knock-in mice and with mice in which MS-ChAT neurons were selectively ablated via caspase expression, the authors found that cholinergic cell lesioning recapitulated most disease phenotypes, suggesting that cholinergic loss, rather than amyloid deposition, is a likely driver of these phenotypes.

      Strengths:

      The study has several notable strengths. First, the experimental design is rigorous and well-controlled, employing three complementary mouse models that enable elegant causal inference. The use of cell-type-specific APP expression is a powerful approach for distinguishing the contributions of MS-ChAT neurons and amyloid deposition. Second, the combination of multiple behavioral assessments, EEG spectral analysis using FOOOF parameterization, and detailed histological quantification strengthens the validity of the conclusions. Third, the finding that caspase-induced cholinergic lesions largely recapitulate the cognitive and REM sleep phenotypes, while amyloid pathology contributes additional features such as epileptiform spikes and astrogliosis, represents an important mechanistic dissection.

      Thank you for this positive assessment and for recognizing the rigor of our experimental design, the value of our multi-modal approach, and the mechanistic significance of our cholinergic lesion comparison. We will address the weaknesses below point by point.

      Weaknesses:

      Despite the overall strength of the study, several limitations warrant consideration. First, the mechanism by which amyloid is "broadcast" from MS-ChAT terminals to distant brain regions remains unclear. The authors do not definitively determine whether the amyloid detected in hippocampal and cortical regions represents released soluble Aβ, transported APP fragments, or amyloid derived from degenerating axons. Second, while the authors demonstrate that MS-ChAT cell loss correlates with cognitive, emotional, and REMS deficits, the causal relationship among these phenomena and the specific circuits involved remains unresolved.

      Regarding amyloid broadcasting, we fully acknowledge that the precise mechanism remains to be elucidated; while this was not a primary objective of the study, it represents a fascinating and unexpected finding that we will discuss more carefully as an open question for future investigation. Regarding the causal relationship between MS<sup>ChAT</sup> cell loss and the observed phenotypes, we agree that the specific circuits involved remain to be fully resolved; however, we would like to emphasize that the convergent evidence from our three complementary models (and in particular the recapitulation of cognitive and REM sleep deficits by selective cholinergic ablation) provides strong causal support for MS<sup>ChAT</sup> neuronal loss as a key driver of these phenotypes, independent of amyloid deposition per se.

      Reviewer #3 (Public review):

      Summary:

      The central idea of the study is strong and potentially important: that the vulnerability of the cholinergic medial-septal population can account for a substantial fraction of prodromal-like AD phenotypes, thereby shifting part of the mechanistic focus from cortex-centered pathology to subcortical neuromodulatory circuit failure. The work has several notable strengths. The authors combine circuit mapping, calcium photometry, longitudinal EEG/EMG sleep phenotyping, histology, behavior, and a caspase-based lesion comparison to build a multi-level case for medial septal cholinergic involvement in REM Sleep and memory phenotypes. The inclusion of both a focal amyloid model and a partial cholinergic ablation model is especially valuable because it attempts to separate effects of Ch-neuronal loss from effects of amyloid itself.

      However, the manuscript has several issues, from manuscript formatting to experimental design, overarching statements, insufficient exclusion of alternative explanations, incomplete quantification details for key histological results, a discussion that often moves beyond the actual data into speculative translational framing, and a discussion that completely ignores the early presence of p-tau in human AD patients and even lacks supplementary materials.

      Strengths:

      (1) The conceptual premise is compelling: cholinergic basal forebrain vulnerability is a real and important feature of AD, and testing whether selective medial septal cholinergic pathology can drive REM sleep and cognitive phenotypes is mechanistically interesting and clinically relevant.

      (2) The experimental framework is broad and generally thoughtful, spanning anatomy, function, sleep architecture, EEG spectral parameterization, behavior, and histopathology.

      (3) The projection mapping and photometry provide a useful systems-level introduction, establishing that MSChAT neurons are Wake/REM sleep-active and project strongly to hippocampal and cortical targets before the disease manipulations are introduced.

      (4) The MSΔChAT comparison group is valuable because it allows the authors to argue that some phenotypes track with cholinergic loss rather than amyloid per se.

      (5) The longitudinal sleep analysis is one of the strongest parts of the study, especially the emphasis on REM sleep quantity and bout architecture over time rather than relying only on an endpoint comparison.

      Thank you for acknowledging the compelling conceptual premise of our study, the thoughtful and broad experimental framework, and the value of our longitudinal sleep analysis and lesion comparison. We will address all concerns raised below point by point.

      Weaknesses:

      (1) The title overreaches in its use of "prodromal phase." In the clinic, "prodromal AD" denotes a biomarker‑positive, pre‑dementia phase with subtle, progressive cognitive decline before widespread neurodegeneration, whereas here the authors demonstrate substantial cholinergic degeneration alongside cognitive impairment, which corresponds to advanced pathology within these models rather than a clinically prodromal stage. Moreover, APP knock‑in mice are amyloid‑centric, lack tau pathology, and don't recapitulate human disease staging; therefore, it would be better to avoid terms used for AD staging in the clinic. A more accurate framing of the title would be "Modeling the prodromal-like phase in an Alzheimer's disease mouse model".

      This is a valid point. We agree that the term “prodromal phase” requires more careful framing in the context of animal models, and we will revise the title and relevant statements accordingly. We would like to note, however, that the REM sleep disturbances we report seem to emerge prior to overt cognitive decline in our longitudinal analysis, which we consider to reflect a prodromal-like feature of the model. Nevertheless, we will adopt more precise terminology throughout the manuscript to avoid conflation with clinical staging criteria.

      (2) The opening statement in the abstract (line no 22) is overstated. Current evidence supports that changes in REM sleep, slow‑wave sleep disruption, and excessive daytime sleepiness are associated with a higher risk of AD and reflect early involvement of brain regions vulnerable to AD proteinopathy. No study indicates that REM sleep changes per se are a strong predictor on their own. For example, Jin et 2025 studied REM latency in AD and concluded that prolonged REM latency may be a marker of early neurodegeneration (PMID: 39868572). Thus, the opening statements need to be modified.

      We appreciate this comment and will carefully nuance our opening statement to better reflect the current state of evidence. However, we respectfully note that several reports, including Pase et al. (Neurology, 2017; PMID: 28835407) and Ibrahim et al. (Sleep, 2024; PMID: 38001022), have demonstrated that REM sleep loss is associated with increased risk of incident neurodegenerative disorders, particularly Alzheimer's disease, supporting the broader validity of our framing. We will revise the statement to more accurately capture the complexity of this relationship while preserving its scientific relevance.

      (3) Line 63: The current phrasing of neuromodulators being also essential for orchestrating sleep/wake states is very simplistic. Sleep/wake regulation is a highly complex process involving several interacting neurotransmitters and neuromodulatory systems. I recommend revising this sentence to reflect the broader, multi‑system nature of sleep/wake control.

      We agree and will revise this sentence to better reflect the multi-system complexity of sleep/wake regulation, acknowledging the interplay between multiple neurotransmitters and neuromodulatory systems.

      (4) Line 64: "ACh is required for the generation of REMS" is incomplete. The sentence implies REM sleep generation depends exclusively on ACh. Instead, the sentence must emphasize that ACh is a crucial component of a broader REM sleep circuitry and explain why it is critical for REM sleep.

      We agree and will revise this sentence to clarify that ACh is a crucial component of a broader REM sleep-generating circuitry, rather than a sole requirement, while better contextualizing its specific contribution to REM sleep regulation.

      (5) Line 65: The sentence "Importantly, reductions and alterations in REMS have emerged as strong predictors of clinical AD onset" (Reference 37) is an overstatement of the evidence; Peas et al. 2017 analyzed a dementia cohort that included AD cases and concluded: "Despite contemporary interest in slow-wave sleep and dementia pathology, our findings implicate REM sleep mechanisms as predictors of clinical dementia." The authors should rephrase this to reflect that the study examined REM sleep changes in a mixed dementia population with AD, rather than to establish REM alterations as strong, standalone predictors of AD onset.

      We will revise this statement to more accurately reflect the evidence. We would like to note, however, that while the Pase et al. (2017) cohort included a mixed dementia population, 75% of incident dementia cases (24 out of 32) were consistent with Alzheimer's disease, lending meaningful support to the relevance of REM sleep alterations specifically in the context of AD. We will ensure this nuance is clearly conveyed in the revised manuscript.

      (6) Lines 73-75 address human Alzheimer's studies and state that basal BF-Ch neurons are vulnerable to Aβ but largely omit the well-established contribution of early tau pathology. In human AD patients, p-tau accumulation in BF is an early event (Braak I-II) and is closely associated with BF-Ch neuronal loss and BF atrophy and has been documented extensively. By relying almost exclusively on Aβ-centric framing, the current text risks implying that BF-Ch degeneration is solely amyloid-driven, which is not accurate. Even though the mouse model used here is "amyloid-heavy" and lacks tau pathology, the introduction should acknowledge the role of p-tau (especially when the paragraph contextualizes human studies) and clarify that in humans, BF-Ch vulnerability reflects converging amyloid and tau insults, so that readers do not infer a purely amyloid-dependent mechanism from the way the background is presented.

      We agree and will revise this section to acknowledge the well-established contribution of tau pathology to BF cholinergic neuronal vulnerability in human AD, including its early accumulation at Braak stages I-II. We wish to clarify, however, that the present study focuses exclusively on amyloid-driven mechanisms, and the introduction will be revised to ensure readers do not infer a purely amyloid-dependent mechanism in the broader human disease context.

      (7) Line 92: and elsewhere in the manuscript, I recommend avoiding the term "prodromal phase" and instead using the phrase "prodromal-like phase in an AD mouse model". The authors should be more precise in describing the disease stage in animal models that don't recapitulate human disease staging and ensure that clinical staging terminology is specific to human studies.

      As noted in our response to weakness (1), we will systematically revise the manuscript to replace “prodromal phase” with more precise terminology that clearly distinguishes our animal model findings from clinical disease staging.

      (8) Age and duration of pathology are major concerns. The different models are not adequately matched for amyloid exposure duration and age at testing. Age is the strongest risk factor for AD, and varying both chronological age and time under pathology across groups is a major design flaw. In MSChAT-AppNL-G-F/GFP mice, AAV injection was delivered at 11-13 weeks of age, and animals were sacrificed at 13-14 months post-injection (roughly 15-16 months old), whereas AppNL-G-F/NL-G-F knock-in mice and APPWT were 13-14 months old at the time of termination. Thereby, there is a difference in the duration of Aβ exposure across models. This mismatch directly weakens comparisons such as the lower epileptiform spike counts in MSChAT-AppNL-G-F versus AppNL-G-F/NL-G-F mice, because differences could simply reflect shorter cumulative pathology exposure rather than a genuinely weaker circuit-specific effect.

      The same issue affects the internal control logic of the MSΔChAT model, which is intended to isolate cholinergic neuron loss from amyloid aggregation. For this comparison to be clean, ages and exposure durations should be aligned as closely as possible. Instead, MSΔChAT mice are tested earlier than the AppNL-G-F/NL-G-F and MSChAT-AppNL-G-F/MSChAT-GFP cohorts, introducing a 4 to 7-month age gap that complicates attribution of phenotypic differences solely to cholinergic loss versus amyloid pathology.

      Finally, the absence of sham-operated controls is a concern, as it prevents separating the effects of the surgical procedure and AAV delivery from those of amyloid expression or cholinergic ablation.

      Thank you for raising these important points. Regarding age matching, we acknowledge that chronological ages are not perfectly aligned across groups; however, we wish to emphasize that the duration of amyloid pathology is carefully matched across models. Indeed, AAV injection in MS<sup>ChAT</sup>-AppNL-G-F mice marks the onset of amyloid expression, directly paralleling the onset of pathology from birth in App<sup>NL-G-F/NL-G-F</sup> knock-in mice. We believe pathology duration represents the most biologically relevant variable for comparison in this context, and we will clarify this in the revised manuscript. Regarding the MS<sup>ΔChAT</sup> cohort, animals were culled upon reaching a comparable degree of REM sleep loss, providing a functionally meaningful matching criterion. Finally, regarding sham-operated controls, we acknowledge this limitation; however, based on our experience, surgical procedure alone has negligible effects on the cellular populations under study, and the inclusion of an additional sham group across all experimental cohorts would have required a prohibitive number of animals, raising significant ethical concerns under the 3R principles. We will address these points more explicitly in the revised manuscript.

      (9) Line 115 through 117: The text cites Figure 2D, but does not refer to Figure 2C for the statement "their phenotypes were then compared in detail with MSChAT-AppNL-G-F and AppNL-G-F/NL-G-F global knock-in mice that were aged at the same time". Figure 2C depicts D54D2 amyloid staining in MSChAT-GFP vs MSChAT-AppNL-G-F mice. For clarity and consistency, I suggest adding a Figure 2C notation to this sentence (e.g., "Figures 2A, 2C").

      Thank you for this observation, we will correct the figure citation accordingly in the revised manuscript.

      (10) In Figure 1C-D, the authors map MSChAT projection targets across a wide range of brain areas, including hippocampal subfields, mPFC, primary cortices, entorhinal cortex, olfactory bulb, thalamus, anterior hypothalamus, amygdala, and medial habenula, and identify several of these as substrates through which MSChAT activity could influence REM sleep and cognition. However, the lateral hypothalamic area (LHA) is conspicuously absent from both the listed projection targets and the tracing panels shown in Figure 1D, despite the anterior hypothalamus being reported as an innervated region.

      This omission is notable given that LHA-MCH neurons are among the best-established REM-sleep-promoting neurons, and the authors themselves cite prior work implicating LHA-MCH neurons in the AppNL-G-F REM sleep phenotype (ref. 49, 107; line 403) as an alternative cell-circuit candidate, a claim they explicitly try to weigh against their own MSChAT-centered model in the discussion.

      a) The MSChAT neurons are reported to be REM sleep- and wake-active (Figure 1A-B), the same vigilance-state profile as LHA-MCH neurons,<br /> b) The Discussion directly engages with LHA-MCH neurons as a competing/complementary REM sleep-generating mechanism, and<br /> c) The reported anterior hypothalamus innervation (Figure 3C) raises the question of whether MSChAT axons specifically innervate LHA, and whether any projections specifically to LHA or LHA-specific amyloid deposition were examined. Clarifying this would help position the proposed MSChAT-hippocampal circuit mechanism relative to the well-established LHA-MCH REM sleep node.

      We appreciate this important observation. We will carefully re-examine our tracing data to determine whether MS<sup>ChAT</sup> axons specifically innervate the LHA, and whether amyloid deposition was detectable in this region in our MS<sup>ChAT</sup>-App<sup>NL-G-F</sup> model. We agree that clarifying the potential anatomical relationship between MS<sup>ChAT</sup> projections and LHA-MCH neurons is important to properly position our proposed circuit mechanism relative to this well-established REM sleep-promoting node, and we will address this in the revised manuscript.

      (11) Line 125: "13- to 14-month-old MSChAT-AppNL-G-F mice immunohistochemical analyses employing the amyloid-specific antibodies....", in the methods section (Line 652) the authors mention MSChAT-AppNL-G-F and MSChAT-GFP mice were perfused 13-14 months after AAV injection (age at the time of injection was 11-13 weeks of age). This leaves the question of how they have 13- to 14-month-old MSChAT-AppNL-G-F mice available to study Amyloid-β load.

      Thank you for catching this inconsistency. We confirm that this is an error in the manuscript: line 125 should read “15-16 month-old” referring to the chronological age of the animals at the time of perfusion, rather than “13-14 months,” which corresponds to the duration of AAV expression. We will correct this in the revised manuscript.

      (12) Line 174-175: As currently written, the sentence could be read as both wild-type and homozygous AppNL-G-F/NL-G-F mice received AAV injections and were then aged 13-14 months post‑injection. In fact, the Methods clearly state that knock‑in mice are simply aged from birth without any AAV manipulation. The sentence should be rephrased to avoid suggesting that global APP knock‑in animals are part of the AAV‑injected cohorts.

      Thank you for flagging this ambiguity. We will revise the sentence to clearly distinguish between AAV-injected and global knock-in cohorts in the revised manuscript.

      (13) Lines 182-183, 196-197, and 209 refer to "Supplementary information" and imply that detailed behavioral data and analyses are provided in that section. However, in the current submission, the supplementary material consists only of Figures S1-S7 (Amyloid marker and cerebral vasculature, Aβ in hippocampus, GABA and glutamatergic neurotransmission, and sleep/wake parameters) and does not include supplementary figures or tables for the behavioral assays described in the main text. This discrepancy makes it impossible to verify the full behavioral dataset and the analyses referred to in the results section. The authors should carefully check the submission package and ensure that all referenced supplementary figures, tables, and detailed behavioral results are included and appropriately labeled.

      The supplementary information referenced in the main text will be provided in full in the revised manuscript, together with analyzed datasets and analysis scripts, in accordance with eLife's data sharing policy.

      (14) The lack of details for histological quantification is a major concern for a manuscript in which major conclusions hinge on Aβ load and MS-Ch neuronal counts. The histological quantification section is severely under-specified. The authors describe a 23% MSChAT loss, differences in regional Aβ burden, and a vascular association; however, the methods section is strangely silent about the quantification pipeline. For Aβ quantification, it is not clear whether "load" reflects percent positive area, plaque counts, or another metric; which Fiji thresholding algorithm(s) were used; how ROIs were defined; how staining batch effects were controlled; and how autofluorescence was normalized. For neuronal counts, the strategy for identifying and counting ChAT-positive neurons, normalization, and blinding are not described. There are no details on section spacing, axis of counting, the number of sections counted per animal, or whether both hemispheres were analyzed. Given that the reported differences are modest and central to the main claims, a more detailed and rigorous description of the image-analysis pipeline is essential.

      We agree that a comprehensive description of our histological quantification pipeline is essential. We will provide full methodological details in the revised manuscript, including: the metric used for Aβ load quantification, ROI definitions, Fiji/ImageJ thresholding algorithms (accounting for batch effects and autofluorescence normalization), as well as the strategy for identifying and counting ChAT-positive neurons, section spacing, number of sections per animal, hemisphere coverage, normalization, and blinding procedures. All ImageJ scripts will be made available to ensure full transparency and reproducibility.

      (15) Statistical annotations in figures: There is inconsistency in how statistical significance is indicated across the figures. For example, in Figure 5C, the significance between MSΔChAT and AAV‑Aβ<sup>-</sup> is indicated by a connecting bracket (**), whereas the comparison between AAV‑Aβ<sup>-</sup> and AAV‑Aβ<sup>+</sup> is marked by asterisks (***) placed above AAV‑Aβ<sup>+</sup>. In addition, the single asterisk above KI-Aβ<sup>+</sup> does not clearly specify which pairwise comparison it refers to (e.g., AAV‑Aβ<sup>-</sup> vs WT‑Aβ<sup>-</sup> or another contrast). This heterogeneity makes it difficult to decipher exactly which group comparisons have been tested and found significant. The notation should be standardized and explicitly linked to the corresponding pairwise comparisons (for example, by using consistent brackets/lines and specifying all contrasts in the figure legend). Figures must be self-explanatory.

      We acknowledge that the current statistical annotations can be difficult to interpret when multiple experimental groups are displayed within a single plot. While the notation is consistent across figures, we agree that clarity can be improved, and we will revise all figure annotations to explicitly link significance indicators to their corresponding pairwise comparisons, using standardized brackets throughout, with all contrasts clearly specified in the figure legends.

      (16) Figure 5D statistical notation and group comparisons: The statistical markings in Figure 5D do not seem to match the results text and are difficult to interpret. The authors state that both MSΔChAT and AAV‑Aβ<sup>+</sup> mice lack a preference for the novel object compared with AAV‑Aβ<sup>-</sup> controls, yet the figure does not clearly indicate significance for MSΔChAT versus AAV‑Aβ<sup>-</sup>, and the notation over AAV‑Aβ<sup>+</sup> is ambiguous. As a result, it is unclear which group differences are being tested and reported. It would be preferable to use the standard convention of placing significance annotations directly over the experimental groups (e.g., AAV‑Aβ<sup>+</sup>, KI-Aβ<sup>++</sup>, MSΔChAT) or use notation above brackets to ensure that the figure labels are fully consistent with the statistical statements in the results.

      We wish to clarify that in Figure 5D, the absence of novel object preference manifests as equal exploration of both objects (approximately 10 seconds each), such that comparisons against the 10-second chance level reflect this lack of preference. We will revise the annotations and figure legend to make the statistical comparisons explicit and fully consistent with the results text.

      (17) The discussion contains many compelling ideas, but it needs pruning and recalibration. The best discussion points are those linking the lesion comparison to REM sleep/cognitive outcomes and those situating MS cholinergic neurons within broader REM sleep circuitry. The least convincing sections are those implying disease-stage equivalence, prion-like spread, and direct therapeutic implications without sufficient evidentiary support.

      This is a constructive feedback, and we agree that the Discussion would benefit from pruning and recalibration. We will streamline it to focus on the most evidentially supported points, particularly those linking cholinergic loss to REM sleep and cognitive outcomes, while toning down or removing speculative statements regarding disease-stage equivalence, prion-like spreading mechanisms, and direct therapeutic implications.

      (18) Line 448: The authors discussing reduced anxiety-like behavior in their model corroborates with the 3xTg mouse model (Ref: 116). Interestingly, they don't consider or include reports of anxiety-like disorders from human cohort studies that indicate the prevalence of higher anxiety and its association with preclinical and prodromal AD stages (SCD, MCI) and progression of AD. This apparent contradiction with the human literature is not discussed in the discussion section. The authors should explicitly address how their anxiolytic-like phenotype fits with clinical data (e.g., species differences, task specificity, disease stage, or model limitations) and clarify whether they view this as a limitation of the model or as evidence for a more complex relationship between amyloid, cholinergic dysfunction, and emotional behavior.

      Thank you for raising this important point. We will expand the Discussion to address this apparent contradiction with the human literature. Indeed, while increased anxiety is reported in early AD stages, it tends to normalize or decrease at later stages (Botto et al., 2022; PMID: 35461471), which may partly reconcile our findings. In addition, anxiety-like phenotypes are highly inconsistent across AD mouse models, varying with model type, age, sex, and behavioral assay (Pentkowski et al., 2021; PMID: 33979573). Anxiety-related changes in human AD may reflect damage to brain regions beyond the MS cholinergic system, involving additional circuits and mechanisms not captured by our model.

      (19) Line 654 states, "Comparable durations of amyloid pathology," but this is not fully substantiated, as the onset and progression of amyloid in the AAV-driven MSChAT-AppNL-G-F model versus the global AppNL-G-F knock-in model are not described. The data support comparison at a similar late-stage amyloid burden, but not necessarily equal duration of pathology.

      We will nuance our wording at line 654 by replacing “comparable durations of amyloid pathology” with “comparable amyloid burden,” acknowledging that while both cohorts were aged for matched durations, the kinetics of amyloid progression may inherently differ between an AAV-driven focal model and a germline knock-in model.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper from Bardossy et al. explores whether viral macrodomains in dual-host viruses contribute to infection in the mosquito vector. Using the CHIKV Caribbean strain, the authors generated nsP3 macrodomain catalytic site mutants (N24A or N24D) and identified a compensatory mutation site at position 31 during virus propagation in Vero cells. They then assessed the impact of these mutations on viral growth kinetics in A549 (human) and U4.4 (Ae albopictus cells), as well as on infectivity and dissemination in vivo in Ae. aegypti and Ae. albopictus. Biochemical and structural analyses of recombinant macrodomain proteins (alone or in combination) revealed effects on stability, catalytic activity, and ADP-ribose binding. Overall, the study demonstrates that CHIKV macrodomain catalytic activity plays an important role in virus infectivity and dissemination within the mosquito vector.

      Strengths:

      A complete set of experimental approaches spanning generation of recombinant viruses, in vitro characterization, in vivo studies in mosquitoes, and detailed biochemical and structural characterization.

      Weaknesses:

      (1) The sequence analysis of the generated stocks revealed the emergence of a second-site mutation at position 31 of the nsP3 macrodomain when (N24A or N24D) CHIKV mutants were generated on Vero cells. However, it is not clear from the text or the experimental design how many independent replicates were performed. Based on the current description, it appears this was done only once, which raises the question of whether mutations at position 31 represent a reproducible outcome of infection. This is particularly important because experiments in A549 cells did not reveal emergence of mutations at position 31. To strengthen this finding, the experiment should be performed at least three independent times.

      We thank the reviewer for raising this important point. The emergence of second-site mutations at residue D31 in Vero cells was actually observed in two independent experiments, each initiated by transfection of viral RNAs encoding either the N24A or N24D mutant. In the second experiment, additional timepoints were sampled for sequencing. Comparison of the two experiments revealed that, in the second replicate, only the D31N variant was recovered, whereas D31H was not detected. We will include the results of both independent experiments in the updated Figure 1 and will modify the text accordingly to improve clarity.

      In addition, we will present new experiments performed in other cell types that further confirm the reproducibility of D31 mutation emergence across different cellular contexts. Specifically, we will include results from new viral RNA transfections in BHK-21 mammalian cells and C6/36 mosquito cells, which will be included in the updated Supplementary Figure 1.

      (2) Based on the primer information used to generate amplicons for sequencing, the amplicons evaluated do not span the full nsP3 gene as stated in the text (Line 105). Instead, they cover only the first 119 amino acids of the macrodomain (160 aa long). Thus, the current data do not rule out the emergence of other compensatory mutations elsewhere in the nsP3 macrodomain or in the full-length protein. Additional sequencing is recommended, or the text should clearly state that only a portion of the macrodomain was sequenced.

      We thank the reviewer for this correction. The sequencing was indeed focused on the region of the macrodomain surrounding the introduced mutations, covering the first 119 amino acids of nsP3. This region was selected to confirm the stability of the introduced mutations at position 24 and to monitor the emergence of potential second-site mutations in its immediate vicinity. We acknowledge that this approach does not rule out the emergence of compensatory mutations elsewhere in the macrodomain or in the full-length nsP3 protein. We will correct the text accordingly to accurately reflect the region that was analyzed.

      (3) Another key question is whether this is a specific feature of the Caribbean strain or a feature conserved across different CHIKV lineages.

      This is an excellent point. To address it, we introduced N24A and N24D mutations into an infectious clone of the Indian Ocean strain, a representative of the ECSA lineage, and assessed the emergence of D31 second-site mutations during viral stock production. As observed with the Caribbean strain, the D31N secondary mutation consistently emerged in both N24A and N24D Indian Ocean mutant viruses. These results will be included in the updated Supplementary Figure 1 and in the main text.

      (4) The use of A549 cells (interferon-competent) to study CHIKV infection is somewhat surprising, as the current literature indicates that this cell line is not efficiently infected by Asian or ECSA lineages of CHIKV (PMID: 17604450) unless the Mxra8 receptor is overexpressed (PMID: 29769725) or IFN signaling is inhibited (PMID: 31682641). The data presented here are compelling and suggest specific features of the Caribbean strain that enable efficient infection of this cell line (Do the authors observe detectable cytopathic effect (CPE) in CHIKV-infected A549 cells?).

      However, to further support the authors' claim related to human immunocompetent cells, it would be important to demonstrate the phenotype in an additional interferon-competent cell line that is well-established as highly permissive to CHIKV, such as human fibroblasts.

      We thank the reviewer for raising this point. We acknowledge that previous studies have reported limited infection of A549 cells by certain CHIKV strains. However, in our hands, with our viral stocks and under our experimental conditions, we observe an increase in viral titers following infection of A549 cells with WT Caribbean strain virus, indicating productive viral replication. We also confirmed that the Indian Ocean strain replicates in A549 cells under the same conditions, and we will include growth curve data for this strain in the updated version of the manuscript. Importantly, we did not observe detectable cytopathic effects in A549 cells infected with either WT or N24 mutant viruses, which is consistent with the notion that CHIKV replicates less efficiently in this cell line compared to other mammalian cell lines. We will include a sentence acknowledging this in the discussion of the revised manuscript.

      Regarding the suggestion to use human fibroblasts, we respectfully note that the primary focus of this study is the role of macrodomain catalytic activity in the mosquito host, and the experiments in A549 cells were performed to confirm the known importance of the macrodomain in interferon-competent mammalian cells. We therefore consider the current data in A549 cells sufficient to support this conclusion within the scope of the manuscript.

      (5) To fully support the conclusion stated in lines 234- 237, the authors should fully sequence the virus stock used to demonstrate that no additional mutations (beyond N24D-D31H/N) are present that could contribute to the enhanced dissemination phenotype. This is especially important if the experiment was performed with only one stock of virus, given justified gain-of-function concerns.

      We acknowledge that the full viral genome was not sequenced, and we cannot rule out the presence of additional mutations elsewhere in the genome that could contribute to the observed phenotype. However, we note that the enhanced dissemination phenotype was also observed in independent experiments performed with the Indian Ocean strain mutant viruses in Ae. albopictus, which were generated independently from the Caribbean strain stocks. The consistency of the phenotype across two independently generated sets of mutant viruses from different CHIKV lineages strongly supports the conclusion that the enhanced dissemination phenotype is linked to the macrodomain mutations. These new data will be included in the updated manuscript.

      (6) The authors did not assess transmission but transmission potential (only viral dissemination to heads was measured). The sentence at line 360 should be modified to accurately reflect the data-supported conclusion.

      We thank the reviewer for pointing this out. The text will be modified accordingly to accurately reflect that we assessed transmission potential, based on viral dissemination to heads, rather than actual transmission.

      Reviewer #2 (Public review):

      Summary:

      To address how the CHIKV macrodomain contributes to replication dynamics in mammalian and insect hosts, the authors initially created two separate mutations in the highly conserved N24 residue, which is known to be critical for the CHIKV macrodomain's ability to erase ADP-ribose from target proteins. Interestingly, they could not produce a virus with a mutation in this residue without second-site mutations in an aspartic acid residue nearby (D31). However, when tested biochemically, these second-site mutations did not enhance the enzymatic activity of the protein, indicating that other enzyme dynamics, such as substrate binding, may be impacting these mutations. Mutations at this residue allowed the CHIKV to replicate in Vero cells and in mosquito cells, but they replicated poorly in IFN-competent human cells, indicating clear IFN-specific impacts on these viruses. Interestingly, they found unique impacts on virus dissemination and replication in live mosquitoes. While the N24A/D31N virus did poorly in vivo in all accounts, the N24D/D31H/N virus tended to infect both the bodies and heads of the mosquitoes better than the WT virus, though titers were reduced. The authors claimed, based on a DSF assay, that there were no real differences in ADP-ribose binding and thus suggested that these differences could be due to changes in substrate specificity, as the D31 residue resides in the substrate exit path, potentially tuning the virus to unique substrates in different species. The authors also produced crystal structures of the mutants to demonstrate the changes in the binding pocket caused by these mutations.

      Strengths:

      The authors have done a rigorous job of evaluating CHIKV macrodomain mutant viruses and the proteins' biochemical activities. The use of live mosquitoes is highly unique and provides important insights into the importance of the macrodomain in different species.

      Weaknesses:

      It is not clear if the interpretation of the ADP-ribose binding data is correct. It appears there are notable differences that could explain the results, though the authors chose to minimize the impact that these differences had on the results. The N24D-D31H/N proteins had at least a 1C degree difference in the thermal shift assay when compared to the N24A/D31N, single D31 mutants, and WT proteins, which is likely significant and could explain the dichotomous results between the two viruses in mosquito cells. Even the single N24D mutant had enhanced binding compared to the WT protein. Furthermore, as this virus has no enzymatic activity, one could hypothesize that enhanced binding to a substrate that is normally cleaved by the protein could certainly lead to alterations in phenotypic effects, whether good or bad. The authors should test the binding activity in a separate assay, such as an ITC assay, to determine if there are, in fact, binding differences or not. Having said this, it is likely that the impacts of these mutations on replication and transmission in human and mosquito cells are multi-factorial and could include both enhanced binding with altered substrate specificity amongst other activities.

      We agree that the mutants may indeed have stronger binding for modified substrates than the WT protein; however, given that DSF is not a quantitative measure of binding affinity and that free ADPr is not the relevant ligand (in fact we do not know the relevant ADPr-modified molecule), we have refrained from speculating further than saying in the discussion:

      “The progressive selection of D31H over D31N in the mosquito host further suggests that subtle differences in ADP-ribose substrate recognition may influence viral fitness in the mosquito environment in ways that are not yet understood.” (Lines 368-370)

      Additionally, as both mutants had no detectable enzymatic activity but had quite different phenotypes in mosquitoes, I don't agree with the title stating that catalytic activity modulates dissemination and transmission potential in mosquitoes. It seems more likely that alterations in binding activity or substrate recognition (even suggested by the authors) impact these phenotypes in mosquitoes.

      Regarding the title, we agree with the reviewer that it could be misleading, as both mutants lack catalytic activity yet show distinct phenotypes in mosquitoes. We will therefore modify the title to: "Loss of macrodomain catalytic activity modulates Chikungunya virus dissemination and transmission potential in Aedes mosquitoes", which more accurately reflects that the observed phenotypes arise as a consequence of the loss of catalytic activity at position N24.

      Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the nsP3 macrodomain catalytic activity in the replication and transmission of CHIKV in mosquito vectors. The conserved dual-host alphavirus catalytic site N24 has previously been shown to be essential for ADP-ribosylhydrolase activity. Despite this, mosquito-specific alphaviruses do not share this catalytic site. To assess whether the macrodomain catalytic activity of a dual-host virus was essential in insect hosts, the authors targeted the N24 site to abolish catalysis while maintaining binding capacity. The loss of ADP-ribosylation led to the emergence of compensatory mutations at site D31 that impact viral infectivity, dissemination, and transmission in Aedes sp. mosquitoes in vivo. The conclusions are well supported by the results and provide insight into the importance of nsP3 macrodomain activity in the mosquito vector, which hasn't been explored before.

      Strengths:

      The main strength of this study is the use of Aedes sp. mosquito models to investigate the selective pressure of macrodomain mutations in vivo. The functional characterization as well as the structural analysis of the mutants provide supporting evidence of a potential role of the compensatory mutations at site D31 in substrate recognition.

      Weaknesses:

      A considerable part of this study relies on the use of N24 mutant viral stocks generated in Vero cells, which yields an additional mutation at site 31 and consequently doesn't allow the authors to properly dissect the effect of mutation of N24 and D31 independently. It would be recommended to generate stocks with individual mutations in both A549 and U4.4 cells, pooling and concentrating them if needed. Replication of the N24A mutant in A549 cells does not lead to mutation at residue 31. Yet surprisingly, there is no reversion from N back to D at site 31 when the double mutant Vero stocks are passaged in A549. Since they are double mutants, it isn't possible to assess whether the defects in the growth of mutants N24A/T-D31N and N24D-D31H/N compared to WT are due to site 24 or 31, or both (Figure 2, panel c). Even though the authors emphasize that the compensatory mutation could have additional roles that impact viral infectivity and transmission in mosquito cells, it would strengthen the work to show that these mutations would spontaneously appear in stocks generated directly in mosquito cells. As a corollary, is it known whether insect-specific alphaviruses that lack macrodomain catalytic activity have corresponding mutations at site 31?

      We thank the reviewer for this important comment. Regarding the generation of viral stocks in A549 cells, transfection of N24A and N24D viral RNAs into A549 cells did not yield sufficient viral titers to produce usable stocks. However, as described in our response to Reviewer #1, we confirmed the reproducible emergence of D31 second-site mutations in C6/36 mosquito cells and BHK-21 mammalian cells, which will be included in the updated Supplementary Figure 1. These results demonstrate that D31 mutations spontaneously emerge in stocks generated directly in mosquito cells, addressing the reviewer's concern.

      Regarding the question about insect-specific alphaviruses, we examined the sequence at position 31 using a multiple sequence alignment of 14 alphaviruses with diverse host ranges, including dual-host, insect-specific, and aquatic alphaviruses. We observed that Yada Yada virus (GenBank: QGR15362.1) has an asparagine (N), Tai Forest alphavirus (GenBank: YP_009333615) has an aspartic acid (D), Mwinilunga alphavirus (GenBank: BBC45634.1) has an aspartic acid (D), Eilat virus (GenBank: QBG67155.1) has an aspartic acid (D), and Agua Salud alphavirus (GenBank: QEV83787.1) has a lysine (K) at this position. These results suggest that insect-specific alphaviruses do not share a conserved residue at position 31, and therefore no clear conclusion can be drawn regarding a direct correspondence with the compensatory mutations observed in our study. This sequence alignment with the corresponding text will be included as supplementary data in the revised manuscript.

      Additionally, there is a lack of consistency in the prevalence of WT virus at days 5 and 7 in in vivo experiments with Ae. albopictus and Ae. aegypti (Figure 3 and Supplementary Figure 2). This raises concern about the reproducibility of these experiments.

      We thank the reviewer for this observation. We acknowledge that the prevalence of WT virus infection shows variability between experiments and timepoints. Based on our experience with infectious blood meal experiments, this variability is sometimes observed between independent experiments even under identical experimental conditions and with the same virus. In this particular case, each set of experiments was performed with independently produced viral stocks. Specifically, for the experiments shown in Figure 3, WT, N24A-D31N, and N24D-D31H/N viral stocks were produced in parallel from transfection of viral RNAs, and the same stocks were used for sequencing, growth curves, and mosquito infections. Subsequently, when we generated single D31H and D31N mutant viruses, a new WT viral stock was produced in parallel with the D31 mutant stocks, and these independently produced stocks were used for the experiments shown in Supplementary Figure 2. Differences in absolute infection rates between experiments are therefore expected, as they reflect both the use of independently produced viral stocks and the inherent variability in the efficiency of midgut infection and dissemination between mosquito cohorts. Importantly, the comparisons between WT and mutant viruses are always made within the same experiment, using stocks produced in parallel.

      The inability to tease apart the roles of N24 and D31 in mosquito hosts partially prevented the authors from fully achieving their aims, but the work is nonetheless of interest to the field and suggests that more work is necessary to fully understand the role of the nsP3 macrodomain and its catalytic activity in the two disparate but obligate hosts for CHIKV and other dual-host alphaviruses.

    1. Author response:

      Reviewer #1 (Public review):

      A previous study from the same team (McDougle & Taylor, 2019) demonstrated that explicit strategies during visuomotor adaptation can be dissociated into retrieval-based and algorithmic strategies. However, whether these distinct forms of explicit processing differentially influence implicit recalibration has remained unresolved, with previous studies providing evidence both for relatively independent explicit and implicit processes and for interactions between them. This study addresses this question through a series of experiments that used Critical and Non-Critical targets to induce distinct strategic modes while maintaining comparable adaptation at the Critical target.

      Experiment 1 replicated previous findings showing broader implicit generalization under algorithmic strategies. However, this broader generalization could be explained by spillover effects arising from adaptation at the Non-Critical targets. Experiment 2 was designed to reduce such spillover effects by increasing the spatial separation between the Critical and Non-Critical targets. Although broader generalization was still observed in the algorithmic condition, this effect was interpreted as reflecting greater variability in reaching behavior at the Critical target. Finally, Experiment 3 introduced additional controls using an error-clamp paradigm, and the difference in generalization width between the two strategies largely disappeared.

      We appreciate the reviewer’s thoughtful and comprehensive summary of our study.  One thing we would like to clarify is that the non-critical targets used in Experiment 1 received the same type of online continuous feedback as the critical target.Therefore the extent of implicit recalibration should be comparable between the critical and non-critical targets in Experiment 1. In contrast, for Experiment 2, we tightened the control for implicit recalibration at the non-critical targets by: 1) delivering delayed endpoint feedback for all non-critical targets while keeping the online continuous feedback for the critical target, which is known to suppress implicit recalibration, and 2) we widened the spatial gap between the critical and non-critical targets, to minimize spillover 2) As a result, the implicit recalibration was diminished at the non-critical targets, contributing to the shrinkage of the generalization curve around the critical target. 

      One other issue that we would like to make clear is that Experiment 1 was not a straight replication of a previous study, at least to our knowledge. We believe that the reviewer is referring to our previous study (McDougle and Taylor 2019), which found broader generalization for algorithmic strategies (Experiment 4). However, in that study cursor feedback was always delayed. As such, the observed broader generalization was most likely due to the strategy itself and not implicit recalibration. 

      Together, these findings led the authors to conclude that implicit recalibration is relatively insensitive to the type of explicit strategy employed and is primarily shaped by the statistics of the movement plans on which learning occurs.

      The experimental design using Critical and Non-Critical targets is particularly interesting and represents a creative approach to manipulating strategy use. Reaction times were generally longer in the algorithmic group, even at the Critical target, suggesting that the manipulation was at least partially successful in biasing participants toward algorithmic versus retrieval-based strategies. The results that the implicit recalibration is independent of the explicit strategy (how you aim) but depends on the aiming point by the explicit strategies (where you aim) are basically reasonable.

      We are glad to know that our primary finding and conclusion appears reasonable. While we acknowledge that the finding doesn’t appear to be particularly exciting at face value, it does speak to larger questions regarding the independence of different learning systems and how just statistical or surface-level differences in training can result in relatively large differences in apparent behavior that could be easily misinterpreted as the result of system interactions.

      I would like the authors to clarify two points.

      First, how reasonable is it to infer the use of distinct explicit strategies primarily from reaction time differences? While longer reaction times in the algorithmic group are consistent with greater computational demands, it remains unclear whether the longer reaction times observed at the Critical target necessarily reflect different strategy implementations at that location. In particular, could the increased cognitive demands associated with the Non-Critical targets in the algorithmic condition have carried over to the Critical target, thereby prolonging reaction times without implying qualitatively different strategies at the Critical target itself?

      This is a fair concern, as RT is an indirect marker of strategy use and, by itself, cannot establish that participants used different strategies at the Critical target. Prior work, however, provided guidance for the design of our experimental manipulations. Algorithmic strategies, in which an aiming solution is computed online, are associated with longer RTs, whereas retrieval of a previously cached stimulus–response association produces substantially shorter RTs (McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Moreover, caching becomes increasingly difficult as the number of target-specific solutions increases, particularly beyond approximately four targets (Velazquez-Vargas and Taylor, 2024; Bejjanki and Taylor 2026). Our manipulation was designed around these findings: participants in the Algorithmic condition learned the 45° rotation across 10 targets, whereas participants in the Retrieval condition repeatedly encountered the 45° rotation only at the Critical target. As expected with this experimental design, RTs at the Critical target were significantly longer in the Algorithmic condition across all three experiments.

      We agree with the reviewer, however, that this RT difference could in principle reflect a more general carryover of cognitive demands from the Non-Critical targets rather than online computation at the Critical target itself. We can address this possibility more directly by asking whether RT at the Critical target exhibits the parametric signature expected of an algorithmic process. A defining feature of mental rotation is that RT scales with the magnitude of the computed aiming solution (Georgopoulos and Massey, 1987; Bhat and Sanes, 1998; McDougle and Taylor, 2019; Velazquez-Vargas and Taylor, 2024). Although rotation magnitude was fixed in the present experiments, participants’ actual reach angles varied naturally from trial to trial. Indeed, McDougle and Taylor (2019) originally demonstrated this relationship using actual reach angle rather than imposed rotation magnitude. We can therefore test whether trial-by-trial RT covaries with reach angle at the Critical target in the Algorithmic condition but not in the Retrieval condition. Such a relationship would be difficult to explain as a nonspecific carryover of cognitive load and would instead provide direct evidence that preparation time at the Critical target reflects the computation of the aiming solution.

      We also observe a second, independent difference at the Critical target: reach angles are consistently more variable in the Algorithmic condition than in the Retrieval condition across all three experiments (Figure S6). This pattern is consistent with repeated online computation producing variability in the selected aiming solution, whereas retrieval of a cached stimulus–response association produces a more stable response. Importantly, a general carryover account based solely on increased cognitive demands does not readily explain why movements to the Critical target should also be systematically more variable. Nor is this pattern easily explained by a speed–accuracy tradeoff: the Algorithmic group had more preparation time yet nevertheless exhibited greater variability. Consistent with this notion, Velázquez-Vargas and Taylor (2024) found that retrieving cached solutions produced less variable and more precise movements than movements that are not cached in the memory trace. 

      Third, we can conduct additional analysis to compare RT variability between algorithmic and retrieval groups at the critical target location. According to Logan instance theory (Logan, 1988), the retrieval of cached stimulus-response associations produces a stable RT profile, in contrast, trial-by-trial algorithmic computation can result in more variable trial-by-trial RT differences. 

      While we agree that RT differences alone should not be taken as definitive evidence of distinct strategies, our specific experimental design, the longer and more variable RTs at the Critical target, the greater trial-by-trial variability at that same target, and, if confirmed, a parametric relationship between RT and reach angle provide converging evidence that participants in the two conditions relied on different strategy implementations when preparing movements to the Critical target.

      Second, the interpretation of Experiment 3 is not entirely clear to me. The manuscript argues that the algorithmic group continued to exhibit greater reaching variability than the retrieval group. If this variability indeed reflects greater variability in movement plans, one might expect a broader implicit generalization function in the algorithmic group. However, the generalization widths were comparable between groups. Could this result instead suggest that the implicit recalibration process itself generalized more narrowly in the algorithmic group, thereby offsetting the broader distribution of movement plans? More generally, I would appreciate further clarification regarding the relationship between reaching variability, movement-plan variability, and the resulting width of the implicit generalization function.

      We appreciate the reviewer’s thoughtful comment and agree that this is an important distinction. First, we would like to clarify the relationship among reaching variability, movement-plan variability, and the width of the implicit generalization function. Previous work has shown that when implicit recalibration occurs at a particular target location without an explicit aiming strategy, its generalization across the workspace can be described by a Gaussian-shaped function centered near the trained target location (Morehead et al., 2017). When an explicit aiming strategy is involved, however, implicit recalibration is centered closer to the planned aiming location rather than the visual target itself (McDougle et al., 2017). Thus, implicit recalibration is greatest near the direction in which the movement is planned, and a broader spatial distribution of movement plans can, in principle, produce a broader aggregate implicit generalization function. In the present study, we therefore use trial-to-trial variability in endpoint hand angle as a behavioral proxy for variability in planned movement direction, while recognizing that endpoint variability may also contain contributions from execution-related noise.

      This framework motivated the progression from Experiments 1 to 3. In Experiment 1, participants in the algorithmic condition exhibited substantially greater reaching variability, consistent with the idea that they sampled a wider range of movement plans across trials. Because error feedback associated with these different movement plans can induce implicit recalibration around each planned direction, greater variability in strategy use could contribute to the broader implicit generalization observed in the algorithmic group. In Experiment 2, we imposed stricter controls on spillover from noncritical targets, which reduced the overall breadth of generalization; nevertheless, model fits still suggested a modestly broader implicit generalization function in the algorithmic group, consistent with the remaining difference in reaching variability.

      Experiment 3 was designed to further reduce the direct influence of strategic variability on the induction of implicit recalibration by using a modified error-clamp paradigm. Error-clamp feedback has been shown to elicit implicit recalibration independently of task success and the participant’s explicit strategy (Morehead et al., 2017). We therefore used error-clamp feedback as an incidental signal to induce implicit recalibration while participants implemented either algorithmic or retrieval-based strategies. Importantly, however, Experiment 3 did not completely eliminate between-group differences in reaching variability: the algorithmic group continued to show greater variability around the critical 45-degree location than the retrieval group. We agree with the reviewer that, in principle, comparable generalization widths could arise if this broader distribution of movement plans were offset by a narrower local generalization of implicit recalibration in the algorithmic group. 

      However, not all variability is equivalent—or well described by a Gaussian distribution. Depending on the direction of the skew, variability in aiming can produce different effects on the implicit recalibration function, appearing as either broader generalization or greater amplitude. These effects are difficult to appreciate in Figures 3 and 5 for Experiments 2 and 3, respectively. We therefore sought to illustrate this more clearly in Figure 6, which shows how differences in the underlying aim distributions can shape the resulting implicit recalibration function in directionally complex ways. What complicates matters further is that implicit recalibration can asymptote (Morehead et al., 2017; Kim et al., 2018; Wilterson & Taylor, 2021). As a result, plan-based generalization can distort the implicit recalibration function in different ways depending on which side of the aim the error falls. The block-by-block analysis, suggested by reviewer 2, may shed light on this issue because we can get a sense if implicit recalibration has reached asymptote. 

      At a minimum, in the revised manuscript, we work to make clearer that subtle changes in the reach distribution may have a corresponding impact on the shape of implicit recalibration’s generalization function.  

      Reviewer #2 (Public review):

      This study addresses an important question in motor learning: whether algorithmic versus retrieval-based explicit strategies differentially shape implicit recalibration. The progressive experimental logic across three experiments is commendable, and the plan-based generalization account is a plausible and interesting interpretation. However, several methodological concerns limit the strength of the conclusions. I recommend the authors temper their claims accordingly, in the results/discussion section.

      Concerns

      (1) The retrieval group received 5 pre-exposure trials before main training began, which the algorithmic group did not. Faster RTs in the retrieval group could therefore reflect task familiarity from extra practice rather than efficient memory retrieval per se. I might have missed this, but I did not see performance data from these pre-exposure trials. The early training advantage in the retrieval group might be confounded with the 5 pre-exposure trials they received. Unless there is a direct comparison between the pre-exposure trials for the caching group and the first 5 trials of the algorithmic group, the claim that "storing and retrieving a memory from a short-term memory cache confers more rapid performance improvements than executing an algorithmic strategy" seems somewhat unwarranted.

      We appreciate the reviewer raising this potential confound. We agree that the five pre-exposure trials in the Retrieval condition introduce a small difference in initial task familiarity. However, this account makes a straightforward prediction: if the shorter RTs in the Retrieval condition simply reflect five additional trials of general task experience, then the RT difference should disappear once the Algorithmic group has received a comparable amount of practice.

      We can test this directly by comparing the five pre-exposure trials in the Retrieval condition with the first five trials of the Algorithmic condition. We will also compare these pre-exposure trials with a later five-trial window from the Algorithmic condition to determine whether additional practice substantially reduces Algorithmic RTs. Assuming the observed pattern is as expected, RTs in the Algorithmic condition remain substantially longer even after considerably more than five trials of practice. Thus, the group difference cannot be explained simply by the Retrieval group having five additional trials of task familiarity. This persistent RT difference, together with our prior work showing characteristic RT differences between algorithmic computation and retrieval of cached aiming solutions, supports our interpretation that the groups relied on different strategy implementations.

      We are less certain what the reviewer means by the “early training advantage.” If this refers to angular error, we agree that the Retrieval group shows somewhat better performance very early in training, but this difference is not a central focus of the present study and largely disappears by the second or third training block. We will clarify the text so that we do not overinterpret this transient difference.

      If instead the reviewer is referring to RT, then the matched-trial analysis directly addresses the concern. Five additional familiarization trials could plausibly produce a brief initial advantage, but such an effect should dissipate within a small number of subsequent trials. In contrast, the RT difference between the Algorithmic and Retrieval conditions remains robust throughout training. We therefore do not think that general task familiarity provides a sufficient explanation for the observed RT differences.

      The algorithmic group also visited the critical target approximately 40% of trials across 356 trials (about 140 trials?). McDougle & Taylor (2019) showed that 300 trials of practice with 2 targets is enough transition from algorithmic to caching strategies. It seems likely that the number of visits to the critical target here was sufficient for caching to develop in the algorithmic condition. This concern about caching in the algorithmic group has implications for the implicit recalibration measurements. As I understand it, the 7 exclusion blocks were distributed throughout training, and so, implicit recalibration was measured across both early and late practice. If caching emerged in the algorithmic group during late practice, then the generalization functions - averaged across all 7 exclusion blocks - conflate early algorithmic strategy and later caching. The broader generalization function observed in the algorithmic group may therefore be driven primarily by early exclusion blocks, while later exclusion blocks may increasingly resemble the retrieval group as caching develops. This is testable in the data: if generalization breadth in the algorithmic group narrows across the 7 exclusion blocks while remaining stable in the retrieval group, that would be consistent with a strategy transition occurring during training. The authors should either report exclusion block-by-block generalization functions separately for each group, or acknowledge that the averaged generalization functions may obscure a strategy transition in the algorithmic group.

      The reviewer raises an interesting possibility. In McDougle and Taylor (2019), however, the transition from algorithmic computation to retrieval occurred in a condition with only two targets in the task set. With repeated practice, participants needed to retain only two target-specific aiming solutions, making it feasible to replace online computation with retrieval of cached stimulus–response associations. By contrast, the Algorithmic condition in the present study contained 10 target locations. Our prior work suggests that caching becomes increasingly difficult once the number of target-specific solutions exceeds approximately four, at least over the timescale of several hundred trials (Velázquez-Vargas and Taylor, 2024; Bejjanki and Taylor, 2026). Thus, our task was designed to maintain pressure toward an algorithmic strategy throughout training.

      The reviewer nevertheless raises a more specific possibility that is not ruled out simply by the size of the target set: participants might selectively cache the aiming solution for the frequently sampled Critical target while continuing to use an algorithmic strategy at the remaining targets. We think the existing behavioral data argue against such a clear transition.

      First, reaction times at the Critical target in the Algorithmic condition remained substantially longer than those in the Retrieval condition throughout training. If participants had cached the aiming solution for the Critical target, we would expect preparation times at that location to be the same as the Retrieval condition. However, the RTs for the Algorithmic and Retrieval conditions are significantly different in the last block of training. 

      Second, within the Algorithmic condition, reaction times at the Critical target remained similar to those at the Non-Critical targets. Selective caching of the Critical target predicts a different pattern: preparation should become faster at the Critical target than at the surrounding locations, where participants would still need to compute the appropriate aiming solution. However, we do not observe a significant difference between RTs at Critical and Non-critical targets for the Algorithmic conditions at the end of the training block. Taken together, these two observations suggest that the Critical target continued to be treated similarly to the other members of the 10-target set rather than becoming a privileged, cached stimulus–response association.

      Third, participants would have to single out the Critical target as being distinct. All targets had the same visual appearance, the Critical target was never presented on consecutive trials, and participants were not informed that it had a special role in the experiment. However, we acknowledge that its higher sampling frequency, its somewhat greater separation from neighboring targets, and the location of the subsequent exclusion trials could nevertheless have made it more salient. Thus, we cannot rule out selective caching solely from the task structure.

      For this reason, we agree that the reviewer’s proposed analysis provides a useful additional test. If the Critical-target strategy progressively transitioned from algorithmic computation to retrieval, one prediction is that the generalization function in the Algorithmic condition should become narrower across successive exclusion blocks and increasingly resemble that of the Retrieval condition. We will therefore attempt to estimate the width of the generalization functions as a function of the training block between the Algorithmic and Retrieval conditions.

      There is, however, an important limitation to interpreting block-by-block generalization functions in this experiment. Implicit recalibration is both plan-based and temporally labile. Generalization is centered around the planned aiming direction (McDougle et al., 2017), and recent work indicates that implicit adaptation can decay over relatively short intervals (Zhou et al 2017; Hadjiosif et al 2023). Consequently, the first trial of an exclusion block provides the cleanest sample of the current state of implicit recalibration. Across later trials in the block, the measured response can be influenced both by temporal decay and by where the sampled target falls relative to the participant’s current aiming direction.

      To minimize systematic sampling bias, the starting exclusion target was randomized across participants. This means that these effects should average out at the group level, but individual exclusion blocks do not provide equally precise samples of the entire generalization function. A fully balanced estimate of every position within each exclusion block would require substantially more participants than were included in the present experiments. We will therefore present the blockwise analysis while interpreting changes in the estimated breadth cautiously.

      (2) The error-clamp paradigm in Experiment 3 introduces two problems. First, it breaks the relationship between planned movement direction and feedback of movement direction, likely reducing the sense of agency over movement feedback (indeed, typical error clamp study instructions tell participants to ignore the movement feedback). Reduced agency may itself suppress differences between algorithmic and caching conditions. First, if strategy type exerts its influence on implicit recalibration via the explicit plan - as the plan-based generalization account predicts - then severing the link between intended movement and feedback might close off the channel through which strategy could shape the implicit system, regardless of which strategy is used. Second, reduced agency could modify the explicit strategies themselves. For caching, the stimulus-response association might be reinforced by a consistent relationship between intended movement and observed outcome; the clamped feedback may make it more difficult to reinforce the cached response, weakening the stimulus-response association. For the algorithmic strategy, effortful mental rotation may depend on the perception that the computation meaningfully determines the outcome; as participants understand that clamped feedback does not depend on their behavior (although yes, the text-based "Excellent/Good Move feedback) does depend on their behavior, they may engage in somewhat less complete mental rotation. Both possibilities could contribute to convergence between groups in generalization. It is noted that the preserved RT difference between groups in Experiment 3 partially argues against a loss of effort under the algorithmic condition, but it does not rule out weakened formation of stimulation-response associations during caching.

      The reviewer raises an important point. By design, the error-clamp manipulation in Experiment 3 decouples the participant’s planned movement from the visual consequence of that movement. While this gives us precise control over the error driving implicit recalibration, it could reduce agency over the cursor and thereby alter the interaction between explicit strategy and implicit learning in ways that are difficult to rule out completely. In particular, as the reviewer notes, reduced agency could potentially weaken either the influence of the explicit plan on implicit recalibration or the strategies themselves. Because Experiment 3 was intended to test for the absence of a strategy-dependent difference in implicit recalibration, we acknowledge that higher-order interactions of this kind represent an inherent limitation of our study if Experiment 3 is taken in isolation. 

      There are nevertheless several observations that make us think that reduced agency is unlikely to provide the primary explanation for the convergence between groups. First, the progression across Experiments 1–3 is informative. In Experiment 1, where participants retained normal control over the cursor, the broader generalization function in the Algorithmic condition closely mirrored the broader distribution of reach directions. This relationship suggests that the apparent difference in implicit generalization could arise from differences in where participants planned their movements rather than from a direct effect of strategy type on the implicit system. In Experiment 2, we sought to reduce the difference in the distribution of planned movements while preserving normal action–outcome contingencies and, importantly, the generalization functions became correspondingly more similar. Experiment 3 then controlled the error signal itself and again produced similar generalization across strategy conditions. Taken together, this progression favors the interpretation that strategy affects the measured generalization function indirectly, through differences in the distribution of movement plans, rather than directly altering the underlying implicit recalibration process.

      We nevertheless agree that these experiments cannot exclude all possible interactions between explicit and implicit learning systems. Indeed, whether these systems interact directly has been an important and persistent question in the sensorimotor adaptation literature. Several studies have reported evidence consistent with direct interactions (e.g., Albert et al., 2022; Maresch and Donchin 2021; t’Hart and Henriques 2024), whereas our own work has generally pointed toward indirect interactions mediated by factors such as movement planning and the current state of implicit adaptation (Taylor et al 2010; Taylor and Ivry 2011; McDougle et al 2017). Indeed, our recent study was designed specifically to distinguish these possibilities under tighter experimental control (Chen and Taylor, 2026), yet we found that the interaction between explicit and implicit processes is more complex than a simple independent-versus-interacting dichotomy. Going forward, we think it is more cautious to first rule out low-level statistical or distributional differences that could account for apparent effects before invoking higher-order interactions between learning systems.

      Finally, one motivation for the present study was that much of the literature on implicit generalization trains participants at a single target location before measuring generalization across the workspace. Under such conditions, participants have ample opportunity to retrieve a stable target-specific aiming solution. If algorithmic and retrieval strategies fundamentally alter implicit generalization, then many existing estimates of generalization may characterize implicit learning under retrieval-like conditions rather than providing a strategy-independent property of the implicit system. Across the present experiments, we find little evidence for such a fundamental difference once the distribution of movement plans and the experienced error are better controlled. We therefore think the most parsimonious interpretation of the current results is that algorithmic and retrieval strategies primarily influence implicit generalization indirectly through how movements are planned. 

      We plan to revise the manuscript to acknowledge that Experiment 3 cannot completely rule out higher-order effects associated with reduced agency under error-clamp feedback. We will also provide additional validation that participants were implementing distinct strategies, beyond the group-level RT differences, by testing whether RT in the Algorithmic condition scales with the instructed rotation magnitude, and whether RT variability shows group-level difference (Logan, 1988). If present, this relationship would provide stronger evidence that participants continued to engage the intended strategy under the clamp. We agree, however, that confirming distinct strategy use would not by itself rule out the possibility that reduced agency altered how those strategies interacted with implicit recalibration.

      Reviewer #3 (Public review):

      Summary:

      This manuscript asks whether two forms of explicit strategy use in visuomotor adaptation, i.e., algorithmic mental rotation and retrieval of a cached aiming solution, differentially influence implicit recalibration. The question is relevant because much prior work treats explicit strategy as a unitary process, whereas the algorithmic/retrieval distinction is theoretically meaningful and grounded in cognitive theory. Across three experiments, the authors report that algorithmic strategy conditions initially produced broader fitted implicit generalization functions than retrieval conditions, but that this difference was reduced or eliminated when reach variability and sensory prediction errors were more tightly controlled.

      Strengths:

      The paper is clearly written, theoretically well-motivated, and employs a commendably transparent and progressive experimental logic. The three-experiment structure, in which confounds are systematically identified and addressed, represents a strong model of cumulative experimental design (I will certainly use it in teaching courses on experimental methods):

      Experiment 1 establishes an apparent difference in implicit generalization breadth. Experiment 2 attempts to reduce error spillover from Non-Critical targets by increasing angular separation and using delayed endpoint feedback. Experiment 3 uses an error-clamp design to decouple variable reaching from error feedback. This sequence is appropriate for testing whether the initial difference reflects a strategy-dependent change in implicit recalibration or instead follows from the distribution of movement plans and error exposure. The authors also provide reaction-time and performance data that are broadly consistent with the intended distinction between algorithmic and retrieval-like task performance.

      We thank the reviewer for this thoughtful and constructive assessment of the manuscript. We especially appreciate their recognition of the progressive experimental logic across the three experiments and of the broader theoretical motivation for distinguishing algorithmic and retrieval-based strategies. We are also grateful for the reviewer’s comments on the clarity and transparency of the work.

      Weaknesses:

      The evidence does not support the strongest claims made in the manuscript, namely that algorithmic and retrieval strategies generally do not reshape implicit recalibration.

      In general, I am skeptical of the authors' interpretation of null results. Several central conclusions depend on non-significant group differences, especially in Experiment 3. Non-significant tests are repeatedly treated as evidence that groups are equivalent or that confounds are absent (e.g., implicit recalibration magnitude (Algorithmic: 11.43 {plus minus} 6.43{degree sign}; Retrieval: 15.49 {plus minus} 8.99{degree sign}; t(38) = −1.65, p = .11), adaptation level before Exclusion probes (F(1,256) = 3.04, p = .08) and Exclusion RT differences (F(1,266) = 3.15, p = .08), whereas a modest model-dependent breadth effect (bootstrap p = .02) is treated as meaningful (for more on the model-dependent breadth effect, see below).

      Without confidence intervals, equivalence tests, or Bayesian analyses, I think that the authors' interpretations comprise an inferential gap. A failure to find a significant difference is not equivalent to evidence of equivalence, particularly given that the implicit recalibration signal gets progressively attenuated across experiments (Experiment 1: ~16-17{degree sign}; Experiment 2: ~11-15{degree sign}; Experiment 3: ~7-8{degree sign}). With a substantially diminished signal in Experiment 3, the null result could partly reflect reduced statistical sensitivity rather than true equivalence.

      We agree with the reviewer that our original interpretation of several non-significant effects was too strong, especially without providing some form of equivalence test. Our central hypothesis predicts little or no difference between algorithmic and retrieval-based strategies under conditions in which movement plans and error exposure are controlled, and we therefore face the inherent difficulty of drawing conclusions from an expected null effect. As the reviewer notes, a non-significant conventional hypothesis test does not by itself provide evidence that two conditions are equivalent.

      We therefore plan to supplement the existing analyses with quantitative assessments of the strength of evidence for the null/equivalence, using Bayesian factor analyses to confirm whether two conditions are equivalent. These analyses will allow us to distinguish between effects that are sufficiently small to support our theoretical interpretation and effects for which the data are simply inconclusive. We will also revise the manuscript throughout to avoid treating p > .05 as evidence of equivalence in the absence of such supporting analyses.

      We also now appreciate that the magnitude of implicit recalibration decreases progressively across experiments. This reduction could diminish our sensitivity to differences between the Algorithmic and Retrieval conditions and therefore represents an important qualification on the null result. Because the experiments used similar trial structures and were conducted with the same experimental equipment, the source of this reduction is not immediately clear. The block-by-block analysis suggested by Reviewer 2 may provide useful insight into how implicit recalibration evolves over the course of training and whether this attenuation emerges gradually within experiments.

      My main technical concern is the analysis of generalization breadth already alluded to. The central claims rely on group-level Gaussian fits to only seven Exclusion probe locations spanning −45{degree sign} to +45{degree sign} around the Critical target. In several cases, the fitted centers and widths are poorly constrained by the sampled range. For example, in Experiment 2 the algorithmic group's fitted center is shifted to approximately 29{degree sign}, meaning that the probe range samples the function asymmetrically relative to its own peak. In Experiment 3, fitted centers are near or outside the sampled range, while estimated widths are very broad. Under these conditions, the width parameter may partly reflect extrapolation or parameter trade-offs between center, amplitude, and width rather than a genuine difference in generalization breadth.

      Based on prior work characterizing implicit generalization in relative isolation from explicit strategy (Morehead et al., 2017; Poh and Taylor, 2019), we expected a relatively narrow generalization function, with a full width at half maximum of approximately 30°. We therefore expected probes spanning −45° to +45° around the Critical target to capture most of the function. At the same time, prior work on plan-based generalization predicts that the function should shift toward the participant’s aiming direction (Day et al., 2016; McDougle et al., 2017; Chen and Taylor, 2026), which complicates the choice of where to center the probes. Expanding the range and density of probe locations is also not cost-free, because additional exclusion trials increase temporal decay (Hajiosif et al., 2023) and begin to overlap with trained locations.

      We nevertheless agree with the reviewer that, when the fitted center approaches the edge of the sampled range, estimates of Gaussian width can become poorly constrained and may partly reflect parameter trade-offs or extrapolation beyond the observed data. We therefore plan to test whether the group differences persist when the Gaussian fits are constrained so that their centers fall within the sampled range. We will also examine complementary nonparametric measures of generalization breadth, such as the area between the group generalization curves across the sampled probe locations. Convergence across these approaches would provide stronger evidence that the reported differences reflect the observed shape of the generalization functions rather than instability in the Gaussian parameter estimates. If the results are not robust across approaches, we will revise the manuscript to qualify the interpretation of the fitted width estimates accordingly.

      Lastly, I think that the authors' use of an error-clamp paradigm is, from an experimental point of view, quite elegant. By controlling the sensory prediction error independently of reach direction, they can isolate implicit recalibration from the confounds identified in Experiments 1 and 2. However, I see a fundamental problem or question concerning construct validity here: In Experiments 1 and 2, the algorithmic strategy was operationalized as participants computing a counterrotated aiming direction in response to a visible cursor rotation. This is a naturalistic context where mental rotation is both required and meaningfully connected to task success. In Experiment 3, however, there is no visuomotor rotation to compensate for. The error-clamp renders the cursor feedback task-irrelevant. Instead, participants are instructed via text commands (e.g., "move towards 45{degree sign}") to reach invisible locations, rendering the "algorithmic strategy" in this context essentially an instructed spatial navigation toward arbitrary angular locations, not genuine visuomotor mental rotation driven by an error signal.

      To put it differently, are we sure that the cognitive process engaged by the algorithmic group in Experiment 3 is the same as the algorithmic mental rotation strategy in Experiments 1 and 2? If not, then the null result in Experiment 3 may not speak to the original question about how algorithmic strategies interact with implicit recalibration after all. Instead, it may reflect the absence of a genuine strategy manipulation.

      This concern is closely related to that raised by Reviewer 2. We are fairly confident that participants in Experiment 3 were nevertheless engaging in the intended strategy manipulation. Participants in the Algorithmic condition showed substantially longer RTs than those in the Retrieval condition, and they were able to accurately generate the instructed angular reach directions across trials. The two conditions also differed in the variability of both RT and reach direction, consistent with online computation of an aiming solution in the Algorithmic condition and retrieval of a more stable cached response in the Retrieval condition. We can provide an additional validation by examining the relationship between RT and reach angle and comparing RT variability across two conditions. If RT scales parametrically with instructed reach angle in the Algorithmic condition but not in the Retrieval condition, this would provide stronger evidence that participants were engaging an online mental-rotation-like computation rather than simply following arbitrary spatial instructions. Moreover, RT yielded by the retrieval strategy would tend to be less variable than the algorithmic strategy.

      We agree, however, that this does not fully address the reviewer’s broader concern. Experiment 3 necessarily changed the context in which the strategy was implemented. In Experiments 1 and 2, mental rotation was used to counteract a visuomotor perturbation and was therefore directly tied to successful control of the cursor. In Experiment 3, the error clamp removed this instrumental relationship: participants still had to compute and execute different angular reach directions, but those computations no longer determined the visual cursor outcome. In that sense, the algorithmic process in Experiment 3 was less naturally embedded in the task and could reasonably be viewed as a somewhat different instantiation of the task.

      We therefore acknowledge that Experiment 3 cannot establish with certainty that the same higher-order cognitive process was engaged in exactly the same way as in Experiments 1 and 2, nor can it rule out the possibility that this change in task relevance altered how explicit strategy interacted with implicit recalibration. At the same time, when considered together with Experiments 1 and 2, we think the overall pattern remains informative. The apparent strategy-dependent difference in generalization was largest when movement plans and error exposure differed most, became smaller when these factors were better controlled while normal action–outcome contingencies were preserved, and was eliminated when the error signal itself was experimentally controlled. This progression is more consistent with an indirect influence of strategy through differences in movement planning and error exposure than with a robust direct effect of strategy type on implicit recalibration.

      Nonetheless, we agree that Experiment 3 should not be interpreted as a definitive test of whether algorithmic strategy, in its more relevant visuomotor adaptation context, can lead to different interactions with implicit recalibration compared to a retrieval strategy. We plan to revise the manuscript to make this limitation explicit.  

      To their credit, the authors report a compelling RT dissociation that mirrors Experiments 1 and 2: The algorithmic group shows slower RT, which is decreasing over training (0.98s → 0.76s), whereas the retrieval group exhibits faster, stable RT (0.52s → 0.45s). While this pattern is consistent with genuine strategy differences persisting in Experiment 3, it could also reflect the greater spatial precision demands of reaching to invisible targets from text instructions, rather than genuine mental rotation per se. Reaching to an invisible location defined by a verbal angular label is inherently more demanding than reaching to a visible target, regardless of strategy type, and this demand is asymmetrically present in the two groups, since Non-Critical targets are invisible for the algorithmic group but visible for the retrieval group.

      Thus, from my point of view, experiment 3 should not be used as definitive evidence that algorithmic and retrieval strategies during standard visuomotor adaptation cannot differentially influence implicit recalibration.

      We agree that the RT difference in Experiment 3, by itself, cannot rule out the possibility that the Algorithmic condition imposed greater spatial precision demands because participants were reaching to invisible locations specified by angular instructions. We can, however, test for a more diagnostic signature of algorithmic computation by examining whether RT scales parametrically with the instructed reach angle. A general cost associated with reaching to invisible targets could increase overall RT, but it would not necessarily predict the characteristic increase in preparation time with the magnitude of the required angular transformation.

      We will therefore examine the relationship between RT and instructed reach angle in Experiment 3. If RT increases systematically with angular displacement in the Algorithmic condition, this would provide additional evidence that the longer RTs reflect online computation of the instructed aiming direction rather than simply the greater difficulty of reaching invisible targets.

      We can also compare this RT–angle relationship across experiments. If participants are engaging the same underlying algorithmic computation in Experiments 1–3, we would expect the slope relating RT to angular displacement to be similar across experiments, even if the overall intercept differs because of differences in task structure and spatial demands. A comparable slope would therefore provide converging evidence that the same computational process was engaged despite the altered task context in Experiment 3.

      We acknowledge, however, that similarity of the slopes would itself require yet another inference from a null difference and should therefore be interpreted cautiously. As with the generalization functions, we will use a Bayes factor analysis to quantify the strength of evidence for the null.

      Overall, the manuscript addresses a meaningful question and the multi-experiment structure is useful. The evidence is incomplete for the broad claim that implicit recalibration is insensitive to strategy type. The study would make a clearer contribution if the authors narrowed the claims, strengthened the generalization analyses, and treated null effects with appropriate inferential tools.

      Based on the reviewers’ comments and the additional analyses they have suggested, we think we will be able to place our conclusions on a firmer empirical footing while also tightening and narrowing them. In the revised manuscript, we will strengthen the generalization analyses, use more appropriate inferential tools for interpreting null effects, and temper our broader claims about the insensitivity of implicit recalibration to strategy type. We will also more explicitly acknowledge the limitations of the present experiments, especially Experiment 3.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary and Strengths:

      Shin et al deepen our understanding of high-frequency oscillations in the frontal cortex during REM in a manner that sheds important light on the roles of these events. In particular, they reveal that cortical HFOs are modulated by theta oscillations, occur in chains and recruit cortical neuronal activation patterns in a manner that is distinct from other high-frequency events during non-REM or in the hippocampus. They also show that these events occur during increased oscillatory cross-talk between hippocampus and cortex and may protect cortical neurons from downregulation of firing during sleep. Overall, this is important work with several novel observations pointing towards an important role for these events that will become increasingly understood over time.

      I also wanted to comment that 2D is a beautiful illustration of separate and essentially exclusive communication channels used during HF events in NREM vs REM. They almost perfectly complement each other's frequencies.

      Weaknesses:

      I have only one major scientific critique: I believe we need to see quantification of how phasic REM theta waves with versus without HFOs differ. What do REM HFOs add to the "normal" theta oscillation? Without this comparison, it is more difficult to interpret the meaning of these events. Given that HFO chains have IEIs around the time of a theta cycle duration, are the repeating spiking activities stronger during HFO repeats than during adjacent theta waves without HFOs?

      Here, we provide additional analyses to demonstrate that the phasic, theta-modulated PFC activity that we observe during HFOs is specifically tied to the occurrence of HFOs and not a strong phenomenon during non-HFO-associated theta periods. In Figure S5 M (middle and right), we find that aligning PFC multiunit activity to theta periods in phasic REM but temporally distant from HFOs does not elicit the same degree of theta-modulated activity as aligning to HFOs (as in Figure 2A and Figure S5 M, left).

      Additionally, we provide analyses of the theta periods adjacent to HFOs (at different temporal thresholds) and demonstrate that this theta-modulated spiking activity is largely absent (Figure S5; compare to Figure 2A). Unlike Figure S5, this analysis was not restricted to putative phasic REM.

      We have now added Supplementary Figures S5L-N to the revised manuscript.

      What percentage of theta waves contain HFOs, and what is the firing rate during those theta waves with vs without HFOs? Is there differential firing rate modulation? The authors may even consider that all REM-HFO-specific quantifications should be shown as differential from phasic theta cycles without HFOs.

      Although theta oscillations are continuously expressed during REM sleep, HFOs occur only intermittently, such that only a small subset of theta cycles contain HFOs. Across all animals and epochs included for analysis, we found that ~7.4% of theta cycles contain HFOs. We present an epoch-level quantification of this in Author response image 1, where the proportions were calculated across all, tonic, and phasic theta cycles. As expected, a higher proportion of putative phasic theta cycles contain HFOs.

      Regarding differential firing rate modulation, we refer reviewer to normalized MUA plots in the manuscript (Figures 2A and 4F). We would like the emphasize that what we show is PFC multiunit activity that is normalized by the mean population firing rate during REM sleep. Thus, these figures, specifically the chain HFO aligned figure, indicate that there are peaks in activity around baseline level in the background of an overall decrease in population activity relative to baseline (i.e. the troughs surrounding the peaks have lower activity compared to baseline, as in Figure 9C). While this suggests that there is an overall decrease in firing rates during HFOs as compared to baseline theta periods without HFOs, this simply provides a qualitative account of this difference. In Figure S5N, we present firing rate comparisons during HFO chains versus theta periods during putative phasic REM bouts at least 4 theta cycles away from HFOs. We find that HFO chain-associated PFC neuron firing rates are lower compared to non-HFO theta periods, supporting our finding of activity suppression during HFOs.

      Author response image 1.

      Proportion of theta cycles with HFOs. (A) Proportion of theta cycles with HFOs across all, tonic, and phasic cycles (***p = 4.90e-05, rank sum test).

      Lastly, we appreciate the reviewer's suggestion that REM HFO quantifications could be framed relative to phasic theta cycles without HFOs. We agree that such comparisons are informative and ensure that the results we present are specific to periods with detected HFOs. In response, we have added additional analyses of theta periods outside of HFOs in Figure S5. Furthermore, while we found that a larger proportion of chain HFOs occurred during bouts of putative phasic REM compared to isolated events (Figure S5 and Figure S3C), most of the analyses that we performed were on events pooled across putative tonic and phasic, since putative phasic REM is relatively scarce (<10%).

      We also refer the Reviewer to our response to Reviewer 2’s major comment #1 below where we reiterate several of our findings that demonstrate the HFO-specificity of the reported dynamics, as well as our extended response to Reviewer 3’s Public Review comment #1 where we show that the dynamics associated with REM HFOs are absent during HFOs detected during awake behavior (comparable theta state) on the W-Track (Figure S11). We hope that the additional control analyses we present as well as our expanded explanations adequately address the Reviewer’s concerns.

      As a non-scientific comment on the manuscript itself: unfortunately, the paper is difficult to read and understand at times, requiring great effort by the reader. This is to an extent that communication is hindered. The paper is dense with changing methods, often from panel to panel. Unfortunately, the panel quantifications are not explained in the results section in a manner that readers can understand without going to read the methods, often for each individual panel. These measures should be explained in a way that lets readers understand the conclusions of each panel and what gross calculations were used to reach those. Instead, too much jargon is used rather than clear descriptions of the overall calculations being done for each panel.

      We have now split and updated the figures in a more logical progression of ideas:

      Figure 1: Prefrontal cortical HFOs in REM sleep using spectral analyses.

      Figure 2: Characteristic spiking modulation in PFC during REM HFOs.

      Figure 3: HFO and gamma distinction in theta cycles, and PFC-CA1 coherence (including chain/ isolated HFOs in phasic and tonic REM stages).

      Figure 4: Differential modulation of PFC spiking activity during REM PFC HFOs vs. NREM PFC ripples.

      Figure 5: Characteristic PFC population activity during REM HFO chains.

      Figure 6: Comparison of PFC reactivation during REM PFC HFOs vs. NREM PFC ripples.

      Figure 7: Differential engagement and excitability modulation of CA1 neurons by REM HFOs.

      Figure 8: REM theta phase shifting CA1 neurons preferentially engaged by REM HFOs.

      Figure 9: (Model) Network model with ACh reproduces spiking modulation during REM PFC HFOs vs NREM PFC ripples.

      Figure 10: (Model) Model reproduces restricted REM coactivity vs. widespread NREM coactivity.

      We have also rearranged the figures to parallel the main figures and results section.

      The authors mention in the discussion section that they see increased functional connectivity between mPFC and CA1, but most data suggesting this seems to be based on LFP rather than spiking. Functional connectivity is best defined by spiking-spiking relationships. And these authors have spiking data. So I believe either the descriptive language should be pulled back to something like "oscillatory coupling" or more analyses should be dedicated to showing spike-spike coordination across regions.

      We have updated the manuscript accordingly. Specifically, we have modified the text on Page 13, Lines 23-24:

      “These chains are associated with increased measures of oscillatory coupling between PFC and CA1…”

      Reviewer #1 (Recommendations for the authors):

      (1) Please ensure that analytical methods are presented in the same order in the methods section as they are in the results section - panel by panel. That said, the methods section is well-written and presented.

      We apologize for any confusion this may have caused. We have now reorganized the methods section to ensure they presented in the same order as the panels in the main figures.

      (2) Please specify whether the recordings/behaviors occur during the animal's light circadian phase.

      We have added a statement in the methods under the “Behavior” section on Page 34, Lines 9-12 indicating that the experiments took place during the light phase:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Please specify how many tetrodes are in mPFC and how many in CA1?

      We have added this clarification on Page 33, Lines 27-28:

      “Tetrodes were split equally between PFC and CA1 (15, 16, or 32 tetrodes in each region).”

      (4) All mentions of "coherence" should have frequency bands specified. "Theta coherence", for example.

      We have now added this information to all relevant mentions of “coherence”.

      (5) The intuitive logic of the phase slope index (2K) should be briefly explained for maybe half a sentence in the results section. The intuition should be explained better in the methods section devoted to it.

      We have added more clarification on the phase slope index method, both in the legend of Figure 3 on Page 21, Lines 22-23:

      “Phase slope index (PSI), which is a measure of phase lag consistency across different frequencies…” and in the methods section under “Phase slope index” on Page 39, Lines 21-26:

      “In practice, PSI is used to assess the consistency of phase lag relationships between two signals across different frequency bands and is a measure that is weighted by oscillatory coherence. We opted to use PSI to estimate the directional flow of information instead of other methods, such as Granger causality, since it has been demonstrated that PSI is less prone to false positives.”

      (6) "Cofiring" and "coactivity" should be defined clearly as measures - preferably in the results section if possible. They sound similar and are somewhat jargon-y without self-explanatory meaning (or difference from each other). How should readers understand and interpret them?

      We apologize for the confusion regarding these two terms, which are both used throughout the manuscript. Here, we specifically used to term “cofiring” to specify the explicit quantification of coincident activity between pairs of neurons (e.g. Figure 4E, left) as described in the Methods section under “Ripple and HFO co-firing”. There were instances where “coactivity” was used to refer to this quantification, and they have been changed to “cofiring”. We have now added a statement to the manuscript to clarify that “cofiring” is a quantification of coincident activity between neurons on Page 7, Lines 34-35:

      “Overall cortical co-firing, which is a measure of coincident activity between neuron pairs during discrete events”

      Additionally, the term “coactivity” in the manuscript is used when describing neuronal activity in the model or when there are mentions of coincident activity other than the explicit quantification described above.

      (7) The temporal threshold for "cofiring" should be stated in the results to enable interpretation of the results.

      We apologize for the lack of clarity regarding the cofiring metric. Here, the temporal threshold that we are imposing is determined by the ripple/HFO event times (see Methods section “LFP event detection”). If both neurons in a pair emit spikes within the defined window of an event, they are considered to have “cofired” (Cheng and Frank 2008, Singer and Frank 2009, Sosa, Joo et al. 2020).

      We have now clarified this in the results section on Page 7, Lines 34-35:

      “Overall cortical cofiring, which is a measure of coincident activity between neuron pairs during discreet events…”

      We have also added a clarifying statement in the Methods section under “Ripple and HFO cofiring” on Page 40, Lines 24-25:

      “Here, cofiring assesses coincident activity between neuron pairs within the start and end times of events, and thus, no explicit temporal threshold was implemented.”

      (8) Please clarify more systematically in which region the NREM ripples were detected. The natural assumption is the hippocampus, but at times it is mentioned that they are detected in the cortex. Are they cortical in all analyses? Readers could easily get confused about this and misinterpret "ripples". To clarify further, if these are always cortical events, I suggest renaming "ripples" in this text to "cortical NREM HFOs".

      We apologize for the confusion. For all mentions of “ripples” in the text, we are referring to PFC ripples specifically in NREM sleep. When hippocampal ripples are mentioned, we differentiate them by explicitly using “sharp-wave ripples” or “SWRs”. Our decision to use “ripples” for NREM sleep was based on previous studies that investigated NREM high-frequency events in cortex (Khodagholy, Gelinas et al. 2017, Helfrich, Lendner et al. 2019, Vaz, Inati et al. 2019, Aleman-Zapata, Morris et al. 2022, Ghosh, Yang et al. 2022, Shin and Jadhav 2024). Furthermore, we elected to use the “HFO” nomenclature for high-frequency events in REM sleep, since it has been used in previous studies to describe these events (Tort, Scheffer-Teixeira et al. 2013, Bueno-Junior, Ruckstuhl et al. 2023). Thus, to remain consistent with the literature, we decided to use these terms to describe these sleep-state-specific events throughout the manuscript:

      Hippocampal sharp-wave ripples (SWRs) in NREM sleep

      PFC ripples in NREM sleep

      PFC HFOs in REM sleep

      We have now added the following statement on Page 5, Lines 1-3:

      “However, to avoid ambiguity, and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout.”

      (9) The analysis performed for 4B is not explained clearly. It is somewhat better explained in the methods. I believe the reader should understand that each HFO is treated as an event, and spiking participation per unit was measured, and then the similarity of that pattern was assessed between HFOs. I also find the x-axis being quartiles makes understanding this graph particularly difficult.

      Why not label by raw lag and show a correlation plot rather than break down by quartiles? Alternatively, labeling the millisecond values of these quartiles on the x-axis labels may make comprehension much easier.

      We apologize for the lack of clarity regarding this method. We have added points to clarify the analytical procedure used for Figure 4B (Now Figure 5B) on Page 8, Lines 25-28:

      “We represented each HFO as a binary vector of PFC neurons active during the event, and computed the Pearson correlation between the vectors of every pair of consecutive HFOs” As well as in the Figure 5 legend on Page 24, Lines 7-10:

      “Here, the PFC spiking activity during each HFO was binarized across all neurons and the Pearson correlation coefficient was calculated between adjacent events as a measure of pattern similarity. Then the relationship between pattern similarity and IEI was reported.”

      In addition to the quartile labels, we have now added the average inter-event interval (IEI) of each quartile to the x-axis labels of Figure 5B to improve comprehension as well as a statement in the figure legend on Page 24, Line 11:

      “Below each quartile is the average IEI of that quartile in milliseconds.”

      (10) "Rank order analysis" should be again defined in the results in a manner that the reader can follow the point of the analysis and figure. For example, the goal is to assess the spike sequence across HFO events by looking at the regularity of spike timing rank for each neuron in each HFO event. Also, could this correlation be more simply calculated and presented as just a standard deviation around the mean of that cell's rank?

      We apologize for not including an explanation of this analysis in the results section. We have now added more detail about the rank order analysis to the Figure 5 legend on Page 24, Lines 18-21:

      “The mean rank-order correlation from the leave-one-out cross-validation procedure. Each event’s rank was correlated with the averaged rank across all other events. The average across all events compared to a distribution of means generated by jittering (n = 1000) spike times is shown.”

      We also provided a short description of the procedure in the results section on Page 8, Lines 32-36:

      “Furthermore, we examined whether the order in which individual PFC neurons fired during HFOs within a chain was preserved across chains. For each chain, we extracted each cell's first-spike rank order and compared it to a leave-one-out template constructed from the average normalized rank across all other chains.”

      Regarding the Reviewer’s second comment about the correlation, if we presented the result as the standard deviation around the mean of the cell’s rank, the result would be similar to the example rank order in Figure 5C, which we provided as a visualization of the rank-order template procedure we used. However, this would not necessarily demonstrate the consistency of population activity across HFO chains. To demonstrate consistency in activity across HFO chains, we used a leave-one-out cross-validated approach where each chain event was assessed separately. For each event, the neurons firing during that event was ranked and normalized 0-1. Then, the average rank of each neuron across all other events was calculated. Lastly, the correlation between the ranks of the left-out event and the average ranks of the template was taken to assess similarity in sequential activity. We opted to use this method since it has been demonstrated to be effective for evaluating similarities in sequential across events (Stark, Roux et al. 2015, Valero, Viney et al. 2021).

      (11) In Figure 5B and the related results section, I gather that "spatial" relates to place field location rather than anatomical/tetrode location of the neuron? This was not my original understanding and should be stated clearly.

      We apologize for the lack of clarity regarding this result. Yes, the term “spatial” refers to spatial rate map correlation between CA1 neuron pairs, specifically during the behavioral W-Track session prior to the sleep session where cofiring was assessed. We have updated the y-axis label of Figure 5B (Now Figure 7B) to include “rate map” and have updated the results section on Page 9, Lines 32-33:

      “…we observed a higher degree of spatial rate map correlation for high cofiring pairs…”

      (12) Figure 5E and its caption are almost totally unable to be explained since axes aren't explained well in either. It is only by inference from the results text that meaning can be assumed.

      We apologize for the lack of clarity regarding Figure 5E (Now Figure 7E). We have now updated the y-axis labels on both updated figures (Figures 7E and 8B) and have reworded and added more detail to the figure legend on Page 28, Lines 1-5:

      “High REM PFC HFO cofiring CA1 neurons exhibited a greater degree of suppression during NREM PFC ripples. For a description of modulation index, see Methods section Ripple/HFO aligned modulation. Here, since CA1 neurons exhibit a robust decrease in activity in response to NREM PFC ripples, we refer to the modulation as suppressive.”

      (13) Figure 5F is interesting, supporting the concept that HFOs "protect" neurons from downscaling. However, how can a "neuron" be cofiring? Would cofiring not be defined in a pairwise manner, and so each unit of the cofiring measure would be a pair of neurons? This question applies to other panels in this figure. Please clarify this.

      We apologize for the lack of clarity regarding the exact cofiring metric that we use. For Figures 7D-F and 8A, since we wanted to relate cofiring to changes in firing rate and modulation state during NREM PFC ripples, we calculated single cofiring values for each CA1 neurons by averaging across all pairings with PFC neurons. Thus, this metric gives us an estimate of the overall cofiring strength of each CA1 neuron. We have now added a new section in the Methods under “Calculation of a single cofiring metric and separation into populations of high and low cofiring neurons” on Page 44, Lines 34-43:

      “Since we wanted to relate the above cofiring metric to other measures, we needed to obtain a single-value cofiring metric for each neuron. To do this, we averaged the cofiring values across all pairings for a neuron (e.g. 1 CA1 neuron paired with all PFC neurons) and reported it as the cell’s cofiring. Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0 (High cofiring > 0; low cofiring < 0). Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions (High cofiring > mean; low cofiring < mean).”

      We have also clarified this in the Figure 7 legend on Page 27, Lines 25-28:

      “For comparisons between cofiring and other metrics (e.g. firing rate), a single cofiring value was calculated for each neuron by averaging the cofiring metric across all neuron pairings. Additionally, high and low cofiring CA1 neurons were split based on average cofiring values > 0 and < 0, respectively.”

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate high-frequency oscillations (HFOs) in the prefrontal cortex during REM sleep. They identify a specific pattern where these HFOs occur in "chains" that are phase-locked to theta oscillations, primarily during the "phasic" periods of REM. The study contrasts these events with isolated HFOs and NREM ripples, suggesting a unique role for these chains in coordinating activity between the prefrontal cortex and the hippocampus. Most notably, the authors report that a specific subset of hippocampal cells-those that co-fire with the prefrontal cortex during these HFOs-increase their firing rates over the course of sleep, suggesting a potential mechanism for selective memory consolidation.

      Strengths:

      The study addresses an under-explored area of sleep physiology: the fine-grained temporal coordination between the cortex and hippocampus during REM sleep. The identification of HFO "chains" and their association with higher theta power provides an interesting framework for understanding how the brain might organize information transfer outside of NREM sleep. The observation that specific hippocampal populations show differential firing rate changes based on their participation in these HFO events is a striking finding that warrants further investigation.

      Weaknesses:

      The primary weakness of the study lies in the lack of a clear distinction between global brain states and the specific events being analyzed. Because the authors compare HFOs across different sleep stages (NREM, tonic REM, and phasic REM) without sufficient controls, it is difficult to determine if the observed differences are intrinsic to the HFOs themselves or simply a reflection of the different physiological states in which they occur.

      We would first like to note in case it was unclear – as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly spiking activity patterns observed in prefrontal cortical-hippocampal circuits during these event-types that occur in the two sleep stages. As noted in response to Reviewer 1’s Comment #8 and Reviewer 3’s Comment #1, to avoid ambiguity and to remain consistent with existing literature, we use the following terms to describe these sleep-state specific events throughout the manuscript: Hippocampal sharp-wave ripples (SWRs) in NREM sleep, PFC ripples in NREM sleep, and PFC HFOs in REM sleep.

      We refer the Reviewer to our Response to Reviewer 1’s first comment where we provide additional analyses of theta periods outside of HFO times (i.e. baseline REM periods), including Figure S5. We also further address these comments as a response to Reviewer 2’s major comment #1 below.

      Furthermore, the evidence for "structured reactivation" is not yet convincing. The temporal alignment of these reactivation events appears inconsistent, with peaks occurring well before the HFO itself, and the analysis does not sufficiently control for pre-existing cellular assembly strengths.

      We have now addressed this as a response to Reviewer 2’s major comments #3 and #8 below, including Figure 6.

      Additionally, some of the sleep architecture presented appears atypical, such as very short REM bouts and direct NREM-to-REM transitions that bypass standard progression, raising questions about the consistency of the sleep detection across animals.

      We have now addressed this as a response to Reviewer 2’s major comment #2 below, including Figure S1 and Figure S3. We expect that addition of these figures will mitigate concerns about our sleep staging procedures and will clarify certain points raised by the Reviewer.

      Finally, the study does not account for potential confounds like baseline firing rates when interpreting the behavior of "high-cofiring" neurons, which may simply be the most active cells in the population.

      We have now addressed this as a response to Reviewer 2’s major comment #6 below.

      Reviewer #2 (Recommendations for the authors):

      In this study, the authors detect activity periods during REM sleep that feature high-frequency oscillations in the prefrontal cortex. They term these REM HFOs and report that they occur in "chains" phase-locked to theta oscillations. They contrast these with HFOs that appear in isolation and with ripples observed during NREM sleep, either in the PFC or in the hippocampus. The data presented is very interesting in places. The authors show that these HFOs are observed primarily during "phasic REM" periods that have higher theta power. It appears that overall firing is substantially lower surrounding these events. Intriguingly, it appears that CA1 cells that co-fire with the PFC during HFOs increase their firing rates over the course of sleep, whereas other neurons show a firing decrease. This seems to be the most striking finding of this study.

      Major:

      (1) The findings are generally intriguing, but the study makes some choices that are hard to understand. They begin comparing HFOs that occur in chains to HFOs that occur in isolation, even though it appears that chained and isolated HFOs likely occur at different times, some during tonic REM, some during phasic REM, and others during NREM. The lack of control for sleep state makes it difficult to determine if the reported differences are intrinsic to the HFO patterns or merely reflections of the underlying global brain state. Overall, I just didn't quite understand the motivation for comparing isolated and chain HFOs, as it seems natural that chains would occur during periods of greater synchrony.

      We appreciate the Reviewer for raising these points and agree that the properties of ripples and HFOs cannot be interpreted independently of the underlying brain state. As we mentioned in the initial response, we do expect that the generation of these ripples/HFOs in NREM and REM sleep are inextricably linked to global brain state (ex., cholinergic tone, as shown in the model in Figures 9-10), which results in differing patterns of activity across sleep states. We noted in the response to the public review comment #1 above that as clearly noted in our manuscript title, state-dependence is an integral part of our results, with the primary comparison in the manuscript being between REM cortical HFOs and NREM cortical ripples, and correspondingly, spiking activity patterns observed in cortical-hippocampal circuits during these event types that occur in the two sleep stages.

      Sleep state: Our primary goal in comparing isolated and chain HFOs (chain vs. isolated HFOs are only compared in REM periods, not NREM periods) was not to suggest that the differences that we observe are state-independent, but rather to provide evidence that temporal clustering of HFOs underlies distinct PFC dynamics, as well as enhanced CA1 engagement. Rather than attempting to dissociate ripple/HFO occurrence and the differential physiology that underlie NREM and REM sleep states, we view them as complementary—both sleep states are permissive to generation of high-frequency oscillations. Similarly, putative tonic and phasic REM substates (low vs high theta power) are differentially permissive for isolated and chain HFOs (Figure S3C). We add additional detail on tonic vs. phasic REM substages in Supplementary Figure S3, showing that the rate of HFOs and HFO chains is significantly elevated in putative phasic REM. While we would have liked to analyze REM HFOs during tonic and phasic states separately to be able to make more concrete conclusions, the scarcity of phasic REM sleep made it difficult to make accurate comparisons between the two, especially for spiking modulation. We have therefore removed the qualifying term “phasic REM” from the abstract.

      Also, we would like to clarify that while we investigated chains of events (ripples) in NREM sleep as a comparison to REM HFOs, we do not refer to them as HFOs in NREM anywhere in the manuscript. In NREM, while there are PFC ripples that are clustered into chains based on our definition (separation of <200 ms), we do not observe a prominent peak in the IEI distribution that would suggest entrainment by other oscillations (e.g. theta or spindles). This analysis was only to show that associated results are specific to REM HFO chains, and not seen during comparable chains of NREM cortical ripple events.

      Chain vs isolated HFOs: Regarding the Reviewers comment about why we chose to compare isolated and chain HFOs, the motivation is clearly demonstrated by differences in these events in spectral properties (Figure 3H-I) and spiking modulation (Figure 4F-G). Our initial motivation to investigate these chains of events in REM sleep came from our observation that PFC population activity aligned to REM HFOs was theta-modulated (multipeaked, suggesting multiple high-frequency events over a short duration). This led us to hypothesize that there may be chaining of events in REM sleep, which in line with previous studies demonstrating the clustering of events, such as spindles in NREM sleep (Darevsky, Kim et al. 2024). In that study (Darevsky, Kim et al. 2024), the authors showed that reactivation of motor patterns was more persistent during trains of spindle events as compared to isolated spindles. Furthermore, trains/chains of multiple hippocampal SWRs have been shown to underlie the replay of extended experience (Davidson, Kloosterman et al. 2009). Thus, the separation of high-frequency events into isolated and chained events has precedence and may have functional significance. Furthermore, since we see that isolated and chain events have a bias for occurring during putative bouts of tonic and phasic REM sleep, respectively (Figure S3C), characterization of both event types is an important step for understanding potential differences in interregional interactions during REM sleep. Furthermore, a recent appreciation for the role of sleep stage sub-states (Chang, Tang et al. 2025) further emphasizes the importance of investigating these events separately. We expect future studies to further dissect the roles of tonic and phasic REM states in memory and cognition, and our finding provides an account of the existence of different events for future reference.

      HFO chains during periods of greater synchrony: While it may seem natural that chaining would occur during periods of high synchrony (theta synchrony here), we show that our result is HFO-specific, especially for spiking activity modulation (Figures 4F-H and Supplementary Figures S5K-N), which makes it novel and important. Previous studies have focused on theta-gamma cross-frequency phase amplitude coupling, which has been proposed as a mechanism where slower theta oscillations temporally organize faster local population activity, thereby synchronizing neural ensembles within and across brain regions (Belluscio, Mizuseki et al. 2012). Here, we present a comparison of HFOs and gamma events and show that while gamma events can also occur in chains (added to Supplementary Figure S4H), possibly due to increased synchrony as the Reviewer stated above, we do not observe theta-modulated population activity aligned to these events (Supplementary Figure S4J), which is a defining property of HFOs that we propose underlies our results. This indicates that our reported results are unique to periods of synchrony associated with HFOs.

      Control analyses for theta periods: Regarding the comment about lack of control for sleep state, we refer the Reviewer to our response to Reviewer 1’s comments. We have performed additional control analyses comparing PFC activity during REM theta periods adjacent to detected HFOs and demonstrate that phasic PFC population activity is largely absent (Figure S5).

      We are aware that it is difficult to dissociate the generation of ripples and HFOs from the underlying brain state, since they are so tightly linked. However, we do expect that our clarifying points in addition to the REM theta state controls provide strong evidence that the results that we present are intrinsic to HFOs and not general reflections of activity during baseline theta activity in REM. Here, we further reiterate a few of the main results that demonstrate this:

      (1) Phasic spiking modulation of PFC population activity associated with HFOs is not present during baseline theta periods (Supplementary Figure S5K).

      (2) Phasic spiking modulation during HFOs is not linked to extracted gamma events (Supplementary Figure S4J).

      (3) Phasic spiking modulation, as well as activity suppression, is strongest during chains of HFOs (Figure 4F, Supplementary Figure S4J).

      (4) PFC-CA1 theta coherence surrounding HFOs increases relative to baseline (Figure 3E, z-scored relative to baseline coherence).

      (5) Assembly peaks are sequentially organized surrounding HFO chains but not during isolated or shuffled (baseline) HFO times (Figure 6C, Supplementary Figure S7E, and Author response image 2).

      (6) A higher proportion of CA1-CA1 pairs are high cofiring during HFOs as compared to baseline periods (Figure 7B), indicating specific CA1 engagement during HFOs (in addition to coherence).

      Lastly, we show that HFOs detected during periods of active behavior (high theta) on the W-Track are not associated with many of the defining features of HFOs in REM sleep, thus demonstrating the specificity of REM HFOs despite similar background theta activity (Figure S11, in response to Reviewer 3 comment #1).

      (2) It is crucial that the study provides REM-specific sleep examples for each of the data sessions, marking tonic and phasic REM and indicating when isolated and chain HFOs are observed. It remains unclear how interspersed these events are. Do some REM episodes just have isolated HFOs and others chains?

      We thank the Reviewer for raising this point and apologize for the lack of clarity. For additional transparency, we now provide example sleep state plots for each animal (Supplementary Figure S1H-I).

      We also provide hypnograms that show bouts of putative phasic REM on top of REM periods (Supplementary Figure S3). Additionally, as a compact way of demonstrating the validity of our separation of putative tonic and phasic bouts, we provide a plot showing the average velocity and spectrogram surrounding putative phasic REM bouts across all putative phasic REM transitions (Supplementary Figure S3A). Regarding the incidence of HFOs in tonic and phasic REM, we refer the Reviewer to Supplementary Figure S3D, where we show that HFO rate is significantly higher during putative bouts of phasic REM, during which they tend to be organized in chains as compared to putative tonic bouts (Supplementary Figure S3E). In line with this, we additionally report that although both isolated and chain HFOs occur in putative tonic and phasic bouts of REM sleep, the proportion of chain HFOs during phasic REM is significantly higher than that of isolated HFOs (Figure S3), indicating a bias for chains to occur in phasic REM sleep, potentially due to stronger theta input. However, since putative phasic REM accounts for <10% of REM sleep in our dataset, consistent with previous reports using similar methods (Mizuseki, Diba et al. 2011), we were unable to restrict spiking analysis to HFOs in phasic REM. We found that chain HFOs during both putative tonic and phasic REM sleep elicited theta modulated population activity in PFC, thus we pooled HFOs across REM states for analysis.

      While a larger proportion of chain events occur in putative bouts of phasic REM sleep as compared to isolated events (Figure S3), chain and isolated events occur in both putative tonic and phasic substates (Proportion of events in putative tonic REM is (1 - proportion in phasic) shown in Figure S3C, right). Thus, the majority of analyses comparing isolated and chain HFOs, especially for spiking data, were pooled across states. Instead of showing example plots for all 22 sleep epochs (22 out of 36 epochs with >5 s of putative phasic REM), we expect that this analysis will be sufficient to illustrate the distributions of isolated and chain HFOs across putative REM sleep substates. We have also removed the qualifying term “phasic REM” from the abstract.

      To clarify this, we have added a statement to the new “Limitations” section of the manuscript on Page 16, Lines 10-20:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states. (Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult. Finally, although our model predicts that distinct cell-type activity profiles shape REM sleep HFO dynamics, we did not record a suDicient number of interneurons to test these predictions directly. Future studies using appropriate behavioral tasks, longitudinal sleep recordings, and cell-type specific opto-tagging will be able to resolve these limitations and further clarify the roles of high-frequency oscillations in REM sleep.”

      We have now added Supplementary Figure S3.

      This is also important because some of the REM sleeps detected and shown in Figure S1H seem unusual. For example, in Animal 1, there are multiple bursts of REM that seem very irregular. In some other sessions, animals occasionally appear to enter REM sleep with very little preceding NREM, which goes against existing literature. It's not clear which ones of these meet the 30 s duration threshold. Are the chain events occurring in these periods?

      In reference to the hypnograms shown in Supplementary Figure S1H (Now Supplementary Figure S1I), these were generated by concatenating all 9 sleep epochs regardless of whether they passed the inclusion criterion of > 30 s of total REM sleep. Furthermore, only a subset of the epochs shown were included for analysis based on a secondary, manual inspection that is performed to confirm inclusion. Sleep state plots (e.g. Supplementary Figure S1H) were visually inspected to further confirm the transition into REM sleep, ensuring absence of noisy T/D ratio or spurious detection due to noisy signals – epochs where microarousals or persistent subthreshold fluctuations in animal movement induced noisy TD ratio increases, and thus inaccurate REM designation, were excluded. We thus used a total of 36 sleep sessions from a possible 90 sleep sessions.

      We apologize for not specifying what portions of data represented by the hypnograms were included. We have now provided updated hypnograms only illustrating the sleep epochs included for analysis (Figure S1I).

      We have now added Supplementary Figure S1.

      Regarding the Reviewer’s point “Are the chain events occurring in these periods?”: Yes, all of the chain events (and isolated events) come from these updated epochs that are now shown. No events in the excluded epochs were included, since we wanted to only analyze the data that came from curated REM epochs that we were confident in.

      (3) The analysis supporting structured reactivation was not generally convincing. Figure 4E does not provide convincing evidence of this. Indeed, reactivation strength is lowest around the time of the event, and appears highest 0.5s before. The example panel seems rather anecdotal. It's also not clear why REM reactivations should be compared to NREM ones here. I could not follow what was done in Figures 4F-H. Why should the first half and second half of an HFO event be correlated?

      We appreciate the Reviewer for raising these concerns. We agree that this section is somewhat dense and at time hard to follow, so we will clarify with further explanations and analyses (see also response to Comment #8 with new reactivation figures).

      In Figure 6B (originally Figure 4E), our intention was to show that if you simply align assembly activation to REM HFOs, a relatively flat response is observed when averaged across all assemblies, in stark contrast to assembly reactivation during NREM cortical ripples. This could suggest a couple of things: 1) There is no real assembly activation in response to REM HFOs or 2) Assembly activation is organized differentially (compared to NREM) surrounding HFOs. What we hypothesized, and then quantified based on observations, is that assembly activity is sequentially organized around HFOs. If this were true, it would suggest a consistent temporal relationship between REM HFOs and assemblies (for example, assembly 1 tends to be active 115 ms after the onset of chains, assembly 2 is active 230 ms after, etc.). Of course, the sequential pattern that we show in Figure 6A may arise trivially, especially since these types of sequential visualization plots can simply arise from noise. Thus, a cross-validation method must be utilized to ensure the sequences are indeed reflective of an underlying computation.

      In the methods section “Assembly sequence detection surrounding HFOs” we explain the splitripple/HFO procedure that we used to compare assembly sequences across two halves of the data, similar to methods used in hippocampal place cell sequence cross-validation (Plitt and Giocomo 2021, Sosa, Plitt et al. 2025). For every split and assembly alignment, a sequence similar to Figure 6A, right is generated based on the first half of aligned data. Then the second half of aligned data is sorted based on the peak indices of the first half of data and the correlation between the peak reactivation indices across all assemblies for the two datasets is calculated. A high correlation indicates high sequence similarity across the two halves of data (not two halves of an HFO as the Reviewer mentioned), suggesting temporal consistency of assembly reactivation. By utilizing this method, we show that assembly reactivation sequences across randomly chosen halves of data are most similar for chained HFOs (Figure 6C and Supplementary Figures S7C-F). We have now provided an additional control analysis that investigates this sequential assembly activity for time-shifted chain events (Author response image 2). As in Figure 6, assembly activity was aligned to the first event in each time-shifted chain. In addition to the analyses in Supplementary Figure S7, this control further indicates that sequential assembly activity is preferentially restricted to HFO chains.

      Author response image 2.

      Assembly sequences surrounding time-shifted HFO chains. (A) Distributions of r values calculated from the Pearson correlation between peak reactivation bins across all assemblies for two randomly chosen halves of the shifted HFO-aligned data. There was no difference between r values for time-shifted chains and shuffled data, indicating no structured assembly activity during periods outside of real HFOs chains.

      To further quantify this, we calculated two additional metrics: 1) the slope difference between fitted lines for the two halves of data for each split and 2) the absolute peak difference between assemblies in the two halves of data. First, for the slope difference metric, the slope of the best-fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data in Figure 6D indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. Secondly, Figure 6E is a quantification of the peak reactivation displacement between the two halves of data. If there is a high probability of small peak differences, as in the real data, this indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data.

      We apologize for the omission of the description of the slope and peak difference metrics that we used in Figures 6D,E. We have added information in the Methods section under “Assembly sequence detection surrounding HFOs” on Page 43, Lines 23-33:

      “Furthermore, the slope and peak differences were calculated as additional metrics of sequence and temporal reactivation consistency between the two halves of data, respectively. For the slope difference metric, the slope of the best fit line between assembly ID and peak reactivation index was taken for the two halves of data and compared. A smaller slope difference compared to shuffled data indicates that assembly reactivation is structured in a more similar manner across the two halves of the real data. For peak reactivation difference, the temporal displacement of the peak reactivation index between the two halves of data was calculated and compared to shuffle. A high probability of small peak differences indicates that the timing of peak reactivation of assemblies relative to HFO chain onset is similar across the two halves of data. Shuffling of assembly strength was carried out as above.”

      (4) The authors argue that the occurrence of HFOs, rather than theta power, is the reason for lower MUA activity, but the analysis for this (Figure S4I) is quite confusing. The left panel actually seems to indicate that MUA is indeed lower when theta power is high.

      We apologize for the confusion regarding this figure and appreciate the Reviewer’s point that periods of high theta power can also appear to be associated with reduced MUA. We agree that the original presentation may not have clearly separated the contributions of theta and HFOs. What we convey with Supplementary Figure S5K (originally Supplementary Figure S4I) and new Supplementary Figures S5L-M is that the observed theta-modulated PFC spiking response (Figure 2A) cannot be solely explained by baseline theta periods (outside of HFOs) in REM sleep. Since theta oscillations are ubiquitous during REM sleep, an important control is to demonstrate that the theta-modulated population activity is not simply a consequence of ongoing theta activity. Thus, we aligned PFC activity to theta oscillations of varied power and show that the fluctuating, theta-modulated population activity is absent, indicating that HFOs associated with the theta oscillation are driving this phasic response.

      We are not solely arguing that the presence of HFOs is the driver of decreased PFC multiunit activity. In both the data and model, we show that the magnitude of theta power detected in PFC (possibly input to PFC) is inversely related to multiunit activity (Figure 9E). Our interpretation is not that theta is unrelated to MUA, but rather that HFO occurrence provides additional explanatory power beyond theta alone. Accordingly, we show directly in Supplementary Figures S5L-M, in response to Reviewer 1’s comment #1, that theta periods in phasic REM not associated with HFOs do not elicit the MUA activity suppression similar to HFOs.

      The phase-alignment performed is hard to follow and is not being applied to HFO periods. If the study is trying to argue that high-theta periods without HFOs in the same recording sessions show lower MUA, then perhaps some sort of shuffle or jitter would be more suitable. For example, in Figure 1B, it seems there are some high-theta periods that don't have HFOs and appear to have higher MUA.

      We expect that the new analyses where we provide additional baseline theta controls for periods adjacent to HFOs in Supplementary Figures S5L-M now clarifies this point. Briefly, alignment to theta phases at different temporal distances from detected HFOs does not exhibit the same fluctuating PFC activity as HFO alignment.

      Regarding the phase alignment procedure that we used for the control analysis, since REM sleep is characterized almost entirely by ongoing theta activity, the control condition was not a separate brain state but rather theta periods outside of HFOs. We therefore needed a systematic way to select comparable theta cycles and phase bins in order to make a valid comparison with HFO-aligned PFC population activity. Since we demonstrated that there is significant phase amplitude coupling between theta and HFOs, thus a theta phase preference of HFOs, we used that specific phase bin across multiple theta cycles to align PFC activity. This phase bin selection was performed separately for each epoch to account for inter epoch and animal variability in phase preference. We reasoned that this procedure would allow for a valid comparison as compared to random alignment, since activity was aligned to similar phases in the baseline theta vs HFO conditions.

      (5) As far as I could tell, the study does not distinguish between putative excitatory and inhibitory neurons in the PFC, but only in the CA1, even though these play very different roles in the model. What is the rationale for not separating these? How are reactivations to be interpreted among interneurons?

      We apologize for the lack of emphasis on this point, which was originally highlighted in Supplementary Figure S2G (now in Supplementary Figure S2F), and for omitting the explanation as to why we did not separately analyze putative excitatory and inhibitory neurons. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with recording primarily from pyramidal neurons (Supplementary Figure S2F, left). We therefore decided to pool and not separate the populations into putative pyramidal cells and interneurons for the spiking analyses presented. To further validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO-aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript, as we did not have enough interneurons to investigate them separately. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      Regarding the interpretation of reactivation in the context of interneurons, we expect reactivation reflects coordinated ensemble activity with excitatory neurons encoding task-relevant information, and inhibitory interneurons shaping timing and neural synchronization. We however did not record enough distinct interneurons to test the predictions of the model, which is now noted in the Limitations on Page 16.

      Relatedly, we find that PFC assemblies detected from pooled data have task relevant representations (Figures 5F-G).

      (6) Are the high-cofiring CA1 neurons generally higher-firing than the other cells? Could this perhaps explain why they behave differently?

      We apologize that this information was not more evident in the manuscript, as it is an important control. In Supplementary Figure S9A, we show that there was no difference in baseline firing rate between low and high cofiring CA1 neurons.

      We have now explicitly referenced this figure in the main text on Page 9, Lines 42-45:

      “Analysis of low and high cofiring CA1 neurons during REM HFOs showed that high cofiring neurons exhibited elevated activity during chained events as compared to low cofiring neurons (Figure 7D), independent of baseline firing rates (Figure S9A).”

      (7) It appears that the decreased firing around HFO's could be a consequence of the stronger firing modulation around these periods, related to time averaging, rather than suppression per se. How does the firing rate compare to other periods with similar modulation that might not have HFOs?

      We thank the Reviewer for raising this important point. We expect that the new analyses, where we provide additional baseline non-HFO-associated theta periods, and theta periods adjacent to HFOs at different temporal distance as controls in also Supplementary Figure S5L-N now clarifies this point. These figures show that the decreased firing rate is specific to HFO chains (Supplementary Figure S5N). Indeed, if the suppression that we observe is related to time averaging, or another analytical artifact, our claims of suppression during HFOs would not be valid. We present a number of results and provide further explanations to support the accuracy of our characterization of PFC population suppression.

      First, event-aligned multiunit activity was quantified as baseline-normalized population firing relative to detected events. For both NREM and REM, activity was normalized by the mean population firing rate during a baseline period within the same sleep state in which events were detected. Values are therefore expressed as deviations from baseline. This normalization allows comparison of relative changes in firing around events within each state. Importantly, values below baseline reflect reductions relative to the state-matched baseline period.

      Second, we show that smoothing activity with a larger gaussian kernel preserves the dip in population activity, consistent with suppression of activity surrounding HFOs (Figure 9C, note that this is a data figure presented in the context of the model). However, as raised by the Reviewer, this normalized measure does not on its own distinguish sustained suppression from transient deviations introduced by event-locked temporal structure in firing.

      Third, to address this, we employed an alternative method to demonstrate that HFO chains tend to occur during periods of PFC suppression (Figure 4G). Briefly, we detected events in PFC where activity fell below a threshold and calculated the probability of HFOs surrounding these “suppression” events. We refer the Reviewer to the Methods section under “Detection of population suppression” where we explain this procedure in more detail. We found that chain ripples, during which the strongest suppression is observed (Figure 4F), are associated with decreases in PFC activity (Figure 4G and Supplementary Figure S6).

      Fourth, we refer the reviewer to Figure S5N, where we compare the firing rates of PFC neurons during HFO chains and theta periods outside of HFOs during putative phasic REM bouts. The observed reduction in firing during true HFO chains compared to non-HFO periods therefore reflects HFO event-specific activity suppression rather than a common occurrence during baseline REM periods.

      Lastly, we performed a control analysis complimentary to Supplementary Figure S5K where we aligned PFC activity to the preferred theta phase of HFOs and investigated how distance from detected HFOs modulates PFC activity (Figure S5). We found that there was no consistent theta-modulated activity aligned to HFO-adjacent theta phases.

      Regarding the final comment, if the Reviewer meant “modulation” as in the theta modulation or suppression observed when PFC population activity is aligned to HFOs (Figures 2A and 4F), we are not aware of any other REM periods where this strong theta modulation or suppression of PFC population activity is present. To our knowledge, we are the first to demonstrate such a modulation of PFC population activity in REM sleep. The closest comparison that we can make is PFC activity aligned to gamma events that are coordinated with HFOs (suppression of PFC), but this is explained only with association with HFOs, as shown in Supplementary Figure S4J.

      (8) The assembly reactivation measure does not control for pre-existing assemblies. The term "activation strength" would therefore be more appropriate.

      We thank the reviewer for this important methodological point. The concern that ICA-based reactivation strength does not, by itself, distinguish behavior-induced reactivation from pre-existing assembly activity/structure is well-taken, and we have implemented several complementary analyses that directly address these concerns.

      First, the interleaved structure of our recordings (8 run epochs interleaved with 9 sleep epochs) allows us to investigate the within-session pre/post assembly strength differences (i.e. each W-Track run session has a preceding (pre) and following (post) sleep session). An increase in assembly strength from pre to post is a hallmark of behaviorally relevant assembly reactivation (Kudrimoti, Barnes et al. 1999, Peyrache, Khamassi et al. 2009). For each run epoch, the same run-derived templates were projected onto the preceding and following sleep epochs, and the pre and post strengths were compared. The distribution of post-minus-pre reactivation differences across all epoch pairs is significantly skewed toward positive values (Figure 6F), indicating that templates from the run epochs are more strongly expressed in the post-sleep epochs. This asymmetry cannot be explained by pre-existing assembly structure, which would predict similar assembly strengths.

      Second, reactivation strength in post-experience sleep increases across the experiment, with templates from later running epochs producing the strongest reactivation in the following post-sleep (Figure 6G). This increase in reactivation strength over time cannot be explained by preexisting assembly structure, which predicts similar assembly expression strength independent of experience.

      Third, the detected assemblies carry behaviorally meaningful structure. Assembly activation maps computed during running exhibit spatially organized "assembly fields" similar to single-cell place fields (Figure 6H), demonstrating that the detected assemblies represent specific spatial locations or task variables rather than behavior-independent states. Pre-existing co-firing structure unrelated to ongoing experience would not be expected to produce spatially tuned, task-locked assembly activation. Furthermore, this spatial tuning of assemblies was verified by comparison with surrogates, where assembly maps were generated using circularly shuffled activation times (1000 shuffles). Assemblies with p < 0.05 (z > 1.65) were considered to have significant spatial structure (Figure 6I).

      These results establish that what we measure is the selective re-expression of behaviorally relevant assemblies in subsequent sleep epochs, consistent with the use of "reactivation" in the established literature (Peyrache, Khamassi et al. 2009, Lopes-dos-Santos, Ribeiro et al. 2013). We have therefore retained the term "reactivation strength" and have added text to the manuscript noting these new results that justify the use of “reactivation” strength.

      We have now added Figure 6.

      We have also added the procedure for the calculation of spatial information to the Methods section under “Spatial information of assembly fields” on Page 44, Lines 8-23.

      (9) Can the study rule out that the rank-ordering in Fig 4C is related to firing rates? Higher-firing rates tend to fire earlier, and lower-firing cells later.

      We thank the Reviewer for raising this interesting point. Here, we assume that the Reviewer meant the baseline firing rates of the neurons, not the intra-HFO firing rates of the neurons. Indeed, when we look at baseline REM firing rates of these PFC neurons, we do find that neurons with higher firing rates tend to fire earlier than low-firing-rate neurons (Author response image 3). This is also true when PFC rank and firing rate are assessed for isolated REM HFOs and NREM PFC ripples (Author response image 3). Similarly, we also observe this relationship in CA1 during SWRs, during which rank order correlation is typically assessed as a method for replay detection. In line with this, a previous study has shown that CA1 neurons with high excitability at animals’ current location tend to initiate replay events (Karlsson and Frank 2009). Furthermore, high-firing-rate, rigid CA1 neurons are more active during SWRs than low-rate, plastic neurons (Grosmark and Buzsaki 2016), and there are distinct populations of neurons in both hippocampus and PFC that are preferentially active during immobility in sleep epochs (Jarosiewicz, McNaughton et al. 2002, Kay, Sosa et al. 2016, Tang, Shin et al. 2017), potentially biasing replay activity during high-frequency events. Similar dynamics may underlie activity during PFC ripples and HFOs in NREM and REM sleep, respectively. The critical point here is in the leave-one-out cross-validation that we implemented to determine sequence similarity—each left out event’s cell rank was correlated with the averaged rank of the template that was generated from all other events. This analysis provides a basis for our claim that there is preserved sequential PFC activity across HFO chains. We did not observe neuron firing consistency during isolated HFOs or during pseudo-HFO chains (coherently shifted chain HFO times), which indicates that REM HFO chains are unique temporal windows during which PFC activity proceeds in a more structured manner.

      Author response image 3.

      Firing rate difference of low and high rank neurons (A) Comparison of baseline firing rates of PFC and CA1 neurons split by average rank across all PFC REM HFOs, PFC NREM ripples, or CA1 SWRs. Baseline rates were calculated separately for NREM and REM sleep.

      Minor:

      (1) It gets confusing that the authors sometimes (but not always) refer to HFOs during NREM as "ripples" but not if they occur during REM. The terminology is inconsistent. When they refer to HFO chains, it seems they now pool between REM and NREM periods, as well as across phasic and tonic REM periods, which is confusing.

      We apologize for the confusion regarding the terminology. In the revised manuscript, we now use NREM ripples exclusively for NREM events and REM HFOs exclusively for REM events. We have removed mixed labels such as “ripple/HFO” except where a collective term is explicitly defined. We also clarified that HFO chains refer to REM events only and revised the relevant text/figure legends to avoid any implication that chain analyses pool NREM and REM events.

      (2) P7 L8: It might be helpful to emphasize "broader temporal distribution".

      We thank the Reviewer for the suggestion. We have updated the text on Page 8, Line 18:

      “Since we observed a broader temporal distribution of activity…”

      (3) P8 L9: What do they mean by spatial? Do they mean the place-fields of these same neurons during a previous task period?

      We apologize for the confusion. The Reviewer is correct. Here, we calculated the spatial rate map correlation between CA1 neurons as a measure of place field similarity during the W-Track session prior to the sleep epoch being assessed.

      For clarification, we have added “rate map” to the text on Page 9, Line 33:

      “…spatial rate map correlation…”

      We have also updated the y-axis label for Figure 7B for clarity.

      (4) P8 L27: What do they mean by "coordinated SWRs"? As opposed to what?

      Here, we are referring to our previous study where we investigated ripples in NREM sleep and showed that ripples and SWRs in PFC and CA1, respectively could either be independent from or coordinated with events in the other region (Shin and Jadhav 2024). A main result in the study showed that CA1 neurons are strongly suppressed during independent PFC ripples and that there was a relationship between activity suppression and reactivation during coordinated SWRs (CA1 SWR-PFC ripple coordination in NREM). We specifically mentioned “coordinated” since these are SWRs that are also coupled with SOs and spindles as compared to SWRs that are independent from PFC ripples (Shin and Jadhav 2024). Overall, we wanted to frame this result in the context of oscillatory coupling and mechanisms of memory consolidation.

      (5) P37 L28 says "we observed a bimodal distribution" but L31 says "unimodal". Which is it?

      We apologize for the confusion. We observed a bimodal distribution for CA1-PFC cofiring in REM sleep, but a unimodal distribution in NREM sleep. Because of these two observations, we decided to split the CA1 population into high and low cofiring neurons based on two different thresholds:

      (1) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) 0.

      (2) Splitting the population by cofiring values greater than (high cofiring) or less than (low cofiring) the mean of the distribution of averaged cofiring values.

      Using two separate thresholds to split high and low cofiring CA1 neurons demonstrates the robustness of the firing rate change result in Figures 7F and Supplementary Figures S9B-D.

      We added a statement that clarifies that the bimodal distribution was seen in REM sleep only on Page 44, Lines 37-43:

      “Furthermore, since we observed a bimodal distribution of CA1-PFC cofiring values in REM sleep, we split the population based on whether the average (across all cell-cell combinations) cofiring value or correlation coefficient of a cell was above or below 0. Also, since the NREM ripple cofiring distribution was unimodal, we additionally split the populations using the mean of the average cofiring or correlation coefficient distributions.”

      Reviewer #3 (Public review):

      Summary:

      Shin et al. examine hippocampal-prefrontal interactions during sleep using simultaneous CA1 and prefrontal cortex recordings in rats performing a spatial memory task. They identify high-frequency oscillation (HFO) events in PFC during REM sleep that occur in theta-modulated chains and are associated with increased CA1-PFC coherence and sequential, sparse reactivation of cortical ensembles. This pattern contrasts with the synchronous reactivation observed during NREM cortical ripples. Together with a simple cholinergic network model, the authors propose that REM HFO chains represent a distinct mechanism for hippocampal-cortical coordination that complements NREM ripple-mediated processing during sleep.

      Strengths:

      A major strength of the work is the extensive electrophysiological dataset, which includes simultaneous recordings of large neuronal populations in both hippocampus and prefrontal cortex across behaviour and subsequent sleep. The analyses linking high-frequency events to population dynamics, interregional coherence, and ensemble reactivation are technically sophisticated and provide an incredibly detailed description of REM-associated cortical activity patterns. In particular, the demonstration that REM HFOs occur in chains aligned to theta phase and organise sequential activation of cortical assemblies represents a potentially important advance in understanding the neural structure of REM sleep activity. The integration of experimental data with a computational model further provides a useful framework for interpreting the observed differences between REM and NREM network states in terms of neuromodulatory influences.

      Weaknesses:

      While overall this study provides a highly valuable body of work, there are two primary limitations, which, if overcome, would provide substantially more significance to the overall characterisation of REM HFOs. Specifically:

      (1) Distinction from wake HFOs

      The results largely support the authors' claim that REM HFO chains represent a distinct pattern of neural coordination compared to NREM cortical ripples. The analyses consistently show differences between REM and NREM events in terms of neuronal modulation, ensemble structure, and interregional coupling. However, similar high-frequency events during wake are not examined. Since REM sleep shares several network features with wakefulness, including strong theta oscillations, evaluating whether comparable PFC HFOs occur during wake would provide clarity on whether these events are specific to REM sleep (and its associated functions) or represent a more general theta-associated phenomenon.

      To investigate PFC high-frequency oscillations during running behavior on the W-Track, events were extracted in the same manner as NREM and REM events (Methods). Events during wake were subset by periods where the animals’ velocity was >4 cm/s to provide a comparison of events during periods of high theta. While we were able to detect HFOs during wake that exhibited a similar spectral profile in the high frequency band, we did not observe 1) strong association with gamma or theta oscillations, 2) prominent HFO chaining, 3) HFO aligned theta modulated PFC activity, 4) comparable levels of theta phase amplitude coupling, 5) association with population suppression, or 6) a relationship between peri-event theta power and multiunit activity (Supplementary Figure S11). Many of the defining features of PFC REM HFOs are absent during wake, indicating REM specificity of the results we present.

      (2) Link to memory consolidation

      The manuscript proposes throughout that REM HFO chains may contribute to memory consolidation by coordinating hippocampal-cortical reactivation, but the evidence for this functional role remains indirect. The authors do highlight this as a limitation of the study - the inability to link their findings to learning - but it is not clear why. Further details of the behaviour results should be included. If no learning occurred across the eight behavioural sessions, this should be reported. If learning did occur, but could not be linked to HFO events, this should also be reported.

      To address these concerns, we have now added an explicit “Limitations” section in the main text of the manuscript that includes a statement about learning. We have also added Supplementary Figure S1 in the manuscript, which illustrates the performance of all 10 animals on the W-Track task. Finally, we have also included Figures 6F G in the manuscript, showing that PFC assembly reactivation strength during sleep epochs increases during learning.

      Reviewer #3 (Recommendations for the authors):

      Most of my specific comments were related to further clarification that will help the reader's understanding.

      (1) I'd recommend simplifying terminology. Open to debate, but would it not be simpler and clearer to just say NREM HFO vs REM HFO? Obviously, there is a need to mention how NREM HFOs have previously been referred to as cortical ripples, but I'm not sure it is such a helpful terminology to continue for the field, given, as you state, how different cortical ripples are from hippocampal SWRs. If not, I'd at least provide a clearer explanation early in the manuscript, distinguishing NREM ripples from REM HFOs but collectively still calling them 'cortical events'.

      We appreciate the Reviewer’s suggestion regarding terminology and agree that it would be simpler and clearer to use a single term, ripple or HFO, to describe these events. Initially, we had used a unified term (ripples across both states) but ultimately decided to switch to state-specific terminology due to previous comments we received on the manuscript and to emphasize the distinctions between NREM and REM events. Ultimately, we decided on calling them ripples in NREM and HFOs in REM since there is precedence for both terms in each respective sleep state (Khodagholy, Gelinas et al. 2017, Vaz, Inati et al. 2019, Bueno-Junior, Ruckstuhl et al. 2023, Shin and Jadhav 2024), but we do agree that this distinction can be confusing if not clearly stated. Thus, we have added an additional statement in the manuscript on Page 4, Line 44 to Page 5, Lines 1-3 for clarity:

      “Similar criteria were used to detect cortical high-frequency events in NREM and REM states; however, to avoid ambiguity and to conform to previous nomenclature, we refer to cortical NREM events as ripples, cortical REM events as HFOs, and hippocampal sharp-wave ripples in NREM as SWRs throughout”

      (2) I assume experiments occurred during the light phase, but it would be good if this could be stated explicitly.

      We have now added text specifying that these experiments took place in the light phase on Page 34, Lines 9-12 of the Methods section under “Behavior”:

      “During the recording day, animals were introduced to the novel W-maze (~80 × 80 cm with ~7 cm wide tracks) for the first time and learned the task rules over eight behavioral sessions during the animals’ light phase between the hours of 9 AM and 6PM.”

      (3) Page 2 Line 14: Rephrase to make clearer, e.g. 'that have a shift in...'.

      We have rephrased the sentence for clarification on Page 2, Lines 13-15:

      “REM HFO chains also preferentially engage CA1 neuronal populations that demonstrate a shift in their preferred theta-phase from behavior to REM sleep.”

      (4) Page 4 Line 36 - 'and find coherent shifts in TD', this is self-fulfilling. I would rephrase to something like 'resulting in...'.

      We have rephrased the sentence on Page 4, Lines 36-38:

      “We separated NREM and REM sleep stages based on theta-to-delta (TD) ratio in CA1, which revealed coherent shifts in TD ratio across CA1 and PFC at the onset and offset of REM sleep…”

      (5) Figure 1H left - clarify how many animals or multiunits this is based on.

      We have now updated Figure 1 legend to specify the number of animals and epochs included on Page 19, Line 4:

      “REM HFO aligned multiunit activity (MUA) in PFC (n = 10 animals, 36 epochs)…”

      (6) Figure 1I (right), it would be useful to see the x-axis frequency start from 0, since you are cutting the peak in power.

      We thank the Reviewer for this suggestion. We had initially set the frequency limits to 4 and 12 to specifically illustrate the absence of theta-modulated activity during NREM ripples. However, as suggested by the Reviewer, it is informative to expand the frequency range to ascertain the location of the peak frequency for NREM. Indeed, the peak frequency of NREM ripple-aligned PFC activity tends to be lower than 4 Hz, which is consistent with a single peak of activity that lasts <1 s.

      We have now updated Figure 2C with these new panels.

      (7) Figure 5 E/I - use of ** is confusing, it looks like a significance comparing e.g. quartile 2 to 1, but I think this is the correlation significance. I'd move ** to the top right corner and ideally include rho values.

      We thank the Reviewer for this suggestion. We have now added the r values for the correlation to Figures 7E and 8B to resolve any ambiguity. In addition, we updated Figure 7F, right and Figure 5B to maintain consistency across main figures.

      (8) Figure 7D - Why are stimulated and non-stimulated cells so different at baseline? This is not the case for the NREM results.

      The y-axes in Figures 10C,D (Formerly Figure 7) show the fraction of cells with at least one spike per 10ms bin, either for all pyramidal cells or all interneurons. We report in the figure legend that the stimulated cells represent only 30% of the network for any given stimulus (Page 32, Line 18), so at baseline in both the NREM and REM simulations, there are roughly 3x as many non-stimulated cells with a spike per 10ms bin compared to stimulated cells, as these populations are active at roughly equal firing rates per neuron outside of stimulation periods.

      (9) Methods - there is limited info on spike sorting procedure - can you provide a reference with further details (I couldn't find)? Is this method equally valid for identifying PFC units?

      Matclust is a MATLAB-based spike sorting graphical user interface that allows for manual curation of neuron clusters through the visualization of spike waveform amplitude, peak-to-trough, and principal components. Polygons or boxes are drawn around spike data points, and single unit clusters are resolved through refinement in multiple dimensions. It was developed by Mattias Karlsson and was first used in a publication reporting replay of remote experiences in the hippocampus (Karlsson and Frank 2009). Although there is no formal reference, it can be found at https://bitbucket.org/mkarlsso/matclust/src/master/. Other labs have used the software for clustering neurons from cortical areas (Yu, Liu et al. 2018, Proskurin, Manakov et al. 2023), which demonstrates its robustness across multiple brain areas.

      Additionally, examples of clustered neurons in PFC over the course of the experimental paradigm used here can be found in our previous publication (Shin, Tang et al. 2019). In addition to the aforementioned references, we show that PFC neurons can be accurately clustered and that neurons are stable over time, according to a number of cluster metrics.

      (10) Why were there no further analyses of pyramidal cells and interneurons beyond Figure S2?

      We thank the Reviewer for bringing up this important point, which is similar to Reviewer 2, comment #5 above. We repeat our response here. When we plot the average waveform peak-to-trough and mean firing rates of the PFC neurons, we observe a large cluster with moderate mean firing rates and peak-to-troughs consistent with primarily recording from pyramidal neurons (Supplementary Figure S2F, left). All results are similar if we exclude putative interneurons. To validate our decision to pool the cells, we separated the population into putative pyramidal cells and interneurons based on peak-to-trough. Putative interneurons were identified as cells with a peak-to-trough <0.3 ms (we obtained similar results when using a hyperplane to separate units based on both peak-to-trough and firing rate, with a smaller subset identified as putative interneurons). When these putative interneurons were excluded from the HFO aligned multiunit PFC plot, we observed very similar activity to Figure 2A (Supplementary Figure S2F, right). We thus decided to pool the cells into a single population for the purpose of this manuscript. We are, however, aware that different cell types may contribute to the phenomenon that we report here and attempt to more thoroughly differentiate the contribution of pyramidal cells and interneurons with our modeling result in Figures 9 and 10.

      We have now added the following to the figure legend on Page 51, Lines 21-24:

      “REM HFO aligned PFC multiunit response when spikes from putative interneurons are excluded (compare to Figure 2A). Due to this similarity of the phasic PFC response when putative interneurons are omitted, we decided to pool PFC neurons for all further analyses.”

      (11) Looking at Figure S1B, the second to last main block of NREM sleep shown has a clear peak passing TD threshold, but oddly not classed as REM - I can only assume this is due to the duration limits on your classification?

      Yes, this is due to the REM duration threshold that we implement in our sleep scoring algorithm. We used a minimum REM bout threshold criterion of 10 s for inclusion, as in previous reports (Rothschild, Eban et al. 2017, Zhang, Zhang et al. 2020). Additionally, we have included example sleep plots for all 10 animals in Supplementary Figure S1.

      (12) Include details of how head speed was calculated - just based on the 30fps video?

      Yes, the head speed of the animal was determined by tracking the animals’ position and calculating the speed based on cm/pixel values. We have now added more detail on this in the Methods section under “Surgical implant and electrophysiology” on Page 33, Line 44 to Page 34, Lines 1-2:

      “Additionally, the animals’ speed was calculated based on predetermined cm/pixel values and the position displacement between frames captured at 30 fps.”

      (13) Page 31 line 7 - With the reference you cite, they didn't really show tonic and phasic REM can be segregated based on theta frequency - they just defined it as such. You've done it for some, but I'd ensure all references to phasic REM are defined as putative - mostly missed within the discussion - I'd also add this as a brief limitation (without recording of eye movements or PGO waves).

      We thank the Reviewer for these suggestions. We have now ensured that all mentions of “phasic REM” are qualified with “putative” and have added the requested limitation in the Limitations section on Page 16, Lines 10-15:

      “Third, we did not record eye movements or ponto-geniculo-occipital (PGO) waves, both of which would have allowed for more accurate segregation of tonic and phasic REM sleep states.(Simor, van der Wijk et al. 2020) Although we observed a bias for isolated and chain HFOs to occur in putative tonic and phasic REM substates, respectively, the scarcity of putative phasic REM bouts made the direct comparison based on substage difficult.”

      We have also modified the wording in the Methods section under “Theta inter-peak intervals during bouts of high and low theta power” on Page 37, Lines 14-15 to indicate that the referenced article simply used the theta frequency-based method to define tonic and phasic REM – not to explicitly separate the two states:

      “Since previous studies have segregated putative tonic and phasic substages of REM sleep based on CA1 theta frequency…”

      References:

      Abdou, K., M. Nomoto, M. H. Aly, A. Z. Ibrahim, K. Choko, R. Okubo-Suzuki, S. I. Muramatsu and K. Inokuchi (2024). "Prefrontal coding of learned and inferred knowledge during REM and NREM sleep." Nat Commun 15(1): 4566.

      Aleman-Zapata, A., R. G. M. Morris and L. Genzel (2022). "Sleep deprivation and hippocampal ripple disruption after one-session learning eliminate memory expression the next day." Proc Natl Acad Sci U S A 119(44): e2123424119.

      Belluscio, M. A., K. Mizuseki, R. Schmidt, R. Kempter and G. Buzsaki (2012). "Cross-frequency phase coupling between theta and gamma oscillations in the hippocampus." J Neurosci 32(2): 423–435.

      Bueno-Junior, L. S., M. S. Ruckstuhl, M. M. Lim and B. O. Watson (2023). "The temporal structure of REM sleep shows minute-scale fluctuations across brain and body in mice and humans." Proc Natl Acad Sci U S A 120(18): e2213438120.

      Cairney, S. A., S. J. Durrant, R. Power and P. A. Lewis (2015). "Complementary roles of slow-wave sleep and rapid eye movement sleep in emotional memory consolidation." Cereb Cortex 25(6): 1565–1575. Chang, H., W. Tang, A. M. Wulf, T. Nyasulu, M. E. Wolf, A. Fernandez-Ruiz and A. Oliva (2025). "Sleep microstructure organizes memory replay." Nature 637(8048): 1161–1169.

      Cheng, S. and L. M. Frank (2008). "New experiences enhance coordinated neural activity in the hippocampus." Neuron 57(2): 303–313.

      Darevsky, D., J. Kim and K. Ganguly (2024). "Coupling of Slow Oscillations in the Prefrontal and Motor Cortex Predicts Onset of Spindle Trains and Persistent Memory Reactivations." J Neurosci 44(43). Davidson, T. J., F. Kloosterman and M. A. Wilson (2009). "Hippocampal replay of extended experience." Neuron 63(4): 497–507.

      Ellenbogen, J. M., P. T. Hu, J. D. Payne, D. Titone and M. P. Walker (2007). "Human relational memory requires time and sleep." Proc Natl Acad Sci U S A 104(18): 7723–7728.

      Foster, D. J. and M. A. Wilson (2006). "Reverse replay of behavioural sequences in hippocampal place cells during the awake state." Nature 440(7084): 680–683.

      Ghosh, M., F. C. Yang, S. P. Rice, V. Hetrick, A. L. Gonzalez, D. Siu, E. K. W. Brennan, T. T. John, A. M. Ahrens and O. J. Ahmed (2022). "Running speed and REM sleep control two distinct modes of rapid interhemispheric communication." Cell Rep 40(1): 111028.

      Grosmark, A. D. and G. Buzsaki (2016). "Diversity in neural firing dynamics supports both rigid and learned hippocampal sequences." Science 351(6280): 1440–1443.

      Helfrich, R. F., J. D. Lendner, B. A. Mander, H. Guillen, M. PaD, L. Mnatsakanyan, S. Vadera, M. P. Walker, J. J. Lin and R. T. Knight (2019). "Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans." Nat Commun 10(1): 3572.

      Jarosiewicz, B., B. L. McNaughton and W. E. Skaggs (2002). "Hippocampal population activity during the small-amplitude irregular activity state in the rat." J Neurosci 22(4): 1373–1384.

      Ji, D. and M. A. Wilson (2007). "Coordinated memory replay in the visual cortex and hippocampus during sleep." Nat Neurosci 10(1): 100–107.

      Karlsson, M. P. and L. M. Frank (2009). "Awake replay of remote experiences in the hippocampus." Nat Neurosci 12(7): 913–918.

      Kay, K., M. Sosa, J. E. Chung, M. P. Karlsson, M. C. Larkin and L. M. Frank (2016). "A hippocampal network for spatial coding during immobility and sleep." Nature 531(7593): 185–190.

      Khodagholy, D., J. N. Gelinas and G. Buzsaki (2017). "Learning-enhanced coupling between ripple oscillations in association cortices and hippocampus." Science 358(6361): 369–372.

      Kudrimoti, H. S., C. A. Barnes and B. L. McNaughton (1999). "Reactivation of hippocampal cell assemblies: effects of behavioral state, experience, and EEG dynamics." J Neurosci 19(10): 4090–4101. Lee, A. K. and M. A. Wilson (2002). "Memory of sequential experience in the hippocampus during slow wave sleep." Neuron 36(6): 1183–1194.

      Leemburg, S., V. V. Vyazovskiy, U. Olcese, C. L. Bassetti, G. Tononi and C. Cirelli (2010). "Sleep homeostasis in the rat is preserved during chronic sleep restriction." Proc Natl Acad Sci U S A 107(36): 15939–15944.

      Lopes-dos-Santos, V., S. Ribeiro and A. B. Tort (2013). "Detecting cell assemblies in large neuronal populations." J Neurosci Methods 220(2): 149–166.

      Mizuseki, K., K. Diba, E. Pastalkova and G. Buzsaki (2011). "Hippocampal CA1 pyramidal cells form functionally distinct sublayers." Nat Neurosci 14(9): 1174–1181.

      Nitsche, M. A., M. Jakoubkova, N. Thirugnanasambandam, L. Schmalfuss, S. Hullemann, K. Sonka, W. Paulus, C. Trenkwalder and S. Happe (2010). "Contribution of the premotor cortex to consolidation of motor sequence learning in humans during sleep." J Neurophysiol 104(5): 2603–2614.

      Peyrache, A., M. Khamassi, K. Benchenane, S. I. Wiener and F. P. Battaglia (2009). "Replay of rule-learning related neural patterns in the prefrontal cortex during sleep." Nat Neurosci 12(7): 919–926.

      Plitt, M. H. and L. M. Giocomo (2021). "Experience-dependent contextual codes in the hippocampus." Nat Neurosci 24(5): 705–714.

      Proskurin, M., M. Manakov and A. Karpova (2023). "ACC neural ensemble dynamics are structured by strategy prevalence." Elife 12.

      Rothschild, G., E. Eban and L. M. Frank (2017). "A cortical-hippocampal-cortical loop of information processing during memory consolidation." Nat Neurosci 20(2): 251–259.

      Shin, J. D. and S. P. Jadhav (2024). "Prefrontal cortical ripples mediate top-down suppression of hippocampal reactivation during sleep memory consolidation." Curr Biol 34(13): 2801–2811 e2809.

      Shin, J. D., W. Tang and S. P. Jadhav (2019). "Dynamics of Awake Hippocampal-Prefrontal Replay for Spatial Learning and Memory-Guided Decision Making." Neuron 104(6): 1110–1125 e1117.

      Siapas, A. G. and M. A. Wilson (1998). "Coordinated interactions between hippocampal ripples and cortical spindles during slow-wave sleep." Neuron 21(5): 1123–1128.

      Simor, P., G. van der Wijk, L. Nobili and P. Peigneux (2020). "The microstructure of REM sleep: Why phasic and tonic?" Sleep Med Rev 52: 101305.

      Singer, A. C. and L. M. Frank (2009). "Rewarded outcomes enhance reactivation of experience in the hippocampus." Neuron 64(6): 910–921.

      Sosa, M., H. R. Joo and L. M. Frank (2020). "Dorsal and Ventral Hippocampal Sharp-Wave Ripples Activate Distinct Nucleus Accumbens Networks." Neuron 105(4): 725–741 e728.

      Sosa, M., M. H. Plitt and L. M. Giocomo (2025). "A flexible hippocampal population code for experience relative to reward." Nat Neurosci 28(7): 1497–1509.

      Stark, E., L. Roux, R. Eichler and G. Buzsaki (2015). "Local generation of multineuronal spike sequences in the hippocampal CA1 region." Proc Natl Acad Sci U S A 112(33): 10521–10526.

      Tang, W., J. D. Shin, L. M. Frank and S. P. Jadhav (2017). "Hippocampal-Prefrontal Reactivation during Learning Is Stronger in Awake Compared with Sleep States." J Neurosci 37(49): 11789–11805.

      Tort, A. B., R. Scheffer-Teixeira, B. C. Souza, A. Draguhn and J. Brankack (2013). "Theta-associated high-frequency oscillations (110-160Hz) in the hippocampus and neocortex." Prog Neurobiol 100: 1–14.

      Valero, M., T. J. Viney, R. Machold, S. Mederos, I. Zutshi, B. Schuman, Y. Senzai, B. Rudy and G. Buzsaki (2021). "Sleep down state-active ID2/Nkx2.1 interneurons in the neocortex." Nat Neurosci 24(3): 401–411.

      van de Ven, G. M., S. Trouche, C. G. McNamara, K. Allen and D. Dupret (2016). "Hippocampal Offline Reactivation Consolidates Recently Formed Cell Assembly Patterns during Sharp Wave-Ripples." Neuron 92(5): 968–974.

      van der Helm, E. and M. P. Walker (2011). "Sleep and Emotional Memory Processing." Sleep Med Clin 6(1): 31–43.

      Vaz, A. P., S. K. Inati, N. Brunel and K. A. Zaghloul (2019). "Coupled ripple oscillations between the medial temporal lobe and neocortex retrieve human memory." Science 363(6430): 975–978.

      Wilson, M. A. and B. L. McNaughton (1994). "Reactivation of hippocampal ensemble memories during sleep." Science 265(5172): 676–679.

      Yang, S. R., H. Sun, Z. L. Huang, M. H. Yao and W. M. Qu (2012). "Repeated sleep restriction in adolescent rats altered sleep patterns and impaired spatial learning/memory ability." Sleep 35(6): 849–859.

      Yu, J. Y., D. F. Liu, A. Loback, I. Grossrubatscher and L. M. Frank (2018). "Specific hippocampal representations are linked to generalized cortical representations in memory." Nat Commun 9(1): 2209.

      Zhang, L. B., J. Zhang, M. J. Sun, H. Chen, J. Yan, F. L. Luo, Z. X. Yao, Y. M. Wu and B. Hu (2020). "Neuronal Activity in the Cerebellum During the Sleep-Wakefulness Transition in Mice." Neurosci Bull 36(8): 919– 931.

    1. Author response:

      We read the Assessment as identifying two decisive gaps: (i) claims of group differences are supported by within-group tests rather than by a test of the group X condition interaction, and (ii) the drift-diffusion modeling is not yet validated or fit in a framework that permits group comparison. We agree with both, and we do not defend the current versions of these analyses. The revision will rebuild them rather than supplement them. We also agree with the reviewers that several of our conclusions about “internal beliefs” are currently stated more strongly than the analyses support, and these will be scaled back to what the modeling can carry.

      Below we first note factual errors and internal inconsistencies in the manuscript that the editors asked us to flag promptly, then summarize the planned revisions, then respond to each public review comment in turn.

      (1) Corrections and clarifications for the record

      Reviewer #2 identified one outright error in our text, and in re-checking the manuscript we found several further inconsistencies. We list them here so that they are on the record alongside the first version of the Reviewed Preprint. In each case the reviewers' reading of the manuscript is correct and the manuscript is at fault.

      (a) Aggressive investors and pupil modulation (Reviewer #2, 4.3). Our Results text states that “aggressive investors showed no significant pupil modulation by ambiguity,”. This is incorrect and contradicts our own Fig. 2d, which reports a significant, ambiguous vs. non-ambiguous difference in all three groups, including aggressive investors (T(31) = 2.74, P = 0.0304, Bonferroni-corrected). The figure and its statistics are correct; the text is wrong. The interpretive claim built on it - that aggressive investors show “blunted” arousal - is therefore unsupported and will be removed. The revised manuscript will state the correct within-group result and will test group differences directly rather than by contrasting significant against non-significant within-group tests.

      (b) Within-group versus between-group claims. Relatedly, our Results state that ambiguity-related pupil differences under the subjective model are “no longer significant relative to zero,” whereas the Discussion describes them as no longer differing across groups. These are different claims, and only the first was tested. This conflation runs through several of our physiological conclusions and is the same problem the Assessment identifies. It will be resolved by replacing these statements with explicit between-group and interaction tests.

      (c) Sample description (Reviewer #1, 5). The numbers are not contradictory but are certainly underspecified, and we accept that as written they cannot be reconciled by a reader. To state them plainly: 57 unique individuals each completed one to three sessions, yielding 108 sessions; 7 sessions were incomplete and 5 showed no response variability, leaving 96 sessions with usable behavior; of these, 95 had usable pupil data and 79 usable EEG. The three strategy groups are tertiles of these 96 sessions (32 each), and the smaller Ns in Figs. 2-4 (N = 31, 29, 25) reflect modality-specific exclusions within each tertile. The revision will include a participant/session flow diagram and a table reporting how many individuals contributed one, two, or three sessions, together with per-figure Ns.

      (d) DDM fitting procedure (Reviewer #1, 11). One clarification: the DDMs were fit separately for each participant, not separately per group (Methods, Section 4.6); group comparisons were then performed on the participant-level parameter estimates. We note this only for accuracy of the record. The reviewer's substantive objection is unaffected and we accept it: comparing parameters estimated in independent per-participant fits does not constitute a test of group differences within a common statistical framework, and it cannot test interactions.

      (e) Inconsistent specification of the EEG regressor. The trial-wise EEG covariate is described in one paragraph of Methods 4.6 as 8-13 Hz power averaged over 0-0.5 s, and in the text following Eq. 3 as a 1315 Hz difference over 0.1-0.4 s. The former corresponds to the analysis actually performed. We will correct Eq. 3's description and report the frequency band and window once, unambiguously.

      (f) Time-locking and figure/caption errors. As Reviewer #1 notes (comment 14), Methods describe stimulus-locked epochs (-0.25 to 1 s) while Figs. 2-3 are labeled relative to decision onset over -1 to 2 s. We will state for each analysis whether it is stimulus- or response-locked and harmonize axis labels accordingly. In addition: the Fig. 3 caption reads "N = 25 for ideal and aggressive; N = 29 for aggressive," where the second instance should read conservative; and Section 2.6 cites Fig. 1b and 1c for the leadership and team-performance results, which are Fig. 4e and 4f.

      (g) Delta/theta claims in the Discussion. Our Discussion attributes early delta- and theta-band enhancements to ideal investors. No such effects appear in our own time-frequency results, which report frontal beta-range and parietal alpha/low-beta clusters. These Discussion statements are not supported by the data presented and will be deleted. This also bears on Reviewer #1's comment 13, since it removes the only interpretation that depended on the low-frequency edge of our filter passband.

      (2) Summary of planned revisions

      (2.1) A single statistical framework with explicit interaction tests. All group comparisons will be replaced by unified models that include group, condition, and their interaction, with sessions nested within participants. Choice will be modeled with a generalized linear mixed model of the form invest ~ ambiguity ⨉ group + trial + (1 + ambiguity | participant/session); decision time with the corresponding linear mixed model including random slopes for ambiguity (Reviewer #1, 15). Time-resolved pupil and time-frequency EEG effects will be evaluated using cluster-based permutation tests on the group-by-condition interaction statistic rather than by aggregating within-group tests. Where our claim is that an effect is absent, most importantly, that ambiguity-related pupil differences vanish under subjective valuation, we will support it with equivalence testing and Bayes factors rather than with a non-significant P value, since a null result is not evidence for the null.

      (2.2) Continuous analyses as primary; grouping justified or abandoned. We accept that trichotomizing a continuous decision tendency requires justification that we did not provide. The revision will (i) define a continuous EV-consistency index and re-run every key analysis with it as a continuous predictor, establishing that the principal conclusions do not depend on the split. The “ideal” label implies a normative optimality we have not demonstrated and will be replaced with descriptive labels (EV-consistent, EV-exceeding, EV-shortfall).

      (2.3) Full specification, validation, and reliability of the belief parameter k. We agree the current description of k is inadequate. The revision will give the generative choice model in full, and explicitly state the estimation procedure and parameter bounds.

      (2.4) Rebuilt and validated drift-diffusion modeling. The boundary will be freed and estimated per participant/session, so that variance is no longer forced into the drift term. We will report whether the group effect on baseline drift survives.

      (2.5) Report why repeated sessions are treated as independent samples. We will conduct additional behavioral analyses to show repeated sessions from the same individual are sufficiently independent to be analyzed as unique samples. In particular, we will examine whether within-participant similarity across sessions is greater than cross-participant similarity. These analyses will provide direct evidence for whether sessions can reasonably be treated as separate samples rather than requiring sessions to be nested within participants.

      (2.6) Sharper framing and appropriately scaled claims. We will remove the framing that positions the field as treating ambiguity aversion as uniform and instead situate the work within literature that already documents heterogeneity in ambiguity attitudes. The distinction between ambiguity aversion and the internal representation of ambiguity will be made explicit, with the latter identified as our actual question. Claims about “internal belief models” will be restated in terms of what is estimated. Physiological hypotheses will be stated as specific, directional, falsifiable predictions rather than by appeal to the broad range of processes EEG has been linked to.

      (2.7) Quality control for pupillometry, EEG, and the VR context. We will report blink rates and counts, proportion of interpolated samples, trials and channels rejected, and ICA components removed. We will also report validate the luminance GLM.

      (2.8) The collaborative task. The Apollo Distributed Control Task preceded the Lottery Choice Task by design, so it must be reported as part of the protocol regardless of the leadership analysis. However, we agree that the leadership and team-performance analyses are not sufficiently motivated to carry the interpretive weight currently given them. They will be moved to a clearly labeled exploratory section, removed from the Abstract and framing, and presented without causal or trait-level interpretation.

      (3) Timeline and next steps

      The revisions above require refitting the behavioral, physiological, and computational analyses rather than adding to them, including hierarchical model fitting, parameter recovery, and new quality-control analyses. We therefore anticipate submitting the revised manuscript within approximately three months, and would welcome guidance if the editors would prefer a different schedule. We are content for the first version of the Reviewed Preprint to be published with this provisional response attached.

      We are grateful to both reviewers for the time invested in this manuscript. Several of the problems they identify are ones we should have caught ourselves, and the paper will be considerably stronger for their having been raised now rather than after publication.

    1. Author response:

      General Statements.

      We thank all reviewers for their careful evaluation and constructive comments.

      Fidel Serrano recently completed a related study on the role of CK1δ in the circadian clock which is published in eLife (https://doi.org/10.7554/eLife.110786.1). During these studies, we became interested in the physiological role of CK1δ autophosphorylation, whose functional significance has remained unclear.

      In the present manuscript, Fidel discovered that the autophosphorylated, auto-inhibited form of CK1δ accumulates specifically during mitosis. These findings provide a physiological context for CK1δ autophosphorylation that has remained elusive for many years. They constitute the central foundation of our model, which further integrates previous findings on APC/C function and activity with the data presented in our eLife study.

      Several suggestions and questions raised by the reviewers concern the regulation of CK1δ in cultured cells, which are predominantly in G1. Many of the corresponding experiments and analyses are already included in our eLife paper. We apologize that this overlap may complicate the review process. However, we are unable to publish identical datasets in both manuscripts.

      We therefore briefly summarize here the findings from the eLife study that are most relevant to the present work.

      - Overexpressed CK1δ is unstable, whereas CK1δ-K38R (catalytically inactive) and CK1δ-R178Q (altered specificity/activity) are comparatively stable (eLife Figs. 3A-E and 5B).

      - Assembly with PER2 stabilizes overexpressed CK1δ (eLife Figs. 4A-D).

      - Treatment with PF670462 similarly stabilizes the kinase (eLife Figs. 3F and 5D).

      - CK1δ binding sites in the centrosomal/Golgi area are present in excess even relative to overexpressed kinase (eLife Fig. 5C).

      The reviewers also raised questions about the PER2-CRY1 nuclear foci.

      Thermodynamically stable/persisting nuclear foci form upon coexpression of PER2 and CRY1. They were preliminarily characterized in the eLife study (eLife Fig. 2). For the purposes of the present work, they constitute a serendipitous and convenient tool for studying interactions of PER with CK1δ. To induce foci formation, we used a stable cell line expressing DOX-inducible mK2-CRY1. mK2-CRY1 accumulates only at relatively low levels because the protein without a binding partner is intrinsically unstable, as confirmed by Western blot analysis (eLife Fig. 4C). Coexpression of PER2 stabilizes mK2-CRY1 and promotes the formation of nuclear PER2-mK2-CRY1 foci.

      These foci contain elevated levels of endogenous CK1δ (eLife Fig. 2E and EV2), which accumulate gradually over a 24-hour period. This observation indicates that, at steady state, a fraction of endogenous CK1δ is continuously degraded in the absence of overexpressed PER2. Because endogenous CK1δ is synthesized at a relatively low rate, the unstable pool is small at any given time and therefore not readily detectable in conventional cycloheximide chase assays against the kinetically stabilized steady-state background of CK1δ.

      In our point-by-point response, we therefore refer the reviewers to the corresponding datasets and analyses presented in the eLife manuscript.

      Point-by-point description of the revisions.

      Reviewer #1 (Evidence, reproducibility and clarity):

      Summary:

      Involved in various cellular pathways and processes, Ser/Thr kinase CK1δ is thought to be constitutively active and its tail autophosphorylation is suggested as a putative inhibitory mechanism of CK1δ kinase activity. Here, the authors investigate CK1δ's phosphorylation status, in relation to its location and dynamics throughout the cell cycle. Immunofluorescence and biochemistry studies showed that the subcellular distribution of CK1δ is dynamic and in equilibrium between centrosomal and nuclear pools. The authors argue that CK1δ phosphorylation protects it from degradation. Finally, the authors show differences in CK1δ phosphorylation status and location throughout the cell cycle. Combining their findings with the existing literature, they propose a model of CK1δ phosphorylation status, location and dynamics throughout the cell cycle.

      Major comments:

      (1) Shortcomings in Immunofluorescence experiments:

      a. Lack of a positive control for centrosomal staining, eg. PLK1 (Fig.1A,C, 2B, 6A-B). This would also allow a colocalisation analyses between CK1δ and a centrosomal marker, which would build a more convincing body of evidence in favour of centrosomal recruitment of CK1δ in baseline conditions. A lack of a negative control for CK1δ IF? Most CK1 antibodies also cross-react with other isoforms.

      As suggested, we show centrosomal staining with a well-studied centrosomal marker pericentrin (PCNT; revised Fig. 1). Consistent with what has been well-described in the literature that CK1δ localizes to the centrosome (Sillibourne et al., 2002; Greer & Rubin, 2011; Greer et al., 2014) and with our observations, we confirm that CK1δ is indeed localized at the pericentral region.

      The CK1δ antibody does not crossreact, this is shown in Fig EV3A of the eLife paper.

      b. Lack of proper image quantification (Fig.1A,C, 2B, 6A-B). To establish a centrosomal recruitment of the CK1δ pools, colocalisation of CK1δ and a centrosomal marker must be quantified using a standard colocalisation quantification method and shown appropriately. Same comments regarding the colocalisation of CK1δ with PER2/mK2-CRY1-positive foci.

      Fig. 1: Colocalization analysis with pericentrin (PCNT) as centrosomal marker is provided.

      Fig. 2: The colocalization analysis of CK1δ (green) and mK-CRY1 (magenta) shows that in all PER2-untransfected cells (50 cells evaluated), CK1δ is concentrated at a single, mK2-CRY-negative spot, corresponding to the pericentrosomal region, which we have established in Figure 1. In contrast, in all PER2-transfected cells (15 cells evaluated) CK1δ localizes to mK2-CRY1 positive nuclear foci. The fields-of-view we have provided are representative of the typical phenotype of CK1δ localization in the presence or absence of PER2/mK2-CRY1-positive foci. Here it should be noted that CRY1 foci formation is strictly dependent of co-expression with PER2, (eLife paper).

      Fig. 6: This figure presents an analysis of fixed cells. Because only a small fraction of asynchronously growing cells is in mitosis at any given time, the number of cells assigned to individual mitotic stages is necessarily low (typically fewer than 10 cells per stage). The purpose of this figure is therefore primarily illustrative: to document and confirm the known subcellular localization of CK1δ during the different stages of mitosis rather than to provide a comprehensive quantitative analysis.

      c. Lack of proper statistical analyses (Fig.1, 2, 6). As immunofluorescence constrains one to only show few images per condition at best, statistical analyses on broader image analysis data are essential to measure the significance of the changes observed on the representative images provided.

      See answer to previous question.

      d. Lack of biological replicate numbers (Fig.1, 2, 6). Was it from 3 independent biological replicates?

      Fig.1 and 2: three independent biological replicates.

      Fig.6 is just illustrative to show the already known localization of CK1δ across mitosis. The localization shown in the four panels is seen in all cells at the respective cell cycle stages.

      e. NOTE: if available, using a confocal microscope would be best to provide optimal Z-axis resolution. This would provide you with more accurate colocalisation data, like CK1δ recruitment to the centrosome or PER2/mK2-CRY1-positive foci.

      Characterization of PER-CRY foci is published in the eLife paper in Fig. 2 and EV2.

      (2) Shortcomings in biochemistry experiments

      Lack of loading controls in Fig.3A-C and 4A-B. Until they are added, no conclusions can be safely interpreted from the experiments. Myc is a good turnover control for CHX across all experiments. Enrichment of phospho-proteins upon CalA treatment?

      Loading controls are provided in this revision. Enrichment of phospho-proteins following CalA treatment has been demonstrated and published in numerous previous studies (Cegielska et al., 1998; Ishihara et al., 1989; Rivers et al., 1998).

      (3) The authors infer from the literature that overexpressed CK1δ is unassembled, without checking it in any experiment. It could be that CK1δ overexpression drives overexpression of its binding partners, in which case most of overexpressed kinase could actually be assembled. Since CK1δ assembly is of great importance to the study's conclusions, it should be confirmed experimentally.

      In our eLife study, we show that overexpressed CK1δ is not fully assembled with stabilizing binding partners such as PER2. In fact, centrosomal/Golgi binding partners remain always available in excess even relative to overexpressed CK1δ, as demonstrated by the MG132-induced increase in centrosomal accumulation (eLife Fig. 5C). However, binding is dynamic and the affinities and concentrations are such that a substantial fraction of overexpressed CK1δ remains unbound and is therefore subject to degradation.

      (4) The authors infer from the literature that phosphorylated CK1δ is inactive - which is not a given. All Western blotting experiments should include confirmed CK1δ substrates like PER2 or DVL3 to confirm that phosphorylated CK1δ is indeed inhibited. As added benefit, blotting for known CK1δ substrates will act as confirmation that the FLAG-tag on overexpressed CK1δ/ε does not impact its kinase function, and that CK1δ/ε inhibitor treatments like PHF670 have indeed worked. Furthermore, because CalA and OA inhibit many phosphatases, inferences on CK1δ/ε activity upon these treatments should be taken catiously.

      The activity of phosphorylated CK1δ is severely attenuated by autoinhibition, although the kinase is not completely inactive. This has been demonstrated repeatedly in the literature and does not require re-establishment in the present study. We previously showed this directly in a PNAS study by Marzoll et al. (https://doi.org/10.1073/pnas.2118286119). At some point, one must rely on established published findings unless there is compelling evidence supporting an alternative interpretation.

      (5) Figure 1 requires further biochemical analyses to support immunofluorescence data. Western blotting analyses showing fluctuating CK1δ levels in nuclear vs. Centrosomal (https://pmc.ncbi.nlm.nih.gov/articles/PMC7618310/) fractions in the different treatment conditions would help illustrate the redistribution of CK1δ pools between nuclear and centrosomal areas. Additionally, adding a cytoplasmic fraction in each treatment condition would help visualise the amounts of unrecruited CK1δ in overexpressed conditions.

      Microscopy is a well-established and widely accepted approach for assessing subcellular localization. In many cases, it is superior to biochemical fractionation assays, which rarely yield completely pure nuclear or centrosomal fractions. The fractionation experiments suggested by the reviewer could provide additional supportive evidence, but they are not required to substantiate the conclusions presented here.

      (6) Figure 2 requires further biochemical analyses to support immunofluorescence data showing colocalisation of CK1δ with PER2, using for example immunoprecipitation to show that pulling down PER2 also pulls down CK1δ, and vice versa. Optimally, an additional Western blotting experiment showing total / phospho-PER2 and CK1δ levels for each condition in cytoplasmic vs. nuclear fractions would consolidate the evidence from immunofluorescence and immunoprecipitation.

      Association of CK1δ with PER2 has been demonstrated extensively in numerous previous studies (Aryal et al., 2017; Narasimamurthy et al., 2018; Philpott et al., 2020; Cao et al., 2021; Marzoll et al., 2022; An et al., 2022) including our own eLife publication.

      (7) The authors consistently interpret data showing CK1δ phosphorylation as CK1δ tail phosphorylation (p.8-11). A band shift in CK1δ signal on a Western blot does not in any way show the tail specifically is phosphorylated. Indeed, although CK1δ tail phosphorylation is widely recognised, several identified sites on other CK1δ domains, e.g.: the kinase domain, can be phosphorylated by CK1δ and other kinases. The authors' conclusions must therefore not mention tail-specific phosphorylation unless domain-specific phosphorylation is established. This could be done by comparing signal from antibodies specifically recognising known phospho-sites on the CK1δ tail against total CK1δ signal in experiments of Fig.3, 4, and 5.

      The electrophoretic mobility shift is caused by phosphorylation of the CK1δ tail. This has been demonstrated in numerous previous studies. For example, we generated a CK1δ variant in which all serine and threonine residues in the C-terminal tail were replaced by alanines. In this mutant, CK1δA, inhibition of phosphatases by CalA no longer induces an electrophoretic shift (Figure 3B).

      (8) The use of anti-FLAG antibody in all figures looking at overexpressed CK1δ is not the most appropriate choice, as the anti-CK1δ antibody works perfectly for Western blotting and immunofluorescence studies. This adds more variables and hinders comparisons with endogenous CK1δ. If CK1δ is successfully overexpressed, wild-type levels should be negligible compared to overexpressed CK1δ levels, so the need for a FLAG antibody is not justified. Using an anti-CK1δ antibody in all experiments instead of anti-FLAG would confer them greater solidity.

      The amount of endogenous CK1δ is limited, and we therefore use endogenous detection only when it is scientifically necessary. FLAG-tagged constructs, in contrast, provide a robust and convenient experimental system. In the absence of evidence that the FLAG tag introduces artifacts or alters the observed behavior of the kinase, we do not consider it necessary to repeat these experiments using endogenous detection alone.

      (9) The authors base themselves off a correlation between CK1δ phosphorylation status and its observed stability to establish a causal relationship between both ('phosphorylation protects the overexpressed kinase from degradation' p. 7; 'tail phosphorylation protects the overexpressed kinase from rapid degradation' p.9). However, no experiments performed suggest there is a causal relationship occurring. This could be done by looking at specific known CK1δ tail phospho-sites using phosphosite-specific antibodies. The authors could assess whether CK1δ stability is impacted when overexpressing wild-type vs. phospho-dead mutant CK1δ.

      In the eLife study, we show that inhibition of kinase activity by PF670462 stabilizes CK1δ and that the overexpressed kinase-dead mutant CK1δ-K38R is stable.

      (10) The authors consistently link APC/CCDH1 with CK1δ's putative degradation, when none of their study touched on APC/CCDH1 activity or its involvement in CK1δ dynamics. To support these conclusions, fig.5 requires further biochemistry studies. To show APC/CCDH1's interaction with CK1δ, I suggest the authors (1) pull down CK1δ by immunoprecipitation and check for APC/CCDH1, phosphorylated CK1δ, and ubiquitin levels in enriched samples for each cell cycle stage in untreated vs. MG132-treated conditions. To consolidate this, the authors could (2) pull down CDH1 and check for phosphorylated and total CK1δ levels in pulled-down samples. Finally, they should investigate whether reducing CDH1 (3) expression (siRNA knock-down) and (4) CDH1 activity (specific E3 ligase inhibitor) have any impact on CK1δ levels at each cell cycle stage.

      These data were previously published in Penas et al. (doi: 10.1016/j.celrep.2015.03.016). For example, Fig. 5C shows that siRNA-mediated depletion of Cdh1, the G1-specific cofactor of APC/C, stabilizes overexpressed CK1δ, as well as the established APC/C-Cdh1 substrate Cyclin B1.

      (11) Methods section lacks a subsection detailing image acquisition, including the type of microscope used (confocal vs. widefield), magnification used, whether acquisition settings were kept consistent throughout conditions / technical/biological replicates.

      Is provided in the revised manuscript.

      (12) Methods section must include a subsection detailing immunofluorescence image analysis parameters and statistical analyses, e.g.: criteria for centrosomal localisation categories in Fig.1, criteria for categorising cells in different stages of mitosis (Fig.6), overall threshold / criteria stringency.

      Is provided in the revised manuscript.

      Minor comments:

      The Reviewer has raised about 50 minor points, counting the individual remarks and sub-point and sub-sub points. We have addressed a number of these comments where they were scientifically relevant or helpful for improving clarity. Most of the questions raised concern published and generally accepted data. We therefore do not believe that a point-by-point response to every individual minor remark is constructive or necessary.

      (1) PF670 was shown to selectively inhibit CK1δ/ε over 42 common kinases (TOCRIS), however it is not well-characterised regarding the remaining 474 kinases encoded by the human genome. There is thus a possibility for PF670 inhibition overlap between CK1δ/ε and other kinases, which should be mentioned, and the authors' conclusions should he more nuanced as a result.

      (2) CHX is a protein synthesis inhibitor, which means it does not selectively target CK1δ expression, but that of every protein within the cell. This should be addressed and controlled for, if possible. A potential way to go about it would be to knock-down CK1δ using siRNA and compare untransfected controls with select post-transfection timepoints to assess the impact of inhibiting CK1δ expression on total CK1δ levels.

      (3) All unshown Western blotting replicates should be included in the supplementary materials.

      (4) Figure 1:

      a. A-B: experiment is missing PF670 alone and CHX alone conditions to control for direct effects of either compound, versus combined.

      b. A-B: in the text, please address the fact that >50% cells show unclear or no centrosomal pattern in untreated conditions.

      c. C: the authors claim that CK1δ-FLAG levels are decreased in PF670-treated conditions because PF670 inhibits excess CK1δ autophosphorylation, inducing its subsequent degradation. Before this claim is made, a proteasome inhibitor like MG132 should be included within the experiment, to show that this decrease in CK1δ expression under PF670 treatment can be rescued with MG132 treatment. This would plead in favour of excess CK1δ degradation and exclude the possibility that 4h of PF670 treatment simply reduces the rate of CK1δ expression. Should you wish to be more specific and confirm that APC/CCDH1 is responsible for CK1δ degradation, using a CDH1-specific inhibitor or siRNA-mediated knock-down of CDH1 could confirm that CK1δ degradation occurs via APC/CCDH1-mediated ubiquitination of CK1δ.

      d. C-D: experiment is missing pre-DOX induction control to confirm that CK1δ-FLAG is indeed being overexpressed.

      (5). Figure 2:

      a. A: should include non-transfected control panels to confirm the success of PER2 overexpression.

      b. B: Although observed in a preprint from the same team, there is no peer-reviewed evidence that establishes mK2-CRY1 as a reliable indicator of PER2 overexpression. 2A shows a correlation between both but does not exclude the fact that PER2 must be included in the imaging of the 2B panels, instead of using mK2-CRY1 as a proxy readout of PER2 expression and location. Appropriate analyses would then be required to show colocalisation of CK1δ with PER2 in the highlighted puncta.

      c. 'PER2-dependent foci': cannot be said of the data unless PER2 dependency has been validated in those images. Please see above point to resolve this.

      (6) Figure 3:

      a. B: the use of kinase-dead CK1δ mutant does not allow to fully separate direct autophosphorylation from phosphorylation by other kinases, unlike what the authors mention: CK1δ kinase activity may be required for the phosphorylation of certain sites by other kinases. In this case, the decrease of phosphorylation in the kinase-dead CK1δ mutant would not only result from inhibited autophosphorylation, but also from reduced phosphorylation by other kinases. Please adjust your conclusions accordingly (p.8).

      b. C: poor visualisation of CK1δ overexpressed condition, especially showing critically reduced signal at the 6min CHX timepoint, compared to its kinase-dead homolog. If this change is present in every replicate performed, please address it in the text. If not, perhaps you may have to display another representative replicate in the figure.

      c. The authors claim that the increased levels of overexpressed with CalA treatment suggest 'that full or partial phosphorylation of the CK1δ tail stabilizes both active and inactive forms of the kinase' (p.8). Importantly, tail phosphorylation was never shown, so this should be corrected. Additionally, it could be that CalA treatment considerably increases the rate of CK1δ expression - which would also match data in fig.4A, since CalA and CalA+PF670 treatments alone drastically increased CK1δ levels. This should be checked by pre-treating cells with CHX before applying CalN, or by treating cells with CHX and CalN simultaneously, and interpreted accordingly.

      (7) Figure 4:

      a. A: CK1δ levels in CHX-free controls from both 1h pre-treatments look much higher compared to untreated controls, which the authors interpret as an indication that 'most of the newly synthetised overexpressed kinase was degraded in untreated cells' (p.9). However other explanations are not explored: since loading controls are not provided, it may be that sample loading in the gel is simply off. Importantly, it is also possible that pre-treatments increased CK1δ expression before CHX application. Please make sure you touch on each

      b. A-B: the authors claim that CK1δ-FLAG and CK1ε-FLAG levels are decreasing with CHX treatment because they are being degraded. Adding a panel with a proteasome inhibitor like MG132 would solidify this argument. Rescue of CK1δ/ε degradation under CHX treatment would show that the loss is indeed mediated by the UPS. Should you wish to be more specific and confirm that APC/CCDH1 is responsible for CK1δ degradation, using a CDH1-specific inhibitor or siRNA-mediated knock-down of CDH1 could confirm that CK1δ/ε degradation occurs via its ubiquitination by APC/CCDH1.

      c. B, D: blot in B does not match the quantification trends in D. E.g.: 60min CHX + CalA + PF670 condition shows clearly lower CK1ε signal compared to its 0min CHX control. Please ensure the biological replicate you display on the figure is indeed representative of your results.

      d. A, C: 'Hyperphosphorylated CK1δ remained stable throughout the CHX chase [...], indicating that tail phosphorylation protects the overexpressed kinase from rapid degradation' (p.9). Meanwhile this is true, unphosphorylated CK1δ in the CalA+PF670 treatment condition was also stabilised, showing that CK1δ phosphorylation may not be required for kinase stabilisation. This is an important point and should be addressed in the data interpretation. On another note - and as mentioned above -, tail phosphorylation specifically is not shown and cannot be inferred unless domain-specific phosphorylation is investigated.

      e. B, D: CK1ε-FLAG levels decrease with CHX treatment compared to its baseline in the CalA+PF670 condition, which is not the case for CK1δ-FLAG (A). Thus, the data shown does not support the authors' conclusions 'CK1ε turnover is regulated in a similar manner to CK1δ'. Please address this issue and adjust your conclusions accordingly.

      f. C-D: performing statistical analyses on protein band intensity in different conditions would be interesting to establish the significance of those changes.

      (8) Figure 5

      a. B: non-arrested controls mentioned in the text are missing from the figure.

      b. B: 'the accumulated CK1δ remained dephosphorylated' (p.10). The experiment is missing important positive / negative controls of CK1δ phosphorylation status to conclude whether CK1δ remained phosphorylated or unphosphorylated across conditions. One or the other cannot be concluded from the blot as it is. Using phospho-specific antibodies may also help to visualise phosphorylation status.

      c. B, D-E: blots should include CDH1 phosphorylation levels (hyperphosphorylated CDH1 is inactive), CK1δ substrates (indicators of CK1δ activity), and relevant phosphatase substrates to create a cohesive picture of changes in mitosis and support fig.7.

      d. D-E: please address the differences in CK1δ profile between overexpressed and endogenous CK1δ in G2/M phase.

      (9) Figure 6

      The authors say: 'similar results were observed when using U2OStx cells and staining for endogenous CK1δ'. However, CK1δ's subcellular distribution pattern is different in endogenous vs. overexpressed conditions in the telophase-cytokinesis and post-mitosis stages. In the telophase-cytokinesis stage, endogenous CK1δ seem to form nuclear hotspots, while overexpressed CK1δ is more diffuse. In the post-mitosis stage, overexpressed CK1δ shows a clear polar pattern in the nuclear periphery, while endogenous CK1δ shows a diffuse pattern similar to that of overexpressed CK1δ described in fig.1 as 'unassembled' by the authors. Please address this and adjust your conclusions appropriately.

      (10) Figure 7

      a. The authors infer APC/CCDH1's activity levels or relationship to CK1δ from existing literature only. Since it is a central mechanism of their study, literature-only components are insufficient for a summary figure. To include these elements in the figure, the authors must include investigations of APC/CCDH1's activity levels and involvement in CK1δ degradation at different stages of the cell cycle in their study. Please refer to point 10.

      b. Similar comment for CK1δ assembly status and activity levels. Please refer to point 4.

      c. Similar comment for phosphatase activity levels in different stages of the cel cycle. By observing phosphorylation status in known substrates of established CK1δ phosphatases, one can easily confirm phosphatase activity levels.

      d. Does not consider the fact that some results were different in endogenous vs. overexpressed CK1δ models. Please nuance your claims.

      (11) Reference missing p.12 paragraph 1: 'Yet, overexpression of CK1δ consistently accelerates the circadian clock, implying that kinase availability can influence clock speed. This finding suggests that CK1δ activity may be regulated not only by catalytic mechanisms but also by its spatial availability within the cell'. The facts stated there are not covered in the study's findings and is not referenced with a published study.

      (12) Grammar mistakes / typos to report:

      p.3 paragraph2: 'the kinases undergo futile cycles of phosphorylation and dephosphorylation'

      p.26 Fig.1A legend: 'Endogenous CK1δ was detected by IF to'? Unfinished formulation

      p.8: 'PP1 was previously suggested as a major PPase of CK1δ/ε'

      p.8: 'fewer phosphorylation sites are targeted by other kinases'

      Figure 4C legend: 'Overexpressed unphosphorylated CK1δ [instead of CK1ε] is degraded with a half-life of about 15 min'

      Figure 4E: Ponceau staining

      (13) Nomenclature inconsistencies to report:

      p.4 paragraph2: 'protein PER2', then p.4 paragraph3 'PERIOD2'.

      'CK1δ' used throughout the article's body text, but 'CSNK1D' is used in IF panel legends. The gene name was never introduced in the main text, nor has it been explained in the figure legends, so perhaps go for CK1δ for all mentions, including in figures.

      'FOV' nomenclature is unclear, please define in the figure legend.

      Figure 2B: what are the arrowheads pointing to? - please clarify in the figure legend.

      Mislabelled figures: figure 3 in the text refers to figure 4 in the figure section, and figure 4 in the text to figure 3 in the figure section.

      Figure 6A-B: abbreviated mitosis stages 'Pro' and 'Meta-Ana' should either be defined in the figure legend or put in full writing within the figure.

      Reviewer #1 (Significance):

      The paper will appeal to those working on CK1 biology, including cell cycle and circadian rhythms.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary:

      In this study, authors aimed to address functional links between CK1 activation, subcellular localization and protein stability, which is an important biological question. This manuscript is a follow up of a recent study by the same team, which is currently deposited at Biorxiv and a fraction of data seems to overlap, which is somewhat confusing. Overall, the concept that the stability of CK1d is dynamically controlled across the cell cycle is interesting. On the other hand, how this is functionally connected to the circadian cycle described in the previous study remains unclear.

      The data presented in this study are highly preliminary and lack a number of essential controls, which weakens an otherwise interesting concept.

      Major issues:

      One of the main conclusions of the study is that CK1d is degraded predominantly in the nucleus while the centrosomal pool is protected from the degradation. This concept is interesting but unfortunately, the experimental evidence supporting this model is very limited. Authors previously showed that neither inhibition of proteasome or treatment of cells with PF670462 significantly influenced levels of endogenous CK1d but both treatments promoted accumulation of the tagged and overexpressed FLAG-CK1d. Absence of the phenotype at the level of endogenous protein clearly raises question whether this may be just an artifact of overexpression, tagging or both combined.

      As shown in our previously published eLife study, CK1δ is synthesized at a relatively low rate from its endogenous locus. Free CK1δ continuously shuttles between the cytosol and the nucleus. Although association of CK1δ with the centrosomal/Golgi area is dynamic, nuclear export followed by binding to centrosomal/Golgi structures constitutes the major sink for the kinase at steady state. Consequently, the fraction of unbound CK1δ that is targeted for degradation in the nucleus at any given time is very small and cannot be detected in cycloheximide chase assays against the much larger background of kinase stabilized by association with centrosomal/Golgi binding sites. To reveal that CK1δ is subject to degradation at all, we had to overexpress the kinase. Under these conditions, the fraction of unbound CK1δ increases substantially, allowing nuclear degradation to be detected experimentally.

      The statement that degradation occurs in the nucleus is based on weak data with CK1d-NES construct that was claimed to accumulate at higher levels compared to the wild type CK1d. However, that experiment used quantification of microscopic data in which two distinct regions (nucleus, cytosol) were compared, which is technically challenging. Including a reporter for normalizing for the transfection efficiency would strengthen conclusions of that experiment. Overall, quantification by immunoblotting may be more accurate.

      As suggested, we performed CHX time-course experiments with NES- and NLS-tagged CK1δ constructs followed by immunoblot analysis (shown in revised Fig. 4E).

      Specificity of the microscopy staining was not validated Fig. 1. Confirmation of the staining specificity by RNAi or KO approaches is essential. This antibody from Abcam has been discontinued which makes it impossible to reproduce this experiment.

      The authors should therefore attempt to demonstrate localization using other available antibodies against CK1d.

      The antibody is available through Thermo Fisher in the United Kingdom: https://www.fishersci.co.uk/shop/products/100-ul-mouse-monoclonal-af12g4-casein-kinase/13070202#

      It is specific for CK1δ and does not recognize CK1ε, as demonstrated in our eLife study.

      Furthermore, FLAG-tagged CK1δ detected with anti-FLAG antibodies, as well as endogenous CK1δ detected with the commercial antibody, localize to the pericentrosomal region and relocalize upon PER2 overexpression to mCRY1-containing nuclear foci. Together, we consider these data compelling evidence for the specificity of the antibody staining.

      The authors also failed to demonstrate that the dots represent centrosomes. To do so, they must perform co-staining with a robust centrosome marker.

      As requested, Pericentrin (PCNT) was used as a centrosomal marker and is now provided in the revised Fig. 1.

      This would help classify approximately 50% of the cases that are currently assessed as "potential" but are in fact inconclusive.

      Centrosomal colocalization with PCNT improved the robustness of the quantification.

      The quantification in the experiment is incorrect. Since the percentages are reported, both columns should add up to 100%. In Panel 1B, approximately 10% of the cells are missing, while in Panel 1D, there appear to be about 20% more cells.

      As indicated in our previous figure, the y-axis represents the number of cells, not percentages. We analyzed approximately 100 cells, which may have led to the misunderstanding that the values refer to percentages. As well, we have replaced this figure with a revised Figure 1 which shows that CK1δ co-localizes with PCNT as a centrosomal marker.

      Fig. 2B shows four different fields based on which authors come to conclusion that co-expression of CRY and PER2 promotes re-localisation of CK1d to nuclear foci. This is an interesting possibility but it is hard to conclude without any quantification. What was the fraction of cells that expressed CRY and PER2 that showed this phenotype? What was the fraction of cells that did not show this phenotype although both of these proteins were expressed?

      It would help to label cells expressing PER2 either by expressing it as a fusion protein or by co-expressing a marker protein ideally from the same plasmid.

      All cells in this stable cell line express mK2-CRY1. The cells were transiently transfected with PER2, and therefore only a fraction of the cells received the PER2 expression construct. In our eLife paper, we showed that all PER2-expressing cells stabilized mK2-CRY1 and formed nuclear foci. We demonstrate here that every cell containing such nuclear foci also showed accumulation of endogenous CK1δ within these structures. We did not observe a single cell with PER2-induced nuclear foci that lacked endogenous CK1δ accumulation.

      Cells that did not express PER2 were identified by the absence of nuclear foci and by low levels of mK2-CRY1, which was homogeneously distributed throughout the nucleus, as also shown in the eLife manuscript. In these cells, endogenous CK1δ was concentrated at a single discrete structure that, based on data shown in Fig. 1, corresponds to the pericentrosomal region. Under these conditions, neither CK1δ nor mK2-CRY1 was detected in nuclear foci.

      Fig. 5 suggests that massive phosphorylation of CK1d, which is responsible for its mobility shift on SDS-PAGE, is most likely linked with mitosis. Authors should use established markers to estimate a fraction of mitotic cells in their G2/M fraction. In principle, there are two possible explanations for the doublet observed with CK1d staining. Ether CK1d exists in two pools with different phosphorylation states in mitosis, or perhaps more likely, this fraction contains G2 cells where CK1d is not yet modified and mitotic cells where CK1d is fully phosphorylated. Performing a shake-off experiment yielding a pure fraction of mitotic cells could help to distinguish between these two options.

      As described in the main text and the methods section, the fractions were prepared by mitotic shake-off to further enrich our sample for rounded cells that are loosely attached during mitosis.

      Degradation of CK1d in telophase/cytokinesis when APC/Cdh1 becomes active is not apparent in Fig 6. The signal at mitotic spindle is missing, but there is still plenty of signal remaining in the cells. It is possible that the signal is just redistributed in the cell and the data shown do not support degradation of the protein. Authors could film cells expressing fluorescently labeled CK1d and quantify the signal during progression through mitosis and mitotic exit. The statement that "Following nuclear envelope reformation and mitotic exit, CK1d localized primarily to the single centrosome in each daughter cell" is incorrect. First, authors cannot deduce from the fixed cells whether they have just formed the nuclear envelope and exited mitosis.

      The reviewer is, of course, absolutely correct in the points raised.

      First, we cannot deduce from fixed cells whether they have only recently exited mitosis. This was not our intention. We merely selected cells in G1 and referred to them as “post-mitotic,” without intending to imply that these cells had just exited mitosis. To clarify this point, we changed the previous statement:

      “Following nuclear envelope reformation and mitotic exit, CK1δ localized primarily to the single centrosome in each daughter cell (Fig. 6A, 4th column)”

      to:

      “In G1, CK1δ localized primarily to the single centrosome (Fig. 6A, 4th column),” and replaced in column 4 of Fig. 6A and B the label “post-mitosis” with “G1.”

      Furthermore, we cannot deduce from fixed cells whether, or to what extent, CK1δ is degraded upon mitotic exit. This was neither the intention nor the conclusion drawn from Fig. 6. Rather, we show in Fig. 3 that phosphorylated CK1δ is not degraded, and in Fig. 5 that CK1δ is predominantly hyperphosphorylated during mitosis, leading us to conclude that this phosphorylated pool of CK1δ is stable. In G1, CK1δ is dephosphorylated and unassembled kinase is degraded. We currently have no data regarding the kinetics of CK1δ dephosphorylation, assembly with centrosomal/Golgi structures, versus degradation of unassembled dephosphorylated CK1δ.

      The data shown in Fig. 6 serve merely to illustrate the subcellular distribution of CK1δ, which is consistent with previous reports. In Fig. 7, we present a model that attempts to integrate the new findings reported here together with the data from our recent eLife paper and the broader body of knowledge regarding both CK1δ biology and cell-cycle regulation. Of course, we do not claim that this model does by no means represents a final verdict, and many important questions remain open. However, we believe that the model provides plausible novel concepts and mechanistic ideas that have not been proposed in previous publications and therefore merit publication, as they provide a basis for further investigation and discussion.

      Live-cell imaging:

      Live-cell imaging of fluorescently tagged CK1δ throughout mitotic progression and mitotic exit could, in principle, provide additional insight, but such experiments are technically extremely challenging. Moreover, the central idea of our model is that as much CK1δ as possible is preserved throughout the cell cycle, whereas degradation selectively targets unassembled and potentially harmful kinase. When CK1δ is expressed at physiological levels, which would require tagging the endogenous locus, the fraction of kinase degraded upon mitotic exit is expected to be very small and therefore likely below the threshold for reliable quantification by fluorescence microscopy. Similarly, although overexpressed CK1δ undergoes substantial degradation in G1, the kinase is simultaneously synthesized at a high rate. Hence, quantitative interpretation of overexpressed CK1δ levels during mitotic exit by microscopy (without CHX) would still be difficult.

      Second, the images of interphase cells constantly show multiple dots (probably surrounding the centrosome), which is a pattern that likely corresponds to Golgi rather than a single centrosome.

      The reviewer is correct. Indeed, CK1δ localization to the Golgi apparatus is well established in the literature. In our original wording, we did not explicitly distinguish between Golgi and centrosomal localization, which may have been somewhat misleading. In the revised version, we therefore refer more cautiously to the “pericentrosomal region” rather than strictly to the centrosome.

      The data shown in Fig. 6 serve to illustrate the known subcellular distribution of CK1δ. In the model presented in Fig. 7, we attempt to integrate the new findings reported here together with the data from our eLife paper and the broader body of knowledge regarding both CK1δ biology and cell-cycle regulation.

      It is unclear to which figure points the paragraph "Tail phosphorylation protects CK1δ/ε from degradation". I assume that one figure is missing.

      Figs. 3 and 4 were accidentally swapped, and we apologize for this error. The paragraph in question refers to Fig. 4, which shows the cycloheximide-induced degradation kinetics of CK1δ/ε.

      Minor points:

      CK1 kinase inhibitor PF670462 should not be named as PF670 as this causes confusion. Authors should either use the full name of the compound or just call it as CK1 inhibitor with providing details in the methods.

      PF670 has been changed to PF670462.

      Fig. 3B is discrepant with the figure legend. Figure shows CK1e but legend says kinase dead CK1D-K38R

      The captions to Figs. 3 and 4, as well as the references to these figures in the text, are correct. However, the actual Figs. 3 and 4 were inadvertently swapped during figure assembly. We apologize for this mix-up.

      The authors` interpretation of CK1 involvement in checkpoint is incorrect. The authors state that CK1 activity decreases p53 function promoting recovery, but Inuzuka et al (ref. 51) showed that inhibition of CK1 leads to this outcome.

      We thank the Reviewer for noting this mistake. We are no experts in p53 regulation, which is rather complex. CK1 decreases MDM2 stability and hence enhances p53 function.

      We corrected the statement and placed it in the right context: “CK1 phosphorylation triggers β-TrCP-mediated degradation of MDM2 and activates p53, thereby enhancing p53-dependent responses involved in checkpoint signaling and DNA repair (Inuzuka et al., 2010; Winter et al., 2004). After DNA repair, CK1δ has…”

      Reviewer #2 (Significance):

      It is generally assumed that CK1 is constitutively active, which is likely an oversimplified view; in a physiological context, some degree of regulation can be expected. Demonstrating that there are several pools of CK1 that are differently regulated at the level of protein stability during the cell cycle would be a significant advance in our understanding of CK1 functions.

      Reviewer #3 (Evidence, reproducibility and clarity):

      In this study, Serrano et al. employed a combination of cell biological and molecular approaches to investigate the localization and regulation of Casein Kinase CK1 during the cell cycle using U2OS cells. They show that CK1 dynamically localizes between the centrosomes and the nucleus but can be sequestered away from the centrosomes upon overexpression of its binding partner PER2. They provide evidence that CK1 strongly accumulates in a hyperphosphorylated form upon inhibition of phosphatases (using Calyculin), and thus conclude that CK1 tail phosphorylation protects the kinase from degradation. Using synchronized cells, they show that CK1 accumulates unphosphorylated in S-phase (APC/Cdh1 inactive) but phosphorylated at the G2-M transition. Immunostaining shows that CK1 localizes to the centrosomes during mitosis.

      Overall, they propose that the activity and abundance of CK1 are regulated during the cell cycle. However, this claim would require several experiments to support it.

      Major comments:

      - As presented, some of the data are inconclusive. Co-staining with a centrosomal marker is required to determine whether or not Ck1 localises to the centrosomes. A large proportion of cells exhibit "potential" (their term) centrosomal staining, so a centrosomal marker is essential before any conclusions can really be drawn.

      In our revised Figure 1, we confirm this centrosomal staining using an antibody against pericentrin (PCNT).

      - Figures 3 and 4 have no loading controls, and these two figures have been mixed up in the text.

      The reviewer is right, we have corrected the Fig. 3 and Fig. 4 mix-up.

      Loading controls are now also provided.

      - A mobility shift on SDS-PAGE does not prove that a protein is phosphorylated. The authors should provide experimental evidence that the mobility shift is really due to phosphorylation. As they are inactivating phosphatases using CalA, it is likely the case, but they should prove it. Furthermore, the authors did not map any phosphorylation sites in this study, so they do not know whether CK1 phosphorylation occurs in the tail (as they assert) or elsewhere.

      We provide data in Fig. 3B showing that a CK1δ variant in which all serine and threonine residues in the C-terminal tail were replaced by alanines does not undergo an electrophoretic mobility shift upon CalA treatment. These results demonstrate that the CalA-induced mobility shift is caused by phosphorylation of the CK1δ C-terminal tail.

      - The figure legends in general are limited and lack crucial information. For instance, in Figures 3C and 3D, how was the half-life of CK1 determined?

      We have adapted the figure caption:

      (C) Densitometric quantification of n=3 Western blots (see A) shown as mean ± SD. Overexpressed unphosphorylated CK1δ is degraded with a half-life of about 15 min. Both CalA and CalA + PF670462 treatments stabilize the kinase. (D) Densitometric quantification of n=3 Western blots (see B) shown as mean ± SD.

      - The CK1 regulatory model presented in Figure 7 is not supported by the data. What experimental evidence, for instance, shows that CK1 is inactive during mitosis? To make this claim the authors should directly assay its activity.

      We show that the majority of CK1δ is phosphorylated and therefore auto-inhibited during mitosis. In the original version, we referred to this fraction as inactive. In the revised manuscript, we refer to the phosphorylated kinase as auto-inhibited.

      Reviewer #3 (Significance):

      This study may be of interest to researchers working on cell cycle regulation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study identifies a non-canonical essential role for acyl carrier protein in maintaining apicoplast metabolism and blood-stage survival in Plasmodium falciparum. The main conclusions are largely supported by strong genetic and biochemical evidence, although some claims regarding the dispensability of fatty acid synthesis pathways remain incomplete. The work provides novel mechanistic insight into ACP-mediated stabilization of pyruvate kinase II and will be of broad interest to the malaria and apicoplast biology communities.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We kindly ask that the editorial assessment be revised to focus on the major conclusions. Alternatively, we would respectfully suggest revising the second sentence to something akin to “The main conclusions are largely supported by strong genetic and biochemical evidence, while function of fatty acid synthesis pathways in low-lipid conditions will require future studies to fully resolve.”

      Public Reviews:

      Reviewer #1 (Public review):

      This study provides evidence that the apicoplast-locaized isoform of acyl-carrier protein (ACP) has acquired important non-enzymatic functions in the malaria parasite. Previous studies have shown that the apicoplast-located FASII-dependent pathway of fatty acid synthesis is not essential in Plasmodium blood stages. In contrast, genome-wide knockout studies suggested that ACP, a key protein in this pathway, is essential in these stages, indicating that it may have additional non-canonical functions. In this study, the authors confirm that ACP is essential in Pf blood stages (using both apicoplast IPP rescue and conditional knockdown); show that this essential function requires modification with 4-phosphopantetheine and use proximity biotinylation and complementary immunoprecipitation pull-down approaches to provide compelling evidence that ACP binds to and stabilizes the apicoplast-located isoform of pyruvate kinase II. Notably, these interactions appear to differ from those associated with the binding of mitochondrial isoforms of ACP to proteins involved in Fe-S biosynthesis. Loss of ACP was shown to lead to a decrease in PKII levels and apicoplast DNA/RNA synthesis, consistent with loss of NTP synthesis in this organelle. The data are clear and very well described, and the findings represent a significant advance in our understanding of metabolic regulatory mechanisms in apicomplexan apicoplast studies.

      Strengths:

      The study uses a variety of complementary genetic approaches to demonstrate the essentiality of ACP and the enzyme involved in its activation with 4-PP in Pf blood stages, demonstrating that the ascribed non-enzymatic function is mediated by holo-ACP. Similarly, a number of complementary biochemical approaches, including proximity biotinylation, immunoprecipitation, and co-expression of PfACP and PK-II in a heterologous bacterial expression system, are used to confirm the physiological significance of the PfACP and PK-II interaction. The study also reports additional findings, such as the independence of P. faciparum blood stages on exogenous (media) fatty acids, indicating that intracellular stages can salvage all of their requirements from the red blood cell.

      Weaknesses:

      Overall, this is a very strong study. While questions remain around the function of other apicoplast ACP-interacting proteins detected in this study, I don't have any suggestions for significant improvements.

      We thank the reviewer for these positive comments.

      Reviewer #2 (Public review):

      This study focuses on revealing the essential divergent function of the Acyl Carrier protein (ACP) in the deadliest human malaria parasite, Plasmodium falciparum. More precisely, using inducible KO, cellular and biochemical approaches, the authors determined that instead of a canonical role for ACP allowing the de novo synthesis of fatty acids in the apicoplast (essential relict plastid) of the parasite, the enzyme couples with pyruvate kinase II to generate nucleoside triphosphate to maintain parasite survival during blood stages. The study is novel, well-designed, providing interesting new data on Plasmodium and apicomplexa biology. The results convincingly support the major claim of the study. However, it is currently incomplete to support some claims on the essentiality of some apicoplast pathways.

      In this study, Geher et al. focused on deciphering the role of the Acyl Carrier Protein (ACP) present in the relict non-photosynthetic plastid, i.e. the apicoplast of the most lethal human malaria parasite, Plasmodium falciparum. More particularly, they determined an essential function of ACP independent of its usual/typical function as the central protein for the normal function of the apicoplast Type II fatty acid synthesis (FASII) pathway. Rather, the protein seems to associate with the apicoplast Pyruvate Kinase II, together generating an essential nucleoside triphosphate (NTPs) source to fuel the apicoplast and parasite survival instead.

      By generating a TetR-DOZY-based inducible KD line for ACP, they confirmed that the protein is indeed essential to maintain apicoplast integrity and parasite survival during asexual blood stages, as previously predicted and experimentally shown. They showed that ACP requires a biochemical modification, typically activating the protein for its function in the FASII pathway, i.e. binding of the 4-PP group by holoACP synthase. Then, they showed that the other enzymes of the FASII pathway are likely dispensable during the blood stage, as they were able to generate a KO line of the first enzyme of the pathway, FabD (which was predicted to be essential in P. falciparum). Based on a cell culture approach in a controlled culture medium, they further claimed that, unlike current evidence-based hypotheses, the FASII pathway (and thus a potentially FASII-linked ACP) has no role/activity during blood stages. Using a proximity biotinylation approach, they determined that ACP associates with the apicoplast pyruvate Kinase II (PKII), previously shown to generate NTPs in the apicoplast for energy and DNA/RNA maintenance (Xia et al. 2019), and not to fuel the FASII pathway as its main function in blood stages. Finally, they showed that the disruption of ACP induces the reduction of the presence/content in PKII in the parasite, as well as the drastic reduction of the apicoplast DNA and RNA content. Together, they concluded that the main function of ACP is indeed the NTP formation via its association with PKII, rather than its canonical role for the generation of fatty acids in the apicoplast.

      To clarify, we conclude that the essential function of ACP in blood-stage P. falciparum parasites includes a critical stabilizing interaction with pyruvate kinase II. Apicoplast ACP presumably still plays a central biochemical role in FASII pathway function, but that role in FASII is dispensable for blood-stage parasites.

      This study is novel and focuses on a topic of particular interest in malaria biology, but also for most of the apicomplexa-related diseases, and beyond for plastid bearing orgnaisms and this unusual role for ACP. The study is well thought out with proper biochemical approaches that convincingly point to this association of ACP with PKII for NTP synthesis as a major function during P. falciparum blood stages. However, there are currently some important experimental issues/flaws, missing experiments that induced wrong interpretations and thus do not support some important claims of the study, notably for the role of FASII and the interaction between ACP and PKII.

      We note that the major and most important conclusion of our manuscript is that apicoplast ACP has an essential stabilizing interaction with pyruvate kinase II that is required for organelle function and biogenesis. This conclusion is entirely independent of our growth experiments with ∆ACP and ∆FabD parasites in low-lipid conditions, which in principle could be removed from the manuscript without weakening the major conclusions. Nonetheless, we feel that these findings in low-lipid conditions have merit, especially since they contrast with a prior study in the literature regarding P. falciparum growth in low-lipid conditions. We hope that these contrasting results and the questions they raise will stimulate future studies to fully test and understand FASII function under different conditions, including low-lipid conditions.

      We elaborate on these points and address the reviewer’s critiques below.

      Therefore, at this point, the study is only partial and would require major additions and/or important text edits/revisions before being considered for acceptance.

      We note that the manuscript has already been accepted for publication in accordance with the current eLife publishing model.

      Major points:

      From the graph of P. falciparum growth, we can see that in the lipid-rich condition, where both FabH KO and ACP KO can survive, the addition of mevalonate was essential for the growth of ACP KO. Along with the other evidence (PKII association, DNA levels...), we therefore agree that PfACP is involved in the mevalonate pathway.

      To clarify, our model is that ACP supports IPP synthesis by the apicoplast nonmevalonate/MEP pathway indirectly by stabilizing and thus supporting function by pyruvate kinase II that supplies the pyruvate and NTPs required for IPP synthesis by the MEP pathway.

      The authors claim that the FASII pathway is inactive/not essential in the P. falciparum blood stage. However, the authors have not shown any evidence on whether ACP is or not involved in the FASII pathway during the asexual blood stage.

      To clarify, there is overwhelming data in the prior published literature that we cite (including refs. 13, 14, and 32) to establish that FASII is dispensable for blood-stage Plasmodium growth in vivo in rodent parasites and in vitro culture in human parasites. Prior studies also strongly support a role for apicoplast ACP as the central scaffold for FASII-mediated acyl chain synthesis. However, our and prior studies support the conclusion that essential ACP function in blood-stage parasites is independent of its role in FASII.

      As currently designed, the experiments presented cannot conclude on that point for several reasons. Indeed, it was previously shown that (i) the expression of the protein from the FASII pathway are all present in blood stages and are significantly upregulated in patients that are under under "nutrient starvation" (Daily et al. Nature 2007), (ii) that, growing parasites under similar low lipid conditions in vitro induces an activation/upregulation of FASII, which can be measured by stable isotope precursor labelling and lipidomics (Botté et al. 2013).

      We are aware of these prior studies and cite and discuss the Botté et al. 2013 reference in our manuscript, which provided isotope-labeling evidence to support FASII activity in low-lipid growth conditions for P. falciparum. We note that neither study addresses whether FASII activity is required for growth in low-lipid conditions.

      (iii) that growing the PfFabI KO line under deprived lipid conditions leads to parasite death (Amiar et al. 2020), indicating that the FASII pathway can become critical, if not essential, depending on the host nutritionnal content together correlating patients' data and metabolic adaptation for the same reasons in the related parastie Toxoplasma gondii (Amiar et al. 2020, Krishnan et al. 2020, Liang et al. 2020, Primo et al. 2021, Charital et al. 2024, Dass et al. 2024, Bitew et al. 2025).

      All of the studies cited by the reviewer focus primarily or exclusively on Toxoplasma gondii parasites. We agree with the reviewer that these and other studies provide strong evidence that FASII activity contributes to growth of T. gondii parasites, including roles for apicoplast ACP that appear to differ from what we have unveiled for P. falciparum malaria parasites. We acknowledge and discuss these differences from T. gondii in the final section of the Discussion section and think that exploring these differences will be a fascinating area for future study.

      The Amiar et al. 2020 paper cited by the reviewer is the only study we are aware of that has directly tested the ability of a ∆FASII parasite (in this case, ∆FabI) to grow in low-lipid conditions. We acknowledge that they observed little to no growth of ∆FabI parasites in these conditions. Our growth assays with ∆ACP and ∆FabD parasites indicated a different outcome in which both WT and ∆FASII parasites grew similarly in low-lipid conditions. Our results thus contrast with the prior study. As noted below, the minimal lipid growth conditions explicitly reported in the methods section of the Amiar et al. paper are identical to those used in our study: fatty acid-free BSA, 30 µM palmitic acid, and 45 µM oleic acid (all sourced from Sigma) with daily media changes. Thus, the basis for these differences is unclear and additional follow-up work will be needed to explore and resolve these differences.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

      Here, the authors are expecting to show that FabH (and thus the FASII pathway) is not essential in an experiment that is not designed to be in low lipid conditions but rather in lipid rich conditions: Such high lipid conditions of culture in this study is granted by daily feedings with high fatty acid supplement (30-90 uM palmitic acid and 30-60 uM oleic acid). These fatty acid concentrations were used previously by Mitamura et al. (2005) and Miichi et al.(2007) to replace non-determined supplements such as Serum or Albumax supplement to grant similar growth by a completely controlled culture medium.

      This means the concentrations above do not represent limited fatty acid concentrations, especially not with daily feeding (representing an excess supplied amount of lipids, unlike regular 48h feedings) that allowed the authors to easily reach very high non-physiological parasitaemia of more than 20%!! Amiar et al. previously showed essentiality of FabI in P. falciparum in the limited fatty acid culture at a lower concentration (<30uM 16:0, <45um 18:1), than the Mi-Ichi et al. controlled medium with regular 48 h culture feeding. Therefore, with the current experimental settings, the FAH KO is placed in high lipid conditions, thus preventing any conclusion on its essentiality under low lipid conditions.

      The basis for the reviewer’s statements here is unclear, as this critique and the conditions it describes do not conform to the published conditions reported in the Amiar et al. 2020 paper. The methods section of that study for “Plasmodium falciparum growth assays” explicitly states (page e7):

      “Media was replaced daily, sub-culturing were performed every 48 h when required, and parasitemia monitored by Giemsa-stained blood smears. Growth assays in lipid-depleted media were performed by synchronizing parasites before transferring trophozoites to lipid-depleted media as previously reported (Botte ´ et al., 2013; Shears et al., 2017). Briefly, lipid-rich AlbuMAX II was replaced by complementing culture media with an equivalent amount of fatty acid-free bovine serum albumin (Sigma), 30 µM palmitic acid (C16:0; Sigma) and 45 µM oleic acid (C18:1; Sigma).”

      We used identical culture conditions to those described above: fatty acid-free BSA in place of lipid-rich AlbuMAX, 30 µM palmitic acid, and 45 µM oleic acid. We thus obtained growth results that contrast with the prior study and suggest that additional, future studies will be required to understand and resolve these differences.

      Furthermore, it is too uncertain to conclude that ACP is only essential for the mevalonate pathway.

      Please see our response above that clarifies our model for ACP function in supporting pyruvate kinase II and the many apicoplast pathways that appear to depend on PKII.

      This would be a similar discussion to the Yeh et al. 2011 and the Swift et al., where induced Apicoplast knockout caused parasites to require IPP to survive, but there were always remnant apicoplast vesicles and thus the putative presence of an active FASII in the parasite, where de novo fatty acid synthesis could be maintained.

      It is extremely unlikely that FASII remains active upon apicoplast disruption and loss of the apicoplast genome. The apicoplast-encoded SufB is lost upon apicoplast disruption and can no longer participate in making Fe-S clusters. Without Fe-S synthesis, the apicoplast lipoate synthase (LipA) cannot make lipoate to activate pyruvate dehydrogenase (E2 subunit) and produce the acetyl-CoA needed for FASII activity. There is no experimental evidence that FASII remains active upon apicoplast disruption and loss of the apicoplast genome.

      Amiar et al. (2020) and Krishnan et al. (2020) showed that disruption of FASII and absence of de novo FA synthesis in T. gondii could be compensated by the exogenous supplementation of myristic acid, C14:0.

      As explained above, we acknowledge that FASII contributes to Toxoplasma gondii growth and includes functions that appear to differ from P. falciparum.

      Here, high fatty acid supplementation using commercially available fatty acids may include unexpected fatty acid species such as myristic acid in palmitic acid or oleic acid, since all commercially available fatty acids guarantee only >99% but not 100%. If P. falciparum requires a very, very low amount of myristic acid to survive, the amount of possible contamination, like 1 nM, may be sufficient to maintain their survival. Thus, ACP and FabH might be very important to generate de novo fatty acids within parasites, but this was not shown by the authors.

      As noted above, we used identical culture conditions and commercial sources of defined fatty acids to those reported in the Amiar et al. study. We do not see a basis for the reviewer’s critique that the two studies utilized differing culture conditions. Nevertheless, we agree that future studies are needed to understand and resolve these differences.

      Therefore, the manuscript currently contains incorrect conclusions on the potential essentiality/use of FASII, against current experimental evidence.

      As explained above, we do not see a basis for the reviewer’s critique here or for viewing one study as more or less definitive than the other, as identical culture conditions were used yet contrasting results were obtained for reasons that remain uncertain. Future studies beyond the scope of the present manuscript will be required to fully understand and resolve these differences.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      We either request more solid experimental evidence showing the absence of fatty acid synthesis at low fatty acid conditions by re-doing the growth assay in the lower fatty acid feeding conditions without daily feeding to clarify if the ACP and FabH are essential in the blood stage growth, or not; as well as showing the absence of fatty acid synthesis at low fatty acid conditions using isotope labelled precursor. Without these, the authors cannot conclude on this important point. Alternatively, toning down the text to acknowledge the possibility of FASII being active and critical under certain conditions would be acceptable.

      We note that the Amiar et. al 2020 study cited by the reviewer reported growth assays for WT and ∆FabI NF54 P. falciparum in low-lipid conditions that similarly lacked direct tests of FASII activity by isotope labeling.

      We agree that our results, which utilized distinct ∆ACP and ∆FabD NF54 (PfMev) lines, contrast with the results and conclusions of the Amiar et al. study. We fully agree with the reviewer that future studies, utilizing tandem growth assays and isotope-labeling metabolic flux assays (e.g., mass spectrometry), will be required to fully test and understand the dependence of FASII activity on the lipid content of the growth medium and the functional dependence of P. falciparum growth on FASII activity in low-lipid conditions.

      We have revised the final paragraph of the second results section of our manuscript to incorporate this perspective:

      “These results contrast with the prior study [49] of ∆FabI parasites and the proposed model that blood-stage P. falciparum requires FASII activity for growth in low-lipid conditions and suggest that parasites can rely on scavenging host-derived fatty acids over a wide range of lipid conditions. Future studies involving tandem growth and isotope-labeling experiments of WT and ∆FASII parasites will be required to fully test and understand FASII function and the dependence of P. falciparum growth on this pathway in low-lipid conditions.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript addresses an important question in cardiac biology: whether distinct cardiomyocyte (CM) subpopulations play specialized roles during heart development and regeneration. Using single-cell RNA sequencing and newly generated genetic tools, the authors identify phlda2 as a specific marker of primordial cardiomyocytes in the adult zebrafish heart. They further show that these primordial CMs function are essential for myocardial morphogenesis and coronary vascularization but are dispensable for myocardial regeneration or revascularization after injury. These findings indicate that heart regeneration doesn't simply recapitulate developmental processes.

      Strengths:

      A major strength of the study is the generation of a phlda2 BAC reporter, which provides a specific and reliable marker for primordial cardiomyocytes. The lack of genetic tools has previously limited functional analysis of this CM population. By using phlda2 regulatory elements to generate reporter and NTR-based ablation lines, the authors can visualize and selectively manipulate primordial CMs in vivo. This enables a direct functional interrogation rather than relying on lineage tracing or correlative evidence. Through genetic ablation, the authors convincingly demonstrate that primordial CMs are essential for myocardial morphogenesis and coronary vascular organization during development but are not necessary for heart regeneration.

      Weaknesses:

      (1) The manuscript would benefit from clarifying whether the primordial cardiomyocytes ablation affects epicardial cell behaviors during heart development, given that the wellestablished role of the epicardium in supporting coronary vessel growth, it is possible that the vascular phenotypes observed after primordial CM ablation may be affected, at least in part, by altered epicardial cells.

      We thank the reviewer for this important suggestion. To address this possibility, we examined epicardial cells in primordial CM-ablated hearts using the epicardial marker tcf21. Surprisingly, we found that epicardial cells rapidly expanded following primordial CM ablation, with an increase already detected at 5 days post-treatment. This increase persisted through at least 30 dpt, when epicardial cell abundance remained higher than that in control hearts. These findings suggest that the coronary vessel defects are unlikely to result from a reduction in epicardial cell number. In addition, we investigated whether loss of the primordial CM layer permits abnormal inward migration of epicardial cells or coronary vessels into the myocardium. However, we observed no evidence of ectopic localization of either cell population following primordial CM ablation. While we cannot exclude the possibility that altered epicardial function or signaling contributes to the vascular phenotype, our results indicate that the vascular defects are not attributable to reduced epicardial cell abundance or abnormal epicardial invasion. We have included these experiments in Page 10 and 11, and Fig. 3I-3J of revised manuscript.

      “Because epicardial cells play critical roles in coronary vessel development, we next asked whether the vascular defects observed following primordial CM ablation were secondary to alterations in the epicardium. We found that tcf21+ epicardial cells revealed an increase rather than a decrease in epicardial cell abundance in primordial CM-ablated hearts (Fig. 3I3J). These findings suggest that the impaired coronary vessel organization is unlikely to result from a reduction in epicardial cell number, although we cannot exclude the possibility that altered epicardial function or signaling contributes to the vascular phenotype.”

      (2) Because primordial cardiomyocytes form a dense, single-cell-thick layer covering the ventricular surface, it would be informative to determine whether their loss alters the spatial distribution or inward migration of coronary endothelial cells or epicardial cells.

      We appreciate the reviewer for this insightful suggestion. Because primordial cardiomyocytes form a continuous single-cell-thick layer at the ventricular surface, we examined whether their ablation affects the spatial distribution or promotes inward migration of epicardial cells or coronary endothelial cells. Using tcf21 and deltaC reporters, we carefully analyzed the localization of these cell populations within the myocardium. We did not observe any evidence of abnormal inward migration or ectopic localization of either epicardial cells or coronary endothelial cells following primordial CM ablation (Fig. 3J–3K). These results indicate that loss of the primordial CM layer does not disrupt tissue compartmentalization or lead to inappropriate cellular invasion into the myocardial interior. We have included these experiments in Page 11, and Fig. 3J-3K of revised manuscript.

      “In addition, because primordial cardiomyocytes form a continuous layer at the ventricular surface, we investigated whether their ablation permits abnormal invasion of epicardial cells or coronary vessels into the myocardium. However, we observed no evidence of ectopic localization of either cell population following primordial CM ablation (Fig. 3J-3K). Thus, loss of the primordial CM layer does not appear to disrupt tissue compartmentalization or permit abnormal cellular invasion into the myocardium.”

      (3) The manuscript carefully examines the relationship between primordial CMs and gata4<sup>+</sup> cardiomyocytes during regeneration. However, their relationship during heart development should be more fully addressed.

      We thank the reviewer for this important suggestion. To further address the relationship between primordial cardiomyocytes and gata4<sup>+</sup> cardiomyocytes during heart development, we examined their spatial and cellular relationship in juvenile zebrafish hearts. Consistent with our observations during regeneration, we did not detect any overlap between phlda2<sup>+</sup> cardiomyocytes and gata4<sup>+</sup> cardiomyocytes in the juvenile heart (7-8 wpf). These results indicate that primordial CMs and gata4<sup>+</sup> proliferative CMs represent distinct cardiomyocyte populations during both heart development and regeneration. We have included these experiments in Page 13, and Fig.5E of revised manuscript.

      “Consistently, no overlap between phlda2<sup>+</sup> and gata4<sup>+</sup> cardiomyocytes was observed in the wpf juvenile zebrafish heart (Fig. 5E).”

      (4) As loss of cardiomyocytes is known to induce gata4:GFP activation during regeneration, it would be important to determine whether ablation of primordial cardiomyocytes alone triggers gata4:GFP expression in neighboring cardiomyocytes. This analysis would further support the conclusion that primordial cardiomyocytes are not required for regenerative responses.

      We appreciate the reviewer for this important suggestion. To determine whether ablation of primordial cardiomyocytes alone is sufficient to activate regenerative signaling, we treated adult phlda2:mCherry-NTR;gata4:EGFP fish and gata4:EGFP siblings with Mtz for 12 hours per day over three consecutive days without ventricular resection. Under these conditions, we did not observe any induction of gata4 expression following primordial CM ablation (Fig. S5). These results indicate that loss of phlda2<sup>+</sup> cardiomyocytes alone is not sufficient to trigger regenerative gata4 activation in the absence of injury, further supporting that primordial CMs are dispensable for activation of the regenerative response. We have included these experiments in Page 11 and Fig. S5 of revised manuscript.

      “To determine whether ablation of primordial cardiomyocytes is sufficient to activate regenerative signaling, we first treated adult phlda2:mCherry-NTR;gata4:EGFP fish and control gata4:EGFP siblings with Mtz for 12 hours per day over three consecutive days without ventricular resection. We did not observe any induction of gata4:EGFP expression following primordial CM ablation (Fig. S5), indicating that loss of phlda2<sup>+</sup> cardiomyocytes is not sufficient to trigger regenerative gata4 activation in the absence of injury.”

      Reviewer #2 (Public review):

      Summary:

      In the manuscript "Primordial Cardiomyocytes orchestrate myocardial morphogenesis and vascularization but are dispensable for regeneration", Sun et al. identify a novel marker of primordial cardiomyocytes and use it to visualize and ablate the population during development and regeneration. The role of the primordial layer has not been investigated because the tools to manipulate this population have not existed. The manuscript is straightforward, easy to understand, and addresses an important question that has not been explored.

      While the manuscript provides important insights into the role of primordial CMs, backed by a convincing methodology, the authors should clarify their requirements for heart development and maturation. Specifically, is the primordial layer required for the fish to survive?

      We thank the reviewer for this important question. We found that efficient ablation of phlda2<sup>+</sup> primordial cardiomyocytes does not affect overall survival of zebrafish under standard laboratory conditions. Although these animals exhibit clear defects in cardiac structure and coronary vascular organization, they remain viable during the experimental period, indicating that the primordial CM layer is not essential for survival. While we did not assess detailed physiological parameters such as cardiac function, swimming behavior, or long-term fitness in this study, the observed structural abnormalities suggest that subtle functional consequences may exist. These aspects will be important directions for future investigation. We have included the description on page 14, paragraph 2.

      “Although primordial CM ablation does not affect survival under laboratory conditions, the observed structural defects may have functional consequences on cardiac performance and overall physiological fitness, which warrant further investigation.”

      Do primordial CMs regenerate when ablated during development, and do the defects observed (in trabecular and compact CMs and coronary vessels) resolve after 10 days posttreatment when they were detected?

      We appreciate the reviewer for this important question. To determine whether primordial cardiomyocytes regenerate following ablation during development, we performed Mtz-mediated ablation in juvenile zebrafish (7–8 wpf) and examined the hearts at extended time points after treatment. We found that phlda2<sup>+</sup> cardiomyocytes did not recover even at 90 days post-treatment, indicating a persistent loss of this population following developmental-stage ablation (Fig. S6). Importantly, we further assessed whether the cardiac defects observed at earlier time points resolve over time. We found that the abnormalities in trabecular and compact myocardium, as well as coronary vessel organization, persisted at 90 days post-treatment and did not show evidence of recovery (Fig. S4C). These findings demonstrate that the observed defects are not transient developmental delays but represent long-lasting structural alterations of the heart following primordial CM ablation.

      Major Comments:

      (1) Figure 1: A more detailed characterization of the three CM populations would be helpful in the text as well as a new Supplemental Excel Data Sheet with the top unique genes expressed in each.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have expanded the description of the three cardiomyocyte populations in the Results section to provide a more detailed characterization. We have revised the description on page 8. In addition, we have added a new Supplemental Excel Data Sheet (Table S1) listing the top differentially expressed genes for each CM cluster.

      “Notably, Cluster 2 showed additional enrichment for pathways involved in mitochondrial respiratory chain assembly, ATP synthesis, TCA cycle, and ribosome biogenesis, suggesting a relatively higher metabolic and biosynthetic activity state. In contrast, Cluster 1 was enriched for GO terms associated with mitochondrial stress responses, protein degradation, and cytoprotective pathways, indicating a stress-adapted cardiomyocyte state. Cluster 3 showed reduced enrichment of metabolic pathways, consistent with an immature metabolic profile, and was further enriched for genes involved in muscle development and epithelial morphogenesis, suggesting a role in cardiac morphogenesis and tissue organization.”

      (2) Figure 3: Lower magnification views of control and MTZ-treated hearts are needed for all reporters shown (cmlc2, gata4, deltaC) at 10 days post-treatment. These "whole heart" views will enable the reader to get a gross sense of how disrupted heart development is following ablation of the primordial layer.

      We appreciate the reviewer for this helpful suggestion. We have added low-magnification whole-heart images for all reported markers (cmlc2, gata4, and deltaC) to better illustrate the overall cardiac morphology following ablation of the primordial layer. We have included these data in Fig. S3 of revised manuscript.

      (3) It should also be stated clearly in the Results section text what day post-fertilization Mtx treatment began. A section on Mtx treatment, including timing (what day was it applied and what day was it washed out) and dose, should be added to the methods section.

      We thank the reviewer for this helpful suggestion. In response, we have now clearly stated the timing of Mtz treatment in the Results section in page 9, 10 and 11 of revised manuscript.

      “To address the role of phlda2<sup>+</sup> cells during heart development, we performed the following experiments using a standardized Mtz treatment protocol (see Methods), with identical treatment conditions applied to 7-8 wpf juvenile zebrafish.”

      “To address the role of primordial cells during heart regeneration, we perform the below experiments using standardized Mtz treatment protocol (see Methods), with identical treatment conditions applied to adult zebrafish (4–6 months post-fertilization).”

      In addition, we have added a Mtz treatment section in the Methods, which now includes detailed information on dosage, duration, and washout schedule to ensure full reproducibility of the experiments.

      “Mtz treatment

      For conditional ablation of phlda2<sup>+</sup> cardiomyocytes, zebrafish expressing phlda2:mCherryNTR were treated with 10 mM metronidazole (Mtz) for 12 hours per day for three consecutive days. Fish were washed out and maintained in fresh system water after each daily treatment. For developmental analyses, juvenile zebrafish (7–8 weeks post-fertilization) were subjected to Mtz treatment as described above. Following completion of Mtz exposure, fish were maintained under standard conditions. Hearts were collected at 10 days post-treatment for assessment of gata4 activation, and at 30 days post-treatment for analysis of myocardial structure and coronary vessel development. For regeneration experiments, adult zebrafish (4–6 months old) were similarly treated with Mtz for three consecutive days with daily washout. Three days after the final Mtz treatment, ventricular apex resection was performed. Hearts were harvested at 7 days post-amputation (dpa) for analysis of gata4 activation and early regenerative responses, and at 30 dpa for evaluation of myocardial regeneration and coronary vessel revascularization.”

      (4) What happens to the heart 2 and 6 months post-treatment? Are there long-term consequences to primordial layer ablation or do the defects seen at 10 days post-treatment eventually resolve?

      We thank the reviewer for this important question. To determine whether the developmental defects observed following primordial CM ablation are transient or persist long term, we performed additional analyses at later time points after Mtz washout. First, we found that phlda2<sup>+</sup> cardiomyocytes failed to recover following Mtz-mediated ablation. In adult phlda2:mCherry-NTR fish, phlda2<sup>+</sup> cells remained absent 30 days after Mtz washout (Fig. 5B). Similarly, when juvenile fish (7–8 wpf) were treated with Mtz and subsequently allowed to recover, phlda2<sup>+</sup> cells were still not restored at 90 days post-treatment (Fig. S6). Second, juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated primordial CM ablation and analyzed 90 days after Mtz washout. We found that vascular abnormalities persisted long after ablation of primordial CMs (Fig. S4C). Coronary vessels remained disorganized and fragmented, indicating that the vascular phenotype does not resolve over time. Together, these findings demonstrate that primordial CM ablation causes long-lasting defects and that the abnormalities observed are not transient developmental delays. Instead, loss of primordial CMs results in persistent cellular and vascular defects that remain evident months after Mtz treatment. We have included these experiments in Fig. S4C and Fig. S6 of revised manuscript.

      “Notably, these vascular abnormalities persisted at 90 days post-treatment, indicating that the defects do not resolve during subsequent cardiac growth and maturation (Fig. S4C).”

      “Similarly, when juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated ablation and analyzed 90 days after treatment, phlda2<sup>+</sup> cells remained absent, demonstrating a persistent failure of primordial CM recovery (Fig. S6).”

      (5) Also, does the primordial layer come back in these animals where the lineage is ablated during development? Or is it permanently lost as shown in Figure 5 when it is ablated during adulthood?

      We appreciate the reviewer for raising this important question. To determine whether the primordial layer can be reestablished following ablation during development, we treated juvenile phlda2:mCherry-NTR fish (7–8 wpf) with Mtz and examined hearts 90 days after treatment. We found that phlda2<sup>+</sup> cardiomyocytes remained absent at this late time point (Fig. S6), indicating that the primordial layer does not recover following developmental-stage ablation. We have included these experiments in Page 12, and Fig. S6 of revised manuscript.

      “Similarly, when juvenile zebrafish (7–8 wpf) were subjected to Mtz-mediated ablation and analyzed 90 days after treatment, phlda2<sup>+</sup> cells remained absent, demonstrating a persistent failure of primordial CM recovery (Fig. S6).”

      (6) Figure 4: Need to show that phlda2 reporter fluorescence in lost/reduced following Mtz treatment during adulthood before apex amputation.

      We thank the reviewer for this important suggestion. We agree that confirming efficient ablation of phlda2<sup>+</sup> cardiomyocytes in adult fish prior to regeneration analysis is essential.

      In our study, we have already demonstrated in Fig.5B that Mtz treatment in adult phlda2:mCherry-NTR fish results in efficient and sustained loss of phlda2<sup>+</sup> cells, with no detectable recovery at 7 and 30 days post-treatment. These data confirm robust ablation of the primordial CM population following Mtz treatment in adults. Therefore, additional redundant imaging prior to apex resection was not performed in Fig 4.

      (7) Figure 5: It is interesting that primordial CMs do not regenerate following apex amputation or genetic ablation. This result suggests that primordial CMs are only important during development and dispensable during adulthood? This result also makes me question whether primordial CMs are actually required for heart development, which is why it is important to address whether the fish recovers.

      We appreciate the reviewer for this insightful comment. Our data indicate that primordial cardiomyocytes are essential for proper heart development, as their ablation during juvenile stages leads to significant structural and vascular abnormalities. Importantly, we further examined long-term outcomes and found that these defects do not resolve over time. Juvenile zebrafish subjected to primordial CM ablation failed to recover phlda2<sup>+</sup> cardiomyocytes even at 90 days post-treatment, and coronary vascular abnormalities also persisted at this late stage (Fig. S4C). These findings indicate that the observed developmental defects are not transient delays but instead reflect long-lasting structural alterations of the heart. In addition, we found that adult zebrafish similarly fail to regenerate primordial cardiomyocytes following genetic ablation (Fig. 5B), further supporting the limited regenerative capacity of this population. Together, these data demonstrate that primordial cardiomyocytes are required for proper cardiac development, and their loss leads to persistent defects that are not reversed during subsequent growth or regeneration.

      Minor:

      Line 234: Did the authors mean to write Cluster 3 (instead of Cluster 2)?

      We thank the reviewer for pointing out this error. We confirm that this was a labeling mistake, and “Cluster 3” is correct. The text has been corrected in the revised manuscript.

      Line 265: There is a typo of some sort in the phrase, "54.7% reduction closed to the ventricular wall".

      We thank the reviewer for pointing out this error. We have changed the description on page 10 of the revised manuscript.

      “The compact myocardium was disorganized compared with controls, and trabecular muscle formation was severely impaired, with an approximately 54.7% reduction in trabecular area, predominantly observed in regions adjacent to the ventricular wall (Fig. 3A, 3B and S3A).”

      Is there a corollary lineage in mammals? This should be addressed in the Introduction or Discussion.

      We thank the reviewer for this insightful suggestion. At present, a direct corollary lineage to zebrafish phlda2<sup>+</sup> primordial cardiomyocytes have not been clearly defined in mammals. However, mammalian hearts also contain heterogeneous cardiomyocyte populations with distinct developmental states and metabolic profiles, including immature or embryonic-like cardiomyocytes that persist in specific regions during development and early postnatal stages. These populations may share functional similarities with the zebrafish primordial CMs in terms of developmental organization and maturation roles. We have now discussed this point in the Discussion and emphasized that whether a comparable lineage exists in mammals remains an important open question for future studies. We have included the discussion on page 15 of the revised manuscript.

      “Although a direct corollary of phlda2<sup>+</sup> primordial cardiomyocytes has not yet been identified in mammals, mammalian hearts contain heterogeneous cardiomyocyte populations with immature states. Whether these populations represent a functional equivalent of zebrafish primordial CMs remains an open question and needs further investigation.”

      Reviewer #3 (Public review):

      Summary:

      The authors performed single-cell RNA sequencing of adult zebrafish hearts and identified markers for distinct cardiomyocyte subpopulations. One marker, phlda2, marks primordial cardiomyocytes. They generated transgenic reporter lines to characterize phlda2 expression patterns and a phlda2-NTR ablation line to determine the functional requirement of primordial cardiomyocytes during heart regeneration. They found that phlda2+ primordial cardiomyocytes are essential for myocardial morphogenesis and coronary vessel development. Interestingly, when phlda2+ primordial cardiomyocytes are ablated during heart regeneration, gata4+ cortical cardiomyocytes, coronary vessel revascularization, and scar tissue formation are not affected.

      Strengths:

      The authors identified a new primordial cardiomyocyte marker, phlda2. They further demonstrated that primordial cardiomyocytes are important for heart morphogenesis but dispensable for heart regeneration. Their findings reveal a potential difference between heart development and regeneration programs.

      Weakness:

      Despite the interesting findings, the authors did not provide supplemental data for their scRNAseq to demonstrate the data quality and support their conclusions, and some results are not well described.

      We appreciate the reviewer for this important suggestion. In the revised manuscript, we have added supplemental data to support the scRNA-seq analysis, including gene expression tables for each cardiomyocyte cluster (Table S1), and full GO-term enrichment results (Table S2). In addition, we have revised the Results section to improve the clarity and description of the scRNA-seq findings. Please see detailed responses below for point-by-point clarification.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors did not provide enough data to demonstrate the quality of their scRNAseq. They only mentioned that they obtained "high-quality" transcriptomics. Specific parameters such as how many total reads and reads per cell should be provided.

      We thank the reviewer for this important suggestion. In the revised manuscript, we have added detailed sequencing quality metrics to the Methods section in Page 6 of the revised manuscript. The dataset contains 136,174,297 total reads with an average sequencing depth of 36,168 reads per cell.

      “The newly generated scRNA-seq data yielded 136,174,297 total reads with an average sequencing depth of 36,168 reads per cell.”

      (2) The authors utilized cmlc2:EGFP fish to perform scRNASeq. It will be helpful to include feature plots of cmlc2 and EGFP transcripts.

      We thank the reviewer for this suggestion. We have now included the description in Page 8, and feature plots of cmlc2 transcripts and EGFP reporter expression in the scRNA-seq dataset as a supplementary figure (Fig. S1A and S1B).

      “The expression of cmlc2 transcripts and EGFP reporter signal in the single-cell dataset further confirmed the enrichment of cardiomyocytes (Fig. S1A and S1B)”

      (3) The authors show that notch 3 is in cluster 3 of cardiomyocytes and suggest that this reflects elevated NOTCH signaling activity. The authors might consider using RNAScope to further validate that Notch 3 is expressed in cardiomyocytes. It will be also helpful to confirm phlda2 expression patterns during zebrafish heart development and regeneration by RNAScope.

      We appreciate the reviewer for this helpful suggestion. To further examine the spatial expression patterns of these genes, we analyzed publicly available spatial transcriptomic data from zebrafish hearts. We found that phlda2 is enriched in the outer region of the heart during both uninjured and regenerating conditions. Similarly, notch3 and actn1 were also predominantly localized to the outer heart region in the uninjured heart. These findings are consistent with our scRNA-seq results and support the spatially restricted signature of Cluster 3 cardiomyocytes. We have included these data in Page 8, and Fig. S2A-S2C of revised manuscript.

      “To further validate their spatial distribution, analysis of previously published spatial transcriptomic data revealed that phlda2, notch3, and actn1 were predominantly expressed in the outer region of the heart (Fig.S2A-S2C).”

      (4) The authors did not provide any data as supplemental tables to support their analyses of scRNAseq and GO-term analysis.

      We thank the reviewer for this suggestion. In the revised manuscript, we have added new supplemental tables providing full support for the scRNA-seq and GO-term analyses (Table S1 and S2), including lists of differentially expressed genes for each cardiomyocyte cluster and the corresponding GO enrichment results.

      (5) The description of the phenotype in Fig. 3A and B is not clear, especially for the sentence "trabecular muscle formation was severely impaired showing an approximately 54.7% reduction close to the ventricular wall.". The authors might consider using a bracket to show the compact muscle and trabecular muscle and the distance to the ventricular wall.

      We appreciate the reviewer for this suggestion. We have added brackets in Fig. 3A to label the compact and trabecular myocardium for improved clarity. Regarding “distance to the ventricular wall,” we found that this measurement varies substantially across different regions within the same heart, making a single distance-based metric unreliable. Therefore, we quantified trabecular muscle using the trabecular area fraction (trabecular area/total ventricular area) within a defined region of interest as a robust and unbiased indicator. The reported 54.7% reduction refers to this area fraction, and we have revised the text accordingly for clarity in Page 10 of the revised manuscript.

      “The compact myocardium was disorganized compared with controls, and trabecular muscle formation was severely impaired, with an approximately 54.7% reduction in trabecular area, predominantly observed in regions adjacent to the ventricular wall (Fig. 3A, 3B and S3A).”

      (6) It is not clear how the authors quantify the vessel "length"/ventricular area and found that there is no difference (Fig. 3G). The vessel length is significantly shorter in the images (Fig. 3F) as the author also indicated that the vessels are fragmented.

      We thank the reviewer for this important comment. Coronary vessel “length” was quantified by selecting a fixed region of interest (ROI) within the ventricular area, followed by skeletonization of deltaC:EGFP<sup>+</sup> vessels using ImageJ. The total vessel length within the ROI was measured and normalized to the ROI area to obtain vessel length density. Although the representative images (Fig. 3F) show a more fragmented vascular pattern, the total summed vessel length within the defined region was not reduced. This indicates that primordial CM ablation primarily affects vascular organization rather than overall vessel length within the ventricular area. We have updated the figure legend for clarity of the revised manuscript

      “Vessel length was measured within a fixed region of interest (ROI) after skeletonization of deltaC:EGFP<sup>+</sup> vessels in ImageJ.”

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is predominantly expressed in CNS and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 interacts with Nav and Cav channels, modulating channel properties and neuronal excitability.

      In this manuscript, Lu et al. demonstrate that PRRT2 is a potent regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect is highly conserved in PRRT2s across species as well as among DspB family members (TRARG1 and TMEM233). The authors further confirmed the interaction between Nav channels and PRRT2 in heterologous expression systems as well as in Prrt2-V5 knock-in mice. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared with WT mice.

      Overall, this is a well-executed and methodologically comprehensive study. This work offers valuable insight into the physiological functions of PRRT2 and reveals a potential pathogenic mechanism underlying PRRT2-associated neurological disorders.

      The revised manuscript has addressed most of the concerns raised by the reviewers and has been substantially strengthened, although I still have several concerns regarding the discussion section.

      Strengths:

      (1) Overall, this is a well-executed and methodologically comprehensive study. The electrophysiological data strongly support the conclusion that PRRT2 is a potent regulator of Nav channel slow inactivation. The observation that this regulation is conserved in PRRT2 across species and among DspB family members raises the possibility that altered regulation of Nav channels may also contribute to the pathogenesis of TRARG1- or TMEM233-associated disorders.

      (2) Co-immunoprecipitation assay performed using brain tissue from genetically modified Prrt2-V5 knock-in mice provides convincing in vivo evidence for the interaction between PRRT2 and Nav1.2 channels.

      (3) Prrt2-V5 KI mice show markedly reduced PRRT2 protein expression and display phenotypes similar to those observed in Prrt2-mutant mice, supporting an important role of PRRT2 in regulating neuronal and network excitability.

      We sincerely thank the reviewer for the meticulous evaluation of our revised manuscript and for the constructive and insightful comments.

      Weaknesses:

      (1) Nav1.6 is also highly expressed in cortical neurons and is widely regarded as a major contributor to action potential initiation and sustained high-frequency firing. Given that PRRT2 similarly regulates the fast and slow inactivation of Nav1.6 and Nav1.2 channels, the potential contribution of Nav1.6 regulation to neuronal and network excitability should be discussed.

      We appreciate the reviewer’s suggestion. In the revised manuscript, we have clarified that PRRT2-mediated regulation of Nav1.2, together with its regulation of Nav1.6, may contribute to neuronal and network excitability in the cortex. Please refer to Page 15, Line 439-440.

      (2) Slow inactivation is generally considered to develop over timescales ranging from hundreds of milliseconds to seconds or longer. Therefore, the statement in Discussion (Page 13, line 381-382) that "slow inactivation develops on a timescale of tens of milliseconds to seconds" may not accurately reflect the conventional kinetic definition of slow inactivation and should be clarified.

      We thank the reviewer for this comment. We have corrected the timescale description in the Discussion accordingly. Please refer to Page 13, Line 382.

      (3) Page 14, line 417-430: "question about how Nav channel slow inactivation is regulated in cells that do not express PRRT2".

      PRRT2 is unlikely to be the sole regulator of Nav channel slow inactivation. Other molecules and signaling pathways may regulate Nav channel and contribute to neuronal excitability. In addition, neuronal excitability can also be regulated through modulating other Nav properties, such as long-term inactivation or slow recovery from inactivation, as well as through modulating the activity of other ion channels, for example, Kv7.2 and Kv7.3 channels. Therefore, PRRT2-negative cells may utilize alternative mechanisms to fine-tune neuronal excitability. In its current form, this paragraph somewhat overstates the role of PRRT2 and would benefit from a more balanced discussion.

      We appreciate the reviewer’s constructive suggestion. We agree that PRRT2 is unlikely to be the sole regulator of neuronal excitability or Nav channel slow inactivation. In the revised Discussion, we have clarified that PRRT2-dependent regulation of Nav channel slow inactivation represents one mechanism among several that fine-tune neuronal excitability. Other mechanisms, including regulation of potassium channels such as Kv7.2/Kv7.3, and other ion channel- or signaling- dependent pathways, may also contribute to excitability control in both PRRT2-positive and PRRT2-negative neurons. We have revised relevant paragraph to provide a more balanced discussion of alternative mechanisms. Please refer to Page 15, Line 430-438.

      (4) Page 50, Figure 7-figure supplement 2: It would be helpful to include representative traces of the 1st and the last compound APs in panels B and C.

      We appreciate the reviewer’s valuable suggestion. We have now added representative traces of the first and last compound action potentials, corresponding to the 1st and the 100th responses, respectively, to panels B and C of Figure 7-figure supplement 2. Please refer to Page 51, Figure 7-figure supplement 2.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Page 19, line 546: a sampling rate of 10kHz was used for recording Nav currents. Because the Nav channel isoforms examined in this manuscript (Nav1.1, Nav1.2, Nav1.4, Nav1.5, and Nav1.6) exhibit extremely rapid activation and inactivation kinetics, the temporal resolution provided by a 10 kHz sampling rate may not be optimal for detailed kinetic analysis. Although I am not requesting additional experiments, the authors may wish to choose a higher sampling rate (e.g., 50 kHz) in future Nav channel studies.

      We thank the reviewer for this helpful suggestion regarding the sampling rate for Nav current recordings. We will adopt higher sampling rates for rapid kinetic analyses of sodium currents in our future work.

      (2) Typo: Page 24, line 712-713: should "a Digidata (Molecular Devices, 1332A)" be "... 1322A"?

      We thank the reviewer for pointing out this typo. We have corrected “Digidata 1332A” to “Digidata 1322A” in the relevant Method section of revised manuscript. Please refer to Page 25, Line 722.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We again thank the editor and reviewers for their detailed attention to our work. In our previous revisions we endeavored to address the principal concerns raised by reviewers that we were capable of addressing. We recognize that asymptomatic pertussis transmission represents a particularly thorny area of epidemiology and public health, where multiple (and sometimes overlapping) mechanisms have been proffered even as empirical evidence remains thin, particularly in low-resource settings such as sub-Saharan Africa. A key finding of our work is that prospective surveillance in such a low-resource setting revealed abundant evidence of otherwise unobserved asymptomatic incidence. Moreover, as we note in our prior revisions, this finding is supported by recent work in South Africa and elsewhere (Kayina et al., 2015; Moosa et al., 2019, 2025). As such, we believe that further prospective surveillance in similar settings would be highly informative, a point that we have sought to emphasize in our present manuscript (and associated commentary).

      While we broadly agree with many of the concerns raised by the reviewers, we believe that we are unable to significantly strengthen the present work through further revisions. We do, however, wish to respond to several points raised in these reviews. Of note, a reviewer raises the possibility of a "breakthrough strain" without reference to existing literature. We agree that we cannot test this hypothesis, and though it is not incompatible with our own findings, it is, however, not consistent with recent molecular surveillance in South Africa (Moosa et al., 2023). The reviewer also raises the potential of low adult vaccination coupled with recent reintroduction. This hypothesis relies on our investigation looking "at the right place and the right time", and further does not explain how immunologically naive adults would have escaped morbidity. We have adopted what we believe is a more parsimonious interpretation of our results (i.e., that asymptomatic infection represents evidence of previous immune exposure), though we agree that a more thorough exploration of this particular issue is warranted, particularly in light of our persistently colonized mothers. In addition, we have noted similar studies in sub-Saharan Africa that also found widespread evidence of asymptomatic pertussis, which we believe is inconsistent with a “right time, right place” interpretation.

      A reviewer also pointed to the work of Warfel et al. (2014) and Althouse and Scarpino (2015). We are familiar with both of these studies and agree with their broad relevance to the field (our apologies for omitting Warfel et al.). The reviewer states that, "If the mechanism underlying the results in Zambia is that either WP or natural infection does not block transmission (in the absence of a breakthrough strain), that would upend many of the assumptions in pertussis research." Critically, we believe that our prospective field study of human patients in a real-world public health system complements previous research, including animal trials and simulation studies. Simply put, given that our study was unable to establish the prior vaccination or exposure status of participants, we do not claim to have shown evidence for transmission despite wP vaccination or prior infection.

      Regarding Althouse & Scarpino (2015), we believe that, for the majority of readers, the most compelling analysis in their paper was the examination of genome sequences that pointed to substantial asymptomatic transmission in the US. This conclusion emerged from their population model, which required that “births” (representing transmission events) exceeded “deaths” (representing recovery of infectious individuals) in order to be consistent with the sequence data. Unfortunately, this paper does not provide a detailed explanation of their methods and data sources, nor is this work directly reproducible through, for example, an open-access code/data repository. We have explored the availability of US genome sequences over the time period of their study and were able to find only 36 sequences: 2 from the pre-vaccine era, 8 from the wP vaccine era, and 26 from the aP vaccine era. Given this notable imbalance in the number of sequences (and thus sequence diversity) that was biased in favour of the most recent time period, is it then surprising that the “birth rate” in their model had to exceed the “death rate” in order to match the genetic diversity in the data? Based on a careful inspection of this work, we do not consider its conclusions to represent a gold standard against which all subsequent studies should be judged. We also note that genomic surveillance and analysis of pertussis remains sparse relative to other fields, though recent works have added dramatically to the corpus of available sequences (Bridel et al., 2022).

      Finally, we note that our previous revisions addressed several concerns raised in the present reviews. For example, we previously sought to address reviewers' about our presentation of the strength of our evidence. In this regard, we broadly agree with the reviewers, and we now state that our results "suggest that pertussis transmission occurs between minimally symptomatic mothers and their newborn infants." We believe this largely addresses a present reviewer's concern that 'the mother-to-infant transmission pathway should be framed as "highly suggestive" rather than "confirmed"'. We also note that our results examine three different threshold Ct values (survival analysis, Fig 4), a point that we believe partially addresses a reviewer's suggestion to "including a sensitivity analysis using a stricter cut-off" and concerns about "the decision to use a Ct<45 threshold, as this is higher than standard clinical cut-offs". Indeed, we discuss the issue of clinical cut-offs (and their appropriateness) at some length in the section, "Test reliability, disease surveillance, and public health where we state, "We recognize that such weak and potentially ambiguous signals may not be appropriate for clinical diagnosis. However, our results demonstrate that they nonetheless contain valuable information about pathogen presence and infection intensity that can (and should) be leveraged for disease surveillance." We have also included in the present work a detailed discussion of qPCR sensitivity and efficiency that we believe should interest others working in pertussis surveillance.

      We do not view our own research as the "last word" in this rather controversial subject. In this spirit, we have attempted to present our work transparently, state our claims carefully, and underscore future activities that we believe would benefit the pertussis research community going forward. For example, we agree that further attention to shared exposure and functional immune data among low-resource communities could provide valuable insights into epidemiology and ecology of pertussis. However, we also believe the trade-offs of including one set of activities over another should be clearly acknowledged by researchers, clinicians, and public health officials. To simply state that we must measure more fails to account for the very real resource constraints that we all face.

      Althouse, B. M., & Scarpino, S. V. (2015). Asymptomatic transmission and the resurgence of Bordetella pertussis. BMC Medicine, 1–12. https://doi.org/10.1186/s12916-015-0382-8

      Bridel, S., Bouchez, V., Brancotte, B., Hauck, S., Armatys, N., Landier, A., Mühle, E., Guillot, S., Toubiana, J., Maiden, M. C. J., Jolley, K. A., & Brisse, S. (2022). A comprehensive resource for Bordetella genomic epidemiology and biodiversity studies. Nature Communications, 13(1), 3807. https://doi.org/10.1038/s41467-022-31517-8

      Kayina, V., Kyobe, S., Katabazi, F. A., Kigozi, E., Okee, M., Odongkara, B., Babikako, H. M., Whalen, C. C., Joloba, M. L., Musoke, P. M., & others. (2015). Pertussis prevalence and its determinants among children with persistent cough in urban Uganda. PLoS One, 10(4), e0123240.

      Moosa, F., du Plessis, M., Weigand, M. R., Peng, Y., Mogale, D., de Gouveia, L., Nunes, M. C., Madhi, S. A., Zar, H. J., Reubenson, G., & others. (2023). Genomic characterization of Bordetella pertussis in South Africa, 2015–2019. Microbial Genomics, 9(12), 001162.

      Moosa, F., du Plessis, M., Wolter, N., Carrim, M., Cohen, C., von Mollendorf, C., Walaza, S., Tempia, S., Dawood, H., Variava, E., & others. (2019). Challenges and clinical relevance of molecular detection of Bordetella pertussis in South Africa. BMC Infectious Diseases, 19, 1–11.

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018: Findings from the PHIRST study. Journal of Infection, 106550.

      Warfel, J. M., Zimmerman, L. I., & Merkel, T. J. (2014). Acellular pertussis vaccines protect against disease but fail to prevent infection and transmission in a nonhuman primate model. Proceedings of the National Academy of Sciences, 111(2), 787–792. https://doi.org/10.1073/pnas.1314688110


      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study investigates the role of asymptomatic pertussis carriage in transmission between mothers and their infants, in particular. The authors used a longitudinal cohort study that involved 1,315 mother-infant dyads in Lusaka, Zambia, and they utilized qPCR-based detection of IS481 to track Bordetella pertussis transmission over time. Insights from the study suggest that minimally symptomatic or asymptomatic mothers may act as a reservoir for B. pertussis transmission in the infants, thus challenging the traditional surveillance methods that focus on symptomatic cases. Additionally, the study also identified a subgroup of persistently colonized individuals where mothers were majorly asymptomatic despite sustained bacterial presence.

      The authors aimed to improve comprehension of pertussis transmission dynamics in high-burden low-resource settings, and they advocated for enhanced molecular surveillance strategies to capture full pertussis infection, including those that might have gone undetected.

      Strengths:

      The strengths are the use of innovative study design, especially the longitudinal approach and routine sampling, rather than symptom-driven testing that minimizes bias in the study. The methodology was also rigorous and transparent by evaluating the IS481 signal strength to classify pertussis detection and conducting retesting to assess qPCR reliability. There were also important epidemiological insights, and the findings challenge the traditional wisdom by suggesting that pertussis transmission may frequently occur outside of symptomatic cases. The findings also showed its relevance to global health and policy by arguing for the incorporation of molecular tools like qPCR for surveillance of pertussis in low-resource settings.

      Weaknesses:

      These include reliability on qPCR-based detection without additional validation measures like confirmatory culture or serology. There are also potential alternate explanations for transmission patterns observed in the study such as shared environmental exposure or household transmission. Additionally, there is limited generalizability as the study was done in a single urban site in Zambia. There is also a lack of functional immune data.

      Reviewer #2 (Public review):

      Summary:

      In this paper, the authors describe the results of a longitudinal study of pertussis infection in mother/infant dyads in Lusaka, Zambia. Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage. This work represents an important contribution to our understanding of the global burden of pertussis. Also, it highlights the still under-appreciated role of asymptomatic transmission across many infectious diseases (including vaccine-preventable ones).

      Strengths:

      Unlike many past studies, the authors assessed the infection status of individuals independently of whether they were symptomatic for a respiratory infection. As a result, this work represents one of the first studies specifically designed to assess asymptomatic transmission of pertussis. Using qPCR, the authors find strong evidence for the role of asymptomatic transmission from mothers to infants and also evidence for long-term bacterial carriage.

      Weaknesses:

      While I am quite enthusiastic about the work, I am concerned that a number of likely relevant confounders were not discussed and that the broader implications of their findings were not well grounded in the existing literature. For example, I could not find information on the vaccination status of the mothers in the study. Given the conclusions about asymptomatic transmission and the durability of immunity, it is important to know the vaccination status of the mothers. Moreover, did the authors have other metadata on the mother/infant dyads, e.g., household size, vaccination status of household members, etc.? Given the potential implications of more widespread asymptomatic transmission associated with pertussis infection, I believe the authors should better couch their results in the context of the broader debate around asymptomatic transmission.

      We appreciate the reviewers' detailed feedback. We provide an overview of our responses here and we address specific recommendations below. In light of reviewers’ comments, we have revised our manuscript in order to improve the clarity of our presentation and to better situate our results within the context of the existing literature. Unfortunately, as the field study has been concluded, many of the reviewers’ recommendations are not possible. These include additional testing (i.e., culture or serology) or sequencing. We have updated the manuscript to more clearly indicate our knowledge regarding maternal vaccine status and immunological immunity of study participants. We have also provided a more comprehensive overview of existing pertussis studies, including genomic surveillance and details regarding sub-Saharan Africa and Zambia in particular. Finally, we have revised the formatting of Figure 4 (survival analysis) to more clearly highlight differences between mothers and infants and to better align with the text, and note that the underlying results are unchanged.

      A particular concern raised in the reviews that we wish to address is the recommendation of culture- or serology-based tests as "confirmatory". We have revised the manuscript in light of this feedback to better reflect our own position on this matter. We believe these recommendations do not adequately account for important trade-offs between testing sensitivity and specificity that are widely recognized in both clinical practice and epidemiology (Enøe et al., 2000; Florkowski, 2008; Swift et al., 2020). When the results of different testing methodology disagree, rarely is one method, a priori, correct. Rather, the disagreement may point to specific test limitations or important biological questions about the study system.

      In the case of pertussis detection, cell culture is recognized for its very low sensitivity, while serological detection is complicated by debate around appropriate threshold levels and time horizons for seroconversion and subsequent decay (Lee et al., 2018; van der Zee et al., 2015). Furthermore, while anti-PT antibodies are a common target of serological detection, these are not reliable correlates of protection (Mills, 2001; Wilk et al., 2019), nor are they reliably generated in response to colonization (de Cellès & Rohani, 2024; Graaf et al., 2020). Overall, the detailed relationship between exposure, carriage, transmissible infection, and the dynamics of anti-PT serology remains poorly characterized (Craig et al., 2020; de Cellès et al., 2025).

      While we agree that these are important questions in epidemiology and public health, we nonetheless wish to highlight that there is no “free lunch": each additional test and protocol comes with additional cost and complexity that should be evaluated based on the specific goals of the intended surveillance. In our case, the repeated sampling of longitudinal surveillance serves as a low-cost "confirmatory" testing regime. We disagree that cell culture would have added value to the present study and would not recommend its addition to future studies (primarily due to low sensitivity). While we agree that before-and-after serology of mothers would have added important context to the present study, we nonetheless expect that significant ambiguity would have surrounded any such results (e.g., Moosa et al. (2025)).

      One area that we strongly agree warrants further attention is the household dynamics in pertussis transmission, particularly in low-resource settings where crowding is common. In our study we were not able to rule out environmental and/or shared transmission events, though our survival analysis did demonstrate a greater impact of mothers on infants than vice versa, results which suggest a causal mechanistic role. In previous studies we detailed the demographics of household size, number of children, and mothers' age (Gill et al., 2021; Gunning et al., 2020), though we have not conducted formal analyses of these important covariates here. We also note that the code and data are freely available, allowing for others to build on our work.

      We believe that future prospective studies are an invaluable tool for directly tracking pertussis disease transmission, including both community and household studies. We have argued here for the value of qPCR-based community surveillance, which could integrate into existing public health activities. Regarding household studies, we note that a key challenge in implementing these studies is selecting an appropriate sampling interval and duration to best capture epidemiological linkages. Our results suggest that qPCR-based real-time population-level surveillance could be used to initiate such a prospective household study during a pertussis outbreak so that a higher sampling frequency (e.g., weekly swabs) could be gainfully employed over a shorter time period.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Enhance Validation of qPCR Findings:

      We address these comments above at greater length. We also note that the term “false positive” is rather ambiguous here, as we lack a clear distinction between carriage and transmissible infection. We note that our manuscript includes a considerable discussion of qPCR validation, including negative controls and sample retesting. We agree that test sensitivity and specificity remains an important question, and have endeavored to clearly indicate these concerns throughout the manuscript.

      (2) Clarify Transmission Dynamics:

      While we agree that this is an important question, we lack the relevant sequence data to test it. Notably, we suspect that any such phylodynamic linkage would require genomic sequencing due to the relatively low genetic diversity observed in pertussis (see, e.g., population-level estimates of time to most common ancestor in Lefranc et al., (2022). However, sequencing pertussis genomes remains resource-intensive, and we expect that the deployment of such sequencing at scale would be cost-prohibitive in low-resource settings.

      We have revised our manuscript to underscore uncertainties around shared/household exposure. We also direct the reviewer’s attention to our survival analysis, where a notable asymmetry exists between mothers and infants. Here, mothers’ prior qPCR signals exhibit a larger impact on their infants than infants on their mothers (Fig 4). If shared exposure was the principal cause of the observed increase in hazard in ego from alter, then we would expect no such asymmetry between mothers and infants.

      (3) Expand Discussion on Public Health Implications:

      As noted above, we have revised our introduction and expanded our discussion to better account for the existing literature. And, while we are hesitant to put forward specific recommendations based on our study, we feel confident in stating that active pertussis surveillance in low-resource settings is A) almost entirely absent B) possible to achieve, and C) necessary to resolve long-standing questions about pertussis epidemiology at regional and national levels. We have endeavored to clarify these points, particularly within the discussion.

      (4) Address the Role of Immunity More Directly:

      While we lack such immune data, we have pointed to recent work from the notable PHIRST study in South Africa, as well as highlighted ambiguities surrounding these data.

      Reviewer #1 (Recommendations for the authors):

      (1) Do we know the vaccination status of the mothers in the study? … If these data are not available, I think that the paper must be re-framed to acknowledge that all the conclusions are statistical in nature, based on publicly available vaccine coverage data from Zambia.

      We do not have information on the immunization status of mothers, though we cite national rates for Zambia across the relevant time period. We have clarified this point in the revised manuscript. We have endeavored to clearly acknowledge that many of our conclusions are statistical in nature and to clearly quantify the strength of evidence.

      We strongly disagree that anomalously low vaccination rates amongst mothers (i.e., relative to national averages) would materially alter the interpretation of our findings.

      Overall, our findings strongly suggest ongoing pertussis transmission in this population. Based on this, we expect that mothers in our study who were not vaccinated would likely have some degree of infection-derived immunity. Indeed, some have argued that the preponderance of mild/asymptomatic infections in mothers is, of itself, evidence of prior immunological exposure (Fine & Clarkson, 1982).

      (2) Do we know anything about rates of pertussis in Zambia, especially in the study site?

      We address this important question in the discussion. In particular, we state that: “As a populous, middle-income and primarily urban country, Zambia offers an evocative example of pertussis surveillance, where no cases have appeared in official WHO reports since 2009”.

      (3) I couldn't find information in the paper related to the severity of infection in the infants. It's mentioned in the section describing results in Figure 5, but I only saw analyses with symptoms (as opposed to severe symptoms). Do you have outcome data from infants testing positive?

      This question was addressed in more detail in our previous work (Gill et al., 2021), which we briefly summarize in the Introduction. We also show the frequency of severe symptoms in Fig 5 (bottom panel), and detail mild versus serious symptoms in our subgroup analysis (Fig 7D).

      (4) Do we know anything about vaccine-resistant strains of pertussis in Zambia?

      We are not aware of any such work. As we noted above (and now address in our Discussion), widespread genomic surveillance and microbiological characterization of pertussis are sorely lacking across Africa.

      (5) While I believe sequencing is beyond the scope of the current study, the authors should comment on the potential utility of sequencing elements of the pertussis genome and use that to demonstrate causality and direction of transmission more strongly.

      We believe that existing literature has addressed the potential of sequence data and phylodynamics to infer transmission, particularly for pathogens with high mutation rates such as RNA viruses. To date, research into the phylodynamics of pertussis has focused exclusively on population-level dynamics (Lefrancq et al., 2022), where estimates of time to most recent common ancestor (TMRCA) are long, indicating low genomic variability at the scale of countries and years. To our knowledge, no work on pertussis has directly inferred transmission chains from sequence data. Given the existing evidence, we expect that any such work would require genome-level sequencing, which would likely be cost-prohibitive in low-resource settings.

      (6) … However, it would be helpful to understand more about how your results fit into the broader story around pertussis resurgence. … if the infant cases were all mild, they might never have been captured in surveillance data sets.

      We believe that a key result of our study is the remarkable mismatch between country-level symptoms-based surveillance and prospective surveillance, which demonstrates that such mild cases have almost certainly not been captured. These findings are mirrored by recent work in South Africa (now addressed in our Discussion, see Moosa et al. (2025)). We believe that prospective surveillance, particularly in under-surveilled regions, is critical to understanding pertussis transmission writ large, which we have attempted to communicate throughout our discussion.

      (7) Relatedly, if there are still high rates of asymptomatic mother-to-infant transmission with whole cell vaccination, then why is there an observed drop in infant pertussis following whole vaccination in most countries?

      In previous work, we demonstrated that some infants in this cohort exhibited asymptomatic infection (Gill et al., 2021). We note that a drop in pertussis incidence amongst infants after the roll-out of the whole-cell vaccine is not contradictory with our findings. We want to clarify that our results, and evidence that mother-to-infant transmission can occur, does not imply that the whole-cell vaccine fails to protect against transmission.

      We have previously used epidemiological evidence to infer the population-level impacts following the roll-out of whole-cell pertussis infant immunization. For example, we observed an increase in the inter-epidemic period that, together with the drop in infant cases, are consistent with a reduction in transmissible infections (Broutin et al., 2010; Rohani et al., 2000).

      References

      Broutin, H., Viboud, C., Grenfell, B. T., Miller, M. A., & Rohani, P. (2010). Impact of vaccination and birth rate on the epidemiology of pertussis: A comparative study in 64 countries. Proceedings of the Royal Society B: Biological Sciences, 277(1698), 3239–3245.  https://doi.org/10.1098/rspb.2010.0994  

      Craig, R., Kunkel, E., Crowcroft, N. S., Fitzpatrick, M. C., Melker, H. de, Althouse, B. M., Merkel, T., Scarpino, S. V., Koelle, K., Friedman, L., Arnold, C., & Bolotin, S. (2020). Asymptomatic Infection and Transmission of Pertussis in Households: A Systematic Review. Clinical Infectious Diseases, 70(1), 152–161. https://doi.org/10.1093/cid/ciz531

      de Cellès, M. D., & Rohani, P. (2024). Pertussis vaccines, epidemiology and evolution. Nature Reviews Microbiology, 1–14. https://doi.org/10.1038/s41579-024-01064-8

      de Cellès, M. D., Wong, A., Dalby, T., & Rohani, P. (2025). Natural immune boosting biases pertussis infection estimates in seroprevalence studies. Nature Communications, 16(1), 8883. 

      Enøe, C., Georgiadis, M. P., & Johnson, W. O. (2000). Estimation of sensitivity and specificity of diagnostic tests and disease prevalence when the true disease state is unknown.  Preventive Veterinary Medicine, 45(1–2), 61–81.

      Fine, P. E. M., & Clarkson, JacquelineA. (1982). The recurrence of whooping cough: Possible implications for assessment of vaccine efficacy. The Lancet, 319(8273), 666–669.  https://doi.org/10.1016/S0140-6736(82)92214-0 

      Florkowski, C. M. (2008). Sensitivity, specificity, receiver-operating characteristic (ROC) curves and likelihood ratios: Communicating the performance of diagnostic tests. The Clinical Biochemist Reviews, 29(Suppl 1), S83.

      Gill, C. J., Gunning, C. E., MacLeod, W. B., Mwananyanda, L., Thea, D. M., Pieciak, R. C., Kwenda, G., Mupila, Z., & Rohani, P. (2021). Asymptomatic Bordetella pertussis infections in a longitudinal cohort of young African infants and their mothers. eLife, 10, e65663. https://doi.org/10.7554/elife.65663

      Graaf, H. de, Ibrahim, M., Hill, A. R., Gbesemete, D., Vaughan, A. T., Gorringe, A., Preston, A.,  Buisman, A. M., Faust, S. N., Kester, K. E., Berbers, G. A. M., Diavatopoulos, D. A., & Read, R. C. (2020). Controlled Human Infection With Bordetella pertussis Induces Asymptomatic, Immunizing Colonization. Clinical Infectious Diseases: An Official Publication of the Infectious Diseases Society of America, 71(2), 403–411.  https://doi.org/10.1093/cid/ciz840

      Gunning, C. E., Mwananyanda, L., MacLeod, W. B., Mwale, M., Thea, D. M., Pieciak, R. C., Rohani, P., & Gill, C. J. (2020). Implementation and adherence of routine pertussis vaccination (DTP) in a low-resource urban birth cohort. BMJ Open, 10(12), e041198.

      Lee, A. D., Cassiday, P. K., Pawloski, L. C., Tatti, K. M., Martin, M. D., Briere, E. C., Tondella, M. L., Martin, S. W., & Group, C. V. S. (2018). Clinical evaluation and validation of laboratory methods for the diagnosis of Bordetella pertussis infection: Culture, polymerase chain reaction (PCR) and anti-pertussis toxin IgG serology (IgG-PT). PLoS One, 13(4), e0195979.

      Lefrancq, N., Bouchez, V., Fernandes, N., Barkoff, A.-M., Bosch, T., Dalby, T., Åkerlund, T.,  Darenberg, J., Fabianova, K., Vestrheim, D. F., Fry, N. K., González-López, J. J.,  Gullsby, K., Habington, A., He, Q., Litt, D., Martini, H., Piérard, D., Stefanelli, P., … Brisse, S. (2022). Global spatial dynamics and vaccine-induced fitness changes of Bordetella pertussis. Science Translational Medicine, 14(642), eabn3253.  https://doi.org/10.1126/scitranslmed.abn3253 

      Mills, K. H. G. (2001). Immunity to Bordetella pertussis. Microbes and Infection, 3(8), 655–677. https://doi.org/10.1016/s1286-4579(01)01421-6 

      Moosa, F., Kleynhans, J., Makhathini, L., du Plessis, M., Tempia, S., McMorrow, M. L., Moyes, J., Buys, A., Maake, L., Smit, S., & others. (2025). Bordetella pertussis infection and antibody dynamics in household cohorts in two South African communities, 2016–2018:  Findings from the PHIRST study. Journal of Infection, 106550. 

      Rohani, P., Earn, D. J., & Grenfell, B. T. (2000). Impact of immunisation on pertussis transmission in England and Wales. The Lancet, 355(9200), 285–286. https://doi.org/10.1016/S0140-6736(99)04482-7 

      Swift, A., Heale, R., & Twycross, A. (2020). What are sensitivity and specificity?  Evidence-Based Nursing, 23(1), 2–4.

      van der Zee, A., Schellekens, J. F., & Mooi, F. R. (2015). Laboratory diagnosis of pertussis.  Clinical Microbiology Reviews, 28(4), 1005–1026.

      Wilk, M. M., Allen, A. C., Misiak, A., Borkner, L., & Mills, K. H. G. (2019). The immunology of Bordetella pertussis infection and vaccination. In Pertussis: Epidemiology, Immunology & Evolution. Oxford University Press.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Fisher et al describes the molecular mechanism underlying how G beta gamma subunits engage with the beta 3 isoform of PLC. The paper used a combination of cryo EM, BRET assays, and biochemical assays of PLC beta activity. A key discovery is that G beta gamma is not sufficient to drive membrane binding by itself, and instead promotes G alpha activation. The work is important, but suffers slightly from some ambiguity in the actual interface that is present in their cryo EM model, as crosslinkers could stabilise a transient and non-native complex. This is somewhat abrogated by the careful mutational analysis, which shows that mutation of any of these three sites does somewhat block PLC beta G beta gamma activation. However, there could be some improvement in the presentation of this data, as well as possible mutant selection. Overall, this paper is a nice complement to the Falzone et al paper, showing the membrane-bound complex of PLCB3 on membranes, with this work building on this work, highlighting the importance this will have in our full understanding of PLC beta activation.

      Thank you for the positive feedback.

      Major concerns:

      My biggest concern is the potential that this interface is artefactual based on the crosslinking strategy utilised. Here are thoughts on how this could be better validated, presented in a more convincing way.

      (1) The authors' main claim is that there is a degree of plasticity of G beta gamma binding to the PLC beta 3 isoform, with three possible binding sites. The main complication of this is, of course, the possibility that the crosslinking stabilises a non-native complex, driven by a mutated cysteine.

      Because of this, any other additional details about this interface are going to be critical for the scientific audience to judge if this is accurate.

      What would greatly help Figure 1 is an evolutionary conservation analysis of the novel Gbg interface in PLC, to see how well this is conserved, and compare this to the conservation of the previously annotated sites. Conservation of these sites on both the G beta gamma and PLC side would help justify this as a native complex.

      This will also help orient the reader to the identity of the mutated residues assayed in Figure 3.

      We agree that crosslinking can capture non-physiologically relevant interfaces. However, because we do not observe any crosslinking between Gβγ and a PLCβ3 variant that retains a cysteine in the X–Y linker or between PLCβ3 and any other cysteine in the Gβγ heterodimer, we believe it is site-specific.

      The question about sequence conservation in the Gβγ–PLCb3 interfaces is interesting and we have included this information in Figures S8 and S9.

      (2) The g beta gamma orientation is also different than what I have observed in previous g beta gamma effector structures. Is there any precedent for this as an effector interface? A supplemental figure comparing this structure to other g beta gamma interfaces from other enzymes, for example recent Tesmer structure with PI3K.

      We agree that the orientation of Gβ in the crosslinked structure is different. We include a comparison of this reconstruction to other published Gβγ–effector complexes as Figure S6.

      (3) The mutational analysis in Figure 2D-G seems to give some strange results, and I have some question why certain residues were chosen rather than others. Mutation of the Gbg side will be more complicated, as of course that can affect any of the three surfaces. My main question is that, from the way Figure 2A is oriented, the main salt bridge in their novel interface to me looks like R199-D228, with K183 being in the wrong orientation to E226, and D167 being far from any charged residues. Why did the authors not make the corresponding R199 to D or E mutation?

      Thank you for pointing this out, and we expanded our analysis to include this residue. The R199A and R199E mutations had no defects in basal or Ga<sub>q</sub>-stimulated activities. However, R199A had 3-fold lower activation by Gβγ, while R199E was not activated in this assay. The R199E mutation also had significantly decreased agonist-dependent BRET with Gβγ and decreased PI(4,5)P2 hydrolysis, while retaining robust recruitment to the plasma membrane by Gα<sub>q</sub>. These data is included in the main text and Figures 2-4.

      (4) To help the reader's interpretation of Figure 2A, I would recommend a supplemental figure showing the density for interfacial residues, as that also would increase confidence in the interface.

      Thank for the suggestion. In revised Figure S3, we show the Gβγ–PLCb3 D892-PH<sub>cys</sub> complexes determined in this study at different contour levels.

      Reviewer #2 (Public review):

      In this manuscript, the authors dissect how Gβγ potentiates PLCβ3 signaling in cells. Using engineered crosslinking to stabilize a Gβγ-PLCβ3 complex, single particle cryo-EM, and cell-based functional assays, they identify and map multiple putative Gβγ interaction surfaces on PLCβ3, including a previously unrecognized binding mode. Structure-guided mutagenesis supports the functional relevance of these interactions and suggests that Gβγ potentiation is not primarily mediated by PLCβ3 membrane recruitment, but instead enhances PLCβ3 activity after the lipase is already at the membrane.

      Previous reconstitution work on the membrane surface (Falzone & MacKinnon, 2023) proposed a recruitment/partitioning-centric model in which Gβγ increases PLCβ3 output largely by elevating its membrane surface concentration, whereas Gαq primarily increases catalytic turnover; under those reconstitution conditions, the two inputs can combine approximately multiplicatively. In receptor-driven cellular signaling, however, PLCβ3 is robustly recruited to the plasma membrane upon Gαq activation, which raises the question of whether Gβγ contributes mainly through additional recruitment or through a post-recruitment mechanism once PLCβ3 is already at the membrane.

      This manuscript helps address that gap by using membrane-anchored PLCβ3 and complementary cellular readouts to separate "getting PLCβ3 to the membrane" from "boosting activity once PLCβ3 is already there." Their results argue that, in cells, membrane recruitment is largely dominated by Gαq·GTP, while Gβγ can further potentiate PIP2 hydrolysis after membrane association, consistent with a modulatory role at the membrane rather than primary recruitment.

      Overall, the work provides a structural and mechanistic framework for Gβγ-PLCβ3 cooperation and helps clarify the basis of Gq pathway amplification. The manuscript is generally strong, but some issues need to be addressed.

      Thank you for the positive comments.

      Major comments:

      (1) BMOE/BM(PEG)2 crosslinking may enforce a non-native docking geometry, potentially compromising the physiological relevance and precision of the Gβγ-PLCβ3 interface as described. Although a >50% 1:1 crosslinked complex is formed and remains active, the solution maps show lower local resolution for Gβγ, consistent with a dynamic, potentially heterogeneous, interface. One interface is captured via a single engineered cysteine pair (PLCβ3 E60C-Gβ C271), which could potentially bias the pose. It would be helpful if the authors could provide additional orthogonal support (e.g., alternative crosslinked sites) and bolster the clarification of its uniqueness and relevance.

      We did attempt to isolate other crosslinked complexes. PLCβ3-D892 self-crosslinked under all reaction conditions, while PLCβ3-D892 XY<sub>Cys</sub>, which retains an endogenous cysteine within the X–Y linker (C516), did not result in any crosslinked product when incubated with Gβγ. Only the PLCβ3-D892 E60C crosslinked to Gβγ. With the exception the C68S mutation at the C-terminus of Gg to eliminate its prenylation site, all endogenous cysteines were retained in both Gβ and Gγ. Indeed, Gβ contains two solvent-exposed cysteines in its canonical effector binding surface (C204 and C271), but we did not observe any crosslinker density involving C204. While we cannot exclude the possibility that crosslinking occurred between PLCβ3-D892 E60C and other residues in Gβγ, we were unable to identify any 2D classes corresponding to these alternative conformations. These observations, together with the high efficiency of crosslinking, are consistent with a stable and persistent interaction.

      (2) In the crosslinked structure, the authors report that GβD228 interacts with PLCβ3 R199 and K183. In Figure 2A, R199 appears closer to Gβ D228 than K183, yet only K183 is functionally tested. Testing R199 (e.g., R199E/R199A) would strengthen the structure-guided validation of this interface.

      We agree, and functional analysis of PLCb3 R199E is included in the revised manuscript (see Figures 2-4).

      (3) The mutagenesis strategy appears inconsistent across figures/assays, which makes it difficult to interpret phenotypes and directly link the functional data to the proposed interfaces. For example, in Figure 2E, we see R185L but R215E, while residue L40 is mutated to Gly in the IP accumulation assays but to Glu/Lys (L40E/K) in the BRET assays (Figures 3B/3D/3F). The authors should (i) clearly justify the rationale for each substitution (conservative vs charge-reversal, interface disruption, etc.) and (ii), where possible, test the same mutants across assays (or provide evidence that alternative substitutions yield consistent conclusions).

      Mutagenesis experiments were initially carried out independently in the Lambert and Lyon Labs. As the study progressed, additional mutants were identified and/or designed based on results from both groups. The residues subject to mutagenesis are overall consistent across the different assays, with differences in the identity of the mutation varying in some cases. The L40G mutation is one such example, where given its modest impact on Gβγ-mediated activation in the IP accumulation assay, more impactful changes were made (L40E and L40K) for the BRET and signaling assays. In the revision, we now state that mutations were designed to maximally disrupt the three observed interfaces, such as by changing the size of the side chain and/or introducing charge reversal mutants.

      Reviewer #3 (Public review):

      Summary:

      PLCβ3 is activated by both Gαq and Gβγ subunits. This paper follows previous solutions and cryoEM studies of PLCβ3 / Gβγ, trying to understand the molecular details of activation using cellular BRET assays and cryoEM.

      Strengths:

      The authors find evidence for multiple binding sites on PLCβ3 for Gβγ and suggest that Gβγ is not bone fide activator per se but enhances Gαq activation by positioning the catalytic site towards substrate, although this is not completely convincing. Although these sites may not naturally be operative, the authors might want to develop the potential role of these sites.

      The authors also find that this activation is not through recruitment of the enzyme to the membrane by Gβγ released upon G protein activation, in accord with other PLCβ enzymes, but not for PLCβ3, and again, the authors might want to develop this point further.

      Thank you for the suggestions. We are investigating whether the other PLCb isoforms contain multiple Gβγ binding sites and the relative importance of preactivation by Ga<sub>q</sub> for a manuscript in preparation.

      Weaknesses:

      (1) I'm confused as to why the authors feel that their mechanism is distinct from the two-state enzyme, the synergistic activation proposed by Ross in 2011, using a primarily thermodynamic argument. As written, the authors appear to be very reliant on structural and BRET studies that do not give the details that would disprove this interpretation. The main issue is that the author's mechanism does not fully explain how Gβγ activation occurs for PLCβ2 in reconstituted systems in the absence of Gαq subunits.

      The reconstitution experiments are under extremely artificial conditions, using nM-µM of purified proteins and liposomes that contain up to 30% PI(4,5)P2. Under these conditions, we think the increased activity is due to interfacial activation promoted by Gβγ binding to the lipase once it is associated with the liposome surface. This would be sufficient to account for the dose-dependent increase in both PLCb2 and PLCb3 activity as a function of Gβγ concentration. Given the higher basal activity of PLCβ2 and its decreased sensitivity to activation by Ga<sub>q</sub>, one possible explanation is that this isoform differs in its autoinhibition and/or structure of its proximal CTD that Ha2’ displacement is not a prerequisite for activation. In addition, Gβγ may also be a direct activator of PLCβ2. Further studies, ideally in cell-based systems, are needed to answer these questions.

      (2) In a recent study, McKinnon presents a model showing that Gαq and Gβγ activate PLCβ3 by two distinct pathways and that activation by Gβγ occurs through membrane recruitment. It is not surprising that the authors find that this is not true since the pelleting method used by McKinnon is subject to error. The authors should directly address the limitations of this previous work and the changes in proteoliposomes with sedimentation that alter partition coefficients. Although the inability of Gβγ to drive membrane binding is in accord with the quantitative studies of Scarlata, showing that the affinity of PLCβ3 to Gβγ is fairly weak as compared to the intrinsic membrane partition coefficient.

      We have added some of the limitations of proteoliposome sedimentation experiments to the discussion.

      (3) It was proposed many years ago that in signaling complexes Gαq - Gβγ may not have to fully dissociate when binding PLCβ, but rather shift their relative orientation when binding to PLCβ to allow activation. Is their model consistent with this? Is it possible that PLCβ3 keeps Gβγ from diffusing to enhance the rate of Gq / Gβγ re-association?

      Our crosslinked complex is compatible with simultaneous binding of a Gα<sub>q</sub>-Gβγ heterotrimer to the PLCb3, without disrupting the observed interface. If Gαq were to interact with the Gβγ molecules bound to the PH or EF hands, the interaction would be mediated by the N-terminal helix of Gα<sub>q</sub>. It is possible Gβγ–PLCβ3 interactions may slow heterotrimer reassociation, but this may be complicated by the intrinsic GAP activity of the lipase.

      (4) The authors find that Gβγ binds multiple sites, and it is clear that the PH domain site is the primary one in accord with previous work. Could these weaker sites be an artifact of the elevated concentrations used in cryoEM and BRET assays?

      While more studies have focused on the PH domain as a Gβγ binding site, our data does confirm the EF hands are also functionally relevant. To our knowledge, the role of the EF hands has not been investigated in this capacity until very recently, and so we hesitate to label them primary or secondary. It is possible the EF hands may be a lower-affinity site for Gβγ and the protein concentrations needed in cryo-EM drive complex formation. However, it is also possible the concentration of free Gβγ adjacent to an activated receptor may be high enough to saturate the PH and EF hand binding sites.

      (5) Although their assays infer differences in binding affinities, it would strengthen the paper if the authors could estimate the association energies of these different binding sites. This estimation would also address the concern stated above.

      We appreciate this suggestion and quantifying the affinities of the Gβγ–PLCβ3 interactions is the subject of future studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please correct PIP2 to the correct PIP2 (many examples throughout).

      These have been corrected.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) Figure S1B: The lane-condition labels above the third gel appear to be incorrect, as both lanes are marked identically (Gβγ +/+, PLCβ variant +/+, BMOE +/+) despite clearly different banding patterns. Please confirm.

      We have confirmed the markings above the gels are correct.

      (2) In Figure S2, the authors show three fitted models, but the helical density for Gβγ cannot be seen in two of them at the displayed contour level. The authors should provide views of the map at different contour levels (thresholds) to better support the model fitting. Otherwise, it is difficult to assess whether the Gβγ subunit could adopt alternative orientations (i.e., whether it may be rotated) within the density.

      We have included a new figure (Figure S3) that provides images of the maps at different contour levels.

      (3) Page 5: "where PLCβ3 is increased by the overexpression of either Gβγ or Gαq" should be revised to "where PLCβ3 activity is ...".

      This sentence has been corrected.

      (4) Figure 2: Please label residue R215 in Figure 2A/2B (or the relevant structural panel), since R215E is tested in 2E but the position is not shown.

      R215 is now included in Figure 2C.

      (5) Page 19, Figure 2 legend: "Changes ... Figure S3" should be "Changes ... Figure S5".

      We have corrected this figure call.

      Reviewer #3 (Recommendations for the authors):

      The studies seem well carried out, although more details regarding the BRET controls and the significance of the values should be included.

      We have revised the captions to provide more details about the experimental controls and a brief description of significance. Individual p-values are included in the supplemental tables.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study reports a novel and potentially impactful role for NINJ2 in maintaining lysosomal integrity and regulating cellular susceptibility to ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes and interacts with LAMP1, a key lysosomal membrane glycoprotein involved in sensing lysosomal stress. Loss of NINJ2 increases lysosomal membrane permeabilization (LMP), resulting in selective leakage of lysosomal contents, including labile iron, into the cytosol. The authors further show that NINJ2 deficiency reduces the expression of ferritin storage proteins, thereby sensitizing cells to ferroptosis induced by RSL3 and erastin. Collectively, the work proposes a mechanistic link between NINJ2-mediated control of LMP, iron homeostasis, and ferroptotic vulnerability, with potential relevance to cancer biology.

      Strengths:

      This study identifies a novel role for NINJ2 in regulating lysosomal integrity and ferroptosis and establishes a mechanistic link between lysosomal membrane permeabilization, iron homeostasis, and ferroptotic sensitivity, with potential translational relevance in cancer.

      Weaknesses:

      (1) The results overall support the authors' conclusions and provide a plausible mechanistic framework; however, additional quantification of Western blot data and further discussion of mechanistic questions would strengthen the study.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      (2) The findings are likely to have a broad impact by linking lysosomal integrity to ferroptosis and iron homeostasis, both of which are relevant to cancer biology and therapeutic targeting.

      We thank the reviewer’s comment. We have discussed the potential implications of these findings for cancer treatment in the “Discussion” section.

      Reviewer #2 (Public review):

      This manuscript, "Nerve Injury-Induced Protein 2 preserves lysosomal membrane integrity to suppress ferroptosis", identifies a previously unrecognized function of NINJ2 as a regulator of lysosomal membrane integrity and iron homeostasis, thereby suppressing ferroptosis. The authors demonstrate that NINJ2 localizes to lysosomes, interacts with LAMP1, limits lysosomal membrane permeabilization (LMP), stabilizes ferritin, and protects cells from ferroptotic cell death. They further extend these mechanistic findings to human cancer datasets, showing cooverexpression and positive correlation of NINJ2 with ferritin genes in iron-addicted cancers.

      Overall, the study is conceptually interesting, technically solid, and integrates cell biology, iron metabolism, and ferroptosis in a coherent framework. The work expands the functional repertoire of the Ninjurin family beyond plasma membrane rupture and inflammation, which will be of interest to researchers in cell death, lysosome biology, and cancer metabolism.

      Strengths:

      (1) The identification of NINJ2 as a lysosome-associated protein that suppresses ferroptosis represents a meaningful advance beyond its previously described roles in inflammation, pyroptosis, and tumorigenesis.

      (2) The work distinguishes NINJ2 functionally from NINJ1, reinforcing the idea that structurally related Ninjurins have divergent membrane-related roles.

      (3) The study presents a logically connected pathway:

      NINJ2 loss → LMP → labile iron increase → ferritin degradation → ferroptosis sensitization, which is well supported by the data.

      (4) The link between LAMP1, ferritin turnover, and ferroptosis is particularly compelling and timely given recent interest in lysosomal contributions to ferroptotic signaling.

      (5) The authors use confocal microscopy, proximity ligation assays, biochemical IPs, iron measurements, protein half-life analyses, ferroptosis assays, and TCGA-based analyses, providing convergent evidence for their model.

      (6) Use of two distinct cell lines (MCF7 and Molt4) strengthens generalizability.

      (7) The integration of cancer expression datasets linking NINJ2 with ferritin expression in hepatocellular and breast carcinomas enhances translational relevance.

      (8) Assigning NINJ2 a lysosomal protective function, distinct from NINJ1-mediated plasma membrane rupture, is novel.

      (9) Linking NINJ2 to ferroptosis regulation via lysosomal iron handling, rather than canonical GPX4 or system Xc</sup>-</sup> pathways, is also novel, along with proposing a NINJ2-LAMP1-ferritin axis as a buffering mechanism against iron-driven lipid peroxidation.

      (10) These insights are not incremental; they reframe how NINJ2 may function at the intersection of membrane biology, iron metabolism, and regulated cell death.

      Areas for improvement:

      While the study is strong, several issues should be addressed for mechanistic depth and general relevance.

      (1) Although NINJ2 is shown to interact with LAMP1 and LAMP1 knockdown rescues ferritin levels, it remains unclear whether the NINJ2-LAMP1 interaction is required for lysosomal protection. The authors could: a) Map the NINJ2 domain required for LAMP1 interaction and test whether an interaction-deficient mutant fails to protect against LMP and ferroptosis. b) Rescue NINJ2 KO cells with wild-type versus mutant NINJ2 to establish causality.

      We thank the reviewer’s comments. Ongoing work in our laboratory is focused on elucidating the molecular mechanism by which the NINJ2-LAMP1 interaction regulates lysosomal membrane integrity and ferroptosis, and we anticipate reporting these findings in a future publication.

      (2) The conclusion that NINJ2 suppresses ferroptosis relies primarily on RSL3 and Erastin sensitivity. A direct assessment of ferroptosis would hence the study, such as:

      (a) Include ferroptosis rescue experiments using ferrostatin 1 or liproxstatin 1.

      (b) Assess lipid peroxidation directly (e.g., C11 BODIPY staining) to strengthen the ferroptosis claim.

      We thank the reviewer for this thoughtful comment. We agree that ferrostatin-1 or liproxstatin-1 rescue experiments, together with direct analysis of lipid peroxidation, would provide complementary evidence for ferroptosis. We will incorporate these additional experiments in future studies to further strengthen the mechanistic basis of our findings.

      (3) The manuscript discusses lysosomal ferritin degradation but does not directly examine NCOA4, a central mediator of ferritinophagy. It would be good to: a) Test whether NCOA4 knockdown rescues ferritin loss and ferroptosis sensitivity in NINJ2 KO cells. b) This would clarify whether NINJ2 acts upstream of canonical ferritinophagy pathways or via an alternative mechanism.

      We appreciate the reviewer's thoughtful suggestion. Defining the contribution of NCOA4 to NINJ2-mediated ferritin degradation is an important question that could further clarify the underlying mechanism. Addressing this issue will require a comprehensive set of additional experiments, which will be addressed in the future studies.

      (4) The study is entirely cell-based, despite references to inflammatory and tumor phenotypes in Ninj2-deficient mice. While not strictly required, even limited in vivo validation (e.g., ferroptosis markers or iron accumulation in existing Ninj2 KO tissues) would substantially strengthen the manuscript.

      We thank the reviewer for this insightful suggestion. We agree that in vivo validation of ferroptosis markers and iron accumulation in Ninj2-deficient tissues would further strengthen our conclusions. However, these experiments will require substantial additional investigation, which will be pursued in the future studies.

      (5) Finally, most imaging data (e.g., Galectin 3/LAMP1 colocalization, PLA signals) and immunoblot data are presented qualitatively. The authors should provide the qualifications of Western blots and other measurements.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) What mechanisms might underlie the regulation of LAMP1 transcript levels by NINJ2?

      A clear mechanism by which NINJ2 regulates LAMP1 transcripts has not been elucidated and warrants further investigation. Nevertheless, several possibilities can be considered. First, LAMP1 transcription is known to be regulated by TFEB (transcription factor EB), a master regulator of the lysosomal–autophagy pathway. Upon lysosomal membrane permeabilization (LMP), TFEB translocates to the nucleus and activates a broad set of lysosome-related genes, including LAMP1. Notably, phosphorylation of TFEB by mTORC1 at Ser211 inhibits its activity by preventing nuclear translocation. Thus, it would be of interest to determine whether NINJ2 modulates TFEB phosphorylation status and subcellular localization. In addition, the tumour suppressor p53 has been reported to engage in complex crosstalk with TFEB in regulating basal autophagy. Interestingly, p53 expression is increased in NINJ2-KO cells. It is therefore plausible that NINJ2 regulates TFEB activity through p53, or alternatively modulates the p53–TFEB signaling axis more broadly to maintain lysosomal integrity.

      (2) Does Ninjurin1 play a similar role in regulating lysosomal membrane permeabilization (LMP)?

      At this moment, it remains unclear whether NINJ1 plays a role similar to that of NINJ2 in regulating LMP. In fact, our previous studies demonstrated that NINJ2 physically interacts with NINJ1 and may antagonize NINJ1-mediated pyroptosis. Furthermore, NINJ1 has recently been reported to promote ferroptosis by interacting with the xCT cystine/glutamate antiporter (PMID: 38464226), a function that contrasts with the protective role of NINJ2 against ferroptosis identified in the present study. These findings suggest that NINJ1 and NINJ2 may have distinct, or even opposing, functions in regulating cell death pathways. Nevertheless, further studies are required to determine whether NINJ1 also participates in the regulation of LMP and to define its relationship with NINJ2 in maintaining lysosomal membrane integrity.

      (3) What are the potential clinical implications of these findings, particularly in the context of cancer progression or therapeutic targeting?

      Targeting NINJ2 may have important clinical implications in cancer therapy. Given its role in maintaining lysosomal membrane integrity, inhibition or loss of NINJ2 could promote lysosomal membrane permeabilization (LMP), thereby sensitizing cancer cells to ferroptosis through increased intracellular labile iron accumulation and disruption of redox homeostasis, ultimately enhancing tumor cell killing. As such, NINJ2 inhibition may represent a strategy to selectively destabilize lysosomal function in cancer cells and improve responsiveness to ferroptosis-inducing agents or other combination therapies that exploit oxidative stress vulnerabilities. Indeed, we previously developed a peptide that targets NINJ2. Whether this NINJ2-targeting peptide can sensitize cancer cells to ferroptosis therefore warrants further investigation.

      (4) The authors should provide quantification of all Western blot data throughout the manuscript to enhance data robustness and reproducibility.

      All the western blot data have been quantified throughout the manuscript in the revised version.

      Reviewer #2 (Recommendations for the authors):

      (1) Controls for knockdown efficiency of NINJ2 (Figure 2D) should be shown.

      NINJ2-KO MCF7 cells were generated previously and published in the article (PMID: 38325550) along with sequencing confirmation.

      (2) In Figure 2A legends, the concentration of LLOMe is 1mM or 1µM - need to be clarified?

      The concentration for LLOME is 1µM. This typo has been corrected.

      (3) In the figure legends section, "Figure 4" is missing.

      We thank the reviewer’s comment. Figure 4 has been added to the Figure legends.

      (4) Some description of NINJ2 ko cells generation should be included in the materials section.

      In the Materials and Methods section, we briefly described how these cell lines were generated and cited the original publication (PMID: 38325550)

      (5) The manuscript would benefit from a schematic model figure summarizing the proposed NINJ2-LAMP1-iron-ferroptosis axis.

      In the revised manuscript, we provided a model to elucidate the role of NINJ2 in modulating lysosomal membrane integrity and iron homeostasis.

      (6) Some sections of the Introduction are lengthy and could be streamlined to focus more directly on lysosomes and ferroptosis.

      We have streamlined the introduction.

      (7) Statistical methods should clarify whether data meet assumptions for Student's t-test and whether multiple comparisons were corrected where applicable.

      Statistical analysis has been added to the figure legends when applicable.

    1. Author response:

      We thank the editors and three reviewers for their careful evaluation and constructive feedback. We are pleased that our identification of IR20a as a multimodal tuning receptor required for both low-salt and arginine sensing in a distinct gustatory neuron population was recognized as a valuable contribution to sensory coding. We agree with the major points raised and outline our planned revisions below, organized thematically.

      (1) Receptor assembly and integration model

      Our genetic, heterologous, and calcium imaging data show functional cooperation among IR20a, IR25a, and IR76b but do not demonstrate physical association. In the revision we will replace terms like “distinct subunit assemblies” and “peripheral integration” with more cautious language such as “functional receptor combinations.” We will state explicitly that our data cannot resolve whether the three IRs form a single heteromeric complex, and we will discuss the alternative possibility that IR76b and the IR20a/IR25a pair function as separate receptors within the same neuron. This interpretation better accounts for the response patterns observed upon co-expression. Throughout the text and in a revised summary figure, we will clearly differentiate elements directly supported by data from those that remain inferential.

      (2) Tarsal calcium imaging versus behavior

      The mismatch between tarsal calcium imaging and behavioral arginine responses will be addressed directly. We will explain that tarsal recordings sample only a small subset of IR20a neurons, whereas proboscis extension and feeding assays predominantly engage the more numerous labellar sensilla, whose neurons may carry different receptor compositions. We will generate labellar imaging where feasible; if additional functional data cannot be obtained, we will acknowledge this limitation rather than overinterpreting the tarsal results.

      (3) Synergy versus additivity

      We will discuss behavioral and cellular data separately. Recognizing that true synergy is difficult to demonstrate in feeding and proboscis extension assays, we will adopt conservative terminology when interpreting those experiments. For cellular data, where mechanistic insight is stronger, we will present the evidence for functional synergy and discuss why the outcomes may differ between the cellular and organismal levels.

      (4) Contextualizing prior IR76b and IR20a literatures

      We will expand the Introduction and Discussion to cite more fully the works establishing IR76b as a low-salt sensor and IR20a’s role in amino-acid sensing, including the earlier report that IR20a overexpression can inhibit IR76b-dependent salt responses. We will clarify how our single-cell imaging and loss-of-function data obtained in the native context refine models derived from ectopic expression. In its endogenous setting, IR20a marks neurons narrowly tuned to amino acids such as arginine, and IR20a is strictly required for low-salt detection. These findings contrast with the earlier view that IR20a functions broadly as an amino-acid sensor or as a salt-response blocker.

      (5) State-dependent modulation and IR56b

      Our data do not identify IR56b as the direct molecular sensor of internal state. We will reframe IR56b as a necessary component for state-dependent modulation of low-salt preference and retract any claim that it is the sensor itself. We will also clarify that IR20a and IR56b define genetically separable peripheral pathways that may converge on downstream circuits.

      (6) Methodological and presentation issues

      We will address the following points raised across reviews:

      (1) Co-localization: Higher-resolution confocal images and co-localization analysis will be provided.

      (2) Summary model figure: A new figure will illustrate the distinct functions of IRs in low-salt and amino-acid taste, clearly indicating which aspects are directly supported and which are inferential.

      (3) Feeding-assay control: We will either include an isosmotic sucrose control to avoid the water confound or explicitly discuss this limitation and temper the interpretation of feeding-preference results.

      (4) S2 cell quantification: Complete details on response criteria, responder fractions, and statistical reporting will be added.

      (5) Figure and supplementary corrections: All noted errors in figure legends, scale bars, citations, and supplementary-file mismatches will be fixed.

      We are confident these revisions will bring our mechanistic claims into close alignment with the evidence and substantially improve the manuscript. We again thank the editors and reviewers for their detailed and helpful comments.

    1. Author response:

      We thank the editors and reviewers for recognizing the originality and potential value of the proposed framework. The reviews rightly ask us to distinguish three things more sharply: what is established directly in human breast tissue, what is inferred from mammary and aging studies, and what remains hypothesis. We agree, and the revision will make that distinction explicit throughout.

      We will define the proposed reserve state operationally and specify the findings that would distinguish active niche maintenance from passive persistence. Throughout, we will treat passive persistence as a legitimate competing hypothesis rather than a settled question. Heterogeneity and immune or inflammatory associations will be presented as consistent with active maintenance and causal directionality, not as establishing them. We will also integrate the mixed epidemiologic evidence more centrally into the argument, rather than treating it as a caveat.

      We will reassess the evidence for senescence in the aging breast and describe it more precisely, correcting or narrowing statements that outrun the data. The figures will be revised so that observed inflammatory and immune-regulatory features are clearly separated from proposed senescence- and SASP-mediated mechanisms. We will also clarify the limits of our analogies: postpartum involution and cross-tissue reserve systems will be presented as sources of candidate mechanisms and testable predictions, not as direct evidence for the proposed mechanism in human ARLI. Finally, we will frame the translational implications as contingent. They depend on first demonstrating that senescent cells are enriched near persistent lobules, identifying the relevant cell types and immune states, and establishing causal relevance.

      We appreciate the reviewers' constructive suggestions. We believe these revisions will preserve the conceptual contribution of the model while making its evidentiary status, the alternative explanations, and its falsifiable predictions substantially clearer.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      I thank the authors for the revised manuscript and for the detailed responses.

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

      Thank you for your critical review and insightful comments.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid, and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei support the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      (b) Published work links PLK to cell division, FAZ elongation, etc... The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc....

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      - The authors have now addressed this question by assessing what % of KING is phosphorylated at T301 and adding commentary on this point in the revised paper.

      - I would suggest that the model (new figure 8) include a dephosphorylation step, as that is proposed by the authors in the text. Also include in the legend some commentary on the role of phosphorylation, which is the center point of this paper, but not currently mentioned.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function. Thank you.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      - The authors have addressed this question by demonstrating that T301 phosphorylation is reduced upon treatment with a PLK inhibitor, thus supporting that PLK phosphorylated T301 in vivo. It is noted that one might consider an alternate kinase is also able to phosphorylate T301 in absence of PLK activity, as that could explain the relatively low (~27%) reduction in phosphorylation by PLK inhibitor treatment.

      Thank you.

      Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      The authors have addressed prior weaknesses in the manuscript through additional experimentation and rewording of the conclusions.

      Thank you for your critical review of our manuscript and for the very constructive comments and suggestions to improve the manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (There is some redundancy below with my comments in the public review, but I've included here for clarity and further explanation.)

      The authors have addressed my primary concern, as treatment with PLK inhibitor reduces phosphorylation of T301, while also providing some comment on relative impact of PLK-mediated KIN-G phosphorylation.

      It is notable that phosphorylation of T301 was reduced by only ~27%, while phosphorylation of S569 was reduced by ~100% in the presence of PLK inhibitor. The authors note that this might be explained by slower dephosphorylation of T301. In the absence of a phenotype, and with cell doubling continuing unabated in presence of the inhibitor, it is intriguing that more loss is not observed. An alternative explanation is that an alternate kinase might also be able to phosphorylate T301 in the absence of PLK activity, and the authors should consider that possibility.

      We added a sentence in the main text to suggest an alternative explanation.

      The model shown in figure 8 should include a dephosphorylation step, per the authors comments in the text regarding the small fraction of T301 that is phosphorylated and proposal of a phosphorylation/dephosphorylation cycle. The Fig 8 legend needs to have some commentary on the role of phosphorylation, as phosphorylation is the center point of this paper.

      We have modified the model in Figure 8 to include dephosphorylation by an unknown protein phosphatase and a statement about the role of TbPLK phosphorylation on KIN-G function.

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Fig 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Fig 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces motility of KIN-G. These statements are contradictory.

      (3) p. 8, and Fig 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      (4) Fig 7B. Please explain labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      (5) Fig 4, 5, and 7: "% Cells" is reported. Please indicate what number of cells total were examined.

      These minor comments have already been addressed in the previous revision.

    1. Author response:

      Reviewer #1 (Public review):

      R1-Q1: The manuscript frequently presents the relationship between spontaneous gamma oscillations and pain perception as established fact. Given the continuing debate regarding the functional significance of EEG gamma oscillations in pain processing, these statements should be moderated.

      We thank the reviewer for raising this important and thoughtful point. We agree that the relationship between spontaneous gamma oscillations and pain perception remains a matter of active debate, and we will moderate these statements throughout the manuscript. We will acknowledge the ongoing debate and present the gamma-pain relationship as an active area of investigation rather than settled fact.

      R1-Q2: The responder analysis is the most serious methodological concern. Participants in the active group were retrospectively classified as 'responders' based on increased gamma power after neurofeedback, and only these participants appear to have been included in the primary analyses and matched to sham participants. As only 23 of 44 participants (52%) met this criterion, the responder rate alone does not demonstrate successful neurofeedback-induced gamma modulation. More importantly, selecting participants based on the outcome variable and subsequently testing that same outcome constitutes circular analysis (double dipping), invalidating the statistical inference. Consequently, the reported effects should be interpreted as an association within a post hoc selected subgroup rather than evidence that neurofeedback increased gamma activity and reduced pain.

      We appreciate this careful critique. We wish to clarify the rationale behind our analytical approach and address the concern.

      A well-established finding in the neurofeedback literature is that a substantial proportion of participants are "non-learners" — individuals who, despite receiving real feedback, fail to achieve effective control over the targeted neural activity. This is not a failure of the intervention, but reflects individual differences in neurofeedback learning capacity. Our core research question is therefore: "Among individuals who can successfully learn to upregulate gamma oscillations, does this upregulation reduce pain perception and nociceptive brain responses?"

      To address the circularity concern and improve transparency, we will make the following revisions:

      - We will reframe the wording from "NFB increases gamma and reduces pain" to "Successful gamma upregulation via NFB is associated with reduced pain in those who achieve it." All causal language will be replaced with appropriately cautious, correlation-based terminology.

      - We will report full-sample results for completeness.

      R1-Q3: The criterion for successful neurofeedback-induced gamma modulation was not prespecified. It is therefore unclear whether successful modulation was defined by the responder classification, the main effect of session, the group × session interaction, or one of the post hoc comparisons.

      We thank the reviewer for requesting this clarification. We will specify the exact criterion in the revised manuscript, i.e., a participant was classified as a responder if their post-intervention gamma power minus pre-intervention gamma power was positive (i.e., an increase in gamma power following the neurofeedback intervention).

      R1-Q4: Several methodological details reduce the reproducibility and replicability of the study. The spectral analysis does not clearly describe how trial-wise power estimates were aggregated within participants before group-level analyses, and the preprocessing pipeline includes manual ICA-based artifact rejection without specifying the criteria used for component selection. In addition, the analysis pipeline and custom neurofeedback software should be made publicly available to enable independent reproduction and verification of the reported findings.

      We thank the reviewer for these constructive suggestions. We will supplement and refine the methodological details in the revised manuscript, and we will make the analysis code and the experimental program (including the custom neurofeedback software) publicly available via an open repository.

      R1-Q5: Updating the feedback only once per second using a 2-s sliding window results in discontinuous visual feedback that may reduce feedback quality and could introduce visually evoked activity. In addition, the viewing distance of approximately 30 cm likely required substantial eye movements while following the moving feedback object.

      We will discuss the limitations of the discontinuous visual feedback and the viewing distance in the revised manuscript. We acknowledge these as valid methodological concerns and will address them as limitations in the Discussion.

      R1-Q6: The muscle-confound analysis is insufficiently documented. EMG recordings were acquired only in the second cohort, but the manuscript does not clearly state how many participants contributed to this analysis or whether responder selection was performed before or after restricting the sample. These details should be explicitly reported.

      We thank the reviewer for pointing out that the description of the muscle-confound analysis was insufficiently detailed. In the revised manuscript, we will clarify the EMG analysis procedures and explicitly report: (a) the exact number of participants contributing to the EMG analysis; (b) the cohort from which they were drawn; and (c) whether responder selection was performed before or after restricting the sample for EMG analysis.

      Reviewer #2 (Public review):

      R2-Q1: The most critical issue is about the exclusion of non-responders from the analysis. I find this problematic, as the reasoning becomes circular (selecting the participants who managed to increase gamma and then asking whether gamma NFB influenced pain), effect sizes are inflated, and the selection itself may introduce biases. For example, the responders may differ from the non-responders with respect to other characteristics (better attention skills, better self-regulation, etc). It would be more principled to present the results for the entire sample and only present the responder analysis as a secondary analysis. In the preregistration, the responder-only analysis was not mentioned.

      As detailed in our response to R1-Q2, the responder analysis reflects a conceptually motivated subgroup defined by successful neurofeedback learning — a well-documented challenge in NFB research where many participants are non-learners. Our central question is whether successful gamma upregulation (among those capable of achieving it) is associated with pain reduction. We will make this rationale explicit in the revised manuscript. We will also: (a) transparently report full-sample results; (b) discuss potential biases introduced by subgroup selection (e.g., differences in attention, self-regulation); and (c) acknowledge the lack of preregistration for the responder analysis.

      R2-Q2: Another critical point is about the causal claims made in the abstract, introduction, and discussion. Given that the current results provide only correlational evidence in a subsample, the language should be revised to avoid overinterpretation. If the authors can demonstrate a significant mediation effect (NFB group → gamma change → pain change), they may be able to argue that increases in gamma activity mediate the observed reduction in pain.

      We will substantially revise the language throughout the manuscript to avoid causal claims. We also plan to conduct a formal mediation analysis (NFB group → gamma change → pain change) to test whether changes in gamma activity statistically mediate the observed pain reduction.

      R2-Q3: For the sham procedure, the authors used the preceding participant's gamma data for feedback. This raises two questions: How was this handled for the first participant? Did the authors check the discrepancy between actual gamma and presented gamma in the sham NFB group?

      We thank the reviewer for raising this point, and we will clarify both points in the revised manuscript. (a) Because group assignment was randomized, the first participant could in principle have been assigned to the sham group. To prepare for this possibility, we collected EEG data from one participant in advance (equivalent to pilot data) to serve as the sham feedback signal, had the first participant been assigned to the sham group. In the actual experiment, however, the first participant was randomly assigned to the active group, so this pre-collected dataset was never used. (b) We will also compare the discrepancy between actual gamma power and the sham feedback signal in the sham group, and report this result in the revised manuscript.

      R2-Q4: Was baseline gamma power comparable between groups?

      We thank the reviewer for this suggestion. We will report and compare baseline gamma power between the active and sham groups in the revised manuscript.

      Reviewer #3 (Public review):

      R3-Q1: Gamma-band oscillations recorded with scalp EEG are difficult to measure reliably, are not observable in all participants, and can be strongly affected by muscle activity. The authors acknowledge this issue and include posterior neck EMG, but the control remains limited. A lack of correlation between one posterior neck EMG channel and Pz gamma power is not sufficient to exclude muscle contamination, especially because gamma-band artifacts can arise from multiple muscle groups and may not be well captured by a single EMG channel. This is particularly important because changes in posture, facial tension, breathing, and arousal could all influence high-frequency scalp activity.

      We thank the reviewer for this important suggestion. We will revise the manuscript to discuss more explicitly the inherent difficulty of recording pure gamma-band oscillations with scalp EEG. We will acknowledge that scalp gamma is not reliably observable in all participants, is vulnerable to contamination from multiple muscle sources, and that a single posterior neck EMG channel provides only limited control. We will also discuss the possibility that changes in posture, facial tension, breathing, and arousal may contribute to high-frequency scalp activity.

      R3-Q2: The evidence for a causal relationship between parieto-occipital gamma activity and pain perception should be interpreted cautiously. The authors show that gamma power increased in approximately half of the active neurofeedback participants and that these responders showed reduced pain ratings and laser-evoked potentials. However, because the main analgesic effect is tied to responder classification, it remains difficult to separate the specific effect of gamma upregulation from broader individual differences in task engagement, suggestibility, relaxation ability, attentional state, or neurofeedback learning capacity.

      We appreciate the reviewer's careful consideration of this point. We will temper our conclusions, presenting the current evidence as demonstrating the feasibility of gamma-band neurofeedback training in a subset of participants and an association between successful training and pain reduction, rather than a demonstrated causal relationship. We will discuss individual differences (attention, suggestibility, relaxation ability, neurofeedback learning capacity) as potential confounds that cannot be fully disentangled from gamma-specific effects.

      R3-Q3: It is not clear whether the experimenters were also blinded during data collection and interaction with participants. This matters because neurofeedback studies are particularly vulnerable to expectancy.

      We will clarify that the study employed a single-blind design: participants were unaware of their group assignment. We will state this clearly in the revised manuscript.

      R3-Q4: The choice of the two neurofeedback scenarios requires clearer justification. The manuscript describes a deep ocean scene followed by a seaside scene with relaxation instructions, but it is not clear why these two scenarios were selected, and whether they were matched for attentional engagement and affective content. This is not a minor point, because both groups showed reductions in pain ratings after the entire neurofeedback procedure.

      We will provide a stronger rationale for the selection of the two neurofeedback video scenarios. We will also place greater emphasis on the pain reduction observed in both groups, acknowledging the substantial nonspecific analgesic effects associated with the procedure.

      R3-Q5: The comparison with tACS is interesting but currently underdeveloped. The authors suggest that neurofeedback may succeed where gamma-frequency tACS failed because it allows real-time, personalized, self-regulatory modulation of ongoing activity. This is plausible, but the manuscript should discuss this distinction more deeply. Neurofeedback may not simply be a different way of modulating gamma; it may recruit volitional control, attentional engagement, immersion, expectation, etc. These mechanisms could be central to the observed pain reduction and may partly explain why neurofeedback effects differ from those of externally applied stimulation.

      We thank the reviewer for this insightful comment. We will expand the discussion of why neurofeedback may produce effects beyond those achieved by gamma-frequency tACS. In particular, we will elaborate on the potential contributions of volitional control, attentional engagement, immersion, expectation, and self-regulatory processes, and discuss how these factors may be central to the observed pain reduction rather than merely incidental to the gamma modulation.

      R3-Q6: Overall, this is an interesting study that introduces a promising neurofeedback approach for experimental pain modulation. The findings are encouraging, especially the convergence between subjective ratings and laser-evoked potentials in responders. However, the conclusions should be tempered. The current evidence supports the feasibility of training gamma-band activity in a subset of participants and suggests that successful training is associated with reduced experimental pain.

      We thank the reviewer for this balanced assessment. We fully agree that the conclusions should be tempered, and we will revise the manuscript accordingly to reflect that the current evidence supports feasibility and association rather than established causal efficacy.

      Summary

      In summary, the planned revisions include:

      (i) full-sample results reported;

      (ii) moderating causal language throughout and adding a formal mediation analysis;

      (iii) clearly specifying the responder criterion and adding this to the preregistration;

      (iv) providing complete methodological documentation and publicly releasing all analysis code;

      (v) expanding the Discussion to address limitations regarding scalp gamma measurement, EMG control, visual feedback, viewing distance, single-blind design, NFB scenario rationale, and nonspecific effects;

      (vi) adding analyses on baseline gamma comparability and sham-feedback discrepancy.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely thank you and the reviewers for the thoughtful evaluation of our manuscript and for the constructive comments and important suggestions. We are encouraged by the recognition that the study is valuable and of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution. We are also grateful that the reviewers acknowledged the integrative approach of our work, including fecal metabarcoding, behavioral assays, electrophysiological recordings, chemical analyses, and field observations.

      We have carefully considered all comments and have revised the manuscript accordingly. In particular, we have made the following major revisions:

      (1) We clarified the biological origin of limonene and expanded the discussion of possible bat-associated sources, including snout secretions and microbial contributions, while avoiding overinterpretation of limonene as an exclusively endogenous mammalian compound.

      (2) We added more detailed descriptions of contamination controls, including instrument cleaning procedures, materials used for housing and handling, and blank-control results, to address the concern that limonene could have originated from human-associated or environmental contamination.

      (3) We added individual-level odor data and revised the presentation of terpenoid profiles to show more clearly where limonene was detected and how its relative contribution varied among samples.

      (4) We clarified the rationale for the concentration choices used in the electrophysiological and field experiments, emphasizing that these assays were designed to test physiological detectability and functional sufficiency rather than to reproduce exact natural emission concentrations or determine response thresholds.

      (5) We revised our interpretation of limonene more cautiously. We now state that limonene is sufficient to trigger avoidance responses, while we acknowledge that its natural ecological specificity, concentration dynamics, and interactions with other bat odor components require further investigation.

      (6) We corrected the terminology related to limonene enantiomers. Because our GC-MS and GC-EAD analyses did not use an enantioselective column, we replaced “(-)-limonene” with “limonene” throughout the manuscript, figures, legends, and supplementary materials.

      (7) We improved the figures and supporting materials by moving the figure illustrating predator-prey relationship into the main text, adding electrophysiological traces from all tested crickets, revising figure legends, clarifying the rationale for comparing bat body odor with air controls, and providing additional chemical-identification details in the supplementary materials.

      (8) We checked and clarified statistical annotations, including the exact adjusted P-values for relevant comparisons.

      We believe these revisions have substantially improved the clarity, rigor, and balance of the manuscript. We hope that the revised manuscript and the detailed point-by-point responses satisfactorily address all concerns raised.

      We are confident that our study represents an important contribution, as it fundamentally expands the traditional acoustic-centred view of bat–insect interactions by demonstrating that crickets can use olfaction as a complementary sensory modality to detect bat odors and initiate avoidance behavior. The multidisciplinary evidence, integrating behavioral assays, electrophysiology, chemical profiling, and field validation, provides convincing support for this novel olfactory pathway.

      eLife Assessment

      This valuable study raises the intriguing possibility that crickets use bat-associated odors as cues of predation risk, extending the classic bat-insect arms race beyond its usual acoustic framework. The authors combine fecal metabarcoding, behavioral assays, electrophysiology, chemical analyses, and field observations to show that Loxoblemmus equestris avoids the odor of the insectivorous bat Scotophilus kuhlii, and that synthetic (-)-limonene can elicit antennal responses, avoidance in the laboratory, and reduced calling activity in the field. However, the evidence is currently incomplete because the identity, biological source, natural concentration, and ecological specificity of limonene as a bat-derived predator cue require stronger support, including clearer quantification, contamination controls, individual-level odor data, and evidence that crickets can distinguish bat-associated limonene from common environmental sources. The work will be of interest to researchers in sensory ecology, chemical ecology, predator-prey interactions, and bat-insect coevolution.

      We sincerely thank the editors for the positive and constructive assessment of our work. We greatly appreciate the recognition of our study’s value and its potential interest to multiple research communities.

      We have carefully considered all comments and have revised the manuscript accordingly. The revisions include clarifications of limonene’s biological origin and ecological specificity, strengthened contamination controls and individual-level odor data, clearer rationales for experimental concentration choices, and more cautious interpretation throughout the manuscript. All changes are addressed in detail in our point-by-point responses below.

      Importantly, the central conclusion of our study—that crickets can detect and avoid bat odor through olfaction, and that limonene is sufficient to trigger avoidance responses in both laboratory and field settings—remains robustly supported by the multidisciplinary evidence we present. We believe the revised manuscript now provides a clearer, more balanced, and scientifically rigorous account of our findings, and we hope it meets the standards of eLife.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript examines whether insects can use bat odor as a cue of predation risk. The authors focus on the insectivorous bat Scotophilus kuhlii and the cricket Loxoblemmus equestris. They first use fecal DNA metabarcoding to show that crickets are part of the bat's diet, and field surveys to show that L. equestris is abundant at local foraging sites. In laboratory Y-tube assays, the authors show that crickets strongly avoid air carrying bat body odor. Gas chromatography coupled with electroantennographic detection showed that cricket antennae respond to components of bat odor. Chemical analyses identified several volatile compounds, with 2,2-dimethylheptane and (−)-limonene associated with antennal responses. Further analyses suggested that snout secretions are likely to contribute to the bat's body odor. The authors then tested individual compounds. Among the commercially available candidates, (−)-limonene elicited a strong antennal response and was sufficient to cause avoidance in the olfactometer. In field plots, spraying (−)-limonene reduced cricket calling activity relative to pre-exposure levels, whereas calling increased in control plots treated with hexane. Overall, the study argues that crickets can detect a vertebrate predator through olfactory cues and that a single bat-associated volatile can trigger antipredator behavior.

      This is an interesting and enjoyable study that addresses an understudied aspect of predator-prey interactions. The manuscript is clearly written, the experiments are presented in a logical sequence, and the figures are crisp and easy to follow. I really appreciated the combination of behavioral assays, electrophysiology, chemical analysis, and field observations.

      We sincerely thank you for your very positive and encouraging evaluation of our work. We are delighted that you found the study interesting and enjoyable. We also appreciate your kind remarks about the clarity of the manuscript, the logical flow of the experiments, and the quality of the figures.

      We are especially grateful that you recognized the value of our integrative approach. Combining behavioral assays, electrophysiology, chemical analysis, and field observations was central to our study design, and we are pleased that this approach resonated with you.

      You provided an accurate and comprehensive summary of our work. You confirmed that our main narrative is clear and logically coherent. Specifically, you followed our progression from establishing the predator–prey relationship, to demonstrating olfactory avoidance, to identifying limonene as an active compound, and finally to validating its behavioral effects in both laboratory and field settings.

      We have carefully considered all your constructive comments and suggestions. We address them in detail in our point-by-point responses below. Your feedback has been extremely helpful, and we believe the revised manuscript is substantially stronger as a result.

      My main issue concerns the identity and biological origin of the proposed bat odor cue, (−)-limonene. Limonene seems like an unusual compound to be emitted endogenously by a mammal, particularly by an insectivorous bat. It would be helpful if the authors could clarify whether mammals are known to synthesize this compound de novo, and, if not, what the likely source of this plant-associated terpene would be in S. kuhlii. Possible sources could include environmental exposure, diet, roosting material, handling, or temporary housing conditions.

      I do not doubt that crickets avoid synthetic (−)-limonene. Indeed, this result is quite plausible given that limonene is widely used in insect repellent or repellent-associated fragrance products. However, this also makes contamination an important issue to address explicitly. How did the authors exclude the possibility that limonene entered the samples from human-associated sources, such as insect repellents, soaps, cleaning products, field equipment, cloth bags, cages, gloves, or other materials used while handling wild-caught bats? It would strengthen the manuscript to report limonene levels for individual bat odor collections, all relevant blanks, and any handling or housing controls.

      More broadly, given the common occurrence of limonene in plants and human-associated products, I am not yet convinced that it would function as a reliable "keystone kairomone" as suggested around line 253. How would crickets distinguish bat-associated limonene from limonene emitted by a mint leaf, citrus peel, pine material, or other non-threatening environmental sources? The authors may wish to soften this interpretation or provide additional evidence that crickets respond to limonene in a bat-specific context, perhaps through concentration, temporal patterning, co-occurring volatiles, or enantiomeric composition.

      We sincerely thank you for your critical and constructive comments. Your questions regarding the identity, biological origin, and ecological specificity of limonene are insightful and have helped us substantially strengthen the manuscript. We address each of your points below.

      On the biological origin of limonene and whether mammals synthesize it de novo

      You raised an important point that limonene seems unusual for a mammal to emit endogenously. We fully agree. We agree that direct evidence for de novo synthesis of limonene in mammals is currently limited.

      However, limonene in bat body odor could still have biological origins. First, it may originate from skin- or gland-associated microbiota. Recent work has shown that skin-associated microorganisms can substantially shape bat volatile odor profiles (Sun et al., 2026, BMC Biology), and some microbes possess enzymes capable of terpene biosynthesis. Second, previous studies have independently reported limonene in the secretions of several bat species (Faulkes et al., 2019, PeerJ; Zhang et al., 2022, Ann. N.Y. Acad. Sci.). This suggests that its presence in bats is not unique to our study. Third, our own analyses detected limonene in hair and snout secretions, but not in faeces or blank controls. This pattern is consistent with a biological source associated with the body surface, rather than diet or environmental deposition.

      We have now expanded our Discussion to cover these possibilities more explicitly. We also emphasize that the exact source, i.e., endogenous, microbial, or otherwise, remains an open question that warrants future investigation. Please see lines 258–273 of the clean version of the revised manuscript, or see the excerpt below:

      “Although limonene reliably induced avoidance behaviour in crickets, two related questions still merit careful consideration. One question is whether the limonene we identified genuinely originates from bats or reflects contamination during sampling. Limonene is common in plants and numerous consumer products (Boncan et al., 2020; Schuman, 2023), making its endogenous production by a mammal seem unusual. Nevertheless, multiple lines of evidence militate against contamination. First, we adhered to rigorous protocols. For example, all instruments were cleaned with ethanol and oven-dried before each use; bats were housed in stainless-steel cages, and cloth bags had been rinsed with purified water. Second, limonene was absent from all blank controls, including empty-chamber air samples and clean swabs, and was not detected in bat faecal samples. In contrast, it was consistently identified in hair and snout-secretion samples from bats. Third, independent studies have similarly identified limonene in the secretions of other bat species (Faulkes et al., 2019; Zhang et al., 2022). Furthermore, emerging evidence indicates skin-associated microbes may contribute to bat volatile profiles, with some taxa possessing enzymes involved in terpene biosynthesis (Sun et al., 2026). Taken together, these observations point towards an endogenous or microbe-mediated source, although the exact biosynthetic pathway remains to be determined.”

      On contamination from human-associated sources

      You asked how we excluded the possibility that limonene entered our samples through handling, equipment, cleaning products, or other human-associated sources. We appreciate this concern and have addressed it in detail.

      We believe contamination is highly unlikely for several reasons. First, we followed strict protocols throughout. All instruments were cleaned with ethanol and oven-dried before and after each use. We used stainless-steel cages and cloth bags made of degreased bleached cotton washed with purified water. These materials are not sources of terpenes. Second, we ran multiple blanks. Limonene was not detected in any empty-chamber air controls or in blank cotton swabs. In contrast, it was consistently found in multiple bat snout-secretion samples. This clear difference between samples and blanks strongly argues against contamination. Third, we now report individual-level odor data (see new Figure 3; Supplementary Table 5.xlsx). These data show that limonene was consistently present across individual bats. It did not appear sporadically, as one would expect from accidental contamination.

      In the revised manuscript, we have added detailed descriptions of our contamination controls in the Methods section (Please see lines 405–407, 443–444, 471–477 of the clean version of the revised manuscript, or see the excerpt below). We have also included the blank-control results (Supplementary Table 5.xlsx), as you suggested.

      Lines 405–407: “Prior to sampling, all glassware was thoroughly rinsed with ethanol and dried in an oven at 120°C, and volatile odor collection was conducted in a dedicated odor-free room to minimize environmental contamination.”

      Lines 443–444: “The empty-chamber controls were used to account for potential background signals from the experimental system and to provide a baseline for comparison with bat odor extracts.”

      Lines 471–477: “Upon capture, bats were placed in clean stainless-steel cages and kept in groups consistent with their natural social associations during the brief interval prior to immediate odor sampling. Hair samples (10 mg per individual) were clipped from dorsal and ventral regions. Snout secretions were collected using sterile cotton swabs (CS15-005, Shenzhen SihuaBo Technology Co., Ltd., China), with two blank swabs as controls. These blank swab controls were included to account for potential volatile contamination from ambient air or the swab material itself (Supplementary Table 5).”

      On how crickets distinguish bat-associated limonene from environmental sources

      You raised a thoughtful question about ecological specificity. Given that limonene is abundant in mint, citrus peel, pine, and other non-threatening plants, how would crickets use it as a reliable indicator of bat presence?

      We agree with you completely. We do not claim that limonene alone serves as an unambiguous bat-specific signal. Instead, our interpretation is more nuanced. We argue that elemental perception represents one effective strategy within a broader olfactory toolkit. It does not exclude the importance of other cues.

      In our revised manuscript, we have softened our interpretation accordingly. We now state explicitly that limonene is sufficient to trigger avoidance under our experimental conditions, but we do not interpret it as a uniquely bat-specific keystone kairomone. We also discuss mechanisms that could help crickets reduce false alarms under natural conditions. These include concentration differences, temporal patterning (bats are active at night), spatial context (specific foraging habitats), co-occurrence with other bat-specific volatiles, and possibly enantiomeric composition. Please see lines 274–292 of the clean version of the revised manuscript, or see the excerpt below.

      Lines 274–292: “The second question is how crickets might distinguish bat-derived limonene from environmental sources of this compound, given its prevalence in mint, citrus peel, pine and other non-threatening plants (Boncan et al., 2020; Schuman, 2023). It seems implausible that crickets could simply rely on limonene per se to differentiate a bat from a leaf. Two non-exclusive mechanisms could help resolve this issue. First, limonene need not be the only olfactory cue mediating risk perception. Our findings establish the sufficiency of limonene as an avoidance trigger, but do not preclude a role for other odor components. The crickets’ antennal responses to other bat volatiles in our GC–EAD analyses suggest more complex peripheral perception. Additional compounds, either alone or in synergistic blends, may modulate the full behavioral response in nature. Therefore, elemental perception via limonene likely represents one effective strategy within a broader olfactory toolkit available to insects. Second, crickets may discriminate bat-derived limonene through context-specific cues (e.g., temporal and spatial patterning, co-occurrence with other bat-specific compounds) to minimize false alarms. Comparative studies on enantiomeric specificity and detection thresholds of cricket olfactory sensory neurons will be essential. Equally critical will be future efforts to quantify natural bat odor composition, limonene release rates, ambient exposure concentrations, and odor-plume dynamics, which together will inform ecologically valid stimulus design in controlled assays. Critically, our field data confirm that limonene exposure in nature robustly triggers an adaptive anti-predator response, irrespective of the precise discrimination mechanism.”

      We acknowledge that fully testing these ideas would require substantial additional work. We have therefore framed this as an important direction for future research, rather than as a resolved issue in the present study.

      We thank you again for these insightful comments. Your feedback has helped us present a more rigorous, balanced, and transparent account of our work.

      Reviewer #2 (Public review):

      Summary:

      Many insects possess extremely sensitive olfactory systems that can detect chemical signals from distances of several kilometers. For decades, the arms race between bats and insects has served as a prime example of acoustic co-evolution. The auditory adaptations of insects to echolocation have been well documented. Cricket has a multi-sensory predator recognition system with keen olfactory, tactile, and auditory senses. However, whether crickets can use the scent of bats to avoid them remains unknown at present. The authors hypothesized that cricket prey (Loxoblemmus equestris) might eavesdrop on predator bat (Scotophilus kuhlii) VOCs as an early warning. L. equestris is one of the prey species of S. kuhlii, and the authors demonstrated that the body odor of the insectivorous bat S. kuhlii triggers robust avoidance and electrophysiological responses in the cricket L. equestris, and that a single compound, (-)-limonene, is sufficient to elicit this avoidance in the laboratory and suppress calling in the field. Overall, this paper has a complete chain of evidence and should be a highly praised study.

      We sincerely thank you for your very positive and encouraging evaluation of our work. We are especially gratified that you recognized our study as having a "complete chain of evidence" and as a "highly praised study." This recognition means a great deal to us, given the multidisciplinary nature of our approach and the effort required to integrate behavioral, electrophysiological, chemical, and field data into a coherent narrative.

      We also appreciate your accurate summary of our work. You correctly highlighted the broader context that while acoustic co-evolution between bats and insects is well documented, whether crickets can use bat scent as an early warning cue has remained unknown. Your summary confirms that our main findings are clear: the body odor of S. kuhlii triggers robust avoidance and electrophysiological responses in L. equestris, and that limonene alone is sufficient to elicit avoidance in the laboratory and suppress calling in the field.

      We are particularly grateful that you acknowledged the multi-sensory nature of cricket predator recognition, with keen olfactory, tactile, and auditory senses. We agree that crickets are an excellent model for studying multimodal predator detection, and we hope our study encourages further exploration of olfaction in this classic predator–prey system.

      We have carefully considered all your specific comments and suggestions. These include questions about the novelty framing of olfactory eavesdropping, the rationale for our concentration choices, and the presentation of electrophysiological comparisons. We address each of these points in detail in our point-by-point responses below. Your thoughtful feedback has been extremely helpful in improving the clarity and rigor of the manuscript.

      Comments:

      (1) Olfactory eavesdropping can transcend the evolutionary divide between vertebrate predators and invertebrate prey, enabling invertebrates to trigger defensive avoidance behaviors in response to predator-derived volatile odors. This phenomenon is empirically well-documented and requires no excessive emphasis.

      Thank you for this comment. You are absolutely right that olfactory eavesdropping across the vertebrate–invertebrate divide is not a new concept in itself, and we acknowledge that this phenomenon has been well documented in previous studies.

      However, we would like to clarify our intended emphasis. In the Introduction and Discussion, we have already stated that empirical examples combining chemical identification, electrophysiological validation, behavioral assays, and field confirmation within a direct predator–prey context remain relatively limited. This is especially true for the bat–insect system, where research has historically focused on acoustic interactions rather than olfaction.

      We did not intend to overstate the novelty of olfactory eavesdropping per se. Instead, our emphasis was on providing a complete chain of evidence in a vertebrate–invertebrate predator–prey system that has traditionally been viewed through an acoustic lens. In that sense, we believe our study adds a complementary olfactory perspective to this classic system, rather than claiming to have discovered olfactory eavesdropping as a novel phenomenon. We hope this clarifies our position, as already stated in the original manuscript (lines 71–78, lines 88–90, lines 235–241 of the clean version of the revised manuscript, or see the excerpt below):

      Lines 71–78: “However, a fundamental gap exists in understanding whether such olfactory eavesdropping can operate across the vast phylogenetic divide separating vertebrate predators and invertebrate prey (Apfelbach et al., 2015; Dicke and Grostal, 2001; Schoeppner and Relyea, 2005). Although olfactory interactions across broad taxonomic boundaries are widespread in nature, such as mosquitoes using host odors to blood-feed, elephants and moths sharing pheromonal components, and aroids chemically mimicking carrion to attract pollinating flies (Kang et al., 2023; Zaremska et al., 2022; Zhao et al., 2022), these interactions are primarily shaped by selective pressures tied to foraging, reproduction, or mutualisms, rather than by predation-related selection.”

      Lines 88–90: “The bat–insect system presents an ideal model to address these questions. Despite the clear importance of olfaction to both taxa and its established role in predator–prey ecology, whether it plays any functional role in the iconic bat–insect arms race remains unexplored.”

      Lines 235–241: “Beyond the specific bat–insect model, our work addresses a central question in sensory ecology: how chemical eavesdropping operates within predator–prey systems between phylogenetically distant taxa with fundamentally divergent olfactory systems (Adams et al., 2020; Emerson and Johnson, 2024; Kaupp, 2010). While intraphyletic kairomone detection is well-established (e.g., rodents avoiding carnivore odors, aphids responding to ladybug chemicals), compelling experimental evidence for such olfaction-mediated recognition across broad phylogenetic divides has been limited (Apfelbach et al., 2005; Ferrari et al., 2007; Tanis et al., 2018).”

      (2) Without quantitative analysis and without knowing the relative content of this key substance limonene, I don't quite understand how to determine the concentration of limonene standard for EAD, as well as the concentration in field experiments. How is the concentration of limonene determined in field spraying, and is this actually the case in the wild environment?

      Thank you for this question. We fully agree that knowing the natural concentrations and relative content of limonene in bat odor would be valuable. However, our experimental aims were not to mimic natural emission levels precisely. Instead, they were designed to answer two distinct questions: First, whether cricket antennae are physiologically capable of detecting limonene at all; and second, whether limonene alone is sufficient to trigger behavioral responses under controlled and semi-natural conditions.

      For the EAG experiments, we selected a concentration gradient (0.001%, 0.01%, 0.1%, 1%, and 10% v/v) following standard practices in insect chemical ecology and referencing a previous study (Tang et al., 2024). The goal was to establish dose-dependent antennal sensitivity, not to match a specific natural concentration. Our data clearly show that cricket antennae respond across a range of concentrations, with stronger responses at higher doses.

      For the field experiment, we used a 10% v/v limonene spray over 25 m<sup>2</sup> plots. We acknowledge that this concentration does not quantitatively reflect natural bat emissions. Natural odor plumes are highly dynamic and depend on airflow, turbulence, temperature, humidity, vegetation structure, and distance from the source. Accurately reconstructing these natural dynamics would require detailed quantitative measurements of bat odor release rates and plume modeling, which were beyond the scope of the present study. Instead, our field experiment was designed for a functional purpose: to test whether limonene could alter cricket calling behavior under semi-natural conditions, using a concentration sufficient to produce a detectable odor stimulus in the field.

      We also note that the field-applied concentration is comparable to what has been used in other chemical ecology studies testing the behavioral effects of single volatile compounds under natural or semi-natural conditions. In that context, our positive result supports the ecological relevance of limonene as an avoidance cue, without requiring that the exact applied concentration matches natural bat emissions.

      We agree that quantitative characterization of natural bat odor composition, limonene release rates, and ambient exposure concentrations is an important direction for future research. According to your comments, we have added this point to the revised Discussion as a clear future direction (please see lines 286–290 of the clean version of the revised manuscript, or see the excerpt below). We have also clarified in the Methods that our assays were designed to test physiological detectability and functional sufficiency, rather than to establish concentration thresholds or mimic natural emissions exactly (please see lines 515–522, 556–558 of the clean version of the revised manuscript, or see the excerpt below).

      Lines 286–290: “Comparative studies on enantiomeric specificity and detection thresholds of cricket olfactory sensory neurons will be essential. Equally critical will be future efforts to quantify natural bat odor composition, limonene release rates, ambient exposure concentrations, and odor-plume dynamics, which together will inform ecologically valid stimulus design in controlled assays.”

      Lines 515–522: “For each antenna, a hexane control was first presented to establish baseline antennal activity. For the initial screening, limonene, undecane, pentadecane, and hexadecane were diluted to 10% (v/v) in hexane and delivered individually in a randomized order. To assess dose-dependent responses, limonene was further tested at five concentrations (0.001%, 0.01%, 0.1%, 1%, and 10%, v/v in n-hexane), following the concentration gradient used in a previous study (Tang et al., 2024). Following the initial hexane control, the five limonene concentrations were tested in a randomized order across trials. Each stimulus lasted 0.5 s, with an inter-stimulus interval of 1 min to allow full recovery of antennal responses.”

      Lines 556–558: “Our assays were designed to test physiological detectability and functional sufficiency, rather than to establish concentration thresholds or mimic natural emissions exactly.”

      (3) Figures 1C and D should compare the GC-EAD response of L. equestris to the odor of bat body and the odor of bat nasal secretions. It should not be compared with the air control group. Figure 1D has the same problem.

      Thank you for this suggestion. We understand your point that comparing GC-EAD responses between bat body odor and snout secretions would be a more direct way to identify the anatomical source of active compounds.

      However, we would like to explain why we did not include this comparison in the current study.

      First, the purpose of Figures 1C and 1D (i.e., Figure 2C and 2D in the revised manuscript) was to answer a more fundamental question: whether bat body odor, as a whole, contains volatile compounds that are detectable by cricket antennae. Comparing with an odor-free air control was therefore the appropriate first step. It established the basic phenomenon of olfactory detection before we moved on to source attribution.

      Second, we did conduct chemical profiling of snout secretions, hair, and faeces using HS-SPME-GC-MS (presented in Figure 3). These analyses showed that limonene was consistently present in hair and snout secretions, but absent from faeces and blanks. This allowed us to identify snout secretions as the most likely source of limonene, without requiring GC-EAD recordings from secretion samples themselves.

      Third, we did not perform GC-EAD on snout secretions for practical reasons. The secretion samples were collected in very small amounts. They were almost entirely consumed during the HS-SPME-GC-MS chemical analyses, leaving insufficient material for additional GC-EAD testing.

      We agree with you that directly comparing GC-EAD responses to snout secretions versus whole-body odor would be an excellent experiment. It would further strengthen the source attribution and provide more direct evidence. We have noted this as a valuable direction for future studies in the revised Discussion.

      We hope this clarifies our rationale. Thank you again for your thoughtful suggestion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) I would suggest moving Supplementary Figure 1 into the main figures. It contains important information about the predator-prey relationship and the ecological relevance of L. equestris, so it seems too central to be placed only in the supplement.

      We agree with your suggestion. The predator–prey relationship and the ecological relevance of L. equestris are indeed central to the biological context of this study. We have therefore moved the original Supplementary Figure 1 into the main text as Figure 1. We have renumbered the remaining figures accordingly, revised the figure legends, and updated the Results text to better highlight this ecological context (revised manuscript, Figure 1 legend, lines 856–864).

      (2) Lines 134 to 136: Only one representative EAD trace is shown. I suggest showing all five traces, either in the main figure or as a supplementary figure, to better illustrate the reproducibility of the antennal responses across individuals.

      We agree. To better illustrate reproducibility across individuals, we have added EAD traces from all five tested crickets as Supplementary Figure 1. The figure legend now describes the sample sizes for both the bat odor treatment and the odor-free control (revised manuscript, Supplementary Figure 1 legend, lines 900–906).

      (3) Lines 144 to 153: It would be helpful if the authors reported which VOC collections contained limonene and in what amounts. Showing individual-level data for the bat odor samples, rather than only pooled or summarized profiles, would strengthen the conclusion that limonene is consistently associated with S. kuhlii body odor.

      We agree. To better show the consistency of limonene detection across individuals, we have added Figure 3B, which displays an individual-level terpenoid profile. Each stacked bar represents one VOC collection, and the limonene segment indicates its presence and relative contribution. We have also revised the Results section to state explicitly that limonene was detected in hair and snout secretions but absent from feces (lines 144–149 in the revised manuscript).

      Lines 144–149: “To identify the biological sources of bat body odor, we analyzed VOCs from hair, faeces, and snout (pararhinal gland) secretions of nine bats using headspace solid–phase microextraction coupled with gas chromatography–mass spectrometry (HS–SPME–GC–MS). Snout secretions and hair shared similar hydrocarbon-rich VOC profiles, whereas faecal volatiles were distinct (Figure 3A). Individual-level terpenoid profiles further showed that limonene was detected in hair and snout secretion VOC collections but was absent from faeces (Figure 3B and Supplementary Table 2).”

      (4) Figure 2A: I recommend adding representative chromatograms or VOC traces for feces, hair, and snout secretions. This would make the source comparison more transparent and would help readers assess the underlying chemical profiles behind the heatmap.

      We agree that representative chromatograms would make the source comparison more transparent. However, due to the analytical workflow, individual chromatograms were not retained in the dataset we received. As an alternative, we added Figure 3B, which shows individual-level terpenoid profiles. This allows readers to assess which samples contained limonene and how its relative contribution varied among sample types and individuals. In addition, we have provided the NIST retention index, quantitative ion, qualitative ion, and molecular weight for each terpenoid compound in Supplementary Table 2.

      (5) The chemical identification of the key compounds would benefit from more detail. The authors state that compound identities were confirmed by matching retention times and mass spectra to authentic standards, including a mixed standard injection. It would be useful to provide the retention times, match and reverse-match scores, blank traces, and, if available, sample-plus-standard co-injection data showing peak augmentation without the appearance of new peaks. For (−)-limonene specifically, the enantiomeric assignment would require an enantioselective method, such as chiral gas chromatography, unless this was already performed and not described.

      We fully agree that comprehensive identification evidence is important for transparency and reproducibility.

      Regarding the identification data: we have already confirmed compound identities by matching retention times and mass spectra to authentic standards, including a mixed standard injection. In the revised version, we will deposit the raw chromatographic data and identification details (retention times, match scores, and blank traces) in Figshare as supporting information.

      Regarding co-injection: we acknowledge that sample-plus-standard co-injection would provide even stronger confirmation. However, our bat odor samples were difficult to obtain and were almost entirely consumed during the GC–EAD and GC–MS analyses. We therefore could not perform additional co-injection validation. We have noted this limitation in the revised Materials and Methods (revised manuscript, lines 463–464).

      Regarding the enantiomer assignment of limonene: you are correct that determining the specific enantiomer requires a chiral GC column, which was not available in this study. To avoid overinterpretation, we have replaced “(-)-limonene” with “limonene” throughout the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) This is a typical study in the field of chemical ecology. As long as the source of the active substances is determined and the biological activity has been detected, this research is complete, so Figure 2 seems to be unnecessary.

      We agree that identifying the source of the active compound and confirming its biological activity are central to this study. However, we believe Figure 2 serves an important purpose. Bat body odor could originate from multiple sources, i.e., hair, faeces, or snout secretions, and comparing VOC profiles across these sources helps us determine which source most likely contributes to the odor cues detected by crickets. This is especially important for limonene, because it is a plant-associated terpenoid and not a typical animal-derived volatile, as the reviewer #2 pointed out. Our analysis showed that hair and snout secretions shared similar VOC profiles, while faecal VOCs were distinct and lacked limonene. This supports snout secretions as the likely primary source. To make this purpose clearer, we have revised the Methods section to state that this analysis was conducted to investigate potential biological sources of bat body odor (revised manuscript, lines 494–495, 567–569).”

      Line 494–495: “To identify the biological source of the characteristic body odor, we compared the VOC profiles from hair, faeces, and snout secretions.”

      Line 567–569: “Principal component analysis (PCA) based on a binary (presence/absence) matrix was performed using the vegan package in R to compare profiles from hair, feces, snout secretions, and bat body odor.”

      (2) Figure 3B seems to be incorrect. The significance of n-hexane and 1% limonene is ***P < 0.001, while for 10% limonene, why is it only two stars, **P < 0.01? Please check it.

      We have rechecked the statistical analysis and confirmed that the original annotation was correct. The significance levels in original Figure 3B (Figure 4B in the revised manuscript) were based on Bonferroni-corrected paired t-tests comparing each limonene concentration with the n-hexane control. The adjusted P value was 0.0006 for 1% limonene (P < 0.001) and 0.00485 for 10% limonene (P < 0.01). The higher adjusted P value for 10% limonene reflects greater among-individual variation at this concentration, which may be due to differential sensitivity of individual antennae at higher doses. We have now added the exact adjusted P values in the revised manuscript, lines 163–167.

      Lines 163–167: “EAG responses to limonene were concentration-dependent (repeated-measures ANOVA, F(5, 25) = 24.95, P < 0.001, η<sup>2</sup>p = 0.83; Figure 4B), with both 1% and 10% limonene solutions eliciting significantly stronger responses than the hexane control (Bonferroni-corrected paired t-tests, 1%: t(5) = 10.77, P < 0.001, Hedges' g = 3.82; 10%: t(5) = 6.92, P = 0.005, Hedges' g = 2.46).”

      Finally, we would like to express our sincere gratitude to the editors and the reviewers for the thoughtful feedback. Your comments have significantly improved the quality and clarity of our manuscript. We hope the revised version satisfactorily addresses all concerns raised.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study systematically investigates repeat expansion in the plant Arabidopsis thaliana using a new k-mer-based method, expanding on smaller studies to more comprehensively identify cis- and trans-acting loci associated with repeat dynamics. The approach is methodologically sound and broadly applicable to large-scale short-read datasets for assessing copy number variation and genomic repeat content. While convincing in its scope and novelty, the findings would be further strengthened with exploratory analyses of datasets from other species with more or fewer repeats in their genomes.

      We agree with the assessment and appreciate the Editor’s handling of our manuscript and careful consideration of our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Overall, this study is an excellent and systematic investigation of the expansion of repeat sequences in Arabidopsis thaliana, and the genetic mechanisms underlying these expansions. Many of the key findings here confirm smaller studies of both repeat sequence variation and the individual genes associated with the expansion of various repeat classes. The authors present a highly effective and practical approach that requires datasets that are far more readily available than the multiple reference genomes used to annotate repeat variation in recent works. Therefore, they provide an approach that shows significant promise in non-model systems in which far less is known of repeat variation and its underlying drivers.

      Thank you for your comments and careful consideration of our work.

      Strengths:

      This is a very methodologically sound study that extends the relatively well-studied Arabidopsis thaliana repeat landscape with more systematic sampling, highlights the loci associated with repeat expansions (many of which were previously identified in a piecemeal manner), and provides some evolutionary inference on these.

      Weaknesses:

      Regarding cis-QTLs: I foresee at least two causes of these associations: non-repetitive cis-acting sequences that promote or permit the expansion of local repeats, and variation in repeat sequences themselves that directly tag the expanding sequence itself. It's arguable whether these are truly two distinct classes, but an attempt to discriminate between them may provide some insight as to the local factors that allow for repeat expansion, beyond the mere presence of a repeat sequence. One way to discriminate these could be to map the ~1300 12-mer frequency profiles on the reference genome, and filter any SNPs with elevated 12-mer frequency from the GWAS (or to categorize them independently).

      While it would be interesting to further investigate the mechanisms underlying cis-QTLs, we do not believe this dataset can distinguish between the two proposed models of cis-variation. As argued in the manuscript, the observed cis-association signals are linked to repeat-associated SNPs and therefore are most likely driven by variation in repeat copy number, supporting the latter hypothesis.

      I also have a question regarding the choice of k=12 in kmer profile analyses. Did the authors perform any GWAS with other values of K? If so, how did the results change? I would expect that as K is increased, the associations would become more specific to individual repeat families, possibly to the point where only cis-acting loci are detected. The authors show convincing evidence that k=12 is appropriate; however, I would be interested to see if/how GWAS results vary among e.g. k=10, 12, 15, 18.

      We attempted to regenerate the primary datasets using K = 14, but found that generating a complete 14-mer matrix for the number of samples analyzed was computationally infeasible, even after filtering low-frequency K-mers such as singletons. While this analysis could likely be done by redesigning the pipeline around a database-backed approach, we considered such development beyond the scope of this revision.

      As a compromise, we regenerated the primary dataset and repeated the GWAS analyses using 10-mers. To assess the impact of K-mer length, we compare GWAS results generated from both 10-mers and 12-mers in the new Figures S19-21 and describe these results in expanded discussion on the impact of K-mer length on our results.

      Reviewer #2 (Public review):

      Summary:

      The authors introduce a K-mer-based method for profiling repeat content within a species, applied here to 1,142 A. thaliana genomes sequenced with short reads. This approach allowed them to bypass the challenges of genome assembly, particularly for repetitive regions, while still quantifying copy number variation. Their analysis identified >50 trans-acting loci regulating repeat abundance, enriched for genes involved in DNA repair, replication, and methylation. They also speculate on the role of selection in shaping genome repeat content, arguing that purifying selection tends to suppress alleles that promote repeat expansion.

      The work presents a scalable way to extract meaningful insights from the large quantities of short-read datasets available. However, I have several concerns regarding the methodology, scope of claims, and interpretation of results.

      Thank you for your comments and careful consideration of our work.

      Strengths:

      The authors leverage a large dataset, >1100 samples, of A. thaliana. The scale of the study is impressive and clearly bolsters their findings. Additionally, this provides a framework for future, large-scale studies and offers a solid foundation for hypothesis generation. The k-mer-based method is generally practical for large-scale analysis and should be transferable to other datasets. Finally, the authors are commendably upfront about many of the project's limitations.

      Weaknesses:

      The decision to use k=12 is loosely justified. While the authors performed a sweep of k-mer lengths (from 5-20) and noted computational constraints, the choice is highly dataset-specific. Benchmarking across different k values with additional datasets (especially including other species) would strengthen confidence in the robustness of the method.

      Our decision to use 12-mers to profile genome content in A. thaliana was based on an empirical evaluation of the sensitivity and specificity of different K-mer lengths for this application. Because our goal was to characterize intraspecific variation in genome content, we did not extend this analysis beyond A. thaliana. Nevertheless, we agree that alternative K-mer lengths may capture additional and potentially useful information and that using the method in other species would be interesting.

      Although we found generating a 14-mer dataset to be computationally infeasible, we regenerated the primary dataset and repeated the GWAS analyses using 10-mers. To assess the impact of K-mer length, we compare GWAS results generated from both 10-mers and 12-mers in the new Figures S19-21 and describe these results in expanded discussion on the impact of K-mer length on our results.

      All analyses rely exclusively on the TAIR10 reference genome, which is incomplete and known to collapse certain repetitive regions. This dependence raises concerns that some repeats (especially recently expanded or highly variable ones) are systematically undercounted. With improved A. thaliana assemblies now available, testing the method against a more complete reference would alleviate these concerns.

      To our knowledge there is not a published gapless A. thaliana assembly, although several of the recent assemblies are pretty close. Although we continue to use the 1001 Genomes SNP dataset that relies on TAIR10 for the GWAS analyses, we now present Figure 2A and Figure S9 using the Col-PEK assembly (Hou, Wang, Cheng, Wang & Jiao 2022 Molecular Plant).

      The manuscript's conclusions are framed in very broad terms (e.g., "shaping genome evolution in plants"). However, the study is restricted to a single species, A. thaliana, which may not represent other plants. While the findings may suggest general principles, the claims in the abstract and conclusion should be moderated to reflect the study system more accurately.

      We concede that A. thaliana does not possess a typical plant genome and thus have moderated our claims throughout the paper.

      The identification of >50 trans-acting loci enriched for DNA repair and replication genes is compelling, but the conclusions remain correlational.

      We agree that the conclusions of our work are correlational, as they are based on GWAS associations and have not been validated through functional experiments. We would love to see tests of the hypotheses we presented, but we believe this is beyond the scope of the current study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The Snakemake workflow Kmer-it had a few minor bugs that prevented use with modern Snakemake versions, which I fixed as part of the review (submitted as a pull request). Please be sure to validate that the whole workflow works with the latest Snakemake version.

      We appreciate the Reviewer’s testing of our software and have updated it accordingly.

      (2) L162: Since you already map samples to a reference genome to filter organellar DNA, consider adding some form of filter or correction for duplication rate in paired-end data, which I have found to be one major source of batch effects as observed here between sequencing centers.

      We appreciate this suggestion. In earlier versions of the analysis, we removed duplicated reads, but doing so substantially weakened the relationship shown in Figure 1C. Because it is difficult to distinguish duplicate reads arising from genuine repeat copy number variation from those resulting from technical artifacts, we ultimately chose not to include duplicate removal in the final pipeline. However, we did remove the effect of the sequencing center using a linear model.

      (3) Have you used this approach with raw long-read data? Do you see any difference between Illumina and low-error-rate long read data (e.g., HiFi, R10 Nanopore with SUP calling)? This could be compared in a directly paired manner using some of the recent papers with HiFi/ONT data on 1001G accessions, e.g., Lian et al, 2024, Wlodzimierz et al, 2023, Teasdale et al, 2025. (I do not believe such a comparison is required to prove the utility of this method, but it could be of interest as such sequencing technologies become increasingly cost-competitive).

      We have not evaluated this approach using raw long-read sequencing data, although we agree that it would be an interesting direction for future work. As noted in the manuscript, we detected significant batch effects attributable to sequencing center, even among datasets generated with the same sequencing technology. Given these observations, we are cautious about comparing K-mer frequencies across sequencing platforms that have distinct error profiles, as technical differences could confound biological signals.

      Reviewer #2 (Recommendations for the authors):

      (1) I would strongly suggest modulating some of the claims made in the paper, especially with regard to "genome evolution in plants", given that the paper focuses on A. thaliana. Alternatively, the authors could test the method in a different dataset from a different species. This would alleviate some concerns regarding the choice of k-mer length and demonstrate robustness.

      We agree that A. thaliana is not representative of most plant genomes and have revised the manuscript to better reflect this limitation. Since the primary goal of this study was to characterize intraspecific variation in genome content within A. thaliana, we believe that extending the analysis to an additional species falls beyond the scope of the present work. We do not believe that our choice of K-mer length undermines the robustness of the approach. To evaluate this concern, we repeated the GWAS analyses using 10-mers and found broadly consistent results, which are presented in Supplementary Figure S19-21. These findings suggest that the major conclusions are not strongly dependent on the specific K-mer length selected.

      (2) A k-mer length of 12 is quite small. There are 4^12 possible 12-mers (~17 million), so you would expect that all 12-mers would occur in the A. thaliana genome by chance at least once. While larger k-mers may be more computationally expensive to compute, it would be worth repeating the analysis with a larger k-mer length.

      We attempted to regenerate the primary datasets using K = 14, but found that generating a complete 14-mer matrix for the number of samples analyzed was computationally infeasible, even after filtering low-frequency K-mers such as singletons. While this analysis could likely be done by redesigning the pipeline around a database-backed approach, we considered such development beyond the scope of this revision.

      (3) Line 689 on page 31 should include the figure number.

      Thanks! Fixed.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Marconcini et al. report results of an ambitious study on the genetic mechanisms that contribute to resistance of Drosophila flies to the toxin octanoic acid (OA). This study was motivated by two observations: first, Drosophila sechellia, a close relative of D. melanogaster, has evolved specialized feeding on fruits of Morinda citrifolia, which contain high concentrations of OA and second, that artificial selection on Drosophila simulans, a sister species of D. melanogaster, can generate higher resistance to OA. Previous studies had performed genetic mapping studies between D. simulans and D. sechellia that implicated certain genomic regions in resistance to OA and, in particular, implicated several Osiris gene paralogs as contributing to resistance, though the molecular mechanisms of resistance remain unclear. In this study, Marconcini et al. performed two major experiments. First, they performed evolution-and-resequence on Drosophila simulans populations exposed to OA for 50 generations and identified candidate regions with excessive shifts in allele frequencies as candidate regions containing OA resistance genes in D. simulans. Second, they performed a CRISPR knock-out screen in a D. melanogaster cell line to identify genes that contribute to OA resistance and susceptibility.

      Evolve-and-resequence yielded many candidate genomic regions with extreme allele frequency shifts, which may be regions containing OA resistance genes, or linked genes, or regions that happen to show a strong shift in all replicate populations by chance. As the authors note, detecting significant shifts in allele frequencies is a challenging problem, and the authors use two measures of allele frequency shifts (the Cochran-Mantel-Haenszel method and Bait-ER) and perform simulations under neutrality to estimate a reasonable significance threshold. I am not entirely convinced by this method of estimating significance levels, because the simulations involve assumptions that may not be met by the real populations. I would think that a permutation test would provide an assumption-free method of estimating significance levels. I have tried to think whether there is something about the design of these experiments that would preclude the use of permutation tests (which are used widely for genome-wide studies, such as QTL), but I can't think of one. Perhaps the authors are aware of a reason permutation tests would be invalid here, and if so, they should state this reason.

      Significance thresholds have been estimated using a variety of approaches in the literature, including false discovery rate (FDR) control, simulations under neutral drift with predefined cut-offs, arbitrary significance thresholds, haplotype-based analyses, permutation tests, amongst other methods. Our choice was guided by a review (doi:10.1186/s13059-019-1770-8), which compared the performance of several of these approaches and found that the relatively simple assumptions underlying the Cochran–Mantel–Haenszel (CMH) test often performed as well as, or better than, more complex alternatives. As with most significance thresholds used in genome-wide analyses, the CMH threshold applied here is ultimately based on a degree of arbitrariness, although we aimed to be as stringent as possible. As a complementary method, we used Bait-ER; here the threshold used followed the recommendations provided in the original publication describing this method (doi:10.1111/jeb.14134).

      One reason that permutation-based approaches may be less widely adopted than theoretical null models is the extensive linkage disequilibrium among SNPs, which results in large blocks of correlated variants that cannot be considered independent observations (and therefore not shuffled). In principle, permutations could be performed at the haplotype-block level; however, defining haplotype blocks itself requires selecting thresholds or criteria that are often user defined or based on significance cut-offs. Although several tools are available for this purpose, our experience has been that the resulting block definitions remain sensitive to these choices and therefore introduce a comparable degree of arbitrariness. In reality, there is not yet a standard method in the field that has emerged as the “best” one.

      There is overlap between regions detected by the two methods, but the methods disagree for many regions. The authors state that a "majority of prominent peaks were found by both methods," but I am unclear on what "prominent" means here. It would be more helpful to be more quantitative about the extent of overlap.

      To quantify the agreement between the two methods, we calculated the overlap in genomic coverage (base pairs) between CMH and Baiter candidate regions. For G25, the two methods shared 1.90 Mb of candidate regions. The overlap encompassed 65.3% of the genomic span identified by CMH and 64.1% of the span identified by Bait-ER. For G50, the methods shared 4.77 Mb of candidate regions. The overlap encompassed 63.6% of the CMH candidate span and 99.9% of the Bait-ER candidate span. We have now replaced the admittedly qualitative statement the reviewer highlighted ("majority of prominent peaks were found by both methods”) with this information.

      The authors hypothesized that the response would be at similar genomic loci in all populations (line 222). It seems at least possible that epistatic interactions would lead to different combinations of alleles evolving in each population. I wonder if it would be possible to test whether there is heterogeneity in the responses across the replicate populations.

      We agree that epistatic interactions could lead to different allelic combinations being favored in different replicate populations. However, both BaitER and the CMH test are designed to detect parallel evolutionary responses across replicates, which was the focus of our study. One approach for testing population-specific responses is the LRT-2 test (doi:10.1534/genetics.118.301824). We did not pursue this analysis because it would likely generate many additional candidate loci, making interpretation more challenging, while providing limited additional insight into the repeatable genomic responses that were the primary focus of this work.

      The evolve-and-resequence method yielded many possible regions contributing to OA resistance in D. simulans, but perhaps too many regions to test directly or even to build sensible hypotheses about the genes involved. Thus, the authors performed a second experiment to try to narrow down the list of possible candidate genes. They performed a CRISPR knockout screen in a D. melanogaster cell line for genes that contribute to resistance or susceptibility to OA. The authors identify several limitations of this experiment, but they nonetheless identified several genes where knockouts contribute to OA susceptibility or resistance. Intersecting top hits with regions that experienced selection identified two "resistance" genes: kraken and Alkbh7. The selection hit at kraken is quite compelling, whereas the evidence at Alkbh7 is less strong because only two SNPs were marginally significant. Further functional assays, including gene knockouts in D. melanogaster and D. sechellia, provide some support for the claim that both of these genes can contribute to resistance to OA in flies.

      Beyond the few issues raised above, I do not have significant questions about methodology or the results. I do think, however, that the authors should be more conservative about the implications and significance of their results. For example, on line 139, the authors claim that this intersection approach provides a "powerful paradigm to investigate ecotoxicology." I am not sure I agree that the identification of two genes that may contribute to OA resistance, after a seemingly heroic selection experiment and CRISPR screen, suggests that this method is all that powerful. It seems that most of the genes that contribute to the selection response remain unidentified.

      We agree with the reviewer and have modified the sentence accordingly. While we believe that integrating the approaches discussed in this paper can provide valuable insights into the genetic basis of ecotoxicological traits, these approaches are not a panacea for traits with highly complex genetic architectures, such as OA resistance. Nevertheless, the identification of two candidate genes with some evidence of contributing to the trait represents a meaningful advance toward understanding its underlying genetic basis.

      Finally, given that one motivation of this project was to identify genes that contribute to evolved resistance to OA, I am surprised that the authors did not generate CRISPR alleles of kraken and Alkbh7 in D. simulans and then use these together with the existing alleles in D. sechellia to perform reciprocal hemizygosity tests to determine if these two genes actually contribute to evolved resistance in D. sechellia. This test is simpler to perform and may be more sensitive than the allelic replacement that the authors propose (lines 446-449).

      While generating null alleles for the candidate genes in D. simulans is beyond the scope of this revision, we note that we did attempt reciprocal hemizygosity tests using the mutants available in D. melanogaster. However, the resulting hybrids were recovered in low numbers, and these animals were rather weak, making them unsuitable for the severe OA exposure conditions employed in our assays. We agree that, where feasible, future studies should incorporate reciprocal hemizygosity tests, as they represent a powerful approach for validating the contribution of candidate genes to the trait of interest. We have revised the final sentence of the Results section accordingly.

      Reviewer #2 (Public review):

      Summary:

      The authors studied the resistance against octanoic acid, a compound of the noni fruit, in D. simulans, using experimental evolution and resistance/susceptibility in D. melanogaster cells. They identified novel candidate genes and performed functional tests.

      Strengths:

      The idea of using experimental evolution of a non-resistant species to develop resistance is interesting, and the idea of narrowing down a large list of candidate loci by CRISPR-based gene knockout in cell culture is innovative. The reviewer also liked the (easy) follow-up experiments to validate the results.

      Weaknesses:

      The reviewer is not convinced of the conceptual idea behind their approach: the intersection of the two approaches implicitly assumes that null alleles (or at least compromised alleles) should be selected during experimental evolution. The reviewer considers this unlikely, and the authors made no attempt to test this implicit hypothesis in their data.

      We respectfully disagree with the reviewer’s interpretation of the conceptual idea behind our approach. Our strategy did not assume that experimental evolution selects for null alleles, but rather for any type of variant that could contribute to the trait being selected for (i.e., increases in OA resistance), pointing to candidate genes contributing to the trait. Like many evolve-and-resequence experiments, this approach identified hundreds of candidate genes. This is why we took an orthogonal, genome-wide CRISPR screening approach, where loss-of-function mutations could lead to increases or decreases in OA tolerance of cultured cells. However, we stress that the naturally selected alleles – which could be gain or loss of function – may have much subtler phenotypic effects than the null alleles used for functional validation.

      Along the same lines, it is not clear how to reconcile an upregulation of candidate genes in resistant flies with the knockout experiments.

      We respectfully disagree that these findings are difficult to reconcile. The observed upregulation of the candidate genes in the selected, resistant D. simulans is consistent with a role in promoting OA resistance, while the knockout experiments (whether in cultured cells or in whole animals) demonstrate that loss of gene function reduces resistance. These observations are complementary: increased expression is associated with enhanced resistance, whereas complete loss of function impairs it. The knockout experiments of kraken and Alkbh7 were intended as functional validation of gene involvement and do not imply that the alleles selected during experimental evolution are loss-of-function alleles (we rather hypothesize that the selected alleles are gain-of-function through some, as yet undetermined, mechanism).

      The experiments to validate the effect of candidate genes did not match the experimental evolution conditions.

      This is correct. As our results suggest that the phenotype is shaped by multiple genes, such that the effect of any individual gene is likely modest compared to their combined contribution. Consequently, detecting and validating the effect of a single gene requires more stringent OA conditions than those needed to observe the overall phenotypic response over the course of several generations. In addition, the shorter-term plate assay was more practical for higher temporal resolution of the analysis of mortality in the presence of OA.

      The statistical analysis suffers from some problems and an insufficient description of the analyses performed.

      Although D. simulans GWAS data are available, the authors did not make an attempt to estimate the effect of selected variants in the candidate genes in the GWAS data set.

      We agree that comparing the experimental evolution and GWAS results (from our previous work, doi:10.1093/g3journal/jkag032) is of considerable interest. We have now expanded the Discussion to explicitly discuss the relationship between the two datasets. Overall, the overlap between the approaches was limited, although two GWAS candidate genes, bez and CG13003, fall within genomic regions exhibiting significant CMH signals at generation 25 (but not generation 50) of the evolve-and-resequence experiment. (We note that these genes could not have been identified in our CRISPR screen, as they are not expressed in S2R+ cells). More broadly, differences between the GWAS and evolve-and-resequence results likely reflect the distinct evolutionary processes captured by each approach: GWAS maps standing phenotypic variation among isofemale lines, whereas experimental evolution tracks allele frequency changes under sustained selection. Understanding why some signals are shared whereas others are not – whether due to effect size, genetic background, epistasis, pleiotropic costs, or the contribution of initially rare variants – remain important open questions.

      The reviewer would have liked to see more connection between the experimental evolution and the GWAS data. As some D. simulans genotypes have similar resistance to D. sechellia, it would have been interesting to test whether this genotype contributed to the observed resistance.

      While D. simulans genotypes displayed a range of OA resistance levels, none approached D. sechellia levels of resistance (see Figure 2e from our previous work, doi:10.1093/g3journal/jkag032) (The reviewer might have conflated the data from our GWAS of D. melanogaster strain, shown in Figure 2b of that paper, where some lines of that species exhibit comparable resistance to D. sechellia under the conditions of that assay). Regardless, our experimental evolution data do not provide sufficient resolution to identify the specific favorable alleles underlying the response to selection. Instead, we detect genomic regions containing many linked variants whose frequencies change under selection. Consequently, a direct comparison between evolved genotypes and GWAS-associated genotypes is currently difficult. We note, however, that expression of the D. sechellia kraken allele in D. melanogaster did not produce a significant effect on resistance, suggesting that even the most promising candidate alleles might not have strong effects in isolation.

      At several places, the authors discuss the challenge of studying a polygenic trait, but at the same time, they claim to have detected and validated candidate genes. It would be helpful if the authors could discuss why they consider that their assays could really detect the contribution of single loci to the polygenic trait. In particular, when GWAS did not detect their candidate genes.

      Our results do not imply that kraken and Alkbh7 are major-effect loci or that they explain a substantial proportion of the phenotypic variation. Rather, our data indicate that these genes make measurable contributions to OA resistance, consistent with the expectation that complex traits are influenced by many loci of individually modest effect. The absence of these genes among the top GWAS candidates does not preclude their involvement, as GWAS and experimental evolution interrogate different aspects of the genetic architecture and differ in their power to detect loci of varying effect sizes and allele frequencies.

      It is not clear to the reviewer why the authors did not pay more attention to the highly significant peaks emerging from the experimental evolution study. Their functional validation would have been biologically more plausible.

      We agree that the significant peaks identified in the evolve-andre sequence experiment represent promising targets for future investigation. However, these peaks typically span large genomic regions containing tens to hundreds of genes, making it difficult to prioritize individual candidates based on the experimental evolution data alone. In this work, we focused our functional validation on genes independently supported by the CRISPR screen, which provided gene-level resolution. We fully acknowledge that additional causal genes are likely to reside within the selected regions and remain to be functionally characterized.

      Impact:

      Given the obvious challenges of functional testing of polygenic traits and the clear limitations of the interpretation of the results, the study will be helpful for future studies aiming to characterize polygenic traits. Unfortunately, the results are just another piece of controversial results regarding resistance against octanoic acid, a trait that is rather easy to evaluate.

      The reviewer appears to imply that a trait being straightforward to phenotype necessarily implies that its genetic basis should also be straightforward to resolve. Many classic complex traits, such as human height, are simple to measure yet have an extraordinarily complex, highly polygenic genetic architecture. We believe that OA resistance represents a similar challenge: while the phenotype is readily assayed, differences in assay conditions, genetic backgrounds, and the contribution of many loci of individually modest effect make its genetic basis difficult to dissect. We have strived to be cautious in our conclusions, in particular the evolutionary interpretations; nevertheless, to our knowledge, this is the first study to provide functional evidence supporting the contribution of specific genes to OA resistance in D. sechellia, combining both loss-of-function phenotypes and expression data. As emphasized by the title of our manuscript, we view the principal contribution of this work as demonstrating how complementary experimental approaches (both of which are fairly novel for study of toxin susceptibility/resistance genetics) can be integrated to prioritize and functionally evaluate candidate genes underlying complex adaptive traits.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers propose to include additional statistical and quantitative measures to strengthen the results. Furthermore, certain parts of the manuscript could be rephrased to make sure that the readers can understand more clearly the implications and significance of the results. To test whether the genes identified in the present manuscript (as contributing to octanoic acid resistance) are also involved in the evolution of the resistance, reciprocal hemizygosity tests using CRISPR alleles of kraken and Alkbh7 in D. simulans would be a plus, but such experiments are not required because the main focus of the paper is on the genetic basis of octanoic acid resistance, and not on evolution.

      We thank the Reviewing Editor for these constructive comments. We have revised the manuscript to clarify the interpretation and significance of our findings and have addressed the reviewers’ comments throughout. Regarding reciprocal hemizygosity tests, we agree that they would provide a valuable means of assessing the evolutionary contribution of candidate genes. However, generating the necessary reagents in D. simulans represents a substantial undertaking beyond the scope of the present study, particularly given the expected modest effects of individual loci underlying this highly polygenic trait. We have nevertheless revised the end of the Results section to mention reciprocal hemizygosity tests as an important direction for future work.

      Reviewer #2 (Recommendations for the authors):

      (1) Provide more details about the selection tests: which sequences were used? Please report p-values. It would also be important to discuss the possibility of false positives caused by a bottleneck in D. sechellia. A genome-wide analysis could help to see if the bottleneck increased the signal of positive selection.

      We are not entirely sure what additional analyses are being suggested. The sequences and methods used for the selection analyses are described in the Methods, and the statistical support for these analyses is reported in the Supplementary Material (MK test p-values and FUBAR posterior probabilities). We are also unclear as to how a genome-wide analysis would address the interpretation of the gene-specific selection analyses presented here. If the reviewer intended a different analysis, we would appreciate further clarification.

      (2) The significance level of the CMH test needs to be determined with the effective population size, not with the census size, as done by the authors. This is important, as it is not clear if the candidate genes remain significant after significance adjustment based on the effective population size.

      We thank the reviewer for this comment. In our analyses, effective population sizes were explicitly incorporated by Bait-ER to model the effects of genetic drift. The CMH significance thresholds, however, were obtained from neutral forward simulations following the recommended workflow for the method, which requires census population sizes rather than effective population sizes as input. To make these simulations as realistic as possible, we therefore used the observed census population sizes at each generation of the experimental evolution. Estimating generation-specific effective population sizes for use in an alternative simulation framework would require substantially more temporal data than are available here (e.g., sequencing many additional time points) and is beyond the scope of the present study.

      (3) Include the allele frequency trajectory across time for the candidate genes.

      We thank the reviewer for this suggestion. However, plotting allele frequency trajectories for the candidate genes is not straightforward because the evolve-and-resequence analysis identified broad linked genomic regions rather than individual causal variants. Each candidate region contains numerous SNPs spanning several kilobases, with each SNP exhibiting its own allele frequency trajectory across the 10 replicate populations. Consequently, there is no single representative trajectory for a given candidate gene, and plotting all SNPs within each region would be difficult to interpret. For kraken, however, we identified seven candidate regulatory SNPs and have plotted their individual allele frequency trajectories, which we include in Author response image 1 to illustrate the diverse patterns observed:

      Author response image 1.

      (4) Estimate the effect of candidate loci in the GWAS data (independent of significance).

      We thank the reviewer for this suggestion. However, we are not entirely sure what analysis is being proposed. In particular, it is unclear which variants the reviewer is referring to, as our candidate genes are associated with multiple linked variants rather than a single causal SNP. Consequently, we are unsure how the effect of a candidate locus should be estimated in the GWAS dataset. Moreover, there is no guarantee that the same variants are represented in both datasets. For example, variants filtered out during the GWAS because of low allele frequency may subsequently have increased in frequency during experimental evolution and therefore contributed to the evolve-and-resequence signals. We would appreciate further clarification of the analysis the reviewer has in mind.

      (5) Discuss the challenge of false positives (see: 10.1016/j.cub.2020.12.023).

      We thank the reviewer for this suggestion. We agree that false positives (as well as false negatives) are an important consideration when studying complex polygenic traits, both through the initial “screening” efforts (e.g., GWAS, experimental evolution) and follow-up functional validation (e.g., RNAi, mutant, overexpression analyses). We believe this issue is already addressed in the manuscript through our discussion of the limitations of the individual approaches, the polygenic nature of OA resistance, and our cautious interpretation of the functional validation results. Our conclusions are limited to identifying candidate genes that contribute to OA resistance, rather than claiming to have identified all of the loci or variants underlying the evolution of this trait. We are therefore not sure what additional discussion the reviewer has in mind and would appreciate further clarification if a specific point from the cited study is intended.

      (6) The figures with the Manhattan plots should be improved to indicate the overlapping genes in the Manhattan plots, rather than in the circle figures below. By using different colors this should be quite easy and clean.

      We thank the reviewer for this suggestion. However, we believe Author response image 2 more appropriately illustrate the overlap between the CMH and BaitER analyses. Although the Manhattan plots display individual SNPs, our candidate loci are defined by broader genomic regions comprising blocks of linked significant SNPs rather than by single variants. Simply highlighting SNPs within overlapping regions would therefore add visual complexity without providing additional biological insight beyond that already captured by the circos plots. As an illustration, we provide here an example of the generation 50 Manhattan plot with overlapping SNPs highlighted in red:

      Author response image 2.

      (7) A more focused discussion of the assumption that functional data from D. melanogaster can explain resistance in D. simulans or D. sechellia. At some places, epistatic interactions are mentioned, but the reviewer feels that the entire screen is based on the idea that similar effects are found across species, hence, this needs to be adequately reflected in the discussion.

      We thank the reviewer for raising this important point. We would like to emphasize that our study was not based on the assumption that functional effects identified in D. melanogaster or D. simulans necessarily explain the evolution of OA resistance in D. sechellia. Rather, our motivation stemmed from the longstanding difficulty of identifying individual genes underlying this highly polygenic trait using mapping approaches alone. We therefore sought to combine orthogonal experimental approaches to identify genes contributing to OA susceptibility (of D. melanogaster and D. simulans) and resistance (D. sechellia). We hoped, but did not assume, that genes supported by multiple independent lines of evidence would be informative for understanding the natural evolution of OA resistance in D. sechellia. Indeed, we acknowledge in the manuscript that the cell-based CRISPR screen, experimental evolution, and the natural evolution of D. sechellia occurred under very different selective contexts and timescales. As reflected in the title of our manuscript, our conclusions are intentionally framed around the identification of novel toxin resistance loci, while remaining cautious about their evolutionary interpretation.

      (8) Discuss that the controls in the RNAi test were quite variable. Could this reflect some problems with the assay?

      We do not believe that the variability among the control lines reflects a problem with the assay. Rather, the different controls (Gal4, UAS-RNAi etc.) represent distinct genetic backgrounds, each of which may exhibit a different baseline level of OA resistance. Indeed, in our recent GWAS of OA resistance (doi:10.1093/g3journal/jkag032), we observed substantial natural variation in OA resistance among D. melanogaster and D. simulans lines. Importantly, each RNAi line was compared with its corresponding genetic background control, so differences among control lines do not affect the interpretation of the individual RNAi experiments.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, the authors investigate the functional consequences of nuclear envelope rupture caused by the depletion of the nucleoporin NPP-3.

      They observe that loss of NPP-3 causes condensed chromosomes to localize to the nuclear periphery. This anchoring is independent of the pathway required to anchor heterochromatin and telomeres, but it depends on spindle assembly checkpoint proteins as well as centromere and kinetochore proteins. While the authors propose that relocalization of chromosomes to the nuclear periphery protects genome stability, they do not demonstrate this.

      Overall, some of the observations are interesting, but several points should be addressed. Furthermore, the manuscript could be much clearer if certain sections were shortened, simplified, or removed.

      Major points:

      (1) The title is misleading because the authors provide no experimental evidence that chromosome relocalisation protects genome stability. They are more cautious in the abstract, where they state that it 'may serve a protective role'. If they could provide stronger experimental evidence that chromosome relocalization protects genome stability, this would significantly strengthen the manuscript.

      We understand the more severe chromosome missegregation and DNA damage phenotypes in the npp-3 mdf-1 or npp-3 mdf-2 double RNAi may suggest that npp-3 RNAi is more sensitized to the loss of spindle assembly checkpoint (SAC) component MDF-1 or MDF-2 for chromosome protection, but may not clearly show that the localization of chromosomes to the nuclear periphery serves a chromosome protective function. Based on our current phenotypic analyses, we will tone down the title to “Prophase Chromosomes Relocalization to Nuclear Periphery in NPP-3/NUP205 Depletion Depends on the Spindle Assembly Checkpoint and Inner Kinetochore Proteins”.

      We have thought about artificially tethering the chromosomes to the nuclear periphery. However, without the same stimulus/defect and activation of the response pathway, the effects could be different. In future, to further clarify the functions of chromosome relocalization, we could analyze the chromosomal missegregation and DNA damage phenotypes in npp-3 lem-2 RNAi, which has a partial chromosomal nuclear periphery phenotype, to see if there is any quantitative relationship between chromosome nuclear periphery location and DNA protection.

      (2) Here, the authors use acute inactivation of npp-3. Do chromosomes also localize to the periphery upon partial npp-3 inactivation? What are the minimal levels of nuclear envelope rupture that cause chromosomes to localize to the periphery? Given that NPP-3 and NPCs have pleiotropic functions, it would be important to analyze conditions where only a few nuclear envelope ruptures are induced. In such conditions, they might be able to explore the link between chromosome localization and genome stability.

      We have not performed partial npp-3 RNAi yet. To explore whether the level of nuclear envelope rupture correlates with chromosome localization to the periphery or if a minimal level of nuclear rupture is required, we have performed the lacI::GFP reporter assay in different NPP RNAi to indicate the nuclear permeability defects (in Fig. S1A and B). npp-2, npp4 or npp-5 RNAi also causes increased permeability, at a comparable level as npp-3, npp-7 and npp-13 RNAi, yet only npp-3, npp-7 and npp-13 RNAi causes chromosome relocalization. This may explain the peripheral chromosome relocalization response may be more closely linked to the specific NPC subcomplex’s (inner ring and nuclear basket) or individual NPP’s function, rather than the general permeability defect or rupture size.

      However, we have also attempted to use other methods to induce targeted nuclear rupture with specific rupture size by laser ablation (Author response image 1, 355 nm UV pulsed laser marked by the white rectangle), and we could occasionally observe all chromosomes relocalizing to all over the nuclear periphery after 660 s upon laser ablation in different strains. However, the chromosome relocalization phenotype is not very consistent (~33%, n =12). Importantly, the laser also introduces DNA damage at the rupture sites, as marked by HUS-1 (Author response image 2), complicating the interpretation, so we did not include this attempt in the manuscript.

      Author response image 1.

      Time-lapse imaging of embryos expressing LEM-2::GFP and mCherry::H2B, a 15*30pixel rectangle was laser-microirradiated to induce nuclear envelope rupture (white rectangle). 0s is the time applying laser ablation. Scale bar 5 μm.

      Author response image 2.

      Time-lapse imaging of embryos expressing HUS-1::GFP and mCherry::H2B, a 15*30pixel rectangle was laser-microirradiated to induce DNA damage (white rectangle). 0 s is the time applying laser ablation. Scale bar 5 μm.

      Another attempt to achieve different nuclear rupture sizes is by observing nucleus in different stages of embryos. By npp-3 feeding RNAi approach, we observed an increase of the proportion of the nuclear circumference with nuclear rupture during embryonic development (Fig. S3D). Yet, the chromosome periphery phenotype is displayed in all embryo stages, suggesting that the localization of chromosomes to the nuclear periphery occurs across a range of rupture sizes.

      Given these observations, it appears that the extent of nuclear rupture or permeability increase may not be the sole determinant of chromosome relocalization. Instead, disruption of certain NUP subcomplexes or NPPs might stimulate this process.

      (3) The authors primarily examined P1 cells. Is the behaviour of the chromosome different between cells of different lineages?

      We observed that chromosomes localize to the nuclear periphery in all cells during embryonic development until the embryo dies, as demonstrated by snapshots and time-lapse imaging (Fig. S1C and Fig. S1E).

      We focused on the P1 blastomere for our analyses because its cell cycle timing and division orientation is well characterized, which facilitates precise examination of chromosome localization and cell cycle dynamics under different perturbation conditions. We will add a sentence at the beginning of that result section to clarify this choice: "The chromosome nuclear periphery localization in npp-3 RNAi is consistent in all cells at different embryonic stages (Fig. S1C and Fig. S1E). For analysis in distinct perturbation conditions, we specifically examined the P1 cells, whose cell cycle timing and division orientation is well characterized and suitable for such studies.”

      (4) The authors mentioned that defective chromosomal localisation does not occur upon npp-2 or npp-4 depletion. How do they explain this? Did they attempt to inactivate other NPPs in the Y complexes, and can they be certain that NPP-2 depletion is complete?

      (see response to point 2) We observed that npp-2 or npp-4 RNAi can cause nuclear rupture with certain permeability defect compared to wildtype (Fig. S1A and S1B), but not chromosomal nuclear periphery localization, suggesting that such chromosome localization to the nuclear periphery does not just depend on the nuclear rupture but could be a more specific phenotype related to individual NUP subcomplex’s (inner ring and nuclear basket) or NPP’s function. 

      For Y complex, we also now tested NPP-5 depletion, which shows permeability defect but did not show nuclear rupture marked by NPP-1:GFP or chromosomal periphery phenotype (Fig. S1A and S1B), which is consistent to the paper cited [1]. 

      To confirm the efficiency of NPP-2 depletion, we have observed the presence of smaller nuclei (as described in phenobank) by NPP-1::GFP marker (Fig. S1A) and a marked reduction in NPP2::GFP signal in npp-2 RNAi-treated embryos (unpublished data). These evidences support that NPP-2 depletion was effective. 

      (5) The section on AIR-1 (line 147) is confusing and could be removed. To my knowledge, air-1 depletion does not cause the appearance of multiple centrosomes, except maybe in a very few embryos. air-1 depletion causes major defects, so it is difficult to draw a parallel with npp-3 depletion.

      We agree with this suggestion and have decided to remove the section on AIR-1 in the manuscript text to avoid confusion. To clarify, air-1(RNAi) causes multiple centrosomes in only a small proportion of embryos (approximately 13%) [2] , primarily by affecting centrosome positioning [3]. and does not routinely lead to significant centrosome amplification.

      Our interest in AIR-1 was inspired by previous findings by Hachet et al., showing that AIR1/Aurora A localizes at sites without NPP-3 during mitotic entry—these sites are believed to correspond to centrosome locations [4]. This relationship prompted us to explore how AIR-1 depletion might influence NPP-3 localization at the nuclear envelope and its potential effects on chromosome positioning. Interestingly, we discovered that following air-1(RNAi) treatment, nuclei exhibit discontinuous NPP-3 localization on the nuclear envelope. Thereby, we could investigate the interplay between NPP-3 and chromosomal dynamics. Chromosomes localize to nuclear periphery specifically without NPP-3 (Author response image 3). This suggests a potential negative correlation of NPP-3 localization with the chromosomes. We did not imply any regulation by AIR-1.

      Author response image 3.

      Condensed chromosomes tend to localize at nuclear envelope sites without NPP-3 or NPP-7. (A) (C) Selected representative confocal images of all chromosomes localizing at the nuclear periphery in the control and air-1(RNAi) embryos expressing H2B::GFP and mCherry::NPP-3 (A) and GFP::NPP-7 and mCherry::H2B (C). The upper right image is the zoom-in view of the nucleus. Scale bar 10 μm. (B) (D) The intensity of GFP and mCherry is normalized to the average intensity along the nuclear envelope and plotted in the control and air- 1(RNAi) nucleus from (A) and (C).

      (6) The authors show that condensed chromosomes tend to localize to the nuclear envelope upon NPP-3 depletion. Do they condense at the nuclear envelope (NE), or do they condense first and then move to the periphery? This is unclear from the data presented in Figure 1D. Also, why do chromosomes condense earlier? This point could be discussed.

      Based on our time-lapse data in Figure 1D, most chromosomes appear to condense at or near the nuclear periphery, but there are still some chromosomes inside the nuclear space at approximately -600 seconds before NEBD, and then subsequently moving to the periphery during condensation (with full condensation at 100 s past NEBD). This suggests that initial condensation may occur both within the nucleus and at the periphery. Over time, chromosomes condense and cluster at the nuclear envelope. However, we did not separately measure the condensation level of individual chromosomes at the nuclear periphery versus in the middle of the nucleus, which could be challenging in live cells. So far, we cannot separate the chromosome nuclear periphery phenotype and the condensation.

      To confirm whether chromosome condensation precedes or follows relocation to the nuclear periphery, we could perform depletion of condensin II component, e.g. hcp-6, and see if lack of condensation affects chromosome relocalization.

      (7) The authors evaluated the consequences of NPP-3 depletion on transcription using RNA sequencing. The relevance of this experiment is questionable, however, as npp-3(RNAi) embryos have significant general defects and not only mislocalised chromosomes.

      We recognize that npp-3 RNAi embryos exhibit broad developmental defects, and therefore global transcriptional changes. Our RNA-seq analysis revealed that differentially expressed genes did not display positional bias within the genome. NPP-3 depletion downregulates many pathways, including pathways related to RNA polymerase II activity and cell cycle regulation, consistent with a global transcriptional downregulation. Thus, while the transcriptomic data are broad, it is consistent with the chromosome condensation phenotype.

      (8) In the co-depletion experiment npp-3(RNAi), X(RNAi) presented in Figure 3B, the levels of NPP-3 depletion seem highly variable. All the images shown are not similarly exposed, so it is difficult to evaluate these data.

      In our co-depletion experiments, the images for npp-3(RNAi) and X(RNAi) (Fig. 4) are displayed side-by-side under the same exposure conditions and scaled the same way with the same intensity thresholds to facilitate comparison. Despite that, some differences in the background intensity can be observed across samples. Thus, the normalised mCherry::NPP-3 signal intensity (subtracting the background) the single and double RNAi samples will be quantified and added to supplementary figure S4.

      To ensure consistency of double RNAi, we used ligation-based RNAi constructs designed to simultaneously target both NPP-3 and X, aiming to achieve comparable knockdown efficiencies. Nonetheless, variability in RNAi efficiency is a recognized limitation, and we interpret our data within this context. To further validate the knockdown, we will also assess the efficiency of the other gene X using the corresponding fluorescent reporter (see response to Reviewer 3, point 2).

      (9) Inactivation of mdf-1/2 suppresses the mislocalization of the chromosomes observed upon npp-3 inactivation. Does it also suppress the premature chromosome condensation phenotype?

      Inactivation of MDF-1 or MDF-2 suppresses both the chromosome nuclear periphery localization and extended duration of interphase to prophase and prometaphase in npp-3 inactivation. We have performed the analysis to assess the timing and extent of chromosome condensation in the single and double depletion conditions (Author response image 4). The chromosome condensation dynamics is similar in the mdf-1 RNAi and npp-3 mdf-1 double RNAi, as well as the control group, indicating that MDF-1 is also involved the premature chromosome condensation phenotype caused by NPP-3 depletion.

      Author response image 4.

      The dynamic changes of chromosome condensation parameter, in which 30% of pixels in the ROI analyzed is below the threshold scaled intensity (<77), in the different groups. The sample size is 5. Error bars show mean ± SEM.

      (10) Figure 5B: The delay induced by npp-3 depletion is not severe, based on the micrographs presented. The authors should show more representative images. The graph shows the elapsed time between NEBD and NER, and not NER to NEBD, as indicated.

      We have aligned the nuclear envelope reassembly (NER) time and highlighted the time point in the images (Fig. 5A). Our data show that in control embryos, this duration is approximately 780 seconds, while in npp-3(RNAi) embryos, it extends to about 930 seconds. This difference is statistically significant, as determined by one-way ANOVA (or appropriate nonparametric/mixed tests). The representative image is consistent with the quantification presented in Fig. 5B. We could add the corresponding videos to the supplementary information.

      (11) The authors observed that depleting mdf-1 slightly enhanced the lethality associated with npp-3 inactivation. Based on this observation, they conclude that loss of chromosome anchoring exacerbates genomic instability and severely impairs embryonic survival. However, the genetic interaction is not strong, as npp-3(RNAi) embryos already present more than 95% embryonic lethality and have defects other than just mislocalized chromosomes (e.g., defects in kinetochore and spindle assembly).

      It is correct that the average embryonic lethality observed in npp-3(RNAi) embryos reachs 95% (Fig. S6). Given the broad developmental defects and such high baseline lethality, the genetic interaction with mdf-1 is modest.

      Nevertheless, our findings highlight that in the npp-3 mdf-1 double RNAi condition, we observe significantly increased rates of lagging chromosomes (Fig. 5D), micronuclei formation (Fig. 7A), and elevated DNA damage (Fig. 7B). These effects support the idea that MDF-1-, MDF-2dependent chromosome anchoring to the nuclear periphery (and condensation) plays a positive role in NPP-3 depleted cells.

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to determine the molecular mechanisms by which nuclear pore component NPP- 3/NUP205 regulates chromosome localization in C. elegans embryos. Previous studies had shown that depletion of NPP-3 caused premature chromosome condensation and movement of chromosomes to the nuclear periphery. Peripheral location of chromosomes is also observed under respiratory stress conditions, suggesting that peripheral chromosome positioning could act as a protective response to stress conditions. How NPP-3 affects chromosome positioning was unknown. Here, the authors conduct a screen to identify factors that promote chromosome relocation to the periphery in npp-3-depleted embryos, identifying an important role for spindle assembly checkpoint components in this process.

      Strengths:

      Using cytological tools to visualise chromosomes and nuclear envelope markers, the authors show that, in addition to the peripheral location of chromosomes, NPP-3 depletion causes partial rupture of the nuclear envelope and premature chromosome condensation. By systematically codepleting NPP-3 and factors required for heterochromatin association with nuclear lamina (CEC4), telomere binding to nuclear envelope (SUN-1 and POT-1), proteins required for the nuclear rupture repair machinery (BAF-1 and LEM-2), kinetochore proteins and components of the spindle assembly checkpoint (SAC) (MDF-1 and MDF-2), the authors convincingly show that SAC components are required for peripheral relocation of chromosomes in absence of NPP-3. The study also provides convincing evidence that peripheral relocation of chromosomes in the absence of NPP-3 has functional implications as it causes transcriptional deregulation and premature relocation of SAC components from the nuclear envelope to chromosomes. Codepletion of NPP-3 and SAC components accelerates progression through miotic prophase and increases the incidence of defects in chromosome segregation during mitosis. These findings demonstrate that SAC proteins play an important role in regulating chromosome positioning during prophase (at least in the absence of NPP-3) and that they can regulate cell cycle progression at earlier stages than previously thought.

      Weaknesses:

      The authors also propose that NPP-3 depletion causes DNA damage; however, the evidence presented to support this claim is not as strong as that presented for the effects mentioned above. Also, the premature condensation of chromosomes appears as a clear consequence of NPP-3 depletion, but this intriguing phenotype remains unexplored.

      The DNA damage evidence is based on lagging chromosomes in 2-cell stages, the HUS-1 reporter, and micronuclei in embryos at the 20-30 cell stage. It is noted that npp-3 RNAi is pleiotropic and also causes DNA damage. There is additional DNA damage caused by loss of MDF-1 and MDF-2 in npp-3 RNAi, but we agree that it is difficult to say whether the effect is additive or not, complicating the interpretation. Thus, we will tone down in our title to describe the dependency of the chromosomal nuclear periphery phenotype (see response to Reviewer 1 point 1).

      As for the premature chromosome condensation phenotype in NPP-3 depletion, we hypothesize it may result from accumulation of factors such as BAF-1 at the nuclear periphery, which could facilitate chromatin condensation. To confirm whether chromosome condensation precedes or follows relocation to the nuclear periphery, we could perform depletion of condensin II component, e.g. hcp-6, and see if lack of condensation affects chromosome relocalization (also see response to Reviewer 1 point 6).

      Reviewer #3 (Public review):

      Summary:

      This manuscript reports that RNAi depletion of the inner-ring nucleoporin NPP-3/NUP205 in Caenorhabditis elegans embryos causes nuclear envelope rupture, premature chromatin condensation, and relocalization of condensed prophase chromosomes to the nuclear periphery. Through a candidate epistasis screen, the authors argue that this relocalization requires spindle assembly checkpoint (SAC) components (MDF-1, MDF-2, SAN-1), inner kinetochore proteins (HCP- 3, HCP-4, and partially KNL-1), and NE rupture-repair factors (BAF-1, LEM-2), but not the CEC-4 heterochromatin- or SUN-1/POT-1 telomere-anchoring pathways. They further show that NPP-3 loss extends prophase and the NEBD-to-anaphase interval in a SAC-dependent manner, redistributes MDF-1/MDF-2, and reduces import of KNL-1/BUB-1/HCP-1. Codepletion of NPP-3 with MDF-1 abolishes both the arrest and the peripheral localization while increasing lagging chromosomes, HUS-1 foci, micronuclei, and lethality, which the authors interpret as evidence that peripheral positioning is protective.

      Weaknesses:

      (1) The "protective" conclusion is largely correlative. The protective claim rests on the observation that co-depleting MDF-1 (or MDF-2) with NPP-3 removes the peripheral localization and simultaneously increases DNA damage, micronuclei, and lethality. However, depleting a SAC component removes at least three things at once: the peripheral localization, the prophase extension, and the NEBD-to-anaphase arrest. , they unrelated to chromosome positioning, the current design cannot separate damage caused by loss of a protective peripheral location from damage caused by checkpoint bypass. As presented, the increased damage is at least as consistent with simple SAC bypass. To support the protective model, the authors should provide a manipulation that disrupts peripheral positioning without abrogating the SACdependent arrest (for example, via the BAF-1/LEM-2 or kinetochore depletion) and show that damage still increases. The LEM-2 co-depletion, which partially suppresses positioning, is a natural place to test whether micronuclei and HUS-1 foci also rise.

      We agree that the current data are largely correlative. In the early C. elegans embryos, single depletion of SAC components MDF-1 or MDF-2 does not affect the mitosis duration or chromosome segregation. When spindles are defective, the functional SAC delays progression through mitosis [5]. Depleting SAC components such as MDF-1 in npp-3 RNAi indeed impacts multiple processes, including prophase and prometaphase duration, chromosome repositioning and condensation, making it challenging to disentangle effects specifically to related chromosome positioning.

      To address this, future experiments involving NPP-3 LEM-2 co-depletion, which has been shown to partially impair peripheral chromosome positioning, will be utilized to assess whether disruption of partial peripheral localization results in increased DNA damage, micronuclei formation, or HUS-1 foci accumulation, independent of SAC function (see response to Reviewer 1 point 1). 

      (2) Knockdown efficiency of the partner gene in double RNAi is not verified. The double depletions are performed by cloning both gene fragments into a single vector. This risks reducing the effective dose of each dsRNA, so an apparent suppression in an npp-3; gene X (RNAi) condition could reflect weaker NPP-3 knockdown rather than a true epistatic relationship. The authors partially address this by showing that NPP-3::mCherry is still reduced in npp-3;mdf-1 (Figure S4A/B), which is helpful, but they do not demonstrate efficient knockdown of the partner genes in any double condition. For the key epistasis conclusions (MDF-1, MDF-2, HCP-3, HCP4 suppressions), the knockdown of the second gene should be independently validated with a reporter strain for the second protein.

      We acknowledge that in our double RNAi experiments, the knockdown efficiency of the npp-3 is checked by imaging (see response to Reviewer 1 point 8), whereas that of the second gene was not validated in each condition. To address this, we confirmed the effectiveness of certain partner gene depletions, e.g. HCP-3, by examining the levels of the respective proteins using available GFP-marked strains or immunofluorescence (Author response image 5). Additionally, for genes like HCP-3, KNL-1, and BUB-1, we also assessed the functional consequences on chromosome segregation, where severe defects observed (Fig. S4C) can support effective depletion. However, for some strains, we do not have GFP makers and will need to check the RNA levels. 

      Author response image 5.

      Representative confocal image of GFP::HCP-3 and mCherry::H2B at the NEBD time point of P1 cell in the control, single and double RNAi. Scale bar, 5 μm.

      (3) Alternative explanations for the transcriptomic and H3K9me3 data are not excluded. NPP-3 depletion blocks nuclear import of molecules smaller than ~70 kDa and arrests development at early gastrulation. Both the RNA-seq changes (30% of genes downregulated) and the increased H3K9me3 signal could therefore be secondary consequences of nucleocytoplasmic transport failure and developmental arrest rather than evidence of position-dependent transcriptional repression. Notably, the authors' own finding that up- and down-regulated genes show no chromosomal positional bias (Figure S2C/D) argues against a model in which peripheral repositioning drives silencing of specific chromatin domains. This section should be reframed more cautiously, with the transport/arrest confound explicitly discussed, and RNA-seq replicate number and differential-expression thresholds reported.

      We agree that these chromatin modifications and transcriptional alterations could be related to the nuclear transport failure and developmental delay in npp-3 disruption. We will discuss this possibility in results and discussion. 

      (4) Evidence for SAC "activation in prophase" is indirect, and the effect is small. The claim of a novel prophase role for the SAC rests on MDF-1/MDF-2 intensity changes that are repeatedly described as "modest," "slight," or "mild," measured with small n and Student's t-tests, together with phenotypic suppression of prophase extension. There is no direct readout of SAC catalytic activity (for example, MCC assembly). The prophase-extension suppression by MDF-1 is the strongest evidence; the intensity data are weak support. I recommend tempering "the SAC is activated in prophase" to a hypothesis, and strengthening it with a more direct assay if feasible.

      We agree that the evidence for SAC activation during prophase is indirect. The prophase extension (100 s) in npp-3 RNAi is a functional assay to support SAC activation, and the suppression in npp-3 mdf-1 double RNAi suggests dependency. The changes in MDF1/MDF-2 intensities are modest. Biochemical analyses of MCC assembly in C. elegans mixed cell cycle stage embryos is challenging to demonstrate SAC activity in prophase. 

      (5) The BAF-1 arm of the model is inferred rather than demonstrated. The authors state that baf- 1(RNAi) and npp-3;baf-1 produced clustering too severe for epistasis, so BAF-1's requirement for peripheral localization is not actually established genetically; it rests on increased BAF-1 accumulation (correlative) plus the LEM-2 partial suppression. The proposed BAF- 1/CENP-C bridge is extrapolated from Drosophila (ref. 71). This is reasonable as a discussion hypothesis but should not be presented in the abstract or summary model as an established dependency.

      We did not include the BAF-1/CENP-C bridge hypothesis in the abstract or the model, and we will discuss this as a speculative mechanism rather than an established dependency.

      References:

      (1) Rodenas, E., Gonzalez-Aguilera, C., Ayuso, C. & Askjaer, P. Dissection of the NUP107 nuclear pore subcomplex reveals a novel interaction with spindle assembly checkpoint protein MAD1 in Caenorhabditis elegans. Mol Biol Cell 23, 930-944 (2012).

      (2) Schumacher, J.M., Ashcroft, N., Donovan, P.J. & Golden, A. A highly conserved centrosomal kinase, AIR-1, is required for accurate cell cycle progression and segregation of developmental factors in Caenorhabditis elegans embryos. Development 125, 4391-4402 (1998).

      (3) Kotak, S., Afshar, K., Busso, C. & Gonczy, P. Aurora A kinase regulates proper spindle positioning in C. elegans and in human cells. J Cell Sci 129, 3015-3025 (2016).

      (4) Hachet, V. et al. The nucleoporin Nup205/NPP-3 is lost near centrosomes at mitotic onset and can modulate the timing of this process in Caenorhabditis elegans embryos. Molecular Biology of the Cell 23, 3111-3121 (2012).

      (5) Encalada, S.E., Willis, J., Lyczak, R. & Bowerman, B. A spindle checkpoint functions during mitosis in the early Caenorhabditis elegans embryo. Molecular Biology of the Cell 16, 1056-1070 (2005).

    1. Author response:

      Reviewer #1:

      We thank Reviewer #1 for the thoughtful critique. We have revised the manuscript to clarify the physiological rationale for the experimental system, the potential relevance of mitochondrial STX–VDAC signaling to neurodegeneration, and the appropriate scope of our conclusions.

      (1) Lack of Rationale

      The study provides no justification for investigating sex-specific aspects of Alzheimer's disease by focusing on VDAC-mediated mitochondrial dysfunction in POMC neurons. These hypothalamic neurons are not recognized as early or primary sites of AD vulnerability, making the biological premise unclear.

      We thank the reviewer for raising this important point. We agree that the original manuscript did not sufficiently explain the rationale for using POMC neurons.

      Our rationale is primarily physiological rather than disease-specific. POMC neurons are a well-characterized, metabolically sensitive, and estrogen-responsive neuronal population in which membrane-initiated estrogen signaling and the actions of STX have been extensively characterized. They therefore provide a physiologically relevant neuronal system for identifying the molecular mechanisms through which STX regulates mitochondrial function.

      The potential relevance to neurodegeneration is supported by evidence that hypothalamic and POMC neuronal function can be disrupted in neurodegenerative disease models. Do and colleagues (2018) reported hypothalamic neurodegeneration, increased inflammatory and apoptotic markers, and reduced POMC neuronal populations in 3xTg-AD mice and further showed that exercise attenuated hypothalamic apoptosis and restored POMC neuronal populations (Do, Laing et al. 2018). In addition, Shen and colleagues (2016) demonstrated disruption of POMC/MC4R signaling in APP/PS1 mice and showed that restoration of this pathway improved synaptic function (Shen, Tian et al. 2016).

      These studies do not establish POMC neurons as a primary site of AD pathology, but they demonstrate that POMC-related neuronal systems can be vulnerable to neurodegenerative processes. This provides a broader biological context for investigating mitochondrial mechanisms in this neuronal population.

      There is also a strong physiological rationale for examining estrogen-sensitive mechanisms in POMC neurons. These neurons are established targets of 17β-estradiol and are highly responsive to metabolic and mitochondrial state. STX is a non-steroidal estrogenic compound that activates membrane-initiated estrogen signaling, and our previous studies demonstrated neuroprotective and mitochondrial effects of STX in a neurodegenerative disease model.

      Thus, POMC neurons were used because they provide a well-defined estrogen-responsive and metabolically sensitive neuronal population in which STX signaling can be mechanistically investigated—not because we consider them an initiating site of AD pathology.

      We have revised the Introduction and Discussion accordingly. The revised manuscript now emphasizes the physiological significance of STX–VDAC signaling for neuronal mitochondrial function and presents its potential relevance to neurodegeneration as an important area for future investigation.

      (2) Weak Link to AD Pathogenesis

      Although mitochondrial dysfunction is well established in AD, the authors do not convincingly demonstrate a mechanistic or pathological connection between VDAC2 and AD. VDACs are not established contributors to AD etiology, and the manuscript does not strengthen this association.

      We agree that the present experiments do not establish VDAC2 as an etiological or pathological driver of AD. This was not the objective of the present study.

      The primary goal was to identify the molecular target(s) through which STX influences mitochondrial function. Using BF-STX chemoproteomic capture and competition experiments together with single-cell gene-expression analysis, planar lipid membrane electrophysiology, and mitochondrial metabolic analyses, we identify VDAC proteins as mitochondrial targets of STX and demonstrate functional effects of STX on VDAC channel properties and mitochondrial bioenergetics.

      Our previous studies demonstrated neuroprotective and mitochondrial effects of STX in the 5xFAD model (Lee, Bostick et al. 2025), providing a neurodegenerative context that motivated the present mechanistic investigation. However, the current findings should not be interpreted as establishing VDAC2 as an AD pathogenic mechanism.

      We have revised the manuscript accordingly. The STX–VDAC interaction is now presented principally as a mitochondrial mechanism with potential relevance to neuronal physiology and neurodegeneration. Whether this pathway contributes to the neuroprotective actions of STX in disease models will require direct experimental testing.

      (3) Unclear Relevance to AD Contexts

      While the data support an interaction between STX and VDAC2 affecting mitochondrial parameters (ATP production, membrane potential, glycolysis, respiration) in POMC neurons, the study does not show whether this mechanism is relevant to mitochondrial dysfunction in AD. No validation is provided in AD-related models or in contexts related to sex-specific AD phenotypes.

      We agree that the present experiments do not directly establish the relevance of the STX–VDAC interaction to mitochondrial dysfunction in AD.

      The current study was designed to identify and characterize the molecular mechanism underlying the mitochondrial actions of STX, rather than to test this pathway in a specific neurodegenerative disease model. Our findings demonstrate that STX interacts with VDAC proteins, modifies VDAC channel properties, and alters mitochondrial bioenergetics.

      The study builds on our previous findings in the 5xFAD model, in which STX reduced amyloid-β-associated pathology and affected mitochondrial function (Lee, Bostick et al. 2025). The identification of VDAC proteins as STX targets therefore provides a mechanistic foundation for future studies examining whether this pathway contributes to neuronal protection under neurodegenerative conditions.

      We have revised the Discussion to acknowledge the absence of direct disease-model validation. Future studies using conditional or neuron-specific manipulation of VDAC isoforms will be important for determining the physiological and neuroprotective significance of STX–VDAC signaling in vivo.

      Accordingly, our conclusions now emphasize what is directly supported by the present experiments: STX interacts with VDAC proteins and modulates VDAC channel function and mitochondrial bioenergetics. The relevance of this mechanism to neurodegenerative disease remains to be established.

      (4) Interpretation of Competitive Binding Data

      The competitive binding results in Figure S4B are not adequately interpreted. The dose-dependent competition observed for VDAC3 suggests it may be a stronger candidate than VDAC2, yet this possibility is not addressed.

      We thank the reviewer for highlighting this important point. We agree that the original manuscript placed too much emphasis on VDAC2 based on the chemoproteomic data.

      Our chemoproteomic experiments identified VDAC1, VDAC2, and VDAC3 as STX-interacting proteins. Competition with unlabeled STX produced a particularly clear reduction in VDAC3 labeling, including complete loss of the VDAC3 signal at three molar equivalents of unlabeled STX. We therefore agree that VDAC3 represents an important candidate STX target.

      We have revised the Results and Discussion so that the competition experiment is no longer interpreted as demonstrating preferential or exclusive binding to VDAC2.

      However, the persistence of VDAC1 and VDAC2 labeling may also be influenced by properties of the BF-STX photoprobe as alkyl diazirine probes can preferentially photolabel membrane proteins (Kleiner, Heydenreuter et al. 2017). We now present this only as a possible technical consideration rather than an explanation established by our data.

      Our subsequent emphasis on VDAC2 was based on the integration of several observations. Single-cell qPCR demonstrated that Vdac2 is the predominant transcript in the native POMC neurons examined (revised Figure 3B), with an approximate expression hierarchy of Vdac2 > Vdac3 >> Vdac1. In addition on a technical note, recombinant VDAC3 is more difficult to reconstitute reliably into artificial membranes because of its lower stability in detergent.

      The revised manuscript therefore recognizes all three VDAC isoforms as candidate STX targets but does not claim that VDAC2 is the exclusive or preferential target. Direct comparisons of STX binding and functional modulation among the three isoforms will be required to establish isoform selectivity.

      Reviewer #2:

      We thank Reviewer #2 for the positive assessment of our chemoproteomic strategy and multidisciplinary characterization of the STX–VDAC interaction. We also appreciate the reviewer’s identification of important limitations and directions for future investigation.

      (1) Physiological Relevance of Cell-Line Experiments

      Most experiments were performed in immortalized cell lines rather than primary neurons or in vivo models, limiting their physiological relevance.

      We agree that the use of immortalized neuronal cell models represents an important limitation.

      However, these models provided the cellular material, reproducibility, and experimental control required for chemoproteomic target capture and Seahorse metabolic measurements. Importantly, we complemented these studies with single-cell analysis of native POMC neurons and electrophysiological characterization of recombinant VDAC channels.

      Nevertheless, these approaches do not substitute for direct demonstration of STX–VDAC signaling in primary POMC neurons or in vivo. We have therefore revised the Discussion to emphasize that establishing the physiological significance of this mitochondrial pathway will require validation in primary neuronal preparations and whole-animal models.

      (2) Structural Binding Site of STX on VDAC

      The exact structural binding site of STX on VDAC remains unresolved.

      We agree. BF-STX chemoproteomics identifies proteins interacting with STX in a cellular environment but does not resolve the amino acid residues or structural pocket responsible for binding.

      We have clarified this limitation in the Discussion. Determining the STX-binding site will require complementary approaches such as targeted mutagenesis, direct binding measurements with purified VDAC proteins, and structural studies in membrane-like environments, including lipid nanodiscs. Such studies should also determine whether STX recognizes a conserved feature among VDAC isoforms or exhibits isoform selectivity.

      (3) Lack of VDAC2 Loss-of-Function Experiments

      No loss-of-function experiments (e.g., VDAC2 knockdown) were performed to establish a direct causal link between VDAC2 and STX's bioenergetic and neuroprotective effects.

      We agree that VDAC loss-of-function experiments would provide an important additional test of causality.

      Our present evidence is convergent: chemoproteomics identifies VDAC proteins as STX-interacting targets; single-cell analyses demonstrate VDAC expression in POMC neurons; electrophysiological studies show that STX modifies VDAC channel properties; and metabolic analyses demonstrate STX-dependent changes in mitochondrial bioenergetics. Together, these findings support a STX–VDAC mechanism but do not establish that VDAC2 alone is necessary for the mitochondrial or neuroprotective actions of STX.

      We have revised the manuscript accordingly and explicitly identify the absence of loss-of-function experiments as a limitation.

      Because whole-body VDAC2 loss-of-function is associated with severe developmental consequences, future studies will require conditional or neuron-specific approaches. Parallel evaluation of VDAC1 and VDAC3 will also be important because all three isoforms were identified by chemoproteomics and potential functional redundancy may complicate single-isoform manipulations.

      (4) Non-linear Dose-Response at Higher STX Concentrations

      The non-linear dose-response at higher STX concentrations also requires further investigation.

      We agree that the non-linear concentration-response warrants further investigation and have revised the Discussion to avoid interpreting the STX response as a simple monotonic concentration-response relationship.

      One possibility is that STX engages multiple mitochondrial targets with different apparent affinities and opposing effects on respiration. At lower concentrations, STX may preferentially engage a higher-affinity target, potentially VDAC, whereas higher concentrations may recruit lower-affinity targets that constrain this response. Consistent with this possibility, our BF-STX dataset (Supplemental Tables) identified several mitochondrial proteins involved in oxidative phosphorylation and metabolite transport, including ATP5F1C, NNT, SLC25A4, SLC25A5 and SLC25A3. ATP5F1C is required for efficient mitochondrial ATP production (Fiorillo, Scatena et al. 2021), and estrogenic regulation of ATP synthase has been reported (Massart, Paolini et al. 2002, Moreno, Moreira et al. 2013). NNT, SLC25A4, SLC25A5 and SLC25A3 also regulate mitochondrial respiration, redox balance, and ATP production (Mayr, Merkel et al. 2007, Lopert and Patel 2014). However, BF-STX enrichment does not establish direct STX binding, relative affinity, or functional modulation of these proteins. We therefore present the multi-target explanation only as a hypothesis. Direct binding and concentration-dependent target-engagement studies will be required to determine whether the non-linear response reflects recruitment of a lower-affinity mitochondrial target, concentration-dependent effects on VDAC itself, or downstream mitochondrial feedback.

      We thank the editors and reviewers again for their constructive comments. The revised manuscript more clearly defines the physiological rationale for the experimental system, establishes the scope of the STX–VDAC mitochondrial mechanism supported by our data, and places its potential relevance to neurodegeneration in an appropriately forward-looking context.

      References cited in the response

      Do, K., B. T. Laing, T. Landry, W. Bunner, N. Mersaud, T. Matsubara, P. Li, Y. Yuan, Q. Lu and H. Huang (2018). "The effects of exercise on hypothalamic neurodegeneration of Alzheimer's disease mouse model." PLoS One 13(1): e0190205.

      Fiorillo, M., C. Scatena, A. G. Naccarato, F. Sotgia and M. P. Lisanti (2021). "Bedaquiline, an FDA-approved drug, inhibits mitochondrial ATP production and metastasis in vivo, by targeting the gamma subunit (ATP5F1C) of the ATP synthase." Cell Death Differ 28(9): 2797-2817.

      Kleiner, P., W. Heydenreuter, M. Stahl, V. S. Korotkov and S. A. Sieber (2017). "A Whole Proteome Inventory of Background Photocrosslinker Binding." Angew Chem Int Ed Engl 56(5): 1396-1401.

      Lee, H.-J., Z. Bostick, J. Doherty, T. L. Swanson, M. J. Kelly, J. F. Quinn, N. E. Gray and P. F. Copenhaver (2025). "Neuroprotection against beta-amyloid toxicity by the novel estrogen receptor modulator STX requires convergent signaling pathways." Frontiers in Molecular Neuroscience Volume 18 - 2025.

      Lopert, P. and M. Patel (2014). "Nicotinamide nucleotide transhydrogenase (Nnt) links the substrate requirement in brain mitochondria for hydrogen peroxide removal to the thioredoxin/peroxiredoxin (Trx/Prx) system." J Biol Chem 289(22): 15611-15620.

      Massart, F., S. Paolini, E. Piscitelli, M. L. Brandi and G. Solaini (2002). "Dose-dependent inhibition of mitochondrial ATP synthase by 17 beta-estradiol." Gynecol Endocrinol 16(5): 373-377.

      Mayr, J. A., O. Merkel, S. D. Kohlwein, B. R. Gebhardt, H. Böhles, U. Fötschl, J. Koch, M. Jaksch, H. Lochmüller, R. Horváth, P. Freisinger and W. Sperl (2007). "Mitochondrial phosphate-carrier deficiency: a novel disorder of oxidative phosphorylation." Am J Hum Genet 80(3): 478-484.

      Moreno, A. J., P. I. Moreira, J. B. Custódio and M. S. Santos (2013). "Mechanism of inhibition of mitochondrial ATP synthase by 17β-estradiol." J Bioenerg Biomembr 45(3): 261-270.

      Shen, Y., M. Tian, Y. Zheng, F. Gong, A. K. Y. Fu and N. Y. Ip (2016). "Stimulation of the Hippocampal POMC/MC4R Circuit Alleviates Synaptic Plasticity Impairment in an Alzheimer's Disease Model." Cell Rep 17(7): 1819-1831.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This revised manuscript represents a partial response to the concerns raised in the first round of review. The authors have made one genuine mechanistic addition in the form of the semi-permeabilized cell reconstitution assay, removed the most overreaching conclusions regarding the contribution of cytoplasmic TDP-43 aggregation to disease, and made several minor presentational improvements. However, the central weaknesses of the original submission remain substantially unaddressed. The exclusive reliance on non-physiological TDP-43 variants, the incompletely resolved mechanism linking XPO1 to TDP-43 phase behavior, and the limited organoid validation continue to limit confidence in the major claims. The authors have, in several instances, responded by removing contested data rather than by providing the additional evidence that was requested.

      (1) The justification for the 2KQ acetylation-mimetic system remains inadequate.

      The authors respond to the concern about the non-physiological nature of the 2KQ mutant by citing published evidence that TDP-43 acetylation occurs in ALS patient spinal cord and is upregulated under oxidative and proteotoxic stress conditions. While these references are real and support the relevance of acetylation as a pathological post-translational modification, they do not resolve the central concern: there is no quantification of how much endogenous TDP-43 is acetylated at the specific lysine residues mimicked by 2KQ in degenerating human neurons, and no evidence that the degree of RNA-binding disruption imposed by the double glutamine substitution is ever achieved by endogenous acetylation in vivo. The 2KQ mutant eliminates RNA binding essentially completely, whereas physiological acetylation events are graded, reversible, and likely partial. The response conflates the existence of TDP-43 acetylation as a phenomenon with validation that 2KQ is a physiologically accurate model of that phenomenon. None of the new experiments address the request to test whether wild-type TDP-43 expressed at near-physiological levels, or a bona fide heterozygous ALS-linked TARDBP mutant in iPSC-derived neurons, responds to XPO1 modulation in a qualitatively similar fashion. Until this is shown, the mechanistic conclusions of this paper remain constrained to a highly artificial overexpression system and cannot be extrapolated to physiological or pathological TDP-43 biology with confidence.

      We agree with the reviewer that the TDP-43 2KQ mutant is a non-physiological variant. To address this concern, we have removed all statements that extrapolate our findings to disease pathogenesis. As reflected in the revised title, we now present this study as an investigation of factors that modulate TDP-43 phase transition and aggregation using a sensitized model system, rather than as a direct disease model. The Abstract and Discussion have also been revised accordingly.

      The choice of the 2KQ mutant was dictated by the requirements of the screening strategy. To identify modulators of TDP-43 phase transition, it was necessary to use a TDP-43 variant that reliably undergoes phase separation in cells within an experimentally practical timeframe. In this context, we believe the 2KQ mutant provides a suitable and justified experimental tool. We have revised the manuscript to clearly distinguish observations made with this engineered construct from conclusions regarding physiological or disease-associated TDP-43.

      (2) The homozygous K181E organoid model is still not adequately justified, and no heterozygous comparison has been provided.

      The authors acknowledge that the homozygous background is "more sensitive for detecting phospho-TDP-43" and argue that homozygous conditions are commonly used in experimental TDP-43 research. However, the critical issue is not whether homozygous models are used in general, but whether the homozygous background specifically alters the relative contribution of cytoplasmic aggregation versus nuclear RNA-processing dysfunction in this study. In a homozygous K181E model, both alleles produce an RNA-binding-defective TDP-43, meaning that every molecule of endogenous TDP-43 in the cell is dysfunctional. This is categorically different from the patient situation in which one wild-type allele is present, and it may substantially exaggerate nuclear loss-of-function relative to cytoplasmic gain-of-function phenotypes. The authors have not performed the requested comparison with heterozygous K181E/+ organoids, nor have they acknowledged that the organoid genotype itself could bias the interpretation of what KPT-276 treatment rescues. Given that the organoid section is now the sole in-disease-model validation of the XPO1 mechanism, this limitation is more consequential than it was in the original submission.

      We agree with the reviewer and have removed all speculative statements regarding the relative contributions of cytoplasmic aggregation and RNA splicing defects to disease pathogenesis. The organoid section has also been revised to focus solely on the experimental findings. Specifically, we only present evidence that inhibition of nuclear export in a sensitized organoid model promotes the accumulation of cytoplasmic phosphorylated TDP-43 without making broader claims regarding its role in disease pathogenesis.

      (3) The new semi-permeabilized cell data is a genuine contribution, but the mechanistic interpretation remains insufficiently constrained.

      The development of the streptolysin O semi-permeabilized cell reconstitution system is the most substantive new addition to this revision. The finding that LMB-stabilized anisosomes resist cytosol washout but dissolve upon RNase T1 treatment is interesting and provides a plausible indirect mechanism: XPO1 inhibition retains nuclear RNA, and this elevated nuclear RNA availability contributes to maintaining the liquid LLPS state of the TDP-43 2KQ condensate. This is a meaningful mechanistic advance and deserves credit. However, several important limitations of this new data are not adequately discussed. First, RNase T1 degrades single-stranded RNA globally during permeabilization, so the experiment does not identify which specific RNA species stabilize the anisosome, nor whether these are pre-mRNA splicing intermediates, mature mRNA, non-coding RNA, or another class. Second, the same nuclear export blockade that retains RNA will also retain the nuclear concentrations of many RNA-binding proteins, splicing factors, and other XPO1-dependent cargos. The RNase T1 experiment does not exclude the possibility that the relevant effect is mediated by an RNA-binding protein whose nuclear concentration increases upon LMB treatment and which, upon RNase digestion, can no longer engage TDP-43 or the anisosome shell. Third, the permeabilized cell system is by definition not intact and has lost cytosolic factors; whether the RNA-dependent stabilization of anisosomes operates in the same way in intact cells during physiological or pathological nuclear export perturbation is an assumption, not a demonstrated fact. The authors should more carefully frame these data as hypothesis-generating and explicitly note these alternative interpretations in the Discussion.

      We have now added some sentences on page 11 to acknowledge the limitation of our experiments. It reads as “However, our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires further validation.”

      (4) The conceptual asymmetry between XPO1 inhibition and XPO1 overexpression phenotypes is not resolved by the new mechanism.

      The paper continues to present two XPO1 perturbation phenotypes that are difficult to reconcile within a single mechanistic model. XPO1 inhibition enlarges anisosomes, maintains their liquid character by FRAP, and retains them in the nucleus. XPO1 overexpression also enlarges TDP-43 puncta, but these are FRAP-impaired, gel-like, and appear in the cytoplasm. The RNA-retention model proposed by the new semi-permeabilized data explains why XPO1 inhibition stabilizes the liquid state, but it does not explain why XPO1 overexpression drives the opposite outcome: gel-like hardening and cytoplasmic redistribution. If increased nuclear RNA availability is the key variable downstream of XPO1 inhibition, then XPO1 overexpression would be expected to decrease nuclear RNA and thereby destabilize anisosomes toward dissolution or hardening. The paper does not test whether nuclear RNA levels are indeed altered by XPO1 overexpression, nor whether the cytoplasmic gel-like puncta seen in XPO1-overexpressing cells are RNA-poor relative to control anisosomes. The revised Discussion does not engage with this asymmetry in a satisfying way, and the figure model remains qualitative. A quantitative or at least semi-quantitative model that accounts for both arms of the XPO1 perturbation is needed.

      We thank the reviewer for this point. We have now explicitly mentioned in the discussion that the effect of XPO-1 on anisosome dynamics is likely mediated by an indirect mechanism. We also acknowledge that we do not fully understand why over-expression of XPO-1 causes TDP-43 to accumulate in gel-like structures in the cytoplasm. Although we did not check whether overexpression of XPO1 increases cargo export, we cited studies showing that over-expressed XPO1 disrupts the normal distribution of cargos between nucleus and cytoplasm on page 7. To avoid confusion, we also revised the result part on page 7, emphasizing on the difference rather than the similar increase in puncta size by opposing manipulations.

      (5) The removal of RNA-seq data weakens rather than strengthens the organoid section.

      The authors have removed the bulk RNA-seq analysis from the revised manuscript in response to concerns that the modest transcriptional rescue was being over-interpreted. While the decision to remove over-interpretation is appropriate, the result is that the organoid section now rests entirely on pTDP-43 immunostaining as its sole readout. The revised paper thus uses reduction in immunofluorescent pTDP-43 puncta in homozygous K181E organoids as the only evidence that nuclear export inhibition mitigates TDP-43 proteinopathy in a disease-relevant context. This is a weaker evidentiary base than before the revision, not an improvement. The originally requested more sensitive orthogonal readouts, including biochemical fractionation for SDS-insoluble TDP-43, filter-trap assays, or RNA aptamer-based detection of TDP-43 aggregates, remain absent. Without at least one additional independent measure confirming that cytoplasmic TDP-43 aggregation is genuinely reduced rather than simply rendered antigenically undetectable, the organoid conclusion is not adequately supported. At minimum, the authors should provide total and cytoplasmic TDP-43 fractionation data from organoid lysates to corroborate the immunostaining result.

      We removed the RNA-seq analysis because both the reviewers and editors agreed that the modest transcriptional rescue should not be overinterpreted. Upon reconsideration, we believe that reinstating these data would not address the reviewer's principal concern, namely whether nuclear export inhibition reduces TDP-43 aggregation in organoids. We have therefore chosen to limit our conclusions to the direct observation supported by the current data, namely a reduction in cytoplasmic phosphorylated TDP-43 immunoreactivity. We also add a sentence to acknowledge that “Whether the reduction in pTDP-43 immunoreactivity reflects a decrease in insoluble TDP-43 aggregates remains to be determined” on page 10.

      (6) No functional neuronal readout has been provided for the organoid model.

      The organoid section now makes the claim that "nuclear export is required for the formation of p-TDP-43-containing aggregates in a disease-relevant organoid model," but no measure of neuronal health, integrity, or function is reported in association with this. Even a simple assessment of neuron survival by TUJ1 or MAP2 quantification, neurite complexity, or cleaved caspase-3 staining before and after KPT-276 treatment would substantially strengthen the biological significance of the pTDP-43 reduction. The current data establish a pharmacological effect on a pathological marker but do not demonstrate that this has any consequence for neuronal biology in the organoid, which is what the disease-relevance framing implies.

      We thank the reviewer for this helpful suggestion. We agree that assessments of neuronal survival or function would be important if the manuscript were claiming that nuclear export inhibition improves neuronal health or rescues disease phenotypes in the organoid model. However, in response to the reviewers' comments regarding the physiological relevance of the homozygous K181E organoids, we have substantially revised both the framing and interpretation of this section.

      Specifically, we have revised the statement to read, "These results imply that maintaining TDP-43 in the nuclear demixed liquid state might diminish p-TDP-43 accumulation but whether the reduction of pTDP-43 immunoreactivity reflects a decrease in insoluble TDP-43 aggregates remains to be determined," thereby limiting our conclusion to the direct experimental observation. We no longer make claims regarding disease modification or functional rescue in the organoid model. Given this revised scope, we believe that additional measurements of neuronal survival or function, while certainly of interest, are not essential to support the conclusions presented in this study. We have also revised the conclusion to explicitly acknowledge that the functional consequences of reducing p-TDP-43-positive puncta (whether this can be translated to reduced aggregation) remain to be determined by future studies (page 10).

      (7) The abstract and title continue to overstate the mechanistic conclusions.

      Despite the stated intent to reframe the study as a screening study and to temper the conclusions, the revised abstract retains the language: "These findings establish nuclear export as a key regulator of TDP-43 phase transitions and define a mechanistic framework that links altered nuclear transport and phase dynamics to TDP-43 aggregation potential." Similarly, the Discussion still states: "a particularly compelling aspect of our study is the discovery that the nuclear export receptor XPO1 governs TDP-43 liquid-to-solid transitions and subcellular localization." The word "governs" and the phrase "establish nuclear export as a key regulator" are not warranted by data that derive entirely from an overexpressed acetylation-mimetic mutant in a colon cancer cell line and a homozygous K181E organoid model. A more accurate framing would describe these findings as identifying nuclear export as one of several cellular processes that modulate TDP-43 phase behavior in a sensitized model system, with an indirect RNA-mediated mechanism that remains to be defined at the molecular level. The title change from "governs" to "modulates" is appreciated but does not extend into the abstract and Discussion, where the strong causal language persists.

      We have revised the title of the paper, reframing it as a screen that reveals modulators of TDP-43 phase separation. The last sentence of the abstract is also revised accordingly. It now reads as “These findings identify multiple modulators of TDP-43 phase transitions in a sensitized model system and establish a framework for further dissecting the link between nuclear transport and TDP-43 phase dynamics.” We also tone down our conclusions and discussions.

      (8) Individual siRNA knockdown validation for XPO1 has not been provided.

      The authors argue that validation with 6 independent siRNAs across two rounds of screening, combined with convergent pharmacological data, is sufficient to establish XPO1 as a genuine hit. While the convergence of chemical and genetic evidence is reassuring, the specific request was for protein-level confirmation of XPO1 knockdown efficiency in the DLD1 TDP-43 2KQ cells used for mechanistic follow-up, together with demonstration that the anisosome phenotype is specifically caused by loss of XPO1 and not by off-target effects. This is a straightforward experiment, and its absence is particularly notable given that the entire mechanistic XPO1 narrative hinges on this specificity. At minimum, an immunoblot confirming XPO1 protein depletion in cells treated with the siRNA pool identified in the screen, in the same cell background and induction conditions as the follow-up experiments, should be provided.

      While we agree with the reviewer that studies relying on siRNA should provide sufficient information regarding knockdown efficiency and specificity, we respectfully disagree that this should be a major concern in the present study. As explained in the manuscript, we deliberately chose not to pursue mechanistic studies using chronic XPO1 knockdown because prolonged depletion of this essential nuclear export factor is likely to produce secondary effects that could complicate data interpretation. Instead, we employed multiple chemically distinct XPO1 inhibitors to achieve acute inhibition, thereby minimizing indirect consequences while providing a more appropriate approach for mechanistic analysis.

      We agree that assessing knockdown efficiency is technically straightforward. However, because our mechanistic conclusions are based primarily on acute pharmacological inhibition rather than siRNA-mediated depletion, we prioritized experiments that directly addressed the central mechanistic questions raised by the reviewers, particularly the semi-permeabilized cell assay. Moreover, the XPO1 inhibitors used in this study are well-characterized, highly specific compounds that have been extensively validated and widely used in the literature. We therefore believe that our experimental strategy provides a reliable basis for the conclusions presented.

      (9) The identity of XPO1-dependent cargos that regulate anisosome dynamics remains entirely unknown.

      The authors acknowledge that XPO1 does not directly bind TDP-43 and that the mechanism is likely indirect. The new RNA data provides one plausible indirect pathway. However, the possibility that one or more specific RNA-binding proteins or splicing factors, whose nuclear levels rise upon XPO1 inhibition, are the proximate drivers of anisosome stabilization has not been addressed. This matters because if the relevant mechanism operates through a specific cargo rather than bulk RNA retention, the model for how nuclear export connects to TDP-43 aggregation in disease would be fundamentally different. The authors decline to pursue adaptor identification on grounds of scope, which is a defensible position for future work. However, the framing should explicitly state that the current data cannot distinguish between bulk RNA retention and cargo-specific effects, and that the conclusion that nuclear export modulates TDP-43 phase behavior via RNA accumulation is a working hypothesis supported by but not proven by the RNase T1 experiment.

      We thank the reviewer for this helpful suggestion. We have now added a sentence on page 11, which state that “our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires validation.”

      Minor remaining issues.

      The number of independent iPSC clones and organoid batches used for the KPT-276 treatment experiment is now stated as two batches per condition, which is minimal for a 3D organoid study and does not fully address the concern about clone-level variability. Ideally, organoids from at least two independently derived isogenic clones per genotype would be used. The mCherry overexpression control added in Supplemental Figure 4 is a useful addition and is acknowledged. The immunoblotting confirmation that drug treatments do not alter total TDP-43 levels addresses a prior concern adequately. The addition of the sentence noting that anisosomes have not been validated in human patient samples is appreciated and appropriate. Statistical detail has been improved in figure legends. These minor improvements are noted positively but do not compensate for the major unresolved concerns above.

      We thank the reviewer for his/her appreciation of our previous revision. We hope that the new changes now satisfactorily address the remaining concerns.

      Reviewer #2 (Public review):

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. They further strengthen their findings by introducing an additional model system, a semi-permeabilized in vitro assay, which provides mechanistic evidence that XPO1 activity prevents anisosome dissolution by retaining nuclear RNAs. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. Some minor concerns remain, mostly about experimental design of the newly added data.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      (6) The authors extend their study by incorporating a semi-permeabilized in vitro system, which provides compelling evidence that inhibition of nuclear export promotes the retention of nuclear anisosomes, an effect driven by the accumulation of nuclear RNAs.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      We thank the reviewer for acknowledging the strength and the potential significance of our study.

      Weaknesses:

      (1) The manuscript has significantly improved with the revisions. Some experimental procedures and method details, as well has statements remain incompletely described:

      (a) What is the smear in Figure S1 after VLX treatment?

      We thank the reviewers for the positive assessment. We do not know why VLX treatment causes a fraction of TDP-43 to migrate slowly. We suspect that it may form detergent-insoluble aggregates. However, we cannot be sure whether this occurred during drug treatment or sample preparation. We now add a sentence in the figure legend to clarify this point.

      (b) The authors state that "The reduction in TDP-43 signal was not due to protein elimination.", however no data is provided to support that statement.

      We reasoned that the reduction in TDP-43 was probably not caused by protein elimination because the puncta could be reformed when permeabilized cells were incubated with exogenously added cytosol and ATP/GTP. We have revised the text to avoid this confusion. The revision on page 8 reads as “The reduction in TDP-43 signal probably resulted from a shift of TDP-43 from a phase-separated high fluorescent state into a soluble state with reduced fluorescence intensity (Zhang et al., 2026). We attributed this phenotype to the depletion of cytosolic factors and ATP during cell permeabilization because it is known that anisosome formation and maintenance require HSP70, a cytosolic ATPase (Yu et al., 2021).”.

      (c) The authors state that "TDP-43 shifts from phase-separated state to a soluble state ...", however no data is provided to support that statement.

      Since TDP-43 protein was apparently still in the nucleus after cell permeabilization (see above) but became invisible, the best interpretation is that the protein is shifted into a soluble state, which reduces the fluorescence intensity substantially. We have revised the text to clarify this point. We also cited a recent study showing that EGFP-alpha-synuclein oligomerization/aggregation enhances its fluorescence intensity in cells.

      (d) Why did the authors choose cow lover cytosol for this study?

      The main reason is because we have access to a large amount of cow liver cytosol that is known to have activities in in vitro reconstitution assays. We now cite several papers from us that reported the use of the same cytosol in other in vitro assays in the method (page 14). 

      (e) The experimental setup for supplementing with cytosol/ATP/GTP is unclear. A more detailed schematic would be helpful to understand at what stage in the experiment these factors were added. Which step of the protocol was performed at 37 {degree sign}C, which is indicated in the figure schematic but not described in the methods.

      We now revise the schematic in Figure 6A and include more details in the method and figure legend.

      (f) In the organoid model, the authors mention that they observe similar levels of total TDP-43, however they do not provide quantification. Instead, they provide a graph that shows highly significant changes in nuclear TDP-43, which was not addressed in the text.

      The total TDP-43 level was shown by immunostaining in green in Figure 7. This was used as a control to show that the increase in p-TDP-43 was not simply caused by an overall increase in its protein level. We have added the quantification to Figure 7C. We also discuss the reduced nuclear TDP-43 in organoids bearing the disease mutation in the main text.  

      Additionally, some questions remain unclear:

      (1) The anisosomes induced by ATP/GTP or cytosol are insufficiently characterized. It remains unclear whether these structures correspond to canonical ring-shaped anisosomes, and whether they exhibit dynamic (liquid-like) or more static (gel-like) properties.

      We agree that the structures reformed after incubating permeabilized cells with cytosol and ATP/GTP are not fully characterized. Due to their small size, we could not see the typical ring-shaped anisosome morphology. FRAP experiment is also tricky. Due to these issues, we have revised the text to acknowledge that we do not know the exact identity of these structures. We speculate that they are anisosome-related because like anisosome formation, it depends on cytosolic factor and energy (page 9). It is worth noting that whether these structures are anisosomes is not the main conclusion of this experiment. We conclude from this experiment that TDP-43 was still in the nucleus after cell permeabilization (not degraded). The fact that we could not see the protein likely because the protein was in a low-fluorescence soluble state.

      (2) The contribution of the cytosol and ATP/GTP supplementation experiments to the overall narrative is unclear. While the findings are intriguing, their interpretation within the context of the study is not well articulated. In particular, the rationale for including cytosol is not sufficiently justified, given that ATP/GTP alone induces a pronounced effect, whereas cytosol alone does not.

      Since the formation of anisosome requires HSP70, a cytosolic chaperone that likely needs to be imported into the nucleus, we included cytosol and ARS/GTP in our in vitro reaction. We revise the description in the result part to improve clarity (page 8-9).

      (3) The authors should address why endogenous XPO1 does not co-localize with anisosomes, whereas overexpressed XPO1 does. This raises the possibility that the observed co-localization may be an artifact of non-physiological protein levels, which should be discussed.

      As discussed in Yu H et al., Science 2021, proteins in anisosomes cannot be stained by antibodies due to an antibody accessibility issue. It was mentioned in our paper as “since antibody staining could not conclusively demonstrate the sequestration of endogenous XPO-1 in anisosomes due to an antibody penetration barrier {Yu, 2021 #937}.” We now revise this section completely to better clarify this point. We could see overexpressed XPO1 in anisosome because it has a mCherry tag.

      (4) The iPSC-based model remains insufficiently characterized. While the authors propose that this system recapitulates the accumulation of liquid and solid aggregates resembling anisosomes, it is unclear whether this phenotype is robustly observed and whether KPT treatment effectively modulates it.

      The full characterization of the iPSC-derived organoids is presented in a second paper that is posted in BioRxiv (https://www.biorxiv.org/content/10.1101/2025.11.09.687455v2), which is cited. This manuscript reports not only the accumulation of p-TDP43, but also other ALS-related phenotypes including cell death, gene transcriptional changes, cryptic exon inclusion etc. in mutant organoids.

      (5) The rationale for the selected treatment durations is unclear, and the timing appears inconsistent across experiments (ranging from 3 to 16 hours), including within experiments involving the same compound. This variability should be justified or standardized.

      The longer treatment (24 h) was used in the chemical genetic screen in which different drugs may act with different efficiency. To maximize our chance of detecting more drug effect, we used a longer treatment scheme. For later follow-up experiments involving Spuatin-1, Bortezomib, TRP, because the phenotype appears quickly. To avoid secondary effects from long treatment, we shortened the treatment to 3-5 hours. For LMB treatment, we used long treatment to reveal the steady-state phenotype (anisosome enlargement in size and reduction in number has reached maximum). This time point was determined in Figure 5A-C. In contrast, shorter treatment (5 h) was to reveal early changes that might be causal to the end-point phenotypes (e.g. the anisosome fusion phenotype could be detected as early as 5 h post-treatment). We have added some explanations in the result section to make this point clear.

      (6) Several figure legends require clarification:

      We thank the reviewer for pointing out the inconsistencies. We have corrected the outstanding issues, as explained below.

      (a) In the section stating “Collectively, our results suggest that the stability and dynamics of anisosomes are modulated by XPO1-mediated nuclear export ...", the cited figure appears to be incorrect. This should refer to Figure 5L rather than Figure 5J.

      Thanks for pointing out this error. This is now corrected.

      (b) Figure 1B: Please specify the number of replicates per concentration, the number of cells analyzed, and the model used for regression analysis. Additionally, the legend indicates a treatment duration of 15 hours, whereas Figure 1A states 24 hours.

      Due to the large sample size, each concentration was analyzed once. We have added other information to the figure legend. We also remove the redundant inaccurate information from the figure legend. The treatment time shown in the figure is correct as it was also indicated in the method.

      (c) Figure 2G: The authors state "7 anisosomes per condition," but the graph displays only 4-6 data points. Please clarify what each data point represents.

      We thank the reviewer for noticing the discrepancy and apologize for the error. We have corrected the figure legend to indicate that 4-6 anisosomes were analyzed for each condition. In Figure 2G, each data point represents the initial fluorescence loss rate averaged from the first 10 sec after reverse photobleaching.

      (d) Figures 3B and 3G: Please clarify whether a defined threshold was used to determine a "reduction in anisosome number."

      In Figure 3B, we used Z score >2 as the threshold. This is now mentioned in the legend and defined in the method. There is no Figure 3G.

      (e) Figure 4B: These do not represent biological replicates, as all samples derive from a single cell line; rather, they constitute independent experimental replicates.

      We have changed the figure legend throughout the paper accordingly.

      (f) Figures 5B and 5H: The legend states "n = 3 biological repeats," but the number of data points shown appears higher. Please clarify.

      In Figure 5B, the graph reflects data collected from 3 independent replicates. To ensure reliable baseline measurement, for each experiment, two independent control samples were included, which is why it has 6 data points. In Figure 5H, each dot represents a randomly selected imaging field. We now mention the total number of fields analyzed.

      (g) Figures 5K, 6C, and 6E: "Mean Fluorescence Intensity (MPI)" should be corrected to "MFI."

      These are all fixed. Thank you for pointing this out.

      (h) Figure 6C: Please include the number of cells analyzed and provide relevant statistical measures (e.g., R<sup>2</sup>, p-value).

      We now include the cell number in the legend and R<sup>2</sup> and p-value in the figure.

      (i) Figure 6D: The experimental timeline is unclear. Please specify the duration of incubation and the timing of each step.

      We now revise the experimental scheme in Figure 6A to better explain the experiment and the sequence of different events. For Figure 6D, permeabilized cells were incubated with cytosol with or without ARS/GTP for 40 min. This information is added to the figure legend.

      (j) Figure 7B: Improved labeling is needed (e.g., clarification of "mean spot volume") to better align with the figure legend.

      To improve clarity, we change mean spot volume to p-TDP-43 puncta mean volume. This refers to the average volume of segmented phosphorylated TDP-43-positive puncta.  

      Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      Comments on revised version.

      The manuscript has been improved with more data and clarification. The RNase T1 treatment experiment suggests that RNA is required for anisosome integrity. However, this does not directly demonstrate LMB increases nuclear RNA availability as changes in protein composition or other RNA-dependent mechanisms may also contribute. The conclusion and discussion need to be edited to consider these alternative scenarios. Overall, as most of the evidence remains indirect, the manuscript should avoid overinterpretation regarding the mechanisms underlying TDP-43 phase transition and aggregation.

      We thank the reviewer for this helpful suggestion. We have added a few sentences in the discussion (page 10) to acknowledge the limitation of the semi-permeabilized cell assay. Specifically, we mentioned that “However, our data does not exclude other RNA species or RNA-binding proteins as anisosome stabilizer. Whether RNA-dependent stabilization of anisosomes operates in the same way in intact cells also requires further validation.” We also revise our manuscript throughout to avoid over-interpretation.  

      Recommendations for the authors:

      Editor's notes:

      The value of the work is clear.

      We also recognise that it may not be possible/practical to get around the 'incomplete' appellation attached to this body of work, by further experiments. However, there may be scope here to retreat from the less well supported mechanistic claims-by editing the title, abstract and discussion and thus earn a 'solid' descriptor on a revised paper that remains a useful addition.

      The authors are best placed to consider creatively how to achieve this.

      We thank the editors for this helpful suggestion. We have revised the manuscript extensively to address every single concerns of the reviewers.

      Reviewer #1 (Recommendations for the authors):

      This revision addresses some minor concerns and adds one mechanistic experiment of genuine value. However, the major deficiencies of the original submission persist: the exclusive reliance on non-physiological TDP-43 model systems without validation in more disease-relevant contexts, the unresolved asymmetry between the two XPO1 perturbation phenotypes, the thin organoid section that now has fewer readouts than before the revision, and overstatement of the mechanistic conclusions in the abstract and Discussion. The manuscript in its current form still does not provide methods, data, and analyses that sufficiently support the primary claim that nuclear export is an established key regulator of TDP-43 phase transitions with mechanistic and disease relevance.

      As mentioned before, we have clarified the interpretation of the data, tempered conclusions where appropriate, and revised the text to explicitly acknowledge the limitations of the current study. We hope that these changes satisfactorily addressed the reviewer’s concern.

      Reviewer #2 (Recommendations for the authors):

      I would suggest adjusting the title to match the data, which shows so much more than just an effect of nuclear export.

      We thank the reviewer for this suggestion. We have changed the title to “Cellular modifiers of TDP-43 phase transition and cytoplasmic aggregation”

      When introducing the semi-permeabilized cell-based in vitro assay, it would be helpful to add a short statement describing what this model resembles, and what advantage it can bring to use this system in the context of the study.

      We have revised this section extensively and hope that improves the clarity. See marked text in page 8-9.

    1. Author response:

      The following is the authors’ response to the previous reviews

      The revisions in this version are minor and primarily include the addition of RT-qPCR validation experiments. In addition, the benchmarking data against commercial systems have now been incorporated into the main manuscript.

      We also sincerely appreciate Reviewer 2 and Reviewer 3 for their highly encouraging evaluations and recognition of our system's robustness. To fully address the remaining mechanistic queries from Reviewer 1 and the benchmarking concerns from the editors, we have performed quantitative RT-qPCR to directly measure transcript levels and have integrated our commercial benchmarking data into the revised manuscript.

      (1) The authors have satisfactorily addressed the concerns raised by the reviewers. However, the mechanistic basis of the observed performance gain remains insufficiently substantiated. The attribution of this improvement to enhanced transcription is currently speculative. This point could be directly tested by quantifying mRNA levels, for example, using real-time PCR, in both the initial and optimized systems. Such analysis would significantly strengthen the mechanistic interpretation of the results.

      To directly validate our claims regarding transcriptional efficiency, we performed quantitative RT-qPCR to determine the transcription levels of the reporter gene in both systems.

      First, we established no-reverse-transcriptase (no-RT) controls to verify complete DNA template removal. The Ct values for these controls remained above 34, confirming the absence of plasmid DNA contamination in our RNA samples.

      Second, transcript levels were calculated using the comparative 2<sup>-ΔΔCt</sup> method, normalized to the standard initial system (100 ng/μL T7) at 30 min. The optimized system achieved a 16.56-fold increase (P < 0.001) in transcript levels. In contrast, supplementing the initial system with high concentrations of T7 RNA polymerase (400 ng/μL) only yielded a 2.87-fold increase (P < 0.01)—which is nearly 6-fold lower than our optimized system.

      These findings perfectly mirror our protein-level titration assays (Figure S3C). Supplementing the initial system with excess T7 RNA polymerase fails to rescue either transcript accumulation or protein expression. This mutual validation confirms that transcription is severely bottlenecked in traditional systems due to rapid nucleotide degradation or inhibitory reaction environments. By streamlining the reaction buffer to seven core components and omitting runoff/dialysis, our system successfully relieves these systemic bottlenecks. We have incorporated these new qPCR findings into Figure 3B, the Methods, and the Results sections of the revised manuscript.

      (2) Despite the study representing an advancement towards simplifying protein expression workflows, the evidence is solid and supports the main claims however minor weakness exists i.e. the efficiency claims about the new system needs to be supported by accurate comparisons with typical cell free expression systems...

      We appreciate the editor’s emphasis on establishing standard performance benchmarks. To address this important point, we would first like to highlight that our manuscript already contains extensive, rigorous benchmarking against typical cell-free platforms widely utilized in the literature. This includes detailed head-to-head comparisons with both our 35-component "initial" system and the classical, widely established "PEP-based" system across multiple expression kinetics and western blot analyses (as shown in Figure 4 and Figures S3–S4).

      To fully embrace the editor's valuable recommendations regarding standard commercial performance, we are very pleased to formally integrate our commercial benchmarking data into the revised manuscript as Figure S3C.

      To maintain technical neutrality, we have omitted specific brand names, presenting it generically as "a high-end commercial cell-free system." The data demonstrate that our optimized system significantly outperforms this commercial alternative in both expression speed and final absolute yield, reaching an absolute productivity of 0.46 mg/mL compared to approximately 0.21 mg/mL for the commercial kit.

      We are grateful for the guidance from the editors and reviewers, which has significantly strengthened the scientific rigor of our work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important study that describes the consequences of the DNMT3A mutation in human neuronal development for the first time. The selective impact of DNMT3A function on GABAergic interneurons is interesting and an important feature of future therapeutics. The claims made in that manuscript are supported by strong evidence for the most part. And the data are of high quality in general and presented well.

      Strengths:

      The strengths of the work include: Characterization of multiple DNMT3A loss-of-function alleles, including two misense variants, R882H, P904L, and a deletion allele. The missense mutation lines both include an ideal control with the same genetic background. The CRISPRi-mediated DNMT3A knockdown has also been included. The study identifies the mTOR-PI3K pathway as a factor of overgrowth issues found in the mutant organoid. In bulk mRNA sequencing and whole-genome bisulfite sequencing, identify hypomethylated genomic regions associated with gene expression repression. Again, this is more pronounced in the ventral organoid compared to the dorsal organoid. In addition, the extensive electrophysiological characterizations with a high-density microelectrode array support the more mature status of mutant interneurons.

      Weaknesses:

      Although a strong study overall, some weaknesses are noted. These include:

      (1) The lack of validation data for the generated iPSCs and hESCs, such as the chromosomal contents, ploidy, and pluripotency states.

      We thank the reviewer for their constructive feedback. We previously validated our 882 models with whole genome sequencing and teratoma formation upon mouse fat pad injection, while the parental human embryonic stem cell line (WA01 hESCs) used for P904L variant knock-in was validated by our Genome Engineering Stem Cell (GESC) core upon derivation of that variant knock-in model. We have now added both karyotyping and pluripotency staining (SOX2/OCT4) for all other hPSC lines as (new) Supplementary Figure S17 and included further description in our Methods section under “hPSC Model Generation and Culture” (pg. 19, lines 6-7).

      (2) Other weaknesses relate to data interpretation and insufficient discussion of related matters, as detailed in the recommendations to the authors.

      We thank the reviewer for their insightful suggestions and have detailed our responses in the “recommendations to the authors” section.

      (3) Also, some errors are noted and detailed in the recommendation section.

      We thank the reviewer for catching these errors and have since corrected them, with detailed responses below.

      Reviewer #2 (Public review):

      Summary:

      Chapman, Determan et al. investigate how pathogenic mutations in DNMT3A, which cause Tatton-Brown-Rahman Syndrome (TBRS), disrupt human cortical developmental processes using a comprehensive panel of human pluripotent stem cell models spanning DNMT3A loss-of-function severity. The authors aim to identify the cellular and molecular mechanisms underlying TBRS-associated brain overgrowth and intellectual disability, and to test whether mechanistic convergence exists between TBRS and other overgrowth-intellectual disability disorders (OGIDs) caused by mutations in EZH2 (Weaver syndrome) or PIK3CA pathway components. Their central conclusion is that GABAergic interneuron development is selectively vulnerable to DNMT3A mutation, where reduced DNA methylation causes premature de-repression of neuronal and synaptic genes, driving precocious neuronal maturation and hyperactivity sufficient to disrupt neuronal network synchrony. This report adds to a growing literature supporting the vulnerability of GABAergic interneurons in NDDs and further provides a mechanistic view of this vulnerability, potentially convergent across OGIDs. The mechanistic claims around H3K27me3 compensation and mTOR-based therapeutic convergence, while promising, rest on more preliminary evidence and would benefit from the distinction between correlation and mechanism being made more explicit in the text. Overall, this is a compelling study with a rigorous experimental design and novel findings with a potential impact on a better understanding of the OGID pathophysiology.

      Strengths:

      (1) A major strength of this work is the breadth and rigor of the disease modeling approach. Four independent TBRS model systems are used in tandem: a patient-derived iPSC line with isogenic CRISPR-corrected control (R882H), a knock-in hESC model (P904L) with its wild-type isogenic, patient deletion iPSC lines (Del1/2), and CRISPRi knockdown models (G1/G2), collectively spanning a range of DNMT3A loss-of-function that correlates with phenotypic severity. This allelic series design substantially strengthens causal inference beyond what any single isogenic pair could provide.

      (2) The multi-omic integration across matched developmental stages provides a strong mechanistic foundation for the cellular phenotyping and provides significantly enhanced novelty. RNA-seq, whole-genome bisulfite sequencing, and H3K27me3 CUT&Tag are combined in the same cell types, and timepoints show that DNMT3A loss reduces CG methylation at neuronal and synaptic gene loci, leading to premature transcriptional activation.

      (3) The selective vulnerability of ventral (GABAergic) versus dorsal (glutamatergic) progenitors is one of the study's most important findings. This lineage specificity is consistently observed across all model systems and in both 2D and organoid formats, where ventral NPCs show increased proliferation, premature neuronal gene expression, and increased neurogenesis, while dorsal NPCs are largely unaffected at the transcriptomic and cellular level despite exhibiting comparable DNA methylation changes. This adds to a body of emerging work showing GABAergic interneuron vulnerability in NDDs where ubiquitously expressed genes such as chromatin modifiers are perturbed, and provides additional molecular insights into potential mechanisms of "resilience" of dorsal populations.

      (4) The functional characterization follows a logical progression from single-neuron electrophysiology (demonstrating GABAergic hyperactivity with increased action potential amplitude and firing rate) to network-level analysis using high-density multi-electrode arrays. The HD-MEA experimental design - pairing TBRS or control GABAergic neurons with a constant background of control iGlut neurons - cleanly isolates GABAergic dysfunction as the driver of network hypersynchrony.

      Weaknesses:

      (1) The concomitant induction of proliferation and differentiation in TBRS V-NPCs is conceptually striking, since these are generally considered antagonistic developmental programs. The authors partially address this tension by noting that DNMT3A LOF alone is insufficient to initiate neuronal differentiation, i.e., V-NPCs upregulate neuronal and synaptic genes while retaining progenitor identity, implying that transcriptomic priming and commitment to differentiation are decoupled. However, the relationship between the proliferative phenotype and the epigenetic priming phenotype remains mechanistically unresolved. The manuscript documents mTOR pathway upregulation at the protein level and identifies shared DEGs that include proliferative regulators, but it does not establish whether mTOR-driven proliferation and mCG-loss-driven neuronal gene de-repression/enhanced differentiation are causally linked or represent two independent consequences of DNMT3A LOF.

      We thank the reviewer for their comment and agree that this phenotype, whereby progenitors exhibited both increased proliferation and hallmarks of gene expression associated with neuronal differentiation is striking and interesting, given that these are typically antagonistic paradigms during normal development.

      We documented that these phenotypes involve upregulated expression of both neuronal/synaptic and proliferative genes in V-NPCs (Figure 2d), with concomitant loss of repressive DNA methylation at regulatory elements associated with these genes (Figure 2f, Supplementary Data 5). In this work, DNMT3A mutation had a more prominent role in de-repressing neuronal and synaptic gene expression to promote hallmarks of neuron differentiation, while playing a relatively less central role in direct regulation of proliferation genes, as seen from the relative prominence of neuronal/synaptic- versus proliferation-related GO terms in our Supplementary Data 5 table (pg. 6, lines 16-19).

      To examine the mechanisms underlying increased V-NPC proliferation in our TBRS models, we assessed a potential relationship with the PIK3/AKT/mTOR pathway, as this is implicated in increased proliferation resulting from DNMT3A-associated mutation in myeloid leukemia (Dai et al., 2017, PMID: 28461508). In our work, DNMT3A mutation increased the expression and/or phosphorylation of mTOR signaling pathway targets specifically in V-NPCs (Figure 1q-r, Supplementary Figure S3a-d). However, while TBRS mutation directly affected repressive DNA methylation at a suite of cell proliferation-related genes, these did not include the PIK3/AKT/mTOR pathway genes themselves, suggesting an indirect relationship between altered DNA methylation and increased mTOR signaling.

      We have since incorporated discussion of how DNMT3A-mediated gene repression and levels of PIK3/AKT/mTOR pathway signaling may be interacting, providing a framework for future studies to identify how these related OGID gene mutations may converge mechanistically (pg. 5, lines 19-21; pg. 16, lines 8-10).

      (2) Relatedly, the rapamycin rescue experiment is a valuable proof-of-concept for the PIK3/AKT/mTOR convergence but is limited to a single dose in a single model (882) with a single readout (Ki67+ proliferation). Given the prominence of mTOR pathway convergence in the manuscript as a potential shared therapeutic avenue across OGIDs, the data supporting this claim are somewhat preliminary. It remains unknown whether mTOR inhibition rescues downstream phenotypes (neurogenesis, gene expression, neuronal maturation) or whether less severe TBRS models respond similarly. This might also help tackle the first comment above. e.g., if mTOR inhibition rescued proliferation but not the transcriptomic priming, that would support two independent mechanisms.

      We thank the reviewer for their comment. We explored both the overall levels and phosphorylation of proteins involved in PIK3/AKT/mTOR signaling in the 882, 904, Del1, Del2, and KO V-NPC models (Figure 1q-r, Supplementary Figure S3a-d), finding specific increases of all proteins. We showed that rapamycin addition reversed the increased proportion of KI67+ proliferating cell nuclei resulting from 882 mutation in V-NPCs in main Figure 1s, while demonstrating that rapamycin also reduced the proportion of KI67+ nuclei observed in both less severe 904 and Del1 V-NPC models (Supplementary Figure S3e-f).

      We agree that understanding whether rapamycin treatment can rescue TBRS neuronal phenotypes would be very interesting, as previous work on Tuberous Sclerosis Complex has utilized rapamycin and other mTOR inhibitors to effectively reverse TSC-related alterations of neuronal morphology and neuronal hyperexcitability (Buttermore et al., 2025, PMID: 40792287). Future studies examining convergent mechanisms and therapeutics for OGIDs should examine how similarly targeting this and related pathways rescues altered neuronal morphology, maturation, and function, as we have demonstrated that TBRS mutation has subsequent consequences for V-IN differentiation, maturation, and function. This point has been detailed in the discussion section on pages 15-16.

      (3) The claim that H3K27me3 compensates for mCG loss is an important mechanistic point, but the current data do not distinguish between active compensation, in which EZH2 is recruited in response to methylation loss, and functional redundancy, in which H3K27me3 is independently established and becomes the dominant repressive mark once DNA methylation is reduced. The EZH2 knockdown/inhibition experiments show that H3K27me3 is sufficient to maintain repression at hypo-DMR sites, but they do not establish that H3K27me3 gain is itself a response to methylation loss. Because H3K27me3 profiling was performed only in the severe 882 model, it is also unclear whether H3K27me3 gain scales with DNMT3A LOF severity, as a compensatory model would predict. Finally, the EZH2 overexpression rescue is performed in V-NPCs, whereas the compensation model is developed primarily in D-NPCs, making it difficult to assess whether the same mechanism operates in the lineage where it was originally inferred.

      We thank the reviewer for the opportunity to clarify our findings and experimental reasoning. A previous study using a conditional Dnmt3a knockout mouse model (Li et al., 2022, PMID: 35604009) demonstrated increased expression of multiple PRC2 components following the loss of Dnmt3a. This study demonstrated that sites which lost DNA methylation gained H3K27me3 in postnatal neurons upon Dnmt3a loss. Therefore, we hypothesize that the gain of H3K27me3 likely occurs in response to loss of DNMT3A methylation.

      While we did not perform CUT&Tag for H3K27me3 in our less severe models, we did validate gene expression changes following EZH2 knockdown and inhibition in both the R882H (Figure 4g-h) and P904L (Supplementary Figure S8b) models, finding that gene expression was unchanged in the model with the less severe DNMT3A mutation (P904L). Based upon these findings, we hypothesized that compensatory H3K27me3 may occur only upon severe DNMT3A loss, as seen in the dominant-negative R882H model. Furthermore, as H3K27me3 compensation was more prominent in D-NPCs, we hypothesized that this might be sufficient to prevent de-repression and aberrant neuronal gene repression upon loss of DNMT3A-mediated repression in D-NPCs. However, since TBRS mutation caused the most prominent de-repression of neuronal gene expression in V-NPCs, we also tested whether EZH2 overexpression could reverse this, finding that it partially suppressed this dysregulated neuronal gene expression. To better clarify this logic and the findings, we have made text edits to this results section and referenced Li et al., 2022 in both the results (pg. 9, lines 16-21; pg. 10, lines 4-7,10-12) and discussion (pg. 16, lines 12-15).

      (4) The narrative framing of dorsal neuron development as unaffected by DNMT3A LOF is somewhat at odds with the data presented. The 882 D-NPCs show substantial DNA methylation changes, and TBRS D-INs exhibit what the authors describe as "substantive transcriptomic differences" involving persistent expression of pluripotency and progenitor genes, which seems to be a distinct but potentially significant phenotype. The impact of DNMT3A loss between ventral and dorsal lineages might be more accurately framed as divergent in nature rather than specific to a certain population.

      We thank the reviewer for their comment. While TBRS mutations appear to have a significantly stronger effect on V-NPCs and subsequently V-INs, both transcriptomic and methylation alterations do also occur upon TBRS mutation in D-NPCs and D-INs, as noted in Supplemental Figure S4d, S11, and Supplemental Data 2. However, we observed substantially greater molecular alterations in V-NPCs/V-INs, a lack of overt cellular phenotypes in D-NPCs where assayed, and a lack of functional consequences in matured D-INs, suggesting a more significant requirement for DNMT3A in regulating the differentiation and subsequent maturation of cortical inhibitory interneurons during embryonic and early pre-natal development, the developmental periods that we can readily model in hPSC-derived neurons.

      It should also be noted that these hPSC differentiation models do not recapitulate post-natal deposition of non-CpG (mCA) DNA methylation, a mechanism disrupted postnatally by TBRS-associated mutations in our prior work in murine models (Harrison Gabel; e.g. Beard et al., 2023, PMID: 37952155), which we have now added in the results section (pg. 7, lines 8-11). Therefore, we hypothesize that if we could sufficiently mature D-INs to a state that modeled postnatal development and recapitulated this non-CpG methylation, we might be able to detect cellular and functional phenotypes in later stage D-INs. To avoid misinterpretation, we have altered the language in the results section to confirm that there are both transcriptomic and methylation changes in our D-NPCs/D-INs, but that these are not accompanied by cellular phenotypes or neuronal dysfunction (pg. 6, lines 1-2; pg. 7 line 3; pg. 7, lines 20-23; pg. 8, lines 1-2; pg. 8, lines 18-23; pg. 9, lines 1-4).

      (5) SST stainings are not entirely convincing. They appear mostly nuclear, and some instances localized to rosettes in organoids, whereas the protein is largely confined to processes and is expected to be found outside progenitor-rich zones like rosettes.

      We agree that the perinuclear SST staining detected in these young ventral telencephalic-patterned organoids at day 30 differs somewhat from the more process-localized and cytosolic signal seen in later stage organoids in other studies. This may be related to the use of different commercial SST antibodies across studies but also likely reflects SST immunoreactivity in newborn neurons near the onset of SST expression. For example, immature SST-immunoreactive neurons in the early postnatal rat cortex exhibit predominant SST staining in perinuclear cytoplasm and short processes (e.g. Fig. 3 in Lee et al, PMID: 9664223) while acquiring more cytosolic and process-localized staining as postnatal neuron maturation occurs. Evaluation of immunopositivity for other markers of neurogenesis (ASCL1) and immature neurons (TUJ1) is also congruent with these findings for SST, with TBRS-associated mutations increasing in the fraction of cells in V-NPCs/V-ORGs that express these three markers.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors investigated TBRS etiology by using new human pluripotent stem cell models, modeling varying levels of TBRS-associated loss of DNMT3A function. They identified increased lineage-specific proliferation of precursors in TBRS ventral MGE-like progenitors, which they propose was related to increased signaling through the PIK3/AKT/mTOR pathway. Furthermore, they show that reduced DNA methylation during MGE-like progenitor differentiation into GABAergic interneurons can cause a premature expression of neuronal and synaptic genes, triggering precocious neuronal maturation. In conclusion, they propose that TBRS-derived GABAergic neurons exhibit hyperactivity that can alters the development and structure of neuronal networks.

      Strengths:

      Overall, the data presented is convincing, from an early developmental point of view, given that the iPSC-derived 2D cultures or organoids used do not get to reach a mature state. Nonetheless, the data clearly show the effects that deleterious mutations in TBRS can cause during the period of neurogenesis, which was missing in the field.

      Weaknesses:

      (1) Li et al., 2022 (referred to in the manuscript) seems to already show the interplay between H3K27me3 and Dnmt3a discussed in this study i.e., that in the absence of DNA methylation, there is an expansion of polycomb-like repression. These data should be better acknowledged in the paragraph 'Repressive H3K27me3 compensates for severe loss of DNA methylation' (page 9), given it supports the data presented in this manuscript and suggests this as a common mechanism in the interplay between these two repressive marks, as it is well established in the literature.

      We thank the reviewer for this suggestion. We have now added Li et al., 2022 to both the results section (pg. 9, lines 16-20) and our discussion section (pg. 16, lines 12-13).

      (2) The authors should acknowledge that the omics data come from a mixed population of cells.

      We thank the reviewer for their comment. We have validated that the established 2-D differentiation methods we used in this study generate cell populations with >85-90% enrichment for the desired progenitor and neuronal cell type, based upon marker expression, but acknowledge that these are bulk -omics data obtained from cells that may represent a mixed population and have now detailed this in the methods section under “Sequencing” (pg. 21, lines 16-18).

      (3) The authors are encouraged to further discuss whether the overgrowth observed in ventral GABAergic cultures or organoids compares to the overgrowth observed in diseased patients. One expects MRIs to have been performed in patients and that these could be harnessed to discern if overgrowth occurs in the cortex or ventral regions of the brain.

      We thank the reviewer for their suggestion and do note that at least one published study documents increased cortical thickness in the MRIs of TBRS patients (Jiménez de la Peña et al., 2024, PMID: 37795572); however, to our knowledge studies have not examined regional or cell type-selective overgrowth of cortical tissue in TBRS patients. Future clinical studies examining the nature of the neuronal progenitor overgrowth and resulting consequences for patient brain imaging would be of interest to better understand TBRS-associated etiology of brain overgrowth and its manifestations.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A potential explanation for the ventral organoid's sensitivity compared to the dorsal can be provided. I understand that the PRC2-mediated methylation loss was more pronounced in the ventral than in the dorsal. However, this raises the question: why is this so? Are the PRC2 components more highly expressed in ventral organoids or in GABAergic neurons than in glutamatergic neurons?

      We thank the reviewer for this question. Upon examining RNA-sequencing data for control V-NPCs and D-NPCs, we found that expression of the PRC2 subunits SUZ12, EED, and EZH2 is significantly elevated in D-NPCs relative to V-NPCs (see Author response image 1). This could contribute to restoring epigenetic repression in D-NPCs in the context of DNMT3A mutation, to yield the more limited transcriptomic changes and greater gain of H3K27me3 we observed in TBRS D-NPCs, despite their comparable levels of global mCG loss). This could also account for our finding that overexpression of EZH2 in V-NPCs could normalize some gene expression changes observed in the TBRS models.

      Author response image 1.

      (2) DNMT3A is a major enzyme that mediates non-CpG methylation during synaptogenesis. There is no discussion of this phenomenon or any related analysis. I suppose the brain organoids may be too immature to detect non-CpG methylation. Even then, this can be described in the results sections.

      We thank the reviewer for this comment. While prior work has examined an important role for DNMT3A in postnatal deposition of non-CpG methylation (Christian et al. 2020; Beard et al. 2023), we found that our hPSC models, as the reviewer suggests, were too immature to detect appreciable levels of non-CpG methylation. Accordingly, we have focused our work here on neurodevelopmental requirements for DNMT3A; this work demonstrated that many of the functional alterations of neuronal network activity originate from altered neurogenesis and neuronal maturation of cortical GABAergic interneurons. We have now included a description of the important role of non CpG methylation by DNMT3A during postnatal development and have indicated that the immaturity of hPSC-derived neurons precludes our ability to study non-CpG methylation using these models in the results (pg. 7, lines 8-11) with present language in the discussion (pg. 15, lines 8-11).

      (3) Given the Figure 4 results showing H3K27 methylation and EZH2 knockdown rescue the DNMT3A mutation, the authors argue that EZH2 and DNMT3A regulate a similar set of genes and that the two diseases are related. To claim this, there should be data showing that gene misregulation upon loss of EZH2 and DNMT3A is similar.

      Given our data, it is premature to draw direct relationships with the dysregulated genes we detected in TBRS versus the molecular basis of Weaver Syndrome; therefore, we have attempted to better reflect the potential but currently unproven relationship between these OGIDs and altered the respective language in the results section (pg. 10, lines 10-12) to reflect this. Understanding convergent mechanisms of OGIDs involving both DNMT3A (TBRS) and EZH2 (Weaver Syndrome) gene mutations will be a compelling topic for future studies, as our work here suggests that they may regulate similar gene suites.

      (4) Stronger phenotypes were observed in the R882H mutant compared to the P904L mutant. However, this cannot be translated to the functionality of the mutant protein or disease severity because one is on iPSCs and another is made in hESCs. The limitation should be noted to avoid misleading.

      We agree that the severity of each pathogenic mutation can be influenced by its presence on a different iPSC vs hESC background; accordingly, our study design indicated the hPSC background for each model and employed paired isogenic controls to clearly define the consequences of each TBRS mutation relative to a model with an identical genetic background but lacking the mutation. Our finding that the R882H mutation resulted in more severe epigenomic and transcriptomic consequences than the P904L mutation is congruent with findings made in prior work (Beard et al., 2023 PMID: 37952155; Russler-Germain et al., 2014 PMID: 24656771).

      (5) It is unclear what measurements were done for DNMT3A level quantification shown in Figure 1e- f.

      Protein quantification for DNMT3A models was performed by western blot, shown in Supplemental Fig. S16. We have now included reference to whole blots in the methods section under “Cellular Phenotyping” (page 20 line 13).

      (6) The results section of the paper for Figures 6 and 7 cites the wrong figure numbers and panels.

      We thank the reviewer for catching these errors and have corrected them in the revised manuscript draft.

      (7) Some typos for PI3K (meaning PIK3) in several places, including the Abstract.

      We thank the reviewer for catching these errors and have corrected them in the revised manuscript draft.

      Reviewer #2 (Recommendations for the authors):

      (1) There is a figure numbering discrepancy in the manuscript - the text references six main figures, but the figure pages include seven, with the MEA network data apparently mislabeled.

      We thank the reviewer for catching these errors and have corrected them in the new manuscript draft.

      (2) The nomenclature "D-IN" for dorsal immature neurons is potentially misleading, as "IN" conventionally denotes interneurons, which these glutamatergic cells are not.

      While we agree that IN could be interpreted as interneurons, this nomenclature is defined at an early point in the manuscript and was used to allow readers to easily identify compare findings made in D-NPCs and D-INs (and, as a counterpart, findings made in V-NPCs versus V-INs).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1.1) The manuscript “Realistic coupling enables flexible macroscopic traveling waves in the mouse cortex” by Sun, Forger, and colleagues presents a novel computational framework for studying macroscopic traveling waves in the mouse cortex by integrating realistic brain connectivity data with large-scale neural simulations.

      The key contributions include: (1) developing an algorithm that combines spatial transcriptomic data (providing detailed neuron positions and molecular properties) with voxelized connectivity data from the Allen Brain Atlas to construct neuron-to-neuron connections across 300,000 cortical neurons; (2) building a GPU-accelerated simulation platform capable of modeling this large-scale network with both excitatory and inhibitory HodgkinHuxley neurons; (3) extending phase-based analysis methods from 2D to 3D to quantify traveling wave activity in the realistic brain geometry; and (4) demonstrating that realistic Allen connectivity generates significantly higher levels of macroscopic traveling waves compared to simplified local or uniform connectivity patterns.

      The study reveals that wave activity depends non-monotonically on coupling strength and that slow oscillations (0.5-4 Hz) are particularly conducive to large-scale wave propagation, providing new insights into how anatomical connectivity enables flexible spatiotemporal dynamics across the cortex.

      The authors leverage two existing dense datasets of spatial transcriptomic data and connection strength between pairwise voxels in the mouse cortex in a novel way, allowing for the computational model to capture molecular and functional properties of neurons as determined by their neurotransmitter profiles, rather than making arbitrary assignments of excitatory/inhibitory roles. Additionally, the author’s expansion of 2D phase dynamics to 3D phase gradient analysis methods is important and can be widely applied to calcium imaging, LFP recordings, and likely other electrophysiological recordings.

      Thank you for the accurate summary of our manuscript and list of strengths.

      (1.2) The model’s Allen connectivity approach overlooks critical aspects of real cortical dynamics. Most importantly, it excludes subcortical structures, especially the thalamus, which drives cortical traveling waves through thalamocortical interactions. The authors’ method of electrically stimulating all layer 4 neurons simultaneously to initiate waves is artificially crude and bears little resemblance to natural wave generation mechanisms.

      We agree that excluding subcortical structures, especially the thalamus, is an important limitation of the current model. Because adding these structures would substantially expand the scope and complexity of the model, we now state this limitation more explicitly in the Discussion and leave it as a future extension:

      “For simulations, we choose to randomly stimulate the total population of layer 4 neurons as a way to mimic subcortical input and generate traveling waves, which can be unrealistic. Subcortical structures, such as the thalamus, are vital to cortical dynamics like slow-wave activity [1] and are known to regulate traveling waves [2]. Therefore, a direct and important future improvement would be adding subcortical structures to the model.”

      We also agree that the original constant-current stimulation was too artificial. We therefore replaced it with a 10Hz Poisson spike train delivered to layer-4 excitatory neurons across the isocortex, which more closely mimics the stochastic input that cortex receives from subcortical regions such as the thalamus. The revised stimulation protocol is described in the Results:

      “Therefore, to mimic the stochastic drive that cortex receives from subcortical regions like thalamus, we deliver a 10Hz Poisson spike train to layer-4 excitatory neurons (Figure 2a), since layer 4 is the canonical thalamocortical input layer [3]; each Poisson event applies a fixed voltage bump V<sub>stim</sub>, and no other external input is applied.”

      as well as in the Methods 4.5 Simulation.

      The revised stimulation protocol allowed us to rerun the parameter sweep under stochastic drive. A direct comparison of alternative wave-generating mechanisms remains an important direction for future work.

      (1.3) The model handles voxel-to-voxel connections crudely when neurons have mixed excitatory/inhibitory properties and varying synaptic strengths. Real connectivity differs dramatically between neuron types (pyramidal cells vs. interneurons, across cortical layers), but the model only distinguishes excitatory and inhibitory neurons. Additionally, uniform synaptic weights ignore natural variations in connection strength based on neuron type, distance, and functional role. Integrating the updated thalamocortical dataset mentioned by the authors, even at regional resolution, would substantially improve the model.

      We thank the reviewer for raising these important points regarding cell-type-specific connectivity and heterogeneity of synaptic weights. We agree that cell-type-specific connectivity and heterogeneous synaptic weights are important limitations. Because the current voxelized projectome is not cell-type specific, we now state these limitations explicitly and outline how future versions of the model could incorporate improved density and synaptic weight assumptions in the Discussion (Construction of the Allen connectivity). Specifically, we now write:

      “First, the voxelized projection data is not cell-type specific. In our model, we distinguish only two neuronal populations: glutamatergic (excitatory) and GABAergic (inhibitory), based on the Zhuang-ABCA-1 transcriptomic dataset. However, real cortical connectivity differs dramatically between more refined cell types: pyramidal neurons and interneurons have distinct connection targets, and connectivity is strongly layer-specific. At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse-density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.

      Also, our model uses uniform synaptic weights within each synapse type: every AMPA synapse has conductance g<sub>AMPA</sub> and every GABA synapse has conductance g<sub>GABA</sub>. In reality, synaptic strength varies with presynaptic and postsynaptic cell type, anatomical distance, and the distribution of synaptic weights is typically heavy-tailed (e.g. log-normal [5]). Extending our algorithm to sample synaptic weights from realistic distributions would be a natural next step and is likely necessary for quantitative comparisons with electrophysiological recordings.”

      These additions clarify which aspects of the current model are constrained by the available voxelized projectome and which extensions would require cell-type-resolved connectivity data, region-specific density information, or more detailed synaptic-weight estimates.

      (1.4) While the authors bridge microscopic (single neuron) and mesoscopic (regional connectivity) data to study macroscopic (whole-cortex) waves, they don’t integrate the distinct mechanisms operating at each scale. The framework demonstrates that realistic connectivity enables macroscopic waves but fails to connect how wave dynamics emerge and interact across spatial scales systematically.

      We thank the reviewer for this insightful comment. In the revision, we added the Kuramoto synchrony analysis as a first step toward connecting scales: changes in the microscopic coupling parameter alter network synchrony, and this synchrony measure is closely associated with the macroscopic PGD observable (new Fig. 5). We agree that a full separation of layer-specific microcircuit mechanisms, region-specific connectivity motifs, and whole-cortex wave propagation remains beyond the scope of the present study. The revised text now frames the synchrony-to-wave relationship as one concrete cross-scale link that can be investigated further with this framework.

      (1.5) Claims that Allen connectivity produces higher phase gradient directionality (PGD) than local connectivity appear limited to delta oscillations at very specific coupling strengths and applied currents. Few parameter combinations show significantly higher PGD for Allen connectivity, and these are generally low PGD values overall.

      We agree that the original comparison did not sufficiently establish whether the Allen-connectivity advantage held beyond a small number of parameter choices.

      In the revised manuscript, we expanded the analysis to a 10 × 12 grid of excitatory coupling strengths and Poisson stimulus magnitudes for Allen, local, and uniform connectivity. Across this full (g<sub>AMPA</sub>,V<sub>stim</sub>) grid, Allen connectivity shows higher per-band maximum PGD than local or uniform connectivity, especially in the theta, alpha, and beta bands (Fig. 5b). The delta-band difference is small, consistent with the reviewer’s observation that the original delta-band result did not clearly separate Allen from local connectivity.

      In the revised manuscript, Fig. 5c and Fig. 5d show how PGD varies with stimulus magnitude and coupling strength individually. The full per-band PGD heatmaps for all three connectivities, together with the per-band Allen-minus-opponent gap bars, are provided in Figure 4-figure supplement 6. These additions show that the Allen-connectivity trend is not limited to a single representative operating point.

      (1.6) Broadly, it’s unclear how this computational framework can study memory, learning, sleep, sensory processing, or disease states, given the disconnect between simulated intracellular voltages and the local field potentials or other electrophysiological measurements typically used to study cortical traveling waves. While computationally impressive, the practical research applications remain vague.

      We thank the reviewer for this important point. To bridge the gap between our simulations and experimentally measured data, such as local field potentials (LFP), we used a post-hoc LFP estimation pipeline and added a dedicated Methods subsection describing it (LFP estimation from intracellular voltage). The key idea is that LFP primarily reflects the net transmembrane synaptic current in a local population, which we can reconstruct directly from the voltage traces and the connectivity used in the simulation. Please see Methods section LFP estimation from intracellular voltage.

      This pipeline allows us to test whether the macroscopic traveling-wave structure identified in the voltage traces is also present in an LFP-like signal. In the revised manuscript, we added a new Results section and show, in Fig. 6, side-by-side voltage- and LFP-based snapshots, the voltage-LFP PGD scatter (Pearson r = 0.78-0.89 per band), and per-band PGD comparisons across Allen, local, and uniform connectivity on the LFP signal. These results indicate that our conclusions are not restricted to intracellular voltage and provide a closer bridge to LFP-based experimental measurements.

      (1.7) The paper needs a clearer explanation for why medium coupling (100%) eliminates waves in Allen connectivity (Figure 6) while stronger coupling (150%) restores them.

      Thank you for requesting this clarification. The revised analysis suggests that the non-monotonic relationship between coupling strength and wave activity reflects an interaction between network synchrony and spatial organization. At weak coupling, the network has enough coordination to support propagating waves. At medium coupling, increased synaptic drive pushes the network into an asynchronous irregular state that disrupts coherent wave fronts. At strong coupling, rhythmic synchronization is re-established and again supports wave propagation.

      To substantiate this explanation quantitatively, we measured the Kuramoto order parameter R(t)=|〈e<sup>iϕ(x,t)</sup>〉<sub>x</sub>| from the generalized phase ϕ(x, t) of the band passed voltage field (see updated Quantitative measurement of neuronal activity in Methods) and reduced it to the maximum over each 1-s recording window. We then swept the same ten g<sub>AMPA</sub> values (0.005-0.050 nS) and twelve stimulus magnitudes used for the PGD sweeps, for Allen, local and uniform connectivity, in all five frequency bands. The new analysis is presented in Fig. 5: panel (e) shows synchrony versus g<sub>AMPA</sub> for the three connectivities, panel (d) shows the matching PGD curve, and panel (f) shows the synchrony-PGD scatter with Pearson r per connectivity.

      The synchrony curve captures the main peak-dip-recovery structure of the PGD curve. For Allen connectivity in the alpha band, mean synchrony peaks at weak coupling, collapses to a 75 % suppression in the medium-coupling window g<sub>AMPA</sub> = 0.020-0.035 nS, and recovers near unity at g<sub>AMPA</sub> ≥ 0.040 nS. Uniform connectivity follows the same U-shape with a slightly earlier dip. Per-band versions of the PGD- and synchrony-vs-coupling curves are provided in Figure 5-figure supplement 1, showing that the peak-dip-recovery profile holds across all five canonical bands for Allen and uniform connectivity. Across the 120- point (g<sub>AMPA</sub>, V<sub>stim</sub>) grid, synchrony and PGD are positively correlated in every band and every connectivity (Pearson 0r ranging from ≈ 0.30 to ≈ 0.92 across band-connectivity combinations; per-band scatters in Figure 5-figure supplement 2). Local connectivity follows a different trajectory: its synchrony is moderate at weak coupling and decays monotonically with gAMPA without recovering at strong coupling. This is consistent with local connectivity not supporting large-scale propagating waves at strong coupling, so the peak-dip-recovery interpretation applies mainly to Allen and uniform connectivity. We have added this synchrony analysis to the revised manuscript:

      “To diagnose the mechanism behind this profile, we measured the Kuramoto order parameter R(t) = |⟨e<sup>iϕ(x,t)</sup>⟩<sub>x</sub>| from the generalized phase field ϕ(x, t) of each frequency band and recorded its maximum over each simulation window (Quantitative measurement of neuronal activity). The resulting synchrony curve (Figure 5e) resembles the trend of the maximum PGD well (Figure 5d). For Allen connectivity in the alpha band, mean synchrony peaks at weak coupling, collapses in the medium-coupling window, and recovers at strong coupling. Across the full 120-point (g<sub>AMPA</sub>, V<sub>stim</sub>) grid, synchrony and PGD are positively correlated in every band and every connectivity (Pearson r = 0.30- 0.92, all p < 10−3; Figure 5f, with per-band scatters in figure Supplement 2). Therefore, the PGD trough at medium coupling may be a synchrony trough: increased synaptic drive pushes the network into an asynchronous irregular state, while strong coupling re-establishes rhythmic synchronization that supports wave propagation. This synchrony-to-wave bottleneck seems more significant to the networks with long-range connectivity (Allen, uniform), partly because long-range connections can augment synchrony in the coupled neuronal network [6]. We also note that although uniform connectivity is able to achieve almost an identical level of synchrony to that of Allen connectivity, the PGD remains much lower due to the loss of spatial organization within.”

      (1.8) Does using a single connectivity parameter (ρ = 300) across all regions miss important regional differences in cortical connectivity density?

      We thank the reviewer for raising this important point. We agree that a single global density parameter misses region-to-region variation in local synaptic density, and we have extended the Discussion (Construction of the Allen connectivity) to state this limitation alongside the cell-type-specific connectivity and synaptic-weight limitations discussed in response to Point 1.3. Specifically, we now write:

      “At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.”

      This paragraph is placed directly after the cell-type and weight-heterogeneity limitations added in response to Point 1.3.

      Reviewer #2 (Public review):

      (2.1) This work presents a spiking network model of traveling waves at the whole-brain scale in the mouse neocortex. The authors use data from the Allen Institute to reconstruct connectivity between different neocortical sites. They then quantify macroscopic traveling waves following stimulation of all layer 4 neurons in the neocortex.

      Overall, the results are interesting and shed new light on the dynamic organization of activity across the neocortex of the mouse. The paper uses realistic neuron models specifically fit to intracellular recordings, demonstrating that traveling waves occur in the mouse neocortex with both realistic connectivity and realistic single-neuron dynamics. The paper is also well-written in general. For these reasons, the authors have generally achieved their aims in this work.

      We thank Reviewer 2 for the positive assessment and accurate summary of our work.

      (2.2) Description of Algorithm 1: While the Methods section clearly explains the density parameter ρ, the statement on line 358 concerning the “ideal” average number of connections is a little unclear. The authors should explicitly clarify that ρ is a free parameter that can be adjusted to balance computational feasibility (for a given set of computational resources) and biological fidelity. The ρ parameter used here results in approximately 300 connections per neuron on average. The authors should state clearly that the number of connections per cell is the key determinant of computational feasibility (cf. Morrison et al., Neural Computation, 2005). The authors should also review neuronal density and synaptic connectivity in the mouse neocortex and clearly reference density and connectivity in their model to the biological scales found in the mouse.

      We thank the reviewer for raising this important point about our connectivity algorithm and simulation. We have clarified in the revised Methods (Use of the voxelized connectivity data, Methods 4.2) that ρ is a free parameter that controls the trade-off between computational feasibility and biological fidelity. Specifically, we now write:

      “The density parameter ρ is a free parameter that controls the trade-off between computational feasibility and biological fidelity: higher values of ρ yield more connections per neuron and thus higher biological realism, at the cost of greater memory and runtime [7].”

      A careful accounting of the biological scales involved (synapse density per neuron, total cortical population) and incorporating region or cell-type-specific connection density is left as a future direction. We have noted this in the Discussion subsection:

      “First, the voxelized projection data is not cell-type specific. In our model, we distinguish only two neuronal populations: glutamatergic (excitatory) and GABAergic (inhibitory), based on the Zhuang-ABCA-1 transcriptomic dataset. However, real cortical connectivity differs dramatically between more refined cell types: pyramidal neurons and interneurons have distinct connection targets, and connectivity is strongly layer-specific. At the same time, we use a single global density parameter ρ to set the average number of connections per neuron across the entire cortex. A region-specific ρ could better capture the known variation in local synaptic density across cortical areas, for example the higher synaptic density of primary sensory regions relative to higher-order association areas. Future versions of the model could allow a region- and cell-type-specific ρ derived from region-level synapse-density atlases together with cell-type-resolved connectivity [4], which would simultaneously address the limitations noted above.”

      (2.3) Line 131: From the plots in Figure 2, it is not clear that the stimulus response is necessarily a rhythmic oscillation, in the sense of a single narrowband frequency.

      The reviewer is correct, and we are grateful for the prompt to be more precise. The Results phrasing around Fig. 2 has been softened to avoid any implication of narrowband rhythmicity, and now reads:

      “Under this protocol, we immediately observe macroscopic traveling waves emerge across the cortex (Figure 2 and Videos). The global mean voltage and the region-sorted raster (Figure 2c, d) reveal oscillatory activity that is well synchronized across regions, while the local-mean intracellular voltage maps over a representative 50ms window (Figure 2b) reveal a coherent wavefront sweeping along the anterior-posterior axis, consistent with previously reported cortex-wide waves [8, 9, 10]. The corresponding single-neuron-resolution view of the same simulation, with no spatial averaging, is shown in figure Supplement 1.”

      Moreover, we have revised the Introduction to describe [11] as demonstrating traveling waves in broadband (5-40Hz) activity, making clear that traveling waves can occur without requiring narrowband oscillations (see also our response to your related point below). In the revised manuscript, we also separate the broadband activity into canonical frequency bands and analyze the wave activity in each band independently.

      (2.4) Line 217: The authors should clarify how these findings relate to the results from Mohajerani et al. (Nature Neuroscience, 2013) or differ from them.

      We thank the reviewer for this suggestion. We agree that [12] is an important experimental benchmark. A direct quantitative comparison is difficult because the original data were not aligned to the Allen Brain Atlas CCF used in our simulations. We therefore revised the Discussion to identify this comparison as a future direction, alongside the visual-cortex bidirectional-wave data of [13]:

      “A more detailed quantitative comparison with experimental cortical-wave studies, such as the cortex-wide voltage-imaging data of [12] or the bidirectional visual-evoked waves reported by [13], is left as a future direction.”

      (2.5) Line 230: Because higher temporal frequency activity also tends to be more spatially localized, a correlation between PGD and temporal frequency could be an inherent consequence of this relationship, rather than a meaningful result.

      We thank the reviewer for raising this important point. The reviewer is correct that higher-frequency oscillations tend to be more spatially localized, which can inherently reduce PGD when measured globally. We therefore revised this analysis by separating the broadband signals into canonical frequency bands and comparing PGD within each band.

      In the revised manuscript, we no longer interpret cross-band PGD differences as evidence for a frequency-to-spatial-scale relationship. Instead, we report the level of macroscopic wave activity within each canonical band. This per-band comparison is summarized in Fig. 5b, which reports the maximum PGD in each band (mean ± SEM across the entire (g<sub>AMPA</sub>,V<sub>stim</sub>) grid) for the three connectivities. Allen connectivity shows higher per-band PGD than local or uniform connectivity in the theta, alpha, and beta bands, without requiring an interpretation of PGD differences across frequency bands.

      For completeness, we also computed the per-band mean(Allen) − mean(opponent) PGD gap across the full parameter grid (10 g<sub>AMPA</sub> × 12 V<sub>stim</sub> = 120 points per connectivity). The result is presented in Figure 4-figure supplement 6: panel (a) gives the per-band maxPGD heatmaps over the (gAMPA, Vstim) grid for each connectivity, and panel (b) gives the per-band Allen-minus-opponent mean-PGD gap. The mean gap against Local is −0.004 in delta, +0.027 in theta, +0.036 in alpha, +0.036 in beta and +0.004 in gamma, and against Uniform is +0.015, +0.044, +0.052, +0.040 and +0.011, respectively. The gap is largest in alpha and broadly concentrated in theta-alpha-beta, rather than in delta as we had originally written. We report these per-band gaps descriptively because they summarize one parameter sweep per connectivity rather than independent biological or simulation replications, and we do not interpret the differences across bands as a meaningful frequency dependence. We have added this analysis to the revised manuscript.

      (2.6)Line 247-248: It is not clear that the algorithm for generating connections between neurons presented here really relates to those for community detection. For example, in the case of the Allen Institute data, the communities are essentially in the data already.

      We agree with the reviewer that the relevant anatomical blocks are already present in the data. Our intent was to relate the sampling procedure to stochastic block models for network generation, not to community-detection algorithms. We have revised this passage in the Discussion subsection Construction of the Allen connectivity: ”In essence, our algorithm belongs to the family of stochastic block models [14], where the block structure is given by the voxelization of the Allen Brain Atlas and the inter-block connection probabilities are set by the voxelized projection strengths.”

      (2.7) Line 284-285: The relationship between conduction delay is more direct than this sentence suggests. Conduction delay is fundamentally determined by the time required for action potentials to propagate along axons, making it intrinsically linked to anatomical distance.

      Thank you for raising this important point. We agree that conduction delay is directly tied to axonal propagation time and therefore to anatomical path length. Our original wording was intended to note that the relevant path length cannot be approximated reliably by Euclidean distance in the 3-D coordinate space. We have revised this passage in the Discussion sub-section Cortical model and simulation to state that conduction delay is linked to the white-matter path length of the connection, while noting that our model lacks the actual axonal path geometry through the cortical manifold:

      “However, we did not include conduction delay in our study. Conduction delay is thought to have a proportional relationship with the white matter path length of the connection between two neurons [15]. In our model, although we have the 3D positions of neurons, we do not have geometric information about the cortical manifold. Two neurons can be very close in the Cartesian coordinates measured by the Euclidean distance, but very far in terms of the length of the actual connection in the brain. Therefore, incorporating accurate conduction delay in the model is an important future direction.”

      (2.8) Lines 294-295: Several methods do exist for detecting and characterizing wave dynamics in three-dimensional data (Budzinski et al., Physical Review Research, 2023).

      Thank you for this reference. We have added a citation to [16] in the Discussion subsection Quantitative measurements of 3-D traveling waves, acknowledging that methods for 3-D wave analysis do exist while noting that most published algorithms are designed for 2-D data:

      “There are many techniques available for identifying and measuring large-scale neuronal spatiotemporal patterns [17, 18]. While methods for detecting wave dynamics in three-dimensional data do exist [16], most published algorithms are designed to analyze 2-D data, so our simulation data, which is intrinsically 3-D, presents new challenges for measurement.”

      (2.9) Line 28: It is important to note that the Davis et al. (2020) reference is not actually in the beta band, but instead in the broadband (5-40 Hz). This distinction is important because it demonstrates that waves can occur in neural data without requiring narrowband oscillations.

      Thank you for this correction. We have moved the [11] citation out of the beta-band group in the Introduction and reframed it as a broadband (5-40Hz) reference, so that the sentence now reads:

      “These waves are observed at different frequencies during various brain activities, ranging from slow-wave activity [8, 9], sleep spindles [19], to faster oscillations in alpha [20, 21], beta [22, 23], and gamma [21, 13] frequency bands, as well as in broadband (5-40Hz) activity [11].”

      This distinction is important because it makes clear that traveling waves do not require narrowband oscillations. The revised manuscript therefore separates the broadband activity into canonical frequency bands and compares wave activity within each band.

      (2.10) Line 46-49: This sentence could be clearer, for example, by specifying “certain dynamics” in more precise terms.

      We apologize for the imprecision in the original manuscript. We have clarified the sentence in the Introduction to specify that the dynamics of interest are coexistence patterns of local and global activity in the network, as described in the cited reference [24]:

      “It has been shown that in a coupled neuronal network, the coexistence of global wave activity with locally asynchronous states only arises when the number of oscillators is high enough, where local and global activities can coexist [24].”

      (2.11) Figure 1b(i): Small typo in the label for this panel.

      This has been corrected. Thank you for catching it.

      (2.12) Lines 121-122: It may be important to note that spiking neural networks can also generate self-sustained activity (Vogels and Abbott, JNeurosci, 2005; Kumar et al., Neural Computation, 2008). This self-sustained activity is a form of internally generated “frozen” noise that is fundamentally different from externally imposed noise sources (such as Poisson external input) (Destexhe and Contreras, Science, 2006). Waves appear in this self-sustained activity, as well (Davis et al., Nature Communications, 2021), supporting the generality of this phenomenon.

      Thank you for pointing out these important references. We have added citations to [25], [26], [27], and [24] in the Results subsection Macroscopic traveling waves emerge from random stimulation through realistic connectivity:

      “Spiking networks of this scale can also generate self-sustained activity in similar regimes [25, 26], which differs fundamentally from externally imposed noise [27], and traveling waves have been reported under such conditions [24].”

      In the revised manuscript, we also changed the stimulation protocol from constant current to Poisson input, which is more similar to the stochastic input that cortical neurons receive in vivo. Traveling waves observed in our model under this Poisson drive therefore complement, rather than depend on, the self-sustained-activity regime emphasized by the cited works.

      (2.13) Lines 139-140: “anterior-posterior macroscopic waves in both directions” and ”in the reverse direction right after each other” could be clearer. In addition, the study from Aggarwal et al. (Nature Communications, 2022) could be relevant to note at this point.

      We thank the reviewer for these wording suggestions and the relevant reference. In the revised manuscript, we reran the simulations and updated Figure 2 accordingly. The new representative simulation shown in Fig. 2 emphasizes a coherent wavefront sweeping along the anterior-to-posterior axis, and the original passages describing consecutive opposite-direction waves have been removed from the Results subsection Macroscopic traveling waves emerge from random stimulation through realistic connectivity, which now reads:

      “Under this protocol, we immediately observe macroscopic traveling waves emerge across the cortex (Figure 2 and Videos). The global mean voltage and the region-sorted raster (Figure 2c, d) reveal oscillatory activity that is well synchronized across regions, while the local-mean intracellular voltage maps over a representative 50ms window (Figure 2b) reveal a coherent wavefront sweeping along the anterior-posterior axis, consistent with previously reported cortex-wide waves [8, 9, 10]. The corresponding single-neuron-resolution view of the same simulation, with no spatial averaging, is shown in figure Supplement 1.”

      Bidirectional propagation can still occur in the model, but we have chosen not to present it as a focal result of this revised manuscript; consequently, the specific phrasings flagged by the reviewer no longer appear in the Results text.

      We agree that [13] is an important reference, and we now discuss it in the Comparing with experimental data subsection as a potential benchmark for future quantitative comparison with our model.

      (2.14) Figure 7: The caption for this figure could be clearer.

      Line 281: typo ”Ermentrou”.

      Line 418: typo ”excitatory”.

      The typos (“Ermentrou” → “Ermentrout”, and the “excitatory” typo at line 418) have been corrected. The original Figure 7 has been removed from the revised manuscript; its content (PGD versus stimulus magnitude and coupling strength, and PGD by dominant frequency bucket) has been reorganized across the new Figs. 4-6.

      References

      (1) Steriade M, Mccormick DA, Sejnowski TJ. Thalamocortical Oscillations in the Sleeping and Aroused Brain. Science. 1993;262(5134):679-85. Available from: <GotoISI>:// WOS:A1993MD95200029.

      (2) Ye Z, Bull MS, Li A, Birman D, Daigle TL, Tasic B, et al. Brain-wide topographic coordination of traveling spiral waves. BioRxiv. 2023:2023-12.

      (3) Rockland KS.What do we know about laminar connectivity? Neuroimage. 2019;197:772-84.

      (4) Harris JA, Mihalas S, Hirokawa KE, Whitesell JD, Choi H, Bernard A, et al. Hierarchical organization of cortical and thalamic connectivity. Nature. 2019;575(7781):195+. Available from: <GotoISI>://WOS:000496159900061https://www.nature.com/articles/ s41586-019-1716-z.pdf.

      (5) Buzs´aki G, Mizuseki K. The log-dynamic brain: how skewed distributions affect network operations. Nature Reviews Neuroscience. 2014;15(4):264-78.

      (6) Bazhenov M, Rulkov NF, Timofeev I. Effect of synaptic connectivity on long-range synchronization of fast cortical oscillations. Journal of neurophysiology. 2008;100(3):156275.

      (7) Morrison A, Aertsen A, Diesmann M. Spike-timing-dependent plasticity in balanced random networks. Neural computation. 2007;19(6):1437-67.

      (8) Massimini M. The Sleep Slow Oscillation as a Traveling Wave. Journal of Neuroscience. 2004;24(31):6862-70. Available from: https://dx.doi.org/10.1523/jneurosci. 1318-04.2004 https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6729597/pdf/ 0246862.pdf.

      (9) Liang Y, Song C, Liu M, Gong P, Zhou C, Kno¨pfel T.Cortex-Wide Dynamics ofIntrinsic Electrical Activities: Propagating Waves and Their Interactions. The Journal of Neuroscience. 2021;41(16):3665-78. Available from: https://www.jneurosci.org/ content/jneuro/41/16/3665.full.pdf.

      (10) Aggarwal A, Luo J, Chung H, Contreras D, Kelz MB, Proekt A. Neural assemblies coordinated by cortical waves are associated with waking and hallucinatory brain states. Cell Reports. 2024;43(4):114017. Available from: https://www.sciencedirect.com/ science/article/pii/S2211124724003450.

      (11) Davis ZW, Muller L, Martinez-Trujillo J, Sejnowski T, Reynolds JH. Spontaneous travelling cortical waves gate perception in behaving primates. Nature. 2020;587(7834):4326. Available from: https://doi.org/10.1038/s41586-020-2802-y.

      (12) Mohajerani MH, Chan AW, Mohsenvand M, LeDue J, Liu R, McVea DA, et al. Spontaneous cortical activity alternates between motifs defined by regional axonal projections. Nature Neuroscience. 2013;16(10):1426-35. Available from: https://doi.org/ 10.1038/nn.3499https://www.nature.com/articles/nn.3499.pdf.

      (13) Aggarwal A, Brennan C, Luo J, Chung H, Contreras D, Kelz MB, et al. Visual evoked feedforward–feedback traveling waves organize neural activity across the cortical hierarchy in mice. Nature Communications. 2022;13(1):4754. Available from: https://doi.org/10.1038/s41467-022-32378-xhttps://www.nature. com/articles/s41467-022-32378-x.pdf.

      (14) Holland PW, Laskey KB, Leinhardt S. Stochastic blockmodels: First steps. Social Networks. 1983;5(2):109-37. Available from: https://www.sciencedirect.com/science/ article/pii/0378873383900217.

      (15) Lemar´echal JD, Jedynak M, Trebaul L, Boyer A, Tadel F, Bhattacharjee M, et al. A brain atlas of axonal and synaptic delays based on modelling of cortico-cortical evoked potentials. Brain. 2022;145(5):1653-67.

      (16) Budzinski RC, Nguyen TT, Min´o-Calero J, Davis ZW, Muller LE. Analyzing transientevoked neural activity in three-dimensional cortical recordings. Physical Review Research. 2023;5(1):013012.

      (17) Townsend RG, Gong P. Detection and analysis of spatiotemporal patterns in brain activity. PLOS Computational Biology. 2018;14(12):e1006643. Available from: https: //doi.org/10.1371/journal.pcbi.1006643.

      (18) Gutzen R, De Bonis G, De Luca C, Pastorelli E, Capone C, Allegra Mascaro AL, et al. A modular and adaptable analysis pipeline to compare slow cerebral rhythms across heterogeneous datasets. Cell Reports Methods. 2024;4(1). Available from: https: //doi.org/10.1016/j.crmeth.2023.100681.

      (19) Muller L, Piantoni G, Koller D, Cash SS, Halgren E, Sejnowski TJ. Rotating waves during human sleep spindles organize global patterns of activity that repeat precisely through the night. Elife. 2016;5. Available from: https://www.ncbi.nlm.nih.gov/pubmed/27855061https://www.ncbi.nlm. nih.gov/pmc/articles/PMC5114016/pdf/elife-17267.pdf.

      (20) Zhang H, Watrous AJ, Patel A, Jacobs J. Theta and alpha oscillations are traveling waves in the human neocortex. Neuron. 2018;98(6):1269-81. e4.

      (21) van Kerkoerle T, Self MW, Dagnino B, Gariel-Mathis MA, Poort J, van der Togt C, et al. Alpha and gamma oscillations characterize feedback and feedforward processing in monkey visual cortex. Proc Natl Acad Sci U S A. 2014;111(40):14332-41. Available from: https://www.ncbi.nlm.nih.gov/pubmed/25205811.

      (22) Bhattacharya S, Brincat SL, Lundqvist M, Miller EK. Traveling waves in the prefrontal cortex during working memory. PLoS Comput Biol. 2022;18(1):e1009827.

      (23) Rubino D, Robbins KA, Hatsopoulos NG. Propagating waves mediate information transfer in the motor cortex. Nature Neuroscience. 2006;9(12):1549-57. Available from: https://doi.org/10.1038/nn1802.

      (24) Davis ZW, Benigno GB, Fletterman C, Desbordes T, Steward C, Sejnowski TJ, et al. Spontaneous travelling waves naturally emerge from horizontal fiber time delays and travel through locally asynchronous-irregular states. Nature Communications. 2021;12(1):6057.

      (25) Vogels TP, Abbott LF. Signal propagation and logic gating in networks of integrate-and-fire neurons. Journal of Neuroscience. 2005;25(46):10786-95.

      (26) Kumar A, Schrader S, Aertsen A, Rotter S. The high-conductance state of cortical networks. Neural Computation. 2008;20(1):1-43.

      (27) Destexhe A, Contreras D. Neuronal computations with stochastic network states. Science. 2006;314(5796):85-90.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines Müller glia (MG) reprogramming in the uninjured mouse retina through a combination of Notch signaling inhibition and AAV-induced proliferation. Building on their prior work showing that Cyclin D1 overexpression and p27^Kip1^ knockdown (CCA) promotes MG proliferation with very limited neurogenesis, the authors now demonstrate that Rbpj deletion alone induces a modest degree of MG-to-neuron conversion without proliferation, in agreement with recent work in the field. However, combining Rbpj deletion with CCA-mediated proliferation substantially enhances MG dedifferentiation and the generation of retinal neuron-like cells. Through genetic lineage tracing, histological analyses, and single-cell transcriptomics, the authors provide evidence that MG-derived cells acquire molecular features of bipolar (ON, OFF, and rod bipolar) and amacrine neurons. Most MG-derived cells appear to survive long-term (up to 9 months).

      Strengths:

      Overall, the study is carefully designed and executed, and the manuscript is clearly written with well-presented figures. While the work does not significantly expand the repertoire of neuronal types generated from mammalian MG beyond what has been previously reported in the field, it provides a valuable and improved strategy for inducing robust MG proliferation and neurogenesis in the mammalian retina.

      Weaknesses:

      (1) It would be better to include a negative control AAV when evaluating the effect of CCA AAV in the Rbpj KO background. This could help distinguish the specific contribution of the CCA construct from potential effects of intravitreal AAV injection itself, which can induce mild inflammation, known to influence MG reprogramming.

      To address this concern, in the revised manuscript we included the result from Rbpj KO eyes injected with a negative control AAV (AAV<sub>7m8</sub>-GFAP-GFP) (Fig. S14a). MG reprogramming efficiency, quantified as the proportion of tdT<sup>+</sup>Otx2<sup>+</sup> cells among total tdT<sup>+</sup> cells, was then compared between the AAV-GFP–treated and Rbpj KO–only eyes. At 4 months post-injection, the percentage of tdT<sup>+</sup>Otx2<sup>+</sup> cells in the AAV-GFP–treated eyes was comparable to that of Rbpj KO alone (Fig. S14b–c), and substantially lower than in the CCA-treated eyes. Together, these results indicate that the enhanced MG reprogramming observed in the Rbpj KO+CCA group is driven by transgenes expressed rather than by nonspecific effects of AAV or injection.

      (2) The extent of MG transduction by the CCA AAV is not clear. As quantifications are normalized to total MG (GFP^+^ or TdTomato^+^) or retinal length, it would be useful to clarify whether near-complete transduction is assumed, or if additional information on transduction efficiency can be provided.

      In our previous study (Wu, Liao, et al., 2025, eLife), we have demonstrated that high-dose (4E10vg/injection) AAV7m8 effectively transduced the whole retina, with near-complete MG transduction observed in the vicinity of the injection site, as evidenced by virtually all MG expressing GFP in these regions. In the revised manuscript, we clarified the transduction efficiency in Line 108-110 on Page 5 and Line 625-626 on Page 27.

      (3) In Figure S10, the reduced MG proliferation observed in the CCA + Rbpj deletion group could also potentially reflect decreased GFAP promoter activity in dedifferentiated MG following Rbpj deletion. Alternatively, MG-derived cells may be more fragile under these conditions.

      We thank the reviewer for these excellent insights. We agree that a down-regulation of GFAP promoter activity following Rbpj-mediated dedifferentiation is a highly plausible explanation for the moderate reduction in proliferation, as lower promoter activity would diminish AAV transgene expression. We have included this possibility in the data interpretation (Line 176-178, page 8). Regarding the alternative possibility of increased cell fragility, we agree that cell death cannot be ruled out, but occasional apoptotic cells over a long period of time are difficult to capture experimentally.

      (4) In the CCA + Rbpj deletion condition, do MG undergo single or multiple rounds of cell division?

      We have previously demonstrated that MG typically undergo a single round of cell division in wild type mouse retina following CCA treatment (Wu, Liao, et al., 2025, eLife). Given our observation that Rbpj deletion suppresses CCA-induced MG proliferation (Fig. S11), it is unlikely that the addition of Rbpj deletion would trigger multiple or continuous rounds of cell division beyond the single-round baseline established by CCA alone. While we did not re-evaluate cell division kinetics in the current study, we reason that CCA similarly drives MG to undergo a single round of division in the Rbpj KO context.

      (5) What fraction of neuron-like cells (bipolar- and amacrine-like) arises from proliferation versus direct transdifferentiation? Quantification of MG-derived cells expressing neuronal markers (e.g., Otx2, HuC/D), with and without EdU labeling, would help distinguish these mechanisms.

      The percentages of MG-derived cells expressing neuronal markers with and without EdU labeling, were shown in Fig 3d-e and Fig S19d-e. In the Rbpj KO-only group, neuron-like cells arise exclusively through direct transdifferentiation without cell division, as no EdU incorporation was detected in Rbpj-deficient MG. In this group, a small fraction of MG-derived cells expressed the neuronal marker Otx2 or HuC/D (Fig 3e, Fig S19e). In contrast, the Rbpj KO+CCA group achieved a substantially higher neurogenesis rate, with a significant proportion of Otx2+ or HuC/D+ MG-derived cells also being EdU+ (Fig. 3d, Fig. S19d), indicating that they arose through de novo neurogenesis. By subtracting the contribution of direct transdifferentiation observed in the Rbpj KO-only group, we estimate that majority of MG-derived neuron-like cells in the Rbpj KO+CCA group were generated through proliferation-mediated de novo neurogenesis.

      (6) In Figure S18a, the authors state that "while the neuron-like clusters were best classified as BC-like and AC-like based on their distinct marker gene expression, they also exhibited mixed expression of genes associated with other retinal neuronal types, including RGC markers (e.g., Tubb3, Myt1l, Grin1) and photoreceptor markers (e.g., Crx, Prom1, Epha10, Gucy2e, Scg3) (Fig. S18a), suggesting that the regenerated cells exist in a hybrid state" and "MG derived neuron like cells also expressed genes characteristic of RGCs and photoreceptors, indicating enhanced lineage". However, many of these genes are not specific to RGCs or photoreceptors and are instead broadly expressed in retinal neurons or enriched in bipolar/amacrine populations. Therefore, it is unclear whether these cells exhibit hybrid RGC or photoreceptor identity.

      We thank the reviewer for this insightful comment and for pointing out the need for greater precision in our terminology regarding these markers. While individual markers may lack absolute, 100% cell-type exclusivity, genes such as Tubb3 and Gucy2e serve as widely accepted lineage-associated genes that characterize RGC and photoreceptor programs, respectively (Soto et al., 2008; Sato et al., 2018; Sotani et al., 2024). We have revised the manuscript to replace terms "RGC-specific genes" and "photoreceptor-specific genes" with "RGC signature genes" and "photoreceptor signature genes", respectively. Furthermore, these RGC- and photoreceptor-signature genes are co-expressed across the entire Otx2+ MG population rather than being segregated into distinct, specialized subpopulations (Fig. 4d, Fig. S20). This uniform distribution indicates that these cells possess a hybrid transcriptional program that concurrently incorporates elements of both RGC and photoreceptor identities.

      (7) The authors provide a thorough molecular characterization of MG-derived cells through immunostaining and single-cell sequencing. However, their morphological features, synaptic connectivity (e.g., synaptic marker expression), and electrophysiological properties remain largely uncharacterized. While these experiments may be technically challenging, this limitation should be discussed.

      We agree with the reviewer that characterizing the precise morphological features, synaptic connectivity, and electrophysiological properties of MG-derived cells is a crucial step for any neuronal regeneration study, and we acknowledge that this represents an important limitation of our current study.

      As demonstrated by snRNA-seq data, the MG-derived neuron-like cells exhibit an incompletely mature state, characterized by hybrid transcriptomic signatures. By immunostaining, we did not observe any MG-derived cells with photoreceptor outer segment or typical RGC morphology. Therefore, it is highly likely that these cells have not established functional synaptic connectivity or acquired mature electrophysiological properties. Performing functional or circuitry assessments at this stage would be premature.

      We have added a comprehensive discussion regarding this limitation, along with future directions for long-term functional validation, in the revised manuscript (Line 502-515 on Page 22).

      (8) The conclusion that CCA + Rbpj deletion induces neurogenesis without compromising MG supportive functions or retinal homeostasis appears somewhat oversold. This claim is primarily based on gross retinal morphology and ZO-1 staining. Given the extent of MG dedifferentiation and ectopic cell generation in the ONL and INL, it is likely that retinal function is affected. Functional assessments (e.g., ERG) would be required to support this conclusion. The authors should consider tempering this statement.

      To address the concern raised by the reviewer, we performed electroretinography (ERG) to evaluate both scotopic and photopic retinal function in the Rbpj KO+CCA-treated eyes compared to contralateral untreated controls (Supplementary figure S23e-h). In addition, we conducted optomotor response testing to assess whether visual behavior is affected following treatment (Supplementary Figure S23d). The results demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal function.

      (9) Regarding the mechanism by which CCA-induced proliferation enhances MG reprogramming in the Rbpj knockout background, one plausible explanation is that chromatin states (e.g., histone modifications and DNA methylation) are transiently reset during DNA replication and cell division. While this alone may be insufficient to activate neurogenic programs, it could synergize with Rbpj deletion to allow neurogenic transcription factors (such as Ascl1, Otx2, NeuroD1, and NeuroD2) to access previously inaccessible chromatin regions, thereby promoting MG reprogramming.

      We thank the reviewer for the insightful suggestion on the model, which aligns well with our experimental findings. Our snATAC-seq data demonstrate that CCA-induced proliferation broadly increases chromatin accessibility at key neurogenic loci, including Neurod2, Dll1, and Otx2, in active MG compared to resting MG (Figure 6f–h). This chromatin remodeling alone is insufficient to drive neurogenesis, as CCA-only treated MG largely revert to a quiescent glial state. However, when combined with Rbpj deletion, which derepresses downstream neurogenic transcription factors such as Ascl1 and Neurog2 by relieving Notch-mediated transcriptional repression, these newly accessible chromatin regions can be effectively occupied and activated by the available neurogenic factors. The concept that cell division facilitates epigenetic resetting to enhance reprogramming efficiency is well established in the somatic cell reprogramming field, where proliferation rate is directly proportional to reprogramming success by promoting the erasure of lineage-restrictive epigenetic marks and the re-establishment of new transcriptional circuits. In the revised manuscript, we incorporated this mechanistic discussion to provide a more comprehensive interpretation of how proliferation and Notch inhibition converge to promote MG neurogenesis in Line 462-477 on Page 20-21.

      Reviewer #2 (Public review):

      Summary:

      The inability of the mammalian retina to regenerate poses a major clinical challenge. Much has been learned about the regenerative potential of the retina from teleost fish, where Müller glia (MG) are able to proliferate and produce new neurons after injury. However, MG do not retain this potential in the mammalian retina. The authors showed previously that forcing MG to re-enter the cell cycle by downregulating p27 and upregulating cyclin D1 could induce MG to dedifferentiate, but the results were transient, and these cells eventually reverted back to MG and did not form neurons. Here, they expand on this to show that in MG, coupling forced cell cycle re-entry with deletion of Rbpj, which inhibits the transcriptional effects of Notch signaling, induces some MG to proliferate and take on features of multiple cell types, including MG precursor cells, amacrine-like cells, and bipolar-like cells. This work lends valuable insight into the regenerative potential of mammalian MG, particularly when Notch signaling is manipulated.

      Strengths:

      The major claims of the authors are well-supported. They show convincingly - and through multiple methods including immunostaining, single-nucleus RNA sequencing, and in situ hybridization - that coupling notch inhibition with cell cycle reactivation induces the expression of neuronal markers in mammalian MG. The snRNA-seq data are particularly valuable in demonstrating the induction of bipolar-cell subtypes. Edu labeling is effective in demonstrating the induction of proliferation, and the long-term viability of the generated neuron-like cells is intriguing.

      Weaknesses:

      Whether the newly generated neurons are functionally integrated remains unclear, and the effect of the manipulation on the function of the retina was not tested. Imaging data suggests that many of the newly generated neurons persist for months, but often appear mislocalized. It is also not clear if the manipulation of MG affects long-term MG function. Cell death was not evaluated, and although the authors evaluated the long-term effect on tight junctions, this data was not quantified, and further analysis on morphology or function was not done. Control eyes were untreated, not vehicle-injected.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The transgenic line may be Glast-CreERT, not Glast-CreERT2.

      We appreciate the reviewer for bringing this to our attention. The formal allele symbol for this transgenic line is Tg(Slc1a3-cre/ERT)1Nat, while this strain is generically classified as "Cre/ERT2" by the Jackson Laboratory. Some published studies referred to this line as Glast-CreERT and others as Glast-CreERT2. To maintain consistency with the formal allele symbol, we have adopted "Glast-CreERT" throughout the revised manuscript.

      (2) For snATAC data in Figure 6 e,f, and Figure 19b. It is most likely gene activity, not gene expression, since these are snATAC, not snRNA data.

      For this inaccurate terminology, we have corrected all relevant figure labels and associated text in the revised manuscript to clearly state "gene activity" instead of "gene expression."

      (3) Some text in Figure 6 is a bit too small to read.

      We have increased the font size of the text elements in Figure 6 to ensure readability and have also reviewed all other figures for consistency. Revised figures with improved legibility have been included in the updated manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) There are multiple instances where further elaboration of methods or tools in the test would improve readability and comprehension by a broader audience. It would be helpful to (early, often, and clearly) explain precisely which cell types are labeled in your mouse line and how. Someone unfamiliar with the mouse line may struggle to understand what is labeled by the tdT or GFP. Likewise, it would help to consistently define what cell types are labeled by tdT+ vs. Sox9+, tdT+, etc.

      We have added a clear and detailed description of the mouse lines and labeling strategy early in the Results section, specifying which cell types are labeled by tdT and GFP and how the labeling is achieved. We have also ensured that the definitions of cell type identifiers (e.g., tdT<sup>+</sup> for MG-derived cells, Sox9<sup>+</sup>/tdT<sup>+</sup> for MG remaining in a glial state) are consistently stated upon first use and maintained throughout the manuscript to improve readability for a broader audience. In addition, we added headings for the quantification graphs to improve readability in all quantification figures.

      (2) It is unclear what the difference is between Figure 1c and S1c, and these should be quantified as the % of positive cells, as described in the text.

      We have removed Figure S1c and moved Figure 1c to supplementary figure 1. The MG labeled by EdU and Sox9 or Otx2 were quantified as % of the EdU+ MG.

      (3) S2e: Clarify what pixel level means, is this pixel intensity?

      Yes, "pixel level" in Figure S2e refers to pixel intensity. We apologize for the ambiguous wording and replaced "pixel level" with "pixel intensity" in the revised figure legend to ensure clarity.

      (4) Figure 2: In the magnified image of the GFP+, Sox9- cell, the GFP is also very faint. Could these cells be dying? Analysis of the expression profile of these cells (or ruling out apoptosis) would better support a dedifferentiation argument.

      The faint GFP signal observed in GFP<sup>+</sup> Sox9<sup>-</sup> cells is a sign of ongoing dedifferentiation rather than cell death. This is likely due to chromatin remodeling during reprogramming. A similar decrease in reporter signal intensity during MG dedifferentiation has been previously reported by Le et al. 2024, 2025, supporting the interpretation that reduced fluorescence is a characteristic feature of this process. It is possible that a small fraction of GFP<sup>+</sup> Sox9<sup>-</sup> cells may undergo cell death over an extended period, which would be difficult to detect using apoptosis assays. Our long-term survival experiments demonstrate that more than 80% of MG-derived neuron-like cells survive for at least 9 months following treatment (Figure 7), indicating that majority of these cells are viable. The discussion is included in line 112-114 on page 5.

      (5) Figure 3: The Crx labeling appears everywhere except the identified cell. This seems the opposite of the point you are making.

      Crx signal of the MG-derived cell (tdT<sup>+</sup> Crx<sup>+</sup>), which is pointed out by arrowhead, is in a ring-like pattern. This pattern is consistent with the euchromatin region in inverted nucleus of rod. Crx labeling appears in other cells in the image as Crx is highly expressed in native photoreceptors.

      (6) I think it would be nice to address, in the discussion, the apparent disorganization and mislocalization of cells in the long-term images.

      We thank the reviewer for highlighting this critical observation. During retinal development, precise laminar positioning of neurons is guided by a coordinated interplay of cell-intrinsic transcriptional programs and extrinsic cues including cell adhesion molecules, guidance factors, and interactions with neighboring cells. In the adult retina, many of these developmental cues are no longer present or active, which likely contributes to the failure of MG-derived neurons to migrate to their appropriate laminar positions. Interestingly, the vast majority of our divided MG cells remained localized within the outer nuclear layer (ONL). Because the ONL is the physiological location of photoreceptors, this preferential position could serve as an advantageous baseline layout for driving targeted photoreceptor differentiation in future work. To address reviewer’s feedback, we have expanded our discussion section (Line 543-559, page 23-34) to cover the mechanisms underlying this structural disorganization and its downstream implications for functional circuit integration.

      (7) I'm not convinced that ZO1 alone is sufficient to suggest MG function normally or that retinal homeostasis is maintained. I suggest tempering that conclusion in the text.

      For the revision, we have performed additional experiments to address this concern. The optical coherence tomography (OCT) images revealed that retinal layer organization and ONL thickness were comparable among the uninjected eyes, GFP AAV-injected control eyes, and CCA-treated eyes, demonstrating that overall retinal architecture was well-preserved (Fig. S23a–c). Optomotor response testing revealed no significant differences in visual acuity across groups, suggesting that visual function remained intact (Fig. S23d). Furthermore, electroretinography (ERG) demonstrated that scotopic and photopic a- and b-wave amplitudes were unaffected by the treatment, confirming that light responses from photoreceptor and inner retinal neuron were preserved (Fig. S23e–h). Taken together, these findings demonstrate that combined Rbpj KO and CCA treatment achieves neurogenesis without compromising retinal structure and functional visual circuitry.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Thank you for the helpful comments and criticisms. We provide exciting additional data, in particular a CRISPR actin-binding motif mutant and FRAP analysis of an exon15e-GFP transgene, both further supporting the importance of the IDR in thin filament stability. We believe that these additional experiments provide compelling evidence supporting our conclusion and substantially advance the current limited body of knowledge surrounding the role of IDRs in structural proteins.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Ho and Schock investigates the role of the Z-disc protein Zasp52 during Drosophila flight muscle development. It was known before, mainly by findings from this group, that Zasp52 is required for normal sarcomere morphogenesis, specifically Z-disc morphogenesis in indirect flight muscles. But the exact molecular mechanism by which Zasp52 contributes, apart from the fact that it is localised there and is somehow involved in multimerization/cross-linking, was not clear. This paper proposes that an intrinsically disordered region (IDR) in Zasp52 is needed for some of its functions, by stabilising Zasp52 localisation at the Z-disc. Specifically, the IDR in Zasp52 is proposed to be required for Z-disc maintenance during the mechanical challenges of flight, while being dispensable for the initial morphogenesis during development. This hypothesis is supported by strong genetic evidence and behavioural tests, deleting Zasp's IDR impairs flight from mid-age onwards, while a block in flight activity lifts the phenotype.

      However, some of the phenotypic analysis, in particular the bending of the sarcomere, likely upon mechanical challenge by muscle contractions, needs more detailed investigations to be fully convincing.

      Strengths:

      (1) The linker in the alternatively spliced exon 15 of Zasp52 was deleted with a state-of-the-art genetic editing strategy. Surprisingly, flies are homozygous viable, showing that this long part of the Zasp52 protein is not essential for animal survival or sarcomere morphogenesis.

      (2) The observed sarcomere phenotypes with age, especially the bending Z-discs, are new and exciting.

      (3) The displayed EM images document interesting phenotypes.

      (4) Most of the observed phenotypes can be rescued by re-expression of the long Zasp52 isoform, which does contain the IDR region, but not by a shorter one without it, suggesting that IDR is important.

      (5) FRAP data measure the local turnover of a short-ZaspGFP and show that this increased in the Zasp mutant lacking the IDR domain, suggesting that Zasp-IDR might stabilise Zasp at the Z-disc.

      (6) Interestingly, flight and sarcomere morphology phenotypes can be rescued by preventing the flies from flying, suggesting that they are mechanically induced.

      Weaknesses:

      (1) The western blot quantifications of Zasp isoform expression are weak. No error bars are indicated in the quantifications; the quantifications appear to be more qualitative than quantitative. According to band intensities, the long Zasp isoforms seem to be less present compared to the shorter ones, even in the flight muscles.

      We have now included quantifications with error bars for the Western blots in our resubmission. It is important to keep in mind that the main point in figure 1B is that there are plenty of exon15e-containing isoforms in IFM, in contrast to other tissues with very limited exon15e-containing isoforms. This is confirmed by the analysis of RNA-seq data in figure 1C, and of course, by the flightless phenotype of the exon15e mutant.

      (2) The phenotypic analysis of the sarcomere appears somewhat superficial throughout the paper. Only Zasp52 and phalloidin are shown; no other Z-disc or thick filament proteins. At least myosin stainings and overview images are important to better judge the phenotypic variations. Are the variants between individuals or regional in the same muscle?

      Our images are representative of the observed phenotypes. Phenotypes are consistently present across all individuals, as reflected in our replicates. Interestingly, they appear not to be randomly interspersed among the sarcomeres but concentrated in certain regions of muscle more than others. Full images are available in the online repository FigShare.

      (3) EM images would benefit from better quantification.

      We do not believe that EM images can be meaningfully quantified, because of the many selection steps preceding image acquisition.

      (4) Other proteins were not analysed with the FRAP-based turnover assay for comparison in wild type and mutant. All Z-proteins might turn over faster in the mutant with the defective Z-disc.

      This is the point we are trying to make. The Zasp52 IDR appears to stabilize the Z-disc and is likely involved in fastening a variety of proteins to it.

      Reviewer #2 (Public review):

      Summary and Strengths:

      This in-depth genetic analysis of Zasp52 function in Drosophila indirect flight muscle (IFM) provides an interesting perspective regarding the role of a partially disordered region (IDR) in exon 15e. This exon seems to be exclusively present in IFM and contributes to the prevention of myofibril disintegration during aging, likely due to interactions of this region with Z-disc insertion and/or stability. The addition of an isoform (PR) that lacks exon 15e serves as a nice control to illustrate the necessity of exon 15e in muscle structure and function. Overall, the manuscript is exceptionally well-written, logical, with nicely controlled experiments and detailed statistical analysis that largely support the conclusions drawn by the authors. While exon 15e is clearly involved in preventing muscle degeneration, a solid role for thin filament stability is not clearly shown (as mentioned in the abstract). In addition, which regions/how the proteins of the IDR may contribute are unclear.

      Weaknesses:

      (1) It is not clear in Figure S1A where exon 15e fits within the Zasp52 locus schematic. This is important as a premise of this paper describes this region to be key, and proof from multiple prediction programs would lend more weight to the prediction of the exon being largely disordered. Inclusion of the discussed short linear motifs, comparison with Canoe or LBD3 for similarities and/or an Alphafold structure would help make the authors' point (colorized with known domains).

      We added a bar below figure S2A to show the region corresponding to exon 15e. We used three disorder prediction programs and one structure (order) prediction program. The majority of exon15e is completely disordered and of very low confidence score, and thus uninformative to display as an AlphaFold structure. Likewise, IDR’s are very difficult to classify, therefore we cannot say much more than that LDB3, Zasp52, and Canoe contain IDRs, with Zasp52 and Canoe both having a putative actin-binding domain within the IDR. We now provide data on the function of the ABD in this resubmission.

      (2) Interesting that immobilization rescues the deterioration phenotypes. The authors should explain in more detail how this was done to avoid dehydration/starvation of the flies.

      We provided more details in materials and methods.

      (3) There is a lot of discussion about the potential function of the IDR region, specifically a putative actin binding motif or other 'ordered' regions that may contain short linear motifs. It would strengthen the findings to show which of these may be essential for Zasp52 function in the IFM. The ability to bind actin could be tested biochemically, and/or smaller deletions could be made to unequivocally test the role of the ABD vs other predicted motifs using genetics. If some of these regions are more ordered, where do they lie within, and do they form a predicted fold or structure that gives insight into function?

      We now provide data on the function of the ABD showing that deleting it has almost no phenotypic defects. That means the IDR is largely/entirely responsible for the observed phenotypes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Western blot in Figure 1B needs proper quantification. A ratio between long and short isoforms in the same muscle type might be informative. Is it known which epitope the antibody recognises? Can a GFP insertion that also labels all isoforms be used as verification? Quantifications are also needed in Figure 2A.

      We have added quantifications of all Western blots (Fig. 1B, 2A, and 2A’). The ratio between exon 15e-containing and total Zasp52 in the same muscle type is included in the lowest bar graph in Fig. 1B. The full-length antibody is polyclonal and was raised against Zasp52-PR which contains all ordered domains; the anti-LIM antibody was raised against the last three LIM domains (both are described or referenced in the materials and methods section). Such a GFP insertion cannot exist due to the complex splicing patterns of Zasp52.

      (2) The name of the deletion allele could be specifically indicated in Figure 1A below the red bar.

      Done.

      (3) It would be useful to indicate the order group names in Figure S2B since species names are hard to read.

      For the version of record we provided high-resolution images, where species names can be read. Drosophilids, Ephemeroptera and Odonata are indicated.

      (4) The inverted spelling of the numbers for the control in Figures 2C and 5H is strange.

      Changed to normal spelling.

      (5) The bending of the myofibril at the Z-disc is a really interesting phenotype. However, it seems it is not always visible; at least it is visible in many myofibrils shown in Figure 3B, but in none in Figure 3E, same genotype, just different staining. Hence, I wonder if this bending could be force-induced by the cutting of the thorax during tissue preparation. It would be useful to display some overview images to allow the reader to judge the quality of the tissue preparation, indicating from where the high magnification view shown was taken. The same is true for Figure 5.

      Overview images are available on FigShare. Note that you can see some “H-zone actin” sarcomeres in Fig. 3B, as well as some mildly bent ones in Fig. 3E. We generally selected images that best demonstrated the phenotype described. Furthermore, neither phenotype is fully penetrant so we cannot expect to see it everywhere. Lastly, it is always possible that phenotypes are affected by preparation, since it is impossible to know what the myofibrils look like in situ. However, all samples were prepared using the same protocol with replicates, and since we see a phenotype in our mutants and not in the control, this indicates that something is different between the two.

      (6) The same applies to the visualisation of the "hyper-contracted" phenotype; again, it seems to be an all-or-nothing phenotype in the zoom shown. An overview image should be shown. The zoom in Figure 4E would benefit from displaying phalloidin in a separate channel. Are actin filaments pulled out of the Z-disc? The latter is often seen in non-perfect cuts in wild-type, but the accumulation at the M is curious. It would be informative to locate the ends of the thick filaments in these cases or quantify thick filament lengths; do these invade the Z-discs? This can easily be done by a myosin staining.

      Is this a regional effect or does it depend on the individual or on the preparation? I am surprised to also see the "hyper-contraction" in 10% of wild-type 5-day adults.

      See previous response where we include overview images. Single-channel images are available; it is visible that actin filaments are not pulled out of the Z-disc. Phenotypes are consistent across individuals as evidenced in our replicates but do tend to be concentrated in certain regions of muscle.

      (7) The EM images would benefit from more overview images. At the moment, we only see a single sarcomere from wild type and mutant, with no quantification of the phenotype. Can the authors see the invading thin filaments into the M-band? The disrupted Z-disc phenotypes are impressive. What is the age of the animal shown in Figure 4?

      We have a panel displaying several mutant sarcomeres. Due to the selectivity and challenges of the EM preparation process, we do not believe we can perform meaningful statistics on them. It sometimes looks like myosin heads are visible in the H zone which may support the presence of thin filaments in the H zone (Fig. 4B and C). However, the quality of these particular EM images is not high enough to identify thin filaments. All phenotypes shown are from 3-week-old animals.

      (8) Is UH-3 GAL4 expressed at the adult stage?

      Yes, from 36 h APF into adulthood (Singh et al. 2014). Now mentioned in the results section.

      (9) Figure 6 would strongly benefit from a myosin staining. Do thin and thick filament lengths scale? It seems that overlap is reduced in the double hets. How can this be envisioned with Z-disc stability? Is myofibril diameter reduced?

      We searched for non-additive differences in myofibril diameter but were unable to detect any.

      (10) What is the FRAP turnover rate of a long Zasp-GFP compared to a short one in wild type? A difference would indicate that it is really the IDR domain that keeps Zasp52 longer at the Z-disc, instead of an indirect effect caused by Z-disc morphology

      We have newly added FRAP data of a GFP-tagged exon 15e construct which displays much lower turnover. This indicates that the IDR does indeed retain Zasp52 at the Z-disc.

      Reviewer #2 (Recommendations for the authors):

      (1) The total protein stain should also be included if it is used for quantitation in Figures 1B and 2A-A'.

      These are available on FigShare.

      (2) It is a bit confusing that the Alphafold plot is inversely correlated with the other 3 prediction programs, although this is explained in the legend. Maybe an Alphafold structure would help make the authors' point (colorized with known domains).

      The AlphaFold structure is almost entirely low-confidence disordered region except for the structured domains so we do not believe it would be helpful to include.

      (3) The title of Figure 8 says 'Certain ex15e defects are rescued by immobilization.' What other defects are not rescued? If true, these should be shown.

      There was a full rescue. We deleted the word “certain”

      (4) Please include a brief explanation of the spatiotemporal expression of UH3-Gal4.

      From 36h APF into adulthood (Singh et al. 2014). Now mentioned in the results section.

      (5) Statistics should be added to Figure 8E.

      Figure 8E (now 9E) has statistics.

      (6) The dark blue color used for integrin staining in Figure S3 is difficult to see. Changing this color may help visualize differences. Also, pointing them out with arrows, etc., will help clarify abnormalities.

      We have described these differences in the figure caption. Single-channel images are available for viewing in any color in FigShare.

    1. Author response:

      We thank the editors and reviewers for their thoughtful and constructive comments on our manuscript. We are pleased that they considered the core behavioral findings important and robust, especially the results showing that the magnitude and variability of others’ donations affected the magnitude and variability of participants' donations, respectively. We also appreciate their acknowledgement of the strengths of the experimental design, large sample sizes, the incentive-compatible and across-domain measures included in Experiment 4, and combined behavioral and computational approaches.

      We agree that the manuscript would benefit from greater clarification in several areas, further analyses, and more cautious interpretations. In the revised manuscript, we plan to clarify the rationale for sequentially presenting social information, the role of prediction responses, the theoretical motivation of examining the variability of others’ donation, the use of the between-subjects design, and the motivation for the RL framework. We also agree that the lack of a non-social repeated-donation control condition limits the interpretation of the phase effects. Our design permits strong inferences about differences in donation changes across different conditions, but it cannot establish that the phase-related changes are exclusively attributable to social information exposure. We will revise the wording accordingly, moderate the causal language, and explicitly discuss the limitations of our design.

      To strengthen the behavioral analyses, we plan to supplement the current mixed-effects linear models with models that reflect the repeated-measure structure of the task, including random slopes for the phase. We will also add statistics in the generalization results section and test the asymmetry between generous vs. stingy social influence. Moreover, the reviewers raised an important concern regarding the associations between psychopathy and the donation change. Because psychopathy is negatively correlated with initial donations in several experiments, absolute donation changes may partly reflect the distance between initial donation and the observed donation mean. We therefore plan to reanalyze the psychopathy effects by using signed donation changes and trial-level discrepancies between participants’ initial donations and the observed social information. These additional analyses will enable a more direct and precise assessment of whether psychopathy is associated with greater susceptibility to social influence.

      We further agree that the comparison and validation of the computational models should be strengthened. In the revised manuscript, we plan to clarify that the learning models are intended to describe the updating beliefs about a group-level donation norm from sequential social information, rather than learning about a single donor. We will expand the candidate model set to include non-learning models, such as models based on the actual social mean, a running average. We will also model the prediction phase and the second donation phase separately. This will help identify the models that provide explanatory values for both predictions of others’ donations and individual donation behaviors. In addition, because the models were not estimated via a Bayesian framework, we agree that the term “posterior predictive checks” is inappropriate. We will rename these analyses. We will also rerun the parameter and model recovery analyses using empirically informed noise levels separately for the prediction and donation phases. In addition, we will implement model-evaluation processes, such as cross-validation, that better reflect prediction for new participants.

      In Experiment 4, we plan to directly report the association between social-information-use measures in the perceptual and the donation task to strengthen the domain-generality effect. We will additionally examine whether social susceptibility in the perceptual task is associated with psychopathy by including all trials, including those in which participants moved away from or beyond the social value.

      Finally, we will correct the reporting and presentation issues identified by the reviewers, including the social information use equation in the perceptual task, the pseudo-SD of individual donations formula, supplementary figure captions, and task duration. We will also provide fuller experimental materials and make the preregistration links more prominent. In addition, we intend to make the analysis code, model-fitting scripts, and data available during the revision process.

      We greatly appreciate the editors’ and reviewers’ thoughtful suggestions, which will help us substantially strengthen the manuscript. We are grateful for the opportunity to address these important points and believe that the planned revisions will enhance the manuscript’s clarity, robustness, and its contribution to the understanding of social influence in donation behaviors.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how IRF4 and BLIMP1 coordinate human plasma cell differentiation. Using a stepwise in vitro culture system starting from primary human naïve B cells, the authors define a developmental window enriched for plasma cell precursors and use stage-specific CRISPR/Cas9 perturbation to examine the roles of IRF4 and PRDM1/BLIMP1 during the transition from plasmablast-like precursors to plasma cells. Single-cell transcriptomic analyses suggest that IRF4 acts early to license plasma cell differentiation, whereas BLIMP1 contributes more prominently to consolidation of the terminal plasma cell program. The authors further combine multiome profiling, CUT&RUN, motif modeling, and EMSA assays to propose the sublet nucleotide variation within ISRE/EICE-like motifs contributes to differential or shared binding by IRF4 and BLIMP1.

      Overall, this is a carefully performed and conceptually interesting study. It provides a useful experimental platform for dissecting human plasma cell differentiation and offers a mechanistic model for how two closely connected transcription factors can exert distinct and coordinated genomic functions during terminal B cell differentiation.

      Strengths:

      A major strength of the study is the establishment and detailed characterization of a human in vitro plasma cell differentiation system. The authors combine phenotypic, functional, and single-cell transcriptomic analyses to define the transition from activated B cells to plasmablast/plasma cell precursor-like cells and then to more mature plasma cells. This system is very useful for future perturbation studies of human plasma cell differentiation.

      A second strength is the stage-specific perturbation strategy. By targeting IRF4 or PRDM1 at the precursor-enriched stage, the authors avoid some of the interpretive limitations associated with earlier perturbations that would affect B cell activation, proliferation, and plasma cell commitment simultaneously. The distinct phenotypes observed after IRF4 versus PRDM1 perturbation provide support for a model in which these two factors act in a temporally ordered manner.

      A third strength is the integration of multiple genomic and biochemical approaches. The combination of single-cell RNA-seq, chromatin accessibility profiling, CUT&RUN, computational motif analysis, and EMSA assays provides a rich dataset and supports the idea that ISRE/EICE sequence variation contributes to differential IRF4 and BLIMP1 occupancy.

      Weaknesses:

      While the multi-omic approach and computational modeling are highly impressive, several major assumptions regarding the cellular differentiation model and genomic linkages require more rigorous validation.

      First, because CRISPR editing was performed on heterogeneous bulk Day 7 cells rather than purified precursor populations, it remains ambiguous whether the observed developmental blocks are truly specific to the prePC window.

      We agree that CRISPR/Cas9 editing of bulk D7 cultures complicates interpretation because this population contains both activated B cells and PB/prePCs. We will therefore revise the text to distinguish phenotypic effects measured across the bulk D7 culture from the downstream single-cell analysis focused on cells along the prePC-to-PC trajectory. In particular, our interpretation of IRF4 and BLIMP1 function in prePCs is based primarily on the D9 scRNA-seq analysis, in which cells arrested in the activated B cell compartment are not used to define the perturbed PC-trajectory states. We will clarify this analytic design in a future revision and temper language implying that all effects arise exclusively within prePCs.

      Second, given that IRF4 and BLIMP1 operate within a mutually reinforcing positive feedback loop, the phenotypic divergence between IRF4 KO and PRDM1 KO may reflect differences in protein degradation kinetics or hierarchical dominance rather than a strictly ordered "sequential function".

      We agree that the divergence between IRF4 and PRDM1 perturbations could reflect differences in protein turnover, or hierarchical dominance, in addition to developmental timing. We will revise the Discussion to state that our data support a temporally ordered model in which IRF4 acts early to license the prePC-to-PC transition and BLIMP1 consolidates the terminal state, but that the current experiments do not exclude alternative explanations related to hierarchical dominance or degradation kinetics. We will also note in the revised Discussion that degron-based perturbations, rescue experiments, and gain-of-function analyses would be needed to resolve the functional ordering of IRF4 and BLIMP1 with higher temporal precision.

      Lastly, the motif-lexicon model is elegant and supported by biochemical DNA-binding assays, but the link between motif variation and gene regulation in cells remains partly correlative. (1) Direct testing of selected regulatory elements would make the causal claim stronger. (2) Alternatively, the authors should temper the language and present the motif lexicon as a predictive model for differential occupancy rather than as a fudlly demonstrated mechanism of gene regulation.

      We agree that the current data support the motif lexicon primarily as a predictive model for differential TF occupancy rather than as a fully causal mechanism of gene regulation. We will therefore revise the relevant text in the Results and Discussion. The EMSA data directly test nucleotide-dependent binding preferences, and the CUT&RUN/multiome analyses show that these motif variants are differentially associated with IRF4- or BLIMP1-bound DEG-linked OCRs. However, direct causal testing of endogenous regulatory elements, for example by base editing of selected ISRE/EICE variants, will be required to determine whether these variants are sufficient to predictably alter gene activity in differentiating plasma cells.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Lau et al. investigates the mechanisms underlying IRF4 and BLIMP1 transcriptional activities during antibody-secreting cell fate decision. Both master regulators of plasma cell differentiation, these two transcription factors have distinct targets and non-overlapping roles. The authors used an in vitro culture system to generate antibody-secreting cells from human naïve B cells, and scRNA-seq, Crispr Cas9 editing, and Cut&Run to dissect the molecular mechanisms defining their specificity.

      Strengths:

      The experiments are overall well executed, and the manuscript is well written. The in vitro culture model appears to generate genuine human antibody-secreting cells. The identification of non-conserved nucleotides within the binding motifs that induce the specific binding of IRF4 or BLIMP1 is convincing, novel, and exciting.

      Weaknesses:

      The authors need to correct some overstatements and flaws to improve the manuscript.

      In Figure 1f, the authors aimed to determine whether in their culture system the plasma cells emerged from the plasmablasts or directly from the activated B cells. First, it is noticeable that the distinction between plasmablasts and plasma cells relies here only on the expression of CD138. It does not include a higher capacity to secrete antibody or their proliferative state. In Figure 1e, the authors could have strengthened their distinction by showing the Ki67 staining at day 21 for both subpopulations.

      Second, this question does not seem to be related to IRF4 or Blimp1 activity, and thus one could wonder if it is relevant to this study.

      Finally, and most importantly, the design of the experiment appears flawed to me. The authors sorted cells at day 7 of culture based on their expression of CD20 and put the two subpopulations back for 14 more days. This culture system is a stepwise system, and it is not specified if the CD20<sup>+</sup> cells were put back in the day 7 condition or the day 0 condition with the CD40L stimulation

      We agree that CD138 alone does not fully define terminal PC maturation. In the revised manuscript, we will clarify that CD138 was interpreted in the context of a broader maturation profile, including CD20 downregulation, ICAM2 upregulation, IRF8 loss, IRF4/BLIMP1 expression, Ki-67 loss, and antibody secretion. The D7 PB population was proliferative and CD138<sup>-</sup>, whereas D21 CD20<sup>-</sup> cells were largely Ki-67<sup>-</sup> and included CD138<sup>+</sup> cells, supporting their progressive maturation. We will include Ki67 analysis in CD138<sup>-</sup> and CD138<sup>+</sup> cells at D21 in the revision.

      We agree that the motivation and culture conditions for this experiment required a clearer explanation. The purpose of the D7 sort-and-reculture experiments was to identify the developmental window enriched for cells competent to generate PCs, thereby defining the stage at which IRF4 and PRDM1 should be perturbed. Sorted D7 CD20<sup>+</sup> actB cells and CD20<sup>-</sup>CD38<sup>+</sup>CD27<sup>+</sup> PBs were both placed into the same D7-D14 differentiation conditions, allowing a direct comparison of their PC-generating competence under identical culture conditions. We will clarify this design in the Results and Methods. We have not tested whether returning D7 CD20<sup>+</sup> cells to D0 conditions involving CD40L stimulation restores PC differentiation, and we will now acknowledge in the revision that the CD20<sup>+</sup> fraction may contain cells with distinct intrinsic differentiation potential, including cells differentiating into non-PC states.

      What if it is not the case and some are anergic or have committed to the memory B cell fate during the first 7 days? Then the day 7 CD20<sup>+</sup> fraction would be enriched in these cells.

      We agree that the CD20<sup>+</sup> D7 cells may contain anergic or memory B cell precursors. However, this does not alter the interpretation that the CD20<sup>-</sup> (CD38<sup>+</sup>/CD27<sup>+</sup>) PBs contain a PC precursor population. Even if memory B cells are generated in the CD20<sup>+</sup> fraction by D7, based on the sorting experiments, their presence would have little-to-no effect on developing PC precursor populations.

      Moreover, this experiment didn't show that the plasma cell derived from the plasmablasts in the strict sense of the term, as the CD138<sup>+</sup>CD20- cells could be a mix of proliferative plasmablasts and immature plasma cells.

      We agree that the heterogeneous nature of CD20<sup>-</sup> cells complicates the interpretation. However, we would like to emphasize that all D7 CD20<sup>-</sup> cells are Ki67<sup>+</sup> whereas all D21 CD20<sup>-</sup> cells are nearly all Ki67<sup>-</sup>. Though we concede these could include recently proliferated PCs at D21, it is consistent with this population becoming quiescent. To better address this question in a future revised version, we will include direct measurements of Ki67 levels in CD138<sup>+</sup> and CD138<sup>-</sup> cells at D21.

      In Figure 3a and thereafter, the authors claimed that IRF4 acted earlier than BLIMP1, but both deletions strongly affected differentiation at day 7. IRF4 might have a stronger effect, but it does not mean that it had an earlier effect. To substantiate their claim, the authors would need to demonstrate that, at an earlier time point, deletion of IRF4, but not BLIMP1, results in defective differentiation.

      We agree that the current data do not by themselves prove that IRF4 acts earlier than BLIMP1 in developmental time. We will revise the text to state that the data are consistent with a temporally ordered model, rather than demonstrating strict sequential action. The latter interpretation is based on the distinct IRF4 KO stunted PC state observed by D9 scRNA-seq (Fig. 3C), together with the stronger early phenotypic effect of IRF4 loss (Fig. 3A). However, because both factors are mutually reinforcing and because perturbations were not performed across multiple time points, alternative explanations remain possible, including differences in editing efficiency, protein stability, and feedback-dependent TF decay. We will modify the text to better explain the rationale behind this interpretation while also acknowledging alternative interpretations that do not involve sequential IRF4-BLIMP1 functions (see response to Reviewer 1).

      In Figure 3b, the authors stated that in each individual KO the expression of the other transcription factor was lower. Given that there were no cells in the gate, it is puzzling to figure out how these expressions were compared.

      In Figure 3c, on the UMAP the bottom right part of the activated B cell cluster does not appear to be attributed to any condition. How can it be? Besides, it is highly surprising that at D9 we cannot see any plasmablast on these UMAP, even in the control. Based on the G1/S and G2/M scores, none of the ASC represented were proliferating. Could the authors explain this strong discrepancy with Figure 1?

      Another discrepancy exists between Figure 3b and c: Figure 3b depicted no IRF4- or BLIMP1-expressing cells in either KO, so what were the stunted PC and the BLIMP-KO PC reported in Figure 3c?

      We thank the reviewer for identifying these points of confusion. We will revise Fig. 3B to display outlier events and frequencies more clearly and revise the figure legend to clarify the donor origin of the displayed UMAPs and corresponding supplemental analyses. We will also clarify that Fig. 3B and Fig. 3C represent distinct readouts: flow cytometry measures IRF4 and BLIMP1 protein abundance, whereas scRNA-seq resolves transcriptional states after perturbation. Thus, the “stunted PC” state in IRF4 KO cells is defined at a transcriptional level as a population positioned between prePCs and PCs, not as a population retaining normal IRF4 or BLIMP1 protein expression. We further clarify that the apparent reduction in proliferative plasmablast-like cells at D9 likely reflects both the later timepoint relative to D7 and differences between transcriptional cell-cycle gene scores and Ki-67 protein persistence.

      What are the signature genes defining pre-PC and the score depicted in Supplementary Figure 3d, as the materials and methods only state that they are intermediate between PC and B cells?

      The signature genes defining the scores in Fig. S3D are listed in Table S2. We will clarify the source of these genes in the figure legend of a future revised version.

      Could the authors show IRF4, BLIMP1 and some of their known target expression in these populations?

      We thank the reviewer for this suggestion to highlight IRF4, BLIMP1 and exemplar target genes in the various populations. We will update Fig. S3 to show transcript levels of IRF4, PRDM1 and an example of one of each of their target genes in unperturbed cells to demarcate their normal expression pattern.

      The authors claim that BLIMP1 is not needed to initiate the transition from pre-PC to PC, but in Figure 1, the intracellular staining showed that at day 7 the antibody secreting cells already expressed BLIMP1. This would rather suggest that BLIMP1, unlike IRF4, does not need to be maintained once the cell reaches a certain point.

      We agree that the data do not exclude the possibility that BLIMP1 is required before the perturbation window but is less continuously required once cells have progressed beyond a defined prePC stage. Our statement that BLIMP1 is not required to initiate the prePC-to-PC transition is based on the observation that PRDM1 KO cells did not accumulate in the prePC or stunted PC intermediate states observed after IRF4 loss. We will revise the text to make this interpretation more precise by stating that BLIMP1 appears less important than IRF4 for progression into a PC-like transcriptional state but is required for efficient consolidation of the mature PC program. We will also acknowledge that differences in editing efficiency, protein persistence, and timing of BLIMP1 action could contribute to the observed differences in phenotypes.

    1. Author response:

      The authors thank the reviewers for their thorough and fair assessment of our manuscript. We are currently working to edit the manuscript based on the critiques and guidance offered by the reviewers. This will consist of fixing grammar and typos, expanding the material and methods section to include more information on the behavioral assays, modifying graphs for clarity between visuals and interpretations, and correcting our mistakes in neuroanatomical labeling.

      Public Reviews:

      Reviewer #1 (Public review):

      In their submitted manuscript, Harkinish-Murray and colleagues from the Kozol lab present convincing evidence for a genetically encoded shift in the odor perception of cavefish compared to their surface ancestors. Surface Astyanax, just as zebrafish, are attracted to food odors and are repelled by death odors and the alarm substance Schreckstoff (released from damaged skin by specialized club cells). Based on the experimental evidence in this manuscript, however, their cavefish counterparts are attracted to these odors as well. This would make sense, in an evolutionary framework, as predation is less likely in cave settings and decaying fish are a valuable source of nutrients for their living counterparts.

      Using an F2 hybrid cross scheme between surface fish and cavefish, authors also provide compelling evidence that genetic factors are behind this behavioral shift. Furthermore, they also show that this behavior (i.e., attraction to skin and decay extracts) can be observed in surface fish given long enough food deprivation. This latter observation also makes sense in the light of evolution and is genuinely interesting as it also provides a plausible roadmap to the shift in behavior through Waddingtonian genetic assimilation.

      The manuscript is generally well written and clear, we have identified only few weaknesses, some regarding the presentation of the data.

      (1) For Figure 3, on the x-axis of panels b, e, and h, supposedly we see surface fish vs. different cavefish populations. This is currently missing and makes the figure harder to interpret. Also, two populations (panel e) show a bimodal distribution upon indirect white light exposure, suggesting that some fish still acted as if they were exposed to direct light, while others acted as if they were in darkness (infrared light). We believe this warrants more consideration as it could tell us something about the existing (and relevant) genetic variance within this population. It is also notable that the third cavefish population also showed increased odor indices under indirect white light and infrared light conditions, suggesting that increasing the number of observations could have yielded a statistically significant result.

      We agree with the reviewer that our light testing data suggests complexity in the response to indirect white light within certain cave populations. In addition, an expanded sample size would likely provide clarity on whether individuals fall within two groups, behavior that looks like direct light or infra-red light, that could relate to genetic variation within cavefish populations. We are currently working to reassess the current data and determining the best course of action for continued studies related to light exposure.

      (2) Some extra details about the methods could also be provided to enhance the reproducibility of the experiments.

      We agree with both reviewers that the methodological section on behavior needs to be expanded. We are currently editing our methods section to include more detail on water exchanges, odor preparation, timing, biological replicates, and binning.

      (3) A more serious concern is about the anatomical designation of particular brain regions in Figure 7d and consequently Figure 7f. Whereas we would agree with the positioning of the medial pallium (Dm), we think the region depicting the thalamus is in fact still part of the telencephalon, and the real thalamus should be more posteriorly. On the other hand, we think that the preoptic areas should be under the pallium and not posterior to it (see PMID: 22586363 for corresponding zebrafish anatomy). We would suggest, therefore, that the authors revisit this issue (a minor one, considering the depth of the results presented in the manuscript), and provide a better anatomical annotation - e.g., the identity of particular brain regions could be backed up by Hybridization Chain Reaction experiments for region-specific transcripts. (Disclaimer: we do not consider ourselves experts in adult cavefish neuroanatomy; therefore, we consulted in this case a colleague with much more knowledge on this topic.)

      We agree that our annotation was incorrect or more accurately mislabeled in our write-up of the preprint and submitted manuscript. Therefore, we have now re-assessed the regions using the tissue cleared and light sheet collected zebrafish atlas, Adult Zebrafish Brain Atlas (AZBA; doi: 10.7554/eLife.69988). We are now editing the resubmission in the following manner: our initial labeling of the ventromedial thalamus will be changed to the lateral olfactory tract (nLOT) of the pallium and the preoptic region to the ventromedial thalamus (VM). We will provide a comparable z-slice of the AZBA segmentation file to illustrate the similarity in position. This would also support a known continuous circuit of olfactory integration, with information flowing from the lateral olfactory tract-to the piriform cortex-to the thalamus. We also agree that a more accurate assessment in Astyanax would require HCR in situ hybridization of markers for those specific brain regions or a neurocomputational brain atlas for adult Astyanax populations. Finally, we assert that this small dataset is preliminary at best and only provides regions of shared activity that could explain anything from perception related processes to relay of odor signaling unrelated to perception. Further work with larger sample sizes and additional populations are currently underway for a follow-up study on the neurobiological basis of olfactory processing and perception in adult cavefish.

      (4) It would also be useful to expand the brain imaging data displaying results for similar tests in surface fish, to see if skin and decay extracts trigger different or similar brain activity in those fish.

      We agree with the reviewer that the brain mapping section lacks a sufficient sample size and no control group for comparison (surface fish). However, we found the variation in pERK intensity (notably the putative nLOT) to be informative and decided to include the dataset in the manuscript. We are currently working to fill in these data gaps by sampling all populations and increasing the Pachon cavefish sample size. This will be a follow-up study as mentioned above in the last rebuttal paragraph.

      Further work will surely be able to discern the more precise genetic changes that made the shift in behavior possible. Once these causative variants (or at least linked markers) are determined, it will be quite revealing to see if these variants are indeed already present in the surface population (as hinted by the authors), and also, if besides the Surface x Tinaja F2 hybrids, crosses between other cave populations and surface fish can be performed, we could also see how much evolutionary convergence happened in the parallel evolution of different cave morphs. Were there multiple possible pathways for similar behaviors in different cave populations, or - as in freshwater stickleback populations - do we see broadly the same genetic playbook repeated each time?

      We agree with the reviewer that the hybrid results setup a promising follow up project to map these traits genetically. We are continuing to test odor perception in other hybrid populations and have started Quantitative Trait Locus mapping experiments.

      Another outstanding question, also demonstrated and discussed, albeit briefly, in this paper relates to the behavior-modulating effect of light in cavefish. What is the physiological relevance for a dark-dwelling animal to have this capacity? Is this just the chance result of occasional gene flow from surface populations, or does it have a genuine evolutionary significance?

      Reviewer #2 (Public review):

      Summary:

      The authors tested whether the olfactory cues that drive attraction or avoidance behavior have diverged between surface‑dwelling and cave‑adapted strains of the Mexican cavefish Astyanax mexicanus. They use high‑throughput odor‑discrimination assays between known attractants and repellents by calculating an "odor index" per fish (=the difference in time spent in an odor zone versus a control zone). Further, hybrid crosses to probe heritability, starvation experiments to assess plasticity of odor perception, and whole‑brain pERK detection/mapping to link behavioral changes with known localized neural activity. The results support the hypothesis that the extreme cave environment has selected for an approach response to stimuli that are ancestrally aversive (like alarm or death odors) but in harsh environments can be used as guidance to the rare food sources in this ecosystem.

      Strengths:

      The odor index analysis is convincing, and the experiments for odor attraction/avoidance are robustly performed. The light-to-darkness shift reflected by avoidance to attraction in cavefish towards skin odors is compelling and carefully analyzed. The analysis of odor indices of three cave-dwelling populations in comparison to surface fish highlights a similar regime, yet with differences among the different populations, suggesting population-specific genetic variation.

      Another strength of the paper is exactly this genetic inheritance study by generating F2 hybrids of cave-dwelling and surface-living individuals. The hybrids displayed a continuous range of odor indices for social, alarm, and death odors, indicating that these traits are heritable and likely based on additive genetic markers. Further, the authors uncovered a sexual dimorphism: only female cavefish exhibited approach behavior to social odors, whereas males remained neutral. This result aligns with known differences in olfactory organ morphology between sexes of other species from harsh environments.

      Although limited in number, the neurophysiological correlation using whole‑brain pERK mapping after 10 min of odor exposure is convincing. The data revealed overlapping activation in the thalamus and pre‑optic region for food and decay odors, suggesting that these brain areas mediate the evolved attraction response to previously repellent stimuli.

      Overall, the manuscript presents a concise story: cavefish have evolved attraction to alarm and death odors as a result of shifting from ancestral avoidance-driven to attraction by genetic changes and physiologically similar activation of specific neural circuits. The evidence is robust, with multiple independent experiments (behavioral assays, hybrid genetics, starvation experiments, and brain mapping) that collectively support the conclusions.

      Furthermore, exposure to unpleasant odors can not only be tolerated but can even serve as a trigger for foraging. This plasticity demonstrates that genetic predispositions can be put into practice through active changes in physiology in species or organisms confronted with (drastically) changing environmental conditions.

      Weaknesses:

      I value that the authors are critical of their own data, indicating low numbers in the pERK/brain experiments. Yet this is a weak point as the statistical power is thus limited. However, their reasoning is careful, based on the results and not over-interpreting.

      We agree with the reviewer and direct their attention to the same critique by reviewer 1. We believe this is predominantly preliminary data that was included due to the conspicuous increase in pERK signal from the putative lateral olfactory tract (nLOT) for food and decay exposed cavefish. We are continuing to work on odor stimulated brain mapping and look forward to publishing a comprehensive dataset across wildtype and hybrid populations.

      The layout/design of the ethograms (bout category plots) for both individual and population-wise are not easy to follow. Reworking these display items to convey the information is necessary.

      We agree with the reviewer that the ethograms are challenging to read, especially due to our use of different colors for odor categories. We are currently preparing alternative graphs for displaying ethograms that reduce confusion and make following bout transitions for individual traces and bout probabilities for populations easier on the eyes.

      Taken together, the manuscript uses odor perception and attraction/avoidance behavior studies to show that environmental changes (light-to-darkness) have an immediate impact on smell perception and behavior. Attraction to otherwise repellent odors is used by cavefish to likely adapt to harsh environments with low food sources. The manuscript convincingly demonstrates this plasticity, which is an interesting idea to follow up for other traits spreading among a population. This also underlines that a genome may be fixed and the blueprint for behavioral traits, but extrinsic cues can readily be adapted to change wired behavior even to the extreme as reported here: changing avoidance to attraction.

    1. Author response:

      We thank the editors and reviewers for their thoughtful evaluation and constructive feedback. We are pleased that all three reviewers recognize the importance of mapping Cas9-induced sister chromatid exchanges (SCE) as a previously invisible repair outcome, and that the RDCP analysis is a notable feature of the study.

      We note that since our manuscript was posted, two companion studies in Science have provided direct biochemical evidence for the TRAIP-dependent pathway we discussed:

      (1) Fujisawa & Labib (Science, 2026; DOI: 10.1126/science.aeh2300) showed that TTF2 bridges CDK1-phosphorylated TRAIP to DNA Polymerase epsilon in the replisome, triggering mitotic CMG helicase disassembly, fork cleavage, and repair via SCE. Loss of this pathway reduced replication stress-induced SCE approximately two-fold in mouse ES cells.

      (2) Can et al. (Science, 2026; DOI: 10.1126/science.aeh1834) independently identified the same CDK1-TTF2-TRAIP axis in Xenopus egg extracts and validated it in HCT116 cells, showing that disrupting the TRAIP-TTF2 interaction reduced common fragile site deletions.

      We will incorporate these references in the revised discussion while still framing our RDCP observations as consistent with, rather than definitive proof of, this pathway.

      Below we briefly address the main points raised in the public reviews.

      Reviewer #1:

      We agree that the mechanistic interpretation of the RDCP signature should be presented more cautiously. We will reframe the URR/TRAIP discussion as a model, replacing language such as "direct genetic evidence" with "consistent with." We will add a summary table of RDCP data. We will also expand the description of rescued SCE calls. We will clarify what the DNA repair gene targeting experiment can and cannot answer (delayed protein loss, essential-gene selection) - we think that there is a notable difference at the bulk vs. at the single-cell level depending on the nature of the assay. Fig.1 fonts, labels, and pileup plot descriptions will be improved.

      We agree that the high-SCE subpopulation is particularly interesting and we cannot currently distinguish higher RNP uptake, a permissive cell-cycle state, altered expression, or stochastic variation. This may be better explored by future co-assays with sci-L3-Strand-seq.

      Reviewer #2:

      We agree with the limitations that Cas9-induced DSBs can be dependent on the cell cycle stage and the number of times cuts are made. We will add a brief discussion on this limitation in extrapolating the mechanisms of DNA instability and DNA repair from the observed genomic rearrangements.

      We agree that "a single Cas9 cut" should be revised to "Cas9 targeting of a single genomic locus" to accurately reflect the experimental design. We will clarify the possibility of multiple rounds of cutting at the same sites. We will also clarify "non-local" by modifying Fig.1 - we used this term specifically to include the possibility of inter-sister NHEJ.

      Reviewer #3:

      We will improve the self-contained nature of the manuscript so that readers need not consult the earlier NAR papers, and the companion preprint to understand the key results.

    1. Author response:

      We thank the editors and reviewers for their careful evaluation and constructive suggestions. During the review process, we identified and corrected several presentation, terminology, citation, and figure-legend errors, and these corrections have been incorporated into the current version of preprint. Following the reviewers’ comment, we have also revised the title of the manuscript. We are now preparing a substantive revision that will distinguish more clearly between experimentally supported conclusions and hypothetical regulatory relationships. We plan to examine gene expression at an earlier time point after Br-c knockdown, further investigate the relationships among Kr-h1, chinmo, and E93 using additional RNAi experiments, characterize the cuticular phenotypes of precocious prepupae at higher magnification, and determine whether severe E93-knockdown individuals exhibit evidence of a repeated pupal developmental program. After completing these experiments, we will cautiously revise the proposed regulatory model and moderate conclusions that are not directly supported by the current evidence.

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife assessment

      This study investigates the role of the bile acid receptor TGR5 in adult hematopoiesis of the mouse model. The findings are potentially useful because the loss of TGR5 leads to dysregulation of bone marrow adipose tissue (BMAT) that has emerging regulatory functions. However, the study is still incomplete because the mechanism of TGR5 is not clear, the stromal cells expressing TGR5 have not been well defined, and there is not strong evidence for the role of TGR5 in recovery from transplant stress.

      We thank the eLife editorial team for handling our manuscript. In our revised version, we took into consideration the suggestions of the reviewers, which we believe have significantly improved the quality of the current study. In summary, our new data provide further evidence that TGR5 is expressed in both hematopoietic cell lineage and stromal cells of the bone marrow (BM), including subpopulation analyses for both. While steady-state hematopoiesis remains intact in TGR5 knockout mice, we demonstrate that loss of TGR5 significantly impacts progenitor reconstitution under stress conditions, alters the BM adipose tissue (BMAT) homeostasis in a sex-specific manner and influences the balance of stromal cell differentiation. In particular, TGR5 deficiency resulted in reduced regulated BMAT and an accumulation of adipocyte progenitors, correlating with improved hematopoietic recovery following BM transplantation. The BMAT decrease is further observed in physiologically and pathophysiologically relevant contexts such as aging and obesity, where TGR5 deficiency is associated with a decrease in the myeloid bias. Collectively, our findings support a previously unrecognized role for TGR5 in maintaining BM niche integrity and highlight its potential as a modulator of hematopoietic support, although the precise molecular mechanisms still need to be elucidated.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      Alonso-Calleja and colleagues explore the role of TGR5 in adult hematopoiesis at both steady state and post-transplantation. The authors utilize two different mouse models including a TGR5-GFP reporter mouse to analyze the expression of TGR5 in various hematopoietic cell subsets. Using germline Tgr5-/- mice it's reported that loss of Tgr5 has no significant impact on steady-state hematopoiesis, with a small decrease in trabecular bone fraction, associated with a reduction in proximal tibia adipose tissue, and an increase in marrow phenotypic adipocytic precursors. The authors further explored the role of stroma TGR5 expression in the hematopoietic recovery upon bone marrow transplantation of wild-type cells, although the studies supporting this claim are weak. Overall, while most of the hematopoietic phenotypes have negative results or small effects, the role of TGR5 in adipose tissue regulation is interesting to the field.

      Strengths:

      This is the first time the role of TGR5 has been examined in the bone marrow.

      This paper supports further exploration of the role of bile acids in bone marrow transplantation and possible therapeutic strategies.

      We thank the reviewer for pinpointing the strengths of our study.

      Weaknesses:

      (1) The authors fail to describe whether niche stroma cells or adipocyte progenitor cells (APCs) express TGR5.

      Using the TGR5:GFP reporter model, we identified GFP<sup>+</sup> cells in the stroma-enriched CD45-Ter119-CD31- population that contains the adipogenic progenitor cells (APC).

      These data, along with the corresponding gating strategy are outline in Figure 6A and B.

      We found the subpopulation analyses within the stroma gate challenging at the individual mouse level given the limited cell numbers for these progenitor populations. We attempted to circumvent this by concatenating the individual files per genotype (WT and TGR5:GFP) for two independent experiments. We were surprised to find that there is little to no GFP expression in the APC population. Nevertheless, we found that the CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup>, multi-potent stem cell-like population (Ambrosi et al., 2017), reproducibly showed a TGR5:GFP positivity comparable to that of Ly6<sup>lo</sup> monocytes. We believe this would be compatible with our results showing an increase in CFU-F in Tgr5<sup>-/-</sup> mice, but we remain cautious about its interpretation. Results are shown in Author response image 1 for the concatenated flow plots and numerically in Error! Reference source not found. considering all individual mice analyzed (i.e., non pooled) for completeness. We kindly request the Reviewer’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section.

      Author response image 1.

      Flow cytometry gating strategy used to identify stroma subpopulations in the stroma enriched CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup> gate and their GFP signal for BM cells in TGR5:GFP mice. Subpopulations were immunophenotypically defined as CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (adipogenic progenitor cells (APC)) and CD45Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (multi-potent stem cell-like populations) (Ambrosi et al., 2017). Results are shown as the concatenated data of all mice in each phenotype for two independent experiments, as indicated.

      Author response table 1.

      Frequencies in the total live cell gate for adipogenic progenitor cells (APC) and CD45Ter119<sup>-</sup>CD31<sup>-</sup>CD24<sup>+</sup>Sca1</sup>+</sup> (multi-potent stem cell-like populations, MPSC-like) (gating as in Ambrosi et al., 2017) in the experiments presented in Author Response Figure 1, expressed as average +/- 95% confidence interval for the two independent experiments. The number of mice per genotype and per experiment is indicated in parenthesis. **Paired, two-tailed Student’s t-test statistical analysis for the combined experiments (n=8 per group) indicates p < 0.01 for GFP expressing cells in APC versus MPSC-like populations.

      (2) Although the authors note a significant reduction in bone marrow adipose tissue in Tgr5-/- mice, they do not address whether this is white or brown adipose tissue especially since BA-TGR5 signaling has been shown to play a role in beiging.

      The nature of BMAT and how it relates to brown, white or brown/beige adipose tissue has been a persisting question in the field. Our understanding is that BMAT is currently considered as a distinct adipose depot that is neither white nor brown/beige (Sebo et al., 2019; Suchacki et al., 2020). BMAT does not express UCP1 to an appreciable extent, with reports showing that detectable expression possibly stems from contamination by tissues surrounding bone (Craft et al., 2019). Beyond this consideration, as the regulated BMAT in Tgr5<sup>-/-</sup> mice is almost absent, determination of the brown/beige vs white nature of the little regulated BMAT that remains would be technically extremely challenging.

      (3) In Figure 1, the authors explore different progenitor subsets but stop short of describing whether TGR5 is expressed in hematopoietic stem cells (HSCs).

      We have added these data to the manuscript as part of Figure 1C.

      We have further expanded our data in Figure 1D and Figure 1–figure supplement 1A with the expression of TGR5:GFP in megakaryocyte progenitors (Lin<sup>-</sup>cKit<sup>+</sup>Sca1CD150<sup>+</sup>CD41<sup>+</sup>) as shown in Author response image 2.

      Author response image 2.

      A, representative flow cytometry gating strategy used to identify megakaryocyte progenitors (MkProg) and GFP positivity in TGR5:GFP mice and their wild-type controls. B, frequencies of GFP<sup>+</sup> cells in MkProg population in the BM of 8-12-week-old male TGR5:GFP mice and their controls (n=3 for wild-type control mice, n=4 for TGR5:GFP mice). Results represent the mean ± s.e.m., n represents biologically independent replicates. Two-tailed Student’s t-test (B) was used for statistical analysis. p-values (exact value) are indicated.

      Finally, we have completed the characterization of BM progenitor populations to include the erythroid lineage, thus covering the main hematopoietic populations. We have added these data to Figure 1E and Figure 1–figure supplement 1B.

      (4) Are there more CD45+ cells in the BM because hematopoietic cells are proliferating more due to a direct effect of the loss of Tgr5 or is it because there is just more space due to less trabecular bone?

      We observe an average 20% increase in CD45<sup>+</sup> cell counts in baseline Tgr5<sup>-/-</sup> mice. The absolute volume of bone and BMAT lost in these animals does not account for 20% of the medullary cavity volume, so we speculate that the increase in CD45<sup>+</sup> counts is not solely due to increased available volume for these cells.

      (5) In Figure 4 no absolute cell counts are provided to support the increase in immunophenotypic APCs (CD45-Ter119-CD31-Sca1+CD24-) in the stroma of Tgr5-/- mice. Accordingly, the absolute number of total stromal cells and other stroma niche cells such as MSCs, ECs are missing.

      These data are now included in the manuscript and in Author response image 3. Although we detect an increase in the relative frequency of APCs (Figure 7A), on a per-leg basis we did not detect an increase in APC numbers per leg (Author response image 3). We did however observe a decrease in total CD45<sup>-</sup>Ter119sup>-</sup>CD31sup>-</sup> stromal-enriched cells in Tgr5<sup>-/-</sup> mice. Nevertheless, given that only 2–5% of stromal cells are recovered in single-cell suspensions compared with native tissue (Coutu et al., 2017; Gomariz et al., 2018), we consider results for absolute stromal cell numbers prone to over interpretation. Our conclusion, therefore, remains one of relative enrichment of immunophenotypic APCs, supported by in vitro findings of increased adipogenesis and CFU-F formation after plating equal cell numbers.

      Endothelial cells (CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup>) were also quantified, with no differences observed between groups.

      Author response image 3.

      Absolute number of adipocyte progenitor cells (APC), stroma cells (CD45-Ter119CD31<sup>-</sup>) and endothelial cells (CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup>) (n=5 for both Tgr5<sup>+/+</sup> and Tgr5<sup>-/-</sup> mice). Results represent the mean ± s.e.m., n represents biologically independent replicates. Unpaired, two-tailed Student’s t-test was used for statistical analysis. p values (exact values) are shown.

      (6) There are issues with the reciprocal transplantation design in Fig 4. Why did the authors choose such a low dose (250 000) of BM cells to transplant? If the effect is true and relevant, the early recovery would be observed independently of the setup and a more robust engraftment dataset would be observed without having lethality post-transplant. On the same note, it's surprising that the authors report ~70% lethality post-transplant from wild-type control mice (Fig 4E), according to the literature 200 000 BM cells should ensure the survival of the recipient post-TBI. Overall, the results even in such a stringent setup still show minimal differences and the study lacks further in-depth analyses to support the main claim.

      We thank the reviewer for this comment. On the one hand, we respectfully disagree on the relevance of the effect size, as Tgr5<sup>-/-</sup> mice recover from low platelet and leukocyte counts significantly faster than wild-type controls. Tgr5<sup>-/-</sup> recipients recovered their platelet levels faster than the Tgr5<sup>+/+</sup> recipients, showing values consistently above 200.000/µL one week sooner; platelet levels below this threshold are considered a risk factor for bleeding events (Morowski et al., 2013; Vannini et al., 2019). In addition, on day 15 post-irradiation, Tgr5<sup>-/-</sup> recipients had higher neutrophil levels, at over the 500 cells/µL-threshold value for infection risk (Morowski et al., 2013; Vannini et al., 2019), but this difference was no longer statistically significant upon stringent multiple-comparison correction (uncorrected p-value 0.028, multiplecomparison corrected p-value 0.108). Underlining the relevance, in a clinical setting, G-CSF is routinely administered to patients daily to enhance myeloid recovery even if the acceleration of recovery is by 1-2 days (Trivedi et al., 2009).

      From the perspective of mortality, we agree that it is higher than expected and constitutes a limitation of our work. Note that we discovered a mistake in data plotting during the preparation of the revised version of the manuscript which renders the mortality curve no longer statistically significant, with a p value that changed from 0.0254 to 0.0676. We have duly changed our interpretation in the text to that of a trend towards higher survival.

      Regarding the mortality in our experiments being higher than expected, we have unfortunately suffered from cases of “swollen muzzles syndrome” in our facilities that have greatly hampered our ability to perform myeloablation experiments (Garrett et al., 2018), as even sublethal doses have resulted in the appearance of side effects that were reasons for euthanasia under Swiss legislation. For example, a strong reduction in mobility requires immediate euthanasia. All experiments were performed blinded to genotype allocation, so we can reasonably exclude experimenter bias. Finally, it could be argued that mice with more marked symptomatology leading to euthanasia are more likely to have hematopoietic deficits, which in our case was mostly seen for WT animals. We have therefore chosen to report mortality alongside the longitudinal assessment of peripheral blood counts to ensure clarity of the coherent effect and full transparency. We strongly believe that this disclosure, when examined as a limitation in the discussion section, aligns with the open science efforts advocated by eLife. Unfortunately, it is beyond the scope of this manuscript and the authors' timeline to rederive the Tgr5<sup>-/-</sup> colony in our new facility and perform new bone marrow transplantation studies with titrating doses of BM.

      Lastly, the choice of 250,000 BM cells serves as the control dose per recommended standards for competitive repopulation assays (Purton and Scadden, 2007). Quoting from this methodological reference paper: “In our experiments, each recipient mouse receives cell doses … together with 2x10<sup>5</sup> competing congenic bone marrow. We have found these cell doses sufficient to detect both reductions (Purton et al., 2006) and increases (Janzen et al., 2006; Walkley et al., 2005) in HSC numbers in different mutant mice… Furthermore, caution should be used when designing competitive repopulation assays, as it has been shown that the reliability of this assay is critically dependent on the numbers of HSCs present in the populations being assessed: when too few or too many HSCs (recipients of <1 1x10<sup>5</sup> or >2x10<sup>7</sup> bone marrow cells each from donor and competing sources) are present, the data may not be meaningful (Harrison et al., 1993).”

      Moreover, our lethal radiation rescue experiments with 2.5x105 cells are designed to deliver a minimal hematopoietic source known to rescue the great majority of animals in standard conditions, and thus to best mimic the situations when enhancement of hematopoietic recovery would be clinically meaningful (as used by our group members in (Naveiras et al., 2009; Tratwal et al., 2020; Vannini et al., 2019)). In our hands and in the absence of “swollen muzzles syndrome”, this approach leads to 80-95% overall survival, which mirrors the clinical setting of autologous transplantation that our transplants are meant to mimic. In clinical practice strategies, enhancing hematopoietic recovery would be most useful to improve morbimortality in either alternative hematopoietic progenitor allotransplants, known associated with delayed engraftment, or in the case of febrile neutropenia associated to bacteriemia, which affects 11-15% of both hematopoietic autologous or allogeneic transplant patients (Gil et al., 2013; Gooley et al., 2010). In summary, septic neutropenia was unfortunately modelled by the rescue experiments presented in Figure 7 C-J with swollen muzzles syndrome concomitant to the reconstitution, but we strongly believe that this complication enhances the clinical relevance of our data.  

      (7) Mechanistically, how does the loss of Tgr5 impact hematopoietic regeneration following sublethal irradiation?

      As delineated in the previous point, we have been seriously conditioned by cases of “swollen muzzles syndrome” (Garrett et al., 2018), which has stopped us from proceeding with more irradiation experiments for this particular study. Mechanistic studies are unfortunately beyond the scope of this manuscript, but the mechanistic basis for the relationship between BM adipocyte differentiation and hematopoiesis is now a focus for one of the involved laboratories and should produce follow-up manuscripts in the near future.

      (8) Only male mice were used throughout this study. It would be beneficial to know whether female mice show similar results.

      We thank Reviewer #1 for this question, as it led us to perform the characterization of steady-state hematopoiesis and morphological bone and BMAT evaluation of young female mice, yielding new findings that strengthen our manuscript. In summary, we have found that the decrease in BMAT is sexually dimorphic, with young females not showing reduced BMAT levels. Conversely, we did find a trend towards lower trabecular bone content in females. We present these data as part of Figure 3.

      Reviewer #2 (Public Review):

      Summary:

      In this manuscript, the authors examined the role of the bile acid receptor TGR5 in the bone marrow under steady-state and stress hematopoiesis. They initially showed the expression of TGR5 in hematopoietic compartments and that loss of TGR5 doesn't impair steady-state hematopoiesis. They further demonstrated that TGR5 knockout significantly decreases BMAT, increases the APC population, and accelerates the recovery upon bone marrow transplantation.

      Strengths:

      The manuscript is well-structured and well-written.

      We thank Reviewer #2 for this comment.

      Weaknesses:

      The mechanism is not clear, and additional studies need to be performed to support the authors' conclusion.

      We agree with Reviewer #2 that more studies are needed to understand the role of TGR5 in the hematopoietic system. We have been hampered in our studies of stress hematopoiesis because of frequent cases of swollen muzzles syndrome (Garrett et al., 2018), which prevented us from conducting additional experiments involving myelosuppression. Furthermore, the identification of the mechanism that links changes in the adipocyte differentiation axis and hematopoietic support, which is more complex than initially thought, has become a top priority for one of the involved laboratories and should produce follow-up manuscripts in the near future.

      Recommendations For The Authors:

      Reviewer #2 (Recommendations For The Authors):

      (1) Figure 1: the authors showed the presence of TGR5 in hematopoietic cells using a GFP report in mice. What's the expression pattern of TGR5 in the nonhematopoietic cells? For example, adipocytes, stromal cells, endothelial cells, etc. In addition to analyze the percentage of TGR5-GFP+ cells, the authors should also quantify the expression levels of TGR5 in various hematopoietic and niche components.

      We thank Reviewer #2 for this question, which we have addressed as points 1 and 3 from Reviewer #1. The expression of TGR5:GFP signal in stromal cells, HSCs, and the various hematopoietic compartments is now presented respectively as new panels in Figure 6A-B, Figure 1C-D, Figure 1–Figure Supplement 1A-B, and Figure 1figure supplement 2A, as well as Author response image 1 for stromal progenitors. Please note that specifically for the stromal compartment, we kindly requested Reviewer #1’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section, as we are concerned about over interpretation for these rare APC and multi-potent stem cell-like subpopulations.

      Consistent with the typical low-abundance expression of G protein-coupled receptors, where ligand-mediated activation is more relevant than absolute receptor abundance, TGR5 is also lowly expressed in most cell types. Although technical limitations prevent us from directly quantifying TGR5 expression in adipocytes (due to the difficulty of isolating these populations from bone marrow), previous studies have reported TGR5 expression in human BMSC-derived adipocytes (Velazquez-Villegas et al., 2018) . For similar reasons, we could not quantify TGR5 expression levels by RT-qPCR in the bone marrow niche, as isolating enough cells from each compartment to reliably detect a low-abundance receptor is technically challenging.

      Regarding endothelial cells, previous studies have shown TGR5 expression in vascular endothelial cells (Kida et al., 2013) as well. In our dataset the number of endothelial cells recovered was limited, largely due to the lack of a dedicated endothelial isolation protocol for the BM. Even so, the data we collected are shown as Author response image 4 and Author response table 2 and suggest that TGR5:GFP level in endothelial cells is lower than in the stromal and hematopoietic compartments. As for Author response image 1 and Author response table 1, we remain cautious about the interpretation of GFP expression in these low frequency stromal and endothelial populations and we kindly request the Reviewer’s opinion on whether these subpopulation analyses should be included in the main manuscript or solely here as part of the public review section.

      Author response image 4.

      Representative flow cytometry gating strategy used to identify endothelial cells in BM, defined as CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup> and their GFP positivity in TGR5:GFP mice.

      Author response table 2.

      Frequency of GFP-expressing cells in the CD45<sup>-</sup>Ter119<sup>-</sup>CD31<sup>+</sup> gate for endothelial cells presented in Author Response Figure 4, expressed as average (standard deviation). The number of mice per genotype and per experiment is indicated in parenthesis.

      (2) Figure 2: Regarding the competitive transplantation, the authors should bleed the mice every 4 weeks and show the dynamics of donor chimerism up to 16 or 20 weeks. The difference at the 3-week time point is very tiny and is this change significant? There is no statistical significance shown in Supplemental Figure 2 Panel C at a 3-week time point.

      The full data for repetitive monthly bleedings in primary, secondary and tertiary transplants is now shown in Figure 2J and Figure 2-figure supplement 1I-K. The statistically significant difference on the first month, which represents a 26% loss in short-time progenitor phenotype, is small but potentially clinically relevant as it is these progenitors that sustain early hematopoietic recovery from severe leucopenia and thrombocytopenia. Indeed, the hematopoietic phenotype described in Figure 7 C-J for accelerated rescue of Tgr5<sup>-/-</sup> recipients with wild-type bone marrow would be coherently associated to this short-term progenitor phenotype.  

      (3) Figure 3: in addition to the irradiation/transplantation model, the authors should also use alternative models for stress hematopoietic, for example, 5-FU, etc.

      This is an excellent suggestion, but unfortunately beyond the scope of this manuscript due to the move of one of the co-senior authors to another institution. Follow-up studies are however planned with ablative chemotherapy models relevant to the hematopoietic transplant setting (non 5-FU).

      (4) To get a better understanding of the role of TGR5 in the bone marrow niche, the authors should get and/or generate the TGR5 floxed mice and this will allow the authors to delete it from specific cell populations using cell-type specific Cre. In addition, it would be very interesting to investigate further the molecular mechanisms downstream of TGR5.

      We much appreciate this comment and the interest it reflects in TGR5 BM biology. Mechanistic studies are unfortunately beyond the scope of this manuscript, but the mechanistic basis for the relationship between BM adipocyte differentiation and hematopoiesis are now a focus for one of the involved laboratories and should produce concrete follow-up manuscripts soon.  

      References

      Ambrosi, T.H., Scialdone, A., Graja, A., Gohlke, S., Jank, A.-M., Bocian, C., Woelk, L., Fan, H., Logan, D.W., Schürmann, A., Saraiva, L.R., Schulz, T.J., 2017. Adipocyte Accumulation in the Bone Marrow during Obesity and Aging Impairs Stem Cell-Based Hematopoietic and Bone Regeneration. Cell Stem Cell 20, 771-784.e6. https://doi.org/10.1016/j.stem.2017.02.009

      Coutu, D.L., Kokkaliaris, K.D., Kunz, L., Schroeder, T., 2017. Three-dimensional map of nonhematopoietic bone and bone-marrow cells and molecules. Nat Biotechnol 35, 1202–1210. https://doi.org/10.1038/nbt.4006

      Craft, C.S., Robles, H., Lorenz, M.R., Hilker, E.D., Magee, K.L., Andersen, T.L., Cawthorn, W.P., MacDougald, O.A., Harris, C.A., Scheller, E.L., 2019. Bone marrow adipose tissue does not express UCP1 during development or adrenergic-induced remodeling. Sci Rep 9, 17427. https://doi.org/10.1038/s41598-019-54036-x

      Garrett, J., Sampson, C.H., Plett, P.A., Crisler, R., Parker, J., Venezia, R., Chua, H.L., Hickman, D.L., Booth, C., MacVittie, T., Orschella, C.M., Dynlachta, J.R., 2018. Characterization and Etiology of Swollen Muzzles in Irradiated Mice. Radiation Research 191, 31. https://doi.org/10.1667/RR14724.1

      Gil, L., Poplawski, D., Mol, A., Nowicki, A., Schneider, A., Komarnicki, M., 2013. Neutropenic enterocolitis after high-dose chemotherapy and autologous stem cell transplantation: incidence, risk factors, and outcome. Transpl Infect Dis 15, 1–7. https://doi.org/10.1111/j.1399-3062.2012.00777.x

      Gomariz, A., Helbling, P.M., Isringhausen, S., Suessbier, U., Becker, A., Boss, A., Nagasawa, T., Paul, G., Goksel, O., Székely, G., Stoma, S., Nørrelykke, S.F., Manz, M.G., Nombela-Arrieta, C., 2018. Quantitative spatial analysis of haematopoiesis-regulating stromal cells in the bone marrow microenvironment by 3D microscopy. Nat Commun 9, 2532. https://doi.org/10.1038/s41467-018-04770-z

      Gooley, T.A., Chien, J.W., Pergam, S.A., Hingorani, S., Sorror, M.L., Boeckh, M., Martin, P.J., Sandmaier, B.M., Marr, K.A., Appelbaum, F.R., Storb, R., McDonald, G.B., 2010. Reduced mortality after allogeneic hematopoietic-cell transplantation. N Engl J Med 363, 2091–2101. https://doi.org/10.1056/NEJMoa1004383

      Harrison, D.E., Jordan, C.T., Zhong, R.K., Astle, C.M., 1993. Primitive hemopoietic stem cells: direct assay of most productive populations by competitive repopulation with simple binomial, correlation and covariance calculations. Exp Hematol 21, 206–219.

      Janzen, V., Forkert, R., Fleming, H.E., Saito, Y., Waring, M.T., Dombkowski, D.M., Cheng, T., DePinho, R.A., Sharpless, N.E., Scadden, D.T., 2006. Stem-cell ageing modified by the cyclin-dependent kinase inhibitor p16INK4a. Nature 443, 421–426. https://doi.org/10.1038/nature05159

      Kida, T., Tsubosaka, Y., Hori, M., Ozaki, H., Murata, T., 2013. Bile acid receptor TGR5 agonism induces NO production and reduces monocyte adhesion in vascular endothelial cells. Arterioscler Thromb Vasc Biol 33, 1663–1669. https://doi.org/10.1161/ATVBAHA.113.301565

      Morowski, M., Vögtle, T., Kraft, P., Kleinschnitz, C., Stoll, G., Nieswandt, B., 2013. Only severe thrombocytopenia results in bleeding and defective thrombus formation in mice. Blood 121, 4938–4947. https://doi.org/10.1182/blood-2012-10-461459

      Naveiras, O., Nardi, V., Wenzel, P.L., Hauschka, P.V., Fahey, F., Daley, G.Q., 2009. Bonemarrow adipocytes as negative regulators of the haematopoietic microenvironment. Nature 460, 259–263. https://doi.org/10.1038/nature08099

      Purton, L.E., Dworkin, S., Olsen, G.H., Walkley, C.R., Fabb, S.A., Collins, S.J., Chambon, P., 2006. RARgamma is critical for maintaining a balance between hematopoietic stem cell self-renewal and differentiation. J Exp Med 203, 1283–1293. https://doi.org/10.1084/jem.20052105

      Purton, L.E., Scadden, D.T., 2007. Limiting factors in murine hematopoietic stem cell assays. Cell Stem Cell 1, 263–270. https://doi.org/10.1016/j.stem.2007.08.016

      Sebo, Z.L., Rendina-Ruedy, E., Ables, G.P., Lindskog, D.M., Rodeheffer, M.S., Fazeli, P.K., Horowitz, M.C., 2019. Bone Marrow Adiposity: Basic and Clinical Implications. Endocr Rev 40, 1187–1206. https://doi.org/10.1210/er.2018-00138

      Suchacki, K.J., Tavares, A.A.S., Mattiucci, D., Scheller, E.L., Papanastasiou, G., Gray, C., Sinton, M.C., Ramage, L.E., McDougald, W.A., Lovdel, A., Sulston, R.J., Thomas, B.J., Nicholson, B.M., Drake, A.J., Alcaide-Corral, C.J., Said, D., Poloni, A., Cinti, S., Macpherson, G.J., Dweck, M.R., Andrews, J.P.M., Williams, M.C., Wallace, R.J., Van Beek, E.J.R., MacDougald, O.A., Morton, N.M., Stimson, R.H., Cawthorn, W.P., 2020. Bone marrow adipose tissue is a unique adipose subtype with distinct roles in glucose homeostasis. Nat Commun 11, 3097. https://doi.org/10.1038/s41467-020-16878-2

      Tratwal, J., Bekri, D., Boussema, C., Sarkis, R., Kunz, N., Koliqi, T., Rojas-Sutterlin, S., Schyrr, F., Tavakol, D.N., Campos, V., Scheller, E.L., Sarro, R., Bárcena, C., Bisig, B., Nardi, V., de Leval, L., Burri, O., Naveiras, O., 2020. MarrowQuant Across Aging and Aplasia: A Digital Pathology Workflow for Quantification of Bone Marrow Compartments in Histological Sections. Front Endocrinol (Lausanne) 11, 480. https://doi.org/10.3389/fendo.2020.00480

      Trivedi, M., Martinez, S., Corringham, S., Medley, K., Ball, E.D., 2009. Optimal use of G-CSF administration after hematopoietic SCT. Bone Marrow Transplant 43, 895–908. https://doi.org/10.1038/bmt.2009.75

      Vannini, N., Campos, V., Girotra, M., Trachsel, V., Rojas-Sutterlin, S., Tratwal, J., Ragusa, S., Stefanidis, E., Ryu, D., Rainer, P.Y., Nikitin, G., Giger, S., Li, T.Y., Semilietof, A., Oggier, A., Yersin, Y., Tauzin, L., Pirinen, E., Cheng, W.-C., Ratajczak, J., Canto, C., Ehrbar, M., Sizzano, F., Petrova, T.V., Vanhecke, D., Zhang, L., Romero, P., Nahimana, A., Cherix, S., Duchosal, M.A., Ho, P.-C., Deplancke, B., Coukos, G., Auwerx, J., Lutolf, M.P., Naveiras, O., 2019. The NAD-Booster Nicotinamide Riboside Potently Stimulates Hematopoiesis through Increased Mitochondrial Clearance. Cell Stem Cell 24, 405-418.e7. https://doi.org/10.1016/j.stem.2019.02.012

      Velazquez-Villegas, L.A., Perino, A., Lemos, V., Zietak, M., Nomura, M., Pols, T.W.H., Schoonjans, K., 2018. TGR5 signalling promotes mitochondrial fission and beige remodelling of white adipose tissue. Nat Commun 9, 245. https://doi.org/10.1038/s41467-017-02068-0

      Walkley, C.R., Fero, M.L., Chien, W.-M., Purton, L.E., McArthur, G.A., 2005. Negative cellcycle regulators cooperatively control self-renewal and differentiation of haematopoietic stem cells. Nat Cell Biol 7, 172–178. https://doi.org/10.1038/ncb1214

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public reviews):

      Weaknesses:

      Reliance on self-reports

      A primary limitation of this study, acknowledged by the authors, is its reliance on self-reports of participants’ emotional states. Although considerable effort was made to minimize expectation effects, further research is needed to confirm that the observed behavioral changes reflect genuine alterations in emotional states. Additionally, the generalizability of the findings to long-term remediation strategies remains an open question.

      We agree with this characterisation and have strengthened the corresponding acknowledgment in the Discussion. We would also note that, while self-report measures are inherently subjective, the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ”genuine” emotional state that supersedes self-report currently exists. We have added a sentence to the Discussion making this point explicit: ”While emotional self-reports are inherently subjective, we note that the regularities governing subjective emotional experience are of primary scientific interest in their own right, and no consensus definition of a ‘genuine’ emotional state that supersedes self-report currently exists.”

      Additionally, we agree that what we have described is limited to a short-term intervention and change. Whether these changes bear on longer-term changes remains to be assessed. Furthermore, the mechanisms or processes that would support such a maintenance are of substantial interest, and will be the focus of future work.

      Statistical analysis and interpretation of the dynamics matrix

      Second, the statistical analysis, particularly the computational approach, sometimes lacks sufficient detail and refinement. While I will not elaborate on specific points here, one notable issue is the interpretation of the intrinsic matrix (A). The model-free analysis reveals correlations between emotions at a given time or within an emotional state across time points. However, it does not provide evidence to support lagged interactions across states that would justify non-diagonal elements in A. The other result concerning the dynamics matrix only highlights a trend in the dominant eigenvalue, which is difficult to interpret in isolation. The absence of a statistically significant group x intervention interaction furthermore makes this finding a little compelling. This weakens the study’s conclusions about the importance of intrinsic dynamics, as claimed in the title.

      We thank the reviewer for raising this important methodological point. We address it in three parts.

      (i) Justification for the full dynamics matrix. In response to this comment, we have added a diagonal-constrained variant of A to the model comparison. The full A model was selected by BIC over the diagonal-constrained version, meaning that cross-emotion lagged interactions contributed to model fit over and above what could be explained by individual emotion autocorrelation alone, thereby justifying the inclusion of off-diagonal elements. We have updated the model comparison results section and figure caption accordingly.

      (ii) Interpretation of the dominant eigenvalue. We agree that the dominant eigenvalue alone is difficult to interpret. Our primary claim regarding intrinsic dynamics rests on model comparison (which identified a change in A as necessary to explain post-intervention data in the distancing group) together with the change in the direction of the dominant eigenvector (tested via Hotelling T<sup>2</sup>, p = 0.019). The eigenvalue magnitude is reported as a complementary, interpretable summary of the stability shift.

      (iii) General analysis of the dynamics matrix. We have added a visualization of the full dynamics matrix A to the Supplementary Materials, reporting per-emotion diagonal elements (persistence) and their variability across participants. This shows that the emotion calm exhibited the highest persistence, consistent with our hypothesis, whereas sadness showed lower-than-expected stability, and other emotions exhibited low mean persistence with substantial individual variability. We have added the following sentence to the Results: ”A general analysis of the dynamics matrix A revealed that while calm exhibited the highest persistence, consistent with our hypothesis (0.59 ± 0.85), sadness did not show the expected stability (0.09 ± 0.58), and other emotions demonstrated low persistence (0.13 to 0.24) alongside substantial individual variability.”

      Terminology (controllability, stability, sensitivity, Gramian)

      Finally, to avoid potential misunderstandings of their work, the authors should be more careful about their use of terms pertaining to the control theory and take the time to properly define them. For example, the ”controllability” of emotional states can either denote that those states are more changeable (control theory definition), or, conversely, more tightly regulated (common interpretation, as used in the abstract). This is true for numerous terms (stability, sensitivity, Gramian, etc.) for which no clear definition nor references are provided. Readers unfamiliar with the framework of control theory will likely be at a loss without more guidance.

      This is an excellent and important point. We have made the following changes throughout the manuscript.

      (i) Terminology table. We have added a new Table 1 in the Methods (Conceptual Definitions) providing a side-by-side mapping of the formal control-theory definition and the psychological interpretation for the three key terms: Controllability, Stability, and Sensitivity.

      (ii) Abstract. The abstract previously used ”controllability” in a way that could be read in the psychological sense. We have replaced the sentence referencing the Controllability Gramian with a clarified version describing the measure as ”continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”

      (iii) Introduction. We have added a paragraph explicitly distinguishing the control-theoretic definition of controllability from the psychological concept of emotion control or emotion regulation, noting that high controllability in the formal sense does not imply tight regulation or reduced variability. Definitions of sensitivity and stability are also provided.

      (iv) Discussion. Occurrences of ”controllability” that could be ambiguous are annotated with brief clarifications of which sense is intended.

      (v) Gramian. We now use the ”controllability matrix” (C) throughout, and its SVD-based analysis is described using ”left singular vectors” rather than ”eigenvectors of the Gramian”.

      Reviewer #2 (Public reviews):

      Online recruitment, selection effects, and remuneration

      Acquiring data online inevitably gives rise to selection and self-selection effects. This needs to be acknowledged clearly. Exacerbating this, participant remuneration seems low at an amount below the minimum or living wage in Western countries (do the authors know where their participants came from?).

      We thank the reviewer for raising this. All participants were recruited from the UK via Prolific Academic and reimbursed at £7.50/h. Remuneration rates were comparable to other experimental settings, in keeping with other online studies, UK living wage recommendations, and ultimately determined according to institutional ethical guidance. We acknowledge that online recruitment via Prolific may introduce self-selection effects and that the sample may not be representative of the general population. We have added a corresponding statement to the Limitations section: ”Online recruitment via Prolific Academic may introduce self-selection effects, and the sample may not be representative of the general population.”

      Intervention ongoing during the second block

      Another concern is that the intervention does not simply take place before the second block begins but is ongoing during the whole of the second block in that it is integrated into the phrasing of the task on each trial. It is therefore somewhat misleading to speak of a period ’after the intervention’, and it would have been interesting to assess the effect of this by including a third group where the phrasing does not change, but the floating leaves intervention takes place.

      This is a valid and important observation. In the distancing group, the trial-by-trial question phrasing during the second block included a reminder, meaning the intervention was reinforced on every trial. We have acknowledged this as a procedural difference and potential confound in the Limitations section, noting that this reminder may have encouraged a form of retrospective reappraisal rather than in-the-moment distancing. The design choice was intentional (the reminder was included to encourage continued application of the strategy, mirroring how such techniques are deployed in practice) but we agree it is an imperfect feature of the current design and have discussed it openly. We have added to Limitations: ”Relatedly, a procedural difference existed in that only the distancing group received a strategy reminder during the second video block. While intended to facilitate real-time regulation, this prompt may have inadvertently encouraged retrospective reappraisal to align with the distancing narrative.”

      Observation noise

      As mentioned in the Limitations section, observation noise was assumed and not estimated. While this is understandable in this case, the effect of this assumption could have been assessed by simulation with varying levels of observation (and process) noise.

      We would like to clarify that both observation noise (Γ) and process noise (Σ) were in fact estimated from the data, constrained to be diagonal. We have extended the parameter recovery analysis (Supplementary Section: Parameter Recovery) in which, for each of 104 subjects and both time periods (N = 208 observations), we generated 100 surrogate trajectories from the fitted model and re-estimated all parameters. Recovery quality is quantified via Pearson correlations between true and recovered parameters (A, C, Σ, and Γ), alignment of dominant eigenvectors and left singular vectors, and bias analysis for dominant eigenvalue and singular value magnitudes. The analysis confirms that A and C matrix parameters and the critical eigenvector-based metrics all exceed our target threshold of r = 0.7. Noise covariances are poorly recovered, consistent with typical Kalman filter behaviour on short time series. We have expanded the Limitations to note that the Gaussian observation noise assumption does not fully capture the bounded nature of 0–100 rating scales, and suggest that future work could use a truncated or censored observation model.

      Reliance on formal model comparison

      Relatedly, the reliance on formal model comparison is unfortunate since the outcome of such comparisons is easily influenced by slight changes to assumptions such as noise levels. An alternative approach would have been to develop a favoured model based on its suitability to address the research question and its ability, established by simulation, to distill relevant changes of behaviour into reliable parameter estimates.

      We appreciate this methodological concern but would argue that formal model comparison is wellsuited to our research question. Our central aim is not simply to fit emotion trajectories, but to determine which components of the dynamical system (intrinsic dynamics (A), input weights (C), or both) are altered by the distancing intervention. This is inherently a model comparison question: without comparing models that do and do not allow each component to change, no principled inference about the mechanism of action is possible. A single favoured model, however well-motivated, would presuppose the answer.

      We also note that the reviewers concern, that outcomes are sensitive to underlying assumptions, applies equally to the favoured-model approach: simulations rely on predefined structures and noise specifications that shape parameter recovery, and a misspecified favoured model risks confounding the parameters intended to capture the intervention effect with artefacts of model structure. By contrast, our approach evaluates a principled, nested set of models of increasing complexity, using BIC to penalise unnecessary parameters, which guards against overfitting. The models are intentionally simple (linear, Gaussian, time-invariant) to limit the degrees of freedom available to absorb noise.

      Critically, the approach is not model comparison alone. We followed established best-practice procedures for computational modelling, including posterior predictive checks (simulated trajectories closely matched observed data; Fig. 4C), and parameter recovery establishing that the key metrics are reliably recovered (see Supplementary Materials F Parameter Recovery). We have also added a Random-Effects Bayesian Model Selection (RFX-BMS) to characterise individual heterogeneity in model preferences. Together, these provide converging evidence for the validity of our inference. We have clarified this reasoning in the revised manuscript.

      Statistical limitations; Bayesian inference

      The statistical analyses clearly show the limitations of classical statistical testing with highly complex models of the kind the authors (commendably) use. Hunting for statistically significant interactions in a multivariate repeated-measures design relying on inputs from time series- derived point estimates is a difficult proposition. While the authors make the best of the bad 3 situation they create by using null-hypothesis significance testing, a more promising approach would have been to estimate parameters using a sampler like Stan or PyMC and then draw conclusions based on posterior predictive simulations.

      We agree that fully Bayesian parameter estimation via Stan or PyMC would be a valuable methodological advance. Implementing this for 104 subjects across two time periods with a 5-dimensional Kalman filter is, however, a substantial undertaking beyond the scope of the current revision. In the interim, the RFX-BMS analysis added in response to Comment 4 above provides a group-level Bayesian perspective on model uncertainty that partially addresses this concern.

      Reviewer #3 (Public reviews):

      Dual meanings of controllability

      An interesting but perhaps at present slightly confusing aspect of their described results relates to the ’controllability’ of emotions, which they define as their susceptibility to external inputs. Readers should note this definition is (as I understand it) quite distinct from, and sometimes even orthogonal to, concepts of emotional control in the emotion literature, which refer to intentional control of emotions (by emotion regulation strategies such as distancing). The authors also use this second meaning in the discussion. Because of the centrality of control/controllability (in both meanings) to this paper, at present it is key for readers to bear these dual meanings in mind for juxtaposed results that distancing ”reduces controllability” while causing ”enhanced emotional control”

      We are grateful for this observation, which echoes Reviewer 1’s concern about terminology. We have addressed this comprehensively; please see our response to Reviewer 1 Comment 3 above.

      Strategy reminder and possible reappraisal

      As above the authors use an active control – a relaxation intervention – which is extremely closely matched with their active intervention (and a major strength). However, there was an additional difference between the groups (as I currently understand it): ”in the group allocated to the distancing intervention, the phrasing of the question about their feelings in the second video block reminded participants about the intervention, stating: ”You observed your emotions and let them pass like the leaves floating by on the stream.” I do wonder if the effects of distancing also have been partially driven by some degree of reappraisal (considered a separate emotion regulation strategy) since this reminder might have evoked retrospective changes in ratings.

      This is a well-founded concern. As noted in our response to Reviewer 2 Comment 2, we have added an explicit acknowledgment of this procedural difference and the potential for retrospective reappraisal to the Limitations section. We note, as the reviewer themselves observe in the Strengths section, that demand effects are unlikely to account for the specific pattern of dynamic changes observed: uniform demand effects would be expected to produce flat reductions across emotions, whereas our findings show emotion-specific changes in eigenmode structure and controllability direction. Nevertheless, a partial contribution of reappraisal cannot be ruled out from the current design.

      Mechanism of distancing effects (eye movement and oculomotor avoidance)

      Not necessarily a weakness, but an unanswered question is exactly how distancing is producing these effects. As the authors point out, there is a possibility that eye-movement avoidance of the more emotionally salient aspects of scenes could be changing participants’ exposure to the emotions somewhat. Not discussed by the authors, but possibly relevant, is the literature on differences between emotion types on oculomotor avoidance, which could have contributed to differential effects on different emotions.

      We thank the reviewer for raising the oculomotor avoidance hypothesis. Research suggests that different emotions elicit distinct patterns of gaze behaviour: disgust is associated with visual avoidance, whereas anxiety and other negative emotions show increased attentional bias following fear conditioning. These emotion-specific oculomotor patterns could have contributed to the differential effects we observe on the input weight matrix C. What would be particularly interesting to examine in future work is whether a distancing intervention induces multiple, emotionally-specific gaze behaviours, or a single undifferentiated avoidance response. We have expanded the Limitations to: ”[...] The literature on emotion-specific oculomotor avoidance suggests that gaze patterns differ across emotion categories, which could contribute to differential effects on the input weight matrix for specific emotions. [...]”

      Recommendations for the Authors:

      Reviewer #1 (Recommendations for the Authors):

      (1) In the procedure description, the authors suggest that some emotions (e.g. disgust) would be more volatile and stimulus driven, while others (eg. sad) would be more stable. Is this hypothesis reflected in the model-based state dynamics, typically in the diagonal elements of A?

      Yes, we added a supplementary figure showing the dynamics matrix A, which confirms that calm exhibited the highest persistence, consistent with our hypothesis. Disgust showed lower persistence, also in line with this expectation. However, sadness did not display the expected stability and instead showed relatively low persistence, contrary to our hypothesis.

      (2) Could the authors elaborate on the emotional space covered by the chosen ratings? If the axes are positive-negative and slow-fast, why 5 and not 4?

      We added a sentence in the Methods clarifying that five emotions were selected based on the specific affective qualities of the video stimuli provided in the validated databaseto and to better capture the high-dimensional, nuanced states elicited by the stimuli rather than to fit a traditional four-axis model.

      (3) The whole methods section crucially lacks references. As an example, the whole derivation of the most important metrics (eigenvalues of the Gramian, energy ellipse, etc) leaves the reader completely on its own.

      References have been added throughout the Methods, including for the controllability matrix, SVDbased analysis, and eigendecomposition.

      (4) Before equation 1, when introducing x and u: a) time appears twice (typo), and b) the 1Tˆ notation is not standard (especially without bold) and unclear until way below when the one-hot encoding is mentioned.

      The typo has been corrected.

      (5) Why use one hot-encoding rather than the original ratings from the video database? Videos must vary if not in spread (as suggested in Figure B.1) at least in intensity. Ignoring this variance surely diminishes the accuracy of the modelling.

      We used one-hot encoding so that the input weight matrix C can directly estimate participantspecific intensity and sensitivity, rather than fixing input magnitudes to database averages. Using database ratings would assume that the emotional intensity of each video clip generalizes perfectly to our sample; any mismatch would be absorbed as error in C, potentially biasing parameter estimates and obscuring individual differences in emotional reactivity, which are central to our analysis.

      (6) Could the authors develop the rationale behind the bias in Equation 1?

      We added a sentence clarifying that the bias term h captures the steady-state baseline of the emotional system, i.e. the mean rating toward which emotions converge in the absence of external inputs.

      (7) While I can understand why the authors included a set of models with a diagonal C matrix, I do not see why they did not do the same with A. While the diagonal elements are necessary to persist emotional states and induce some autocorrelation in the ratings, as observed empirically, the influence of the non-diagonal elements is not justified (and Figure 5G seems to confirm that). This is critical as, in the end, the controllability metrics will highly depend on those non-diagonal elements which remain very obscure throughout the manuscript.

      We now included both diagonal and full variants of A in the model comparison; still the full A was selected by BIC.

      (8) Concerning the model comparison, I am not sure what the authors mean by using the BIC at the group level. Did they just sum them across participants? This approach is known to be highly susceptible to outliers and cannot be relied upon in general. So-called ”random effect analysis” tends to be regarded as the gold standard and can be easily implemented by taking - 0.5 * BIC as an approximation to the model evidence. Such an approach would also allow to properly test for group differences (cf. Rigoux et al. 2014).

      We have added a Random-Effects Bayesian Model Selection analysis as a new Supplementary Section; see Comment 4 of the public review response above.

      (9) The sentence ”proportion of the total amount of predictive power provided by the full set of models contained in the model being assessed” does not make any sense to me. Please rephrase.

      The sentence has been rephrased: ”Cumulative model weights (w<sub>j</sub>) normalize raw BIC scores so they can be interpreted as the relative probability that a specific model is the best one among the set being compared:”

      (10) ”the largest eigenvalue of the dynamics matrix A identifies the most stable combination of emotions” is only true if the eigenvalues are below 1, which is not granted.

      We have added the qualifier that this holds provided the dominant eigenvalue lies within the unit circle (|λ| < 1), indicating a system that converges to a steady state.

      (11) Equation 3 does not define the Gramian but the controllability matrix, a confusion that goes through the manuscript. The Gramian is formally defined as W = P(AkBB′A′k). Luckily, for discrete systems, it can be approximated by CC′ and therefore the singular values of C can be used to approximate the eigenvalues of W, which are the usual metrics used to define the energy ellipse and so on. While the results reported in the manuscript are correct (up the the approximation), the general description is wrong or misleading.

      We thank the reviewer for this important correction; we now consistently refer to C as the ”controllability matrix” throughout, and its SVD-based analysis is described in terms of left singular vectors rather than eigenvectors of the Gramian.

      (12) Figure 2: what is the matrix V? If it’s from the singular value decomposition W = USV, then (if I am not mistaken) the direction of the ellipsoid is defined by U. Again, a reference would help.

      We have corrected the figure and caption to refer to ”left singular vectors” throughout. We retain the variable name V rather than adopting the standard SVD convention of U to avoid confusion with the input vector u, which appears throughout the model equations.

      <(13) Correction for multiple comparisons is mentioned as a way to correct for the number of conducted tests. However, later on, some post-hoc analyses are reported with the mention that the correction is done across emotions (so p/5), while multiple tests are run for each emotion. This is critical when all the pairs across the cells of an ANOVA are tested and no correction seems to be applied, which is inducing a huge risk of false positives.

      Along the same line: the correct way to demonstrate the effect of the intervention is to first do an ANOVA to reveal an interaction between group and time, and then only to do post-hoc tests to pinpoint where the interaction is coming from, and not the other way around as reported in the manuscript. Further, a difference in significance is not equivalent to a significant difference, and showing that a time effect is significant in one group but not in the other does not imply that the intervention differs between groups, only testing the interaction can confirm this.

      We have added reporting of the significant group × time interaction effects (F(5,208) = 2.6, p = 0.026 for mean ratings; F(5,200) = 2.5, p = 0.03 for the most controllable direction) and flagged these in the figure captions. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      (14) The notation DV = b0 + b1IV*b2G is confusing as a full model (interaction + main effects) should have 3 parameters in addition to the intercept. Also, why use different models for testing the main effect and the interaction?

      The regression equation has been corrected to DV = β<sub>0</sub> + β<sub>1</sub>IV + β<sub>2</sub>G + β<sub>3</sub>(IV × G) + ϵ, making the interaction term explicit.

      (15) Figure 4: the control subject in panel C seems to rate close to 0 in all emotions except for ”calm”. How was this subject fitted? Does model selection (at the subject level) correctly identify a change of dynamics in this case? I don’t see how the behaviour after the intervention could be realistically fitted with a 65-parameters dynamical system. What type of checks were operated to ensure the quality of the fit beyond the recovery analysis (see below)?

      The top participant in Figure 4C was from the distancing group and the participant rating close to zero on most emotions except calm shown at the bottom was from the control group. For the distancing participant, model selection correctly identified a change in dynamics and input weights (BIC = 4806) over the same-parameters model (BIC = 4821). For the control participant, model selection similarly favoured a change in dynamics and input weights (BIC = 3723) over the sameparameters model (BIC = 4011), suggesting that the relaxation intervention produced a comparable effect on emotional dynamics to the distancing intervention. Note that these two participants were selected randomly to illustrate the visual quality of model fit (i.e. that simulated trajectories closely resemble the empirical rating curves) and are not intended to be representative examples of group differences.

      Regarding the data-to-parameter ratio: the 65 parameters are estimated from 55 observations per emotion per block, giving a more favourable ratio than it might appear.

      Beyond the visual trajectory overlays in Figure 4C, we have now added R<sup>2</sup> and peak cross-correlation as a quantitative measure of individual fit quality. Mean R<sup>2</sup> across all subjects and emotions was 0.6 and mean temporal correlation was r = 0.74−0.80, confirming that both the timing and magnitude of emotional responses were well reproduced. Notably, for the specific control participant shown in Figure 4C, R<sup>2</sup> values were 0.74, 0.83, 0.81, 0.83, and 0.83 for disgusted, amused, calm, anxious, and sad respectively, confirming that even for this visually striking participant the model fit was adequate across all five emotion dimensions.

      (16) What do the authors mean by ”eigenmodes”? In the following sentence, what does ”This component” refer to?

      Eigenmodes is defined in the Stability section as the independently evolving combinations of emotions obtained by projecting the state vector onto the eigenvectors of A, and ”This component” has now an explicit referent.

      (17) Figure 5: see above the comment about the necessity for testing the interactions, which should also be reported in the figures.

      Interaction effects are now included in the relevant figure caption (Figure 5).

      (18) When looking for the relationships between questionnaires and controllability, looking only at the most controllable direction seems rather inefficient due to the multiple comparisons correction. Why not compute the angle (or other measure of similarity) with an ideal ”calm” unit vector?

      Along the same line, it’s unlikely that the most controllable direction will smoothly rotate as a function of symptoms. More likely, the winning (highest eigenvalue) direction will switch from one to another, creating some discontinuity in the summary statistic used for the correlation with clinical scores. How could one work around this issue?

      This is an interesting suggestion; we have not implemented it in the current revision, but we acknowledge it as a promising analysis for future work.

      (19) More generally, it would be interesting to see if there are some regularities in the dynamics across participants. If this is the case, one could construct a canonical emotional dynamics and project all participants on this eigenspace. Emotional trajectories would then differ only in their controllability (eigenvalues) in this common space, making a comparison across participants more straightforward. Could the authors comment on this?

      This is a valuable suggestion that we have not pursued in the current revision, as constructing a common eigenspace across participants requires additional methodological choices.

      (20) I am a bit puzzled by the hypothesis that questionnaires should mediate the intervention effect. Shouldn’t questionnaires be related to before-intervention controllability only? Similarly, could one use initial controllability to predict the intervention response (irrespective or not of the clinical score)?

      Our hypothesis was that participants with greater difficulties in emotion regulation (high DERS-18) would show a smaller intervention effect, as their emotional system might be less amenable to brief distancing. However, this was not the case. We also note that DERS-18 scores were not significantly related to the overall magnitude of pre-intervention controllability (norm of the controllability matrix), but were related to its direction: participants with higher DERS-18 scores showed a most controllable direction pointing toward disgust and away from amusement and calmness, suggesting that trait-level regulation difficulties are linked to the specific emotional configuration of the system rather than its overall controllability.

      Regarding using initial controllability to predict the intervention response: we agree this is a mechanistically appealing question, but it is unfortunately not straightforward to address here. The intervention effect would be quantified as the change between pre- and post-intervention controllability, and since pre-intervention controllability is a constituent of that change score, any correlation between the two would be partially circular by construction.

      (21) The recovery procedure should be way more detailed. How many surrogates were run, etc.? Which kind of quality checks were used to ensure the recoverability was sufficient at the subject level, especially as parameter recovery seems relatively low for some subjects?

      The parameter recovery section now reports that 100 surrogate trajectories were generated per subject (N> = 104) per time period (before and after intervention: N = 208 observations total), with per-subject recovery quality reported across simulations together with across-subject variability; see Supplementary Materials F Parameter Recovery.

      (22) As the result of the eigen decomposition is the endpoint of the analysis, it would be a nice addition to test the recoverability of those measures (eg. correlation between simulated and inferred eigenvalues).

      Recovery of the dominant eigenvalue/vector and dominant singular value/left singular vector is now reported explicitly, including bias analysis and scatter plots of true versus recovered values. Beyond subject-level recovery, we also assessed whether the observed group differences in emotional dynamics and controllability could be reliably recovered at the group level. See Supplementary Materials F Parameter Recovery.

      (23) Although this comment comes close to last, this is a major concern of mine. I am not convinced that the recoverability procedure is sufficient to prove that the inference is working. The model assumes that the observation noise is Gaussian, which is clearly not the case in the data. By simulating surrogate time series with normally distributed noise, the authors do not account for any saturating effects that could destroy a large part of the behavioural information necessary for a successful inversion (eg Figure 4.C showing that simulated data contains a lot of negative ratings). A workaround would be to bind the surrogate time series to mimic the saturation caused by the rating scale, and then run the model estimation on those capped time series.

      We acknowledge this important limitation: the Gaussian assumption does not capture the bounded 0–100 scale, and we have added this to the Limitations with a suggestion that future work use a truncated or censored observation model.

      (24) Supplementary tables with placeholders (v1, v2) that can vary in meaning depending on the line are extremely hard to decipher.

      The supplementary tables have been restructured to a hierarchical format.

      (25) Table I9 is not referenced in the manuscript.

      A reference to this table has been added in the appropriate Results section.

      (26) It’s a shame that neither data nor analysis code has been made available.

      Fully anonymised data and analysis code are now publicly available on GitHub (https://github.com/huyslab/emotioncon public).

      Reviewer #2 (Recommendations for the Authors):

      (1) Abstract: By some definitions, controllability is binary, present or absent, according to whether the controllability Gramian is positive definite. Mention that you use a continuous definition, otherwise ’quantified’ leads to confusion.

      The abstract now describes the measure as: ”Controllability was assessed using continuous measures derived from the controllability matrix, capturing relative differences in how emotional states can be driven by external inputs.”, making the non-binary usage explicit.

      (2) p 5: ’on [not in] the recruiting platform’.

      Corrected.

      (3) p 8: Clarify notation of x<sub>t</sub>t = 1<sup>T</sup> (and same for u). What does this mean?

      The notation has been corrected.

      (4) p 8: Define h in the paragraph following Eq 1, don’t wait until the next section.

      The definition of the bias h (steady-state baseline) has been moved to immediately after Equation 1.

      (5) p 9: Give clear references for your methods here. There are different definitions of controllability etc. than the ones you use.

      References have been added at each key definition.

      (6) p 28: ’20 videos per emotion category were chosen resulting in 50 videos per sequence’ - doesn’t make sense.

      This has been clarified: 20 videos per category across 5 categories yield 100 videos in total. These were split into two matched sequences of 50 videos each. Including 2 videos that were repeated twice resulted in 54 videos per sequence and 108 videos in total.

      (7) Figure B2: 54 videos are listed, and categorized into five categories. How does this relate to the remark right above?

      A clarifying note has been added to the supplementary explaining that each sequence of 54 clips includes repeated videos and is drawn from the pool of 100, with emotion-category sequences matched between blocks.

      (8) p 29: What were the process noise Σ and observation noise Γ assumed in the parameter recovery exercise? What were the consequences of that assumption as assessed by simulation?

      Both Σ and Γ were estimated from the data constrained to be diagonal; the parameter recovery section now explicitly reports their recovery quality.

      (9) p 30: How are results affected by including the excluded participants? The level required to pass attention checks seems arbitrary. How was it chosen?

      We have added a sensitivity analysis including the four excluded outliers, showing results remained qualitatively and statistically similar. With 10 binary attention checks, chance performance is 50%, meaning a participant scoring below 70% is performing only marginally above chance and likely not attending consistently. At the same time, 70% is permissive enough to retain participants who may have missed one or two checks due to momentary distraction.

      (10) Figure F4: Typographically distinguish capital letters referring to panels in the figure from those referring to matrices.

      Panel labels in Figure F4 are formatted in bold to distinguish them from italicised matrix notation.

      (11) Table G1: Showing that differences between groups were non-significant before the intervention but significant after is not enough, you need to show that there was a significant interaction between time point and intervention. [I wrote this after reading the supplementary but before reading the main text. It turns out you know what I’m telling you here. You should mention it more prominently though, including in the abstract and the discussion, because in your chosen null-hypothesis significance testing framework, this is the crucial test of your study. I don’t think there’s any harm at all in being up-front about this - certainly much better than making excuses like the one about randomization at the top of page 11, which I recommend removing].

      This is very right. The significant group × time interaction effects are now reported prominently in the main text Results and figure captions. We also wish to be transparent: the interaction tests were conducted post-hoc rather than as the primary analysis, which is the reverse of the correct order. In the interest of transparency we have left the structure of the results section intact rather than retrospectively reframing it.

      Reviewer #3 (Recommendations for the Authors):

      (1) I would encourage the authors to re-add some basic details regarding their power analyses from the supplement to the main text so the reader can immediately reference the intended effect size, which analysis/analyses were considered primary for the power analysis, etc.

      More information about the power analysis have been added to the Participants section of the main text.

      (2) Similarly, I wondered if there could be a little extra information on how the test-retest reliability was calculated (page 13) - on the first and last views of a video pre-intervention? I wasn’t sure - why do the authors only present ICCs for amusement/joy and disgust/horror? Seems useful to present all (particularly because there may be individual differences in habituation to some emotions).

      We have expanded the test-retest section to clarify that reliability was assessed using duplicate videos shown three times pre-intervention, and now report ICCs with confidence intervals, Cronbach’s α, and habituation/sensitization tests for both disgust and amusement; the selection of these two categories reflects which videos were repeated in the design, they were chosen at random during experimental design.

      (3) Regarding my comments about controllability/emotional control, I think the authors probably have two choices - address this head-on (e.g. with a note describing the relationship/distinction between these two concepts of controllability), or else avoid using it in one of the senses (I would suggest the mathematical sense since overriding the concept of cognitive control seems harder - the authors could use phrases like ’impact of emotional inputs’ instead of ’controllable’). In particular, the abstract could be clearer about the nature of controllability as implemented by the authors - this seems critical for communicability. Because of the high relevance of both of these ’control’ concepts to the paper, if the authors agree with my concern, I would also suggest changes throughout, such as in the results section phrasing: ”In those participants with high DERS-18 scores, the most controllable direction pointed towards disgust ( = 0.26, p = 0.006), and away from amusement ( = 0.26, p = 0.005) and calmness ( = 0.24, p = 0.011; though this did not survive Bonferroni correction).”

      Thank you for this comment. Please see our response to the public review comments above, which we hope address this.

      (4) Regarding the control intervention, which is great, is it possible the follow-up question/reminder affected the results – e.g., is there reason to believe that prospective regulation was the primary difference in the distancing group and not a retrospective effect via this question?

      See our response to Reviewer 2 public review Comment 2 and the corresponding Limitations addition.

      (5) Lastly, purely for interest, the authors could consider elaborating on their brief interpretation as to why difficulties in regulating emotions were specifically linked to the controllability of disgust, amusement, and calmness, but not other emotions (anxiety/sadness) (page 20). I wonder if there is a brief space to discuss the emotional specificity of these results further given the relevance to the wider literature on specific emotion types, e.g. fear vs disgust.

      We have expanded the Discussion to elaborate on the emotional specificity of these findings. Amusement and disgust are strongly influenced by external events, suggesting that stimulus-driven controllability is particularly relevant for these emotions. By contrast, anxiety and sadness are maintained through internally generated processes such as rumination and anticipatory cognition, exhibiting greater emotional inertia over time, and their regulation may therefore be less sensitive to momentary stimulus controllability. This provides a mechanistic account of why controllability effects emerged selectively for disgust, amusement, and calmness, and aligns with growing evidence that emotion regulation is emotion-specific rather than domain-general.

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This important study partially fills the gap in the knowledge of olfaction at the level of the Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) with functional magnetic resonance imaging, electrophysiology, and modeling. The methods used are convincing. Some of the findings confirm ongoing hypotheses, such as the behavioral importance of AON for odor source discrimination. Other results shed light on the dynamics of the connection between the olfactory system and the rest of the brain.

      We sincerely thank the editors and reviewers for the thorough review of our manuscript. We appreciate the insightful comments, which have significantly contributed to improving our work. In this revision, we addressed all the concerns posed by the reviewers, including conducting additional analyses and providing the data generated and codes used in this study.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript combined rat fMRI, optogenetics, and electrophysiology to examine the large-scale functional network of the olfactory system as well as its alteration in an aged rat model.

      Strengths:

      Overall methodology is very solid and the results provided an interesting perspective on large-scale functional network perturbation of the olfactory system.

      Weaknesses:

      The biological relevance and validation of the current results can be improved.

      We thank the reviewer for the comments and suggestions regarding our manuscript. They have been invaluable and instrumental in enhancing our work. Please see R1-1 to R1-8 below for our corresponding responses and revisions.

      Comments:

      (R1-1) Figure 1A, on the top of the figure, ChR2 may be replaced by ChR2-mCherry, as only mCherry is fluorescent. And also, it’s somewhat surprising that in AON and Pir regions (where only axon fibers should be labelled as red), most fluorescence appeared dot-like and looked more similar to cell body instead of typical fiber. The authors may want to double-check this.

      (a) Thank you for pointing out the missing fluorescent marker for ChR2 expression in the label of Figure 1A. We distinguished the transfection of ChR2 in olfactory bulb (OB) neurons from axonal terminals in AON and Pir by identifying the expression patterns in these three regions (see Figure 1). In OB, the cellular nucleus (i.e., labelled in blue by DAPI) was surrounded evenly by ChR2-mCherry expression (red fluorescent, indicated by green arrows in Figure 1), which clearly traced the neuronal cell body shape. Meanwhile, the expression of mCherry in AON and Pir consisted of concentrated red fluorescent dots that were sparsely assembled near the cell nucleus, indicating the presence of synaptic boutons at the axonal terminals. Further, the patterns of expression here did not trace the shape of the neuronal cell body.

      (b) In this revision, we amended the label of Fig. 1A to “ChR2-mCherry”.

      In this revision, changes were made to the confocal images of AON and Pir, based on Figure 1, for clarity and the figure caption was also edited accordingly.

      (R1-2) The authors primarily presented 1 Hz stimulation results. What is the most biologically relevant frequency (e.g., perhaps firing frequency under natural odor stimulation) among all frequencies that were used?

      (a) There are various firing rate and oscillatory frequency of neural activities within the olfactory system (i.e., OB, Pir) in rodents during olfaction. For example in OB, neural oscillations ranging from 1 to 12 Hz (i.e., slow to theta) are driven by sensory stimulation and are closely linked to respiration [1]. Specifically, odor-evoked responsive excitatory bursts of mitral and tufted cells in OB are coupled with the low-frequency respiration rhythm (1-4 Hz) [2,3]. Such coupling has also been documented during light anaesthesia [4]. Meanwhile, higher frequencies such as beta oscillations (15 ~ 30 Hz) have been associated with odor learning and sensitization [5,6], and gamma oscillations (40 ~ 80 Hz) evoked by sensory stimulation are associated with fine olfactory discrimination and odor learning [6-8].

      The piriform cortex (Pir) also exhibits natural oscillatory activity across various frequency bands, including slow, theta, beta, and gamma oscillations [9-11]. As in OB mitral and tufted cells, slow oscillations (< 1.5 Hz) in Pir are also correlated with the respiratory rhythm, indicating the intrinsic respiration-related oscillatory properties within the olfactory system [12]. However, while it has been shown that the primary burst of firing in the anterior Pir is locked to respiration, its coding strategies differed from OB during olfaction [13].

      Despite not being able to entirely mimic the evoked neural activities under natural odor stimulation due to the synchronizations induced by our optogenetic stimulations, our frequencies were chosen to be within the range of firing rates of neurons in the olfactory system under natural circumstances. Our selection of 1 Hz frequency for optogenetic stimulation was based on the balance of experimental simplicity of inducing neural excitation and the biological significance of low-frequency neural oscillations in OB, Pir and AON. The robustness of evoked brain-wide long-range activations is one of our criteria for the selection of 1 Hz stimulation frequency at OB, considering that the goal of this study is to examine long-range olfactory networks. Furthermore, we also examined other frequencies of stimulation covering theta, beta and gamma frequency bands, which provided insights into the response characteristics within the olfactory networks.

      (b) In this revision, we added a brief statement in the Results section that 1 Hz was the most biologically relevant frequency based on discussions in R1-2a above.

      (R1-3) In Figure 2, the statistical thresholding is confusing: in the figure legend, it was stated that “t > 3.1 corresponding to P < 0.001” but later “further corrected for multiple comparisons with thresholdfree cluster enhancement with family-wise error rate (TFCE-FWE) at P < 0.05”? Regardless of the statistical thresholding, such BOLD activation seemed to be widespread (almost whole-brain activation). Does such activation remain specific to the optogenetic stimulation, or something more general (e.g., arousal level change)? Furthermore, how those results (I assume they are group-level results) were obtained was not described very clearly. Is it just a simple average of individual-level results, or (more conventionally) second-level analysis?

      (a) We thank the reviewer for drawing our attention to clarity issues regarding the generation of group-level BOLD activation maps and statistical thresholding methods used.

      In brief, we first generate the BOLD activation map of individual animal through conventional general linear model (GLM) analysis. These individual activation maps then underwent a two-step statistical thresholding method to generate the group-level results. Firstly, uncorrected one-sample t-tests were conducted with a threshold of P < 0.001. Secondly, these thresholded activation maps were further corrected using nonparametric inference with threshold-free cluster enhancement multiple comparison correction of family-wise error rate (TFCE-FWE, P < 0.05) before they were averaged. Hence, the BOLD activation maps presented in our manuscript were the result of group-level analysis instead of individual-level analysis.

      (b) We note the reviewer’s concern about whether the observed widespread BOLD activations were caused by the animal’s general brain state (e.g., arousal levels) rather than the specificity of optogenetic stimulation. First, such widespread activations were only specific to 1 Hz stimulation of OB neurons and OB afferents at AON (Figs. 2A, B and 3A, B) whereas activations were localized to regions in the primary olfactory network with increasing stimulation frequencies. Furthermore, varied neural activity adaptation properties were observed (Fig. 3A, B) following repeated 1 Hz optogenetic excitation of OB neurons and OB afferents at AON indicating that such widespread propagation of evoked neural activity was not driven by the animal’s general brain state, which would be relatively random across animals as we interleaved the presentation of each stimulation frequencies. While we cannot discount that certain characteristics of brain states differ under anaesthetized and awake conditions, we showed that light (1.0% isoflurane) anaesthesia minimally affects the BOLD fMRI activations and/or the propagation of optogenetically-evoked neural activity across thalamo-cortical, hippocampal-cortical and vestibulo-cortical networks [14-20].

      Second, the neural representations in OB have been shown to be brain state-independent to ensure the high fidelity of olfactory inputs to higher olfactory cortices, despite differences in spontaneous baseline activity across states [21-24]. In fact, OB neural activity (i.e., slow to delta oscillations) remains highly coupled to respiration rhythms under low anaesthesia as in awake conditions [4]. Further, low-dose isoflurane (i.e., 1% as in the present study) does not affect cortical gamma oscillations (40 Hz) [25-28]. Although odorant-evoked responses in olfactory cortices (e.g., Pir and olfactory tubercle, Tu) can be modulated by overall brain state, various interactions between such cortical circuits to process odor inputs persist under anaesthesia [29-31].

      (c) In this revision, we clarified the two-step statistical thresholding method that was applied to generate the group-level BOLD activation maps, as discussed in R1-3a above, in the captions of Fig. 2 and the Methods section of the manuscript.

      In this revision, we also included a statement in the Discussion section that the widespread BOLD fMRI activations were specific to the 1 Hz optogenetic stimulation and not the animal’s general brain state, as discussed in R1-3b above.

      (R1-4) In Figure 2, why use AUC to quantify the activation, not the more conventional beta value in the GLM analysis?

      (a) We thank the reviewer for raising this concern. We are aware of the more conventional beta value, b, in GLM analysis and have used them for quantification and comparison of the amplitudes of fMRI activations in our previous rodent fMRI studies [18,32,33]. However, in this study, we chose to utilize area under curve (AUC) as it offers a more comprehensive measure of BOLD signal change over time, including shape, duration, and magnitude, thereby capturing the bulk of neural activities and their dynamics throughout the stimulation period. b primarily represents the peak amplitude of BOLD responses (i.e., the % BOLD signal change) [34] and can be constrained by the assumptions and limitations of the GLM analysis, such as the shape of the canonical hemodynamic response function (HRF). Therefore, AUC provides greater accuracy in capturing different aspects of neural responses across various brain regions, such as transient peaks and/or sustained responses.

      (b) In this revision, we have included the justifications for using AUC over the more conventional b value to quantify BOLD fMRI activations, as in R1-4a above, in the Methods section.

      (R1-5) For Figure 2D, the way that it was quantified can be better described as “relative” activation within one condition, and I don’t know how to interpret the comparison among the relative fraction of activated regions. Perhaps comparison using percentage change (i.e., beta values) is more straightforward.

      (a) We thank the reviewer’s comment and suggestion here. “BOLD activation strength” is the AUCs normalized to the respective sum of BOLD activations of their respective stimulation target. We are of the opinion that “strength” can accurately reflect our purpose for conducting this comparison, which is to distinguish the brain networks that were predominantly recruited by AON- or Pir-driven neural activities compared to OB stimulation. 

      (b) We conducted a comparison using relative fraction for fair comparison considering the differences in absolute BOLD signal amplitudes upon different stimulation targets. As we found the overall amplitudes of activations evoked by Pir stimulation were appreciably weaker, our normalization step ensures that the contribution of Pir is not underrepresented. Without this normalization process, the Pir stimulation seems not to activate most of the downstream targets based on the relatively weak activations. As discussed earlier in R1-4a above, comparison using percentage change/b value would be skewed towards the absolute amplitude of BOLD responses and would be erroneous for subsequent interpretation of brain networks that are primarily recruited by OB vs. AON vs. Pir.

      (b) In this revision, we chose not to include additional statements as we have already described in the Results section the reasons for normalizing AUC to reflect and subsequently compare the BOLD activation strength across various networks that were recruited by optogenetic stimulation of three distinct olfactory regions (i.e., OB, AON and Pir). 

      (R1-6) For Figure 3, it may be more convenient for readers to include the results of 1st activation for direct comparison. The current layout makes it difficult to make direct, visual comparisons among all 3 activations. Again, I think using beta values (instead of AUC) may be more conventional.

      (a) We agree that it will be more convenient for readers if we include the 1st activation maps in Figure 3 for direct comparison. Please refer to R1-4 for the usage of AUC rather than beta values.

      (b) In this revision, we modified Fig. 3 to include activation maps from the 1st fMRI session for easy comparison with the 2nd and 3rd sessions.

      (R1-7) Can the DCM results (at least part of it) be verified using the current electrophysiological data? For example, the long-range inhibitory effective connectivity of AON is rather intriguing. If that can be verified using the electrophysiology data, it would be really great. In the current form, the DCM and electrophysiology results seem to be totally unrelated.

      (a) We thank the reviewer for raising this concern, and it’s a great suggestion to causally link the outcomes from dynamic causal modeling (DCM) analysis and electrophysiology findings.

      (b) In principle, the recorded local field potentials (LFPs) can be treated as the neural dynamic component in DCM analysis under specific conditions. They are as follows:

      (1) LFPs are recorded at the identical brain state as the optogenetic fMRI experiments with identical stimulation paradigms.

      (2) The electrode locations of LFP recordings must match the nodes of the a priori matrix defined for the existing DCM analysis (Fig. 4A).

      However, in the present study, we did not have LFP recordings at the entorhinal cortex (Ent), which will likely influence the modeling of effective connectivity strength as one of the nodes defined is dropped from the analysis. Previous studies have showed that changes in the a priori matrix will result in different connectivity estimations [35-38]. We chose to model the connectivity involving Ent due to its documented interactions with the primary olfactory cortices and hippocampal regions [39,40]. Hence, it would be incomplete without modelling the contributions from Ent.

      (c) In this revision, we decided not to redefine the a priori matrix by removing Ent as a node in the DCM analysis to estimate new effective connectivities. Instead, we acknowledge the absence of Ent recordings.

      (R1-8) In Figure 6, it would be great if the adaptation of BOLD and electrophysiology signals can be correlated at the brain region level. The current figure only demonstrated there is adaptation in the electrophysiology recording, but did not show if such adaptation is related to the BOLD adaptation.

      (a) We thank the reviewer for pointing out that we should associate the electrophysiology and fMRI data to validate the olfactory adaptation. As such, we conducted additional correlation analysis between LFP power and BOLD signal profiles. The method for computing the correlation coefficient is described below:

      (1) For each 1 Hz optogenetic stimulation target (i.e., OB, AON, and Pir), regions of interest (ROIs) that have both LFP recordings and BOLD activations were selected. These ROIs were AON, Pir, ventral caudate putamen (vCPu), visual cortex (V1), and ventral hippocampus (vHP) for OB stimulation; OB, Pir, vCPu, V1, amygdala (Amg), and vHP for AON stimulation; and OB, AON, vCPu, V1, Amg, and vHP for Pir stimulation. Note that we excluded the respective stimulated region as a ROI because of BOLD fMRI signal dropout caused by the implanted optical fibre.

      (2) We first calculate the individual animal AUC differences of 2nd stimulation session vs. 1st session and 3rd session vs. 1st session in LFP power and BOLD signal profiles, respectively.

      Subsequently, the mean of the differences for each ROI across animals was then calculated.

      (3) Correlation coefficients and P values of AUC differences between LFP power and BOLD signal profiles were computed using Spearman nonparametric correlation.

      (b) The relationship between the differences of LFP power and BOLD signal profiles showing neural adaptation was described by the correlation coefficient, r, and the corresponding P value. The r and P values upon the three distinct stimulations were OB: r = 0.77 (P = 0.01), AON: r = 0.46 (P = 0.13), and Pir: r = - 0.06 (P = 0.85), respectively. This indicates that the decrease of LFP power is highly correlated with the decrease of BOLD activations upon OB and AON stimulation compared to those upon Pir stimulation.

      (c) In this revision, we added the description of the correlation analysis conducted, as in R1-8a above, in the Methods section.

      In this revision, we also added several statements in the Results section to indicate that neural adaptation observed is significantly related to the decreased BOLD activations when stimulating OB excitatory neurons or OB afferents at AON, as discussed in R1-8b above. Additionally, we also included the outcomes of the correlation analysis as a new Supplementary Fig. 7.

      Reviewer #2 (Public review):

      Summary:

      Ma and colleagues presented a study on the characterization of brain-wide spatio-temporal impact of olfactory cortical outputs. They take advantage of multi-modal techniques on rats: fMRI, optogenetics, and electrophysiology. In addition, they used cutting-edge analytical techniques and modeling to support and interpret their data. The main findings of the study are:

      (1) The neurons in the Olfactory Bulb (OB) predominantly activate primary olfactory network regions, while stimulation of OB afferents in Anterior Olfactory Nucleus (AON) and Piriform Cortex (Pir) primarily orthodromically activates hippocampal/striatal and limbic networks, respectively.

      (2) Non-specified adaptation or habituation mechanisms may play a significant role in modulating olfactory outputs over subsequent fMRI sessions.

      (3) Artificially induced aging in rats induces profound modification in the functional interaction between olfactory cortices and multiple brain regions.

      The results on AON are of particular interest because of the lack of functional information on this region, despite its recognized importance in shaping OB output and behavior (odor localization tasks).

      Strengths:

      The manuscript is very accurate. The figures are well-crafted, and clear and provide much information with the most appropriate plots and graphics. The study’s amount and data quality are remarkable, and the experimental size adequately addresses the scientific questions. I particularly appreciated the details in the description of the methods regarding the missing data and the size of the different animal groups. The supplementary data complete the leading figures and provide information at a single animal level.

      We are grateful for the reviewer’s appreciation of our study’s breadth and depth, and the quality and quantity of our data. The recognition of the strengths in our methods and the inclusion of single-animal-level data in the supplementary information is highly encouraging. We are also appreciative for the reviewer’s constructive comments below, which has help us to address specific weaknesses of the study and thereby improving the manuscript text and flow. We have read through the comments carefully and have made the necessary corrections. We hope that our responses to R2-1 to R2-11 below has sufficiently addressed all the reviewer’s concerns.

      Weaknesses:

      (R2-1) One of the main reasons the Piriform Cx is understudied in rodents is because of the proximity to air, which creates artifacts in fMRI images. This issue becomes more critical at ultra-high magnetic fields, but I would expect it also at 7T. One main achievement of this study is, indeed, the acquisition of fMRI data from Piriform, and this point should be highlighted by showing raw functional data from a rat. The best would be if an fMRI data sample for a rat, no matter which stimulation, is shared on a public repository, like Zenodo or similar. I am curious to check the quality of the BOLD data from such an ‘enormous’ field of view, particularly in the OB, with a single-shot sequence. Also, the visual inspection of raw data is essential to appreciate how many 0.5 x 0.5 x 1 mm voxels fit into AON, and others analyzed small brain structures, like the amygdala, etc. Was the amygdala entirely visible in BOLD, or did the air in the ear channel make an artifact partially shadowing it?

      (a) As per the reviewer’s suggestion, we displayed the representative EPI images after preprocessing at a matrix size of 128 x 128 with a pixel size of 0.25 x 0.25 mm in Supplementary Figure 2. Up-sampling was not applied out-of-plane. Note that the atlas-based region-of-interest (ROIs) used to extract BOLD signal profiles (i.e., indicated by colored overlays in Supplementary Figure 2) were drawn based on the up-sampled EPI images (from the acquired 64 x 64 to 128 x 128). Notably, the piriform cortex (Pir) and amygdala (Amg) are visible with no appreciable signal dropout at these regions. To clarify, the Pir in our study represented the anterior Pir.

      Signal dropout is pronounced in the posterior ventral regions of the brain (e.g., Bregma- 4 mm, below the Ent) and as expected in a localized region where the implanted optical fiber region that targeted the AON.

      (b) In this revision, we included as a new Supplementary Figure 8 and added a statement in the Results section to describe the absence of appreciable signal dropouts at OB, Pir and Amg, which could affect subsequent quantitative analyses.

      (R2-2) Surprisingly, the only information missing in the methods is the post-surgery period and the time between two consecutive fMRI sessions. How much time was accorded to rats to recover from the surgeries, and what time interval between two scans? This information is crucial for interpreting the decrease in most BOLD responses in subsequent recordings. The supposed adaptation should fit into the known time frames for odor adaptation. Usually, fast adaptation does not last for days (and it should be measured within a single experiment: is it the case?), while for long-lasting adaptation the stimulus (odor or opto) should be maintained constantly ON. This does not seem to be the case in this study. The hypothesis, alternative to adaptation, of a less efficient light activation, for example, due to gliosis around the fiber tips, should be discarded with more evidence than the preservation of OB > Pir responses or acknowledged in the manuscript.

      (a) We thank the reviewer for drawing our attention to the missing information regarding the timeline of each optogenetic fMRI experiment. We conducted fMRI experiments immediately after the surgical procedure for implanting the optical fiber cannula. Such an implantation procedure typically last for 45mins under 1.2-2.0% isoflurane. The animal is then moved from the surgical table to the magnet to begin the optogenetic fMRI experiment. Intervals between two fMRI scans were within one minute. As the stimulation frequencies (i.e., 1 Hz, 5 Hz, 10 Hz, 20 Hz and 40 Hz) were pseudorandomized, the interval between any two identical stimulation frequency was on average 30 minutes. In this case, we are measuring fast adaptation (on the order of tens of minutes) rather than long-lasting adaptation. Further, as the fMRI experiment was conducted immediately following fiber implantation, the risk of gliosis around the fiber tips is expected to be low.

      (b) The main representation of olfactory adaptation is the attenuation of neural responses upon continuous or repeated stimuli [41,42]. Several fMRI studies in rodents have reported that olfactory adaptation was detected in primary olfactory cortices (i.e., OB, AON, and Pir) upon repeated odor stimulation [43,44]. Specifically, these studies showed robust decreased responses in primary olfactory cortices under repeated odor stimulation with an odor interval of 30 minutes, which were comparable to those observed in our study upon repeated 1 Hz optogenetic stimulation. As such, the pronounced neural adaptation that we demonstrated when stimulating OB excitatory neurons or OB afferents at AON is unlikely to be caused by less efficient light activation.

      (c) In this revision, we added statements in the Methods section, as discussed in R2-2a above, to clarify the timeline of our optogenetic fMRI experiment.

      (R2-3) The D-galactose experiments were conducted only after administering the aging molecule, with no baseline/reference data on the same animals. Then, comparisons were made with healthy rats, but the two groups not only can be discriminated with respect to D-galactose administration but also with age (10 VS 18 weeks). A control group for 18-weeks-old rats with no D-galactose treatment would better compare the D-galactose effect and avoid any potential bias from group comparisons of rats at different ages. Do you confirm that D-galactose was injected into each rat 56 times/day in a row, or am I mistaken?

      (a) We appreciate the reviewer’s comment here and understand the concerns regarding the absence of an age-matched control group with saline administration instead of D-galactose. We acknowledge that the inclusion of a control group would have provided a more direct comparison to evaluate the differences between healthy and aged animals. However, we were unable to include this group in the present study due to logistical constraints.

      To clarify, our experiments were conducted on healthy rats at the median age of 14 weeks to ensure the animals had reached adulthood. The median age of the D-galactose injected animal was 18 weeks. In terms of the lifespan of rats, they reach sexual maturity at approximately 6 weeks of age [45], at which point we conducted the optogenetic viral vector injection. The age of rats at 14 and 18 weeks can both be regarded as teenage adult [45,46]. The primary purpose of the experiments on the aged animal model and corresponding analyses was to provide some potential clues into the dysfunction of olfactory networks at the system level as animals aged. Despite the age difference between the two groups, we believe that the comparison between healthy and aged rats still offers valuable insights into the ageing-related dysfunctions within the olfactory system. 

      (b) Note that D-galactose was injected into each rat 56 times in total over the whole experiment (i.e., one injection per day over the course of 8 weeks).

      (c) In this revision, we acknowledge that the comparisons made were not with age-matched healthy controls in the Discussion section.

      In this revision, we also clarified that D-galactose was administered once daily for 8 weeks (i.e., a total of 56 injections) in the Methods section.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Minor:

      (R2-4) The look of the activation maps will greatly improve if the displayed brain slices are coronal (as in the supplementary figures) with no angle.

      We thank the reviewer for this suggestion. The purpose of such a display was to maintain the consistency of displaying the overlay of BOLD fMRI activation maps on the 3D renders of rat brain anatomical images. In addition, such 3D display can better impress readers that the optogenetically evoked BOLD activations were indeed long-range and brain-wide. 

      As such, in this revision, we decided to maintain the existing displays of BOLD activation maps in the main figures.

      (R2-5) The choice of testing different stimulation frequencies deserves more justification. Why was it performed in the OB, which is naturally activated on the breathing rhythm?

      (a) We thank the reviewer’s comment here to seek further clarification. OB is widely recognized as the first stage of olfactory information processing in the brain. By activating the OB excitatory neurons, we can ensure that our optogenetically-evoked activations are olfactory-related.

      The choice of various frequencies ranging from low (i.e., 1 Hz) to high (e.g., 40 Hz) was made to cover a range of firing rate and oscillatory frequency of neural activities neural oscillations that exist within the rodent olfactory system during olfaction. For example in OB, neural oscillations ranging from 1 to 12 Hz (i.e., slow to theta) are driven by sensory stimulation and are closely linked to respiration [1]. Specifically, odor-evoked responsive excitatory bursts of mitral and tufted cells in OB are coupled with the low-frequency respiration rhythm (1-4 Hz) [2,3]. Such coupling has also been documented during light anaesthesia [4]. Meanwhile, higher frequencies such as beta oscillations (15 ~ 30 Hz) have been associated with odor learning and sensitization [5,6], and gamma oscillations (40 ~ 80 Hz) evoked by sensory stimulation are associated with fine olfactory discrimination and odor learning [6-8].

      (b) In this revision, we added a statement in the Results section to further describe our justifications for testing various stimulation frequencies.

      (R2-6) The absence of natural (odor) stimulation should be acknowledged as a limitation for the findings of this study in the Discussion.

      We thank the reviewer for the suggestion. In this revision, we made the acknowledgement in the Discussion section. 

      (R2-7) Line 46: overstatement, the role of the Piriform cortex at the system level is well known and was not discovered by this study. (see Gordon Shepherd’s book: The Synaptic Organization of the Brain, Figure 10.6)

      As per the reviewer’s suggestion, we decided to use “distinguishes” rather than “uncovers” for an accurate description of our study’s findings, which demonstrated the differences between the piriform cortex and anterior olfactory nucleus in driving the downstream activations of olfactory and non-olfactory targets.

      (R2-8) Line 75: please rephrase for an easier understanding.

      As per the reviewer’s suggestion, we edited this statement to “Further, the need to present a multitude of odor combinations also makes it challenging to efficiently and reliably interrogate long-range olfactory networks and their properties”.

      (R2-9) Line 313: the intermediary region between the Piriform and Entorhinal cortices is the Perirhinal Cortex. No doubt about that.

      As per the reviewer’s comment, we modified the corresponding statement by adding “such as the perirhinal cortex” and cited the appropriate references [47,48].

      (R2-10) Line 391: there is no evidence from this study to exclude that the downstream targets would be less activated because of a decreased OB activation in favor of broad inhibition.

      As per the reviewer’s suggestion, we modified the statement by replacing “rather than” with “in addition to” for a more precise statement.

      (R2-11) Line 585: which was/were the regressor/s used in GLM? Any convolution with common HRFs?

      We regarded the block-designed optogenetic stimulation as a task. The regressors are then obtained by convolving the expected task-evoked neural activity (i.e., boxcar function for block design) with the canonical hemodynamic response function, HRF (SPM12, Wellcome Department of Imaging Neuroscience, University College London, UK).

      In this revision, we added a statement in the Methods section to provide clarification on the regressors used in GLM and the convolution step with canonical HRF.

      References

      (1) Kay, L. M. & Stopfer, M. Information processing in the olfactory systems of insects and vertebrates. Semin Cell Dev Biol 17, 433-442 (2006). https://doi.org/10.1016/j.semcdb.2006.04.012

      (2) Cang, J. & Isaacson, J. S. In vivo whole-cell recording of odor-evoked synaptic transmission in the rat olfactory bulb. J Neurosci 23, 4108-4116 (2003). 

      (3) Margrie, T. W. & Schaefer, A. T. Theta oscillation coupled spike latencies yield computational vigour in a mammalian sensory system. J Physiol 546, 363-374 (2003). https://doi.org/10.1113/jphysiol.2002.031245

      (4) Fontanini, A. & Bower, J. M. Variable coupling between olfactory system activity and respiration in ketamine/xylazine anesthetized rats. Journal of Neurophysiology 93, 3573-3581 (2005). https://doi.org/10.1152/jn.01320.2004

      (5) Kay, L. M. et al. Olfactory oscillations: the what, how and what for. Trends in Neurosciences 32, 207-214 (2009). https://doi.org/10.1016/j.tins.2008.11.008

      (6) Gervais, R., Buonviso, N., Martin, C. & Ravel, N. What do electrophysiological studies tell us about processing at the olfactory bulb level? Journal of physiology, Paris 101, 40-45 (2007). https://doi.org/10.1016/j.jphysparis.2007.10.006

      (7) Beshel, J., Kopell, N. & Kay, L. M. Olfactory bulb gamma oscillations are enhanced with task demands. J Neurosci 27, 8358-8365 (2007). https://doi.org/10.1523/JNEUROSCI.119907.2007

      (8) Martin, C., Beshel, J. & Kay, L. M. An olfacto-hippocampal network is dynamically involved in odor-discrimination learning. J Neurophysiol 98, 2196-2205 (2007). https://doi.org/10.1152/jn.00524.2007

      (9) Kay, L. M. Theta oscillations and sensorimotor performance. Proc Natl Acad Sci U S A 102, 3863-3868 (2005). https://doi.org/10.1073/pnas.0407920102

      (10) Lowry, C. A. & Kay, L. M. Chemical factors determine olfactory system beta oscillations in waking rats. J Neurophysiol 98, 394-404 (2007). https://doi.org/10.1152/jn.00124.2007

      (11) Vanderwolf, C. H. & Zibrowski, E. M. Pyriform cortex beta-waves: odor-specific sensitization following repeated olfactory stimulation. Brain Res 892, 301-308 (2001). https://doi.org/10.1016/s0006-8993(00)03263-7

      (12) Fontanini, A., Spano, P. & Bower, J. M. Ketamine-xylazine-induced slow (< 1.5 Hz) oscillations in the rat piriform (olfactory) cortex are functionally correlated with respiration. J Neurosci 23, 7993-8001 (2003). https://doi.org/10.1523/JNEUROSCI.23-22-07993.2003

      (13) Miura, K., Mainen, Z. F. & Uchida, N. Odor representations in olfactory cortex: distributed rate coding and decorrelated population activity. Neuron 74, 1087-1098 (2012). https://doi.org/10.1016/j.neuron.2012.04.021

      (14) Leong, A. T. et al. Long-range projections coordinate distributed brain-wide neural activity with a specific spatiotemporal profile. Proc Natl Acad Sci U S A 113, E8306-E8315 (2016). https://doi.org/10.1073/pnas.1616361113

      (15) Chan, R. W. et al. Low-frequency hippocampal–cortical activity drives brain-wide resting-state functional MRI connectivity. Proc Natl Acad Sci U S A 114, E6972-E6981 (2017). https://doi.org/10.1073/pnas.1703309114

      (16) Leong, A. T. L. et al. Optogenetic fMRI interrogation of brain-wide central vestibular pathways. Proceedings of the National Academy of Sciences 116, 10122-10129 (2019). https://doi.org/10.1073/pnas.1812453116

      (17) Leong, A. T. L., Wang, X., Wong, E. C., Dong, C. M. & Wu, E. X. Neural activity temporal pattern dictates long-range propagation targets. Neuroimage 235, 118032 (2021). https://doi.org/10.1016/j.neuroimage.2021.118032

      (18) Leong, A. T. L., Wong, E. C., Wang, X. & Wu, E. X. Hippocampus Modulates Vocalizations Responses at Early Auditory Centers. Neuroimage 270, 119943 (2023). https://doi.org/10.1016/j.neuroimage.2023.119943

      (19) Wang, X. et al. Functional MRI reveals brain-wide actions of thalamically-initiated oscillatory activities on associative memory consolidation. Nat Commun 14, 2195 (2023). https://doi.org/10.1038/s41467-023-37682-8

      (20) Xie, L. et al. Brain-wide resting-state fMRI network dynamics elicited by activation of single thalamic input. Nat Commun 16, 11247 (2025). https://doi.org/10.1038/s41467-025-66104-0

      (21) Li, A., Gong, L. & Xu, F. Brain-state-independent neural representation of peripheral stimulation in rat olfactory bulb. Proc Natl Acad Sci U S A 108, 5087-5092 (2011). https://doi.org/10.1073/pnas.1013814108

      (22) Lang, J. et al. Odor representation in the olfactory bulb under different brain states revealed by intrinsic optical signals imaging. Neuroscience 243, 54-63 (2013). https://doi.org/10.1016/j.neuroscience.2013.03.057

      (23) Chery, R., Gurden, H. & Martin, C. Anesthetic regimes modulate the temporal dynamics of local field potential in the mouse olfactory bulb. J Neurophysiol 111, 908-917 (2014). https://doi.org/10.1152/jn.00261.2013

      (24) Wachowiak, M. et al. Optical dissection of odor information processing in vivo using GCaMPs expressed in specified cell types of the olfactory bulb. J Neurosci 33, 5285-5300 (2013). https://doi.org/10.1523/JNEUROSCI.4824-12.2013

      (25) Hudetz, A. G., Vizuete, J. A. & Pillay, S. Differential effects of isoflurane on high-frequency and low-frequency gamma oscillations in the cerebral cortex and hippocampus in freely moving rats. Anesthesiology 114, 588-595 (2011). https://doi.org/10.1097/ALN.0b013e31820ad3f9

      (26) Joliot, M., Ribary, U. & Llinas, R. Human oscillatory brain activity near 40 Hz coexists with cognitive temporal binding. Proc Natl Acad Sci U S A 91, 11748-11751 (1994). https://doi.org/10.1073/pnas.91.24.11748

      (27) Murthy, V. N. & Fetz, E. E. Oscillatory activity in sensorimotor cortex of awake monkeys: synchronization of local field potentials and relation to behavior. J Neurophysiol 76, 3949-3967 (1996). https://doi.org/10.1152/jn.1996.76.6.3949

      (28) Tallon-Baudry, C., Bertrand, O., Delpuech, C. & Pernier, J. Stimulus specificity of phaselocked and non-phase-locked 40 Hz visual responses in human. J Neurosci 16, 4240-4249 (1996). https://doi.org/10.1523/JNEUROSCI.16-13-04240.1996

      (29) Murakami, M., Kashiwadani, H., Kirino, Y. & Mori, K. State-dependent sensory gating in olfactory cortex. Neuron 46, 285-296 (2005). https://doi.org/10.1016/j.neuron.2005.02.025

      (30) Wilson, D. A. & Yan, X. Sleep-like states modulate functional connectivity in the rat olfactory system. J Neurophysiol 104, 3231-3239 (2010). https://doi.org/10.1152/jn.00711.2010

      (31) Schreck, M. R. et al. State-dependent olfactory processing in freely behaving mice. Cell Rep 38, 110450 (2022). https://doi.org/10.1016/j.celrep.2022.110450

      (32) Gao, P. P., Zhang, J. W., Chan, R. W., Leong, A. T. L. & Wu, E. X. BOLD fMRI study of ultrahigh frequency encoding in the inferior colliculus. Neuroimage 114, 427-437 (2015). https://doi.org/10.1016/j.neuroimage.2015.04.007

      (33) Gao, P. P., Zhang, J. W., Fan, S. J., Sanes, D. H. & Wu, E. X. Auditory midbrain processing is differentially modulated by auditory and visual cortices: An auditory fMRI study. Neuroimage 123, 22-32 (2015). https://doi.org/10.1016/j.neuroimage.2015.08.040

      (34) Goddard, E. & Mullen, K. T. fMRI representational similarity analysis reveals graded preferences for chromatic and achromatic stimulus contrast across human visual cortex. Neuroimage 215, 116780 (2020). https://doi.org/10.1016/j.neuroimage.2020.116780

      (35) Friston, K. J., Harrison, L. & Penny, W. Dynamic causal modelling. Neuroimage 19, 12731302 (2003). https://doi.org/10.1016/s1053-8119(03)00202-7

      (36) Friston, K. J., Kahan, J., Biswal, B. & Razi, A. A DCM for resting state fMRI. Neuroimage 94, 396-407 (2014). https://doi.org/10.1016/j.neuroimage.2013.12.009

      (37) Bernal-Casas, D., Lee, H. J., Weitz, A. J. & Lee, J. H. Studying Brain Circuit Function with Dynamic Causal Modeling for Optogenetic fMRI. Neuron 93, 522-532 e525 (2017). https://doi.org/10.1016/j.neuron.2016.12.035

      (38) Zeidman, P. et al. A guide to group effective connectivity analysis, part 1: First level analysis with DCM for fMRI. Neuroimage 200, 174-190 (2019). https://doi.org/10.1016/j.neuroimage.2019.06.031

      (39) Salimi, M. et al. Disrupted connectivity in the olfactory bulb-entorhinal cortex-dorsal hippocampus circuit is associated with recognition memory deficit in Alzheimer's disease model. Sci Rep 12, 4394 (2022). https://doi.org/10.1038/s41598-022-08528-y

      (40) Chen, Y. N., Kostka, J. K., Bitzenhofer, S. H. & Hanganu-Opatz, I. L. Olfactory bulb activity shapes the development of entorhinal-hippocampal coupling and associated cognitive abilities. Curr Biol 33, 4353-4366 e4355 (2023). https://doi.org/10.1016/j.cub.2023.08.072

      (41) Pellegrino, R., Sinding, C., de Wijk, R. A. & Hummel, T. Habituation and adaptation to odors in humans. Physiol Behav 177, 13-19 (2017). https://doi.org/10.1016/j.physbeh.2017.04.006

      (42) Sinding, C. et al. New determinants of olfactory habituation. Sci Rep 7, 41047 (2017). https://doi.org/10.1038/srep41047

      (43) Zhao, F. et al. fMRI study of olfaction in the olfactory bulb and high olfactory structures of rats: Insight into their roles in habituation. Neuroimage 127, 445-455 (2016). https://doi.org/10.1016/j.neuroimage.2015.10.080

      (44) Zhao, F. et al. fMRI study of the role of glutamate NMDA receptor in the olfactory adaptation in rats: Insights into cellular and molecular mechanisms of olfactory adaptation. Neuroimage 149, 348-360 (2017). https://doi.org/10.1016/j.neuroimage.2017.01.068

      (45) Sengupta, P. The Laboratory Rat: Relating Its Age With Human's. Int J Prev Med 4, 624-630 (2013). 

      (46) Quinn, R. Comparing rat’s to human’s age: How old is my rat in people years? Nutrition 21, 775-777 (2005). https://doi.org/10.1016/j.nut.2005.04.002

      (47) Burwell, R. D. & Amaral, D. G. Perirhinal and postrhinal cortices of the rat: interconnectivity and connections with the entorhinal cortex. J Comp Neurol 391, 293-321 (1998).

      (48) Kajiwara, R., Takashima, I., Mimura, Y., Witter, M. P. & Iijima, T. Amygdala input promotes spread of excitatory neural activity from perirhinal cortex to the entorhinal-hippocampal circuit. J Neurophysiol 89, 2176-2184 (2003). https://doi.org/10.1152/jn.01033.2002

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPES.

      We thank the reviewer for noting these strengths of the work along with the below encouragement to revise the manuscript to better highlight the key findings and their implications.

      Weaknesses:

      The technical proficiency and complexity of the study and analysis also present a clear limitation and challenge for interpretation. As a reader, even those who are quite knowledgeable about the methods, constructs, and questions being addressed will often struggle (as this reviewer did) to keep the large set of findings in mind and gain an understanding of how they all fit together.

      Indeed, it seems like there are many threads running together in the paper, which makes it challenging to find the through-line of the key findings, or to understand how they might relate to some pre-existing hypotheses, rather than merely interesting patterns detected in the data. In the Introduction and Discussion, it seems as if the key question is to understand the pathways by which MPEs impact cognition, but this is a rather broad topic, so it is not clear exactly what the authors are aiming at with this question and study design.

      As an example, authors operationalize frontal theta power as an index of cognitive control demand, and one of the pathways by which MPEs impact cognition. But this point becomes somewhat circular, since it is not clear how or why the Mismatch x Strength interaction in frontal theta reflects that demand. It would have been better to set this pattern up in the Introduction as a theoretically driven hypothesis, since it currently appears more like a post-hoc interpretation. This is mirrored by how the issue is first brought up in the Introduction, where it states somewhat vaguely: "whether MPEs are followed by an increase in frontal theta... warrants closer examination".

      Again, we appreciate the reviewer’s thoughtful feedback on where the manuscript can be clearer, especially given the rich set of results it reports. Following the reviewer’s guidance, we restructured and revised the Introduction to further motivate the hypotheses that (a) MPEs increase both attention/arousal (grounded in studies of Event Segmentation Theory) and cognitive control (given findings on reward prediction errors and other types of prediction errors), and (b) there are greater increases in these processes triggered by strong compared to weak MPEs. To better link these hypotheses to resulting statistical tests, we note that hypothesis (a) was tested in our trial-level regression models in Fig. 1 by examining main effects of Strength, whereas hypothesis (b) was tested in our models in the Mismatch x Strength interactions. On point (b) and potentially circularity, we note that previous work indicates that frontal theta scales with negative reward prediction errors; as such, we hypothesized that stronger MPEs would elicit more frontal theta, as evidenced by a robust Mismatch x Strength interaction during the probe period.

      Later in the results, there are findings relating frontal theta to pupil dilation, posterior alpha suppression and then subsequent memory. It was hard to understand how all the findings might be linked together functionally or conceptually. Are the authors potentially postulating a mediating or mechanistic pathway, in which the MPE leads to increased cognitive control (frontal theta), which then leads to enhanced subsequent memory of those events? If this is the case, then maybe a formal path analysis would be the best way to test or state this hypothesis. It would also be useful to specify more clearly how the pupil components and alpha suppression factor into this mediating path, since it was not clear.

      Relatedly, the authors suggest that internal attention and arousal also play relevant roles in this pathway, but these are also not clear. In some cases, it is stated as if this is a distinct pathway from the cognitive control one, since there is a focus in the results on the independence of frontal theta and posterior alpha, but elsewhere they seem to be treated as two aspects, or distinct steps, within a single pathway. Again, these different threads of the findings were quite challenging for the reader to follow. Pathway analyses, such as with multiple mediation or moderated mediation, could be a useful way to address this question. For example, it seems as if readiness-to-remember is another behavioral outcome (like subsequent memory) that could be used in the search for mediators.

      We thank the reviewer for highlighting these ambiguities in the original submission and for the thoughtful encouragement to leverage mediation models to more formally test the hypothesized relationships between MPE magnitude and constructs of control, attention, and arousal. In the revision, we now more clearly hypothesize that the effects of strong MPE-driven increases in attention and arousal might be explained, in part, by cognitive control (as indexed by frontal theta) upregulating attention and arousal. To more explicitly test this model of the relationships between our measures, as recommended by Reviewer #1, we now include multivariate mediation analyses to assess whether, at a trial-level, changes in posterior alpha and the immediate pupil effect PC3 are explained in part by increases in frontal theta. Because changes in posterior alpha following MPEs and PC3 scores did not predict subsequent memory, mediation analyses addressing the hypothesis that our attention/arousal measures mediate the effect of frontal theta on subsequent memory were not conducted. Following insights from a cross-correlation analysis, as recommended by Reviewer #2, we tested an additional model to examine whether the effects of MPE magnitude on frontal theta were explained in part by changes in posterior alpha. Examination of the posterior distributions of the indirect effects did not favor our hypothesized model, nor the alternative model that attention upregulates control. Altogether, these outcomes suggest that increases in control, attention, and arousal following strong MPEs may be elicited independently. Yet, we also note that the current set of experiments may be underpowered for these mediation analyses; future work can further investigate the directionality between these effects. Finally, with respect to the readiness-to-remember findings, they were removed in the interest of space, as recommended by Reviewer #2.

      At the minimum, it would be quite helpful to have diagrammatic figures that specify the hypothesized and observed relationships between independent variables (Strength, Mismatch), physiological indices (pupil dilation components, frontal theta, posterior alpha) and key outcome measures (accuracy, RT, next-trial retrieval success, subsequent memory), so that the reader can refer back to them as each component of the analyses is conducted.

      To further increase conceptual clarity, we also followed this helpful suggestion, adding diagrammatic figures to illustrate our hypotheses regarding interactions between the effects (Fig. 3a) and to summarize the observed relationships between MPEs, control, attention, and arousal (Fig. 6).

      Minor Points:

      Many figures had x-axes showing a pupil component or EEG power metric broken down by quartile or quintile. Yet nowhere is it ever explained why this graphical (or analytic?) approach is used and what it reflects, or how it is decided which break down to use (quartile/quintile). If the data are analyzed as a correlation, why is a scatterplot not shown instead?

      In the linear mixed effects models, continuous values were used to assess relationships between variables. In the figures, the continuous variables were binned into quartiles or quintiles for ease of visualization. We opted to visualize the data using this approach, rather than with a scatterplot, given the large number of trials. We updated the figure captions to clarify the approach.

      It was surprising that, unlike readiness-to-remember, which was analyzed via logistic regression and odds-ratio, subsequent memory was not analyzed in the same fashion (i.e., as a binary outcome variable predicted by frontal theta), rather than in a reverse chronological one (subsequent memory predicting frontal theta). Historically, it was the case that subsequent memory was analyzed in this manner, but that was before the era in which trial-level linear mixed-effect models were in wide usage, as they are implemented in this study. Thus, the choice seems like a wasted opportunity or a step backwards analytically.

      We thank the reviewer for this encouragement and agree with the point. In the revision, we note that the readiness-to-remember results were removed in the interest of space and clarity, as was recommended by Reviewer #2. With respect to the subsequent memory analyses, they are now analyzed via logistic regression.

      Reviewer #2 (Public review):

      Strengths:

      The study has a clear behavioral paradigm with multiple measures - behavioral, EEG, and pupillometry that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

      We thank the reviewer for noting these strengths of the work along with the below encouragement to revise the manuscript to better highlight the key findings and their implications.

      Weaknesses:

      The methods are rigorous, and the claims are mostly supported by the data, but there are a few weaknesses or places that could be improved:

      (1) The authors conduct PCA analysis to identify different components of the pupillary response to MPE and relate them to behavior. Specifically, the authors identify components PC3 and PC4, which they interpret as related to MPE. However, some parts of the interpretation could be clearer or better justified:

      (a) The authors refer to PC4 as "post-decision cognitive processing". But, given that RT was between .5-.7s, and PC3 peaked after more than 1s, wouldn't it be cautious to interpret PC3 as postdecision as well?

      Thank you for raising this point. Given that pupil is a relatively sluggish response, it is possible that both components reflect post-decision cognitive processing, even if PC3 peaks before PC4. Following the reviewer’s guidance to adopt more cautious language, we replaced “post-decision” with “post-MPE”.

      (b) MPEs overall elicit longer RTs in this study, suggesting that long RT is a behavioral marker of MPE. Nonetheless, the authors argue on p. 12: "Altogether, these findings indicate that when stronger mnemonic predictions (as indexed by shorter RTs) were violated." And, PC3 is correlated with shorter RTs for mismatches, meaning that behaviorally, these trials were more similar to matches. Thus, how do the authors interpret shorter versus longer RTs for MPEs, and what processes do these RT reflect?

      We thank the reviewer for stressing the need for greater clarity regarding the relationships between RT and the constructs of interest. With respect to RT, we interpret the condition-level difference in RTs between mismatch and match trials as a behavioral marker of an MPE. However, when comparing mismatch trials within a given strength condition to each other (i.e., an analysis at the trial-level), shorter RTs may reflect a stronger prediction, greater certainty that the probe is a mismatch, and therefore the experience of a stronger MPE. Note that while larger PC3 scores were associated with shorter RTs for mismatches (Fig. 2b, right), the mismatch RTs in the largest PC3 quartile were still longer than those on match trials in the corresponding quartile (in other words, there was still a condition-level difference in RTs in the largest PC3 quartile, suggesting that the mismatch trials in this bin are behaviorally still likely to be different from match trials in the corresponding bin).

      To clarify these relationships and our interpretation, we modified the referred to text: “(as indexed by shorter strong mismatch RTs).” Moreover, we added text to the Discussion, further delineating our reasoning here and the implications of our findings for understanding the mechanisms giving rise to and triggering by MPEs of varying strengths. This includes adding an explicit summary of the logic and findings that notes that, at the trial-level, shorter RTs may reflect stronger predictions; at the match/mismatch condition-level, longer RTs may reflect the experience of a MPE. For PC3, shorter RTs (trials with a stronger prediction) in the Strong Mismatch condition were associated with a larger pupillary response. The added text notes that “while longer mean RTs for mismatches compared to matches are a behavioral marker of a MPE, within-condition differences in RTs (i.e., between mismatch trials) may reflect more subtle differences in MPE magnitude, with shorter RTs reflecting stronger predictions and thus stronger MPEs. Strong mismatch RTs were used in the mediation models as a proxy measure of MPE magnitude; this estimate may be noisy because RTs in this experiment are likely sensitive to factors independent of mnemonic prediction strength (e.g., preparatory attention (Supplementary Fig. 6) or memory strength of the mismatch probe). This limitation may have additionally reduced sensitivity to detecting indirect effects.”

      (2) The brain to pupil relationship (p. 13-14): If I understand correctly, this was done on a trial-by-trial basis, but the high temporal resolution allows doing the analysis in a time-resolved manner - does brain activity at a certain time point preceding/following the pupil response correlate with the pupil response? It might be that cognitive control influences attention mechanisms or vice versa (because there is some overlap in the response). Although not testing causality, this temporally resolved correlation would be an interesting way to start probing how signals might influence each other.

      Thank you for this suggestion. We now report a cross-correlation analysis (Fig. 3d) that suggests that cognitive control increases precede attention decreases at retrieval, whereas in response to a strong MPE, cognitive control increases follow attention increases. There were no significant clusters for the temporal relationships between frontal theta and pupil, nor for posterior alpha and pupil.

      (3) The relationships the authors find between brain measures and pupil components were largely not specific to mismatches/matches. However, are they specific to this task? I think it would benefit the paper to show that these relationships are potentially specific to making match/mismatch memory decisions, versus, e.g., any stimulus processing. For example, the authors could run the same analyses locked to stimuli in the study phase, anticipating a different pattern, if indeed these findings are specific to the associative memory task.

      Many of the associations between our measures indeed did not show an interaction with Mismatch. Due to jitter in the ISI in the study phase, some of the analyses in the retrieval phase cannot be performed in the exact same way for the study phase. We will leave these questions to be addressed in future research. We agree that an important question for future research is to address whether these responses depend on making match/mismatch decisions and now include consideration of this point in the Discussion: “Finally, the magnitude of observed increases in control, attention, and arousal following strong MPEs may be influenced by the decision-making process engaged when making match/mismatch judgments. Not all MPEs necessitate behavioral responses. Whether similar magnitudes in neurocognitive responses and consequences for learning are observed upon detection of an MPE, but in the absence of a decision remains unclear.”

      (4) During memory retrieval (i.e., before the probe), the authors find that frontal theta, a marker of cognitive control, was associated on a trial-by-trial basis with more posterior alpha (i.e., less alpha suppression, potentially reflecting less attention), and that this association was stronger for weaker predictions. The authors interpreted this as weaker predictions necessitating more cognitive control, and that more cognitive control was recruited specifically in trials where retrieval included less content (memory reinstatement) to attend to. Generally, cognitive control is recruited to facilitate memory retrieval. If so, one possible interpretation is that this correlation reflects cognitive control effort that has failed to produce enough memory reinstatement. The other possibility is that this correlation reflects more specific retrieval of the correct probe, without retrieval of interfering items (i.e., overall less content). I believe that the former explanation predicts that this correlation would be associated with longer RTs (more difficult decisions), while the latter predicts shorter RTs (easier decisions due to successful retrieval), at least for matches.

      Thank you for these insightful comments. Because this analysis is not key to the main questions about MPEs and given both reviewers’ concerns that the manuscript can be overwhelming for the reader given the sheer number of findings reported, we opted to move this point from the main text to the Supplement. However, following the reviewer’s guidance here, we conducted the proposed analyses and tested these alternative accounting by modeling RTs. The results favour the former interpretation:

      “Greater control being associated with less attention could reflect failure in controlled retrieval efforts to reinstate sufficient memory evidence of the probe and thus fewer retrieval products to which attention is allocated. Alternatively, greater cognitive control could increase the likelihood of retrieval success, eliciting selective retrieval of the correct probe and inhibition of interfering items. To address these alternatives, we examined how the association between frontal theta and posterior alpha related to the difficulty of a trial, as assayed by RTs. The former failure-of-control account would predict that a stronger positive frontal theta- posterior alpha association would relate to longer RTs, whereas the greater-retrieval-specificity account would predict that a stronger positive association would relate to shorter RTs. In a model predicting RTs as a function of frontal theta, posterior alpha, Strength, Mismatch, and their interactions, we found a two-way frontal theta × posterior alpha interaction (β=0.016, CI=[0.001, 0.031], p=0.033), such that a stronger positive association between frontal theta and posterior alpha predicted longer RTs. This relationship did not differ as a function of Strength (no frontal theta × posterior alpha × Strength interaction: β=-0.015, CI=[-0.033, 0.004], p=0.124), Mismatch (no frontal theta × posterior alpha × Mismatch interaction: β=-0.016, CI=[-0.050, 0.018], p=0.342), or interact with Strength and Mismatch (no frontal theta × posterior alpha × Strength × Mismatch interaction: β=0.024, CI=[-0.014, 0.062], p=0.218). Together, these outcomes support the idea that during memory retrieval, positive coupling between frontal theta and posterior alpha may reflect failure or inefficiency of cognitive control efforts to rapidly accumulate mnemonic evidence to which to attend in support of a memory decision.”

      (5) In section 3, the authors found a positive relationship between alpha during memory retrieval and PC3 during MPE. If I understood correctly, this means that less attention during retrieval (less suppression) is correlated with a stronger PC3 response. How do the authors interpret this? Maybe along the same lines as in (5), specifically retrieving the correct information (i.e., less retrieved content to attend to) means a stronger prediction, leading to a stronger MPE, and a stronger MPE response, as reflected by PC3?

      We appreciate this comment, as it highlights a need for greater clarity here. The observed relationship was actually negative (Fig. 4f), meaning that more attention during retrieval was associated with a stronger PC3 response. We suspect that the lack of clarity here may be due to the original statement that “there was a positive relationship between posterior alpha suppression during memory retrieval [and PC3 scores]”. To increase clarity, we have modified this statement to “there was a negative relationship between posterior alpha during memory retrieval [and PC3 scores]”. We interpret the negative relationship with posterior alpha (i.e., positive relationship with posterior alpha suppression) to indicate “that greater attentional allocation during memory retrieval, which occurs when memories are stronger and more retrieval products can be reinstated and attended to (Fig. 1e; Fig. 4c), predicts the magnitude of immediate pupil responses to MPEs.”

      (6) The results with subsequent memory are important and address a major gap in the field that largely did not relate neural effects of MPE to subsequent memory. However, one major limitation of the study is that the authors did not test memory for matches. I understand the logic of avoiding testing matches. Because matches were repeated more times in the study, it's not a fair comparison, and could change participants' overall criterion for old/new decisions. However, one possibility would have been to test only the weak prediction; this could have given some specificity to the neural subsequent memory findings.

      We were indeed concerned about the change in decision criterion and did not include match items for this reason. Nonetheless, this is a useful suggestion that future work could include the weak match items for a better comparison of subsequent recognition memory. We now comment on this in the Discussion: “Whether control-associated enhancements in learning are specific to learning from MPEs can be further tested in future work by including a test of subsequent memory for weak match probes as an additional control condition for comparison.”

      (7) The authors nicely characterized the different PC of pupillary MPE response. But, with respect to subsequent memory, they only present pupil size. Unless there is some methodological reason that prevents testing subsequent memory on the PC, I think this will be very informative about the potential mechanisms underlying memory of MPE.

      The pupil PCs were not associated with subsequent memory, though there are some interesting trends in the 48-delay condition which could be explored in future work. These findings are now reported in the Supplement (Supplementary Fig. 16).

      (8) This paper includes many interesting findings, and I am not sure how they all come together into a cohesive mechanistic understanding of MPE response and subsequent memory. I think the paper would benefit from either a conceptual mechanism figure or, in the Discussion, have a summary of a proposed mechanism integrating the findings together.

      We thank the reviewer for stressing this point, which also was raised by Reviewer #1. To better emphasize the novel contributions of the work and to assist the reader’s understanding of the key findings, we: (a) revised the Introduction to more explicitly describe our hypotheses about the relationships among our measures as they relate to MPE responses and subsequent memory; (b) now report mediation models to more directly test these relationships and include diagrams of the hypothesized relationships; and (c) include a schematic figure at the end of the Results to highlight the key mechanistic relationships supported by the data.

      (9) Relatedly, the section "Immediate, strength-sensitive neurocognitive impacts of MPEs" does not link the arguments to specific data points, so it's hard to follow which data specifically the authors are interpreting.

      The discussion in this paragraph rests on the Strength × Mismatch interactions observed in Figures 1d-f, summarized in the first sentence of the section. To increase clarity, we changed the title of this section to “Neurocognitive impacts of MPEs are strength-sensitive”.

      (10) If I understand correctly, the authors did not find improved memory for strong compared to weak MPE. First, I think this behavioral result should be incorporated in the main paper and in the interpretation of the results. Second, given that the neural effects the authors tested either correlated with memory for strong MPE or did not show a relationship with memory, what neural/pupil response could explain memory for weak MPE?

      Thank you for raising these points. The behavioral result and the frontal theta trial-level regression model for subsequent memory are now described in the main results and included in the Discussion. As noted in the Discussion, the trial-level regression model indicates pre-probe and post-MPE frontal theta effects may explain memory for weak (and strong) mismatch probes. While the magnitude of probe-period MPE frontal theta is additionally predictive of memory for strong mismatch probes (as indicated by the Subsequent Memory × Strength interaction in the trial-level regression model in Fig. 5c and the logistic regression in Fig. 5b), frontal theta during this period does not additionally enhance memory for strong mismatch probes above that of weak mismatch probes (Fig. 5a). We added more discussion on why memory for strong and weak mismatch probes in the current experiment did not differ. We also discuss potential directions for future research to further probe mechanisms underlying MPE-driven learning.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It is recommended that the authors determine whether formal path analyses, testing for mediation and moderation, would provide a useful approach from which to better integrate the disparate set of findings and make clear their causal/functional implications.

      At a minimum, adding a diagrammatic figure is recommended to visually depict the key components of the study (independent variables, physiological indices, outcome measures) and how they relate to each other both conceptually (ideally in a theoretically hypothesized manner) and in terms of the observed findings. Such a figure will help the reader keep track of the many types of findings and results threads, and with the goal of better organizing the results into a clearer narrative through-line.

      Thank you for the suggestions to add mediation analyses and diagrammatic figures. To more formally address our hypothesis that increases in attention/arousal following MPEs are explained in part by increases in cognitive control, we tested two mediation models: one with posterior alpha as the measure of attention and the other with pupil PC3 scores. We now diagram this hypothesis in Fig. 3a. Given the outcomes of the cross-correlation analysis suggested by Reviewer #2, we also tested an additional model where posterior alpha might explain the impact of strong MPEs on frontal theta (now diagrammed in Fig. 3e). We did not find credible evidence for an indirect effect in any of the models, suggesting that strong MPE-driven increases in control, attention, and arousal may be elicited independently. Finally, in a newly added summary diagram (Fig. 6), we highlight the observed effects of strong vs. weak MPEs on RTs, control, attention and arousal; differences in control, attention, and arousal for strong vs. weak predictions preceding the MPE; trial-level mismatch-specific or strong-specific effects on subsequent memory; and mismatch-specific interactions between processes during retrieval and responses to MPEs. We hope these revisions address Reviewer #1’s concerns and that these figures help the reader keep track of the key hypothesized relationships and main findings.

      Reviewer #2 (Recommendations for the authors):

      (1) The relationship between event segmentation and prediction errors has been reviewed recently in two papers (Nolden et al., 2024, Neuroscience & Biobehavioral Reviews; Rouhani et al., 2024, JOCN for a potentially relevant computational model). I wonder if insights from these papers can inform the Introduction/Discussion of the current manuscript.

      Thank you for these suggestions. These papers are now incorporated into the Introduction and the Discussion.

      (2) Brod et al. (2022, Psych. Bull. Rev.) have previously reported increased pupil dilations for MPE correlating with subsequent memory, specifically for strong prediction errors. I think it's worth including this paper in the Introduction as the finding is highly relevant. The Brod paper might provide more direct evidence of "MPE-related increases in pupil size" than the evidence the authors provide (p. 3).

      Thank you for this suggestion. This paper is now incorporated into the Introduction and Discussion.

      (3) I'm confused about the temporal analysis: "Trial-level regression analyses were conducted on the frontal theta, posterior alpha, and pupil time series from the associative retrieval test to identify temporal clusters that were sensitive to the factors of Mismatch (i.e., mismatch vs. match probes) and/or Strength (i.e., strong vs. weak associative pairs). For each participant and each time point, a linear regression model testing main effects of Mismatch and Strength, and a Mismatch × Strength interaction was run using R to compute beta weights for each regressor." What regression exactly was run? A separate model for each participant and time point? Across trials, then? Later, the authors mention that beta weights were averaged across participants and t-tests and permutation tests were conducted. However, if the data were averaged, what t-test was conducted? And how was the permutation test conducted? It's also unclear what the authors mean by "the sign of each participant's beta weights" - what sign?

      Thank you for raising this point. A regression was run separately for each participant and each time point, across trials: neurocognitive measure ~ Mismatch + Strength + Mismatch: Strength. Beta weights were averaged only for visualization; t-tests were conducted on each set of beta weights (across participants), separately for each time point. We modified the text of the Methods to describe our procedure more clearly.

      (4) The associative memory and recognition accuracy data are presented as d'. In addition, the authors should provide hits and false alarms to facilitate a better interpretation of the results.

      We now report these outcomes in Table 1 and refer to them in the main text.

      (5) The authors argue regarding the frontal theta that "Qualitatively, the main effect of Strength emerged later than the main effect of Mismatch, suggesting that the increase in frontal theta evoked by the probe was more sustained for weak compared to strong trials (or, as a corollary, that the greater control elicited by strong MPEs enabled more rapid resolution of conflict and ultimate choice selection)." It was unclear to me how that stems from the data.

      This statement has been removed altogether.

      (6) Especially in Figure 1, I think clarity can be improved if the authors would indicate the specific subsection they are referring to in the text, because even within, e.g., 1d, there are different graphs, so mentioning which graph is relevant for which statement would be helpful to the reader.

      Thank you for this suggestion for improving the clarity of the manuscript. Fig. 1d-f includes multiple parts because we wanted to show the raw data (the mean time series on the left) as well as the model coefficients (time series on the right). We are hopeful that, given that each subplot is titled and that the main text refers to both the subplots on the left and on the right, this will be clear to the reader as is. We welcome further guidance if this remains a concern.

      (7) This seems highly speculative to me: "Elevated PC4 scores on weak hits, where no error was made, may instead reflect retrieval practice and thus internally oriented attention. On these trials, when cueelicited retrieval may have been weaker, the probe may have provided additional support for pattern completion of the learning episode for that association, and this engagement in memory retrieval may have elicited pupil dilation (c.f., Strength effect in Figure 1f). Overall, PC4 may therefore reflect an attentional orienting response that, depending on the relative success of memory retrieval and probe identity, may direct attention internally or externally. " (p.13). In my opinion, the authors make a lot of assumptions about underlying processes. I'd consider removing.

      We removed this text.

      (8) In Figure 3c, should the x-axis be PC3 quintile? And (c) is not in the figure caption.

      Thank you for this note. We addressed these issues.

      (9) The carryover effects the authors report are interesting, but they seem detached, and it is unclear how they fit with the additional findings. In a paper that already includes many findings, I'd recommend either integrating better or removing.

      Following this guidance, we removed the readiness-to-remember findings in the interest of space and clarity.

    1. Author response:

      We thank the reviewers for their time and insightful comments. We are also grateful for their appreciation of the unusual circumstances that led to the publication of the manuscript in its current form.

      In response to reviewer #2’s question about replication, we provide additional details here. The error in our retracted original publication affected only the imaging results. Despite this, we reproduced key behavioural experiments by generating additional datasets (rather than simply rechecking records and authenticating results) and therefore have full confidence in our behavioural findings. We are very happy to share some of these replication experiments below:

      Author response image 1.

      Data showing replication of key experiments. From L-R these data replicate those shown in Figure 2e, Figure 4h, Figure 4c, Figure 4d.

      The conceptual framework of the original study remains valid. It was actually formulated based on the behavioural data and before any physiological recordings were made. We believe that it still represents the most parsimonious explanation for the observed behavioural results. Multiple behavioural findings support a model in which multisensory training leads to the recruitment of visual γd Kenyon cells into an otherwise olfactory memory trace. These include: (1) the requirement for γd KC output during olfactory retrieval following multisensory training; and (2) the sequential learning experiments, which were originally designed to test this model and provide independent evidence for its predictions. We nevertheless agree that the loss of the physiological data reduces the amount of evidence supporting the proposed mechanism. The nature of the physiological changes following multisensory learning remains an important question that we intend to address in future work.

      We also thank the reviewers for identifying the unfortunate typo in the Abstract, which we believe contributed to the confusion over the dopamine receptors tested in this circuit. We selected these receptors based on our in-house single-cell transcriptomic expression data, together with published evidence indicating their specific expression in the neurons of interest. We have also responded to reviewer comments about the clarity of the figures and added additional labels to figure 2, to clarify the experimental paradigm in each case.

    1. Author response:

      We greatly thank the Reviewing Editor, Senior Editor and the reviewers for their constructive and thoughtful feedback, as well as for recognizing the significance and strengths of our study. We are very encouraged by overall positive assessment and appreciate these insights to strengthen the manuscript. Below, we outline our plans to address the key points raised in reviewer#1’s public review. We also note that reviewer#2 did not identify any weaknesses in the study and are thankful for this positive evaluation.

      Point 1: Relating in vivo and in vitro findings and discussing the relevance of myelin clearance in AD. As suggested, we will elaborate our discussion to cover the relationship between our in vivo and in vitro findings and present a more unified picture of GPR34 function. We will also highlight the relevance of myelin clearance in Alzheimer’s disease independent of amyloid plaque burden.

      Point 2: in vivo assessments and phenotypes. While we recognize the value of expanding our in vitro and cellular findings to in vivo and cognitive measures, we believe this additional assessment is beyond the scope of this current study. At least, we will expand our discussion to relate our findings of GPR34-medilated myelin pathology in the contexts of neurodegeneration and cognitive deficiency in AD.

      Point 3: Myelin engulfment versus lysosomal degradation. We agree that the clear distinction between impaired myelin engulfment and altered lysosomal degradation is an important point. Although additional experiments to address these scenarios are beyond the scope of the study, we will perform targeted pathway analysis of existing transcriptomic datasets focusing on the phagosome and lysosomal degradation pathways along with the in vivo findings related to CD68, which could offer an interpretation.

      Point 4: Comparison of transcriptional signatures. As suggested, we will compare the transcriptional signatures induced by myelin exposure in WT iMGLs with our 5xFAD RNA-seq datasets, as well as publicly available datasets from human Alzheimer’s disease microglia, to unravel the potential converged and distinct features.

      Point 5: Reconciling our findings with previous GPR34 studies. We will expand our discussion to compare and clarify our current findings with the previous studies and cover potential reasons for those different observations. We will also cover the current limitations of existing GPR34 pharmacological tools compounds, including their limited blood-brain permeability, which precludes the effective use of those tools for proposed in vivo studies.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      "Learning is a fundamental source of individuality," by Manna and colleagues, interrogates different sources of variation in individual behavior. The authors place individual flies in a Y-shaped arena, which is a common design in the field, and illuminate the arms of the Y with blue versus green light. They track the color preference of individual animals and also perform operant conditioning, meaning that they teach the fly to avoid a particular color/arm by generating a foot shock when the fly enters that arm. There are a number of things that are impressive about this setup: The authors are able to collect data on thousands of individual flies of many different strain backgrounds, and they demonstrate a strong change in color preference after conditioning. This is nice, because in past papers, visual learning ability has been modest and difficult to study. To put a number on it, in this paper, animals on average don't show a color preference at the start of the assay, spending around 30% of their time in the one arm illuminated green, and the remaining time in the two arms illuminated blue. After conditioning, the average animal spends only 23% of its time in the green arm.

      The authors run 64 animals through the assay for each of 88 wild-type strains (maybe? see Major Point 1 below) and see considerable strain-specific (genetic) variation in the change in time spent in the shocked color after conditioning. Some strains show no learning, while others spend <10% of their time in the shocked color after conditioning. They also, I believe, see that some strains have more variability across individuals, which would suggest that some strains have stronger canalization at the development or circuit function level than others, i.e., some genotypes produce more consistent copies of the individual, others less consistent copies. (Or, some genotypes produce robust circuits, and others produce noisy circuits.)

      Finally, the authors argue statistically that learning itself increases variability in individual performance. This makes a lot of sense to me intuitively. Learning changes the physical/chemical properties of circuits in the brain, and because it evolves over time and interacts with environmental variables, it seems like it should send different animals down different channels. Or, at a conceptual level, if I learn to play the piano and my sister doesn't (because of some genetic difference between us or something stochastic), this learning experience will cause all sorts of other differences in our behavior as time passes. I also think the authors do have enough data to be able to make this finding. However, the presentation of the argument in this portion of the paper is hard for me to understand, and I am not an expert in statistics, so the strength of the result is difficult for me to evaluate.

      Major points

      (1) It's difficult to track through the paper the number of animals tested for different assays. At the beginning, it says N=5632, which works out to 64 flies for each of the 88 DGRP strains. 64 happens to be the number of parallel Y arenas they have. Later in the methods, there's a description of more variation within the set of 64 for each strain, two different parent sets per strain, different sexes, conditioned and unconditioned. And, while the results text focuses on the color learning, the methods discuss additional assays (place learning, multi-day learning).

      Given the numbers, does each run of the 64 mazes include all the tested flies of one strain, or are flies of many strains included in each batch? Do different flies do different assays (color, place, multi-day), or do they all do all the assays? Perhaps there is a table including this information already in the supplement, but I recommend making it much clearer in the main results text and methods. While the dataset is large, if it is split over many conditions and/or if batch and genotype confound each other, this will affect the robustness of the results and how strong the conclusions can be.

      (2) The data presentation in Figure 1 is elegant and easy to follow, but getting into Figure 2 and subsequently, I get lost in the statistics and have trouble understanding what is being measured. My understanding of the big picture is that while genetics and individual randomness contribute a lot to behavior, the evidence for learning as an amplifier of individuality is that variance in behavior among animals of the same strain increases over time in the conditioned group (i.e., the group that is doing the most learning, or a specific kind of learning), but not in the control group. This idea is illustrated in the flattening distributions in the cartoons in Figure 1A. The authors should include graphs of the real data that use the same format as in that cartoon. Instead, the graphs present "residuals," and I don't know what those are. I suspect it's "variation left over after accounting for effects of strain and individual stochasticity." I see the residuals being tracked per strain over time in Figure 2H, but I don't see the change over time in other graphs. I'm looking for something simple like, "variation within the strain at the beginning of learning and at later time points in learning." (But I'm not sure exactly what instantaneous measurement would be the focus in longitudinal analyses of learning behavior.)

      (3) Figure 3 is a cool stab at tracking down the precise mechanism by which a stochastic environment interacts with learning to send individuals along different behavioral routes. But again, like in Figure 2, I don't have the sophisticated understanding of statistics to understand exactly what the graphs are telling me, or how they relate to the underlying measurements. I'm relying on the results text alone to reach a conceptual understanding, and just taking the graphs on trust.

      So, overall, the authors have a very nice body of work here, and with the potential to add a new facet to our understanding of the origins of diversity in animal behavior. In addition to the interpretations they focus on here, this dataset also represents an advance in studying visual associative learning in general, and quite an amazing ability to make longitudinal measurements of many behavioral decisions within the same animals. Improving the data presentation to make it easier to follow for a larger swathe of researchers, especially in figures 2 and 3, will increase its potential impact.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to test the extent to which differences in learning capacity and experience contribute to behavioural variation in a genetically identical population under identical environmental conditions.

      Strengths:

      The authors developed and used a scaled-up version of a simple two-choice behavioural paradigm, allowing them to test thousands of individuals across multiple genotypes. They then deployed clever and powerful statistical analysis methods and provided compelling evidence for a role of variability in learning in the expression of behavioural variation.

      Weaknesses:

      There are no major weaknesses, although some level of longitudinal analysis to strengthen the evidence for a strict definition of individuality would be a welcome extension of a future study. In addition, it would have been very interesting, although understandably beyond the current scope, to delineate a potential source of learning variability in the brain.

      Following our provisional response to the reviewers, we have implemented these additions to the manuscript:

      (1) We have added 7 additional tables (Table 1-6 and table 8) to the supplementary that detail how many individual flies were used in which of the seven separate experiments, how the individuals were distributed across genotypes, replicates and sexes, and how many were filtered out before the final analysis. At the bottom of each table, we added a short description of the type of experiment and a brief explanation of the filtering. The four smaller experiments where we tested the two mutant lines and one wild-type DGRP line were used primarily to test and validate the experimental platform and the behavioural paradigm used for the main experiment. In these four experiments we tested green place learning, blue place learning, green colour learning, blue colour learning in four separate batches of flies. In each of these experiments we used 192 individuals (64 individuals x 3 genotypes x 4 experiments = 768 individuals in total). In the multiday experiment we used 64 flies per genotype per each of the four groups of sequences of learning paradigms, in total 512 individual flies. Here, each of these 512 individuals were retested in different learning paradigms over 4 days (Table 5). The main experiment was the green place learning (Table 6) where we tested all 88 DGRP lines and again the two mutant lines was used to obtain the majority of the main results and conclusions (64 individuals x 90 genotypes = 5760 individuals). Lastly, additional 896 individuals were measured in the experiment using blue place learning paradigm to test the consistency of learning behaviour as opposed to colour bias within genotype (Table 8). In summary, in all experiments, we have always measured behaviour in 64 flies per genotype (full loading of the behavioural platform), and they were distributed almost entirely evenly across replicates, sexes, and conditions (control vs conditioned). No individual was reused across experiments. In most cases, after filtering the data, 60 or fewer individuals were used in final analyses. For the very few deviations from this experimental design (which occurred due to unforeseen events such as dropped/sick vials, flies flying away or accidentally squished during setup, skewed number of males and females etc.) we added a short explanation in the text below the tables. In total, across all reported experiments in this study, we measured behaviour in 7936 individuals.

      (2) We have added a schematic visual representation of classical measurement of individuality (variance of the distribution of behaviour within genotype where genetically identical individuals are raised in the same environment), entropy-based measurement of individuality (residual individuality) and the change in residual individuality, as we use them in this study (Figure 2D). We also provide a list of different DH<sub>resid</sub> measures and what distributions are being compared across the DH<sub>resid</sub> in the same figure. We hope this will serve as a more intuitive explanation of individuality and help readers interpret and follow more easily the results that we report after this figure.

      (3) In the same vein, we added another schematic visual representation to Figure 3 (Figure 3F) where we depict how distributions of individual behaviour may change with every decision and how this change translates to (or can be read out from) the change in residual individuality. We have also renamed the X axis of Figure 3E to “DH<sub>resid</sub> Start”, so that it is clearer what is measured here and matches the explanation in Figure 2D.

      (4) We have added two additional supplementary figures where the reader can inspect in more detail how the distributions of individual behaviour change longitudinally across time for each genotype in control and conditioned (Figure 3 – Figure supplement 2 and  Figure 3 – Figure supplement 3). From these figures one can glance how variance as well as the shapes of the distributions change as the flies learn in the conditioned setting, and how they remain largely the same in the control where flies behave spontaneously. We have added a sentence in the main text to introduce these figures: “We found that the distributions of individual behaviour were broader and their shapes changed substantially over the course of the experiment for the conditioned flies, and not for the control flies Figure 3 - figure supplement 2, Figure 3 - figure supplement 3).”

      (5) As noted in the first provisional response to reviewers, we changed the sentence “In every individual, behaviour is shaped by deterministic, genetic factors and by environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.” to “In every individual, behaviour is shaped by fixed genetic factors and by variable environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.”

      (6) Some sentences were edited in the results so that we can correctly refer to the newly added tables and figures. Context, meaning or interpretation of the results in these sentences was not altered.

      (7) While adding the new table references to the text, we noticed a typo that propagated in the previous version where the number of flies used in the main experiment was stated to be N= 5238, when in fact it should have been N=5239. This is now fixed.

      We once again thank the reviewers for their comments and suggestions – we believe their suggestions helped us improve the presentation and interpretability of our study and we hope the reviewers and readers will agree with this as well.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors demonstrate an innovative approach to investigate the effect of cone dropout on visual acuity using their newly developed olo system. By systematically reducing the coverage of real-world input to the cone photoreceptor mosaic ("cone dropout condition"), the authors are able to assess how having fewer cones leads to reduced vision, in comparison to existing approaches ("pixel dropout condition").

      The capture of a rich dataset, including cone imaging and eye motion, is valuable. Benchmarking with the prior literature, suggesting that good visual acuity can be maintained despite a 50% loss in cone density, is impressive. However, it is known that cone density varies dramatically from the peak cone density location in the foveal center to even a location a few degrees outside of the fovea. In addition, there is a high degree of subject-to-subject variation in peak cone density. Given that the C stimulus is hollow in the middle, the stimulus does not actually hit the location of the peak cone density but must land slightly outside of it. Therefore, considering the actual cone density of where the stimulus lands will be important to discuss and/or analyze.

      The reviewer is correct that the cone density will vary dramatically with distance from the foveal center. However, importantly, in our experiment the Landolt C stimulus is fixed in the world and the eye is free to move across it. Therefore, the subject can direct their gaze to any part of the letter, rather than it being fixed to the hollow center of the letter. In the worst case, if the subject kept their gaze fixed at the center of the letter and were viewing the largest letter corresponding to the worst acuity measured in our experiments (20/100), the cone density on average would be 13% lower at the edge of the letter than at the center. However, it is unlikely that a subject would have fixated in this manner, and the vast majority of letters shown during the experiments were much smaller than 20/100.

      For completeness, we have calculated the peak cone densities for each of our subjects using the cone density centroid method described by Reiniger et al (2021). We have added these numbers and a description of the method to the Subjects section in Methods and Materials on lines 339-344.

      The observation of visual acuity maintenance with cone dropout has been a longstanding mystery since the 2013/2018 papers by Ratnam and Foote. The authors should be commended for their approach to addressing this important question. However, there are some simplifications and assumptions being applied to make this jump (i.e., that a 50% reduction in cone stimulation in a healthy eye is comparable to a 50% reduction in cone density in a patient). It seems unlikely that, in a patient's eye, with cone dropout, there will be gaps in the mosaic. Not considering any other non-photoreceptor-related reasons for visual acuity loss, which can occur in patients, the cone aperture acceptance angle may be different due to changes in cone size or packing; the sensitivity of individual cones may also be reduced due to deficits in the visual cycle recovery, which could be affected in disease. Some of these limitations could be addressed and acknowledged more explicitly.

      Cone loss does manifest differently in different retinal degenerative diseases, and in this work we implement dropout on a cone-by-cone level. To address the reviewer’s points, we have added a description of the range of spatial manifestations of cone loss across a range of diseases to the Discussion section on lines 320-328, and emphasize that we focus on one particular manifestation in this paper.

      Overall, this is an impressive study incorporating state-of-the-art technology to probe the fundamental limits of human vision.

      We thank the reviewer for their helpful comments and constructive feedback.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The patient recruitment limitation seems to be a bit artificial here. This is indeed a limitation, but perhaps not the primary limitation or motivating factor. Consider removing/rephrasing this motivation.

      We agree with the reviewer’s comment, and have removed it from the abstract and removed its framing as a limitation in the Introduction section in two instances on lines 31 and 36-37.

      Was the peak cone density quantified, and the location of the peak cone density determined? Reporting the range of eccentricities over which the C stimulus lands relative to the peak cone density location, as well as the actual cone density that is being used to sample the C stimulus on the retina, seems to be important for contextualizing this study. It is a bit too simple to only consider the percentage of cones that are reduced.

      We have added the peak cone density for each subject to the Subjects section in Methods and Materials. This experiment did not require fixation; rather, the Landolt C stimulus was fixed in space and the subject could move their eye freely across it, meaning that different parts of the fovea may have sampled the letter on different trials. At the highest dropout percentage, where acuity was the worst, the letter size was 20/100 at threshold, or 25 arcmin. If the subject were to fixate with their peak cone density at the center of the letter, we have computed that the average decrease in cone density at the edge of the letter (12.5 arcmin away) would be 13%. The majority of trials in the experiment showed letters that were much smaller than this, and would have been subject to even less variation in cone density.

      Do the authors have any idea about the approximate size of the cones in healthy subjects compared to diseased eyes? Importantly, if the cones in patients are larger due to the dropout of their neighbors, then the retinal coverage area would be larger due to their larger size, and the amount of light that can be coupled into larger cones may also be larger. Can this be modeled or discussed?

      In our implementation, we did not emulate a change in cone size, and rather modeled the loss as discrete holes in an otherwise intact retina. We have added text to the Discussion section on lines 320-328 to make the distinction between this form of cone loss and other forms where cones appear to fill in for their neighbors resulting in a contiguous mosaic of lower density overall.

      Acknowledging some of the shortcomings of this approach for simulating the patient condition could be improved. It may be worthwhile to tone down the premise of this paper if these cannot be adequately explained.

      In order to tone down the premise of the paper, we have made the following changes to the text.

      We now emphasize on lines 65-69 in the Introduction section that we focus specifically on the impact of cone loss on acuity without modeling downstream factors.

      In addition to the description of other diseases that we added in response to a previous comment, we have also added the following text on lines 313-318 of the Discussion section:

      “... factors beyond the photoreceptors also play a role in shaping vision under retinal degeneration. In this work, we did not model any downstream factors such as shorter outer segments (Foote 2018), retinal rewiring (Jones 2016, Lee 2021), or ganglion cell hyperactivity (Kramer 2023). Instead, we sought to characterize vision in the presence of cone loss at the lowest possible level, considering only the decrease in sampling power at the retinal input.”

      In the methods, it is not completely clear the rationale for determining the appropriate size of the C stimulus. How is visual acuity determined if the C stimulus size is not changed?

      A more careful explanation of how the C stimulus size is set is warranted.

      In the experiments measuring visual acuity, the C stimulus size was selected by a QUEST staircase on each trial. For each dropout condition, we ran 4 interleaved QUEST staircases with 20 trials each. This is described in the “Acuity Threshold Experiment” section in the main text (lines 99-100) and in Materials and Methods (line 407). To clarify further, we have updated the following sentence on line 423:

      “For each condition, we ran 4 interleaved QUEST staircase procedures (Watson and Pelli, 1983) with 20 trials per staircase, which varied the size of the Landolt C on each trial.”

      What is the clinical visual acuity of the subjects being tested? It seems important to report this if the authors want to use their C stimulus as a proxy for clinical visual acuity.

      The subjects being tested have excellent acuity. In Figure 1, we can see that their adaptive-optics-corrected acuity for the baseline 0% dropout condition ranges from approximately 20/10 to 20/12.5 across the 4 subjects. We have added the following statement to the “Subjects” section in Materials and Methods (line 339):

      “All subjects self-reported to have normal vision.”

      Given that the title of the paper emphasizes the role of eye motion, it seems that a more careful analysis of the magnitude and type(s) of eye motion could be added. There are eye motion data provided in the supplemental figure, but it is not completely clear how this eye motion data is actually being used to derive meaningful information about visual acuity.

      We performed analyses to determine whether there seemed to be a significant difference in eye motion patterns between the cone and pixel dropout conditions, which was described in the section “Analysis of Eye Motion Data” and in Supplementary Figure S1. In that figure, we show that for all 4 subjects there is no significant difference in the iso-density contour area containing 68% of their eye motion data. We suggest in the paper that due to the pseudorandom presentation of trials and the limited duration of those trials, subjects were unlikely to adapt or adjust their eye movement, and that instead their natural eye motion served as a data collector that improved acuity.

      What is the accuracy of the eye motion and cone dropout stimulation delivery in the fovea? Given the small size of the cones, it seems that this is one of the most challenging locations of the eye to test with this new olo technology.

      Eye tracking and targeted light delivery are crucial in the AOSLO system and the reviewer is correct to point out that this is most difficult to achieve at the foveal center. To address this concern, we have done some simple modeling and have added the following text to the Cone-by-Cone Stimulation section in the Methods and Materials.

      “This latency, combined with other factors such as diffraction and residual aberrations, limit the ability to restrict the light to only the targeted cone. Considering a 543-nm focus through a 7.2 mm pupil, a random tracking error with a full-width-at-half maximum (FWHM) of 0.5 arcminutes (Harmening et al. (2014)), a 0.0125 diopter residual defocus error (maximum error given the step sizes of 0.025 diopters in the AOSLO defocus controller), an average cone spacing of 0.5 arcminutes (Wang et al. (2019)), and a Gaussian cone acceptance aperture with a FWHM that is 0.5 times the inner segment diameter (Macleod et al. (1992)), we estimate that each targeted cone receives 5.41 times more light than its nearest neighbor. This means that the ’dead’ cones cannot be fully excluded from the visual processing. Furthermore, the light leakage reported in Fong et al. (2025) further adds to the signal of non-targeted cones.

      Nevertheless, it is important to point out that the information about the stimulus (Landolt C in our case) is sampled at the targeted cone’s location and so, although nearby stimulated cones might detect light, they do not contribute to any increases in the sampling process. This is analogous to adding defocus blur to letters in the pixel dropout condition as neither situation will improve the spatial information.”

      References

      Reiniger, J.L., Domdei, N., Holz, F.G., Harmening, W.M.: Human gaze is systematically offset from the center of cone topography. Current Biology 31(18), 4188–4193 (2021)

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is solid, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes, although some claims about density dependence, competition, and stress-gradient mechanisms are not fully supported by the experimental design. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management on the Qinghai-Tibetan Plateau.

      Thank you very much for these positive assessments of our work. Below, we provide point-by-point responses to the comments from the two peer reviewers, and we hope these revisions are satisfied by you and the reviewers.

      Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where two herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      The study was not designed to investigate the reasons that pikas clip Stellera chamaejasme. That said, based on other studies and general knowledge of the ecology of these pikas, it is likely that they clip (although do not eat) this plant because its relatively large size hinders predator detection. This species of pika does better where vegetation height is low than where it is higher.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section.)

      We appreciate these positive comments on our work.

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications of the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists.

      Thank you for this suggestion. We have acknowledged this limitation in the Discussion by adding the paragraph below (see the Third point):

      “Despite of these, several questions deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if the documented facilitation of yak nutrition by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition? Third, our study site is located at approximately 3,200 m elevation, relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 272-286 in the Discussion section.

      Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015).

      Citation:

      Harris RB, Wang, WY, Badinqiuying , Smith AT, Bedunah DJ (2015) Herbivory and Competition of Tibetan Steppe Vegetation in Winter Pasture: Effects of Livestock Exclosure and Plateau Pika Reduction. PLoS ONE 10(7): e0132897. doi:10.1371/journal.pone.0132897

      Thank you for this suggestion. We have added more details about the study site, particularly regarding wild ungulates, in the Methods section. Specifically, we have included the sentence of “Wild ungulates, such as Tibetan gazelles (Procapra picticaudata) (Harris et al., 2015), and other small mammals such as rabbits and zokors, occur rarely in the area.”

      See these revisions in Line 333-335 in the Methods section.

      It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats).

      Thank you for the suggestion. We have acknowledged this limitation in the Discussion, by adding a paragraph as: “Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 284-286 in the Discussion section.

      Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

      Thank you for your suggestion. We have acknowledged this limitation in the Discussion, by adding the paragraph as “Second, if the documented facilitation of yaks by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition (Yang et al., 2026)?”

      See these revisions in Line 276-279 in the Discussion section.

      Reviewer #1 (Recommendations for the authors):

      Although no doubt a bit sensitive, it would have been better to reveal a bit more about how pikas were removed.

      We have provided more details about how pikas were removed in the no-pika treatment, by adding “For the no-pika treatment, pikas were trapped once every two weeks using 30 live traps (25 cm high × 25 cm wide × 40 cm long) within each plot and relocated elsewhere in the study site.” in the Methods section. We didn’t recorded how many pikas were removed from the corresponding plots, so no data were available for this point.

      See these revisions in Line 411-413 in the Methods section.

      The authors also missed a few relevant papers worth citing, including Badingqiuying, R. B. Harris, and A. T. Smith. 2018. Summer habitat use of plateau pikas (Ochotona curzoniae) in response to winter livestock grazing in the alpine steppe Qinghai-Tibetan Plateau. Arctic, Antarctic, and Alpine Research 50 (1): e1447190

      We have cited this key paper in Line 75 in the Introduction section.

      Reviewer #2 (Public review):

      Summary:

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via the use of shared resources.

      Strengths:

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good job of demonstrating the potential conservation impacts of their research.

      We appreciate these positive comments on our work.

      Weaknesses:

      What the authors tested is really cool, but their claims go far beyond what they can say based on their experimental design. For example, the authors claim to show that pika impacts on yaks display density-dependent transitions from competition to facilitation. However, their experiment only looked at the presence (at moderate densities) and absence of pikas, and they only tested for facilitation, not competition.

      The paper would also benefit from changes to the framing in the introduction and discussion. For example, the authors pitch the work as a test of the stress-gradient hypothesis. However, there is no abiotic stress gradient in the study, which is an essential component of the SGH. They also pitch the work in terms of density dependence, but there is no significant variation in population densities beyond the presence-absence binary. The paper would be stronger if they focused their framing around the literature on facilitative interactions across mammals of different size classes, especially indirect facilitation via use of shared resources, which is what this paper is really about.

      We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the Stress Gradient Hypothesis (SGH). Thus, we deleted the description on SGH. However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate densities of pikas has the best beneficial effects on yak growth. We added a separate paragraph in discussion to have a clear discussion.

      We have made the major revisions below to address these concerns.

      (1) We have revised the title into “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands ”.

      (2) We have deleted all the statements about facilitation and competition predicted by the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We added a paragraph in discussion (Line 248-259) to have a clear discussion on the humped relation between yak weight gains and pika burrow densities as “Because of the natural variations in pika density in the pika-present treatment, we were able to obtain a hump-shaped relationship between yak weight gains and pika burrow densities in these plots. Compared with the absence of pikas, the facilitative effect reached its maximum at approximately 200 burrows/ha but became competitive at densities exceeding 400 burrows/ha (Figure 3C). This result reveals that pika density modulates the net outcome for yak weight gain, with a facilitation peak at ~200 burrows/ha and a competition onset above 400 burrows/ha. Our findings offer empirical evidence for the non-monotonicity theory, under which the competition-facilitation balance varies with population density: facilitation dominates at low densities, competition at high densities, and these density-dependent shifts may underpin community stability and productivity (Zhang, 2003; Zhang et al., 2015). The theory further holds that the facilitation threshold, not the competition-facilitation transition, is the critical factor governing the stability of interacting species or communities (Zhang et al., 2015).”

      Most importantly, there are inconsistencies in what is visualized in the figures compared to the model results. For example, the results section in several places notes a lack of significant interaction terms in the model but shows interactions in the p-values on the figures.

      In the Results section, there are only two places where we discussed non-significant interactions: Line 175–177 “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Figure 3F, I and Appendix 1—figure 1, table 5,8).” and Line 190–192 “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (Appendix 1—figure 2, table 10).”.

      We have cross-checked both the Results section and the Figures sections mentioned above, and confirmed that they are consistent now.

      The authors also plot smoothed lines rather than their model results and then draw interpretations from those lines that cannot be tested in the models that they used.

      Thank you for the suggestion. Now we have added the Appendix 1—table 3 and Appendix 1—table 7 for the model results of generalized additive models (GAMs) for Figure 2C and Figure 3C that plotted with smoothed lines in Appendix 1.

      There are also missing details that are important for model interpretation, including the distributions used and the sample sizes.

      We have provided the Appendix 1—table 13 to summarize all statistical models used in the study, including the distributions used and the sample sizes in the Appendix 1.

      We have also added a sentence of “A summary of all statistical models used in the study is available in Appendix 1 table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Another major concern with experimental design is in the forage nutrient analyses. The authors picked plants along a grazing trail, then measured nutrient content without standardizing based on plant species, so any differences across treatments could be because of what they happened to grab rather than overall forage quality.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g). We have revised this section as below.

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Figure 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) Introduction

      Line 53 - I wouldn't describe small mammals like rodents as keystone species. They can have strong impacts on ecosystems, but not disproportionate relative to population size, which is a key part of that definition. It may be true when you talk specifically about pikas later on, but not small mammals as a general category.

      We have replaced this term with “consumers” here, see Line 63.

      Lines 58-61 - Good hook

      Thank you for this positive comment.

      Lines 69-74 - I don't think the stress gradient hypothesis is the right pitch for this work. The SGH posits that facilitation increases with abiotic stress, but no abiotic stressors were measured in this study. Population density interacts with abiotic factors in the SGH, but population density in and of itself is not an abiotic stressor. The papers you cite here all look at the interaction between population density and abiotic stressors (e.g., water availability). So these lines set me up to expect a stress gradient in your experiment that didn't exist, then left me confused later on. It would be better to highlight the strength of your work (lines 70-76 pose interesting questions and predictions) rather than trying to make it fit within the SGH.

      Thank you for pointing out this problem. We have deleted the statements about SGH here.

      See these revisions in Line 88-91 in the Introduction section.

      (2) Materials and Methods

      Lines 498-509 - Did you verify beforehand that no plants within the enclosures had been grazed on? How did you know that the consumption or clipping was specifically from those pikas?

      We have clarified here by adding “Before cage installation, we carefully checked the plants within each plot and removed those that had been previously grazed or damaged by herbivores.” in the Methods section. In this case, we made it sure that the consumption or clipping was specifically from those pikas with the cages.

      See the revisions in Line 349-350 in the Methods section.

      Lines 509-512 - I would be careful calling this preference. It's really just a record of what they consumed along paths they were walking, which could be about accessibility and convenience as much as preference.

      We have replaced “diet preferences” with “diet composition” here, see Line 359 in the Methods section.

      Line 514 - Clarify the specific question or hypothesis you're testing with this field survey, beyond just generally testing associations.

      We have modified the sentences here as “In July 2021, we investigated the potential facilitation of pikas on yaks mediated by the poisonous Stellera forbs under unmanipulated field conditions in the study site.”.

      See Line 365-366 in the Methods section.

      Lines 532-534 - The intro for this paper sets it up to be about density-dependent movement from competition to facilitation, but the experiment here is set up to compare pika presence/absence. The framing of the paper needs to be adjusted to better align with this experimental design.

      The same issue as mentioned above. We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the balance between competition and facilitation as predicted by the Stress Gradient Hypothesis (SGH). However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate density of pika has the best benefical effect on yak. We added an separate paragraph in discussion to have a clear discussion about this point in the Discussion section.

      We have made the major revisions below to address this concern.

      (1) We have revised the title as “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands”.

      (2) We have deleted the statements about facilitation and competition and the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We kept the discussion on the humped relation between yak weight gains and pika burrow densities (Figure 3C). We added an separate paragraph in discussion to have a clear discussion about this point. For details, see Line 248-259 in the Discussion section.

      Lines 566-569 - Why did you need to simulate pika clipping when you already had pika presence/absence treatments?

      We conducted Stellera removal treatment by simulating poisonous plant clipping behaviors of pikas because we want to confirm that the shifts in abundance of this dominant poisonous plant species is the key mechanism in driving pika-yak facilitation in our system. If we simply looked the differences in yak weight gain in the pika presence/absence treatments, it should be difficult to secure the underlying mechanism. In addition to the reduction in abundance of the poisonous Stellera, pikas may cause a variety of shifts vegetation properties including plant productivity and diversity, and soil disturbances that can exert direct and indirect effects on yak foraging activities, and thus their weight gains.

      Lines 566-569 - Is there another citation you can give showing that Stellera forbs taller than 20cm are both preferred by pikas and exert greater impacts on plant/animal communities? Those are big assumptions that need to be better supported or explained. If there was a logistical reason that you didn't remove all of the smaller forbs, that needs to be laid out as well.

      We have added one new citation here to support this method here. We have revised this section as “To simulate the clipping behavior of pikas, we clipped only those Stellera forbs exceeding 20 cm in height. This threshold was chosen based on previous observations that pikas preferentially target large forbs of this size (Liu et al., 2009).”

      Liu W, Zhang Y, Wang X, Zhao JZ, Xu QM, Zhou L. 2009. The relationship of the harvesting behavior of plateau pikas with the plant community (In Chinese). Acta Theriologica Sinica 29:40-49.

      Also, we have deleted the sentence of “and can exert significant impacts on the plant community and on yak grazing behaviors (Z.Z., field observations)” mentioned above, because these patterns were observed only by the authors in the field and lack supporting data.

      See these revisions in Line 419-422 in the Methods section.

      Line 587 - sample size per treatment? Was it consistently one species that you measured, and if so, which one? If you collected different species or a mix of species across treatments, then you can't really compare the nutrient values because you haven't accounted for interspecific variation.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g).

      We have revised this section as below:

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Fig. 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Lines 603-613 - Please provide the dependent variables in each model, as well as any interaction terms, in addition to the random effects. Please also state explicitly what you were trying to test with each of these models.

      There are two models that use the tweedie family (forb and sedge bite rate). Indeed, we need to include the tweedie power parameter to help understand the mixture of the three families. We have included the p-value (power) in the Appendix 1—table 13 in the Appendix 1. We chose to use tweedie because a normal gaussian family model fitted the results poorly.

      We have added the Appendix 1—table 13 in the Appendix 1, which provides the summary of all statistical models used in the study, including response variables, model type, distribution family (with Tweedie power parameter where applicable), interaction terms, random effects structure, and sample sizes.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Lines 613-614 - Tweedie is a category of distributions that includes quite a few different options, including Gaussian, Poisson, and Gamma distributions, some of which are normal and some of which are not. So, justifying the use of Tweedie distributions in your model structure doesn't really make sense, and it doesn't really tell me which distribution each model pulled from. Please clarify specifically which distributions you used for each model and why.

      We have now added the reason why we used Tweedie distributions by adding the sentence of “There were two models that used the tweedie family (forb and sedge bite rate). We chose to use tweedie because a normal gaussian family model fitted the results poorly” in Line 468-470 in the Statistical analyses section.

      We have also included the p-value (power) for forb and sedge bite rate in the Appendix 1—table 13 in the Appendix 1.

      Lines 615-618 - Provide citations for R packages described in the text.

      We have provided all the related citations for all R packages described in the text, as listed below.

      glmmTMB: Brooks, M. E., Kristensen, K., van Benthem, K. J., Magnusson, A., Berg, C. W., Nielsen, A., Skaug, H. J., Maechler, M., & Bolker, B. M. (2017). glmmTMB Balances Speed and Flexibility Among Packages for Zero-inflated Generalized Linear Mixed Modeling. The R Journal, 9(2), 378–400. https://doi.org/10.32614/RJ-2017-066

      mgcv: Wood, S.N. (2017). Generalized Additive Models: An Introduction with R (2nd edition). CRC Press. AND Wood, S.N. (2011). Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 73(1), 3–36. https://doi.org/10.1111/j.1467-9868.2010.00749.x

      DHARMa: Hartig, F. (2022). DHARMa: Residual Diagnostics for Hierarchical (Multi-Level / Mixed) Regression Models. R package version 0.4.6. https://CRAN.R-project.org/package=DHARMa

      tidyverse: Wickham, H., Averick, M., Bryan, J., Chang, W., McGowan, L.D., François, R., Grolemund, G., Hayes, A., Henry, L., Hester, J., Kuhn, M., Pedersen, T.L., Miller, E., Bache, S.M., Müller, K., Ooms, J., Robinson, D., Seidel, D.P., Spinu, V., Takahashi, K., Vaughan, D., Wilke, C., Woo, K., & Yutani, H. (2019). Welcome to the tidyverse. Journal of Open Source Software, 4(43), 1686. https://doi.org/10.21105/joss.01686

      See these revisions in Line 472-475 in the Statistical Analyses section.

      (3) Results

      Line 127 - Somewhere in the intro or methods, describe what clipping is and why the pikas do it.

      We have provided more details about the clipping behaviors of pikas by adding “Notably, pikas often clip (although do not eat) the wolf poison S. chamaejasmehas because its relatively large size hinders predator detection (Fan et al., 1998).” in Line 330-332 in the Methods section.

      Line 128 - Change from "preferred" to "consumed greater proportions of"

      See the correction in Line 146 in the Results section.

      Line 136 - You measured yak weight once a month, so give the result in monthly weight gain rather than daily.

      Here we preferred to keep the unit of daily weight gains (converted from monthly ones), as this is the standard presentation for livestock growth performance, also see Fig. 1 in Odadi et al., 2011 Science’s paper.

      W. O. Odadi, M. K. Karachi, S. A. Abdulrazak, T. P. Young, Science 333, 1753–1755 (2011).

      Lines 138-140 - The linear models, as you described them in the methods (lines 603-620), don't test for hump-shaped relationships. Please update the methods to explain how you tested the density relationship and how you got this interpretation.

      The description for Figure 3C here showed estimates from a GAM (family: gaussian), and we did not use a linear model in this figure.

      We have now added a new model summary Appendix 1—table 7 for this Figure 3C in the Appendix 1.

      Line 145 - I don't think you can claim that the total available forage was more nutritious for yaks, as you picked plants that yaks happened to be chewing along a grazing path. You would need to take samples from a consistent set of plant species at random locations to make this claim.

      Sorry for this confusion. We modified “the total available forage” as “the major available forage” here (see Line 172) because we collected the same forage plant species and analyzed their nutrients.

      The same as mentioned above, to clarify the sampling methods, we have also revised this point in the Methods section to provide more detail on how forage samples were collected and their quality were analyzed. See these revisions in Line 439-447 in the Methods section.

      (4) Discussion

      Line 166 - Need to address inconsistency in how you talk about pika density. Here you talk about the impacts of pikas at moderate densities, which I think is a fair claim. Elsewhere, you talk about density-dependence, which I don't think you really measured, given that all your sampling was either in the absence of pikas or within a narrow window of densities that can all be categorized as moderate.

      Done! As mentioned above, we have removed the term of “density-dependence” in the whole manuscript, but keep the term of “moderate density” in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph was removed here), and the References section.

      Line 176-178 - This is a really cool finding.

      Thank you for this positive comment.

      Lines 179-181 - You didn't test competition between plant species, and you didn't measure light, soil moisture, or soil nutrients. So you can suggest competition as a potential mechanism, but you can't say definitively that's what is happening.

      We now have lowered our tone here as “We speculate that these improvements in food availability and nutrition for yaks may be due to the release of grasses and sedges from competition with the forbs for limiting above- and below-ground resources”.

      See the revision in Line 208-211 in the Discussion section.

      Lines 184-186 - You did a good job of it here, suggesting a likely potential mechanism at play without claiming it is for sure happening when it hasn't been measured.

      Thank you for this positive comment!

      Lines 188-199 - Strong paragraph. The impacts of large herbivores on smaller animals are well-studied, but reciprocal impacts are often overlooked.

      Thank you for this positive comment!

      Lines 201-205 - You didn't test the stress gradient hypothesis because there was no abiotic gradient. You also did not take any measurements during outbreaks, so you cannot claim to have compared low-moderate to outbreak pika densities. I think the paper would be much stronger if you removed the stress gradient hypothesis and instead focused more on the literature around facilitation between mammals of different body sizes, as you do in lines 205-209.

      We agreed that our work didn’t specifically design to test the stress gradient hypothesis (SGH) between pikas and yaks, so we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Line 209-215 - Again, you didn't test a competition-facilitation balance because you never tested or demonstrated competition. One of the main strengths of this paper is demonstrating facilitation, so build on that strength rather than referencing things you didn't measure.

      The same as mentioned above, we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Lines 218-222 - Not an accurate description of the relationship between herbivore diet and body size. Larger herbivores typically tolerate lower-quality plants in order to consume sufficient calories, but plenty of them do this via mixed feeding. Grazing in large herbivores and livestock is usually due to specifics of the digestive tract (e.g., hindgut fermentation) rather than specifically about body size.

      We have deleted this description of the relationship between herbivore diet and body size here. Instead, we have modified this statement as “The coexistence of a diverse of herbivore species with different diet selections and size classes can lead to an “compensatory effect” on grass and forb biomass that helps to maintain a balance and diverse plant community” in the Discussion section.

      See these revisions in Line 234-237 in the Discussion section.

      Lines 245-251 - Paragraph addresses an important point. Lines 247-249, though, overstate what you measured. There's no measurement of livestock production or biodiversity in the study.

      We have replaced the term of “livestock production and biodiversity” with “livestock growth performance” here. See Line 288-298 in the Discussion section.

      (5) Figures

      Figure 2C-D - This applies to all figures, but you need to plot the best-fit line generated from your model instead of using geom_smooth, which is what these lines look like. You can do this using functions like predict or ggpredict. Using these smoothed lines implied non-linear relationships that you didn't actually test for.

      We have redrawn Figures 2C and 2D to use model estimates directly and have included the related model summaries as Appendix 1—table 3 and table 4 in the Appendix 1.

      Figure 3B - This figure doesn't show yak weight gain in the presence of pikas. Instead, it shows weight loss when pikas are absent. It's a subtle difference, but very important for interpretation. Yak can maintain weight just fine without pikas as long as Stellera are absent, too. Your results consistently show no Pika x Stellera interactions, but that doesn't match your significance values here. Need to double-check and explain that.

      We have revised the descriptions for Figure 3 and 4 in the Results section, by emphasizing that the absence of pikas REDUCED weight gains of yaks, INCREASED toxic plant abundance, and REDUCED the quantity and quality of palatable grasses and sedges. We have revised these sections as below:

      Abstract (see Line 51-53)

      “Compared to the pika-present treatment, pika removal dramatically increased cover of the poisonous Stellera forbs by two-fold, reducing the abundance and protein content of palatable grasses and sedges, yak foraging efficiency, and yak weight gain by up to 42%.”

      Results (see Line 153-167, Line 169-175, Line 184-192)

      Also, we did find significant Pika x Stellera interactions for yak weight gains, we have provided these details in the Appendix 1—table 5 and table 6 in the Appendix 1.

      (6) Recommendations

      Lines 245-257 - Would recommend combining the last two paragraphs into one.

      We have combined the last two paragraphs, see Line 288-298 in the Discussion section.

      Line 530 - Can you replace large with a more precise measure of area?

      We are unable to provide a more precise measure of area here, so we have deleted the description of “in a large area”, but we have also added the note of “in the study site” by the end of the sentence to better describe the location of the plots.

      See the revision in Line 413 in the Methods section.

      Line 541 - Does this mean +/- 7.8 standard deviations? If so, how big is that range in kg?

      Here should be “115±7.8 kg”, we have done this correction in Line 392 in the Methods section.

      Figure 3 - Would be helpful to use "Stellera/pika present" and "Stellera/pika absent" rather than saying "No Stellera/pika" since you also use the No. abbreviation for numbers a lot in this figure.

      We have revised all the related Figures for this issue in Figure 3, 4, and Appendix 1—figure 1,2.

      Figure 3F and I - Show letters for significance on these two plots as well, even if it is just a row of a's.

      We have redrawn Figure 3F,I to address this issue.

      Figure S1 - Applies to all boxplots. Be consistent about showing significance, even if it is a row of a's indicating no difference between treatments.

      We have re-drawn Appendix 1—figure 1,2 to address this issue.

      Table S1 - Something happened with the line numbers, so they are in the table instead of on the left side.

      We have fixed this problem for Appendix 1—table 1 in the Appendix 1.

      Table S1 - Applies to all tables. Include the type of model that you ran (including distribution if not Gaussian) in the table legend.

      We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Table S9 - This legend has a good description, including the type of model you used and what you were testing. Apply this more detailed legend to the rest of the tables.

      Again! We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides new insight into the regulation of cell organization and division in Trypanosoma brucei through the control of a kinesin motor protein by a polo-like kinase. The authors present solid evidence from rigorous biochemical and imaging analyses showing that phosphorylation modulates kinesin function and cellular organization. However, direct in vivo evidence that PLK phosphorylates kinesin-G is lacking.

      We performed experiments to investigate the effect of PLK inhibition on the phosphorylation of KIN-G in vivo in trypanosome cells by immunoprecipitation and mass spectrometry. The new results showed that treatment of trypanosome cells with GW843682X, a potent TbPLK inhibitor validated previously in procyclic trypanosomes, reduced the phosphorylation levels on Thr301 and Ser569 of KIN-G by ~27% and 100%, respectively. The partial reduction in Thr301 phosphorylation after GW843682X treatment could be attributed to slower dephosphorylation of phosphorylated Thr301 after GW843682X was added to the cell culture. Nonetheless, these new results demonstrated that KIN-G is an in vivo substrate of PLK.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript identifies the orphan kinesin KIN-G as a substrate of Polo-like kinase (TbPLK) in Trypanosoma brucei and demonstrates that phosphorylation of Thr301 inhibits KIN-G microtubule binding and disrupts its cellular function. Using a combination of in vitro kinase assays, phosphosite mapping, microtubule binding and gliding assays, and in vivo complementation with phosphomimetic and phosphodeficient mutants, the authors link TbPLK-mediated regulation of KIN-G to defects in centrin arm integrity, FAZ elongation, Golgi organization, flagellum positioning, and division plane placement. The study provides a mechanistic advance in understanding how TbPLK regulates centrin arm biogenesis and integrates KIN-G into the growing regulatory network controlling hook complex and FAZ assembly. Overall, the work is technically strong, internally consistent, and builds logically on previous studies from this group and others.

      Strengths:

      A major strength of the manuscript is the clear mechanistic link between phosphoryltion of Thr301 and loss of microtubule binding activity. The use of phosphomimetic (T301D) and phosphodeficient (T301A) mutants in an RNAi-rescue framework provides a clean and convincing demonstration of functional relevance in vivo. The integration of biochemical assays with detailed cell biological phenotyping (centrin arm length, FAZ elongation, basal body segregation, and cytokinesis markers) is particularly effective and makes the central conclusion robust. The observed phenotypic cascade from centrin arm defects to FAZ and division plane abnormalities is also well aligned with existing models of trypanosome morphogenesis.

      Weaknesses:

      My (more or less main) concern relates to the interpretation of the Golgi phenotype. The conclusion that phosphorylation of KIN-G "impairs Golgi biogenesis" is currently based on fluorescence microscopy using TbGRASP and Sec13 markers and on quantification of the number and distribution of Golgi/ERES puncta in binucleated cells. While these data convincingly demonstrate altered Golgi/ERES number and spatial organization, they do not distinguish between true defects in Golgi biogenesis or duplication and alternative possibilities such as fragmentation, vesiculation, or mislocalization of Golgi membranes. Given the central role of Golgi-centrin arm organization in the proposed model, ultrastructural analysis (for example, by EM or electron tomography) would greatly strengthen this aspect of the study by providing direct evidence for structural alterations of the Golgi and its association with the centrin arm and ERES. Such data would elevate this part of the manuscript from a descriptive fluorescence phenotype to a true structural cell biological insight. I appreciate that this experiment goes beyond the current dataset, but it would substantially enhance the mechanistic depth of the Golgi-related conclusions and strengthen the causal chain linking centrin arm defects to Golgi abnormalities. However, I have to confess, the inclusion of such data would make this reviewer particularly enthusiastic about the work. If this is not feasible, I would recommend tempering the wording of "Golgi biogenesis" to a more conservative description, such as altered Golgi organization or duplication, and explicitly acknowledging the limitations of fluorescence-based analysis for this conclusion.

      Thanks for these very constructive comments, which are very well taken. We totally agree with this reviewer on these points. Since it is not feasible for us to perform EM or electron tomography, we have revised the manuscript to describe the effect of KIN-G phosphorylation on the Golgi as “altered Golgi duplication” rather than “Golgi biogenesis”. We also explicitly acknowledge the limitations of fluorescence-based analysis of the Golgi for this conclusion and suggest that further characterization with EM or electron tomography would allow one to reveal the potential structural alterations of the Golgi and its association with the centrin arm.

      An additional conceptual point concerns the dual role of TbPLK in centrin arm regulation. TbPLK is known to promote centrin arm biogenesis through phosphorylation of TbCentrin2, yet in this study, TbPLK phosphorylation of KIN-G negatively regulates centrin arm assembly. This dual positive and negative regulatory role is intriguing but could be discussed more explicitly. The manuscript would benefit from a clearer conceptual framework addressing how phosphorylation of KIN-G might serve as a temporal or spatial switch to restrain KIN-G activity at specific stages of centrin arm assembly.

      This is a great point. However, we are not sure whether the previous work on TbPLK phosphorylation of TbCentrin2 could lead to the conclusion that TbPLK promotes centrin arm biogenesis through phosphorylation of TbCentrin2. In the published work (de Graffenried et al., MBoC, 2013), trypanosome cells expressing the phospho-deficient mutant TbCentrin2-S54A only showed minor growth defects, exhibiting growth defects after 5 days (de Graffenried et al., MBoC, 2013). Cells expressing the phosphomimic mutant TbCentrin2S54D, however, showed very strong growth defects (de Graffenried et al., MBoC, 2013). The effects of TbPLK phosphorylation on TbCentrin2 appear to be quite similar to that of TbPLK phosphorylation on KIN-G, although the KIN-G-T301A mutant does not have growth defects (up to 5 days in our experiments). It appears that the primary role of TbPLK in regulating TbCentrin2 and KIN-G is negative regulation. Nonetheless, we have added more discussion on these regulatory roles of TbPLK in the revised manuscript.

      Finally, a schematic model summarizing the proposed regulatory pathway from TbPLK phosphorylation of KIN-G to centrin arm assembly, FAZ elongation, division plane placement, and Golgi organization would aid the reader.

      Thanks for this suggestion. We made a schematic model to summarize the roles of KIN-G and its regulation by TbPLK. This is included in Figure 8.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as a misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding, and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei supports the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do not formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      This is a great point that is very well taken. We also had been puzzled by the observation of no growth defects of T301A mutant. This comment enlightened us. From the new experiments we performed to compare the phosphorylated peptides of KIN-G in cells treated and non-treated with the PLK inhibitor GW843682X, we calculated the percentage of phosphorylated Thr301 in non-GW843682X-treated cells. We found that the percentage of peptides containing the phosphorylated Thr301 is ~14% of the total Thr301-containing peptides (Fig. 1H). This result indicates that T301-P is indeed a small minority of the population in the asynchronous trypanosome cells. It is possible that T301 phosphorylation may occur at a specific cell cycle stage such as early S-phase, during which PLK and KIN-G co-localize at the centrin arm.

      (b) Published work links PLK to cell division, FAZ elongation, etc.. The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc..

      Yes, previous work discovered essential roles of TbPLK in basal body segregation, centrin arm biogenesis, FAZ elongation, and cytokinesis. These functions of TbPLK correlate with TbPLK’s localization to multiple subcellular structures, the basal body, the centrin arm, and the new FAZ tip, and are attributed to the regulation of its substrates at these structures. At the basal body, TbPLK phosphorylates SPBB1, which is required for basal body segregation. Defects in basal body segregation can lead to defective flagellum positioning and FAZ elongation. At the new FAZ tip, TbPLK regulates the cytokinesis regulator CIF1, which is required for cytokinesis. At the centrin arm, TbPLK phosphorylates TbCentrin2 at S54 and KIN-G at T301 (and S569, which was newly identified as an in vivo TbPLK site and has not yet been characterized). However, cells expressing TbCentrin2-S54A have very weak growth defects, and cells expressing KIN-G-T301A have no detectable growth defects. In contrast, cells expressing TbCentrin2-S54D and cells expressing KIN-G-T301D have strong growth defects. Therefore, the essential role of TbPLK in centrin arm biogenesis apparently is not attributed to the phosphorylation of TbCentrin2 and KIN-G. It is possible that phosphorylation of other centrin arm-localized protein(s) by TbPLK may be essential for centrin arm biogenesis, but this possibility remains to be explored.

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      We performed experiments and presented the data in Fig. 1H. We also included commentary in the revised manuscript on the points about the potential role of T301 phosphorylation. Thanks for these great comments that significantly improved the manuscript.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      We treated trypanosome cells with a potent PLK inhibitor GW843682X, which was previously demonstrated to inhibit TbPLK activity in vitro and mimic TbPLK knockdown in vivo in trypanosomes, and immunoprecipitated KIN-G for mass spectrometry. We compared the KIN-G peptides identified by mass spectrometry from trypanosome cells treated with or without GW843682X, and found that two phosphosites (T301 and S569) were reduced by ~27% and 100%, respectively, after GW843682X treatment (Fig. 1G). These results provided evidence to support that TbPLK phosphorylates KIN-G in vivo.

      Reviewer #3 (Public review):

      Summary:

      Here, the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during the early S-phase of the cell cycle. Centrin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescence to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      Some of the broader conclusions are not directly supported by the data. For example, the title states "Polo-like kinase phosphorylation of the orphan kinesin KIN-G negatively regulates centrin arm biogenesis in Trypanosoma brucei," but the data do not directly address the specific role of TbPLK in phosphorylating KIN-G in cells. Moreover, some of the more specific conclusions in the paper, for example, that "phosphorylation of KIN-G" causes various cellular defects, are a bit of an overstatement. The supporting data rely on the expression of a phospho-mimetic mutant of KIN-G. Presumably, phosphorylation in cells is a normal part of KIN-G regulation, and it is not just phosphorylation, but rather hyperphosphorylation that is being mimicked by the mutant. Some rewording of the specific conclusions is warranted, and the broader conclusion would be better supported with additional experimental evidence.

      This is a great point that is very well taken. We performed new experiments to address the in vivo phosphorylation of KIN-G by TbPLK. We treated cells with a potent PLK inhibitor, GW843682X, and then immunoprecipitated KIN-G for mass spectrometry to identify changes in phosphorylation. We found that the phosphorylation levels of T301 and S569 were reduced by ~27% and 100%, respectively, confirming that these two sites are in vivo TbPLK phosphosites.

      We also calculated the ratio of phospho-T301 versus non-phospho-T301 in non-treated cells and found that phospho-T301 accounts for ~14% of the total KIN-G protein. This new result suggests that it is the hyperphosphorylation that causes growth defects. We have revised the manuscript accordingly to reflect this point.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Several statements use rather strong causal language (for example, "thereby impairing Golgi biogenesis, FAZ elongation, and division plane placement"). While the phenotypic correlations are convincing, direct causality is largely inferred from prior literature. Slightly tempering this wording could improve precision. It would also be helpful to state explicitly in figure legends the number of cells analyzed per condition and the number of independent experiments for each quantified phenotype.

      Thanks for these constructive comments, which we agree and appreciate greatly. We have revised the manuscript accordingly to improve precision. The total number of cells analyzed, and the number of independent experiments were included in the figure legends.

      Reviewer #2 (Recommendations for the authors):

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      Thanks for this comment. We have deleted the discussion about TbPLK activities that were published previously.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Figure 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Figure 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces the motility of KIN-G. These statements are contradictory.

      We meant to say that there was a slight but insignificant effect. We agree that such statements are somewhat contradictory and, hence, have been deleted. Thanks.

      (3) p. 8, and Figure 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      We have deleted the wording “ventral side”, as it is not necessary. Thanks.

      (4) Figure 7B. Please explain the labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      Thanks very much. We included two sentences in the revised manuscript to explain this point.

      The sentences read as follows: “The nascent posterior is formed near the mid-portion of the NFD cell through microtubule bundling and cytoskeleton remodeling during late stages of the cell cycle (Wheeler et al., 2013). Consequently, the NFD cell inherits the old, existing cell posterior, whereas the OFD cell inherits the newly formed or nascent cell posterior.”

      (5) Figure 4, 5, and 7: "% Cells" is reported. Please indicate the total number of cells that were examined.

      The total number of cells were included in the figure legends.

      Reviewer #3 (Recommendations for the authors):

      (1) The manuscript should be carefully edited for minor grammatical errors.

      Thanks. We have carefully proofread the manuscript and corrected the grammatical errors.

      (2) A general conclusion is that TbPLK phosphorylation of KIN-G in cells is critical for regulating its motor activity. However, this relies on the expression of phospho-mimetic mutants, which bypass TbPLK. Thus, there really is no direct evidence provided to support the specific role of TbPLK other than the in vitro phosphorylation data. Some additional experiments to assess the specific role of TbPLK in phosphorylating KIN-G in cells would lend support for the general conclusion. Is it possible to deplete or inhibit TbPLK and show that this impacts the phosphorylation of KIN-G in cells?

      This is a great point that is very well taken. We performed a new experiment by inhibiting TbPLK with a potent PLK inhibitor, GW843682X, and then immunoprecipitating KIN-G for mass spectrometry to identify the phosphorylation levels before and after GW843682X treatment. We were able to confirm that T301 and S569 phosphorylation was reduced after treatment. This confirms that TbPLK phosphorylates T301 and S569 of KIN-G in vivo in trypanosome cells.

      Specific:

      (1) Figure 2: For microtubule gliding assays, representative videos should be included as supplementary data. Also, when the data do not show a significant difference between KIN-G and KIN-G + TbPLD-K70R, then the authors should not state there is a slight difference, as this is not supported by the data.

      We have deleted the statement. We have included the representative videos for all the experiments presented in Figures 2 and 3. Thanks.

      (2) Figure 3: As stated above, for microtubule gliding assays, representative videos should be included. Moreover, for KIN-G-T301A, the authors say that the gliding activity was "moderately, but insignificantly, reduced." Again, if the difference is not significant, it cannot be concluded that there is a difference compared with the wild-type protein.

      We have deleted the statement. Thanks.

      (3) Some of the headings in the Results section are not accurate. Regarding the in vivo results, the heading "Phosphorylation of Thr301 in KIN-G by TbPLK causes defective cell proliferation" seems to be an overstatement. It may be more accurate to state that "Expression of a phospho-mimetic mutant of KIN-G causes defective cell proliferation." The same comment applies to the other headings that follow this one. Presumably, there is a population of phosphorylated KIN-G in cells, and phosphorylation/dephosphorylation is a normal part of its regulation. In the text, the authors might more accurately conclude that hyperphosphorylation causes the defects they are seeing.

      This is a great point. We agree and as we responded above, we have revised the manuscript, from the title to the main text, to reflect the point that hyperphosphorylation of T301 by TbPLK causes the defects. Thanks very much for these great comments!

    1. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback. We are grateful that the experimental design and the evidence for genetically canalized nutrient resorption efficiency (NuRE) were recognized as compelling and important, with clear implications for predicting wetland nutrient cycling under salinization. We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them, with particular attention to the weaknesses noted.

      Reviewer #1 raised two important concerns regarding the scope and generality of our conclusions. First, the salinity treatment spanned only one growing season, so the genetic canalization we document specifically refers to the absence of a plastic response to an acute salt shock; whether long-term, chronic or multigenerational salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question. Second, the test of nutrient limitation control relies on resorbed N:P and N: K ratios as proxies rather than direct nutrient manipulation, and the metabolomic analysis is used primarily to validate stress effectiveness. We accept these criticisms and will address them as follows.

      Regarding the temporal scope of the treatment, we will explicitly state in the Discussion that our conclusion of canalization pertains to short-term acclimation to an acute salt shock, and we will discuss the scenarios under which chronic, more severe or multigenerational exposure could trigger plastic, acclimatory or transgenerational responses. Given our team’s prior work on parental and transgenerational effects in clonal plants, we will frame these as testable hypotheses for future research. We will also acknowledge that the physiological mechanisms underlying the lack of a plastic increase in NuRE (for example, phloem loading or senescence-associated gene expression) are not directly resolved in the present study and will propose targeted molecular investigations as a natural next step.

      Regarding the nutrient limitation tests, we will clearly acknowledge in the Discussion that the resorbed N:P and N:K ratios provide an established but indirect proxy for nutrient limitation, and we will discuss how direct nutrient addition experiments could provide stronger causal evidence, while deepening the integration of the metabolomic profiles with genotype-level NuRE variation where feasible. We will also quantitatively assess the potential collinearity between ecotype and phylogeographic lineage among the Chinese populations.

      Reviewer #2 raised three substantive issues regarding the interpretation and presentation of our results. First, although latitude emerged as a significant predictor in the variation partitioning, the relatively low R² values indicate that much variance remains unexplained, and our interpretation should be more cautious; in addition, the functional significance of the “inverted” nutrient limitation (slopes greater than 1 for resorbed versus green N:P) needs further explanation. Second, the single moderate salinity level (10 ppt) does not allow assessment of whether more extreme stress might trigger a plastic response. Third, the ecotype analysis applies only to the Chinese populations, as ecotype classification was not available for non-Chinese populations, and this should be stated explicitly in the Results. We fully agree with these points and will revise accordingly.

      To address the concern about latitude, we will temper the language in the Results and Discussion, explicitly noting the limited proportion of variance explained by latitude and acknowledging the role of unmeasured factors, while retaining the study’s central message that genetic and phylogeographic origin outweigh short-term plasticity in shaping NuRE. We will also expand the Discussion to explain the functional significance of the “inverted” nutrient limitation, that is, why P and K are resorbed more completely relative to N, whether as a strategy to maintain optimal N:P:K ratios or a reflection of the higher costs and lower availability of N.

      To address the salinity gradient concern, we will acknowledge in the Discussion that the single moderate salinity level may not have been severe enough to trigger a plastic response and will justify future dose-response experiments to identify the threshold at which canalization of NuRE might be overcome.

      To address the ecotype limitation, we will explicitly state in the Results that the ecotype analysis in Figure 4 is based only on the Chinese populations, for which ecotype classification was available.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study uses technically compelling long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The authors further propose the provocative model that post-translational mechanisms, rather than the transcriptional-translational processes, may contribute to circadian regulation of neuronal excitability.

      We are pleased that our study is recognized as valuable and that our technically very challenging long-term in vivo recordings and computational modeling are appreciated. We agree that we are proposing a provocative model that opposes the current hypothesis in chronobiology, which suggests that all observed biological circadian rhythms are outputs of a transcriptional-translational feedback loop (TTFL) clock. Instead, we suggest that a cell comprises, in addition to the TTFL clock, other posttranslational feedback loop (PTFL) clocks without the need of daily transcription and daily degradation of its core elements. While the circadian TTFL clock is entrained to the daily light-dark cycle, the circadian PTFL clocks are suggested to be entrained to other daily cues such as to the availability of pheromone, or the daily changing levels of hormones and second messenger levels. Our novel hypothesis proposes that the TTFL and PTFL clocks are coupled and linked, constituting an adaptive, plastic network that can tune and phase-lock to different external and internal Zeitgeber signals. However, we certainly do not claim that the TTFL circadian clock is not at all involved in the circadian control of the ORN’s circadian membrane potential rhythms. We clarified our manuscript accordingly.

      However, the evidence for circadian firing in these neurons […] remains incomplete.

      As requested by the reviewers during the previous round of review, we had provided the results of RAIN analysis (Thaben and Westermark, 2014) of individual animals in our first revision (Fig. 4A), which clearly shows that two-thirds of the population express circadian rhythms in key attributes that we used to characterize the spontaneous spiking activity. In the initial review it was assumed by the reviewers that phase alignment of dispersed rhythms would bias interpretations of rhythmicity. After having established with RAIN that individuals show circadian rhythms, albeit dispersed across the population due to the lack of a zeitgeber in DD conditions, phase-alignment of recordings from DD animals is a valid next step to prepare the data for statistical analysis across the population. Phase-alignment of desynchronized rhythms is a generally accepted and necessary method employed in chronobiology (e.g., for insect ORNs: Gosh et al., 2024). It is proven as prerequisite to find and statistically analyze rhythmicity in complex, desynchronized data.

      As we explained in the previous rebuttal and in our first revision, rhythms in electrical activity of insect ORNs cannot be easily synchronized by the light-dark cycle alone, but appear to require daily cycles of pheromone, as shown in other moth species (Gosh et al., 2024). Please be aware that the animals that we used here have never been exposed to pheromone, as we state in the Methods. Therefore, this lack of pheromone exposure can explain why about one third of our experimental population is not expressing any daily or circadian rhythmicity in spiking attributes. This is an important result of our manuscript, reported for the first time for Manduca sexta, providing evidence for our hypothesis that it is not the LD-entrained circadian TTFL clock that governs electrical activity rhythms in ORNs. As we explained in the first revision, and clarified further here in the second revision, we cannot phase-synchronize our animals with cycles of pheromone application in our experimental paradigm because we are researching circadian rhythms in spontaneous spiking activity and not pheromone responses. We failed to obtain phase-alignment with a single pheromone pulse the night before the experiments started. These data were added as supplementary Figure to Fig. 3 in the first revision. Here, we further revised our manuscript to clarify this important finding.

      Thus, as requested by the reviewers in the initial review, we could successfully confirm our previous results of circadian firing in ORNs and the disruption of these circadian rhythms with Orco antagonist OLC15 with RAIN. In the current review, the reviewers raise no further specific critical points or comments that would doubt our careful rhythm analysis of our long-term recordings. Thus, we conclude that we provided clear evidence for our central, exciting new finding. For the first time we demonstrated an unexpected new task for Orco: Orco controls the circadian firing pattern in the spontaneous activity, and thus, of the ORN’s membrane potential, via its property as leak/pacemaker channel.

      However, the evidence […] for post-translational modification of Orco as the underlying mechanism remains incomplete.

      We agree with the reviewers that there are many more experiments and combined efforts of biochemists, structural biologists, and electrophysiologists required to provide complete evidence for post-translational modification of Orco and to reveal the underlying mechanism of its circadian control. It is beyond the scope of the current manuscript to provide all details of post-translational control of Orco.

      The reviewers asked previously for additional evidence that Orco transcription is not controlled via the TTFL clock. As requested, we provided extended qPCR–based evidence that Orco, in contrast to timeless, is not controlled by the TTFL clock on the transcriptional level (Fig. 6 in Revision 1). Furthermore, we added a new result in Revision 1 to demonstrate cAMP-dependent post-translational modulation of Orco open-time probability (Fig 9 in Revision 1) at a ZT at which antennal cAMP levels are low (Schendzielorz et al., 2015). We already showed in Flecke et al., 2010, that the addition of cAMP at different ZTs increases the spontaneous spiking activity only at specific ZTs. Here, we show that the effect of cAMP depends on Orco. Since Orco’s circadian role is not mediated via TTFL control, it can be concluded that post-translational mechanisms provide daily/circadian temporal control. In this additional Figure we provide statistically significant proof that, in agreement with our model-prediction, Orco’s circadian control of the ORN spontaneous activity could be mediated via the second messenger cAMP. ZT-dependent input for Orco would be provided via daily changes in cAMP levels (Schendzielorz et al., 2015).

      In contrast, the study does provide strong evidence that the application of cyclic nucleotides can modulate Orco-dependent activity at a single time point, and reports that the temporal pattern of Orco transcript abundance is not circadian.

      We appreciate that the reviewer confirms that our revised manuscript with additional experiments now provides strong evidence that cAMP modulates Orco-dependent spontaneous activity of M. sexta ORNs. Since we already published that cAMP levels expresses daily rhythms in M. sexta antennae (Schendzielorz et al., 2015), and in vivo cAMP infusion increases spontaneous activity and sensitizes pheromone detection (Flecke et al., 2010), and our computational model here proves that circadian modulation of open time probability of Orco is sufficient to explain our experimental data, it is sufficient for our conclusions to test just the one specific zeitgeber time when endogenous cAMP levels are low and pharmacological cAMP increase has the strongest impact. To further reveal complete ZT-dependence of cAMP modulation of Orco´s control of spontaneous activity is beyond the scope of the current manuscript and not part of the current research question. The structure of Orco is extraordinarily conserved during evolution, thus, the cited experimental results from other laboratories and other species showing that Orco is a hub for posttranslational modification are very likely generalizable to different insect species. We clarified the manuscript accordingly. In Drosophila, Orco has at least 5 phosphorylation sites for protein kinase C (PKC), is cAMP-dependently sensitized, and has a Ca<sup>2+</sup>/calmodulin binding site that orchestrates the localization of the OR-Orco heteromer to the cilia. However, in fruit fly and other insects, so far, it can only be speculated how circadian control is provided for Orco, since there are no other publications that examine the circadian regulation of Orco in detail. We clarified our manuscript accordingly.

      To summarize, the logical conclusion based on our newly provided data is that the current hierarchical hypothesis in chronobiology based solely on a circadian TTFL clock that controls Orco transcription does not explain our findings in hawkmoth ORNs. Therefore, we suggest a new systemic hypothesis based upon coupled TTFL and PTFL circadian clocks that can also reconcile otherwise inconsistent data published for insect and mammalian circadian clocks (please see reviewed data in: Stengl and Schneider, 2024). We clarified our manuscript in the second revision and added a new Figure 10 to illustrate our novel hypothesis.

      However, the findings are incomplete to exclude a role for transcriptional-translational mechanisms and their associated multi-layered controls in circadian regulation.

      We certainly do not imply excluding a role for the TTFL clock in (indirectly) affecting circadian control of the membrane potential of ORNs. The new qPCR experiments added in the first revision clearly show that the circadian control mediated via Orco is not an output of the TTFL clock via transcriptional control of Orco. Instead, we predict links between a PTFL membrane clock comprising Orco as hub to integrate posttranslational control and the TTFL nuclear clock. We clarified our manuscript accordingly, adding a new Figure 10 to further illustrate and visualize our hypothesis. The predictions of this systemic hypothesis will be challenged in further experiments that, however, are beyond the scope of the current manuscript.

      Joint Public Review:

      This manuscript puts forward the provocative idea that a posttranslational feedback loop regulates daily and ultradian rhythms in neuronal excitability. The authors used in vivo long-term tip recordings of the long trichoid sensilla of male hawkmoths to analyze spontaneous spiking activity indicative of the ORNs' endogenous membrane potential oscillations. This firing pattern was disrupted by pharmacological blockade of the Orco receptor. They then use these recordings together with computational modeling to predict that Orco receptor neuron (ORN) activity is required for circadian, not ultradian, firing patterns. Orco did not show a circadian expression pattern in a qPCR experiment, and its conductance was proposed to be regulated by cyclic nucleotide levels. This evidence led the authors to conclude that a post-translational feedback loop (PTFL) clockwork, associated with the ORN plasma membrane, allows for temporal control of pheromone detection via the generation of multi-scale endogenous membrane potential oscillations. The findings will interest researchers in neurophysiology, circadian rhythms, and sensory biology. However, the manuscript has limited experimental evidence to support its central hypothesis and is undermined by several assumptions that underlie their data analysis and model builds, as well as insufficient biological data including critical controls to validate and/or fully justify the model the authors are proposing.

      We want to remind our reviewers that we used “ORN” as abbreviation for olfactory receptor neuron (= sensory receptor neurons, a.k.a. olfactory sensory neuron (OSN)) and not for Orco receptor neuron, although we focus on the function of Orco. Accordingly, our central finding is that Orco as ion channel is required for the daily/circadian modulation of spontaneous action potential activity generated by the olfactory receptor neurons in the absence of pheromone stimulation.

      We do not understand the specific basis for the conclusions of the reviewers. Therefore, we ask to please specify what experimental evidence is missing to support our central hypothesis that Orco is not directly TTFL- but PTFL clock-controlled, and to name specifically what the “several assumptions” are that undermine our careful data analysis and model builds. Which specific argument in our previous rebuttal was wrong, was not conclusive? Furthermore, please specify your claim that “critical controls are missing”. Which controls are missing for which experiments? We did add a new figure panel in the first revision to demonstrate that neither the addition of DMSO (the OLC15 solvent) nor the repeated attachment of the recording electrode, which was necessary to obtain paired datasets, altered the spontaneous spiking activity (Fig 1B in Revision 1). Furthermore, we expanded the time series of qPCR data and added tim as positive control to Orco (Fig 7), strengthening our argument that Orco expression is not under TTFL control. Dose-response curves of various Orco agonists and antagonists have been published before (see our references in Revision 1) and are therefore not repeated here.

      Our newly added data confirm what our modeling predicted: cAMP increases spontaneous ORN activity dependent on Orco. Previous publications provide evidence for daily rhythms in cAMP concentrations in hawkmoth antennae (Schendzielorz et al., 2012).

      As is true for any other hypothesis, a hypothesis can only be falsified but not validated and needs to be tested by many experiments from many laboratories over a long time until it will be replaced by the next hypothesis that better explains accumulating contradicting evidence. We are very much looking forward to experimental challenges of our provocative new hypothesis by colleagues in the field of olfaction and of chronobiology. We are convinced that our manuscript will greatly stimulate the field, possibly provoking a paradigm switch in chronobiology and in olfactory research.

      Strengths:

      The authors raise several intriguing model-based hypotheses regarding the mechanisms that underlie the generation of olfactory rhythms. The electrophysiological approach and the long-term recording paradigm are elegant and technically impressive. In the revised version, the authors have added additional qPCR data supporting the lack of rhythmic Orco transcript expression and included a new figure suggesting that cAMP can modulate Orco conductance.

      We thank the reviewers for their acknowledgement of our careful work and hope that our further revisions and clarifications help to argue our case.

      Major weaknesses:

      (1) The cAMP experiment was only conducted at one time-point, which is insufficient to support the central claim that "AMP and cGMP may have ZT-dependent effects on Orco conductivity".

      We agree with the reviewers and revised our discussion accordingly to clarify that in this manuscript it is not our central claim that cAMP and cGMP may have ZT-dependent effects on Orco conductivity. Instead, our data show for the first time that Orco controls circadian rhythms of spontaneous activity of ORNs and that the circadian rhythmicity of spontaneous activity is lost when Orco is blocked. Therefore, we provide novel experimental evidence that Orco is a prerequisite to the circadian rhythmicity of spontaneous activity and thus, to circadian rhythms in the membrane potential of ORNs. Furthermore, as requested by the reviewers we provided clear evidence in the first revision that Orco is not controlled at the transcriptional level by the TTFL clock, in contrast to the TTFL clock protein TIMELESS. Thus, it follows logically that Orco is under post-transcriptional control. Since cAMP levels show circadian oscillations and Orco is gated by cAMP (Fig 9 in Revision 1), we used our computational model to show that a cAMP-dependent increase in Orco conductance alone, via daily oscillating concentrations of cAMP, is sufficient to explain our findings. Therefore, we propose here that daily/circadian oscillations of cAMP modify spontaneous spiking activity via Orco on a posttranscriptional level. But it certainly does not provide all evidence for respective mechanisms of how this cAMP modulation of Orco is obtained, since this is beyond the scope of the current manuscript.

      Since we realized that it is difficult for our readers to visualize a circadian PTFL membrane clock we added a new hypothesis-Figure (Figure 10) and considerably focused and clarified especially the discussion of our manuscript. We pointed out that a membrane-associated signalosome that comprises delayed negative feedback mechanisms, and, thus, constitutes an oscillator, a “membrane clock” that generates oscillations. Based upon our data we propose a membrane-associated signalosome constituting a PTFL circadian clock with Orco as central element. This PTFL membrane clock generates superimposed ultradian and circadian rhythms in its outputs: rhythms in the membrane potential, Ca<sup>2+</sup>, and cAMP levels. The PTFL clock comprises positive feedforward elements that upregulate its outputs, resulting in more depolarization, higher Ca<sup>2+</sup>- and higher cAMP levels. Via the clock’s delayed negative feedback mechanisms these outputs are downregulated, again, resulting in hyperpolarization, decreasing Ca<sup>2+</sup>- and cAMP levels. This signalosome comprises the pacemaker channel Orco as a central hub that is controlled via changes in voltage, Ca<sup>2+</sup>, and cAMP levels. Nevertheless, we predict coupling between the multiscale PTFL membrane clock and the TTFL circadian clock in the nucleus to obtain stable circadian rhythms. As likely mechanism of coupling we predict that Ca<sup>2+</sup>- and cAMP-dependent kinases interlink both types of clocks, thereby obtaining robust and at the same time flexible interlinked cellular rhythms.

      We hope to now successfully clarify and to visualize our central hypothesis of our manuscript that Orco is not directly controlled by a TTFL circadian clock but is a central element of a membrane-associated posttranslational feedback loop clock (PTFL) clock that is linked to but not forced by the TTFL clock which is predicted to control intracellular Ca<sup>2+</sup> homeostasis in a circadian rhythm.

      (2) The revised manuscript continues to rely heavily on prior publications or defers key mechanistic questions (or important manipulations) to future studies. In its current form, the evidence presented remains insufficient to support the central claim that a PTFL constitutes the primary underlying circadian clock mechanism. The proposed model is intriguing, but the data provided do not yet directly demonstrate the novel mechanism.

      We do not understand why the reviewers considers it to be problematic that we “continue to rely heavily on prior publications”. Certainly, we built upon previous publications of our lab as well as on manifold experimental data published by other laboratories in the field of insect olfaction. Our ample citations demonstrate that we have an overview both of the current state of literature and relevant previous literature, dating back to the very first experiments that pioneered pheromone transduction in insects. Based upon our extensive knowledge and experimental data collected in insect olfaction and based on very careful, critical, rigorous analysis of our data and data published by others, we were able to come up with a novel interpretation of the current literature about insect olfaction that differs considerably from the current main views. We consider this to be our strength and judge it as good scientific practice and not a flaw of our work. However, since we do not focus on OR-Orco heteromers and their function in pheromone/odor transduction in the cilia in the current manuscript, we considerably shortened this part of the discussion, avoiding pointing out that highly sensitive moth pheromone transduction greatly differs from less sensitive general odor transduction in Drosophila. Furthermore, since here we focus on cAMP, but not on cGMP-dependent modulation of Orco, we also deleted/considerably shortened this part of our discussion.

      We certainly agree with the reviewers that, while the data provided in the current manuscript are a logical basis for developing our novel hypothesis, they are not a direct and sufficient demonstration of proof and we are not able yet to directly demonstrate and explain the novel mechanism predicted. We would like to point out that if we provided this final proof, it would not be any more a novel hypothesis, but only a novel finding.

      We agree with the reviewers that our provocative hypothesis requires rigorous testing by many further experiments, hopefully not only by our laboratory, but hopefully stimulating new experimental challenges by other laboratories employing different species. But certainly, these experiments with proof-of-principle will take many years and are beyond the scope of our current research paper.

      As per eLife’s assessment system we would like to ask the reviewers to provide detailed feedback as to which experiments/results within the scope of this manuscript would complete this work, or how they think this study should be framed in the light of the results that we obtained. Nevertheless, we hope that with our current careful review the reviewers will be more convinced by our arguments and experiments as valid basis for our provocative new hypothesis.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We appreciate that the reviewers provided an overall positive assessment of our manuscript and offered constructive suggestions for improvement. All three reviewers noted that a key strength of our study is the implementation of a gut microbiome model for the characterization of interbacterial antagonism pathways such as the type VI secretion system (T6SS) that approaches natural complexity. They note our work represents a significant advance in microbiome research, and generates resources that will be of use to many researchers in the field. Two of the reviewers point out that the complexity of our model limits the nature of measurements we can make, and suggest we temper the strength of the some of the conclusions we draw. As noted in more detail below, in our revised manuscript, we have used more precise wording to characterize our findings, and we are more explicit about the connection between the measurements we made and what we can conclude about the physiological role of the T6SS in the gut microbiome.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate the physiological role of the Type VI secretion system (T6SS) in a naturally evolved gut microbiome derived from wild mice (the WildR microbiome). Focusing on Bacteroides acidifaciens, the authors use newly developed genetic tools and strain-replacement strategies to test how T6SS-mediated antagonism influences colonization, persistence, and fitness within a complex gut community. They further show that the T6SS resides on an integrative and conjugative element (ICE), is distributed among select community members, and can be horizontally transferred, with context-dependent effects on colonization and persistence. The authors conclude that the T6SS stabilizes strain presence in the gut microbiome while imposing ecological and physiological constraints that shape its value across contexts.

      This study is likely to have a significant impact on the microbiome field by moving experimental tests of T6SS function out of simplified systems and into a naturally coevolved gut community. The WildR system, together with the strain replacement strategy, ICE-seq approach, and genetic toolkit, represents a powerful and reusable platform for future mechanistic studies of microbial antagonism and mobile genetic elements in vivo.

      The datasets, including isolate genomes, metagenomes, and ICE distribution maps, will be a valuable community resource, particularly for researchers interested in strainresolved dynamics, horizontal gene transfer, and ecological context dependence. Even where mechanistic resolution is incomplete, the work provides a strong experimental foundation upon which such questions can be directly addressed.

      Overall, this study occupies a space between system building and mechanistic dissection. The authors demonstrate that the T6SS influences persistence and community structure in vivo, but the physiological basis of these effects remains unresolved. Interpreting the results as evidence of fitness costs or selective advantage, therefore, requires caution, as multiple ecological and host-mediated processes could produce similar abundance trajectories.

      Placing the findings within the broader literature on microbial antagonism, particularly work emphasizing measurable costs, benefits, and tradeoffs, would help readers better contextualize what is directly demonstrated here versus what remains an open question. Viewed in this light, the principal contribution of the study is to show that such questions can now be addressed experimentally in a realistic gut ecosystem.

      We thank the reviewer for this thoughtful summary of our study. We were glad to read they conclude our work will have a significant impact on the microbiome field and that the resources we have developed will be of value to the community.

      Strengths:

      A major strength of this study is that it directly interrogates the physiological role of the T6SS in a naturally evolved gut microbiome, rather than relying on simplified pairwise or in vitro systems. By working within the WildR community, the authors advance beyond descriptive surveys of T6SS prevalence and address function in an ecologically relevant context.

      The authors provide clear genetic evidence that Bacteroides acidifaciens uses a T6SS to antagonize co-resident Bacteroidales, and that loss of T6SS function specifically compromises long-term persistence without affecting initial colonization. This temporal separation is well designed and supports the conclusion that the T6SS contributes to maintenance rather than establishment within the community.

      Another strength is the identification of the T6SS on an integrative and conjugative element (ICE) and the demonstration that this element is distributed among, and exchanged between, community members. The use of ICE-seq to track distribution and transfer provides strong support for horizontal mobility and adds mechanistic depth to the study.

      Finally, the transfer of the T6SS-ICE into Phocaeicola vulgatus and the observation of context-dependent colonization benefits followed by decline is a compelling result that moves the study beyond simple "T6SS is beneficial" narratives and highlights ecological contingency.

      We appreciate this detailed and nuanced characterization of the strengths of our study.

      Weaknesses:

      Despite these strengths, there is a mismatch between the precision of the claims and the precision of the measurements, particularly regarding fitness costs, physiological burden, and the mechanistic role of the T6SS.

      We acknowledge that in some places, our manuscript could benefit from greater precision in the language we use when linking the outcomes we observe in our study to their potential underlying causes. Specific revisions we made to address this concern are described below.

      First, while the authors conclude that the T6SS "stabilizes strain presence" and that its value is constrained by fitness costs, these costs are not directly measured. Persistence, abundance trajectories, and eventual loss are informative outcomes, but they do not uniquely identify fitness tradeoffs. Decline could arise from multiple nonexclusive mechanisms, including community restructuring, host-mediated effects, incompatibilities of the ICE in new hosts, or ecological retaliation, none of which are disentangled here.

      We agree that multiple mechanisms could explain why populations of certain species carrying a T6SS decline over time, and why for others, the T6SS contributes to long-term persistence. Our use of the term “fitness cost” to describe the phenomenon of decline observed for P. vulgatus carrying the T6SS was not meant to imply any particular underlying mechanism, but was rather our attempt to characterize the phenotypic outcome we observed in simplified terms. We note that ecological context is an important determinant of the fitness cost or benefit of any given trait, and our study sheds light on the importance of the presence of the WildR community and the mouse intestinal environment to the fitness contribution of the T6SS to B. acidifaciens and P. vulgatus. Nonetheless, to avoid implying an overly simplistic interpretation of our results, we have modified our language in the manuscript in several places when describing the role of the T6SS in species persistence in mice colonized with the WildR community.

      Second, the manuscript frames the T6SS as having a defined physiological role, yet the data do not resolve which physiological processes are under selection. The experiments demonstrate that T6SS activity affects persistence, but they do not distinguish whether this occurs via direct killing, resource release, niche modification, or higher-order community effects. As a result, "physiological role" remains underspecified and risks being conflated with ecological outcome.

      We acknowledge that our study does not fully resolve the physiological processes under selection that mediate role of the T6SS in maintaining B. acidifaciens populations in WildR-colonized mice. Indeed, several of the outcomes of T6SS activity the reviewer lists, such as target cell killing and nutrient release, are inextricably linked and thus inherently difficult to disentangle. We note that we did attempt to measure higher-order community effects of T6SS activity with metagenomic sequencing, but acknowledge that this approach may not have been sufficiently sensitive to detect small community shifts mediated by a relatively low-abundance species. To address the concern that our current framing implies more of a mechanistic understanding that our study achieves, we have substituted “ecological” for “physiological” where appropriate throughout the manuscript.

      Third, although the authors emphasize context dependence, the study offers limited quantitative insight into what aspects of context matter. Differences between native and recipient hosts, or between early and late colonization phases, are described but not mechanistically interrogated, making it difficult to generalize beyond the specific cases examined.

      We are not entirely clear what the reviewer means by “differences between native and recipient hosts”, but we agree that additional quantitative studies will be needed to address the generalizability of our findings. Future studies are also needed to address the mechanistic basis for the difference in the benefit conferred by the T6SS that we observed between P. vulgatus and B. acidifaciens.

      Fourth is the lack of engagement with recent experimental literature demonstrating functional roles of the T6SS beyond simple interference competition. While the authors focus on persistence and competitive outcomes, they do not adequately situate their findings within recent work demonstrating that T6SS-mediated antagonism can serve additional physiological functions, including resource acquisition and DNA uptake, thereby linking killing to measurable benefits and tradeoffs. The absence of this literature makes it difficult to place the authors' conclusions about physiological role and fitness cost within the current conceptual framework of the field. Without this context, the physiological interpretation of the results remains incomplete, and alternative functional explanations for the observed dynamics are underexplored.

      We thank the reviewer for specifically highlighting the potential pertinence of this literature to our study. Indeed, we did not cite studies indicating a link between T6SS activity and the uptake of DNA and other resources released by targeted cells. As we note above, the release of intracellular contents from target cells is an inevitable consequence of the delivery of lytic effectors. Thus, distinguishing between fitness benefits conferred from the elimination of competitor species and those arising from scavenging the nutrients released during this process is not straightforward. Measuring the benefits deriving from the uptake of certain released molecules, such as DNA, was not immediately feasible in the system employed in this study and instead we focused on the direct lytic consequences of the effectors delivered via the T6SS. We revised the Discussion to include reference to these possible downstream benefits of T6SS activity (Lines 476-479).

      A further limitation concerns the taxonomic scope of the functional analysis. The authors state that the role of the T6SS in the murine environment is functionally investigated using genetically tractable Bacteroides species, citing the lack of genetic tools for Mucispirillum schaedleri. While this is a reasonable, practical choice, it means that a substantial fraction of T6SS-encoding species in the WildR community are not experimentally interrogated. Consequently, conclusions about the role of the T6SS in the murine gut necessarily reflect the subset of taxa that are genetically accessible and may not fully capture community-level or niche-specific functions of T6SS activity. Given that M. schaedleri is represented as a metagenome-assembled genome, its isolation and genetic manipulation would be technically challenging. Nonetheless, explicitly acknowledging this limitation and slightly tempering claims of generality would strengthen the manuscript.

      The reviewer points out that studying the T6SS activity in M. schadleri would potentially expand the generality of our claims. We agree that having an isolate of this species along with genetic tools for its manipulation would allow us to probe the importance of the T6SS in the gut microbiome more broadly. At the suggestion of the reviewer, we have added explicit mention of the potential benefit of studying the T6SS in this organism to the Discussion (lines 538-539), an endeavor that lies outside of the scope of the current study.

      Finally, several interpretations would benefit from more cautious language. In particular, claims invoking fitness costs, selective advantage, or physiological burden should be explicitly framed as inferences from persistence dynamics, rather than as direct measurements, unless supported by additional quantitative fitness or growth assays.

      We agree with the reviewer that invoking fitness costs, selective advantages or physiological burdens should be done cautiously, and have made revisions to our manuscript where we acknowledge that more precise language was needed (line 43, 416, line 417). However, we would also argue invoking fitness costs and benefits when describe strain persistence dynamics in mice has substantial precedent in the literature (Feng et al. 2020, Brown et al. 2021, Park et al. 2022, Segura Munoz et al. 2022), to list a handful of representative examples published by different groups). It is unclear to us what additional in vivo growth measurements could be taken to substantiate our claim that the T6SS provides a fitness benefit to B. acidifaciens during prolonged gut colonization, or that carrying the ICE imposes a fitness cost on P. vulgatus during longterm colonization. Our in vitro experiments evaluating the competitiveness conferred by T6SS activity provide a measure of insight into its fitness benefits, but as our in vivo strain persistence data and the work of many others show, in vitro measurements do not necessarily capture in vivo parameters.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to determine how a contact-dependent bacterial antagonistic system contributes to the ability of specific bacterial strains to persist within a complex, native gut community derived from wild animals. Rather than focusing on simplified or artificial models, the authors aimed to examine this system in a biologically realistic setting that captures the ecological complexity of the gut environment. To achieve this, they combined controlled laboratory experiments with animal colonization studies and sequencing-based tracking approaches that allow individual strains and mobile genetic elements to be followed over time.

      Strengths:

      A major strength of the work is the integration of multiple complementary approaches to address the same biological question. The use of defined but complex communities, together with in vivo experiments, provides a strong ecological context for interpreting the results. The data consistently show that the antagonistic system is not required for initial establishment but plays a critical role in long-term strain persistence. This insight that moves beyond traditional invasion-based views of microbial competition. The observation that transferable genetic elements can confer only temporary advantages, and may impose longer-term costs depending on community context, adds important nuance to current understanding of microbial fitness.

      We thank the reviewer for the positive feedback and are glad they agree our study provides new insight into the role of interbacterial antagonism in natural communities.

      Weaknesses:

      Overall, there is not a lack of evidence, but a deliberate trade-off between ecological realism and mechanistic resolution, which leaves some causal pathways open to interpretation.

      The reviewer makes a good point that the complexity of the experimental system we employ precludes some lines of experimentation that would yield more mechanistic information. As the reviewer notes, we were aware of the tradeoff between mechanistic resolution and ecological realism when selecting our experimental system. Our deliberate choice to favor biological complexity over mechanistic clarity in this study stemmed from our perception that a major gap in understanding of the T6SS and other antagonism pathways lies in defining their ecological function in complex microbial communities.

      Reviewer #3 (Public review):

      Summary:

      Shen et al. investigate the contribution of the type VI secretion system of Bacteroidales in the gut microbiome assembly and targeting of closely related species. They demonstrate that B. acidifaciens relies on T6SS-mediated antagonism to prevent displacement by co-resident Bacteroidales and other members of the microbiome, allowing B. acidifaciens to persist in the gut.

      Strengths:

      Using a gnotobiotic model colonized with a wild-mouse microbiome is a significant strength of this study. This approach allows tracking of microbiome changes over time and directly examining targeting by Bacteroidales carrying T6SS in a more natural setting. The development of ICE-seq for mapping the distribution of the T6SS in the microbiome is remarkable, enabling the study of how this bacterial weapon is transferred between microbiome members without requiring long-read metagenomics methods.

      We thank the reviewer for their enthusiasm toward our study.

      Weaknesses:

      Some conclusions are based on only four mice per condition. The author should consider increasing the sample size.

      We agree that in some experiments it would be beneficial to increase the sample size from four mice. However, the experiments we performed for this study are time and resource-intensive. Additionally, the experiments on which we base our primary conclusions were all independently replicated with similar results. Given these factors, we determined that the extra confidence that might be afforded by increasing our sample size did not merit the delay in publication and investment in resources that would be required.

      Overall, the authors successfully achieved their objectives, and their experimental design and results support their findings. As mentioned in the discussion, it would be important to investigate the role of the T6SS in resilience to disturbances in the microbiome, such as antibiotics, diet, or pathogen invasion. This work represents a step forward in understanding how contact-dependent competition influences the gut microbiome in relevant ecological contexts.

      We agree that investigating the role of the T6SS during perturbations of the microbiome is a key next step for this work and thank the reviewer for highlighting this important future direction.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Beth A. Shen et al. present a comprehensive and carefully executed study investigating the ecological role of the type VI secretion system (T6SS) in maintaining bacterial strains within a native, complex gut microbiome derived from wild mice. By integrating genetic manipulation, metagenomic and sequencing-based tracking approaches, in vitro competition assays, and gnotobiotic colonization experiments, the authors provide compelling evidence that the T6SS functions primarily as a persistence factor rather than a determinant of initial colonization.

      The study is conceptually strong and addresses an important gap in our understanding of how interbacterial antagonistic systems operate in complex, native microbial communities. The manuscript is generally well organized, the data are clearly presented, and the main conclusions are supported by robust experimental evidence. In particular, the demonstration that T6SS-encoding ICEs confer context- and hostdependent fitness effects, including transient benefits and potential long-term costs, adds important nuance to prevailing models of microbial competition.

      That said, several aspects of the study would benefit from clarification and deeper mechanistic discussion. Addressing the points below would further strengthen the rigor and interpretability of the work. Overall, this is a strong and interesting manuscript, requiring some revisions.

      We greatly appreciate the positive summary of our work by the reviewer, which highlights the multi-faceted approach we took to address gaps in our understanding of interbacterial antagonism in the microbiome. It is our hope that the reviewer agrees that our revisions of the manuscript, based on their feedback, clarify our methods and strengthen the interpretability of our work.

      Major comments

      (1) The competition assays in Figure 2B suggest that T6SS-dependent fitness effects are most pronounced among members of the order Bacteroidales. However, these experiments primarily measure population-level competitive outcomes rather than direct T6SS-mediated targeting events. In addition, the limited number of non-Bacteroidales strains included in the assay makes it difficult to conclude that T6SS activity is strictly restricted to closely related taxa.

      The authors should either temper their conclusions regarding target specificity or clarify that these data reflect competitive outcomes rather than direct evidence of targeting. Expanding the discussion to acknowledge these limitations would improve interpretative accuracy.

      With regards to the measurement we employed for assessing T6SS-mediated targeting, we acknowledge that this is, to a degree, an indirect way of determining T6SS targeting. However, there is extensive precedent in the literature for the use of similar assays in assessing targeting by many contact-dependent antagonism systems including the T6SS in Bacteroidales (Russell et al. 2014, Chatzidaki-Livanis et al. 2016, Wexler et al. 2016) and many Proteobacteria (e.g. (Hood et al. 2010)), the T4SS in Xanthomonas citri (Souza et al. 2015), the CDI system in Escherichia coli (Aoki et al. 2005), and the Esx system in Streptococcus intermedius (Whitney et al. 2017). In these studies, targeting was demonstrated by specific depletion of the competitor strain in the presence of a strain encoding an active antagonism system. We acknowledge that the competitive index we report Figure 2B reflects the relative population levels of both species in the assay, and thus does not directly show target species depletion. We opted to use this metric to display the data in the manuscript as a way of efficiently encapsulating and comparing many strain combinations in a single figure, and because the competitive index differences we observed in these derive from differences in target species growth yields (see Author response image 1, indicating growth yields from a representative strain pairing).

      Author response image 1.

      The T6SS of B. acidifaciens targets a WildR-derived P. vulgatus strain. CFUs indicate populations of the indicated strains after co-culture of wild-type or T6SS-inactivated B. acidifaciens with P. vulgatus. Data represent means and standard errors (n=3, *P<0.01, t-test with log ><0.01, t- test with log transformed data)

      We additionally acknowledge that more extensive testing is needed to fully understand the target range of the Bacteroidales T6SS. In our study, we assessed targeting of every WildR species that was readily culturable, which to the best of our knowledge, represents the broadest panel of targets for the Bacteroidales T6SS to be tested to date. We limited our testing to these strains, as the goal of these experiments was to gain insight into which co-residents of the WildR could be targeted by B. acidifaciens. We agree that testing of a broader cross-section of potential targets has merits, but this would require targeted cultivation strategies to obtain these organisms, and lies outside the scope of the current study. We have revised the manuscript to clarify that the target range testing encompassed the diversity of isolates available (p. 10, lines 231-236).

      (2) The bae1 gene encoded in Bacteroides caecimuris F12 contains a frameshift mutation. It would be valuable for the authors to comment on whether such frameshift mutations are a common genomic feature among gut-associated Bacteroides species in murine models. In addition, comparative analysis of human gut metagenomic datasets could reveal whether homologous effector proteins are present in commensal Bacteroides populations, and whether these homologs exhibit similar disruptive mutations.

      More broadly, the manuscript would benefit from a discussion of whether expression of a fully functional bae1 effector might impose a fitness cost on Bacteroidales members, for example, through metabolic burden or altered resource allocation. This is particularly relevant in light of recent studies demonstrating that T6SS effectors can drive physiological trade-offs by modulating metabolic dynamics (PMID: 40592326). Integrating this perspective would strengthen the evolutionary interpretation of effector mutagenesis.

      We agree with the reviewer that the functional and evolutionary significance of the point mutation in bae1 merits further investigation. Following the reviewer's suggestion, we looked in our own datasets and available public datasets from mouse and human microbiomes for evidence of bae1 inactivation. Unfortunately, the gene is present at a low enough frequency that these analyses were inconclusive. In our own metagenomic data from WildR mice, we did not obtain sufficient sequencing depth to assess the frequency at which bae1 is inactivated across genomes. We found a single complete copy of bae1 identical to that of B. acidifaciens in one published mouse microbiome-derived MAG, and detected fragments of the gene in a number of publicly available isolate and MAG genomes, but these were too low of quality to assess whether or not the gene was intact.

      As to whether or not bae1 expression imposes a fitness cost in the producing organism, we think this is unlikely to be significant, given that the impacts of Bae1 will be neutralized by the accompanying immunity protein. We speculate that the point mutation in the B. caecimuris gene is more likely to have arisen through genetic drift than as a result of selection.

      (3) Quantification and tracking of ICE transfer in vivo. In Figure 4D, the authors assess the abundance of resident P. vulgatus populations in germ-free mice co-gavaged with wild-type strains and derivatives carrying either the intact ICE or ICE ΔtssC. Because both ICE variants are capable of horizontal transfer, it is essential to clearly describe how the authors distinguish between (i) the original wild-type strain, (ii) engineered donor strains, and (iii) recipient strains that have newly acquired the ICE or ICE ΔtssC.

      Clarification of the specific molecular or sequencing-based strategies used to discriminate these populations is necessary to ensure accurate interpretation of the colonization dynamics.

      In this experiment, the P. vulgatus strains we introduced which carried the ICE (either the wild-type version or ICE DtssC) also contained an erythromycin resistance cassette (ermG) inserted distal to the ICE insertion site. Populations of the ICE-containing strain were quantified by either qPCR targeting the ermG gene (Fig. 4D, Supplemental Fig. 4D) or by plating on erythromycin-containing media (Fig. 4F). Endogenous P. vulgatus populations were quantified by qPCR targeting the ermG insertion site, which is disrupted in the marked strain. These methodological details have been added to the figure legend for clarity. We acknowledge that transfer of the ICE between introduced and endogenous populations is possible, and would not be detected by these metrics. To assess whether this occurs, we performed ICE-seq analyses on samples collected from mice colonized by the WildR and P. vulgatus ICE at early (7 days) and late (56 days) time points. These analyses revealed that overall, ICE distribution in this experiment was similar to that observed in mice colonized with the WildR alone (Figure 4A and Author response image 2). They additionally provided corroborating evidence that the population of ICE-containing P. vulgatus declined over the course of the experiment. Importantly, the only ICE insertion site we detected in P. vulgatus in these samples was that found in the introduced P. vulgatus strain. Previous studies show that GA1-containing ICE can insert at numerous locations in Bacteroides sp. genomes, a finding supported by our mapping of the ICE insertion sites from in vitro transfer experiments (Supplemental Fig. 4C) (Garcia-Bayona et al. 2021). Thus, our ICE-seq detection of a sole P. vulgatus ICE insertion site indicates that transfer of the element between P. vulgatus populations is likely not occurring in our experiments.

      Author response image 2.

      ICE-seq analysis indicates that introduction of P. vulgatus ICE into WildR-colonized mice has little impact on ICE distribution among endogenous strains. Graphs show frequency of mapped ICE junctions deriving from the indicated species as determined by 5¢ or 3¢ ICE-Seq analysis of DNA extracted from fecal samples collected either 7 or 56 days post-gavage of the WildR and P. vulgatus ICE into germ-free mice.

      (4) The analysis of fitness trade-offs associated with ICE acquisition in P. vulgatus convincingly demonstrates that the benefits of ICE transfer are transient and contextdependent. However, the mechanistic basis of these trade-offs remains underexplored. While the study primarily attributes both benefits and costs to T6SS-mediated antagonism, the ICE likely encodes additional genes that could influence metabolism, regulation, or stress responses.

      We agree with the reviewer that there are many mechanistic questions remaining regarding the benefits and costs associated with ICE acquisition, and acknowledge that we have not investigated the fitness contributions of ICE-encoded genes other than the T6SS. Indeed, as we noted in our discussion of the results from introducing P. vulgatus carrying the ICE into WildR-carrying mice, our data suggest that ICE genes outside the T6SS may be beneficial (lines 463-465). At the reviewer’s suggestion, we have reiterated the importance of considering the fitness contribution of genes beyond the T6SS in determining ICE distribution in the WildR community (line 527).

      Minor comments:

      (1) In lines 319 and 333, the manuscript refers to "Supplemental Figure 3F" and "Supplemental Figure 3G," respectively. However, the provided Supplemental Figure 3 appears to end at panel E. Please clarify or correct these references.

      We have modified the text to reference the correct figure panels.

      (2) Line 1043: The notation for "OD600" should be corrected for consistency and accuracy.

      The notation for OD600 has been updated to be consistent throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      Minor comments:

      (1) Line 144. I would be careful of using "strong correlation, in this sentence. Although it shows a higher correlation than lab mice. Also, the labels in Figure 1A for mouse WildRF7 are confusing and not well explained in the figure legend.

      We modified line 147 (new line in edited manuscript) to say “positive correlation” rather than “strong correlation” to better represent the result. We also revised the legend for Figure 1A to better explain the samples of WildR F7 that were analyzed.

      (2) Line 155. It's unclear which strains were isolated from the WildR community, and the reason for isolating only 15 strains. Also, Supplemental Figure 1 shows 17 isolates, not 15.

      We apologize for the confusion here. We isolated 17 strains, which is the number of distinct strains we were able to readily culture from this community. We obtained genome sequences for 15 of these, and were able to assemble a genome for one more of the strains from metagenomic data.

      (3) Line 236. Is it known what makes B. uniformis resistant to B. acidifaciens carrying a T6SSS? Does it have an orphan immunity protein?

      We do not know why B. uniformis is not targeted by B. acidifaciens under the conditions of our experiments. It does not encode homologs of the immunity genes bai1 or bai2, and does not appear to be intrinsically resistant to targeting by this T6SS given that it is effectively targeted by P. vulgatus carrying the ICE (Fig.4b).

      (4) Line 295. There is a consistent decline in C. acid abundance after 27 days in Figures 3B and 3C. How do you explain this? Is the endogenous B. acid expanding to outcompete C. acid exo since the total C. acid exo remains constant when gavaging 100x B. acid exo?

      We believe that the eventual decline in the introduced population of B. acidifaciens is likely due to a fitness cost imposed by the erm resistance marker we employed. We noted this phenomenon when describing the results depicted in Fig. 3F-H, but neglected to include this explanation earlier. This oversight has been corrected (lines 315-317).

      References

      Aoki, S. K., R. Pamma, A. D. Hernday, J. E. Bickham, B. A. Braaten and D. A. Low (2005). "Contact-dependent inhibition of growth in Escherichia coli." Science 309(5738): 1245–1248.

      Brown, E. M., H. Arellano-Santoyo, E. R. Temple, Z. A. Costliow, M. Pichaud, A. B. Hall, K. Liu, M. A. Durney, X. Gu, D. R. Plichta, C. A. Clish, J. A. Porter, H. Vlamakis and R. J. Xavier (2021). "Gut microbiome ADP-ribosyltransferases are widespread phage-encoded fitness factors." Cell Host Microbe 29(9): 1351-1365 e1311.

      Chatzidaki-Livanis, M., N. Geva-Zatorsky and L. E. Comstock (2016). "Bacteroides fragilis type VI secretion systems use novel effector and immunity proteins to antagonize human gut Bacteroidales species." Proc Natl Acad Sci U S A 113(13): 3627– 3632.

      Feng, L., A. S. Raman, M. C. Hibberd, J. Cheng, N. W. Griffin, Y. Peng, S. A. Leyn, D. A. Rodionov, A. L. Osterman and J. I. Gordon (2020). "Identifying determinants of bacterial fitness in a model of human gut microbial succession." Proc Natl Acad Sci U S A 117(5): 2622-2633.

      Garcia-Bayona, L., M. J. Coyne and L. E. Comstock (2021). "Mobile Type VI secretion system loci of the gut Bacteroidales display extensive intra-ecosystem transfer, multispecies spread and geographical clustering." PLoS Genet 17(4): e1009541.

      Hood, R. D., P. Singh, F. Hsu, T. Guvener, M. A. Carl, R. R. Trinidad, J. M. Silverman, B. B. Ohlson, K. G. Hicks, R. L. Plemel, M. Li, S. Schwarz, W. Y. Wang, A. J. Merz, D. R. Goodlett and J. D. Mougous (2010). "A type VI secretion system of Pseudomonas aeruginosa targets a toxin to bacteria." Cell Host Microbe 7(1): 25–37.

      Park, S. Y., C. Rao, K. Z. Coyte, G. A. Kuziel, Y. Zhang, W. Huang, E. A. Franzosa, J. K. Weng, C. Huttenhower and S. Rakoff-Nahoum (2022). "Strain-level fitness in the gut microbiome is an emergent property of glycans and a single metabolite." Cell 185(3): 513-529 e521.

      Russell, A. B., A. G. Wexler, B. N. Harding, J. C. Whitney, A. J. Bohn, Y. A. Goo, B. Q. Tran, N. A. Barry, H. Zheng, S. B. Peterson, S. Chou, T. Gonen, D. R. Goodlett, A. L. Goodman and J. D. Mougous (2014). "A type VI secretion-related pathway in Bacteroidetes mediates interbacterial antagonism." Cell Host Microbe 16(2): 227–236.

      Segura Munoz, R. R., S. Mantz, I. Martinez, F. Li, R. J. Schmaltz, N. A. Pudlo, K. Urs, E. C. Martens, J. Walter and A. E. Ramer-Tait (2022). "Experimental evaluation of ecological principles to understand and modulate the outcome of bacterial strain competition in gut microbiomes." ISME J 16(6): 1594-1604.

      Souza, D. P., G. U. Oka, C. E. Alvarez-Martinez, A. W. Bisson-Filho, G. Dunger, L. Hobeika, N. S. Cavalcante, M. C. Alegria, L. R. Barbosa, R. K. Salinas, C. R. Guzzo and C. S. Farah (2015). "Bacterial killing via a type IV secretion system." Nat Commun 6: 6453.

      Wexler, A. G., Y. Bao, J. C. Whitney, L. M. Bobay, J. B. Xavier, W. B. Schofield, N. A. Barry, A. B. Russell, B. Q. Tran, Y. A. Goo, D. R. Goodlett, H. Ochman, J. D. Mougous and A. L. Goodman (2016). "Human symbionts inject and neutralize antibacterial toxins to persist in the gut." Proc Natl Acad Sci 113(13): 3639–3644.

      Whitney, J. C., S. B. Peterson, J. Kim, M. Pazos, A. J. Verster, M. C. Radey, H. D. Kulasekara, M. Q. Ching, N. P. Bullen, D. Bryant, Y. A. Goo, M. G. Surette, E. Borenstein, W. Vollmer and J. D. Mougous (2017). "A broadly distributed toxin family mediates contact-dependent antagonism between gram-positive bacteria." Elife 6(Jul 11): e26938.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The study used a WNK643 inhibitor as the only tool to manipulate WNK1-4 activity. This inhibitor seems selective; however, it has been reported that it exhibits different efficiency in inhibiting the individual WNK kinases among each other (e.g. PMID: 31017050, PMID: 36712947). Additionally, the authors do not analyze nor report the expression profiles or activity levels of WNK1, WNK2, WNK3, and WNK4 within the relevant brain regions (i.e. hippocampus, cortex, amygdala). Combined, these weaknesses raise concerns about the direct involvement of WNK kinases within the selected brain regions and behavior circuits. It would be beneficial if the authors provided gene profiling for WNK1, 2, 3, and -4 (e.g. using Allen brain atlas). To confirm the observations, the authors should either add results from using other WNK inhibitors or, preferentially, analyze knock-down or knock-out animals/tissue targeting the single kinases.

      Revisions 1: The authors added Fig. S1A during the revisions to show expression of Wnk1-4. While the expression data from humans is interesting, the experimental part of the study is performed in mice. It would be more informative for the authors to add expression profiles from mice or overview the expression pattern with suitable references in the introduction to address this point. The authors did not add data from knock down or knockout tissue targeting the single kinases.

      Thank you for the excellent suggestion. We have added mouse in situ hybridization data curated from Allen Brain Atlas and found mRNA encoding WNK1 and WNK2 highly expressed in the hippocampus compared to WNK3 and WNK4. We also have included WNK1 knockdown data from cell lines (Figure S7A-F).

      Whole body WNK1 knockout is embryonically lethal, and we do not have access to brain tissue specific WNK knockout animal models. In addition, knockout of WNKs from brain tissue samples from animals is not very efficient from our experience and therefore, we included data from cell lines. In other non-neuronal cell lines, WNK1 knockdown replicates the effect of WNK463 (Figure S7A-D). However, in SHSY5Y cells, WNK1 knockdown did not replicate the effects of WNK463 on pAKT levels (Figure S7EF). This suggests tissue-specific effects of WNKs, and it also supports our suggestion that cooperativity among WNK family members is required in neuronal cells. This further supports our conclusion that WNK463 is an ideal tool to test our hypothesis in this study as it targets all 4 WNKs (WNK1-4) and furthermore, WNK463 was reported in the literature to inhibit only the four WNKs out of more than 400 kinases tested, indicating more selectivity than many small molecules used to target other enzymes.

      (2) The authors do not report any data on whether the global inhibition of WNKs affects insulin levels as such. Since the authors demonstrate the synergistic effect of simultaneous insulin treatment and WNK1-4 inhibition, such data are missing.

      Revisions 1: The authors added Fig. S5A to address this point. It is appreciated that authors performed the needed experiment. Unfortunately, no significant change was found, therefore, the authors still cannot conclude that they demonstrate a synergistic effect of simultaneous insulin treatment and WNT1-4 inhibition. It is a missed opportunity that the authors did not measure insulin in the CSF or tissue lysate to support the data.

      Thank you for the comment. As suggested, we tried to measure insulin in mouse hippocampal tissue lysate, and the levels fell way below the detectable range (78 - 5000 pg/mL) of the mouse insulin detection kit (Abcam: AB285341) used.

      (3) The study discovered that the Sortilin receptor binds to OSR1, leading the authors to speculate that Sortilin may be involved in the insulin-dependent GLUT4 surface trafficking. The authors conclude in the result section that "WNK/OSR1/SPAK influences insulin-sensitive GLUT4 trafficking by balancing GLUT4 sequestration in the TGN via regulation of Sortilin with GLUT4 release from these vesicles upon insulin stimulation via regulation of AS160." However, the authors do not provide any evidence supporting Sortilin's involvement in such regulation, thus, this conclusion should be removed from the section. Accordingly, the first paragraph of the discussion should be also rephrased or removed.

      Revisions 1: The authors added Fig. 5M-N to address this point. The new experiment is appreciated. However, the authors still do not show that sortilin is involved in insulin or WNK-dependent GLUT4 trafficking in their set up since the authors do not demonstrate any changes in GLUT4 sorting or binding. The conclusions should therefore be rephrased or included purely in the discussion. Moreover, the discussion was not adjusted either, leading to over interpretation based on the available data.

      Thank you for the suggestion. The conclusion has been rephrased as suggested.

      (4) The background relevant to Figure 5, as well as the results and conclusions presented in Figure 5 are quite challenging to follow due to the lack of a clear introduction to the signaling pathways. Consequently, understanding the conclusions drawn from the data is also difficult. It would be beneficial if the authors addressed this issue with either reformulations or additional sections in the introduction. Furthermore, the pulldown experiments in this figure lack some of the necessary controls.

      Revisions 1: The Authors insufficiently addressed this point during the revisions and did not rewrite the introduction as suggested.

      The background information related to figure 5 has been simplified as suggested. Response regarding the controls used is provided in the response to critique 5 as below.

      (5) The authors lack proper independent loading controls (e.g. GAPDH levels) in their immunoblots throughout the paper, and thus their quantifications lack this important normalization step. The authors also did not add knock-out or knock-down controls in their co-IPs. This is disappointing since these improvements were central and suggested during the revision process.

      GAPDH has been used as a loading control wherever applicable for Western blots on lysates (see Figures: 5D, 3F, 4E, 4C) In other cases, such as in Figure 5E, the analysis of pAS160 is normalized to total AS160 as this is more appropriate compared to GAPDH. For IP experiments such as Figure 5G, GAPDH is not an applicable control as we are using purified protein fragments in this case. For IP experiments (Figure 5K, 5L, 5C), IP proteins have been normalized to the input protein levels serving as a loading control for the IP because GAPDH is an intracellular protein which necessarily is not pulled down along with the proteins being IP’ed. Therefore, in this case, GAPDH is not a valid loading control. The choice of our loading controls used are very well supported by previous publications from our lab and other labs working on WNK pathways.

      (6) The schemes that represent only hypotheses (Fig. 1K, 4A) are unnecessary and confusing and thus should be omitted or placed at the end of each figure if the conclusions align.

      Thank you for the suggestion. Figure 1K is already at the end of the figure 1 and it shows the conclusion of that figure. Figure 4A have been placed at the end of the figures as suggested. Other schemes are only added at the end of the figures as suggested.

      (7) Low-quality images, such as Fig. 5H should be replaced with high-resolution photos, moved to the supplementary, or omitted.

      Thank you for your comment. The suggested images have been replaced with higher resolution ones.

      Reviewer #2 (Public review):

      This study by Jaykumar and colleagues seeks to expand the field's appreciation of insulin responses in the brain, specifically by implicating WNK kinase function in various neuronal responses, ranging from behavioral / memory changes to GLUT4 trafficking to the cell surface with subsequent glucose uptake. This revised study is now comprehensive and presents a logical and reasonably documented cascade of molecular interactions responsible in part for GLUT4 trafficking under the regulation of WKK and insulin. Additional data allow the authors to dissect a plausible WNK/OSR1/SPAK-sortilin pathway for the modulation of GLUT4 trafficking, in part by capitalizing on a overlay of various techniques and systems. The data - much of it in vivo or ex vivo - showing a potential role for WNK function in brain glucose utilization remains a compelling part of the story, with the dissection of the signaling cascade and a potential role for sortilin in mediating WNK function via effects on GLUT4 cellular localization now more convincing.

      Initially, the group shows that oral WNK463 treatment - an inhibitor of WNKs broadly - in mice augments a number of memory readouts. These findings fit within the context of the overall story the authors present: that WNK function is critical to brain glucose utilization, which impacts learning. Multiple approaches are used to show that WNK463 treatment, i.e. inhibition of WNKs, increases glucose uptake, including labeled 2deoxyglucose uptake in vivo in the brain and in isolated synaptosome, and uptake in ex vivo hippocampal slices. These findings are solid and consistent. With the exception of some relatively minor comments regarding the data presentation made to the authors and now fully addressed, the findings showing that WNK463 treatment increases GLUT4-mediated glucose uptake and surface localization of GLUT4 are reasonable, with the hippocampal slice data being particularly relevant.

      While the details of the WNK signaling cascade is dense, in the revised application one clearly appreciates the molecular interrogation and interactions the group is dissecting, supported by the use of multiple models. With the additional findings, these systems and the data now reinforce each other, presenting a strongly documented overall story.

      A limitation of the study with the initial submission was the authors' reliance upon a single pharmacological tool (WNK463) to inhibit WNK kinases. WNK463 apparently has substantial specificity for WNKs and WNK463 treatment lessened OSR1 phosphorylation (a WNK substrate). Nevertheless, the cohesiveness of the findings in terms of the broader pathway engagement (GLUT4 trafficking, glucose uptake) is consistent with the author's proposed mechanisms and conclusions. The authors have additionally addressed this concern in the revised manuscript with more information supporting the specificity of WNK463 as well as the multiple approaches to confirm the effect of WNK463 on the WNK signaling pathway of interest.

      The final few paragraphs of the discussion that weave the author's findings into the field more broadly, including Sortilin function and neurological disorders, are appreciated. Additional clarity in the Methods section is also helpful.

      Thank you for the positive response and acknowledging that we have satisfactorily addressed all of your critiques.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      (1) The principal weakness of the manuscript lies in the interpretation of biological robustness. The authors identify network topologies that sustain oscillatory behaviour despite perturbations to the system or parameters. However, in many cases, this persistence is due to the presence of partially redundant oscillatory motifs within the network. While this observation is interesting and of clear value for circuit design, framing it as evidence of evolutionary robustness may be misleading. The “mutant” systems frequently exhibit altered oscillatory properties, such as changes in frequency or amplitude. From a functional cellular perspective, mere oscillation is insufficient — preservation of specific oscillation characteristics is often essential. This is particularly true in systems like circadian clocks, where misalignment with environmental cycles can have deleterious effects. Robustness, from an evolutionary standpoint, should therefore be framed as the capacity to maintain the functional phenotype, not merely the qualitative behaviour. 

      We agree with the reviewers that our framing conflated qualitative robustness (continued oscillation) with functional robustness (oscillation with preservation of properties such as frequency), and that the latter is a more meaningful definition of robustness in an evolutionary context. We have edited the manuscript to remove any suggestions that our results can explain the multiple-oscillator architecture of circadian clocks. We have also added a paragraph in the discussion that highlights the distinction between qualitative and functional robustness, as well as the other ways in which our training environment differs from an evolutionary context.

      Locations of changes:

      Abstract, second-to-last sentence

      Results, final section, first paragraph

      Discussion, first paragraph

      Discussion, new paragraph

      (2) A secondary limitation is that, despite the methodological advances, the scale of the systems explored remains modest. While moving from 3- to 5-node systems is non-trivial, five elements still represent a relatively small network. It is somewhat surprising that the algorithm does not scale further, particularly when considering the performance of MCTS in other domains — for instance, modern chess engines routinely explore far larger decision trees. A discussion on current performance bottlenecks and potential avenues for improving scalability would be valuable.

      We thank the reviewers for raising this important point. We have edited the manuscript to specify that we faced two distinct bottlenecks in scaling our experiments. The first is the runtime and scaling of the underlying Gillespie simulations, which become much more expensive as circuit size increases. The second is our use of the original (“vanilla”) MCTS algorithm, without the deep-learning value and policy networks that have driven the dramatic gains in domains such as Go and chess. We have added a new Discussion paragraph that identifies these two bottlenecks, followed by a paragraph on future methodological enhancements that goes into more detail about potential improvements to the algorithm. Using a power-law extrapolation of the current scaling, we estimate that without further methodological improvements the largest tractable network is approximately 7 nodes, while AlphaZero-style scaling could plausibly extend the approach to roughly 19-node circuits. We have also added an order-of-magnitude estimate of the 5-node search space (≈2×10<sup>9</sup> topologies) to give the reader a more concrete picture of the current scale.

      Locations of changes:

      Results, final section, end of paragraph 1 (Gillespie runtime as the practical bottleneck)

      Discussion, new paragraph 4 (bottlenecks and possible improvements)

      Discussion, new paragraph 5 (projected scaling under deep-learning-based extensions)

      Introduction, paragraph 4, and Discussion, paragraph 1 (search-space size)

      Methods, new section “Estimation of search space size”

      (3) It is worth noting that the emergence of oscillations in a model often depends not only on the topology but also critically on parameter choices and the nature of the nonlinearities. The use of Hill functions and high Hill coefficients is a common strategy to induce oscillatory dynamics. Thus, the reported results should be interpreted within the context of the modelling assumptions and parameter regimes employed in the simulations.

      We agree that the modeling assumptions substantially impact the interpretation of the results, and we have expanded the description of our modeling framework to make these assumptions explicit. To clarify, our model does not use Hill equations directly. Instead, cooperative binding is represented as a sequential, mass-action binding process. In addition to a new Methods section explaining our model in more detail, we have added a Methods section to show analytically that the effective Hill coefficient in our system is always ≤2. It also mentions an important practical benefit of using sequential binding rather than explicit TF dimerization, which is that it improves the size and scaling behavior of the reaction system. 

      Locations of changes:

      Results, section 2, paragraph 1 (clarification that no Hill function is imposed; effective Hill coefficient ≤2)

      Methods, section 1, new subsection “Sequential binding yields Hill coefficients ≤2”

      Recommendations for the authors:

      (4) It would be helpful to include the explicit reaction equations and corresponding reaction rates used in the simulations, to facilitate reproducibility and better understanding of the modelling assumptions.

      We have added the explicit reaction equations, the corresponding rate parameters, and the bounds used during random parameter sampling. The Results section describing our model now reports the rate parameters used in the simulations shown in the paper. Additionally the new subsection at the beginning of the Methods presents the full stochastic model. Finally, the bounds used for parameter sampling are now included in-line in the corresponding Methods subsection. The original tables of parameter values and sampling bounds (Tables 1 and 2) have been retained.

      Locations of changes:

      Results, section 2, paragraph 1 (rate parameters in main text; Table 1 retained)

      Methods, section 1, new subsection “Stochastic model of a transcription factor network”

      Methods, section “Random sampling” (explicit parameter bounds; Table 2 retained)

      (5) Sustained oscillations are notoriously difficult to observe in Gillespie simulations due to stochastic noise. Could the authors comment on whether simulation times were sufficiently long to distinguish sustained oscillations from transient dynamics?

      We agree this is an important methodological point. We have clarified that simulations were run for 11.1 hours. Because nearly all the oscillators we found had periods below 100 minutes and most oscillators had periods below 12 minutes, almost all were observed for at least 6.6 cycles, which we believe is sufficient to distinguish sustained oscillations from transient dynamics.

      Locations of changes:

      Results, section 2, paragraph 2

      (6) While the manuscript alludes to broader applications of the proposed method, it would be beneficial to elaborate on these possibilities. Clarifying how this tool could be extended to other types of dynamical behaviours or biological questions would strengthen the impact.

      We have rewritten the final paragraph of the Discussion to give concrete illustrative examples of how the framework can be extended beyond oscillator design, including the discovery of more complex design principles and the design of synthetic multicellular circuits such as multi-cell type cancer therapies and morphogenetic patterning circuits.

      Locations of changes:

      Discussion, final paragraph

      (7) It remains somewhat unclear whether the contribution lies primarily in the novel application of reinforcement learning or whether there are methodological innovations within the algorithm itself. Are there existing tools with similar objectives, and how does this work improve upon them in terms of performance or capabilities?

      We have edited the Discussion to state more directly that the principal contribution of this work is the novel application of reinforcement learning to the problem of network topology design problem, rather than a fundamentally new RL algorithm. We also contextualize CircuiTree relative to other topology-search approaches and articulate where the use of RL provides a concrete advantage in navigating large combinatorial search spaces.

      Locations of changes:

      Discussion, end of paragraph 3

      (8) Further details on the training of the reinforcement learning algorithm would be appreciated. Was it trained a priori, and if so, how much data was required?

      CircuiTree is not pre-trained. Like AlphaZero, it learns exclusively from simulations performed during the search itself, with no externally curated dataset or a priori domain knowledge. We have clarified this in two places in our manuscript.

      Locations of changes:

      Introduction, beginning of first paragraph

      Results, section 1, paragraph 4

      (9) Some discussion on how computational time scales with increasing network size would be valuable.

      As mentioned above, we have added (i) an order-of-magnitude estimate of the size of the 5-node search space (≈2×10⁹ topologies) to make the current scale concrete; (ii) a power-law extrapolation of the current algorithm’s scaling that suggests ≈7 nodes is the largest tractable network without further improvements; and (iii) a Discussion paragraph projecting that AlphaZero-style scaling, combined with faster simulators, could plausibly extend the approach to ≈19-node circuits. The methodology for estimating the search-space size is described in a new Methods section.

      Locations of changes:

      Introduction, paragraph 4 (search-space size)

      Discussion, paragraph 1 (search-space size)

      Discussion, new paragraph 4 (extrapolated scaling and 7-node bound)

      Discussion, new paragraph 5 (projected scaling under deep-learning extensions)

      Methods, new section “Estimation of search space size”

      (10) The discussion of knockouts that dampen or restore oscillations is interesting. Was this analysis performed systematically, or were the examples selected through observation? Clarifying the methodology here would add rigour to the interpretation.

      We have clarified that the topologies highlighted in this analysis were selected by observing the fault-tolerant properties of a few topologies that exhibited high mutational robustness in our screen, rather than via a systematic analysis of the screen results. The text now states this explicitly so the reader can interpret the examples accordingly.

      Locations of changes:

      Results, final section, paragraph 3

      Other changes made by the authors

      In addition to the changes above, we have made several minor edits to improve clarity and presentation. We corrected spelling and formatting errors throughout, and improved the clarity of language in the Discussion section. No changes were made to the Figures, Tables, Algorithms, or Supplementary Information.

      We are grateful to the reviewers for their input and believe the manuscript has been substantially improved as a result. We hope the revised version meets with the editors’ and reviewers’ approval.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      The ciliary photoreceptor cells and its downstream neurons of larval annelid must be orchestrated in a specific pattern to promote downward swimming in response to long duration of UV exposure. The authors first conducted neuroanatomical examination of the circuit to identify NOS expression neurons (INNOS) that are immediately downstream to the ciliary photoreceptor cells. The INNOS is activated by UV and produces NO. The NOS is required for UV avoidance by Platynereis larvae and neural dynamics of the photoreceptor cells and their downstream circuit. Following up the RNA-seq data with in situ hybridization experiments, the authors found that two unconventional guanylate cyclases, NIT-GC1 and NIT-GC2, are expressed and localized in different subcellular domain of the photoreceptor cells. Experiments using the culture cells and genetically encoded sensors demonstrated that NIT-GC1 can generate cGMP in response to nitric oxide. Finally, authors build a mathematical model that fit the live imaging data and used it to predict how the magnitude of the photoreceptor activation varied by intensity and duration of UV light.

      Strengths:

      The authors conducted comprehensive interrogations of the UV avoidance pathway at the molecular and circuit levels, and constructed a mathematical model. The main conclusions are supported by layers of evidence from different assays.

      Weaknesses:

      Statistics are missing in both figure legends and methods. The perturbations of genes and molecules were not cell-type-specific and therefore the observed behavioral defect could be attributed to the malfunction of the circuit elsewhere not examined in this study. I suggest adding more explanation about the functions of other NOS-expressing cells and conducting a control experiment to test behavioral response to a non-visual stimulus.

      Thank you for this assessment of our work. We have now added additional panels with statistical tests to the figures and included explanatory text in the figure legends.

      Regarding the cell-type-specific effects, we would like to offer a more nuanced view of this. Some of the genes we studied (NIT-GC1 and NIT-GC2) are only expressed in 4 cells in an organism of ~10,000 cells (as we demonstrated by HCR, immunostainings and the analysis of single-cell RNAseq data) and we knocked-down these genes (validated by antibody staining) with two independent morpholinos (to be able to rule out off-target effects). This is as cell-type specific as it gets. NOS is also expressed in a very limited number of cells. In the larval stages, we detected NOS expression only in the four INNOS cells and the pigmented eyes. We could previously show that UV avoidance is only mediated by the cPRC circuit (including INNOS) whereas phototaxis is exclusively mediated by the pigmented eyes Verasztó et al. (2018). Due to this behavioural specificity, we are therefore confident that the effects of the NOS mutations on UV avoidance are due to the lack of NOS from the INNOS cells and not the pigmented eyes. Besides UV avoidance, we have characterised the speed of ciliary swimming without a light stimulus, as well as phototaxis []. In addition, we also tested phototaxis and detected a reduced phototactic reaction in NOS mutants, but not after chemically inhibition of NO production (Figure 3B and 3F). The additional effects of NO are thus well documented in the paper and independent of the function of NO in the UV reaction. In addition, we have now did further quantifications and added new data (Figure 3 – figure supplement 4) to show that NOS mutants show normal lunar periodicity of sexual maturation, similar to wild-type animals. NOS mutants are viable and fertile, are feeding, building tubes and are mating as wild-type animals (these behaviours were not quantified here, we only show the data for lunar periodicity).

      Reviewer #2 (Public Review):

      Summary:

      This study is quite thorough, tackling this NO-dependent UV avoidance circuit with both breadth and depth. There are several novel discoveries throughout, but the whole package represents perhaps even more than the sum of these parts.

      Strengths:

      The presentation of the work is compelling. The introduction sets up the question and the state of the field very nicely. The discovery of the non-canonical NO receptor pathway in the ciliary photoreceptors is fascinating and will likely open up new avenues for future research into NO pathways in different species. The use of genetic and pharmacological manipulations of circuit components was well thought-out. The authors applied different experimental techniques expertly throughout the study so that they could develop a comprehensive view from the molecular to the behavioral levels.

      Weaknesses:

      While I appreciate the intent of bringing together a large set of measurements from connectomics and calcium imaging in the framework of a model, the model seemed rather poorly constrained. How many parameters are in the model shown in Figure 6A? How many of them are well constrained by experimental measurements? The authors also don’t perform sensitivity analysis on the parameters of the model. And ultimately, the conclusion over the model in Figure 7 is somewhat trivial within the unitless construction: larger amplitude and longer duration stimuli lead to increased activation of the downstream neuron thought to lead to the downward swim behavior. I could imagine that a large family of models would arrive at this same result, and without units, there is no way to really test it with new behavioral experiments.

      We thank the reviewer for these comments. We have now thoroughly revised the model based on new experiments and carried out a sensitivity analysis. We are also more positive about the usefulness of the model though, for the following reasons.

      General usefulness of the model: With the model we can now reproduce all the qualitative dynamics of the circuit. The modelling also completely changed the way we were thinking about the system. For example, we needed to include a time-limited step in the cPRC transduction cascade leading to NOS activation to capture the time-invariance of the peak. In the future, we can also use this model to generate prediction e.g. about what the response to repeated stimulations can be. In the revised version, we included a new series of measurements of calcium dynamics in the cPRCs under varying duration and intensity of UV light. These important new results gave us further insights and necessitated a revision of the model.

      Unitless model: The model is indeed unitless, since already our input data from calcium imaging represent normalised data and the model was fit to these data. We would need a lot more information to build up a proper ground-up biophysical model (e.g. including capacitance, ionic concentrations etc.). This does not mean that the model is not useful.

      Sensitivity analysis: We have created a pipeline to carry out local and global sensitivity analysis and carried out a variance-based sensitivity and identifiability analysis. These data and code are included in the revised version. In the model, most parameters are identifiable. If we fix only two of the parameters all other parameters can be identified.

      Constrains and lots of possible models: In terms of the constrains, since we don’t have dimensioned quantitites, our constraints are bounds on the parameters. However, the structural identifibiality analysis tells us where we can identify parameters and can guard against sloppiness and overfitting. We only fitted the model on a subset of the data and can reproduce dynamics on a larger set of data. For important cellular interactions, the signs of the interactions are constrained. The model thus also allows us to rule out a large number of models - e.g. we could rule out a very simple model of progression from step to step as it was not possible to fit the data to such a model.

      We have updated the text to reflect these changes, e.g.: “Many of the couplings in our model are constrained (e.g. UV leads to INNOS activation) and e.g. reversing the sign of some of these couplings would not arrive at the same result. Several earlier variants of the model could not be fit to the data, The model is thus well constrained by our physiological experiments and the 106 circuit map.”

      Reviewer #3 (Public Review):

      The transition from planktonic to benthic depends upon several physical and chemical cues. Nitric oxide (NO) is known as a critical player in the induction of larval metamorphosis in several invertebrates. Although NO is a widespread signalling molecule in a broad range of organisms regulating key physiological processes, internal regulatory mechanisms studies are scarce. While the UV sensing in larvae of the annelid Platynereis dumerilii using ciliary photoreceptors has been studied, the neuronal signalling mechanism remains unknown. In this study, Kei Jokura et al. investigated how annelid Platynereis dumerilii larvae detect UV sensing and modulate swimming behaviour through nitric oxide feedback. Using existing resources of Platynereis larval connectome/volume EM data, they identified NOS-expressing interneurons within the ciliary photoreceptors circuit (cPRCs). They demonstrated that NO is produced in cPRCs during UV/violet stimulation by using a fluorescent NO-reporter line. Further, they demonstrated that Nitric oxide signalling mediates UV-avoidance behaviour by using NOS-mutant larvae. Finally, they mapped out the signalled mechanisms of the cPRC circuit using published spatially mapped single-cell transcriptome data of Platynereis larvae, the Ca sensor lines, in situ HCR, and immunostaining. Additionally, by using their findings from Ca imagining data of cPRC, INNOS and INRGWa cells collected in wild-type, NOS knockout and NIT-GC2 morphant larvae, Kei Jokura et al. developed a mixed cellular-circuit-level mathematical model. However, my expertise in mathematical modelling is limited, so I cannot comment on this section.

      No doubt, the study has been conducted extensively. However, I have a few comments, please see below.

      Page 4: “In contrast, both two- and three-day-old homozygous NOS-mutant larvae showed a strongly diminished UV avoidance response (Figure 3A, B and Figure 3-figure supplement 1B, C).” Instead of using subjective terms like “strongly,” it would be more relevant to provide statistical values. However, I could not locate any means of statistical analysis on larval behaviour. Can the authors indicate the statistical values for all behaviour studies?

      We thank the reviewer for these comments. We have changed the wording and also added the results of statistical analyses to the figures and explanations to the figure legends.

      Page 5: “(D) Vertical displacement in 30 sec bins of wild type and mutant (NOSΔ11/Δ11 and NOSΔ23/Δ23) three-day-old larvae stimulated with 395 nm light from the side, 488 nm light from the top and 395 nm light from the top.” The error bars for WT are too long at the end of the experiment. It is not clear how the authors decided to use this time frame. Did the authors try carrying this out for an extended time period? How did the authors decide on 120 seconds as the time frame for exposure? Authors should provide data on larval behaviour for an extended time.

      The 120 seconds time frame of exposure takes into account the reaction time and swimming speed of the larvae as well as the size of the assay chamber (160 mm water height).

      By the end of a 120 stimulation many larvae tend to accumulate at the bottom of the chamber due to downward swimming and cannot be further tracked. This effect leads to higher variability in the data towards the end of the experiment in the wild-type batches. We have showed both continuous vertical data as well as the binned data. The 30 sec bin was chosen for convenience and for better comparison with our previous paper on UV avoidance behaviour (Verasztó et al. 2018)

      Page 13: “During the UV response, prototroch cilia beat slower than trunk cilia, resulting in a head down stable state (‘rear-wheel drive’). In contrast, during the pressure response prototroch cilia beat faster than trunk cilia, leading to a head-up orientation (‘front-wheel drive’). Testing this hypothesis will require biophysical experiments and mathematical modelling.” Authors should carry out ciliary beating analysis under UV light in the current study with NOS mutant larvae. Since the pressure and UV detection systems are closely related, comparing the difference in ciliary beating is important to 155 demonstrate this hypothesis. Further, did the authors check the Ca sensor GCaMP6s under pressure conditions?

      We thank the reviewer for this suggestion. We have carried out further experiments and analysed the ciliary beat frequency (CBF) of larvae exposed to UV stimulation. We added these data to Figure 3—figure supplement 3. The results (increase of CBF under UV in wild-type but not NOS mutant larvae) were quite surprising to us and falsified our initial hypothesis. We have rewritten the discussion to reflect this important new finding.

      The response of cPRC cilia to changes in hydrostatic pressure has been extensively documented in our recent paper on the mechanism of barotaxis in the Platynereis larva (see Bezares-Calderón et al., https://doi.org/10.1101/2023.02.28.530398).

      Page 18: “strips. One strip contained UV (395 nm) LEDs (SMB1W-395, Roithner Lasertechnik) and the other infrared (810 nm) LEDs (SMB1W-810NR-I, Roithner Lasertechnik).” Authors should test larval swimming behaviour at different wavelengths. Even though they are performed in previous work, the experiment with different wavelengths is necessary to be conducted in NOS mutant larvae in parallel with a control. This will confirm that NOS is principally associated with UV. Further, to demonstrate that this mechanism is associated with ciliary movement, authors need to provide this evidence.

      The diving reaction by non-directional light can only be induced by UV/cyan light, as we have shown previously, and it is mediated by a single UV-opsin photopigment (c-opsin1). The avoidance experiments can thus only be done with UV/cyan light. We also measured swimming 174 behaviour with 480 nm directional light to test phototaxis (Figure 3D and Figure 3—figure 175 supplement 1F).

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      (1) The current introduction focuses on the nitric oxide signaling. It would be helpful for readers to have an introductory section about Platynereis larvae (a total number of neurons, etc) and rationales to use its nervous system as a model to study mechanisms of NO signaling and gating of visual response.

      We added an extra paragraph to give more detail on the number of cells in the larva and why Platynereis larvae are interesting to study to understand how synaptic and volume signalling interact. “The 3-day-old larva has over 9,000 cells classified into 202 neuronal and 92 non neuronal cell types (Verasztó et al., 2025). The synapticly connected subset of the cells in the body form a connectome of over 2,000 cells. Besides synapses, neurons in the larva also signal by volume transmission mediated by a rich repertoire of neuropeptides (Williams et al., 2017) and other modulators (Bauknecht and Jékely, 2017). The transparent and experimentally accessible Platynereis larvae could therefore inform how synaptic and volume signalling interact to mediate behaviour (Jékely and Yuste, 2024).”

      (2) The readers would want to know why these larvae swim downward in response to UV and upward to 480nm light. Is that for maintaining the certain depth from the surface of the water?

      It has been suggested that the ciliary photoreceptor circuit, which senses UV light, and the rhabdomeric photoreceptor circuit, which senses blue light, exchange messages with each other and that the two work together to form a depth gauge. By allowing larvae to swim at their preferred depth, the depth gauge influences where they end up when they become adults.

      (3) Are there splicing isoforms of NOS in the Platynereis dumerilii? If so, do antibodies and probes for in situ distinguish them? In Drosophila, truncated isoforms can inhibit the function of full length isoform, and therefore it was important to use methods to distinguish spicing isoforms.

      We did not identify any alternatively spliced forms of Platynereis NOS in our published transcriptome resources.

      (4) “INRGWs” acronym appears without explanation in the first paragraph of the results.

      Corrected: “the INRGWa neurons (cholinergic interneurons that express an RGW neuropeptide)”

      (5) In Figure 1C, it is difficult to see the projection patterns of individual cell types. Figure supplement 3 can be combined with the current Figure 1.

      We have moved one panel from Figure supplement 3 to the main Figure 1 (panel D) to show the INNOS projections more clearly.

      (6) Add more explanation about NOSp::palmi-3xHA reporter. Is it membrane-targeted reporter with the upstream promotor sequence of NOS? Or is it endogenous NOS that was tagged with palmi-3xHA?

      It is a membrane-targeted reporter driven by the upstream promoter sequence of the NOS gene. The construct is delivered by plasmid injection. We have clarified the description in the text and the figure legend (“The four apical organ cells, but not the eyes, were also labelled with a transiently expressed NOS-reporter transgene. This transgene contains the upstream promoter sequence of the NOS gene that drives a membrane-targeted palmitoylated tdTomato reporter (Figure 1F).” and “Expression of a membrane-targeted reporter driven by the NOS regulatory region (NOSp::palmi-3xHA-Tomato; magenta”). More details are in the Methods section under Transient transgenesis.

      (7) Show lack of anti-NOS immunostaining in NOS mutant to warrant specificity of the antibody. The subcellular localization of NOS in the dendritic arbors of INNOS is essential for the proposed model. The immunostaining image in Figure 4 -figure supplement 3D can be in the main figure. Is NOS also in the axons of INNOS?

      We further optimised the immunostaining with the NOS antibodies in WT and NOS mutant and added the staining data to the main Figure 1G and Figure 3—supplement 1. NOS was observed to be clearly localised in a region in the neuropil corresponding to the dendritic site of the INNOS cells. The staining was completely absent in larvae of both NOS knockout alleles. We have also added these explanations to the text.

      (8) “that was defective in NOS mutant (Figure 5I)” should be corrected as “Figure 5H”.

      We have restructured this part of the text and figure, the data from the Ser-h1 cells are now in Figure 5 - figure supplement 2.

      (9) Related to Figure 7, how do the larvae respond to a sequence of UV light (e.g. ten times 0.5s ON and 0.5 OFF) or ramping up/down UV light? Can the model make any predictions?

      We carried out new experiments and generated new model predictions. In Figure 6 – figure supplement 7 we show how changing the amplitude and duration of a single UV stimulation influences the response and the model output. In Figure 6 – figure supplement 8 we show how the model behaves when we provide repeated stimulations.

      (10) What are the knock down efficiency of NOS and NIT-GC morpholnio?

      We have now quantified the fluorescence intensity by immunostaining in wild-type and morphant larvae. We show the data for NIT-GC1 and NIT-GC2 in Figure 4—supplement 3F. The knock downs are very efficient. For NOS, we did not do morpholino experiments since we have two null alleles (Figure 3 – figure supplement 1).

      Reviewer #2 (Recommendations For The Authors):

      Why is the behavior of the control animals so different in panels B and C of Figure 3? One group reaches only 15 mm vertical position and the other reaches ~60 mm. Is this just batch variation? Am I missing an experimental variable here? If this is indeed batch variation, then some additional text and statistical analyses might help the reader interpret the behavioral data.

      Thank you for pointing this out. This was a mistake in the original Figure 3C of the unit on the Y axis. We have now corrected this. Additionally we have added statistical tests.

      Really, that’s the only, relatively minor issue I could find. This was a pleasure to read, and I learned a lot. Congratulations on an excellent study.

      Thanks a lot for these comments.

      Reviewer #3 (Recommendations For The Authors):

      Page 3, Figure 1 A: The authors reconstructed the cPRC circuit in 3-day-old larvae and detected NOS gene expression in 2-day-old larvae in Figure 1D&E. Can the authors provide a cPRC circuit reconstruction for 2-day-old larvae?

      We do not have a full connectome of the 2-day-old larva. We also show NOS gene expression in three-day-old larvae (Figure 1 – figure supplement 2).

      Page 4, Figure 2: NO produced by UV/violet stimulation to cPRCs: Including a diagram of the larvae would enhance reader understanding.

      We have added a schematic diagram to Figure 2.

      Page 4: Two Platynereis NOS knockout lines (NOSΔ11/Δ11 and Δ23/Δ23) using the CRISPR/Cas9: Could you direct me to the knockout conformation results for NOS knockout lines NOSΔ11/Δ11 and Δ23/Δ23?

      This is shown in Figure 3 – figure supplement 1 (genetic deletion and loss of antibody signal).

      Page 4: “Three day-old but not two-day-old NOS-mutant larvae also showed reduced phototactic behaviour, suggesting a function for NOS in the visual eyes that mediate three-day-old phototaxis” This sentence is unclear. Why is it only three days old but not two days old?

      2-day-old and 3-day-old larvae have very different type of phototaxis. 2-day-old larvae use their eyspots and show non-visual helical phototaxis. 3-day-old larvae use their visual (‘adult’) eyes for visual phototaxis. We have clarified this sentence: “Given that phototaxis in 1 and 2-day-old trochophore larvae is mediated by their non-visual eyespots (Jékely et al., 2008) and in 3-day-old nectochaete larvae by the visual eyes (Randel et al., 2014), these data suggest a function for NOS in the visual eyes (Figure 3D and Figure 3—figure supplement 1E).”

      Page 5: “Figure 3. NOS is required for UV avoidance in Platynereis larvae.” Detailed experiment setup, if possible, schematic or real setup images would help to replicate the experiments.

      We have added a schematic diagram to Figure 3. 

      Page 5: “All trajectories start at 0 x and y position and time 0 corresponding to 10 sec after the onset of 395 nm stimulation from the side.” How did the authors determine the 10-second duration? Are there any reasons for this choice?

      In our experimental setup we had a limit of tracking individual larvae of approximately 40 seconds. We therefore restricted our analysis to a 10 sec pre and 30 sec post-stimulus interval.

      “(A) Swimming trajectories of wild type (WT, n=32) and NOS mutant (NOSΔ11/Δ11, n=26 and NOSΔ23/ Δ23, n=47) three-day-old larvae.” Authors keep shifting between 2-day and 3-day-old larval data in Figures 1, 2, and 3, causing inconsistency.

      Due to their elongated shape and active muscular contractions it is difficult to carry out calcium imaging experiments with 3-day-old larvae. These experiments were therefore done with 2-day-old larvae. However, we did our behavioural experiments with both 2- and 3-day-old larvae. We observed similar patterns of UV-avoidance behaviour, NOS gene expression, ciliary activity, and mutant phenotype. We are therefore confident that the cPRC responses and circuit activity are similar across these two stages.

      Page 5: “Figure 3. NOS is required for UV avoidance in Platynereis larvae.” Why didn’t the authors present any statistics on the plots? A statistical test is required to prove that the difference is significant.

      We have added statistical tests to the data in Figure 3.

      Page 5: “Analysis of sGCs in Platynereis indicated that these genes are not expressed in any of the cells of the cPRC circuit (not shown and (Verasztó et al., 2017)).” Hence the data is relevant; please provide this data in supplementary.

      We re-analysed previous single-cell data (Achim et al., 2018) and found no sGC homologues detected in cPRC and INNOS. However, we detected an sGCβ subunit in the INRGWa cells. We added these data to the source data and amended the figure and the text. “Analysis of sGCs in Platynereis showed that these genes were not expressed in cPRC or INNOS cells, we only detected expression of an sGCβ subunit in the INRGWa cells (Figure 6B).”

      Page 10: “Diagram of the mathematical model with the components, interactions, parameters and equations used to model Ca dynamics.” It would be helpful to include a legend explaining the meaning of each term (e.g. what is K, Co, Cp etc) and arrow color. Additionally, the main text should provide detailed descriptions to aid understanding.

      We have added a new diagram of the model and its parameters to Figure 6. In addition, in the Methods section we have an extensive description of the model and all its parameters.

      Page 12: “This activated state is maintained for several tens of second.” Can authors point out the data?

      We point to the relevant figure now. “The high-Ca2+-state is then maintained for several tens of second (Figure 4A, B).”

      Page 13: “In the Platynereis circuit, our mathematical model indicates that the magnitude of the NO-dependent signal depends on the intensity and duration of the UV/violet stimulus.” Did the authors conduct experiments of various intensity and duration apart from modelling? If not, such experiments should be carried out to validate the mathematical model.

      We have carried out new calcium imaging experiment with varying duration and intensity of the stimulation. These experiments were key to the revision of the model because they clearly pointed to the time-invariance of the NO-dependent peak in the cPRCs. These data and their model fits are summarised in Figure 6 – figure supplement 7.

      References

      Bauknecht P, Jékely G. 2017. Ancient coexistence of norepinephrine, tyramine, and octopamine signaling in bilaterians. BMC Biology 15. doi:10.1186/s12915-016-0341-7

      Jékely G, Colombelli J, Hausen H, Guy K, Stelzer E, Nédélec F, Arendt D. 2008. Mechanism of phototaxis in marine zooplankton. Nature 456:395–399. doi:10.1038/nature07590

      Jékely G, Yuste R. 2024. Nonsynaptic encoding of behavior by neuropeptides. Current Opinion in Behavioral Sciences 60:101456. doi:10.1016/j.cobeha.2024.101456

      Randel N, Asadulina A, Bezares-Calderón LA, Verasztó C, Williams EA, Conzelmann M, Shahidi R, Jékely G. 2014. Neuronal connectome of a sensory-motor circuit for visual navigation. eLife 3. doi:10.7554/elife.02730

      Verasztó C, Gühmann M, Jia H, Rajan VBV, Bezares-Calderón LA, Piñeiro-Lopez C, Randel N, Shahidi R, Michiels NK, Yokoyama S, Tessmar-Raible K, Jékely G. 2018. Ciliary and rhabdomeric photoreceptor-cell circuits form a spectral depth gauge in marine zooplankton. eLife 7. doi:10.7554/elife.36440

      Verasztó C, Jasek S, Gühmann M, Bezares-Calderón LA, Williams EA, Shahidi R, Jékely G. 2025. Whole-body connectome of a segmented annelid larva. eLife 13. doi:10.7554/elife.97964.3

      Williams EA, Verasztó C, Jasek S, Conzelmann M, Shahidi R, Bauknecht P, Mirabeau O, Jékely G. 2017. Synaptic and peptidergic connectome of a neurosecretory center in the annelid brain. eLife 6. doi:10.7554/elife.26349

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank both editors and the three reviewers for their positive feedback on the revised manuscript. Following suggestions from Reviewer #1, we cite additional literature to clarify the use of ’fitness potential’ and we have revised the caption of Figure 3 to better explain how we generate the panel of mutants. We have also added a sentence to emphasize that the selection coefficient should match the time-scale of bottleneck effects in the evolution environment.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors ask whether a simple whole-head spectral power analysis of human magnetoencephalography data recorded at rest in a large cohort of adults shows robust effects of age, and their results provide compelling evidence that it does. The relative simplicity of the analysis is a major strength of the paper, and the authors are careful to control for many different confounds - although perhaps highly correlated factors like brain anatomy still pose a slight issue. The paper provides a valuable power analysis framework that should inform researchers across the broader neuroimaging community

      Many thanks to the reviewers and editorial team. This is an insightful and engaging set of reviews with a range of productive suggestions. We’re pleased that the strengths of this approach show through and that this can be a positive contribution to the community.

      We have implemented the large majority of suggestions and believe that the paper is greatly improved with them in place.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a careful, well-powered treatment of age effects in resting-state MEG. Rather than extracting (say) complex connectivity measures, the authors look at the 'simplest possible thing': changes in the overall power spectrum across age.

      Strengths:

      They find significant age-related changes at different frequency bands: broadly, attenuation at low-frequency (alpha) and increased beta. These patterns are identified in a large dataset (CamCAN) and then verified in other public data.

      Weaknesses:

      Some secondary interpretations (what is "unique" to age vs global anatomy) may go beyond what the statistics strictly warrant in the current form, but these can be tightened with (I think, fairly quick) additions already foreshadowed by the authors' own analyses.

      Aims:

      The authors set out to replace piecemeal, band-by-band ageing claims with t-maps, and Cohen's f2 over sensors×frequency ("GLM-Spectrum").

      On CamCAN, six spatio-spectral peaks survive relatively strict statistical controls. The larger effects are in low-frequency and upper-alpha/beta ranges (f2 approx 0.2-0.3), while lower-alpha and gamma reach significance but with small practical impact (f2 < 0.075). A nice finding is that the same qualitative profile appears in three additional independent datasets.

      Two analyses are especially interesting. First, the authors show a difference between absolute and relative spectral magnitude (basically, within-subject normalization). Relative scaling sharpens the spectral specificity of the spatial maps, while absolute magnitude is dominated by a broad spatial mode that correlates positively across frequencies, likely reflecting head-position/field-spread factors. The replication of the main age profile is robust to preprocessing decisions (e.g., SSS movement compensation choices) - the bigger determinant of the effect is whether they apply sensor normalization (relative vs absolute).

      Second, lots of brain-related things might be related to age, and the authors spend some time trying to back out confounds/covariates. This section is handled transparently (in general, I found the writing style very clear throughout) - they examine single covariates (sex, BP, GGMV, etc.) and compare simple vs partial age effects. For example, aging is correlated with reductions in global grey-matter volume (GGMV), but it would be nice to find a measure that is independent of this: controlling for GGMV (via a linear model) reduces age-related effect sizes heterogeneously across space/frequency but does not eliminate them, a nuance the authors treat carefully.

      Thank you for this concise summary of the work. We’re glad that the strengths of the approach come through clearly.

      This is a nice paper, and I have only a few concrete suggestions:

      (1) High-gamma 

      There can be a lot of EMG / eye movement contamination (I know these were RS eyes closed data, but still..) above 30-40 Hz, and these effects are the weakest anyway. Could you add an analysis (e.g., ICA/label-based muscle component removal) and show the gamma band's sensitivity to that step? Or just note this point more clearly?

      Thanks for this suggestion. We agree that there is concern about eye movements for these gamma band analyses. The ICA preprocessing we conducted was relatively thorough, and we were able to remove EOG-related components from the majority of datasets, even though the experimental protocol involved eyes closed at rest. It is possible that some components were missed, but they would require more advanced labelling tools or manual intervention to identify.

      There are alternative automated ICA labelling tools that could be used, but these are either not suitable for MEG (ICLabel; https://labeling.ucsd.edu/tutorial, https://mne.tools/mne-icalabel/dev/api/iclabel.html) or optimised for data from CTF/4D systems (MEGNet; https://doi.org/10.1016/j.neuroimage.2021.118402). It is beyond our capacity to modify one of these tools for CamCAN for the current analyses.

      It was much more straightforward to rerun the analysis without the ICA step to see the impact of reintroducing all ocular artefacts removed from the v1 analysis. These results are shown in Supplemental section XXX and replicated below.

      We have added the following Figure 11 and text to the paper.

      Main Text

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged and low-gamma effects are reduced (see supplemental section A.3). “

      Supplemental materials

      “The age results may be contaminated in some way by residual cardiac or ocular artefacts that are not removed during preprocessing. Though ICA denoising was applied, it is possible that some artefactual components were not identified and removed from the dataset. The results at low frequencies and in the gamma range are most likely to be directly impacted by this contamination.

      To explore the impact this has on our analysis, we reran the core GLM effect of age on the data with no ICA artefact rejection at all, allowing all eye movements and heart rate components to remain in the data (Figure 10).

      This no-ICA analysis has three differences to the original in the publication. The low-frequency decrease with age is stronger and more widespread in the analysis that removes ocular artefacts with ICA. A large negative effect is visible in both analyses, though without ICA several frontal and temporal sensors no longer show significant effects. Similarly, the effect size of the low-frequency effect of age is substantially larger when ICA denoising is applied.

      The low-alpha effect was strongly reduced in the analyses that do not remove artefacts with ICA. A large central-occipital group of sensors shows an effect between 7 and 8.5 Hz in the ICA analysis, but this is reduced to a single sensor at 8 Hz when ICA is not computed.

      At high frequencies, the age effect in frontal sensors is larger with ICA, and the age effect in posterior sensors is larger without ICA, though the position and frequencies of significant effects are largely unchanged. The remaining effects in the high alpha and beta ranges are unchanged by application of ICA.

      Overall, ICA either improves the estimation of age effects (low-frequency, low-alpha, high-gamma) or has negligible effects (high-alpha, beta). Only the posterior low-gamma effect is reduced by ICA. Together, we take this as evidence that our core results are robust to interference by eye movements and that the ICA denoising is working effectively to reduce noise in the analysis.”

      (2) GGMV confound control 

      Controlling for GGMV reduces, but does not eliminate, age effects. I have a few questions about this: a) Could we see the residuals as a function of age? I wonder if there are non-linear effects or something else that the regression is not accounting for. Also, b) GGMV and age are highly colinear - is this an issue? Can regression really split them apart robustly? I think by some cunning orthogonalisation, you can compute the effect of age independent of GGVM. I don't think this is the same as the effect 'adjusted' for GGMV (which is what is shown here if I'm reading it correctly). Finally, of course, GGMV might actually be the thing you want to look at (because it might more accurately reflect clinical issues) - so strong correlations are not really a problem: I think really the focus might even be on using MEG to predict GGMV and controlling for age.

      This is an interesting area with some tricky interpretation. Thanks for the nudge to help us clarify further.

      We have added the following text and Figures 17 & 18 to the paper to clarify these points.

      Main Text Section 2.8

      “It is important to note that the correlation between age and GGMV does impact the interpretation of the GLM results, but does not prevent the model fit. We explore the model validation and diagnostics in detail in Supplemental Section A.6. In brief, the model is able to separate the unique contributions of age and GGMV. However, the correlation between factors leads to an inflation in the standard error of the estimates. A hypothesis test on these estimates is valid. However, the inflated variance reduces our ability to detect statistically significant partial effects.”

      Supplemental Section A.6

      “There is a strong correlation between age and Global Grey Matter Volume (Pearson’s r=-0.75). The shared variance arising from this collinearity adds nuance to the interpretation of the results, which we explore in more detail in this section.

      Firstly, the sum-square residuals for the group-level model fit including both Age and GGMV are shown as a function of frequency in Figure 17. We see that the residuals broadly follow the overall distribution of variance in the data, peaking at low frequencies and in the alpha range. These frequency ranges are where the strongest signal is visible, but also the highest variability between participants. We would expect that the group model would not perform so well in the points of greatest variability. Importantly, though the residuals are relatively high in the alpha, this is still in the context of a very well-performing model with R2 values of around 80%. “

      “Secondly, the correlation between age and GGMV is not inherently problematic for the GLM, but it does add complexity and nuance to the interpretation of the results. Some additional model validation statistics are shown in Figure 18. “

      “The singular value spectrum of the design matrix indicates whether a design is low-rank, the smallest singular value in this case in 0.37 which indicates that there isn’t a rank deficiency which would prevent us from estimating the model. Though we can estimate the model, correlated regressors can reduce its efficiency. The variance inflation factors for the joint AGE-GGMV model are above 1 for both parametric regressors, indicating that the standard errors of their estimates are inflated. The VIF of 2.35 indicates that the standard errors of this joint model are around sqrt(2.35) = 1.533 times greater than they would be in a separate or uncorrelated model. Though there is no hard rule for this, the literature generally suggests that a VIF above 5 (or sometimes 10) indicates severe multicollinearity.”

      Including additional regressors in the model can change the age estimate by ‘partialling’ out the variance that can be attributed to the other variables and by inflating standard errors. In this specific case of GGMV, the partialled estimates are reduced heterogeneously across space and frequency, and the amount of inflation is at a tolerable level. Overall, the regression is able to separate the unique effects of age and GGMV, at the cost of this inflation in the associated standard errors.”

      Reviewer #2 (Public review):

      This paper describes the application of the "GLM-Spectrum" mass univariate approach to examine the effects of age on M/EEG power spectra. Its strengths include promotion of the unbiased approach, suitable for future meta/mega-analyses, and the provision of effect sizes for powering future studies. These are useful contributions to the literature. What is perhaps lacking is a discussion of the limitations of this approach, in comparison to other methods.

      Thank you for this summary and the thoughtful review. The emphasis for this paper is exactly on the points you highlighted, and we’re glad that this has come across well. We agree that a broader comparison to other methods would be a useful addition and have included a series of additional discussion points to address this.

      We will take the opportunity to reclarify that our intention for this method is not to replace other, more complex or targeted approaches, but to establish a more generalisable foundation for their development. We’re fully supportive of other approaches and are working on their application ourselves. On reflection, this was not clear enough in the first submission, and we have added the following text to the introduction to clarify.

      “Reporting of whole-head and full-frequency spectra of effect estimates would make it straightforward to aggregate across studies and eventually enable identification of sub-threshold effects that may be missed in single analyses but are consistent across studies. We argue that this approach provides a generalisable foundation that can support more complex analyses with a frequency component (such as aperiodic slopes, burst detection, and dynamic functional networks) that require more researcher degrees of freedom.”

      An analogy is the mass univariate approach to spatial localisation of effects in fMRI/PET images. This approach is unbiased by prior assumptions about the organisation of the brain, but potentially also less sensitive, by ignoring that prior knowledge. For example, a voxelwise univariate approach is less sensitive to detecting effects in functionally homogeneous brain regions, where SNR can be increased by averaging over voxels.

      In the context of power spectra, the authors' approach deliberately ignores knowledge about the dominant frequency bands/oscillations in human power spectra. This is in contrast to approaches like FOOOF and IRASA, which explicitly parametrise frequency components. I am not saying these methods are better; I just think that the authors should acknowledge that these approaches have advantages over their mass univariate approach (in sensitivity and interpretation; see below). I guess it is a type of bias-sensitivity trade-off: the authors want to avoid bias, but they should acknowledge the corresponding loss of sensitivity, as well as loss of interpretation compared to model-based approaches (i.e, models that parameterise frequency; I don't mean the statistical models for each frequency separately).

      This is an important point, and we are in complete agreement about the importance of giving a balanced description of how this approach fits within the broader literature. We have added the following paragraph to the discussion on limitations to lay this out more clearly.

      “Our approach promotes an exploratory and unbiased approach to quantifying the age effect on neuronal power spectra, which is intended to complement more focused analyses. This has the benefit of reducing researchers’ degrees of freedom and of being broadly generalisable. These come at the cost of a loss in specificity and in sensitivity. Our approach does not specifically quantify features derived from the power spectrum, such as alpha-peak frequency or the aperiodic component of the spectrum. These features are mixed into our full-spectrum estimates but not directly quantified. Thus, they can be challenging to interpret from our approach. Secondly, the mass-univariate approach suffers from a potential loss in sensitivity compared to results that aggregate across spatial or spectral regions that contain consistent results. Where a region or frequency band of interest can be supported from the literature, an approach focusing on a single region has the benefit of reduced noise by averaging estimates from a larger range of observations. Finally, models that consider the whole shape of the spectrum [Donoghue et al. 2021] would also be able to combine information across a range of frequencies rather than depending on a single frequency bin for each estimate. These models have the additional benefit that their parameters are often directly interpretable as features of interest, such as spectral slope or peak frequency.”

      An example of the interpretational loss can be seen in the authors' observation of opposite-signed effects of age around the alpha peak. While the authors acknowledge that this pattern can arise from a reduction in alpha frequency with age, this is an indirect inference, and a direct (and likely much more sensitive) approach would be to parametrise and estimate the peak alpha frequency directly for each participant, as done with FOOOF for example (possibly with group priors, as in Medrano et al, 2025, EJN). The authors emphasise the nonlinear effects of age in Figure 2A, but their approach cannot test this directly (e.g., in terms of plotting effects of age on frequency, magnitude, and width for each participant), so for me, this figure illustrates a weakness of their approach, not a strength.

      We agree that this point might be misleading within its own figure and have moved the result to the supplemental material with a more lightly phrased wording in the main text. Figure 3 on effect sizes has moved to Figure 2, and a new Figure 3 shows the quadratic effect of age as suggested later in the review.

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged, and low-gamma effects are reduced (see supplemental section A.3).“

      This supplemental section contains some additional content relevant to the next comment.

      Then I think the section "Two dissociable and opposite effects in the alpha range" in the Discussion section is confusing, because if there is a single reduction in alpha peak frequency and magnitude with age, then there is only one "effect", not "two dissociable" ones. If the authors do want to claim that there are two dissociable age effects within the alpha range, then they need to do a statistical test, e.g., that the topographies of low and high alpha are significantly different. This then reveals another limitation of the mass univariate approach - that space (channel) is not parametrised either - so one cannot test for significant channel x effect interactions within this framework, as necessary to really claim a dissociation (e.g., in underlying neural generators).

      As above, we agree that this can be misleading and that our intention got muddled in the heading and writing of this subsection. We are not intending to argue that there are definitely two distinct and different effects. We clarify this in the main text of our first version, but not clearly enough:

      “The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak”.

      We choose to present the first paragraph of this section as a discussion of two effects because this is what is already reported in the literature on changes in alpha power with age. Whilst the decrease in alpha frequency with age is well replicated, the decrease in alpha power is less consistently reported, and the literature shows a highly variable picture of the spatial pattern of this effect. We do not claim a strong dissociation based on these results, but both are reported within the literature.

      Our intention is to provide some clarity to this literature by taking a step back and looking at the ‘lay of the land’ in an unbiased way. With this approach, we see different effects of age on magnitude within a canonical alpha range that are completely separated in frequency. With this perspective, it is feasible that different publications taking different regions of interest, frequency band definitions, and processing options could lead to mixed reports of increases and/or decreases in alpha power with age.

      When single papers that take focused but inconsistent approaches report inconsistent results, we would argue that our approach leads to a substantial interpretational gain on the level of collections of papers in the literature.

      We have added the following content to supplemental section A.1 to clarify and Table 3 illustrates the issue. We have retained the paragraph discussing the possibility that these two effects combine to represent a shift in frequency of a single peak.

      Main text section 3.1, replacing the section on ‘Two dissociable effects…’

      “Reconciling conflicting reports of the ageing effect on alpha power.

      The literature exploring how ageing changes alpha power is heterogeneous. Papers that report results in resting alpha power, either from a canonical band or from an individual peak frequency, include reports of a variety of contrasting age effects. This includes positive correlations with age [Rempe et al., 2023, Stier et al., 2023], negative correlations [Thuwal et al., 2021, Lodder and van Putten, 2011, Medrano et al., 2025, Park et al., 2024], both positive and negative effects separated by space [Hoshi and Shigihara, 2020, Pathak et al., 2022], or null results when correcting for individual frequencies and aperiodic slopes [Scally et al., 2018, Merkin et al., 2023] (see supplemental section A.1 more detailed summary). There is broad variability in methodological approaches, which could account for the variety of results. Stier et al. [2023] suggest that analyses in sensor or source space may lead to different effects. Critically for our work, this variability also prevents formal aggregation of results and meta-analyses that could clarify the picture.

      We have proposed that, by taking a step back and tolerating a reduction in sensitivity, we can map out the whole-head whole-frequency structure of the age effect with minimal researcher degrees of freedom and bias. Investigating age effects as a complete spectrum shows that two contrasting age effects on alpha magnitude coexist in close proximity in space and frequency: an increase with a small effect size in central occipital sensors around 7-8.5 Hz and a decrease with a large effect size across a broad set of occipital, temporal and frontal sensors between 9.5-12.5 Hz. Different data samples and different data analysis choices, particularly the selection of regions of interest, source reconstruction, sensor normalisation, or correction of aperiodic components, might emphasise one effect or the other in each analysis. Whilst we have not conclusively explored all possible variants of these analyses, we have provided a framework that would allow future studies to perform formal comparisons and meta-analyses to resolve this bottleneck.”

      “Relationship between effects on alpha power and alpha individual frequency.

      The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak (Seen qualitatively in Figure 1A). This change in alpha peak frequency is highly replicable [Cesnaite et al., 2023, Dustman et al., 1993, Sahoo et al., 2020, Scally et al., 2018, Pathak et al., 2022, Zibrandtsen and Kjaer, 2021] and is a highly predictive spectral marker of ageing [Stier et al., 2024]. Decreases in alpha peak frequency have been linked to a decline in cognitive performance in healthy ageing [Cesnaite et al., 2023, Finley et al., 2024] and MCI [Garc´es et al., 2013, L´opez-Sanz et al., 2016, Puttaert et al., 2021].

      This compelling possibility that the age effect on alpha is a shift in a single peak is complicated by strong evidence for presence of multiple alpha peaks within individuals [Lodder and van Putten, 2011, Chiang et al., 2011, 2008, Klimesch, 1999], with distinct generators and functional relevance [Sokoliuk et al., 2019]. A complex pattern of changes in power, frequency, and spatial distribution likely underlies age-related change in alpha oscillations. Future work will need to explore all three features at the individual level to clearly illuminate the change.”

      Supplemental section A.1

      “Part of our motivation for this method is that variability in the methodological choices in different publications makes it difficult to aggregate varying results across the literature. For example, though a decrease in alpha peak frequency with increasing age is reported highly consistently, the effect of age on alpha power is much more variable. Table 3 shows a representative sample of publications over the last 20 years that report a change in alpha power with age (note that this is intended to be a representative rather than an exhaustive list). Over half of publications (9/15) report a decrease in alpha power, whilst the remaining publications report an increase (2/15), both increases and decreases (1/15), a decrease but only without correcting for aperiodic slope (1/15), a quadratic effect (1/15), and no effect (1/15). Stier et al. [2023] suggest that the choice of analysis space (sensor space or source reconstruction) is likely a source of discrepancies between studies.

      Critically, it is difficult to reconcile these findings with the information reported in the publications. For example, even within the 9 publications that report a decrease, there is little correspondence in the spatial location of the effect. We argue that focused approaches cannot resolve this issue alone, as targeted analyses are more specific to each dataset and less generalisable.

      In the specific case of the mixed literature on change in alpha power with age, our results show that both effects are present and separated in frequency. It is feasible that the different methodological choices and datasets used by each study in our survey means that one or other of these two effects were emphasised. As a result, the literature may not be mixed in scientific terms, but that a rich pattern of results is obscured by methodological variability.”

      While the authors show that normalisation of each person's power spectra by the sum across frequencies helps improve some statistics, they might want to say more about disadvantages of this approach, e.g., loss of sensitivity to any effects (eg of age) that are broadly distributed across majority of frequencies, loss of real SI units (absolute effect sizes) (as well as problems if normalisation were used for techniques like FOOOF, where the 1/f exponent would be affected).

      This is an important point, and we have added the following text to clarify, with one exception. Firstly, the normalisation we applied scales the whole spectrum linearly and would change the 1/f intercept but not the 1/f^alpha exponent.

      Main text section 3.3

      “Both absolute and relative power measures are used throughout the literature, but there is little consensus about their interpretation [Sandre and Troller-Renfree, 2026]. We focus on relative power for most of our results. By normalising each participant’s power spectrum by the sum across frequencies, we found that the results gained specificity in frequency band and reduced concern about wide intersubject differences in overall variance. Though this improved some analyses, relative power has important drawbacks. In particular, it can reduce sensitivity to effects that are broadly distributed across the spectrum and uses arbitrary scaling rather than meaningful physical units. We support calls in the literature to report both relative and absolute power measures [Rempe et al., 2023, Sandre and Troller-Renfree, 2026].”

      The authors should give more information on how artifactual ICs were defined. This may be important for cardiac artefacts, since Schmidt et al (2004, eLife) have pointed out how "standard" ICA thresholds can fail to remove all cardiac effects. This is very important for the effects of age, given that age affects cardiac dynamics (even though the focus of Schmidt et al is the 1/f exponent, could residual cardiac effects cause artifactual age effects in current results, even above ~1Hz?).

      Artefactual components were estimated using standard tools in MNE python. Specifically:

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_ecg

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_eog

      We have clarified the text in the methods to make the overall process clearer.

      We believe that this process has been broadly effective, and we have rejected an average of 2.25 ECG components within each dataset. There remains a strong possibility that residual ECG artefact is present in the data.

      We have added the following text to methods section 4.2

      “Artefactual components relating to eye movements or the heart rate were automatically identified by correlation with the simultaneous EOG and ECG channels. ECG artefacts were identified using cross-trial phase statistics [Dammers et al., 2008] and an automatic threshold based on the sample rate of the data, as implemented in the mne.preprocessing.ICA.find_bads_ecg function in MNE Python. Between 0 and 3 EOG components were rejected in each dataset, with an average of 0.99 (standard deviation: 0.79) across all datasets. EOG artefacts were identified by correlation with the HEOG and VEOG channels, with a threshold set to r = 0.35, as implemented in the mne.preprocessing.ICA.find_bads_eog function in MNE Python. Between 0 and 5 ECG components were rejected in each dataset, with an average of 2.25 (standard deviation: 0.84) across all datasets. The continuous sensor data were then reconstructed without the influence of the components labelled as artefacts.”

      Similar to the response to Reviewer 1, we have not been able to rerun a more advanced ICA algorithm on the data but have repeated the analysis without any ICA to see if including all ocular and cardiac artefacts influences the results. This change does not introduce any new signal components to the results but does attenuate the low-frequency and low-alpha effects. With the assumption that our initial ICA analysis captured the majority of the largest ECG components, we are confident that our core findings are not compromised by cardiac artefacts.

      Please see the response to comments to Reviewer 1 for additional text in the manuscript on this point.

      The authors should clarify the precise maxfilter arguments, and explain what "reference" was used for the "trans" option - e.g., did the authors consider transforming the data to match a sphere at the centre of the helmet, which might not only remove some of the global power differences due to different head positions, but also be best for generalisation of the effect sizes they report to future studies (assuming the centre of the helmet is the most likely location on average)? And on that matter, did head positions actually differ by age at all?

      We have used the maxfilter files as provided by the CamCAN team and an equivalent implementation defined in OSL-ephys for the Oxford and Cambridge MEGUK data. We have clarified the text in section 4.2

      “All MEG data pre-processing was carried out using MNE-Python [Gramfort, 2013] and OSL-ephys [Quinn et al., 2022, van Es et al., 2025] using the OSL batch pre-processing tools. For the CamCAN data, we proceeded with analysis on post-maxfilter processed data provided by the CamCAN team. Briefly, the data were processed using AA [Cusack et al. 2015] with automatic bad channel detection (limited to 7 channels) with the origin set to the centre of a sphere fitted to the individual’s Polhemus head shape points. Maxfilter signal-space separation was performed with the temporal extension enabled (temporal window of 10 second and correlation threshold of r = 0.98). Head position was continuously estimated and compensated for during periods where the HPI coils were on. After maxfilter processing, head positions were translated into a default head position defined as a point relative to each individual’s origin in a head coordinate frame. Data with the full maxfilter processing and with the head position translation or movement compensation were extracted from the CamCAN database. An equivalent pipeline was implemented in OSL-Ephys and applied to the data from the MEG-UK datasets from Oxford and Cambridge. Files from the MEG-UK Nottingham dataset were processed with third-order gradiometry applied.”

      We found that head position in CamCAN does change with age. We have included the following in Supplemental section A.5

      “The head position of participants within the MEG sensor dewar is an important consideration that has the potential to change the signal-to-noise level of each individual data recording. There are significant differences in head position as a function of age in the CamCAN dataset in Y (front-back) direction indicating that older participants are seated further forward in the dewar than younger participants.”

      As a point of reference, this pattern is consistent with the largest change with age estimated using the un-normalised raw power spectra in Figure 5. A 1-95Hz region shows a change with age that is consistent with older participants sitting further forward in the dewar. This effect is absent in the relative power contrasts. 

      We have not fully explored this final point so have not included it in the main paper, but believe it is a useful addition to the discussion of the reviews.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, both reviewers are enthusiastic about your paper, indicating that it provides compelling empirical support for its key claims and that it represents a valuable theoretical advance for the field. They provide a number of comments that you may want to consider prior to finalizing the manuscript for publication. All the best, Redmond O'Connell

      Reviewer #1 (Recommendations for the authors):

      (1) It would be handy to see a single table listing each tested "analysis family" (e.g., sensors×frequency, source parcels×frequency), the multiple controls used, and the permutation count. I kept wanting to see this as I was reading to compare back and forth.

      This is a helpful suggestion, we have included the table in a new supplemental section which is referenced from the main text at the start of the results

      Main text section

      “A summary of all GLM analyses carried out in this work is available in supplemental section A.7.”

      With the following table included in supplemental section A.7

      (2) I loved the power-planning content (section 3.2, the table with peak f2, CIs, contour plot). I think you could somehow make this even more explicit because people will use it a lot - both for this age/MEG domain and more generally as a template for other types of power planning in the field. Perhaps a "How to use this paper to plan N" guide in a paragraph? Power analysis is surely both "important and difficult" - but also not impossible. A flowchart?

      We’re very glad that this section is working well and agree that the practical planning content should have been more constructive! We have added the following text as a guide

      Main text section 3.2

      “We propose the following steps as a practical guide for researchers looking to plan a data sample with a reasonable chance of correctly identifying a particular effect.

      (1) Define research question and identify previous results: Your research question must be defined in advance and well specified. There should be relevant data or literature that can be used to guide your decision.

      (2) Define the smallest effect size of interest and the decision criterion for the sample decision: It is critical to define the parameters of how you will make your sample size decision ahead of time. We recommend considering what the ’smallest effect size of interest’ [Anvari and Lakens, 2021] would be for your question. This is specific to your question and is about more than statistical significance. What effect size would indicate that there is a practically meaningful effect for the future literature to consider?

      (3) Estimate effect sizes to inform your decision: Either by aggregating reported statistics from the literature, or by dedicated processing of previous data, compute an estimate of the effect size. Effect sizes are only estimates, so it is important to compute confidence intervals around your estimate to get a measure of variability.

      (4) Compute power/precision curves assuming these results: Using the estimated effect sizes, compute power curves [Baker et al., 2021] that visualise the relationship between effect size, statistical power, and sample size.

      (6) Select the sample size that meets your pre-defined criteria: Using your definitions from step 2, and when considering the whole power curve, make a decision about what sample size would give you a reasonable chance of replicating the effect of interest.

      (7) Make note of any differences or deviations from this plan during your data collection: There are many practical reasons why your planned sample might not match the data acquired in practice. Such deviations from a plan are ok but should be acknowledged, and any mitigating steps explained [Lakens, 2024].

      (8) Document your process for inclusion in a preregistration or publication: Include details on how a sample size decision was made in your research outputs including preregistrations, preprints, and publications. These details will be useful for future researchers to understand your process and to implement their own.”

      Reviewer #2 (Recommendations for the authors):

      (1) Though they mention in the Discussion, the authors could have noted earlier in Section 2.1 that one does not need to assume a linear effect of age - one could use a polynomial expansion or even local splines within the same GLM framework. Indeed, it seems unlikely a priori that effects of age across the wide range of ages in the CamCAN data are linear for all frequencies.

      This is an important point, and one that was straightforward to implement in our model. Based on feedback on other sections, we have removed the previous Figure 2, moved Figure 3 on effect sizes to the second position, and added a new Figure 3 containing a model of the quadratic age effects. We believe that this is a substantial improvement in the paper.

      Main text section 2.3

      “(2.3) Quadratic effects of age

      The linear effect of age is a convenient and simple regression model. However, it makes a strong assumption that change with age is uniform across the whole age range. A second group-level GLM was computed with an additional regressor to quantify quadratic effects of age, which have been reported in the ageing literature [G´omez et al., 2013, Rempe et al., 2023, Stier et al., 2023]. The spectrum of t-values for the quadratic age predictor (Figure 3) shows significant effects for a U-shaped change with increasing age in the low frequency (1-5 Hz) range in central sensors. Significant inverted-U shaped effects are present in two frequency ranges in the beta band. A low-frequency beta effect is present in occipital and temporal sensors between 16 Hz and 20 Hz, whilst a second high-beta effect is present in central sensors between 24 Hz and 30.5 Hz. The low-frequency and low-beta effects overlap in space and frequency with linear effects, suggesting that the low-frequency change with age has both a linear increase with age and a U-shaped component, whilst the low-beta effect has a decrease with age and an inverted-U shaped component. The high-beta effect does not overlap in frequency with any of the reported linear effects. The effect sizes for quadratic effects range between Cohen’s F 2 values of 0.033 for low frequency to 0.076 for high beta and are generally lower than the effect sizes for linear change.”

      Main text section 2.4

      “(2.4) Sample size planning for effects of age on the neuronal power spectrum

      We use the 95% confidence intervals around the effect sizes to make recommendations for future sample sizes for future samples that plan to replicate these results. The observed power calculations have no bearing on the interpretation of the present results. Instead, they should be used as a general guideline for planning future studies. Table 1 gives a full summary of the peak statistics and future sample size range for the six age effects identified in Figure 1B. These results have implications for future sample planning for resting-state electrophysiology studies of ageing. Using the upper bound of the sample size estimates as a conservative estimate, the linear changes with age that have relatively large effect sizes would have well-powered replications with sample sizes of around 50-60 participants. However, the smaller linear effects and all the quadratic effects would require samples of 200 or more participants to have the same probability of detecting the effect if it is indeed present (Table 1). This indicates that study samples should be planned with the smallest effect of interest in mind and that ageing effects in different frequency bands may not all be well powered within the same sample.”

      Discussion section 3.1

      “Linear increase and inverted-U effects in the beta band. The literature consistently reports an increase in low-beta power with age [Gomez et al., 2013, Heinrichs-Graham and Wilson, 2016, Heinrichs-Graham et al., 2018, Hubner et al., 2018, Koyama et al., 1997, Rempe et al., 2023, Stier et al., 2023, Veldhuizen et al., 1993, Xifra-Porxas et al., 2019]. We observed significant inverted-U-shaped quadratic effects of age in two frequency ranges in the beta band, a posterior low-beta component (centred around 18 Hz) and an anterior high-beta component (centred around 25 Hz). This is broadly consistent with reports of quadratic ageing effects in the beta band in the literature [Rempe et al., 2023], though, to our knowledge, our results are the first to suggest a separation of effects into different parts of the beta range. These spectrum changes may be associated with age-related changes in underlying bursting dynamics [Brady et al., 2020, Power et al., 2023].”

      “Methods section 4.7

      Age plus quadratic age models: To move beyond linear change and explore U and inverted-U shaped effects of age, we fitted a group model with three regressors. One constant term, one z-transformed age regressor, and one z-transformed quadratic age regressor (age − mean(age))<sup>2</sup>.”

      (2) The authors could point out an obvious extension of their approach to mixed-effects models, which could properly model within- and between-participant effects (repeated measures), e.g., for longitudinal effects of ageing, GAMMs, etc, including hierarchical linear models that combine trials and participants in the same model, and potentially model trial-specific/stimulus effects, etc.

      This is an important point; we have added the following paragraph to the discussion section.

      Discussion section 3.5

      “Future extensions

      A clear future extension for this work is to formally incorporate estimates of within-subject variability into a mixed-effect model. These powerful models would enable modelling of both fixed effects and random effects, allowing researchers to account for variation within individuals over time and between individuals. LMMs also provide improved approaches for handling missing observations and unbalanced designs, making them especially useful for longitudinal and hierarchical data. A second expansion of this work could use Generalised Additive Mixed Models (GAMMs) to model non-linear relationships using smooth functions while also accounting for random effects. This may allow for more realistic representations of complex patterns of change across age. Linear Mixed Modelling comes with a substantial increase in researcher degrees of freedom and can be challenging to implement and report accurately [Meteyard and Davies, 2020]. We have used fixed-effects modelling in this work in line with our objectives of maintaining generalisability and minimising researcher degrees of freedom.”

      (3) Do the authors want to comment on why effect sizes in Figure 4Ci are so much higher for the Oxford sample? I know the authors' main point is that smaller samples lead to more variable estimates of the true effect size, but there are also other interpretations for these sample differences, e.g., recruitment differences, scanner differences, etc. Have the authors considered implementing empirically Bayesian approaches (like COMBAT, a type of mixed-effects model) to adjust for site differences in both offset and scaling, which I think should be a fairly simple extension of GLM-spectrum?

      We agree that our explanation of here is somewhat lacking. To be transparent, we have thought long and hard about this difference and cannot identify a clear reason why this difference is so striking. There is no apparent difference in the other covariates, such as head position, age, or gender, which could explain why the Oxford dataset has larger effect sizes in this specific frequency range (the results in the low-frequency and alpha ranges are consistent).

      We have added the following to the main text to highlight dedicated data harmonisation strategies that could be applied in this situation.

      Main text section 2.5

      “This may arise from relatively poor estimates of the population level variability from smaller data samples. It is possible that more systematic differences in the participant sampling, recruitment, and data acquisition process could contribute to between-site differences. At present, our analyses can show that the core effects of ageing are replicable across several datasets, though we have not formally combined these datasets into a single analysis. Formal methods for data harmonisation, such as ComBat [Johnson et al., 2006], could correct for additive and multiplicative differences in data scaling across sites to improve site comparisons.

      (4) There are a number of formatting problems - at least in the PDF provided to reviewers - for some symbols, e.g., "below ¡7 Hz" (I presume "j" was originally "~" or something?). Also, sometimes a figure is cited by a single, hyperlinked number, without the word "Figure".

      (5) Line 402: "influence" = "influenced"

      (6) Line 431: date for Gelman & Loken?

      Thank you for highlighting these, we have done a thorough proofread and fixed a large number of spelling and formatting issues.

    1. Author response:

      We thank the editors and the two reviewers for their careful evaluation of our study, as well as for their positive assessment of the research question, model reproducibility, and the results concerning temporal-order representations. We also agree with the core issues raised in the reviews: the current manuscript has not yet sufficiently distinguished descriptive changes in recurrent dynamics from the functional mechanisms underlying the behavioral advantage; the behavioral effect itself is small and close to the performance ceiling; the phase analyses require more stringent controls; and the specific role of short-term synaptic plasticity in the rhythmic advantage has not yet been directly tested. We plan to add the corresponding analyses in the revised manuscript and to temper several mechanistic claims.

      First, we would like to clarify three aspects of the study design and analysis pipeline. First, the rhythmic and arrhythmic conditions were not performed by two separately trained groups of networks. Each network was jointly trained using balanced batches containing rhythmic-match, rhythmic-non-match, arrhythmic-match, and arrhythmic-non-match trials. Therefore, the behavioral differences were compared within the same independently initialized network. We will revise the relevant descriptions in the Abstract, Results, and Methods to make this joint-training procedure more explicit. Second, Figures 2 and 3–4 used different dimensionality-reduction approaches because they addressed different analytical questions. In Figure 2, dPCA was applied to time-aligned population activity to separate task-related variance into temporal, ordinal-position, and stimulus-related components, allowing us to examine how rhythmicity affected each representational component. In contrast, Figures 3 and 4 used standard PCA to construct a low-dimensional population signal for spectral and phase analyses without explicitly demixing task variables. We will clarify this distinction and the rationale for the two analysis pipelines in the revised Methods. Third, STSP was not introduced as an auxiliary module. Previous theoretical and computational studies have suggested that STSP can contribute to working-memory maintenance (Mongillo et al., 2008; Masse et al., 2019). Based on previous studies, we incorporated STSP alongside persistent neural activity to examine how synaptic and consistent neuronal activities jointly support sequential working memory. Under the current architecture and training settings, our ablation experiments showed that RNNs without STSP had difficulty successfully learning the task. We will emphasize this result and further distinguish the roles of STSP in task learning, delay-period maintenance, and the rhythmic advantage. We will also explore whether vanilla RNNs can successfully learn the same task under alternative hyperparameter settings.

      Our core hypothesis is that temporal regularity improves the encoding of sequential information by organizing recurrent population dynamics and phase structure during the encoding period, and that this organization subsequently influences information maintenance during the delay period and behavioral performance. The current results establish a relationship between temporal regularity and encoding-period dynamics, but direct validation of how this organization influences subsequent working-memory maintenance remains insufficient. In the revision, we will focus on strengthening the link between the encoding and delay periods and test whether population dynamics and phase organization during encoding are associated with subsequent information maintenance and behavioral performance.

      To provide a more direct view of the network dynamics, we plan to add intuitive visualizations of network activity and phase structure, allowing readers to evaluate the reported temporal organization with less dependence on dimensionality reduction, filtering, and complex statistical processing.

      The current behavioral advantage of approximately 0.4 percentage points is small, and network performance in both conditions is close to ceiling. Following the reviewers’ suggestions, we plan to evaluate the stability of this effect across learning, random initializations, and representative hyperparameter settings, and to examine whether the rhythmic advantage becomes more pronounced under higher memory load when ceiling effects are reduced. We will also avoid equating statistical significance directly with functional importance.

      Although the manuscript already includes decoding and perturbation analyses of neural activity and synaptic efficacy during the delay period, we will perform additional delay-period analyses to more directly examine whether the temporal organization established during encoding is associated with subsequent working-memory maintenance. These analyses will help distinguish the contributions of rhythmic input to stimulus encoding and memory maintenance.

      We will also test the role of STSP more directly. The current delay-period shuffle results show that synaptic efficacy makes a functional contribution to delay-period maintenance, but this does not yet demonstrate that STSP specifically supports the rhythmic advantage. We will further illustrate the roles of STSP during encoding and delay-period maintenance. In the revision, we plan to increase the number of independent networks in the perturbation analysis and test the interaction between temporal regularity and STSP disruption.

      Finally, we will further tighten the conceptual framing of the manuscript by more clearly distinguishing stimulus-driven phase organization from self-sustained oscillations, and by clarifying that the analyzed signals are model population signals derived from firing-rate population activity rather than LFP or EEG field potentials. Where the current evidence is insufficient to support interpretations in terms of entrainment or self-sustained oscillations, we will adopt more cautious terminology and revise the corresponding conclusions accordingly. We will also clarify the relationship between our phase-organization results and the findings of Liebe et al. (2025) in the revised Results.

      Overall, the revised manuscript will more precisely frame the study around how temporal regularity improves sequential working memory and how this behavioral advantage is associated with the organization of recurrent population dynamics and synaptic-state representations during encoding. After completing the additional training and analyses, we will also make the corresponding model weights, training configurations, and necessary analysis files publicly available to improve reproducibility.

      We again thank the editors and the two reviewers for their detailed and constructive comments. We will revise the manuscript carefully on the basis of these suggestions.

      References

      Mongillo, G., Barak, O., & Tsodyks, M. (2008). Synaptic theory of working memory. Science, 319(5869), 1543–1546. https://doi.org/10.1126/science.1150769

      Masse, N. Y., Yang, G. R., Song, H. F., Wang, X.-J., & Freedman, D. J. (2019). Circuit mechanisms for the maintenance and manipulation of information in working memory. Nature Neuroscience, 22(7), 1159–1167. https://doi.org/10.1038/s41593-019-0414-3

      Liebe, S., Niediek, J., Pals, M., Reber, T. P., Faber, J., Boström, J., Elger, C. E., Macke, J. H., & Mormann, F. (2025). Phase of firing does not reflect temporal order in sequence memory of humans and recurrent neural networks. Nature Neuroscience, 28, 873–882. https://doi.org/10.1038/s41593-025-01893-7

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      The authors then use circular dichroism to show that aSyn with acK at position 43 has less alpha-helical content. From this result, they deduce that "only this site could potentially perturb aS function in neurotransmitter trafficking", but no experiments on neurotransmitter trafficking were performed.

      We agree with the reviewer that neurotransmitter trafficking studies would be interesting, but they would presumably require the use of acetylation mimic mutants (Lys-to-Gln mutations), which we would want to validate by comparison to our semi-synthetic proteins with authentic AcK. Such experiments are planned for a follow-up manuscript, and we will investigate the reviewer’s suggested experiment at that time. Thus, we have not modified the manuscript to address this issue.

      Subsequently, they measure the aggregation speed of the variants in seeded aggregation experiments with preformed fibrils (PFFs) from WT aSyn, and conclude that acK at positions 12, 43, and 80 yields slower aggregation. They reach similar conclusions when measuring seeded aggregation in primary cultures. As far as I understand it, the seeding experiments in cells use seeds that are assembled from partially acetylated alpha-synuclein, but that are made of non-acetylated wildtype alpha-synuclein, and the alpha-synuclein that is endogenous in the cells is also non-acetylated (or at least not beyond what happens in these cells at endogenous levels). It is therefore unclear how the cellular seeding experiments relate to the in vitro aggregation assays with (partially) acetylated substrates.

      We understand the reviewer’s concerns and have modified the manuscript to clarify that the method of in vitro seeding really reports on the impact of acetylation on the elongation phase of aggregation. We have also clarified that this is different than the role that acetylation plays in seeding cellular aggregation with pre-acetylated fibrils. We note that having the monomer population acetylated in cells presents technical challenges that might also be addressed with Gln mutant mimics, and we plan to pursue such experiments in the follow-up manuscript described above.

      Anyway, both aggregation experiments ignore that the structures of aSyn filaments in Parkinson's disease (PD) or multiple system atrophy (MSA) are different from those formed in these experiments, and that, therefore, the observed aggregation kinetics are likely irrelevant for the speed with which disease-relevant filaments form in the brain.

      Finally, the authors describe the cryo-EM structure of mixtures of acK80:WT aSyn filaments, which are predominantly made of WT aSyn, with a previously described structure. Filaments made of only acK80 aSyn have a modified arrangement of this structure, where the now neutral side chain of residue 80 packs inside a hydrophobic pocket. The authors discuss differences between the acK80 structures and those of other structures from in vitro assembled aSyn filaments, none of which are the same as those observed from PD or MSA brains, nor are any attempts made to transfer observations from the in vitro experiments to the structures of disease. The relevance of the cryo-EM structures for human disease, therefore, remains unclear.

      The Conclusion on p.20 mentions an interesting and valuable result: the authors used the acetylated recombinant proteins to determine the extent of acetylation within human protein samples by quantitative liquid chromatography MS (SI, Figures S41-S49). Their conclusion is that "The level of acetylation was variable - no clear trend was observed between healthy control and patients - nor between patients of different diseases (SI, Table S4, Supplementary Data 1)" This result implies that acetylation of aS is not directly related to its pathogenicity, which again adds doubts on the disease-relevance of the results described in the rest of the paper.

      We acknowledge the concerns raised in the above paragraphs and believe that they can all be addressed by clarifying our purpose. The different fibril polymorphs adopted in PD and MSA are likely the result of an interplay of many PTMs and non-proteinaceous cofactors. Therefore, we are not necessarily trying to claim that our AcK80 fold is populated in health or disease, but that by driving Lys80 acetylation, one could push fibrils to adopt this conformation, which is less aggregation-prone. A similar argument has been made in investigations of alpha-synuclein glycosylation and phosphorylation. Our results in Figure 9 imply that Lys80 acetylation could be increased with HDAC8 inhibition. We have revised the manuscript to make these ideas clearer, while being sure to acknowledge the limitations noted by Reviewer #1.

      Reviewer #2 (Public review):

      Weaknesses:

      SDS is not a good membrane model to investigate the effect of lysine acetylation on aS membrane binding because it is a harsh detergent and solubilizes membranes. Negatively charged vesicles or vesicles made of a mixture of lipids mimicking the lipid composition of synaptic vesicles are more accepted in the field to study aS-membrane interactions. The authors used such vesicles for the FCS experiments, and they could be used for the initial screening of the 12 lysine acetylated variants of aS.

      We have noted this shortcoming in revisions of our manuscript, but have not performed new experiments as we do not believe that using vesicles instead would change the conclusions of these experiments (that only AcK43 produces an effect, and a modest one at that).

      It would help the reader to have the experimental details (e.g., buffer, protein/lipid concentrations) for the different assays written in the figure legend.

      We have added additional detail to the figure captions.

      The authors use an assay consisting of mixing 10% fibrils + 90% monomer to investigate the effect of lysine acetylation on aS. However, the assay only probes fibril elongation and/or secondary processes. The current wording can be misleading, and the term aggregation could be replaced by seeding capacity for clarity. For example, the authors state that lysine acetylation at sites 12, 43, and 80 each inhibits aggregation, but this statement is not supported by the data. Instead, the data show that the acetylation at these sites slows down the fibril elongation and thus decreases the seeding capacity of aS fibrils. In order to state that lysine acetylation has an effect on aS aggregation, fibril formation, the author should use an assay where the de novo formation of fibrils is assessed, such as in the presence of lipid vesicles or under shaking conditions.

      As noted in our response to Reviewer #1, we have clarified which phase of aggregation we were investigating in our in vitro experiments.

      It is not clear from the EM data that the structures of the different lysine acetylated variants are different, unlike what is stated in the text.

      We feel that it is clear from structures in Figure 8 and the EM density maps in Figure S38 that the AcK80 fold is indeed different. Although the overall polymorphs are somewhat similar to WT, the position of K80 clearly changes upon acetylation, altering the local fold significantly and the global fold more moderately. We have added backbone RMSD calculations to quantify the differences in WT and AcK80 folds in Figure SX. They differ by ~5 Å in the fibril core region.

      Reviewer #3 (Public review):

      Weaknesses:

      The challenges the authors report regarding semi-synthetic access to aSyn are somewhat surprising, as this protein has been made by a variety of different semi-synthesis strategies in satisfactory yields and without similar problems being reported.

      We understand the reviewer’s surprise and have edited the manuscript to clarify that the NCL yields were not unusually low, but were comparable to ncAA yields, and since it is significantly easier to scan AcK positions using ncAAs, we felt that ncAAs are the method of choice in this case.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Cryo-EM data processing: particles were pre-processed in cryosparc and ChatGPT-generated scripts. Are priors on tilt and psi angles defined through this procedure? And are segments from individual filaments kept strictly in the same half-sets for gold-standard estimation of resolution? Absence of psi and tilt priors will lead to worse refinements than otherwise possible. Worse, a mixture of segments from different filaments into the two half-sets may lead to overestimated resolution estimates.

      We thank the reviewer for their thoughtful analysis of our cryo-EM data processing approach. In response, we have modified our approach and added additional commentary on processing to the Materials and Methods section.

      “CryoSPARC (.cs) files were then converted to RELION STAR files using the PyEM csparc2star.py script. Following this data conversion, a custom Python script developed with ChatGPT precisely determined the start and end coordinates of each fibril. This script processed the csparc2star.py star file output to create coordinate pairs based on the cryoSPARC Fibril ID, notably without transferring the original tilt and psi angles from CryoSPARC to RELION. The output of this custom script served as the input for the autopick RELION extraction step.

      To handle curved fibrils, a special segmentation strategy was implemented: the script traced the coordinates along the fibril, generating a new start and end coordinate pair for individal segments, with the segment's end coordinate assigned either after spanning 10 particles (around 50 nm in length) or when the end of the fibril was reached. This resulted in shorter, straighter segments for processing. Finally, the createAutopick function from cryolo_boxmanager_tools.py in crYOLO was utilized to generate a STAR file linking the final particle coordinates to their movie files, which was then used to perform particle extraction in RELION. The standard RELION image processing pipeline then followed, including 2D classification, refinement, and 3D classification.

      This approach resolves the issue with curved fibrils, but also results in all picked fibrils being 50 nm or shorter in RELION. Consequently, segments from the same fibril receive unique fibril IDs in RELION and may be split into different half-maps. Although this avoids random splitting of individual particles without regard to their origin from the same fibril, it could still cause an overestimation of resolution.

      To assess any potential overestimation of resolution, another script was created to reassign the fibril ID of each particle in the RELION STAR file to that of the closest particles in the cryoSPRAC data, effectively ensuring that all particles from an individual fibril have the same fibril ID and are not assigned to different half-maps. This revised STAR file was then used as the input images STAR file for 3D auto-refinement in RELION. A comparison of maps before and after reassigning the fibril IDs shows negligible effects on the resulting structures and on their estimated resolution, indicating that the original procedure did not results in any significant overestimation of the resolution.”

      Author response image 1.

      RELION Cryo-EM maps before (gray) and after (yellow) fibril ID reassignment for (A) WT-A (B) WT-B (C) <sup>Ac</sup>K<sub>80</sub>-A (D) <sup>Ac</sup>K<sub>80</sub>-B, showing that the new processing approaches did not significantly change the fibril structures.

      (2) The 25% acK80 structure in S52 is understood to be wildtype-only. The authors mention in the main text that this is because of the strong density for the K80 side chain. An additional argument would be that an acK80 would leave an unshielded negative charge on the neighbouring E46, as K80 and E46 form a salt bridge in this structure.

      We appreciate the reviewer’s idea and have included a comment on the salt bridge impact.

      (3) It would be valuable to include side views of the density for all reported reconstructions to assess to what extent the beta-rungs are separated.

      The requested side views have been included in Figure SX.

      (4) Methods sections should be moved into the main text of the paper.

      This Materials and Methods portion of Supporting Information has been moved to the main text.

      Reviewer #3 (Recommendations for the authors):

      (1) We suggest removing "all" from the manuscript title as this claim might not hold up in the future.

      We understand the reviewer’s concern. Our title was meant to imply “all currently known” rather than “all” forever, but as this wording is awkward, we have deleted “all” as suggested.

      (2) The white font in Figure 1A is sometimes hard to read, especially on the yellow-green background between amino acids 70-80.

      We have changed this to black font.

      Figure 1 panels C-E are not referenced in the manuscript text?

      References to the Figure 1 panels have been added.

      In panel C, a structure is predicted, but based on what data? Why is there both a small and a big structure in panel C?

      Explanations of the Figure 1C images have been added to the caption.

      (3) Typo ε-acetyllysine -> Nε-acetyllysine

      This has been corrected throughout.

      (4) You report solubility issues during NCL, and these are typically alleviated by the use of chaotropes during ligation. Please specify what you mean by "standard NCL conditions" and include parameters such as guanidine concentration, pH, concentrations, volumes, and temperature.

      These details have now been added to the methods section.

      (5) According to Figure 2D, the desulfurization was incomplete. Please explain.

      The small peak observed next to the product peak is an adduct with sinapic acid (+206 Da), the matrix mixed in for MALDI acquisition. We do not see a sign of +32/64Da peak which would correspond to incomplete desulfurization

      Author response image 2.

      (6) Please include sequences of your constructs for ncAA mutagenesis. Especially, which intein was attached at what position. This is important for other groups to fully understand the production of acetylated aSyn variants.

      The full DNA sequence has been added to Supporting Information. This plasmid has also been reported previously in the referenced publications.

      (7) You state ncAA mutagenesis "yielded 0.11-1.5 mg" aSyn. I guess this is per-liter culture expression medium?

      This has been corrected.

      (8) The resolution/DPI of Figures 4, 5, S14-16, and S18 should be increased. They look blurred compared to Figures 6 or S19.

      These figures have been updated.

      (9) You provided only summaries in Figures 4 and 5 because of space restrictions. However, it would be nice to have at least some selected individual experiments (WT, acetylation K12/43/80) right next to it without the need to switch to S14, S15, or S16.

      Select experimental data has been added to main text Figures 4 and 5 as requested.

      (10) Please increase the size of the microscopy images in Figure 6 (in Figure S19, it looks much better).

      This size of the images in Figure 6 has been increased.

      (11) The labelling A1-A3 in Figure S33 was somehow unclear to me.

      These three panels show different sections of the TEM grid, illustrating heterogeneity in the 25% <sup>Ac</sup>K<sub>12</sub> fibrils that was not observed for fibrils of the other acetylation variants. We have clarified this in the figure caption.

      (12) In Figure 9, you present preliminary results regarding site-specific deacetylation by HDAC8. These results are interesting, but considering the exploratory nature of this in vitro experiment, I suggest toning down the highly enthusiastic discussion of potential in vivo effects.

      We understand the reviewer’s concern and have mitigated the claims of potential impact from these in vitro results.

    1. Author response:

      Reviewer #1 (Public review):

      My key concern and question is whether the cells presented in the manuscript are tanycytes. Tanycytes are specialized ependymoglial cells located in the circumventricular organs and are known to express specific markers. Importantly, they are not myelinated cells, which is a crucial distinction that the authors do not address.   

      We agree with the reviewer that tanycytes that have been described in the third ventricle have not been reported to be myelinated. However, our study was conducted on the hippocampal formation that borders the ventral horn of the lateral ventricles. We will include images of the myelin-forming ependymal cells that we refer to as tanycytes  

      Additionally, the methodologies described in the manuscript lack clarity and controls.

      We will expand on our methods section and include controls.

      For instance, the use of Cdh5-GCaMP882 mice is not adequately justified. It is unclear what these mice contribute to the study's objectives, particularly concerning the aim of investigating waste removal processes in the brain. Moreover, the rationale behind the purported "fluorophore uptake experiments" is unclear and appears to involve the uptake of fluorophore-labeled goat anti-rabbit secondary antibody, which seems implausible to me.

      Most experiments described in this manuscript were carried out on human brain. However, functional studies will have to be carried out on rodent brain. We will thus process rodent tissue as well to test whether our observed findings are consistent between human and rodent.  The animals used for these experiments were raised for bladder research and are wild type regarding neuronal and glial cells, particularly using the Cy3 channel. Utilizing these brains for our experiments has allowed us to test rodent tissue at both light- and electron-microscopic levels without having to sacrifice additional animals. The consistency of our findings between human and rodent brain further supports that the calcium indicator in the vascular system of these mice did not affect neurons or glial cells.

      Regarding the uptake experiment: When we initially discovered that myelin-forming macroglia form waste-internalizing glial canals within neuronal in spider brain it was unclear where the AQP4-immunoreactive cells were located. The somata of the myelin-forming cells lacked AQP4 immunoreactivity. Suspecting a synergistic interaction between the myelin-forming and AQP4 expressing cells whereby the myelinforming cells create the canal structure that sequesters waste from the neuron and the AQP4-expressing cells create a convective flow toward the waste-internalizing structures. However, unable to locate the somata of these cells, we submerged a freshly dissected spider brain with the attached surrounding tissue intact in physiological spider saline and slowly added blue vital dye solution to test which cells would internalize the dye. We then identified the (blue) cells in the lining of the dorsally located tubular system that we routinely detached from our brain preparations explaining why we were unable to locate these cells.  Immunolabeling of this tubular system revealed the cells that reside in the lining of this tubular system (see Author response images 1 and 2) the original (Figure 8) shows their long slender processes. Interestingly, this system is continuous with the stomatogastric system. 

      Author response image 1.

      Shows the proposed canal system in spiders

      Author response image 2.

      AQP4-immunoreactive cells in the spider primitive ventricular system that we localized due to similar uptake experiments we have conducted in mouse brain.

      We have utilized this method in mouse brain to test the validity of our postulation, that ependymal tanycytes internalize the presented fluorochrome from extracellular spaces and test which areas and structures may be involved in this uptake. As demonstrated in this experiment, the alveus, and a fine network of cell processes within the brain parenchyma show fluorescence, indicative that they internalize the fluorochrome from extracellular spaces. We used goat-coupled fluorochrome to further test with a FITCcoupled secondary antibody that the observed fluorescence is indeed due to uptake of the goat-coupled secondary antibody and not due to intrinsic autofluorescence. Control preparations lacked this fluorescence.  To further test our postulation that the uptake is indeed AQP4-mediated we have applied an AQP4-blocker, which showed a significantly reduced fluorochrome uptake compared to the controls without this blocker.

      As we state in the text, we are aware of the limitations of this experimental design, however, like in our spider experiments we consider these findings helpful as they likely show an overview of the cellular network in the hippocampus that governs waste-uptake and may help identify suitable target areas for similar studies on organotypic tissue cultures utilizing two-photon microscopy.

      The hypotheses and claims presented in this manuscript are not sufficiently substantiated and are conceptually unclear. The notion that amyloid beta and tau proteins play structural roles in a hypothesized "tanycytes"-derived canal network is not sufficiently supported by the evidence. Furthermore, the study lacks rigorous data to convincingly establish the proposed interactions between these proteins and the processes of waste internalization. In conclusion, due to conceptual and methodological issues, I consider the current evidence as inadequate to support the primary claims.

      We respectfully disagree with this comment and hope that the inclusion of additional evidence together with the clear visibility of this canal system in the spider brain will encourage the reviewer to investigate this possibility themselves. We cannot ignore large amounts of amyloid beta-immunolabeled receptacles emanating from tanysomes in swell-bodies and declare them fixation artifacts, particularly when we demonstrate the expression of Presenilin 1 and APP in swell bodies. We furthermore encourage the reviewer to revisit myelinated cells in the brain in both depictions in the available literature and actual brain preparations. We have not been able to locate actual electronmicrographs of longitudinal sections through neurons that show myelination past the axon hillock at the EM-level consistent with our current understanding of myelination. The only depiction of this form of myelination we found were schematic drawings. We will include several new images that show such longitudinal sections through neurons that are easily obtained and we have numerous additional images that we are happy to share. In all our preparations (mouse, rat and human) the myelination pattern is consistent with the images we will depict in new figures 1 and 2.  Not to bring attention to this inconsistency would be dishonest scientific conduct.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors propose the existence of an AQP4-positive tanycyte-associated canal system in the hippocampus and suggest that this system participates in waste clearance and contributes to Alzheimer's disease pathology. Using histological, ultrastructural, immunohistochemical, and RNA-based approaches, the manuscript attempts to reinterpret amyloid-β plaques and tau-associated structures as components of a tanycyte-derived waste-internalization system. The work is conceptually ambitious and raises observations that may stimulate discussion regarding glial organization and waste clearance in the diseased brain.

      Strengths:

      A strength of the manuscript is the combination of imaging modalities and anatomical observations across mouse and human tissue. Some of the reported morphological features are intriguing and may warrant additional investigation. The study also attempts to integrate structural observations with broader hypotheses regarding neurodegeneration and Alzheimer's disease.

      Weaknesses:

      The central interpretation depends almost entirely on identifying the observed hippocampal structures as tanycytes, and the evidence supporting this conclusion remains insufficient. Tanycytes are classically associated with ventricular regions in circumventricular organs, particularly in the third ventricle and median eminence region, yet the manuscript does not provide sufficiently specific anatomical or molecular evidence to convincingly distinguish the described structures from astrocytic, ependymal, radial glial-like, oligodendroglial, myelin-associated, vascular-associated, or degenerative elements. The marker profile used throughout the study, particularly the reliance on AQP4 labeling and Luxol-positive structures, is not sufficiently selective to establish tanycyte identity, especially in pathological tissue where reactive glial changes may occur.

      As mentioned in our response to reviewer 1 we have now included additional experimental evidence that demonstrates the myelinated ependymal cells and additional gene expression experiments. 

      This becomes particularly important because the manuscript repeatedly interprets Luxolpositive and myelin-associated structures as tanycytic processes or "myelin-derived tanycyte protrusions," despite tanycytes not being known to produce myelin. Alternative explanations are not sufficiently explored. Some of the canal-like structures shown in Figure 4 also resemble vascular profiles, and additional vessel markers would be necessary to exclude this possibility.

      Several of the proposed structures and mechanisms are also difficult to reconcile with established cell biology and neuroanatomy. The introduction of new terminology such as "tanysomes," "waste receptacles," and "toroids" further extends the interpretation beyond what is currently demonstrated experimentally.

      The discussion and integration of the existing literature on tanycytes are also insufficient. Tanycytes themselves are not clearly introduced; the manuscript does not adequately discuss what is currently established regarding tanycyte anatomy, ventricular localization, morphology, and function. Foundational literature defining tanycyte biology, including work from the Prévot group or others, is largely absent despite its central importance to the field. Because the manuscript proposes a substantial departure from established neurobiological concepts, it is particularly important that previous literature be discussed comprehensively and critically. The current version does not sufficiently contextualize the proposed model within the existing literature on tanycyte, AQP4, glymphatic, and Alzheimer's disease, making it difficult to evaluate what is genuinely novel versus what is merely being reinterpreted. It is also not entirely clear what is genuinely new here compared with the authors' previous work, particularly reference 11, which appears to present a highly similar conceptual framework.

      More broadly, several of the manuscript's mechanistic conclusions extend well beyond the available evidence. The proposal that amyloid-β plaques and tau pathology represent hypertrophic tanycyte-derived waste structures is provocative and potentially interesting, but currently remains largely correlative and speculative. At several points, it becomes difficult to distinguish direct observations from broader mechanistic interpretation. The manuscript itself acknowledges that the proposed glial-canal hypothesis contradicts the current understanding of nervous system organization and states that ultrastructural serialsection analysis would be required to unambiguously determine the origin of the myelinated profiles described. This point is critical because the study's central conclusions depend on the assumption that these structures are tanycyte-derived. At present, this interpretation remains insufficiently demonstrated, which substantially limits the strength of the broader pathological and mechanistic conclusions proposed throughout the manuscript.

      Although access to human material is understandably limited, the study appears to include only one male and one female AD patient, making it difficult to assess the reproducibility or frequent these structures are across individuals and pathological conditions. The manuscript would benefit from clearer characterization of prevalence, reproducibility, and variability across samples.

      Overall, the manuscript presents an unconventional and thought-provoking model that may stimulate discussion. However, the evidence currently provided does not convincingly establish tanycyte identity for the described hippocampal structures, and several of the broader disease-related interpretations would require substantially stronger anatomical and molecular evidence before the proposed model can be convincingly supported.

      We agree with the reviewer that it is important to correctly investigate and describe cellular structure. The first author of this manuscript is a 30-year veteran of published cellular ultrastructure and the three-dimensional reconstruction of cells and entire cell networks. We have spent the last five years to try and confirm our current understanding of cellular structure, and we are unable to reproduce our current identification of cell structure. One example is mentioned in response to reviewer 1. We are unable to find electonmicrographs in publications that show myelination of neurons consistent with the countless schematic depictions available online and in the literature. This includes publications about myelination. Our findings are all consistent with the new images we will include in our revised manuscript. Longitudinal sections through neurons are easily obtained and neurons can be followed well beyond the axon hillock.

      A second example that is inconsistent with our current understanding of cells and biochemical processes in cells are ‘astrocytes’ and ‘reactive astrocytes’. We will show in the revised manuscript swell-bodies have no defined cytoplasm that every cell requires to fulfil basic cellular functions required for survival. It is very apparent to a structural expert that swell-bodies lack cytoplasm. The ‘consistency’ of this lacking cytoplasm is ‘inconsistent’ with fixation artifacts. In Author response image 3 we demonstrate that immunolabeling for AQP4 shows two types of structures. (1) Immunolabeled structures that are void of immunolabeling in their lumina and show receptacle-like immunoreactivity along the outside (panels I,J,L,M image below) consistent with swellbodies as indicated by the provided amyloid beta-stained and Luxol H&E-stained swellbodies that are associated with immunolabeled receptacle-shaped structures on the outside (panels K,N). A structurally trained eye quickly recognizes that these are not immunolabeled cells but are consistent with swell-bodies and referred to as ‘reactive astrocytes’ in the literature. We show in panel O what an immunolabeled cell looks like, with slender processes and immunolabeled cytoplasm. In this image in panels G1-4 we demonstrate that swell-bodies contain AQP4 mRNA explaining why they are immunoreactive for this protein. It is important that we bring attention to these details and do not randomly describe structure just based on a signal. This is exactly the point we make and I hope that the reviewer recognizes our expertise in cellular neuroscience and in particular in recognizing cellular structure. 

      Again, I urge the scientific community not to dismiss our findings but to actually study the validity of our findings. 

      Author response image 3.

      In the revised manuscript we will include additional experiments that we have carried out, that have also increased the sample number of tested human brain. We have in total so far investigated 13 different human brain samples, six of which are AD-affected. 

      We will revise our discussion to explain our observations in context with the literature better and include so much evidence in support of our hypothesis that it would be unreasonable to dismiss all this compelling and logical evidence.

      This research was not planned; it resulted from our accidental discovery in the spider system when our animals struggled with early onset neurodegeneration that we needed to address. This is when we recognized the waste-internalizing role of myelin in giant spider neurons. As we will discuss in the revised manuscript, such systems are highly conserved throughout evolution, and this is what made us realize that the only images of myelinated neurons we could find were either schematic drawings, single cross sections through myelinated cell profiles or very high magnification insets that also did not show the actual neuron that is myelinated.

      One can argue that there are both myelinated and unmyelinated axons. However, the varicose projections clearly originate in the ependymal lining and double-label for AQP4. This is consistent with our postulation and inconsistent with our current understanding of myelination. 

      I sincerely urge the neuroscience community to re-visit myelination in the brain, we have done this for the past five years with extensive experience in this field and the only hypothesis that is supported by our findings is presented in this manuscript. Please do not dismiss these findings, the spiders have uncovered a waste canal system in the brain and putting this system in place will help us to gain a better understanding regarding neurodegenerative diseases. Lastly, and maybe most importantly, understanding how this system works in spiders and how the myelin is structurally anchored to microtubule that were missing in our degenerating spiders has allowed us to identify the cause for this sudden neurodegeneration and rescue our tropical, cold-blooded spider colony by installing a new heating system and raising the room temperature so that the coldsensitive microtubules no not dissociate anymore.  

      We would like to thank the reviewers to strengthen the content of this manuscript with their critical comments, we hope that our revision will help clarify some doubts.

      (1) Pasquettaz R, Kolotuev I, Rohrbach A, Gouelle C, Pellerin L, Langlet F. Peculiar protrusions along tanycyte processes face diverse neural and nonneural cell types in the hypothalamic parenchyma. Journal of Comparative Neurology. 2021;529(3):553. doi: 10.1002/cne.24965. PubMed PMID: edsgcl.646848213.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editors for their positive assessment of our manuscript, and all the referees for their constructive comments, which have substantially improved this work. We welcome the opportunity to address referee #2's points for the public record, as they highlight key theoretical nuances and valuable future research directions.

      (1) We appreciate the reviewer's point that, in a noisy biological system, the increased motion parallax provided by a rich 3D depth structure naturally aids in separating translation from rotation. We fully agree on this point. Our argument aimed at highlighting a fundamental theoretical distinction. Pure algebraic decomposition algorithms are mathematically capable of solving heading on flat planes. The fact that human perception often shows biases in these zero-depth conditions, unless extra-retinal cues are present, suggests that the visual system does not rely on a generalized, global de-rotation algorithm. Instead, it relies on heuristic, depth-dependent structural signals (like motion parallax and retinal curl). We maintain that while depth certainly reduces noise, its strict necessity points toward an ecologically grounded control strategy rather than a noisy global decomposition process.

      (2) We think the reviewer raises a valid point regarding the exact neural locus of retinal curl encoding. It is true that Graziano et al. (1994) utilized centred spiral stimuli rather than the spatially offset curl geometries defined in our task. We view the spiral tuning of MSTd not as a direct, one-to-one mapping of full-field retinal curl, but rather as the foundational computational building block required to extract such a signal. We fully agree with the reviewer that identifying exactly where and how this population response is decoded into a unified, gaze-relative retinal curl signal remains an exciting and open empirical question for future neurophysiological research.

      (3) We acknowledge the reviewer's call for transparency here. The transition from well-documented multiplicative gain fields (which modulate response amplitude based on eye position) to a direct, localized gaze-centered inhibitory drive is indeed a theoretical abstraction in our model. We utilized this localized inhibition as a functional mechanism to demonstrate how sensory evidence and spatial priors might competitively interact within a standard Mexican-hat recurrent architecture. While gain fields clearly establish that parietal networks integrate gaze position, we agree that the exact local-circuit implementation mapping these gain fields to the specific inhibitory dynamics we modeled has yet to be empirically established.

      (4) The reviewer smartly questions whether the 3-5 second behavioral biases delay emerges from the 2.4-second smoothing window used in our computational flow manipulation. It is important to clarify that this 2.4-second window was used solely to stabilize the computed curl signal against high-frequency gait oscillations. This smoothing was restricted strictly to the modeling phase of the controller and neural model and was not applied to the participants' responses. The gradual build-up of their perceptual bias over 3-5 seconds represents their own intrinsic temporal integration of this trajectory, independent of the smoothing parameters used to smooth the curl used in the fitting of the controller and neural modelling.

      (5) We concede the reviewer's point that citing a computational model (Layton & Browning, 2014) does not constitute direct empirical evidence for a relationship between MSTd heading preferences and their receptive field locations. Our intention was to highlight a successful theoretical framework that elegantly organizes known properties of MSTd into a system capable of bypassing global de-rotation. We readily acknowledge that direct, single-cell empirical validation of this specific topographic relationship is currently lacking in the literature, and we appreciate the reviewer ensuring this distinction is clearly noted for the record.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides an important and biologically plausible account of how human perceptual judgments of heading direction are influenced by a specific pattern of motion in optic flow fields known as retinal curl. By combining psychophysical experiments and neural modeling, the authors demonstrate that what was previously considered an incidental "nuisance" signal actually serves as a functional control signal for estimating heading and steering toward a fixated target. While the evidence for the role of curl signals is convincing and advances our understanding of vision-based navigation, the work's impact would be strengthened by situating these findings among other cues that contribute to heading estimation, and by clarifying both the time course of these computations and their generalizability across different navigational contexts.

      We thank the editors and reviewers for their insightful feedback and positive assessment of our study. In this revised version, we have made substantial modifications to better situate our findings within the broader landscape of cues contributing to heading estimation, while also clarifying the time course of these computations and their generalizability across different navigational contexts. These points are included in new sections in the revised discussion.

      In addition, and following eLife guidelines, we have moved the methods to the end and make sure that the manuscript reads well without needing to go through methods first.

      Next, we address all the concerns raised by the reviewers.

      Reviewer #1 (Public review):

      We appreciate Reviewer #1’s very positive feedback. Incorporating the perspective of ‘incidental’ sensory signals is a valuable suggestion that aligns perfectly with our findings. We agree that this perspective significantly strengthens the impact of our paper.

      In the revised version we have added a last section in the Discussion (Generalizability and Testable Predictions) to comment on the functional utility of 'incidental' signals and incorporated the suggested references. In addition, in the same heading, we briefly elaborate on the predictions and generalizability of the model and possible manipulations that might affect the integration between sensory evidence (curl signal) and straight-ahead prior.

      Reviewer #1 (Recommendations for the authors):

      It would be great if the authors could discuss the implications and predictions of their model.

      First, from a broader perspective, the study forms an important piece in the emerging recognition that incidental sensory signals are not a nuisance to the sensorimotor system, but contain functionally relevant and effectively used visual signals (Rolfs & Schweitzer, 2022). The authors may want to appreciate their contribution to this perspective in the discussion of the impact of their results. Indeed, a similar shift in perspective has been realized in the recognition that saccade-induced motion signals are not entirely suppressed but play a functional role for gaze correction (Schweitzer & Rolfs, 2021).

      Second, the authors could spell out additional predictions of their proposal: What are experimental manipulations that could shift the balance between relying on a straight-ahead prior and sensory estimation of curl? What would happen in extreme cases of such sensory evidence? When would priors become overwhelmingly influential?

      As commented in the public review we have now included these two aspects in the discussion.

      Minor point: After equation 11, the authors may want to specify that, like position, gaze g is also coded as {x,y}, just like position p.

      While the neural model details have been moved to Appendix 2 (including this equation), we added text before Equation 22 (previously Eq. 11) clarifying that gaze is encoded in image coordinates. We omitted point index i because the equation applies to all image points relative to a given gaze g.

      Reviewer #2 (Public review):

      We appreciate the reviewer’s feedback regarding the formalization of our reference frames. We agree that certain definitions were implicitly assumed rather than explicitly stated. We have revised the manuscript to provide all necessary self-contained information, ensuring that the geometry of the task response and the definition of heading are unambiguous. In the last paragraph of the revised introduction, we make clear the response frame of reference which is also included in the caption of fig 1. Also, we have addressed the gap between the task response (in world coordinates) and the functional role of the controller. This is particularly discussed in the discussion (section: reference frames) in which we also provide (and rule out) potential alternatives to our response biases. We also address all the other points raised by the reviewer.

      Major issues:

      (1) The manuscript contains inconsistent, if not misleading, messaging about what information retinal curl does, and does not, provide regarding heading estimation. In the Abstract, the authors state: "We propose an alternative: the visual system utilizes retinal curl directly to estimate heading, rendering the explicit recovery of the FOE unnecessary." Based on my understanding of the rest of the manuscript, I find this statement to be a misrepresentation for two main reasons:

      (a) To "directly estimate heading" relative to what? When not qualified, most people interpret "heading" to mean an observer's heading relative to the world (or some allocentric reference frame). But retinal curl only gives information about an observer's heading relative to the point on which their eyes are fixated. Moreover, that point of fixation will change every few hundred milliseconds in natural viewing, so the retinal curl will change with each new fixation even as heading relative to the world remains unchanged. So I think most readers would grossly misinterpret the claim that retinal curl can be used "directly to estimate heading". Indeed, in the authors' controller model, the initial heading needs to be given, and then the controller can work. But from where does the visual system get the initial heading, since it does not come from curl? These issues are left hanging. Thus, while curl can provide a very useful input for steering toward a fixated target, other signals are needed to estimate heading relative to the world. This has to be made much clearer early on, and a conceptual schematic diagram might help. Also, the authors generally do not specify the reference frame of the variables they are talking about, leaving lots of room for misinterpretations. It should be clear each time they are talking about a variable, such as heading, whether it is relative to the fixation target, body, world, etc.

      In our study, participants were instructed to report their “perceived direction of self-motion” by aligning a rotational encoder (steering wheel) with the direction they felt they were moving within the 3D simulated scene. Consequently, participants reported their instantaneous heading in a world-centered reference frame, from which the 3D trajectories were reconstructed. Since the reviewer had to infer this information, we have clarified this point at the end of the introduction, legend of figure 1, methods (now at the end of the ms.) and discussion to ensure it is immediately evident.

      Participants were informed that the initial heading (i.e. θ<sub>0</sub> in our controller nomenclature) was oriented “straight ahead” relative to their body which was aligned longitudinally with the experimental room. We have modified Figure 1B and revised the Methods section to explicitly clarify this initial alignment and the instructions provided to participants.

      In the revised manuscript, we have clarified that while the participant’s report is world-centered, the retinal curl provides a gaze-relative heading signal. Although this was already mentioned, we emphasize this point. In natural navigation toward a fixated target, a world-centered vector is often unnecessary; an error signal indicating heading relative to fixation is sufficient (as the reviewer also notes). However, the initial alignment of the heading within the 3D scene allows the brain to “calibrate” this internal controller, mapping the retinal curl signal onto the 3D world coordinates required for the task. Ad commented above, a new section in the discussion addresses and hopefully clarifies the relation with the controller.

      The reviewer also asks how we can be certain that participants were reporting in world coordinates rather than an alternative frame, such as “heading relative to the fixation target.” We believe our “Cancelled Curl” (and over-cancelled) conditions provide the most compelling evidence to rule out this alternative. In these conditions, the physical position of the fixation target in the scene remained identical to the unaltered flow condition. If participants were simply reporting heading relative to the fixation target’s spatial location, the observed biases should have persisted regardless of the flow manipulation. Instead, the bias vanished when the curl was removed. This causal evidence proves that the bias is driven by the retinal motion signal (curl) rather than the spatial orientation of the eyes or the target’s position in the scene. Furthermore, the temporal evolution of the response supports a world-centered integration (in agreement with Warren et 2001 Nat Neuro.). For simulated straight paths, the perceived heading remains straight for the first few seconds (consistent with the initial world-centred alignment), with biases only emerging after approximately 3 seconds of integration (a point we elaborate on in our response to Reviewer #3). Had participants been responding based on a simple gaze-relative reference frame from the onset, these biases would have manifested significantly earlier. We have incorporated these points into the revised Discussion to better frame our findings alongside other cues, such as the Focus of Expansion (FOE) and egocentric visual direction that contribute to heading estimation.

      Finally, we have rephrased the abstract sentence for clarity. However, we maintain that the original premise remains valid once the world-centered initial heading is aligned with the gaze-centered reference frame.

      (b) It seems to me that retinal curl will depend on other variables, in addition to heading relative to the fixation target. For example, it seems to me that the magnitude of retinal curl will depend on self-motion speed, the depth structure of the scene, the angle of elevation of the fixated target, and perhaps others. This is not discussed at all, and many readers would get the misguided impression that there is a 1:1 mapping from curl to heading (relative to fixation). If I am right that this is not correct, it means that retinal curl can tell the observer whether to steer right or left to move toward the fixated target, but it cannot tell them how much to steer. Indeed, in the authors' controller model, there is a free parameter that calibrates curl to angle. It makes sense that this works to fit trajectory data that are given from a fixed environment, but it is unclear how the brain would use retinal curl to control steering when these other variables are uncertain or changing unpredictably. Moreover, how does the system change the mapping from curl to steering command as the location of fixation changes relative to the current heading? These are issues that need to be brought up in framing the problem and discussed at some length. If the authors can show mathematically that retinal curl is only dependent on heading (relative to fixation) and not any of these other variables, it would be very valuable to show the equations for this relationship.

      The reviewer notes that we must be clear about the relationship between curl and heading (relative to fixation) and the variables that affect curl. We also thank the reviewer for encouraging to add the equations that show the relation of curl with additional variables. We have now included these equations in appendix 1.

      Beyond the discrepancy between heading (θ) and gaze (ψ), curl is geometrically determined by translational self-motion speed (v), eye height (h), and pitch (α). More specifically, curl = (v.sinψcosα)/h. The derivation is now included in appendix 1. Since h = dsinα, where d is the 3D distance to the fixation point, we could express cos α as a function of distance. Certainly, there is not a 1:1 map from curl signal to heading relative to gaze (e.g. θ-ψ). Participant would need to know v and eye height plus extra-retinal information. Frenz et al (2003, Vis Res.) showed that people can estimate self-motion directly from optic flow, across different simulated eye height and gaze angle; extra-retinal information can, in addition, provide knowledge to ψ and α. It is then plausible that the visual system can use and transform the curl signal from a qualitative directional cue (i.e. steering left or right of fixation) into a quantitative steering command. By combining curl with knowledge of gaze orientation and eye height, the visual system can resolve ambiguities in the flow field and utilize curl as a more precise error signal for locomotor control. These aspects are now included in the new version of the discussion.

      (2B) I also feel that there is a mismatch between what the behavioral task requires and what the controller model does. Subjects are apparently asked to report their heading relative to the world, but the controller model only controls their heading relative to the point that they are fixating. I understand how this is resolved in the model, but I think this type of distinction is buried and will not be apparent to most readers. Again, the reference frames of what is being measured and controlled need to be specified explicitly in all parts of the paper, and the authors need to explain how the system would combine curl-based control with some other measures of (at least initial) heading for world-centered heading to be computed. All of the assumptions need to be clearly specified.

      We thank the reviewer for this point. We have addressed the alignment of the reference frames in our response to Issues 1a and 2a. Once the initial orientation (θ<sub>0</sub>) is established in the world frame, the controller model generates steering adjustments that directly translate into heading predictions within that same world reference frame. By treating the perceptual report as an output of the locomotor controller, we resolve the discrepancy between the steering task and the reported heading.

      (2c) In addition, I found it frustrating that the authors never present raw perceptual data from the observers. Rather, in Figure 2, we see reconstructed trajectories that are perfectly smooth with no indications of noise whatsoever. Since these paths are computed from the perceptual reports, there must be some noise inherent in them. The figures should represent this uncertainty somehow, and it should be explained how these perfectly smooth trajectories are obtained.

      We respectfully disagree with the reviewer’s interpretation regarding data smoothing. The thin lines in Figure 2 represent the mean 3D paths derived directly from the response variable (θ<sub>t</sub>) across trials of identical conditions for each participant (as detailed in the ‘Computation of Perceived Path’ section). No smoothing or filtering has been applied to these plotted trajectories other than computing the mean across trials. We also wish to remind the reviewer that the raw data and analysis code remain publicly accessible for further inspection. Having said that, we include now a supplementary figure showing an example of raw data responses as a function of time. This figure will be a supplemental figure of main Figure 2 (now provisionally included in the Suppl Information).

      Regarding the visual representation: in earlier versions of the manuscript, we included shaded 95% Confidence Intervals (CIs) in Figure 2. However, this addition rendered the plot overly cluttered and obscured the individual trajectories. We therefore chose to present individual participant means (thin lines) alongside group averages (thick lines) to emphasize inter-subject variability. For clarity, the 95% CIs are explicitly displayed in Figure 3, where the data density is more conducive to shaded areas.

      (3) “...the magnitude of retinal curl in the fovea can specify the body trajectory relative to gaze (Matthis et al., 2022)." The main idea put forward by the authors here seems to overlap heavily with this statement that they attribute to Matthis et al. 2022. While I think this paper still adds importantly to the topic, the authors do not discuss how their findings are different from those of Matthis et al. 2022, why they are an important extension, etc. Readers should not have to go read this other paper to have any idea how the present findings are placed in importance relative to the literature.

      We have updated the Discussion to more specifically align our findings with Matthis et al. (2022). We emphasize that our study provides the perceptual validation for their ecological observation that the FOE is often too unstable for reliable use, whereas foveal curl remains a robust signal for path estimation. Our paper provides the causal link, since we manipulate curl in real-time (the ‘cancelled & over cancelled curl’ condition) providing the critical evidence that perceived heading is affected by this signal. The relation with this previous study is made clear in the revised discussion.

      (4) The analysis and treatment of eye movements is extremely weak. The authors discarded trials for which gaze deviated from the fixation point by more than 3 degrees (which is a LOT given that the eye speeds are generally in the neighborhood of 0.5 deg/sec), and they provide basic stats on the distribution of positions. But this largely misses the point: it is not small position errors that are likely to matter, but rather velocity errors. Even a small amount of retinal slip of the target while it is being pursued will cause image motion that is going to alter the optic flow field around the fixation target. So, for example, the retinal curl field may no longer be centered on the fixation target. How do we know that some of the perceptual biases are not influenced by image motion resulting from imperfect tracking of the fixation target? This needs to be analyzed and discussed.

      We thank the reviewer for noting that retinal slip (velocity error) is a more critical metric than positional gaze error. We agree that tracking inaccuracies can introduce translational noise into the flow field. The 3° threshold was established based on the eye tracker’s specifications and the naturalistic setup (1-meter viewing distance without head stabilization). Across all participants, the mean positional error ranged from 1.016° to 1.5° (1 deg is 2.08 cm in our setup). We also calculated retinal slip values, which ranged from 0.12 to 0.27 deg/s (X dimension) and 0.12 to 0.23 deg/s (Y dimension). These values are comparable to natural oculomotor drift (Kowler et al., 1979) and are understandably small given the low velocity of the fixation target. We have added this information about retinal sleep at the beginning of the results section.

      Consequently, it is highly unlikely that retinal slip influenced the results. Furthermore, assuming that tracking error remained consistent across fixation conditions, any present retinal slip cannot explain why the bias followed the retinal curl manipulation as predicted by the controller. We therefore consider retinal slip to be an unlikely confounding factor.

      (5) I found the sections of text comparing the separate and joined fits (starting line 287) to be a bit too rosy. The authors show the separate fits in the main text, and it is not very surprising that these fits are good, given that the model has 30 parameters, and these data are pretty low-dimensional. The authors only show the joined fits in the supplement, and they say that they are almost as good as the separate fits (indeed, they are better in a model comparison sense, but this is 30 parameters vs. 2 parameters). However, when I look at the fits of the joined model in the supplement, I don't find them to be very impressive. In particular, the model grossly misses the data for the straight paths for several subjects (e.g., id5, id6, id8, id10). And fitting the straight paths would presumably be easiest. This implies that the joined model is really missing something and that fitting the curved paths interacts strongly with fitting the data for different fixation target locations on the straight path. I think that the authors should discuss the results a bit more soberly and tone down their conclusions here.

      We thank the reviewer for the opportunity to clarify the logic behind our modeling choices. We acknowledge that the “separate fits” are inherently less informative due to the high number of free parameters relative to the data. Our primary scientific goal was not to achieve perfect descriptive accuracy via 30 parameters, but to test a specific functional hypothesis through the “joint fit.”

      The Logic of the Joint Fit:

      We agree with the reviewer that the joint fit misses some paths in some conditions. Of course, the joint fit reflects a significant compromise. The “Gain” (the weighting of the curl signal) is likely not a static constant but is dynamically tuned based on task demands, confidence in the visual signal, simulated speed, and so on. By using a single Gain parameter, we intentionally ignore this contextual variability to see how much of the behavior can be explained by a “minimalist” controller. In this sense, the 2-parameter joint model is a deliberate attempt to test this limit. By forcing a single Gain parameter to account for all conditions across both straight and curved paths within one flow manipulation (e.g. unaltered flow) we are asking if a single, fixed linear relationship between retinal curl and steering effort/gain can explain the results. We view the joint fit not as a “perfect” model, but as a stronger test of the curl-based control theory. The fact that a 2-parameter model can capture the direction and scale of biases across such a diverse set of conditions (straight/curved paths, five fixation eccentricities) suggests that retinal curl is a robust signal. Upon closer analysis, these discrepancies between the joint model and the data are most pronounced in the over-cancelled condition which is the one when sensory evidence becomes more ecologically inconsistent with the extra-retinal information (gaze direction). While the joint fit successfully demonstrates that a single parameter can capture the general functional role of curl, it fails to account for the complex sensory re-weighting that occurs in ecologically inconsistent conditions (like ‘over-cancelled’ flow). We have updated the manuscript to discuss these limitations in the “fitting the controller” section, framing the model as a parsimonious first-order approximation rather than a complete description of human heading perception based on a minimal set of parameters.

      (6) The section of the paper on neural simulations (starting line 387) has a few weaknesses. First, why are only straight paths simulated here? This does not seem to provide a very rigorous test of the model. Second, it is awkward that the simulation results are presented in units of pixels, rather than degrees. Third, the authors seem to downplay the fact that the neural estimates of heading seem to oscillate rather wildly (over a range of hundreds of pixels, whatever that means, see especially Figure S16). It was far from clear to me how an estimate of heading with these large oscillations is useful. It would seem to require that heading estimates are integrated over substantial lengths of time to be reliable. It was therefore unclear how the model produces such smooth paths from these oscillating estimates.

      We acknowledge that the presentation of the neural model requires more clarity regarding its objectives and its relationship to the behavioral data.

      We first wish to clarify the intended scope of the neural ring-attractor model. Our primary goal was not to provide a comprehensive account of behavioral performance across all conditions (which is the role of the controller model), but rather to demonstrate a biologically plausible mechanism that explains the emergence of the “Opposite-to-Gaze” bias. While the controller demonstrates that the bias follows a specific control law, the neural model shows how such a law can emerge from known primate neurophysiology, specifically, spiral-tuned MSTd neurons, gaze-contingent inhibition, and an egocentric “straight-ahead” prior.

      Why Straight Paths are Sufficient for this Objective. The reviewer asks why only straight paths were simulated. In our study, the straight-path condition with eccentric gaze is the purest test of the bias mechanism. Simulating the straight paths allowed us to isolate the interaction between foveal inhibition and the straight-ahead prior without the confounding variable of path-curvature flow. Given the complexity of the neural network’s parameter space, we focused on these conditions to provide a clear neuro-plausible explanation. We have added text when introducing the model (Neural simulations in the Results section) to make clear why we model straight paths only.

      Units: Pixels vs. Degrees. We acknowledge that the use of “pixels” in the plots of internal neural dynamics may appear awkward. The neural network operates on input stimuli that are defined by the pixel resolution of the videos used in the simulations, we used pixels as the native coordinate system to describe the movement of activity peaks within the network’s internal “map.” We have decided to keep the pixel units in these figures.

      Behavioral Output (Meters): Importantly, the final heading estimates produced by the network are not left in pixels. We use a pinhole camera model to reconstruct the 3D trajectories from the neural activity. These results are expressed in meters, allowing for a direct comparison with the human behavioral data.

      Addressing Wild Oscillations and Smooth Paths. The oscillations observed in the instantaneous heading estimates reflect the stochastic nature of the population peak when tracking high-frequency sensory inputs. In our model, the synaptic time constant (τ) was kept relatively small to ensure a fast, low-latency response to changes in self-motion. While increasing τ would have produced smoother internal dynamics, it would also have introduced delays into the control loop. Instead, we chose to maintain this high sensory responsiveness and applied a temporal moving average later to the network’s decoding to reconstruct the 3D trajectories. This is explicitly stated in the section “Heading Estimation and 3D path reconstruction” in the new appendix 2.

      In addition, the neural activity over time is shown in two ways: the heatmap shows the neuron with preferred heading (one can see more oscillations, specially when the fixation point is closer to the centre (eccentricities -2 and 2), due to larger competition between the sensory evidence and the straight-ahead prior. The other way is the decoded heading. In the ring-attractor model, the decoded heading (φ̂) is not determined by a single neuron but is calculated using a population vector average (equation 19). By summing across the entire population, the decoder effectively integrates sensory evidence from many neurons simultaneously. One can appreciate (see e.g. Fig. 5B) that averaged decoding, leads to a smoother resulting estimate (the white dashed line, whose visibility had been improved in the revised version). Behavioral work by Burr and Santoro (2001) suggests that global motion signals (divergence and rotation in optic flow) are integrated over much longer timescales—roughly 1000ms to 3000ms—compared to local motion units (~200 ms).

      In the previous manuscript, we discussed this aspect in lines 424-426. In the new version, we have added text in the Heading estimation and 3D path reconstruction section (now in appendix 2) stating that we smoothed the decoded signal in agreement with this psychophysical evidence before applying the camera model.

      See also our comment on temporal integration in the responses to reviewer #3

      Reviewer #2 (Recommendations for the authors):

      (7) Line 51: "...a functional role of rotational flow components has been largely neglected in both theoretical and experimental work on heading perception." I feel like this statement is too strong and that the authors try too hard to "sell" their findings by underrepresenting previous work. There are numerous studies (many not cited), both behavioral and electrophysiological, that have examined how heading perception depends on pursuit eye movements, either physical movements or visually simulated ones. These studies directly involve rotational flow components, and several of them have concluded that rotational flow components contribute to estimating heading in the presence of eye movements (just one example is Grigo and Lappe 1999). Because these studies generally involved horizontal pursuit of a target on the horizon, rather than tracking a point in the ground plane (like the authors' work), these studies generally did not involve flow fields with retinal curl around the fixation point. But I consider these older studies just a special case of the more general geometry, and they still involve rotational flow components. Moreover, various previous studies have used stimuli for which there was no FOE present in the visible display (either due to simulated rotation or masking out the FOE), and the authors do not seem to give credit to these works either. In addition, several studies have implicated a role of extraretinal signals in perceiving heading during eye movements, so retinal curl cannot explain everything. Rather than overemphasizing the limitations of previous work, the authors would be better served to explain how their findings extend and generalize from these previous studies.

      We thank the reviewer for pointing out this oversight; it was not our intention to overlook previous work. While our original version cited studies considering rotation-related cue, we have now substantially revised the introduction to include previous work and better acknowledge the informative role of rotation. Our central aim remains to distinguish between models that compensate for rotation to recover a heading vector and our proposal that the visual system exploits retinal curl directly as a primary, functional signal for locomotor control.

      We have now updated the Introduction and Discussion to better situate our work within the context of studies (including Grigo & Lappe, 1999 and some additional ones we have included in the new version) that have investigated the informative role of rotational flow. We now clarify that our study extends these findings by investigating the non-uniform rotational patterns (curl) that emerge during ground-plane fixation, representing a more general and biologically ubiquitous case of locomotor control, while acknowledging previous studies that have also considered the potential role of curl generated by gaze fixation.

      (8) Figure 3: I did not understand why there are two purple and two blue curves in the graphs of the middle column. And the caption does not explain this.

      This a very good observation. This was explained in lines 268-273 (previous version). When the gaze is straight-ahead (same direction as heading), there is no curl. However, we introduced positive or negative curl in the altered conditions. These purple and blue lines refer to these trials and show that the bias re-appears in the expected direction when curl is (unexpectedly) added. We think this adds additional evidence to the curl contributing to heading. Even though this was extensively explained we have added text in the caption of figure 3.

      (9) Line 331: What makes the authors think that retinal curl is computed in area MSTd? They should cite studies to support this idea if it has been shown in physiology.

      Evidence was cited in the introduction (Graziano et al 1994) of sensitivity to spiral motion in addition to neuro-computational models that implement this activity also cited (e.g. work of Leyton et al.)

      (10) The neural network model for computing heading from curl requires a "gaze-centered inhibitory drive" that inhibits activity around where the eyes are looking. This is probably a biologically plausible thing, but is there any evidence to support the idea that this signal exists in the parts of the brain where the authors believe these computations to be happening? They simply posit the existence of this gaze-centered inhibition as though it is common knowledge, but they provide no citations nor discuss any previous evidence for its existence.

      While neurophysiological evidence primarily describes this as gain-field modulation, this process frequently involves localized suppression of neural activity to facilitate coordinate transformations. In parietal areas such as LIP and 7a, eye-position signals do not just enhance responses but can also suppress them, effectively shifting the 'center of gravity' of a population response (Read et al 1997; Born et al 2005, cited in the discussion in the revised section re-evaluating the Focus of Expansion). In the context of our ring-attractor model, this functional modulation is most parsimoniously implemented as a gaze-centered inhibitory drive.

      (11) Lines 482-483: Why should perceptual biases related to retinal curl take seconds to show up?? The curl information itself must be present very quickly, perhaps requiring just a few video frames. So what does this imply about mechanisms? The authors throw out this assertion, but it is left hanging without any further analysis or support.

      The time course reflects the integration requirements of complex motion processing. While local flow is processed rapidly, global patterns like retinal curl require longer temporal windows to reach a stable estimate (Burr et al 2001). In our study, this integration is functionally necessary to filter the higher-frequency 'wobble' induced by gait-cycle oscillations. We now discuss the temporal integration aspects under a new heading in the discussion.

      (12) Line 508: "This suggests that the "bias" observed in our perceived headings may reflect the operation of a control law optimized for action rather than a failure of a perceptual system designed for passive estimation." The authors make this statement to justify why perceptual biases are present with unaltered curl. But I don't fully understand the logic. Are they saying that it is not possible to have a set of computations that can do both things accurately? Is it possible to show this theoretically? Moreover, if it is not possible to rule out other possible sources of the biases, such as those described above (reference frame of judgments, eye movements, etc), then is it necessary to invoke this logic?

      Our logic is that the observed 'bias' is not a representational failure, but a functional byproduct of a control law optimized for active steering. In a closed-loop system, the objective is to null the error signal (retinal curl) to maintain a stable path. When observers are asked to make an open-loop heading report, they likely utilize this same control signal, which manifests as a systematic bias toward the 'null' point of the controller as a result of a sustained fixation in discrepancy with the simulated translation/heading.

      We do not suggest that accurate perception and control are theoretically incompatible; rather, we suggest that perception and action rely in the same underlying information (e.g. work of Brenner & Smeets). While other factors, such as coordinate transformations between reference frames, certainly can contribute to the reporting process, our interpretation provides a parsimonious link between the psychophysical data and the underlying steering mechanism. By framing the bias as a consequence of a 'nulling' strategy, we explain not just the existence of the error, but its specific direction and magnitude relative to the fixated target.

      (13) Line 546: "...MSTd would simultaneously code curvature for trajectory estimation and heading across the neural population, with curvature encoded through the spirality of the most active cell and heading through the visuotopic location of its receptive field center." The latter part of this argument seems to imply a relationship between the heading preferences of MSTd neurons and the locations of their receptive fields. I am not aware of any evidence for such a relationship, so the authors should indicate whether this is based on some experimental data or just a speculation.

      We thank the reviewer for this observation. The proposal that heading is signaled by the visuotopic location of active MSTd populations is a core architectural feature of our model and is supported by several lines of evidence.In the Layton and Browning (2014) framework, MSTd is modeled as a visuotopic map of functional 'hypercolumns'. Each hypercolumn contains neurons tuned to a continuum of spiral patterns, but all neurons in a given hypercolumn share a receptive field center at a specific location in visual space. Consequently, the visuotopic location ($x, y$ coordinates) of the maximally active hypercolumn represents the center of motion (heading), while the spirality (the tuning dimension within that hypercolumn) represents path curvature. We have clarified this in the revised discussion (re-evaluating the FoE) to emphasize that this dual-coding scheme arises from the simultaneous representation of 'where' (population map location) and 'what' (spiral tuning) in MSTd.

      Reviewer #3 (Public review):

      The primary limitation of the paper is that it avoids discussion of some of the inevitable complexities of heading perception. The main issue is what exactly is meant by heading. Different behaviors evolve over different timescales. The geometry of retinal motion defines instantaneous heading, which varies widely through the gait cycle. Time-varying information like this is known to be important in the momentary control of balance. Heading can also be thought of as steering the body toward a distant goal, which evolves over longer timescales. The current manuscript appears to be concerned with heading information integrated over a few seconds and seems to provide evidence that heading is indeed integrated over the gait cycle. The issue of the time scale of the computation is touched on, but it is not related to how it might be used in normal walking or what situations it might apply to. Steering toward a distant goal during walking is not a very difficult problem and may not require evaluation of retinal motion, but control of balance is more challenging and may depend critically on curl. Consequently, the timescale of the computation needs to be considered in order to understand what is meant by heading.

      We thank Reviewer #3 the comments regarding the definition of heading at different time scales, the role of the gait cycle, and the temporal integration of the curl signal. These comments have helped us refine the manuscript’s core arguments.

      We agree that “heading” must be precisely defined within the context of the differing temporal demands of balance and steering. While instantaneous heading provides the high-frequency feedback necessary for momentary postural adjustments and balance, our study is concerned with heading as a gaze-relative signal used for the continuous control of a locomotor trajectory. As such, we have revised the manuscript to specify that the perceived heading measured in our task reflects a signal integrated over the gait cycle to filter out the oscillatory noise induced by head bob and sway (mainly in the Discussion section).

      The reviewer correctly notes that gait-induced head bob and sway produce high-frequency oscillations in the curl signal, yet our behavioral results show smooth, slowly evolving biases. The visual system does not react to “instantaneous” curl, which would lead to jittery, unstable heading estimates. Instead, it integrates flow over a timescale roughly commensurate with a full gait cycle (~500–1000ms). This implies a significant temporal integration process. This temporal integration is consistent with evidence (Burr and Santoro,2001, Vis Res) indicating that optic flow signals (radial and rotational components) are integrated over windows of approximately up to 3 seconds to ensure perceptual stability. Neurally, this likely involves the projection from area MSTd to the Ventral Intraparietal area (VIP), a pathway where fast, eye-centered sensory inputs are transformed into stable, body-centered representations suitable for guiding long-term steering behavior (Chen et al. 2011, JNeurosci.). By grounding our definition of heading in these specific temporal and neural constraints, we tried to clarify how the visual system exploits retinal curl for goal-directed action in natural, dynamic environments and relate our findings to recent studies addressing the role of retinal motion on balance (Powell et al. 2026 Bioarx).

      In our implementation, we explicitly address the high-frequency noise introduced by gait dynamics by smoothing the retinal curl signals computed from the stimulus videos before they are fed into the controller. This temporal filtering allows the fit of the controller’s prediction to the response data while remaining robust to the rapid fluctuations of head bob and sway. In contrast, the neural ring-attractor model would not require an external smoothing step; instead, the integration is an emergent property of the system’s architecture that can be controlled with different parameters, as commented above in a response to Reviewer #2. The dynamics of the synaptic weights and the characteristic “leak” in the population activity naturally implement a leaky integration of sensory evidence, ensuring that the decoded heading reflects a sustained estimate rather than an instantaneous response to visual noise.

      We also agree that we avoided discussing some complexities of the heading perception. In the new version, we also include and integrate the distinction between instant heading and future path in different parts of the ms (introduction) and mainly discussion (temporal integration and steering section) which have been revised substantially.

      Reviewer #3 (Recommendations for the authors):

      There are a number of points that require clarification.

      (1) Head bob and sway were included in the stimulus and need to be addressed in both the analysis of the data and the interpretation. The curl signal in the stimulus varied over time, commensurate with normal gait. However, the results don't reflect the same level of variability that would be produced from curl over the gait cycle. This means that the information must be integrated over some longer timescale. It is not clear from the data analysis what this integration is. Is there an implicit integration with the manipulation of the steering wheel? If subjects indeed appear to be able to use curl to evaluate heading over timescales of seconds, this needs to be explicitly addressed, as it is a novel result. This would require parts of the discussion to be changed/expanded to maintain consistency. For example, line 482 talks about the buildup of biases over time.

      As commented above in the public response, the curl estimated from the optic flow algorithm was smoothed before being input into the controller (path fitting and predictions). The smoothing was only applied to the curl signal, not to the participants responses. Also, as mentioned before, the time course of the bias is consistent with integration times of optic flow reported in the literature. All these aspects are now explicitly included in the new display and conditions section (Flow manipulation conditions).

      (2) Since the experiment included curl variability resulting from the gait cycle, some discussion is needed about the role of retinal motion in the control of balance and posture. There is a large literature about the role of flow in controlling gait and momentary adjustments of the body while walking. Additionally, it should be noted that in the task, head bob and sway from 1 prerecorded individual was shown to all subjects. It is known that gait varies significantly between different individuals, and it should be acknowledged that this may lead to differences at the individual subject level for perceiving heading.

      We have included a point in the discussion addressing the different time scales for different use of optic flow signals (postural control vs locomotion).

      We agree with the reviewer that utilizing a single gait profile for all participants may introduce individual differences in perceived heading, as the simulated head motion might not perfectly match each participant’s unique biological gait signature. However, we prioritized stimulus consistency over idiosyncratic accuracy. By ensuring that every participant viewed the exact same motion profile, we could be certain that the systematic 'opposite-gaze' biases observed across the population were driven by our experimental manipulations of gaze and retinal curl, rather than being confounded by variability in head-motion kinematics. We have added an acknowledgement of this point at the first paragraph of the displays and conditions section in the Methods.

      (3) More information is required about the use of the rotating wheel for the measurement of heading. How easy was it to use? What about time delay, and how does this deal with the bob and sway? Does the wheel impose a de facto integration on the perceptual measurement?

      The rotary encoder provided an intuitive steering-wheel interface that participants found easy to operate. To ensure minimal latency (1–5 ms), the device was interfaced via an Arduino Uno and sampled by a dedicated background Python thread, isolated from the visual rendering loop. We have incorporated these technical details into the Methods (Procedure) section.

      (4) Restructuring the description of the models It remains unclear why the dynamics of the neural network are a necessary inclusion in this paper. It seems interesting, but there is no comparison to actual neural data or other related work. Instead, this appears to be a description of what the network is doing, which is already defined by the equations. This needs to be clarified for its exact interpretation with respect to real neural data, and its importance here for understanding the biases that emerge in heading judgments. The paper would flow better if this section were included as supplementary material or omitted from the paper entirely, as it seems to detract from the other points. If this is a description of why the biases are seen in the controller, then the supplementary material is a good place for it.

      We thank the reviewer for this suggestion. We clarify that the neural model is not intended to simulate specific empirical neural data, but rather to provide a biologically plausible implementation of the controller. This allows us to demonstrate how the observed biases emerge from the dynamics of standard cortical architectures, such as ring attractors. This is now mentioned when introducing the neural simulation results.

      Following the reviewer's suggestion, we have moved the neural model equations to Appendix 2 while retaining the simulation results in the main text (Results). We believe it is essential to present not just the abstract controller, but also its functional implementation, as this provides a mechanistic bridge between retinal signals and locomotor behavior.

      (5) For the modeling approaches, the math would be more appropriate for supplementary materials.

      To ensure a better flow of the paper, we have moved the neural model equations to Appendix 2, while Appendix 1 now details the relationship between measured curl and other variables (speed, yaw, pitch, etc.). We have retained the controller model in the Methods section, consistent with the eLife layout where Methods follows the Discussion.

      (6) How do the models and data analysis deal with the influence of the gait cycle in the input? Do they integrate the information over that timescale? If so, the integration time needs to be specified.

      As commented above, to mitigate gait-cycle fluctuations, we smoothed the computed curl signal before inputting it into the controller and applied a similar smoothing process to the neural model’s readout. Using a LOESS filter, the effective integration window was 2.4 seconds. Close to the integration time reported in Burr et al. 2001. These parameters have now been explicitly specified in the Methods section. For the empirical data, we just utilized trial-averaging.

      (7) What does biologically plausible mean in terms of the neural network model? Especially when control wasn't explicitly a variable in the measured behavior of the subjects.

      By biologically plausible, we mean that our model is constrained by neural architectures documented in the primate brain—specifically ring-attractor dynamics, population coding, and gaze-centered gain-fields. Crucially, the network utilizes recurrent connectivity with a 'Mexican-hat' profile (local excitation combined with lateral inhibition). This is a standard and widely accepted motif in computational neuroscience, representing the consensus on how cortical circuits maintain a stable "bump" of activity to represent spatial variables. Rather than introducing ad-hoc mechanisms, we demonstrate that the observed behavioral biases emerge naturally from these established neural components when they are tasked with maintaining locomotor stability. Since we think this aspect was already emphasized, we haven’t added any additional detail.

      (8) There should be more extensive acknowledgement of the body of literature that has challenged the use of the focus of expansion. That section should also include references to work that has investigated extraretinal signals, as they may also be important.

      We have expanded the Introduction and Discussion to more thoroughly acknowledge research challenging FOE-based models and the critical role of extraretinal signals. These updates, which also align with our response to Reviewer #2, provide a more comprehensive context for our model within the existing body of heading and self-motion literature.

      Minor points:

      (1) Line 69 - In self-generated motion, spiral patterns are almost always centered on the fovea, but many physiological experiments present spirals in the peripheral retina. This is incompatible with the motion generated during self-motion. Therefore, clarify whether the type of spiral motion Graziano investigated was centered on the fovea.

      In the experiments conducted by Graziano et al. (1994), spiral stimuli were centered on the receptive field (RF) of the individual neuron being recorded to accurately characterize its tuning. While the reviewer correctly notes that spiral centers often align with the fovea during active steering (due to fixation on a goal), MSTd neurons possess large RFs that provide a comprehensive 'template' system across the visual field. This population-level representation allows the brain to recover trajectory information even when the focus of motion shifts relative to the fovea—for example, during pursuit eye movements or when fixating on landmarks off the direct path of travel. To keep this part of the text brief, we haven’t add more details concerning this study.

      (2) Line 72 - "Magnitude" instead of "amount".

      This has been changed.

      (3) Line 115 - State explicitly whether the scale of the visual stimulus was matched to the scale of the actual natural images shown in VR.

      We have updated the Methods (Displays and conditions) section to explicitly state that the visual scale was veridical. The virtual camera’s parameters were calibrated such that its field of view (91°) matched the physical dimensions of the projection screen (2.03 m × 1.16 m) at the 1.0 m viewing distance. This ensures that the angular size of the objects and motion gradients in the stimulus were 1:1 with the scale of the simulated natural environment.

      (4) Line 144 - While Farneback is a good dense flow estimation algorithm, it is noisy and may impose biases/variability in the calculation of curl. This should be acknowledged.

      We acknowledge that the Farneback algorithm can introduce variability in curl estimation. To mitigate this, we utilized 10 independent renderings of each experimental trial to compute the flow fields. Although this methodology was reflected in the data previously uploaded to our OSF repository, we have now explicitly added this detail to the manuscript (Flow (curl) manipulation conditions). The computed curl used for the modeling was derived from the aggregate of these different runs, ensuring a robust and stable signal that accounts for potential algorithmic noise.

      (5) Line 206 - "Join fits". Is this a technical term? It sounds awkward. Would "Joint fits" make more sense?

      The referee is right. We have corrected this.

      (6) Line 280 - "Consistent with" (typo).

      This has been corrected.

      (7) Line 286 - Typo in title.

      Also corrected to Fitting the controller.

      (8) Lines 482-493 - There should be more discussion on the time course of integrating the stimulus, and the relationship/generalizability to more natural stimuli.

      This part of the discussion (related to integration time) has been changed considerably to include discussion of postural control in addition to locomotion.

      (9) Lines 516-517 - It is mentioned that retinal flow dynamics override the visual direction cue. This may not be generally true, as the reweighting of the cues might depend on things like task demands or actual stimulus context. In the present experiment, the subject only has access to a large moving textured ground plane, and the body is stationary.

      We agree with the reviewer that cue reweighting is highly context-dependent. However, as noted in the original manuscript (Lines 516-517), we specifically stated that retinal flow dynamics 'can' override the visual direction cue, rather than asserting a universal rule. This phrasing was intentional to acknowledge that while flow is a potent signal—especially in the presence of a large, textured ground plane as used in our paradigm—the relative weighting of these cues remains contingent on the specific sensory and task conditions. We believe this remains a fair and cautious interpretation of our findings.

      (10) Lines 524-256 - It is unclear why Matthis et. al. 2022 is cited for this point. Some of the steering literature, like Wilkie Wann & Allison 2006 or Lappi & Mole 2018 (and some of their other work), would be more relevant for the definition of a control law under these circumstances.

      We agree and this part has been changed substantially.

      (11) Lines 528-532 - Warren et. al. 2001 should be cited in this section because their results were interpreted in terms of focus of expansion, but may result from the curl signal (Powell et. al., 2026. The Role of Retinal Flow in Walking. bioRxiv, 2026-02.).

      We agree with this suggestion and the citation has been added.

      (12) Lines 544-546 - Layton and Browning are focused more on path perception from the implemented spiral tuned cells. Because of this, it wouldn't be appropriate to say it is a shift away from FOE-based heading based on this citation alone. There are more models/psychophysical results that would strengthen this claim.

      We agree with the reviewer that Layton and Browning focus specifically on path perception. Our original intention in citing this work was to emphasize the neurophysiological continuum from radial to circular motion (spiral tuning) rather than discrete expansion/rotation channels. However, we have revised this section to clarify that the reliance on retinal curl represents a mechanism for determining the future path (locomotor trajectory) rather than merely instantaneous heading. This distinction acknowledges that while heading is a momentary vector, the integration of curl signals allows the system to anticipate and control the intended path over time—a framing that better aligns with both the cited literature and our proposed controller model. As commented above this is now extensively discussed in the revised version.

      (13) Figures - There are minor visibility issues for some of the figures. In Figure 2, the thick line is unreadable, and in Figure 1, the x-axis labels are crowded.

      The axis in Fig 1 has been modified to avoid crowdedness. In Fig. 2, the thick (average line) has been modified. We hope they are more visible now.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Poh and colleagues investigate dopamine signaling in the nucleus accumbens (ventromedial striatum) in rats engaged in several forms of Go/No Go tasks, which differed in reward controllability (self-initiated reward seeking or cue-evoked/quasi-pavlovian), and in the specific timing of the action-reward contingencies. They analyze dopamine recordings made with fast scan cyclic voltammetry, and find that dopamine signals vary most consistently to cues that signal a required action (Go cues) vs cues signaling action withholding (No Go cues). Through various analyses, they report that dopamine signals align most clearly with action initiation and with the approach to the reward-delivery location. Collectively, these data support aspects of a variety of frameworks related to accumbens dopamine signaling in movement, action vigor, approach, etc.

      Strengths:

      These studies use several task variants that consolidate a few different components of dopamine signal functions and allow for a broad comparison of many psychological and behavioral aspects. The behavioral analysis is detailed. These results touch on many previous findings, largely showing consistent results with past studies.

      Weaknesses:

      The paper could heavily benefit from some revision to 1) increase clarity of the figures, the methods, and the analysis. 2) The inclusion of many tasks is a strength, but also somewhat overshadows specific points in the data, which could be improved with some revision/reworking. 3) Some conclusions are not fully justified. As shown, support for the conclusion "dopamine reflects action initiation but not controllability or effort" is lacking without more analyses and additional context. 4) Further, the notion that the dopamine signals reported here reflect spatial information could be justified more strongly.

      We thank the reviewer for their detailed evaluation and constructive feedback. We have made substantial revisions to address each concern raised:

      (1) Clarity of figures, methods and analyses

      We have revised the organization of the panels in Figure 1 for clarity.

      We have revised Figure 2 and its caption: we have labelled all comparisons depicted in the figure, and now included a line, “Subject-wise comparisons of dopamine data were made for all alignments”, to clearly show that all statistical tests shown in Figure 2c-e were performed between subjects.

      In the caption of Figure 3, we have now added, “... and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” to clearly show that statistical tests were performed on the trial-wise level for Figure 3c.

      We have now added a table to the Methods section (Table 1), detailing the sample size in each task variant and the number of trials within each No-go classification.

      We have added more information in the Methods section: we now include Videos to show the classified No-go behaviors and other trial types (Go and Free); and provide a schematic of the DLC workflow in the supplementary materials (Supplementary Figure S9).

      (2) Strengthening specific points in the data

      To improve clarity of our main findings, we have revised the layout of the Results section such as including more descriptive headers:

      “Behavioral performance in Go/No-go (“short task”) was unaffected by controllability of reward pursuit”

      “Behavioral performance in Go/No-go/Free (“long task”) was unaffected by controllability of reward pursuit”

      “Motivation to approach the reward magazine was similar between Go and No-go trials”

      “VMS dopamine release encodes reward-related action initiation”

      “VMS dopamine release does not only reflect reward-related action initiation”

      “Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards”

      “Motivational state reflected by No-go behavioral strategy correlates with dopamine signal size during reward approach”

      (3) (4) Conclusions drawn from data

      We have carefully revised our conclusions to accurately reflect our experimental design and analyses, emphasising our core finding that VMS dopamine was consistently increased in Go versus No-go trials, throughout manipulation of the type of trial start (self- and cue-initiated) and effort manipulation (short and long task variants).

      We thank the reviewer for the feedback regarding our VMS dopamine signals reflecting spatial proximity to reward. We performed additional analyses, which we present in Supplementary Figure S10, S11 and S12, to support our interpretation that VMS dopamine encodes spatial proximity to reward.

      We appreciate the reviewer comment relating to the statement, “Dopamine reflects action initiation but not controllability or effort". We have revised the wording of our conclusion to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). This subheading is now changed in the Discussion to “Reward-related dopamine depends on action initiation irrespective of controllability and effort”. The additional analyses that we have performed are shown in Author response image 1entitled: “Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long).

      Additional details on subjects used in each study, analysis details on trialwise vs subjects-wise data, and other context would be helpful for improving the paper.

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures in the original submission of this manuscript.

      To improve the paper, we now include a table in Methods Section 6: Statistical Analysis (Table 1), detailing the sample size in each task variant (with FSCV recordings) and the number of trials within each No-go classification.

      To give more context, we made Author response image 1 to illustrate the number of subjects in each task variant (and their overlap):

      Author response image 1.

      Number of subjects included in each Go/No-go task variant (total n = 21). Values (n) depict overlap of each subject between task variants. One animal was recorded in self-initiated Go/No-go and Cue-initiated Go/No-go/Free (dotted line with arrowheads).

      Reviewer #2 (Public review):

      Here, the authors record dopamine release using fast-scan cyclic voltammetry in the nucleus accumbens/ ventromedial striatum (VMS) while rats perform variants of a Go/No Go task. Two versions are self-paced, in that the rat can initiate a trial by nosepoking at the odor port at any time once the ITI has elapsed, whereas the other two require the rat to wait for a cue-light before responding. Two "long" variants also require either more lever-presses on Go trials, or a longer nosepoke time for No Go trials, and also incorporate "free" trials in which the rat is rewarded for just heading straight to the food tray. The authors find that dopamine levels increase more during the response requirement for Go than No Go trials, indicating a role for invigorating to-be-rewarded actions. Dopamine levels also steadily increased as rats approached the site of reward delivery, and the authors demonstrate quite elegantly that this was not due to orientation to the food tray, or time-to-reward, or action initiation, but instead reflects spatial proximity to the rewarded location. Contrary to previous reports, the authors did not discern any differences in dopamine dynamics depending on whether the trials were cue- or self-paced, and dopamine release did not scale with effort requirements.

      The manuscript is well-written, and the authors use figures to great effect to explain what could otherwise be a hard-to-parse set of data. The authors make good use of the richness of their behavioral data to justify or negate potential conclusions. I have the following comments.

      Re: The lack of relationship between effort to acquire reward in the current study and the magnitude of dopamine release, 1) can the authors unpack this a bit more? 2) Why the difference between the Walton and Bouret studies? Were the shifts in effort requirements comparable across the behavioral tasks? 3) What else could be different between the methodologies?

      We thank the reviewer for the feedback and have responded to each of the three questions below (see points 1-3).

      Firstly, we tried to improve the clarity of our research aims. Our primary comparison throughout the manuscript is between Go versus No-go within each task variant. We ask whether the Go/No-go difference in dopamine signaling persists across different response demands. Thus, testing effort was not central to this main question, but rather a feature of the task that did not affect the Go-No-go dopamine difference.

      (1) Consistent with this aim, we show that VMS dopamine differs between Go and No-go actions persistently across all task variants despite differences in response requirements (action was always accompanied by greater dopamine release compared to action suppression). Our behavioral-training data suggest that the ability to perform short and long tasks differed: rats were first trained to criterion on either a short (∼2 s) or long (∼3 s) Go/No-go variant, with the longer variant requiring substantially more training sessions (short: 18.5 ± 7.6 sessions vs long: 41.1 ± 7.9 sessions; see Author response image 2), indicating behavioral demands were higher for the long-task.

      Author response image 2.

      (2) Regarding the apparent discrepancy with Walton and Bouret (2019), we acknowledge that our original description was imprecise (We wrote: “Previous studies have shown that dopamine signals are influenced by the effort required to obtain rewards”). Our intent was not to suggest a direct contradiction, but rather to emphasize that our findings are consistent with the paper’s broader conclusion that effort encoding by dopamine is limited and highly context-dependent. We have now adjusted the manuscript to better reflect our intent by changing the sentence to, “Previous studies have shown that dopamine signals may be influenced by the effort required to obtain reward but only for particular task conditions (Cousins et al., 1996; Gan et al., 2010; see for reviews, Salamone and Correa, 2024; Walton and Bouret, 2019).”

      For added clarity, these were the main results highlighted in the Walton and Bouret review: Gan et al. (2010) demonstrated that VMS dopamine sensitivity to low-effort costs is prominent early in training (≤ 2 training sessions) and diminishes after extended experience (> 9 sessions). Similarly, Hollon et al. (2014) reported that cue-evoked VMS dopamine primarily tracks reward magnitude with minimal modulation by effort. In line with this literature, our rats were highly trained (≥ 9 sessions until the first recording), and exhibited no difference of average Go minus No-go dopamine between short and long task variants within controllability type (see Author response image 3), supporting the idea that extended training exhibits minimal effort-related modulation of VMS dopamine.

      Author response image 3.

      Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of controllability and effort showed moderate evidence of absence (controllability: BF<sub>10</sub> = 0.324; effort: BF<sub>10</sub> = 0.309). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.101), and the full model including a controllability × effort interaction was the least supported of all models examined (BF<sub>10</sub> = 0.043). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither controllability, effort, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (3) With respect to task comparability and methodological differences, our behavioral paradigm differs in important ways from those highlighted by Walton and Bouret, where effort was often manipulated by training animals to associate cues with different numbers of lever presses within the same session, and typically involved only action initiation. In contrast, our task required both action initiation and action suppression, and changes in response contingencies occurred across separate recording sessions rather than within-session cue-based manipulations. Although these paradigms are not directly comparable, and only had the same dopamine recording technique in common (FSCV), a key takeaway of our results is that regardless of effort differences, VMS dopamine during action initiation is consistently higher than during action suppression.

      I would argue that the cue- vs self-initiated distinction was pretty minor, given that there was a fixed ITI of 5s. How does this task modification compare to those used previously to show that dopamine release corresponds to behavioral controllability? It would help the reader if the authors could spend more time discussing these disparate findings and looking for points of methodological divergence/commonality.

      We agree that clarifying how our manipulation of controllability compares to prior work improves the manuscript, and we have made the necessary adjustments. However, we would first like to correct an incomplete characterization of the task design.

      While the short-task variant used a fixed 5 s inter-trial interval (ITI), the long-task variant employed a variable ITI ranging from 15–25 s. In the long-task variant, the timing of trial onset was less predictable, and we believe this manipulation reduced animals’ ability to precisely estimate when reward pursuit could begin. Under these conditions, whether trials were Self-initiated or Cue-initiated had a substantial impact on animals’ control over the initiation of reward pursuit. That said, we agree that the Self- versus Cue-initiated distinction overall represents a moderate manipulation of controllability compared to those used in studies that focus on controllability.

      A key source of divergence across studies lies in the definition of controllability. We defined controllability as the animals’ ability to choose the time point of beginning the reward pursuit, rather than whether an action was required, and have now added the following sentence in the:

      - Introduction section: “... controllability of reward seeking, defined as the ability to determine when to initiate reward pursuit (Self- vs Cue-initiated trials)...”;

      - Results section: “We defined controllability as the rats’ ability to choose the time point of reward pursuit. In Cue-initiated trials, the time point at which trials could be started was dictated by a cue light, whereas in Self-initiated trials rats were able to choose intrinsically (control) when to attempt a trial start.”;

      - Discussion section

      Importantly, the action requirements for Go, No-go, and Free trials were identical across these trial-start conditions. We found that the degree to which controllability was manipulated in our task was insufficient to modulate the action-specific VMS dopamine signal (Go vs No-go difference), which remained robust across conditions.

      In contrast, controllability has been defined by others as the presence versus absence of an operant action requirement for reward. For example, Goedhoop et al. (2023) directly contrasted operant (lever press required) and Pavlovian (no action required) conditions, removing action execution as a prerequisite for reward. In that context, cues signaling operant control elicited sustained VMS dopamine release, which was interpreted as reflecting anticipation or preparation for executing a learned action. Similarly, Hamid et al. (2021) demonstrated that dopamine “wave” directionality across striatal regions depends on controllability defined by operant versus Pavlovian conditioning.

      Taken together, these comparisons (results from the present study and in the literature) suggest that dopamine sensitivity to controllability may depend on how it is manipulated. We have clarified these methodological distinctions in the revised Introduction, Results and Discussion, and emphasized that more extreme manipulations (such as removing action requirements entirely or increasing uncertainty over trial timing) may be necessary to reveal controllability-dependent changes in VMS dopamine signaling. Alternatively, the apparent discrepancies across studies may primarily reflect differences in the underlying definitions of controllability rather than conflicting results.

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Poh et al. investigated whether dopamine release in the ventral medial striatum integrates information about action selection, controllability of reward pursuit, effort, and reward approach. Rats were implanted with FSCV probes and trained in four Go/No Go task variants:

      (1) trials were self-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (2) trials were cue-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (3) trials were self-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued, and effort was increased,

      (4) trials were cue-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued.

      The authors report that dopamine levels rose during Go trials and slowly rose in No Go trials, but this pattern did not differ across task variants that modified effort and whether trials were cued or initiated. They also report that dopamine levels rose as rats approached the reward location and were greater in rats that bit the noseport while holding during the No Go response.

      Strengths:

      (1) Interesting task and variants within the task paradigm that would allow the authors to isolate specific behavioral metrics.

      (2) The goal of determining precisely what VMS dopamine signals do is highly significant and would be of interest to many researchers.

      Weaknesses:

      (1) This Go/No-Go procedure is different from the traditional tasks, and this leads to several problems with interpreting the results:

      (a) Go/No Go tasks typically require subjects to refrain from doing any action. In this task, a response is still required for the No Go trials (e.g., continue holding the nosepoke). The problem with this modified design is that failure to withhold a response on No Go trials could be because i) rats could not continue holding the response, as holding responses are difficult for rodents, or ii) rats could not suppress the prepotent go response. This makes interpreting the behavior and the dopamine signal in No Go trials very difficult.

      We appreciate the reviewer raising this important methodological consideration. We acknowledge that our Go/No-go task differs from traditional paradigms used in humans and primates (e.g. Raud et al. 2020, 10.1016/j.neuroimage.2020.11658; Eagle, Bari & Robbins 2008, 10.1007/s00213-008-1127-6; Roitman & Loriaux 2013, 10.1152/jn.00350.2013).

      However, our design addresses the specific constraints of studying dynamics in freely moving rodents while maintaining the core feature of Go/No-go tasks: requiring suppression of a prepotent response. Our task accomplishes the primary aim of our study, which is to compare VMS dopamine dynamics during action initiation and action suppression, and below we list the reasons why. Therefore, we do not believe that this difference compromises the validity and interpretation of our results.

      It has been suggested for decades that the two main processes governed by mesolimbic dopamine are reward learning and motivated action, and our study aimed to better understand how VMS dopamine integrates reward-related information and motivated action, rather than studying them in isolation. To do so, we trained rats in a modified Go/No-go task.

      More recent work (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019) demonstrates that VMS dopamine signaling incorporates both action and reward-related information, rather than either of the two alone. Importantly, in freely-moving rodents, examining this relationship requires preventing the approach response that occurs when reward delivery is anticipated (Pavlovian bias, go for rewards). Traditional Go/No-go designs that simply require "doing nothing" would not achieve this control in freely-moving rats, as animals immediately approach the reward magazine as soon as reward is inferred (as seen in our Free trials). Thus, we require a No-go condition, as we and others have defined (Syed et al. 2016), whereby animals have to actively suppress the ‘initiation’ response. Action initiation is defined at the beginning of the Discussion section: “... at two distinct points after trial start: 1) when rats began lever pressing (Go), and 2) when rats walked to the reward magazine, either without action requirement (Free) or after successful trial completion (Go and No-go)”.

      To further strengthen our interpretation that we compare action initiation and suppression, and to facilitate cross-species translation of our results (i.e., rodent to human), we also include Free trials, where reward delivery requires no specific action (which are essentially like “doing nothing” trials in traditional tasks). This addition allowed us to directly compare No-go and Free trials, where animals must actively suppress responding while maintaining task engagement, to a condition where no overt action is required for a reward, respectively. The dramatic difference in VMS dopamine between No-go and Free trials demonstrates that VMS dopamine reflects active action suppression during No-go trials, rather than merely the absence of action requirements. This has now been discussed.

      Finally, we only report correct Go, No-go, and Free trials, which differs from that of human go/no-go studies that focus on the failure of appetitive no-go trials (i.e., inhibiting the pre-potent response). In the present study, the dopamine signals that we interpret are restricted to successful trials only: action initiation (moving the lever press), action suppression (i.e., suppressing the prepotent Go response while maintaining their position in the nose-poke port), or no action (no lever press, not staying in the port). While this design differs from human Go/No-go paradigms, we believe our study of correctly performed Go, No-go, and Free trials are necessary for isolating action-dependent components of dopamine signaling in freely moving rats (action initiation vs action suppression vs action free).

      (b) Most Go/No Go tasks bias or overrepresent Go trials so that the Go response is prepotent, and consequently, successful suppression of the Go response is challenging. 1) I didn't see any information in the manuscript about how often each trial type was presented or 2) how the authors ensured that No Go responses (or lack thereof) were reflecting a suppression of the Go response.

      We appreciate the reviewer's attention to this important methodological consideration. The originally submitted version of the manuscript already addressed both concerns raised.

      Trial type presentation frequencies

      The Methods section describes our trial presentation approach: "On recording days, the trial types were counterbalanced. Within a session, Go left, Go right, and No-go trials were presented with 33% probability each, without replacement. For sessions with Free trials, trials were presented with a 25% chance without replacement."

      This design results in overrepresentation of Go trials overall (66% in the short-task; 50% in the long-task), which establishes the prepotent Go response as intended in standard Go/No-go paradigms. To improve clarity, this detail has now been included in the Methods section.

      Ensuring No-go responses reflect suppression of Go response

      Our paradigm incorporates multiple features that ensure successful No-go performance reflects suppression of the prepotent Go response:

      First, the overrepresentation of Go trials (addressed above) establishes response prepotency. Second, during No-go trials, rats must maintain their snout in the nose-poke port for the duration of the action cue, which creates the requirement to suppress the natural tendency to approach rewards (i.e., Pavlovian bias; Jones et al. 2017, 10.1016/j.bbr.2017.05.044; Guitart-Masip et al. 2014, 10.1007/s00213-013-3313-4; Dayan et al. 2006; 10.1016/j.neunet.2006.03.002). In Go trials, such natural bias does not require suppression as the lever can be approached and pressed during the action-cue period. This conflict between the instrumental No-go requirement and the Pavlovian-instrumental bias toward action makes action suppression particularly challenging (consistent with computational accounts of similar paradigms; Lloyd & Dayan 2023, 10.1371/journal.pcbi.1011569; Jones et al. 2017, Guitart-Masip et al. 2014, Dayan et al. 2006).

      Figure 3 provides behavioral evidence of this challenge: animals frequently left the nose-poke port and developed spontaneous motor strategies (such as biting and digging) to stay in the port, suggesting Pavlovian bias interfering with response suppression for rewards. Importantly, all reported No-go data include only correct trials (i.e., those without lever presses), ensuring that the dopamine signal reflects successful response suppression rather than failed Go attempts.

      (2) The authors observe relatively consistent differences in the DA signal between Go and No Go trials after the action-cue onset. However, the response type was not randomized between trial type, so there is a confound between trial type (Go/No Go) and response (lever/nosepoke). The difference in DA signal may have nothing to do with the cue type, but reflects differences in DA signal elicited by levers vs. nosepokes.

      As stated in the Introduction section and discussed in our rebuttal to point 1a, the focus of our investigation is how VMS dopamine signals differ during action initiation versus action suppression for rewards, as this is a central unanswered question in the dopamine field. More recent work demonstrates that dopamine incorporates not only RPE but also action initiation (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019), and our goal is to further our understanding of action-dependent VMS signals during reward pursuit.

      The reviewer suggests that dopamine differences may reflect differences in lever vs. nosepoke rather than cue type (Go vs No-go). We respectfully suggest this concern reflects a misunderstanding by the reviewer of our experimental question. The cue-action relationship is the experimental manipulation itself. It is not possible to study how dopamine encodes instructed action initiation versus suppression without linking specific cues to specific actions. The suggestion to 'randomize' action type across cue types would eliminate the very phenomenon we are investigating: how dopamine signals differ when cues instruct different action requirements.

      Our experimental design specifically compares reward pursuit with action requirements (Go trials: lever press; No-go trials: sustained hold) to reward pursuit without action requirements (Free trials: direct magazine approach). This design allows us to isolate how action initiation and action suppression influence reward-related dopamine signaling, which can reveal how the timing of action initiation influences RPE-dopamine. And which is the point of the study: to show how actions influence RPE dopamine signaling.

      Supporting this interpretation:

      Firstly, trial types were randomly interleaved, and each auditory cue explicitly instructed a specific behavioral response. Our design directly follows established methods demonstrating that VMS dopamine encodes whether actions are initiated or suppressed following action cues (Syed et al. 2016). That study, like ours, intentionally linked cue identity to a specific action requirement to assess how dopamine reflects instructed behavioral control. Thus, the fact that Go and No-go cues map onto different actions is inherent to the question being addressed, not an unintended confound.

      Second, as discussed in our response to point 1b, the asymmetry between Go and No-go trials is theoretically essential. Go trials align with Pavlovian approach tendencies (action initiation to reward), while No-go trials create conflict with this bias by requiring action suppression despite the cue being associated with a reward. This Pavlovian-instrumental conflict makes suppression particularly challenging (Lloyd & Dayan 2023, PLoS Comput Biol 10.1371/journal.pcbi.1011569) and allows us to examine the role dopamine in overriding prepotent responses.

      Third, the inclusion of Free trials (discussed in point 1a) demonstrates that our findings reflect instructed action control rather than simply motor execution. Free trials require neither lever pressing nor nose poke maintenance, yet show dopamine dynamics distinct from both Go and No-go trials, confirming that dopamine signals encode action requirements beyond motor output per se.

      Finally, we demonstrate that VMS dopamine differs in the same trial type (No-go) and can be classified based on different movement patterns (Biting, Digging, Calm). Importantly, the difference in VMS dopamine only appeared after the action was completed, particularly during reward approach (Figure 3). This data argues against the idea that VMS dopamine is particularly tied to the specific operant manipulanda as suggested by the reviewer, but rather, may reflect an internal motivational state for reward.

      Together, the aim of the present study is not to redefine Go/No-go paradigms for rodents, but to utilize this task structure to investigate action-dependent dopamine signalling for rewards, which cannot be answered without the cue-action mapping that we have used.

      (3) Both Go and No Go trials start with the rat having their nose in the noseport. One cue (Go cue) signals the rat to remove their nose from the noseport and make two lever responses in 5 seconds, whereas the other cue (No Go cue) signals the rat to keep their nose in the noseport for an additional 1.7-1.9 s. The authors state that the time between cue onset and reward delivery was kept the same for all trial types, and Figure 1 suggests this is 2 s, so was reward delivered before rats completed the two lever presses? I would imagine reward was only delivered if rats completed the FR requirement, but again, the descriptions in the text and figures are incongruent.

      The reviewer asks whether reward was delivered before rats completed the two lever presses and notes incongruence between text and figures. We respectfully note that these details were stated in the originally submitted version of the manuscript (see below).

      Reward delivery timing

      The reviewer asks whether reward was delivered before rats completed the two lever presses, which refers to the short-task variant. No - reward was always delivered immediately after the second lever press for all Go trials. This is described in Methods Section 3: Behavioral procedures - Self-initiated task variant. For added clarity, we have now added the term “immediately”: “... food pellet dispensed into the reward-magazine immediately.”

      In the Results section, we report that the average latency to complete two lever presses was 1.8s, which closely matches the 1.7-1.9s nose-poke hold maintenance required for No-go trials in the “short” variant. Thus, the time point of reward delivery was matched between Go and No-go trial types.

      Representation of reward delivery timing in figure and text

      The reviewer's confusion appears to stem from the schematic representation in Figure 1 and the task variant structure. There were two overarching task variants with different trial requirements:

      “Short-task” variants: Go trials required two lever presses (completed on average in 1.8s);

      No-go trials required 1.7-1.9s nosepoke maintenance

      “Long-task” variants: Go trials required a ‘rewarded’ press to occur 2.7-3.2s after cue onset (completed within ~3s); No-go trials required 2.7-3s nosepoke maintenance.

      For simplicity in depicting action-cue onset in Figure 2c, we used grey shading with a speaker icon at approximately 0-2s and 0-3s to represent these two variants. This schematic representation was not intended to indicate the precise reward delivery time, which (as stated in the Methods) occurred only upon successful completion of trial requirements in the short-task variant, or 2s after successful completion of trial requirements in the long-task variant. To improve clarity, we have adjusted the legend of Fig. 2 for more clarity, adding “Shaded gray area depicts approximate duration of action-cue onset for “short” and “long” task variants.”, and included more information under Results: “In the short-task variants, action-cues switched off after trial completion and a reward was delivered immediately” and “n the long-task variants, action-cues switched off after trial completion, or in the case of Free trials after 3s, and reward was delivered 2s later (Figure 1a).”

      (4) The manuscript is difficult to understand because key details are not in the main text or are not mentioned at all. I've outlined several points below:

      (a) The author's description in the manuscript makes it appear as a discrimination task versus a Go/No Go task. I suggest including more details in the main text that clarify what is required at each step in the task. Additionally, providing clarity regarding what task events the voltammetry traces are aligned to would be very useful.

      We respectfully note that the requested details were already present in the originally submitted version of the manuscript (see below). However, we acknowledge that the task design is complex and may benefit from additional clarity in the main text to aid reader comprehension.

      Behavioral task

      The reviewer suggests our task appears more like a discrimination task than a Go/No-go task. We acknowledge that our paradigm differs from traditional Go/No-go tasks used in humans and primates, as discussed in our responses to points 1a and 2. However, we classify this as a Go/No-go task because it shares the defining feature: requiring action initiation (Go) and suppression (No-go). This classification is consistent with established rodent literature examining action initiation versus suppression (Syed et al. 2016).

      Moreover, as discussed in our response to point 1a, we included Free trials specifically to demonstrate that No-go trials require active suppression rather than discrimination alone. The distinct dopamine dynamics across Go, No-go, and Free trials confirm that our task captures action initiation, action suppression, and action-free states, which is an important contrast that we needed to address our research question about action-dependent dopamine signaling.

      The key requirements for each trial type and task variant are described at the beginning of the Results section. A full description of each step required in the task was provided in the Methods Section 3: Behavioral procedures, to avoid repetition in the Results. Specifically:

      Self-initiated Go/No-go task variant (“short”)

      Self-initiated Go/No-go/Free task variant (“long”)

      Cue-initiated Go/No-go and Go/No-go/Free task variant

      To improve clarity, we have added a reference in the Results section to the Methods: Behavioral procedures for additional procedural details.

      Voltammetry trace alignment

      The events to which voltammetry traces are aligned were stated in the legend:

      “... when traces were aligned to action-cue onset…”

      “... aligned to the time when animals departed the nose-poke port…”

      “... we realigned traces to the moment animals arrived at the reward magazine… “

      Figure 2c legend: "c) Traces aligned to action-cue onset"

      Figure 2d-e legend: “d) Traces aligned to nose-poke exit and e) magazine arrival….”

      However, to improve clarity, we have now added “aligned to action-cue onset.. “ to make the trace alignment more immediately apparent when results are first presented, and added: “Dopamine data were aligned to events of interest: action-cue onset, nose-poke exit, and magazine arrival.”

      (b) How many subjects were included in each task variant? The text makes it seem like all rats complete each task variant, but the behavioral data suggest otherwise. Moreover, it appears that some rats did more than one version. Was the order counterbalanced? If not, might this influence the DA signal?

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures, where each task variant description includes the corresponding sample size (“Self-initiated Go/No-go task (‘short’; n =9)”, “Self-initiated Go/No-go/Free task (“long”; n = 5)”, “A total of n = 11 and n = 15 were included in the Cue-initiated Go/No-go and Cue-initiated Go/No-go/Free tasks, respectively.”

      For added clarity, we have also added a table for the separation of No-go trials and the number of subjects that it has come from in the Methods.

      Task variant completion

      Not all animals completed all task variants. As stated in Methods Section 4: Real-time dopamine recordings and analysis, animals had to achieve >60% success rate for each trial type on at least two consecutive training sessions to proceed to recording. Other reasons include electrode degradation before all recordings could be completed (See Author response image 1).

      Training order

      We trained four cohorts of animals. One cohort was trained first in the short-task variant

      (Cue-initiated Go/No-go), and the remaining three cohorts were trained first in Cue-initiated Go/No-go (3s) long-task variant (i.e., without Free trials, behavioral and FSCV data not presented in the manuscript). Training order was not fully counterbalanced due to constraints described above.

      Following additional analyses, our data suggest that training order did not influence our core comparison of Go minus No-go dopamine. To directly address whether training order influenced dopamine signals, we separated animals based on whether they were first trained in the Cue-initiated Go/No-go (2s) or the Cue-initiated Go/No-go/Free (3s). We calculated the average dopamine of each rat during the action-cue period and then calculated the difference between them (Author response image 4). We observed absence of evidence of a difference between the groups.

      Author response image 4.

      Average Go minus No-go dopamine reveals no effect of the initial training variant (2s-first vs 3s-first) or test task version (Go/No-go vs. Go/No-go/Free). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of the initial training variant and test task version both showed moderate evidence of absence (initial training variant: BF<sub>10</sub> = 0.378; test task version: BF<sub>10</sub> = 0.367). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.134), and the full model including an initial training variant × test task version interaction was the least supported of all models examined (BF<sub>10</sub> = 0.069). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither the variant animals were first trained on, the task version administered at test, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (5) There is a major challenge in their design and interpretation of the dopamine signal. Both trial types (Go and No Go) start with the rat having their nose in the noseport. An auditory cue is presented for 2-3 s signaling to the rat to either leave the noseport and make a lever response (Go trial) or to stay in the noseport (No Go trial). The timing of these actions and/or decisions is entirely independent, so it is not clear to me how the authors would ever align these traces to the exact decision point for each trial type. They attempt to do this with the nose-port exit analysis, but exiting the noseport for a Go trial (a rat needs to make 2 lever presses and then get a reward) versus a No Go trial (a rat needs to go retrieve the reward) is very different and not comparable.

      We respectfully disagree with the reviewer’s assertion that our data alignment approach is problematic. Aligning neural activity to specific behavioral epochs that occur at different times and across conditions is a widely used method to investigate the relationship between neural activity and behavior. Just to mention some examples: data collected with fiber photometry (e.g., Tan et al. 2026, doi: 10.1038/s41386-026-02368-4; Hart et al. 2024, doi: 10.1016/j.celrep.2024.113828) and voltammetry (Hamid et al. 2016, doi: 10.1038/nn.4173; Syed et al. 2016, doi:10.1038/nn.4187).

      Alignment method

      We intentionally designed the task so that overall action timing is matched between trial types (as described in our response to point 3), while specific behavioral epochs occur at different times. This allows us to compare dopamine dynamics during comparable behavioral events across Go and No-go trials (e.g., nose-poke exit).

      We align data to three critical behavioral epochs, stated in the Methods Section 4: Real-time dopamine recordings and analysis - FSCV measurement and analysis: action-cue onset, nose-poke exit, and magazine arrival. Each alignment addresses a specific aspect of our research question:

      Action-cue onset alignment captures VMS dopamine dynamics when animals have explicit knowledge of trial type and the required action. This allows us to characterize how dopamine evolves following correct action selection, which is central to our research question about how dopamine differs during successful Go versus No-go action execution, as well as no overt action (Free) in the long-task variant.

      Nose-poke exit alignment captures dopamine dynamics at the moment animals initiate movement. The reviewer suggests that exiting for Go versus No-go trials is "very different and not comparable" because subsequent actions differ (two lever presses vs. direct reward retrieval). However, this is precisely our experimental manipulation: we compare dopamine signals when animals exit the nose-poke port to perform different actions. This comparison is both valid and necessary to address our research question (does dopamine encode action initiation?).

      Magazine arrival alignment captures dopamine dynamics at reward approach. This allows us to differentiate between spatial proximity to reward from other concepts including temporal proximity (how soon is reward) and action requirements.

      We acknowledge that we cannot identify the precise moment of decision formation. In fact, the precise moment of decision formation is irrelevant for our question. However, our aim is to characterize VMS dopamine dynamics during successful action execution for rewards, following cues associated with specific actions.

      (6) The voltammetry analysis did not appear to test the hypotheses the authors outlined in the intro. All comparisons were done within task variants (DA dynamics in Go vs. No Go trials, aligned to different task events), but there were no comparisons across task variants to determine if the DA signal differed in cued vs self-initiated trials.

      Our aim was to investigate whether VMS dopamine signals consistently differed between action initiation and action suppression during reward pursuit. To test this within-variant contrast (Go > No-go), we manipulated how reward pursuit is initiated (self- vs cue-initiated) and the “effort” requirements (short vs long task variants), and our results show that they did not affect the differential between Go and No-go.

      The consistent Go > No-go dopamine that we observed across all task variants, together with the consistent increase during magazine approach, supports our conclusion that VMS dopamine integrates motivated action and reward.

      The reviewer suggests that we should have compared dopamine signals across self- vs cue-initiated task variants. We acknowledge this is an interesting, but entirely different, question and have addressed it in the Discussion section. We note that differences in controllability altered the time course of increased VMS dopamine, presumably by triggering earlier positive RPEs in Cue-initiated tasks as compared to Self-initiated tasks (illumination of the nose-poke light being the earliest predictor of reward). However, since our primary research question relates to the difference in VMS dopamine between action initiation and suppression, our results show that this difference was unaffected in two variants (short and long), strengthening our conclusions about the relationship between action and reward-related dopamine signaling.

      (7) Classification of No Go behaviors was interesting, but was not well integrated with the rest of the paper and was underdeveloped. It also raised more questions for me than answers. For example:

      (a) Was the behavior classification consistent across rats for all No Go trials? If not, did the DA signal change within subjects between biting vs digging vs calm?

      (b) If "biting rats" were not always biting rats on every No Go trial, then is it fair to collapse animals into a single measure (Figure 3C).

      (c) Some of the classification groups only had 2 or fewer rats in them, making any statistical comparison and inference difficult.

      Behavioral classification for each rat was consistent across trials (i.e., 100%, see Author response image 5). Upon reviewing the consistency of classifications within individual animals, we found that “Biting” animals exhibited biting behavior across the majority of their No-go trials. Only one animal (in the Self-initiated Go/No-go "short" variant) showed mixed classifications across trials, occasionally exhibiting digging or calm behavior. For all other animals, the predominant behavioral classification was highly consistent within subjects across sessions. The occasional trials where “biting” animals did not bite were too infrequent to permit meaningful within-animal comparisons. Therefore, we believe collapsing animals by their predominant behavioral phenotype in Figure 3C is appropriate and accurately represents stable individual differences in No-go response strategies.

      As stated in the Methods Section 6: Statistical Analysis - Clustering No-go behaviors and regrouping animals (last-line), we specifically avoided between-subjects statistical comparisons for groups with n≤2, as this would be inappropriate (see Author response image 5), and reported qualitative observations only. These exploratory findings at the individual-trial level suggest behavioral heterogeneity during No-go trials, that others may use for future investigation, but do not form primary conclusions.

      The behavioral classification is integrated with our central findings on VMS dopamine encoding spatial proximity. Our results demonstrate that individual variation in action suppression strategy, in particular Biting behaviors, consistently manipulates the timing of max dopamine release during subsequent reward approach, but not during the action itself (Figure 3b-c). This links our observations of action-dependent dopamine (Figure 2c) with spatial reward approach (Figure 2e). We believe that our findings shed new light into the understanding of how action modulates reward-related dopamine dynamics at the individual level. This has been discussed in Discussion section: Dopamine dynamics are linked to motivated action.

      Author response image 5.

      Rats predominantly stick to a particular strategy to perform No-go trials. Each bar represents an individual animal, and colours represent the % of each classification type.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figures: It would be helpful to have more panel labels on the figures - 1C, for example, labels 4 different dataset panels, similar to a few other cases. This would really help to improve readability, as there are a lot of task types, trial types, behavioral measures, and labels to sift through. The figures overall are very busy, and it is a bit of a challenge to process the tasks and trial comparisons, as well as interpret what the quantification insets mean.

      We thank the reviewer for their suggestion. We have added more panel labels to Figure 1 to improve the ability to follow with the main text. We hope that the reviewer finds it acceptable.

      (2) Figures: A bit more specifically, in Figure 2, the boxplot insets are pretty hard to see, and it's not clear what scale they are on or what data they reflect. Similarly, it's unclear what the horizontal bars reflect in terms of which conditions are being compared. Why are box plots used for some comparisons, and why are some comparisons based on time series bootstrapping, but others are not clear? I would consider broadly reworking this figure and its description for clarity. Figure 3, by comparison, is easier to understand - the quantifications are clearer and labeled.

      We have now labeled the comparisons being depicted by the horizontal bars in Fig 2c. We have also clarified the boxplot analysis in Fig 2d and Fig 2e by adding 'Max dopamine’ labels to the figure, and have made these analysis methods more explicit in the figure legend.

      We used time-series bootstrap analysis to identify when dopamine signals diverged between trial types, which depicts dopamine differences during distinct action requirements. The latency-to-max quantification provides a summary measure to test specific encoding hypotheses (temporal vs. spatial proximity to reward).

      (3) Broadly, more clarity on the FSCV analysis is warranted.

      (a) Targeting: It looks like the dataset contains a mix of medial shell and mostly core accumbens placements. The paper treats VMS as a uniform dopamine region, but it is more standard to separate core and shell (and also other parts of the shell) into subregions. Many of the reported encoding profiles here are known to differ across the accumbens. So, some consideration of this seems appropriate - a minimal signal-behavior comparison for the shell vs core subgroups, for example.

      We thank the reviewer for raising this point. Upon careful re-examination of our histological analysis, we identified an error that occurred when we made the overlay of electrode placements across rostrocaudal planes (to project placements onto a single plane for the sake of simplicity; Fig 2a): We incorrectly assigned some recordings to the nucleus accumbens shell. We corrected this error, which shows that the vast majority of recordings were in the core, with only 2 animals in the shell (black stars). We have adjusted Fig 2a to reflect this, and added an anatomically more complete illustration of electrode placements across rostrocaudal planes as Supplementary Figure S8.

      While we acknowledge reported differences between core and shell dopamine in some contexts, the small number of shell placements precludes meaningful statistical comparison. However, to determine whether average dopamine concentrations differed between core and shell during the action-cue period, we have plotted the values in Author response image 6. The fact that shell data mostly falls centrally into the overall core-data distribution suggests no consistent difference in dopamine release.

      Furthermore, in our experience (and that of colleagues (personal communication)) with appetitive operant tasks, core and shell FSCV dopamine signals do not substantially differ for action-selective encoding and reward approach. Given the sample distribution and our focus on general VMS function in Go/No-go behavior, we believe pooling these regions is appropriate.

      Author response image 6.

      Average dopamine release during the action cue in nucleus accumbens core (circles) and shell (stars) animals showed no distinct separation between regions. Each symbol represents an animal.

      (b) Design: In my understanding of the design, the main distinction between the short and long task variants is a 2-second versus a 3-second required nose poke hold. 3 seconds here is "long" and more "difficult". I'm not sure I agree that a 1-sec distinction really reflects a difference in task difficulty or effort. Can the authors point to a past paper that demonstrates this variation is sufficient to engage a behavioral difference and/or a neural encoding difference? Broadly, some justification of the validity of this manipulation is needed, I think.

      While we lack direct citations for this specific manipulation, we believe that the 2s and 3s hold requirements represent meaningful differences in difficulty based on our extensive rat-behavior experience and behavioral evidence in Author response image 2.

      Importantly, the difficulty of No-go trials does not stem merely from the required time to hold their snouts in the nose-poke port, but from suppressing the motivational/Pavlovian bias to approach reward-associated cues. In No-go trials, subjects must suppress the prepotent tendency to immediately approach reward-related stimuli and instead maintain active suppression of this approach behavior. Even the 2s hold is challenging as animals tend to perform better on Go trials compared to No-go trials (Fig 1b and 1c), demonstrating the inherent difficulty of response suppression even at the shorter duration. The additional 1-second substantially increases this demand, as it represents a 50% increase in hold duration.

      Our training data clearly demonstrate the difficulty in reaching task criterion when increasing the required action (for Go and No-go trials) from 2s to 3s. Across four cohorts of animals trained in Go/No-go task variants, one cohort that was trained first in the 2s task variant, and the remaining three cohorts were trained in the 3s task variant of Go vs No-go. Animals required 41.1 ± 7.9 (n = 27) sessions to learn the 3s hold (approximately 8 weeks), versus

      18.5 ± 7.6 sessions for the 2s hold (approximately 4 weeks; mean ± SEM). This indicates that despite only a 1-second difference in required action performance, animals needed more than double the number of training days to reach criterion, clearly indicating differential effort demands.

      (c) Figure 2 results: the authors state that because there is a greater DA signal to Go vs No Go cues in all the task variants, this means that controllability of reward pursuit and increased task effort do not affect VMS dopamine. But the magnitude of the signals looks different across the task variants - it looks clearly stronger overall in the self-initiated tasks, for example. Given that dopamine signals are not compared across task variants (I think the tasks are all between-subjects?), I don't think the above conclusion is justified.

      We respectfully clarify that our conclusion does not claim controllability and effort have no effect on dopamine magnitude, but rather that these manipulations do not affect the action-selective difference in dopamine (Go > No-go). Our central finding is that the relative difference between Go and No-go remains consistent across all task variants (within-subjects comparison).

      We did not perform across-variant comparisons of absolute dopamine magnitudes because that was not our primary research question. Our focus was to understand whether dopamine differs between action initiation and suppression, and whether this difference can be modulated by controllability or effort.

      We acknowledge the reviewer’s observation that absolute magnitudes appear larger in self-initiated vs cue-initiated task variants. We believe that this likely reflects differences in RPE timing rather than controllability per se: in cue-initiated tasks, the nose-poke light provides an early trial-start signal, distributing RPE temporally across the trial. In self-initiated tasks, trial-initiation and action requirements are temporally integrated. Though understanding how controllability affects absolute dopamine magnitude is an interesting question for future work (e.g., using sophisticated regression-based encoding models), it was beyond the scope of our current investigation, which focuses on action-selective encoding.

      (4) Broadly, I don't think these data, as shown, support the conclusion "dopamine reflects action initiation but not controllability or effort" without more analysis and additional context.

      We have revised the wording of our conclusions throughout the manuscript to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). It is now “Reward-related dopamine depends on action initiation irrespective of controllability and effort”.

      (a) Figure 3 - more description of the classified behaviors would be helpful for interpreting this part of the data. When are the behaviors occurring - during the hold cue? Or is the classification related to what they do immediately after holding? Or something in between> I guess I'm not sure what digging and biting are in the context of a nose poke hold. As described, it's not clear what the signal differences relate to - movement differences? Generally, it's not clear what to make of the behaviors. They seem to emerge spontaneously, but it's not clear whether the specific actions mean anything, so it's a bit difficult to know what to glean from the dopamine is greater during "biting". It's a very different movement pattern, so perhaps this result relates to that, rather than task engagement or motivational drive per se?

      We thank the reviewer for the comment. We have added relevant information in the figure caption and in the Results section to clarify that classified behaviors occurred during the action-cue period (for No-go trials, the hold cue; Figure 3a caption). In addition, we have included Videos to better depict the classified No-go behaviors during the action-cue period.

      We agree with the comment that these classified behaviors, such as biting, seem to emerge spontaneously. Our interpretation of these behaviors is that they may represent the motivational state of each subject. Most importantly, whereas the behavioral differences occurred during the action-cue period (while animals had to suppress actions and stay within the nose-poke port), the difference in VMS dopamine was only observable after this behavior was completed. Thus, the movement pattern per se is likely not relevant to the dopamine release occurring after its completion. This has been discussed in the Discussion section: Dopamine dynamics are linked to motivation action.

      Based on our videos, it appears as though Digging could be perceived as more vigorous (i.e., more general movement in the nose-poke port). However, we did not observe more dopamine during the action-cue period of Digging trials as compared to Biting trials. Furthermore, more vigor during the action-cue period (e.g. Digging trials) did not result in more dopamine during the reward approach period. Together, the data suggest that another process may underlie the large increase in VMS dopamine in Biting trials during reward approach, such as varying attribution of incentive salience.

      (b) In some cases, but not all, dopamine measurement comparisons are done on a total trial basis, and in others, it seems to be subject averages. It's not clear why different approaches are used for different parts of the data. But also, for the trialwise analysis, what statistical steps were taken to incorporate the subject as a random factor in the analysis? If that is not done, then a trial-wise analysis artificially increases the power for the stat (n=trial#).

      We used different analytical approaches depending on sample size and data structure. To compute differences in Go vs No-go dopamine within each animal, as intended by our experimental design, we performed subject-level comparisons (Figure 2).

      For the behavioral classification analysis (No-go, Figure 3), we performed trial-level analyses to increase the statistical power and better characterize this unexpected and interesting phenomenon. We explicitly chose not to perform subject-level group comparisons because

      (1) some groups had only n=2-3 animals, making subject-level statistics underpowered, and (2) behavioral classifications were highly stable within individual animals (see Author response image 5). We acknowledge that formal between-group comparisons (across subjects) are underpowered due to small n, but the stability of within-subject No-go behavioral strategy and qualitatively distinct VMS dopamine profile suggest that these differences may be biologically meaningful and worthy of future investigation in larger samples. We have made these limitations more explicit in the Results.

      This relates to Figure 3, where all trial data are shown next to individual subjects - the subject-wise group comparisons are between 2-5 or so rats, which is quite low. In Figure 2, a subject n of 27 is listed, so it's not clear why this analysis is on such a small set of rats. Generally, it's not clear how many rats/subjects are in each data bit. The methods say only 5 rats are in the long self-initiated task, but 15 in the cue-initiated task. Clarity in all this is needed, including in the figures/captions.

      We have now added detail about the statistical test performed in Figure 3’s caption:

      “After action-cue offset: No-go (trials)’: Individual No-go trials classified by No-go behavior, and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” We have also added in a table in Methods (Table 1), showing the number of trials in each No-go classification, and animals regrouped based on their predominant No-go strategy (see Methods section: Statistical analysis - Clustering No-go behaviors and regrouping animals).

      For the small sample sizes based on the regrouping of animals based on their predominant No-go strategy, we have now added in the caption, “... Rats classified based on their predominant No-go strategy, with no statistical tests performed.”

      (5) I'm also a little confused about the paper's narrative that the dopamine data reflect spatial (but not temporal) proximity to reward - it seems that this conclusion is based on the dopamine signal peaking at magazine entry, but that is different, I think, than a spatial signal per se (space is not manipulated in this study). I think more analysis of the signals during the magazine approach behaviors would be helpful, and possibly comparing rewarded vs unrewarded approaches. The emphasis, including in the title, that a major take-home of the data is that dopamine encodes reward proximity, is not really borne out by the current analyses. Reward expectation is not manipulated independently of the approach action, so it's hard to pin this on "space" vs "reward is soon". This is admittedly a general complexity in characterizing dopamine ramps.

      As the reviewer notes, we acknowledge that 'spatial proximity' and 'reward is soon' are challenging to fully dissociate in appetitive approach paradigms. However, we believe that our data and new additional analyses, which is now included in Results: Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards and Supplementary Figures S10-12, provide compelling evidence that VMS dopamine primarily reflects spatial proximity to the expected reward location, rather than temporal proximity to reward delivery.

      Evidence against full temporal encoding:

      (1) Max VMS dopamine occurred up to 3s before reward delivery in Free trials (Fig 2e, open circles vs triangles). Furthermore, when we calculated the max values of individual trials of realigned traces, maximum dopamine does not consistently coincide with reward delivery across trial types (new Supplementary Figure S11).

      (2) If dopamine encoded temporal proximity from the earliest reward-predictive cue, we would expect consistent accumulation from cue onset (action cue for self-initiated, nose-poke light for cue-initiated). However, realigned trials also did not show a consistent accumulation around these events (new Supplementary Figure S12).

      Evidence for spatial encoding:

      (1) Max dopamine consistently occurred when animals arrived at the magazine across all trial types and task variants, regardless of when reward was actually delivered (Fig. 2e).

      (2) Across individual trials, max dopamine values concentrated around the magazine-panel, with a striking accumulation when animals were in close proximity to it (new Supplementary Figure S10) showing a distance-dependent distribution.

      (3) Assessment of unrewarded magazine approaches during the intertrial interval (ITI) revealed no increase in dopamine release (Fig 2f), indicating that VMS dopamine requires task-relevant reward expectation.

      (6) Examples of the DLC workflow in a supplement would be appropriate. Also, video examples of the 3 kinds of behaviors from the clustering analysis could be useful for understanding what they are/what they mean.

      We have added supplementary figures showing the DeepLabCut workflow (Supplementary Figure S9) and Videos 1-3 demonstrating the three behavioral classifications (biting, digging, calm) during No-go trials.

      (7) Referencing/scholarship: I would suggest broadening the citation pool for the paper to include more older work that has established the notion that dopamine signaling and the accumbens act as a motivation-action interface, as this has been a longstanding notion since at least the 1980s. There is also a sizable literature on dopamine signaling of effort, some of which would be appropriate to cite here.

      We thank the reviewer for the recommendation and have now expanded our citations to include foundational literature to work from the 1980s-90s. These can be found in the Introduction, Results, and Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 353-357- This came as a surprise, as the relevant results are only featured in supplementary information. These should be moved to the main manuscript. As both the biting behavior and faster lever-press completion lead to larger peak dopamine, does this represent response vigor?

      We appreciate the reviewer’s interest in these data. However, we believe that the data that we report in “Supplementary Fig 7: Quartile analysis of Go trials shows coordinated changes in last lever-press timing and VMS dopamine, does not warrant movement to the main manuscript.”

      The purpose of the lever-press timing analysis was to demonstrate that maximum dopamine release coincides with the moment animals arrived at the magazine, rather than with the action period itself. By sorting Go trials based on last lever-press latency, we show temporal coordination between action completion and dopamine timing but critically, dopamine peaks after action completion, not during it.

      This temporal dissociation argues against a 'response vigour' interpretation. If dopamine encoded motor vigour, we would expect the signal to coincide with or precede the vigorous action. Instead, both the lever-press data (Supplementary Figure S7) and the classified No-go trial data (Figure 3, biting behavior) show that changes in max dopamine occur after the actions themselves, during the subsequent approach to reward.

      Together, these findings demonstrate that VMS dopamine reflects spatial approach to the reward location following action completion, not the vigour of the required actions per se. The lever-press analysis serves as supporting evidence for this temporal relationship but does not introduce a novel finding that warrants main figure emphasis. We mentioned this temporal coordination in Results to ensure readers are aware of the converging evidence while maintaining focus on our central findings regarding action initiation versus suppression.

      (2) Line 13- the experimental work cited refers to midbrain dopamine neurons, rather than dopamine release within the VMS. Please correct.

      We thank the reviewer for pointing out this error. We have now corrected the citation to refer to studies measuring striatal dopamine rather than midbrain dopamine neurons.

      (3) Line 166 - Shouldn't this say "consistently delayed for No Go trials"? Looks like peak dopamine occurs later for these trial types.

      This statement refers to traces aligned to nose-poke exit (not action-cue onset). The peak Go dopamine occurred later than other trial types following exit (Green arrows in Figure 2d).

      (4) Lines 391-392 - however however

      Thank you for the comment, we have adjusted the text.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In their important manuscript, Gangadharan, Kober and Rice focus on how Stu2/XMAP215-family microtubule polymerases use their TOG domains to catalytically promote microtubule growth, testing whether their mechanism follows an enzyme-like kinetic model similar to that of actin polymerases. The authors integrate measurements including microtubule polymerization rates and TOG-tubulin binding kinetics to convincingly show that Stu2 follows an enzyme-like model where tight tubulin binding enables efficient polymerization, revealing a shared mechanism with actin polymerases despite their evolutionary divergence. This work will be of general interest to the cell biology and biophysics communities.

      Thank you for the favorable assessment of our manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Gangadharan and colleagues provides significant progress towards a quantitative biochemical mechanism for Stu2 polymerase activity. A key conceptual advance is the novel application of an enzyme-like model, initially developed for the actin polymerase Ena/VASP, to Stu2.

      New refined affinity measurements for a Stu2 TOG domain using Bio-layer interferometry show more than an order of magnitude higher affinity of TOG domains to tubulin compared to previously published reports.

      The findings reinforce the "concentrating reactants" or, more specifically, for TOGdomain proteins, the "tubulin-shuttling antenna" model, compared to the "polarized unfurling" model, a more speculative structural hypothesis.

      The manuscript builds upon a series of previous manuscripts that showcase the profound intellectual engagement with microtubule polymerization mechanisms by TOG-domain proteins from the Rice lab, a thought leader in microtubule polymerization for over a decade.

      Minor remarks:

      (1) A major new experimental finding of this paper is the affinity of TOG domains, which is more than an order of magnitude lower (10 nM) than previous measurements from the same lab (~200 nM). The authors attribute this change to ionic strength differences between buffer conditions, citing the lab's previous work (Ayaz et al., 2014). This argument left me contemplating what the buffer conditions are in both experiments, and I wonder if other readers would feel the same. After going down the rabbit hole, I believe the difference in ionic strength is ~2.3 fold, and at least on the back of my envelope, this works out beautifully with the measured differences in affinities. A short version of this argument may strengthen the manuscript.

      This is a good comment. We should have been clearer about the different buffer conditions. The revised manuscript now explicitly states how the two buffers in question differ in pH and ionic strength. (Page 8, ‘Tubulin binds rapidly …’ section). We tried to perform comparative measurements of TOG:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is stated in the revised manuscript (Page 9, final paragraph before the ‘Unifying measurements …’ section. Along the lines of the reviewer’s ionic strength calculation, and consistent with the increase in affinity we observed with lower ionic strength, we now also state that prior measurements from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with higher ionic strength (100 mM KCl vs 200 mM KCl).

      (2) I am wondering if there may be an alternative explanation to tubulin binding by TOG being the kinetically rate-limiting step for polymerase function:

      TOG + Tubulin ⇌ TOG:Tubulin (fast binding rate, high-affinity binding)

      TOG:Tubulin + MT_end → TOG:MT (tubulin is incorporated into MT, fast transfer rate)

      The binding rate is 3/s, and the transfer rate is 5/s.

      I was wondering if the following step should be considered, which involves a conformational change of tubulin (e.g., straightening) TOG:MT → TOG + MT (ratelimiting straightening and unbinding of TOG from the lattice).

      This is an interesting thought that highlights a gap in the understanding of microtubule dynamics.

      Presumably, the affinity of TOGs for straight tubulin is practically zero for the purpose of this discussion, as there is no lattice binding, which means unbinding is likely very rapid; however, straightening may be the rate-limiting factor here.

      In theory, straightening should also be rapid; however, we lack measurements of how fast or slow this step occurs within the context of a TOG domain, which presumably skews the process towards curved tubulin.

      We agree (based on prior observations) that the affinity of these TOGs for straight tubulin is negligeable in this context. There is much less data about the timescale of tubulin straightening, with or without a TOG domain bound, or even about how tubulin interactions with the microtubule end affect the balance of preferred conformations and/or the rate of conformational change. It’s an extremely interesting topic. Because the straightening process the reviewer envisions is zero-order, the transfer rate in our model could in principle reflect slow straightening (in this view the ‘delivery’ step would need to be very fast, i.e. not rate-limiting). Because there is so little data about this, and because there are not yet methods to study or perturb the timescale of straightening on the microtubule, we prefer not to engage too deeply. We added a sentence to acknowledge this alternative possibility in the revised manuscript (bottom of Page 4 and top of Page 5).

      A hypothetical Stu2, when bound to the microtubule end and with the TOG domain not disengaged from tubulin, would not permit the processivity of that molecule or the binding of a new molecule.

      To emphasize the importance of unbinding, when it is not efficient, as reported for the T238 mutant that results in Stu2 lattice binding (Geyer et al., 2018), the polymerase becomes inefficient.

      The mechanism of polymerase processivity has not been conclusively determined (the Geyer et al. 2018 eLife paper took a step in that direction, though). The model used in this paper is only concerned with how many polymerases are at the microtubule end at steady-state (as opposed to how long a particular polymerase acts before dissociating), so while we appreciate and are interested in these questions, we think it would be better to leave them for future work.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from the Rice lab by Gangadharan et al. investigates the polymerization mechanism of the yeast microtubule polymerase Stu2. The lab has published a number of articles demonstrating the structural basis by which the two TOG domains of Stu2 each bind free tubulin heterodimers, and has developed a tethered polymerization model by which the TOG domains drive polymerization by shuttling those tubulin subunits onto the microtubule plus end. A second model was proposed by Nithianantham et al. (eLife, 2018) based on a closed-to-open transitional state in which Stu2 unfurls and loads two longitudinally associated tubulin heterodimers onto the microtubule plus end. While the second model is not directly tested, the current work aims to further characterize/model the tethered polymerization model using a kinetic framework developed by Breitsprecher et al. for Ena/VASP actin polymerization activity, using a model that is enzymatic (EMBO J., 2011). The general architecture and function of Ena/VASP on actin polymerization versus Stu2 on microtubule polymerization is a reasonable relation and hits upon, as the authors note, potential convergent mechanistic evolution across distinct cytoskeletal networks. The model effectively treats tubulin as the substrate, and the polymerized microtubule plus end as the product. If Stu2 is "enzymatic" in this framework, the model predicts it would behave with Michaelis-Menten kinetics, that there would a Vmax, and polymerase activity would either be "affinity limited" by TOG:tubulin affinity (KD) and/or "kinetically limited" by TOG:tubulin association (Kon) and transfer of tubulin to the microtubule plus end (Kt). The authors find that the Brietsprecher model works well for Stu2 activity, and that Stu2 best aligns with a "kinetically limited" model. The work is interesting and adds to the growing elucidation of the Stu2 microtubule polymerase model. While yeast microtubule polymerases are somewhat distinct in their architecture, there is significant overlap that findings from the manuscript can be utilized to inform the mechanisms of larger, more complex microtubule polymerases such as human ch-TOG.

      Thank you for the nice summary and favorable comments.

      Strengths:

      The manuscript invokes the enzymatic model of Breitsprecher et al. used for Ena/VASP and conducts an elegant series of (mostly established) experiments to determine whether Stu2 microtubule polymerase activity aligns with the model, which they conclude does align, supported by the data/results obtained.

      Weaknesses:

      The authors used biolayer interferometry to measure TOG:tubulin affinity. The affinities obtained were significantly higher than the lab obtained in an earlier publication using analytical ultracentrifugation. While differences in buffer and salt conditions may underlie these differences, additional runs using comparable buffer systems, or the use of a third independent assay to measure affinities, would have added rigor.

      This is a good question that was also raised by reviewer #1. We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is now stated in the revised manuscript (page 9, final paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been independently shown to depend on ionic strength in a way that seems consistent with what we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, final paragraph before the ‘Unifying measurements …’ section).

      The discussion could be expanded to better compare and contrast the results with both existing polymerase models introduced in the introduction, as well as expanded to look at reversible enzymatic activity (microtubule depolymerization at low to zero tubulin concentrations) and microtubule plus versus minus end activity.

      Thank you for the push to be more explicit about the two contrasting models. We made small changes to the introduction (top paragraph on page 3) and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (page 12, penultimate paragraph of the main text).

      The ‘transfer’ reaction is treated as irreversible (analogous to catalysis by an enzyme), so the biochemical model we use for the polymerase cannot account for polymeraseinduced microtubule depolymerization at low to zero tubulin concentration. We added text to state that the model is limited to the growth reaction (page 4, last paragraph) but otherwise prefer to not engage too deeply in questions about the reverse reaction.

      These polymerases are thought to be plus-end specific because of the domain organization of the protein: TOGs bind tubulin such that the N- to C-terminal polarity of the TOG corresponds to the plus- to minus-end polarity of the tubulin, and the basic region used to make a ‘slippery’ connection to the microtubule is located C-terminal to the TOGs. These two factors mean that it is only at the plus-end that TOGs can engage αβ-tubulins with the basic region contacting surfaces ‘deeper’ in the polymer. We added text about these issues, citing prior work, to the legend of Figure 5 (page 11). We chose to not elaborate much since it is not the primary focus of the paper.

      Reviewer #3 (Public review):

      Summary:

      This study by Gangadharan and colleagues seeks to establish a quantitative biochemical model for the microtubule polymerase activity of Stu2. Stu2 is the budding yeast member of the XMAP215 protein family, which is broadly conserved across eukaryotes. XMAP215 proteins play a wide variety of important roles in cells, and these are attributed to effects on microtubule dynamics. Many studies over the last ~20 years have shown that XMA215 proteins selectively associate with microtubule ends, where they increase rates of microtubule assembly and disassembly. More recently, structural biology and biochemical studies by the authors and other groups have shown that the multiple TOG domains on XMAP215 proteins are tubulin-binding domains that selectively bind to curved tubulin, which is present in solution and at microtubule ends, but not to straight tubulin which is present in the walls of the microtubule lattice. This has led to the general model that XMAP215 proteins promote polymerization by delivering soluble tubulin to the growing plus end, and two distinct models have been proposed to explain the mechanism. The 'concentrating reactants' model proposed previously by the authors suggests that TOG domains grab hold of tubulin in solution and concentrate at the microtubule end. The 'polarized unfurling' model proposed by the Al Bassam lab suggests that XMAP215 delivers multiple tubulins to the end, using a step-wise mechanism involving different roles for each TOG domain. The current study seeks to improve our understanding of the mechanism by developing a quantitative model to explain the binding and release of tubulins, the number of Stu2 molecules at the end, and the overall rate of tubulin addition. The authors accomplish this goal using new experimental data. The final model fills in new details of the mechanism. The authors draw a comparison between Stu2 and the actin polymerase, which bears similarity to Ena/VASP, and suggest a convergent strategy for cytoskeletal polymerases.

      Thank you for the good summary and favorable comments.

      Strengths:

      This is a focused and clearly written study that incorporates prior knowledge of XMAP215 and draws inspiration from the actin field. The data are clear and convincing, and the study accomplishes its goal of generating a new, quantitative model for Stu2. The model will be important for microtubule researchers to predict and test key points for altering XMAP215 activity across different organisms and potentially for different tubulin substrates. The comparison to Ena/VASP may also inspire similar comparisons across other microtubule and actin regulators, which could lead to new insights across the cytoskeletal fields.

      Thank you for these comments.

      Weaknesses:

      The study is without major weaknesses, but there are several minor weaknesses worth noting. One is that the final model provides new details regarding the Stu2 mechanism, but does not provide a major new advance in our understanding of how the polymerase works. For example, the discussion does not clearly argue for whether the new results and model rule out either of the prior models. This appears consistent with the 'concentrating reactants' model, but does it clearly rule out the 'polarized unfurling' model?

      Thank you for pointing out what in retrospect was an obvious ‘loose end’ in our discussion. The other referees raised the same point. We made small changes to the introduction and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (top paragraph of page 3 and new penultimate paragraph of the manuscript on page 12).

      A second minor weakness is that the comparison to Ena/VASP is not developed at a deep level based on the final model. I found these ideas exciting and want more critical consideration here, but perhaps it is better suited for a commentary piece to follow.

      We appreciate the enthusiasm and understand where this comment is coming from. Because there has been a fair amount of recent movement in the understanding of TOG domains and what they can do, and because some of the mechanistically interesting parallels entail speculation, we agree with the referee’s suggestion that a future commentary will provide a better venue.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Other minor remarks:

      (1) Figure 2 C is missing the label for what should probably be TOG1*-TOG2.

      Fixed

      (2) Figure 5, lower left, is oddly cropped, showing residuals that are slightly distracting from the beauty of the model.

      Apologies that the figure did not look good in the initial submission. We adjusted it and it looks much better in the revised submission.

      (3) The dynamics assay buffer composition stated in the protein purification Method section is not the same as the PEM buffer used for the dynamics assay. And both are different from BRB80, which, with the chambers, are rinsed. This may very well be accurate, but it raises the question of why not stick to one version, as they are virtually the same.

      Thanks for asking these questions, and sorry for the confusion. First, we should have used different names for the (barely) different buffers. This has now been corrected. Second, the PIPES concentration was not 90 mM, it was 100 mM as in our prior work and this discrepancy failed to get caught in proofreading. Why the other small differences? It’s a good question. The differences reflect an arbitrary decision made at the beginning of the work, there is not a deeper rationale.

      (4) Out of curiosity: Why 90 mM PIPES and not 80?

      Why not 80 mM PIPES? This is just a historical difference. The early measurements of yeast microtubule dynamics (e.g. Gupta … Himes MBoC 2002 and Bode … Himes, EMBO Rep 2003) that partly inspired us to use yeast as a model system used 100 mM PIPES as the working concentration, and we never deviated from that.

      (5) Please state the source of PIPES.

      Sorry for the oversight, we have added the source of PIPES (it is Millipore Sigma P6757).

      Reviewer #2 (Recommendations for the authors):

      (1) Page 4, last paragraph, the authors call out "Fig 1C" which I believe should be "Fig 1D".

      We fixed this, thank for catching this error

      (2) Figure 2C: The authors subtract basal tubulin polymerization (growth rate) from the rates measured in the presence of Stu2 constructs. One assumption in doing this is that Stu2 polymerization activity does not compete for the ability of tubulin (not bound to Stu2) to polymerize on the plus end. I think this is a logical assumption, but it would be beneficial for the authors to state this assumption.

      We said this explicitly in the revised submission (first full paragraph on page 7), borrowing from the reviewer’s phrasing.

      (3) Figure 2C: the label for the last bar is missing - presumably: " Stu2 (TOG1*-TOG2)".

      Fixed.

      (4) Figure 2D: Many of the KM values determined are at the border for points measured, or in one case, beyond the concentration of tubulin sampled. In this regard, the authors should discuss how well the fitted curves correlated with their data. Also, as Vmax and KM are calculated, it would be beneficial if another panel were produced (e.g., Figure 2E) in which the data were presented as a Lineweaver-Burk plot. Doing so, the authors would be able to test their enzymatic model by doping the system with their nonpolymerizable tubulin mutants, which should yield competitive inhibitor behavior but not change Vmax.

      Thanks for pointing this out, we should have been more explicit about this point. We incorporated into the results section an explicit statement about this limitation (first full paragraph on page 7). The suggestion to use blocked mutants and Lineweaver-Burke plots is an interesting one that we hope to pursue in future work using blocked or other mechanism-specific mutants. But we think to do so is complicated enough to be beyond the scope of the present study.

      (5) Figure 3: The authors quantitate the amount of Stu2-GFP fluorescence at microtubule plus ends using line scan analysis of the kymograph. Since the kymographs are processed images, it is more appropriate to integrate intensity from the original frames collected using a circular area. i.e., ID points on the kymograph, and return to the respective position in the corresponding frame to calculate background-subtracted GFP intensity at the plus end.

      This is a fair point. We chose to stick with the kymographbased analysis because the symmetry of the point-spread function and lack of rapid variation in GFP and/or background intensity means that the kymograph analysis is adequate for the intensity-based comparison we were doing.

      (6) Page 7, last line, the authors call out "Fig 2D" which I believe should be "Fig 1D".

      Sorry for the error, we have corrected it.

      (7) Figure 4A: The authors discuss the "sortase epitope," but technically, an epitope is the binding site specifically for an antibody, not to be used in general terms for proteinprotein interaction sites. As such, the authors should describe this as the "sortase recognition sequence" or something similar.

      Thank you for noticing this; we had indeed used ‘sortase recognition sequence’ elsewhere in the paper but we did not catch this instance during proofreading. We have now used that same language in the legend for Fig. 4.

      (8) Page 8, the authors state "KM is approximately equal to Kt/Kon and KM negligeable," but I think they mean "...and KD negligeable".

      Thank you for noticing this typo. We corrected it.

      (9) Page 9, first paragraph last sentence: the readership would be aided by modifying the sentence as follows (adding "Kon" and "Kt"): " ... must be kinetically limited by either the rate of TOG:tubulin binding (Kon) or by the rate of TOG-mediated transfer to tubulin to the microtubule end (kt)."

      Very good suggestion, we implemented it.

      (10) Page 9, second paragraph, the authors call out "Fig 2D", but perhaps they intended to call out "Fig 1D"?

      Sorry for the error, the reviewer is correct and we fixed this.

      (11) Page 9, second paragraph: The authors mention that the transfer rate of tubulin to the plus end via a TOG domain is close to the transfer rate of free tubulin to the growing plus end. Can the authors expand on why they are mentioning this comparison?

      Thanks for the push to be clearer about this. The basic idea is that each ‘delivering’ TOG contributes 50% of the background (uncatalyzed) polymerization rate. So the presence of multiple TOGs (in a single polymerase or from multiple end-resident polymerases) can substantially increase the rate of polymerization. We added brief text to try to make this clearer (first paragraph on page 10).

      (12) Page 9, second paragraph: "TOG-TOG2 polymerases" would be better phrased as "TOG2-TOG2 dimeric polymerases". Noting as well that "2" is missing from the first "TOG".

      This is indeed better phrasing and we have adopted it (also corrected the missing ‘2’) (first paragraph on page 10).

      (13) Page 13, BLI methods: The authors should list the final pH for the PIPES buffer (was it pH 6.9 as in the polymerization assay?).

      Sorry for the oversight, we have added the pH and it was indeed 6.9.

      (14) Page 13, BLI methods: What is "LR1-457"?

      LR1-457 is lab-notebook-speak that did not get purged in editing; it refers to the polymerization-blocked tubulin mutant that also carried a sortase recognition sequence. We replaced ‘LR1-457’ with more evocative phrasing and corrected another typo we found there.

      (15) In Ayaz et al. (eLife, 2014) Stu2 TOG1 and TOG2 affinities for tubulin were measured using AUC, for which the fitted curves appeared to correlate with the data quite well. As the authors note, the values were KD = 70 nM and 160 nM, respectively. This contrasts with the BLI measurement for TOG2-tubulin (~10 nM), which suggests that at least one of the experiments was off the mark - or, as the authors do note, that different buffer and salt condition was used could account for the differences, but that the BLI conditions align with the polymerization conditions (though not exactly) and thus are more appropriate to use. In a supplemental discussion, the authors should run the AUC values through the equation for their model and state what types of differences these values could imply for Stu2 mechanism. If the differences are significant for the Stu2 model derived, the authors should give thought as to whether a third assay should be employed to determine TOG-tubulin affinity. Based on the BLI reagents, it appears the authors would be well-positioned to conduct an assay using SPR. As a potential alternative, the authors could repeat the BLI experiment using the buffer conditions from the Ayaz et al., AUC work (25 mM Tris pH 7.5, 1 mM MgCl2, 1 mM EGTA, 100 mM NaCl, 20 μM GTP) - noting that BSA and Triton X-100 may need to be added as well. If the authors are able to replicate the ~160 nM affinity for TOG2:tubulin, this would be a reasonable way to bootstrap to the conclusion that the BLI is measuring affinity correctly and that the current PIPES-based BLI experiments yielded accurate data.

      We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems. Unexpected challenges and personnel turnover made this slower than anticipated. The reviewer’s suggestions are completely reasonable, but ultimately it was not possible for us to get side-by-side results for TOG:tubulin affinity using the same measurement technique, whether it was biolayer interferometry, analytical ultracentrifugation, or isothermal titration calorimetry. For each technique one or the other buffer caused aggregation, non-specific binding, or some other artifact that prevented analysis. The fact that we were unable to compare the buffer conditions in this way is stated in the revised manuscript (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been shown to depend on ionic strength in a way that is consistent with the changes we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added some text to address the comment about affinity and whether/when the shuttle model would hold (first paragraph on page 11).

      (16) A sentence or two in the discussion, relating how their data aligns (or not) with the Nithiantham model would be beneficial, especially as discussing the two models in the introduction was a central point.

      We completely agree and have now added a paragraph to the discussion to explicitly address the two models and how are or are not supported by the new model and observations (page 12, penultimate paragraph of the main text).

      (17) Discussion: The model in Figure 5 depicts Stu2 engaged with the microtubule, perhaps using its basic linker region (?). The authors could note this in the figure caption for 5A. The authors do not discuss the basis for plus-end polymerization activity versus polymerization activity at both the plus and minus ends. Do the authors propose that this is due to differential Kt values for the two ends and/or differential localization via the basic region to the two ends?

      Thanks for bring this up. The plus-end selectivity of these polymerases is thought to result from the polarity of TOG:tubulin engagement and the positioning of TOG domains relative to the basic region that provides ‘slippery’ binding to the microtubule lattice. We have partially addressed these issues in the legend to Figure 5 (page 11).

      (18) Brouhard (Cell, 2008) demonstrated that XMAP215 can catalyze the depolymerization of GMPCPP microtubules when no free tubulin, or very low levels of free tubulin, are present. This is interesting in that it indicates that the enzymatic activity is reversible. Can the authors comment on how their model would behave in the low-tozero free tubulin concentration regime? Would a different model have to be invoked?

      This is an interesting comment. Because the enzyme-like model treats the transfer step as irreversible, the model cannot account for the kind of ‘depolymerase’ activity Brouhard and others have noted. A more general model that could also encompass the depolymerase activity at low-to-no free tubulin would need to explicitly model the step(s) involved in microtubule association and dissociation. These steps remain a major open question in the field and while it would be quite interesting, trying to address this in a model is beyond the scope of what we can confidently do given the data we have. To be more explicit about this assumption/limitation, we now point this out in the results section where the model is introduced (page 4, last paragraph).

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1D, legend. "...and a transfer rate constant kf that describes how fast...". Should kf be replaced with kt?

      We made this correction, thanks for pointing the problem out

      (2) Figure 2C. The x-axis label under the blue bar is missing. Also, I find the arrows to the left of the bars confusing and unnecessary.

      We fixed the legend problem. We sympathize with the dislike of the arrows but respectfully prefer to keep them in the hopes of avoiding confusion about the fact we are fitting ‘growth rate attributable to Stu2’, not simply growth rate. The figure legend has been expanded to hopefully smooth this over.

      (3) Figure 3. The kymographs are convincing, but it may be helpful for future studies to state here what the polymerization rates are for 0.6 µM and 1.4 µM yeast tubulin. These values are probably different than what one might expect for mammalian tubulin at those concentrations, and the authors could simply determine them from the slopes in the kymographs.

      Good suggestion, we added the growth rates to the legend (as the reviewer expected, they differ from expectations based on mammalian tubulin).

      (4) Results, page 9, line 11: "...yielded a value of 9.6 nM...". Should this be 8.9 nM, which is that value stated in Figure 4C?

      Actually, these different values are correct. We just wanted to point out that whether we used response amplitudes or measured on- and off-rates, we get very similar values for K<sub>D</sub>. We changed wording to hopefully make this clearer: “Calculating the dissociation constant K<sub>D</sub> from the measured rate constants (K<sub>D</sub> = k<sub>off</sub>/k<sub>on</sub>) instead of from the amplitudes yielded a value of 9.6 nM, in good agreement with the amplitude-based determination of 8.9 nM.” (page 9).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The factors that create and maintain diversity in host-associated microbiomes remain poorly understood. A better understanding of these factors will help in the efforts to leverage the adaptive potential of the microbiome to help solve pressing problems in health and agriculture.

      Experimental evolution provides a promising path forward as we can track the causes and consequences in the emergence of novel variants, but experimental evolution remains underutilized in host-microbiome interactions. Here, Gracia-Alvira utilizes a long-term experimental evolution study in Drosophila simulans under hot and cold temperature regimes to identify strain-level variation in an important fly bacterium, Lactiplantibacillus plantarum. They identify three strains of L. plantarum, which are most prevalent in their respective three temperature regimes, suggesting that these are locally adapted bacteria. Then, using a combination of genomics, in vitro, and in vivo, Gracia-Alvira et al attempt to understand the factors that led to the differentiation of the hot and cold L. plantarum and their impacts on the fly host.

      Strengths:

      This is an excellent use of experimental evolution to track the emergence of novelty in the microbiome. The genomic analyses are all solid and appropriate for the data sets. It is especially striking that the comparisons with the other, independent experimental evolution studies in different labs (and across continents between Portugal and South Africa) show a consistent response to temperature. Many have disregarded the microbiome as it is something that is too sensitive to seemingly innocuous variables (particularly in the fly microbiome), such that we cannot find generalities. However, this finding highlights the potential for experimental evolution to uncover these dynamics. The question of how strains emerge and are maintained is timely and is one of the key open questions in host-microbiome evolution currently.

      Weaknesses:

      (1) The framing in the title and throughout the discussion about "subspecies competition" does not match the data that was collected. The subspecies competition requires actually tracking the competitive outcomes between the hot, cold, and unevolved L. plantarum. In the in vivo work, I can see that mixes of the strains were made, but they did not track whether the cold strain outcompeted the hot strain in vivo under cold conditions, for example.

      We thank the reviewer for the honest concern and take this opportunity to defend our claim of "subspecies competition used across the manuscript. As the reviewer states, subspecies competition requires tracking the competitive outcomes between the three clades, and this is what we did by sampling and sequencing across ten years of experimental evolution (Figures 4 and S3). For this reason, we point that the subspecies competition assessment comes from the direct observation of changes in relative abundance across the time series, and not from the follow-up experiments in vivo or in vitro.

      While Figure 4 is suggestive that there is ongoing competition in the hot temperature regime, this is not necessarily shown in the cold, which is dominated by the C clade. It could also be that the bacteria cannot survive in the flies at the different temperatures. The growth curve assays hint that the bacteria can grow, but the plate reader couldn't actually maintain the 18 {degree sign}C temperature (line 455). So all of this evidence is very indirect and insufficient to say that strain competition is driving these patterns.

      We thank the reviewer for the alternative hypothesis that could explain the observed subspecies dynamic. We rule out that dominance of clade C in the cold occurs because the other two clades cannot grow in this regime based on three pieces of evidence:

      (1) In the time series, clades H and U decrease, but never disappear (Figures 4 and S3), even showing some peaks of abundance in specific replicate populations (Figure S3).

      (2) We isolated individuals belonging to clade H in the cold-evolved populations, as shown in figure 2. This is a direct evidence that clade H prevails in the cold-evolved populations, although in low abundance.

      (3) We did grow the three taxa in fly food Petri dishes incubated at both temperature regimes, observing growth in all cases.

      We will include the food growth experiment in the revised manuscript as further supporting evidence for growth in both regimes.

      (2) The in vivo results are interesting in that there appears to be a fitness cost of clade C, but the explanation is underdeveloped. I say under-developed because in Figure 4, the cold L. plantarum remains much higher throughout adaptation to the hot temperature regime than the hot L. plantarum in the cold regime. The hot L. plantarum is low abundance throughout the cold regime. I felt like this observation was not explained, but it seems relevant to understanding the strain dynamics.

      We acknowledge that a strong fitness cost of clade C is observed in axenic D. melanogaster. In the native host, D. simulans, with reduced microbiome, we observed delayed development that could even be an advantage depending on the situation, as pointed out by reviewer 3 in the recommendations.

      Even if we assume that flies colonized with clade C are less fit in the experimental evolution, another caveat is whether the flies can actively select for the L. plantarum clade. Under this assumption, a clade that imposes a fitness cost to the fly (clade C) should be selected against over time because the flies colonized by this clade will have less offspring or develop later than the rest. Alternatively, as the microbiome is shared among all the individuals in the population, the host might not be able to “purge” the pernicious clade, and L. plantarum dynamics might be controlled solely by the relative fitness between clades in the given experimental treatment. We will discuss this hypothesis in the revision as a way to explain the relationship between the abundance of each clade and the effect on the host.

      I will also note that this is not the first time that L. plantarum or other Lactobacillus have been shown to exert fitness costs to Drosophila. Gould, PNAS, 2018, shows that both Lactobacillus plantarum and Lactobacillus brevis in mono-association have lower fitness (measured through Leslie matrix projections using lifespan and fecundity) than axenic flies. Many studies of wild Drosophila fail to find Lactobacillus, or it is low abundance (e.g., Chandler, PLoS Genetics, 2014; Wang, Environmental Microbiology Reports, 2018; Henry & Ayroles, Molecular Ecology, 2022; Gale, AEM, 2025). This might help provide useful context for the in vivo results.

      We thank the reviewer for the references. These observations are compared to our phenotypic results and discussed in the revised version of the manuscript.

      (3) The data in Figure 4 are compelling to focus on the L. plantarum variants. However, I can see from the methods that the competitive mapping included only other strains of Wolbachia.

      We appreciate the thorough reading of the methods by the reviewer. The competitive mapping comprised two steps: first we discarded the reads that mapped to Drosophila, Wolbachia and additional potential contaminants from sequencing facitilies (human, dog...). This step leaves the reads originated from whole the external microbiome of the flies, including L. plantarum. The second competitive mapping step recruits the reads that map any clade of L. plantarum.

      It is not clear how other members of the microbiome changed in response to the temperature regimes. As I note in point #2, given that Lactobacillus is often rare, it is not clear what the rest of the microbiome looks like over the course of adaptation. Indeed, it seems like Mazzucco & Schlotterer, PRSB, 2021 did a broader analysis of the microbiome and found that Acetobacter is by far the most common bacterium (I think this data is also part of the data shown here?). Expanding on why or why not in this context is important and will improve this study, particularly if the focus is on connecting these evolutionary dynamics to ecological competition to explain the emergence of strain diversity.

      We acknowledge that the rest of the Drosophila microbiome is not addressed in this study, as we wanted to focus the storyline around the intraspecific dynamics found in L. plantarum. We consider that a complete characterization of the whole Drosophila microbiome would unnecessarily elongate the paper and thus we treat it as a constant biotic factor.

      We must point out that our dataset is not the one reported by Mazzucco & Schlötterer, which was done in D. melanogaster, rather than D. simulans. Nevertheless, both experiments share the same infrastructure, temperature regimes and fly maintenance.

      We have included a list of taxa that were isolated from the populations, as well as to report L. plantarum prevalence and abundance across the experiment in order to provide context of the microbiome, beyond L. plantarum, to the readership.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gracia-Alvira et al. investigated how environmental temperature affects competition among members of the microbiome, with a focus on intraspecific diversity, using the Drosophila model.

      Notably, the authors identified three clades of Lactiplantibacillus plantarum from a natural population of Drosophila simulans collected in Florida. They tracked the dynamics of these three bacterial clades under two temperature conditions over the course of more than ten years. Using comparative genomics and phylogeny, they showed that these three bacterial clades likely adapted to their host independently in a temperature-specific manner. Further, by combining in vitro culture and in vivo mono-association assays, they demonstrated the functional divergence of these three bacterial clades phenotypically, including their growth dynamics and effects on host fitness. Lastly, they performed pathway analysis and speculated on key genomic variance supporting such functional divergence.

      Strengths:

      The laboratory evolutionary experiment in response to cold or hot environmental temperature is impressive, given its more than ten years of experimental time period. This collection of achieved microbiome samples paired with the fly host data can be a valuable resource for the field.

      Weaknesses:

      The laboratory evolutionary experiment can be limited due to its artificial experimental setup. For example, wild flies rely on a more diverse set of food sources and are constantly exposed to new bacterial inoculations, whereas under laboratory conditions, flies live in a more restricted ecosystem. In addition, environmental temperatures differ among different locations, but they also involve seasonal changes within the same region. This manuscript can be strengthened with further discussions that elaborate on these limitations.

      As the reviewer has correctly noted, our experimental setting is not exempt from limitations. Lab-reared flies are fed with a defined standard diet. Furthermore, although the system is not completely closed to bacterial migration, this is limited as replicate populations are not allowed to mix during the maintenance of the flies. For this reason, we consider our laboratory setting as a compromise between observing wild populations, which undergo all biotic and abiotic stresses but cannot be manipulated, and evolving the bacteria in absence of the host, or in gnobiotic hosts, in which biotic interactions are not fully considered. We will extend on this in the new version of the manuscript.

      Moreover, the extent of host effects involved in these experiments remains ambiguous, because it is unclear whether these Lactiplantibacillus plantarum mostly reside within fly guts or on Drosophila medium. The laboratory evolutionary experiment possibly favored better colonizers on Drosophila medium under either cold or hot temperatures, which subsequently can saturate fly guts. As fully dissociating these variables can be experimentally tedious, the authors may want to comment more on these aspects in the discussion. Or they may want to consider some measurements. For example, measuring the growth rate of these bacteria on Drosophila medium under different temperatures, in addition to the current MRS culture experiments, or measuring the portion of the Lactiplantibacillus on Drosophila medium versus these stably colonizing fly guts.

      The reviewer's point was briefly addressed in the Results chapter: "Phenotypic differences in liquid culture".

      Reviewer #3 (Public review):

      Summary:

      The study presents an analysis of 297 pangenomes derived from 20 populations of Drosophila simulans, at 19 time points for fast-reproducing individuals in a hot environment, or at 10 time points for slow-reproducing individuals in a cold environment, over a period of more than 10 years. The authors select a particular microbial component of the pangenomes and study the dynamics of Lactiplantibacillus plantarum strains in two environments. They discover that the revealed operational taxonomic units could be divided into three phylogenetic clades, which have their own genomic and genetic features, different adaptive capabilities that depend on the environment, and have a distinct impact on the fitness of the host.

      Strengths:

      The authors prove that bacterial microbiome components are sensitive to the environment and could rapidly (years) be fixed in eukaryotic populations. This study establishes a tractable model that potentially enables the study of variability of the physiological influence of distinct strains of an important commensal species, Lactiplantibacillus plantarum, on the Drosophila host. It is clearly shown that this single species consists of several phylogenetically and functionally diverse strains. The authors did not limit their interest to their own model, but rather they have integrated a comparative approach by analysing phylogenetic relationships among 92 described L. plantarum strains.

      Overall, the study is novel and delivers important discoveries of a longitudinal, well replicated experiment, generating a substantial amount of genomic data. It highlights an important dimension of research that environmental selection operates at the subspecies level.

      Weaknesses:

      Even though the authors show only one particular example by conducting their longitudinal experiment, they honestly acknowledge failures important for interpretation of the biological significance of the results (gnotobiotic mono-association experiments was done with D. melanogaster, but not D. simulans) and therefore they state limitations of their conclusions (weaker effects in the non-axenic flies are due to the presence of other taxa or to higher-order interactions with other members of the microbiome). These interactions could significantly affect bacterial growth, metabolism, and physiological influence on the host.

      We agree with the reviewer in that the use gnobiotic animals is a limitation, as by "tuning" the flies' microbiome we are modifying the interactions between members, which can potentially change the phenotypic outcome. Nevertheless, we use it as a complementary approach, rather than the only inference in our study.

      The authors exploit the results of their experiment to speculate about a wide range of evolutionary phenomena, like within-species competition, ecological adaptation and evolution of the host, fitness advantage of bacteria to the host, the benefits of parasitism or mutualism, the domestication of the microbiome, etc. At the end, they conclude that their study "highlights that even subspecies diversity plays a key role in adaptation to environmental temperature". However, the potential mechanisms of such adaptation are barely discussed, so that the focus of the study shifts from the temperature-induced changes in microbial population structures toward metabolism-related adaptations of clade representatives that enable them to diversify their carbon and nitrogen sources. The role of the temperature factor remains elusive.

      We acknowledge that our study does not fully resolve the mechanism by which a different clade ends up dominating each temperature regime. The MRS liquid experiment was an attempt to answer whether differences in optimal growth temperature could explain the temperature-specific abundance of the two clades. Our experiments showed, however, that this was not the case. Beyond this point, it is hard to disentangle the role of the temperature, as it could also act indirectly on the bacteria, for example, through the host or the food.

      A second observation in our time series was that a third clade, U, was unfit in both regimes despite starting the experiment in high abundance. For this reason we also studied what made this clade less fit. Based on our analyses, we propose that the decrease of clade U was driven by the shift to a laboratory diet, shared by all experimental populations.

      In addition to that, the paper has a clearly minimalistic experimental approach to address functional properties of the revealed L. plantarum strains, so that their own fitness, or their relationship with the Drosophila host, is characterised superficially. Therefore, the authors' discourse can be speculative rather than factual (especially when the authors use the expression "likely" to share their guesses in the "Results" section). Nevertheless, these minor drawbacks do not underscore the novelty of the discovered phenotypes and the importance of their further investigation.

      We consider the reviewer's concern and toned down the phrasing when reporting our findings in the revised version of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) One solution to resolve the "competition" issue would be to check that the L. plantarum strains are established at similar or different titers in the in vivo work. Fly phenotypes can be sensitive to microbial load (Keebaugh, iScience, 2018), which might explain some of the counterintuitive in vivo results. In line 227, the authors mention that "bacterial load" is contributing to the magnitude of the effect, but I don't see the data reported anywhere. If this is from the in vitro assays, then the authors need to show that in vitro predicts in vivo L. plantarum abundance.

      Bacterial load inoculated in the in vivo experiment was normalized to OD=0.05 (~5*10<sup>6</sup> CFUs/ml) for the three clades at the beginning of the experiment. Thus, all vials were inoculated with the same titer of L. plantarum. Only the genotype varied between treatments. However, it is possible that, once inoculated, each clade grew to a different titer (as they have different growth rate and carrying capacity).

      Statement in line 227 comes from the differing results in transfers 1 and 2. In transfer 1 we inoculated a fixed load of ~2.5*10<sup>5</sup> CFUs. In transfer 2, however, we inoculated no bacteria to the food, and the flies seeded the vial. Our statement comes from the assumption that bacterial load in transfer 2 has to be lower than in transfer 1 as bacteria seed the vial solely by defecation of the parents.

      Following the reviewer's suggestion, in the revised version of the manuscript we have included a new experiment in which we quantified the bacterial load of each clade in individual flies.

      (2) Tracking the competitive outcomes is tricky, though it could be done with whole genome sequencing. An alternative would be to label the strains with fluorescent proteins (e.g., Obadia Current Biology 2018 has done this in Lactobacillus) and track fluorescence to better understand the results of the "mix" treatment in Figure 6.

      We appreciate the feedback of the reviewer, but consider this rather labour-intensive approach as an interesting option for future follow-up experiments.

      (3) That being said, my main concern with this is the "competition" claim. If the paper were reframed appropriately, this paper could still make an important contribution to the evolution of host-microbiome interactions, but the authors would need to consider what they can and cannot do with this interesting dataset.

      The "competition" claim comes from the changes in relative abundance observed in the time-series data, not from any of the follow-up experiments. Thus, we consider the use of the term "competition" appropriate.

      (4) The text on the figures is very small and hard to read.

      We increased the size of the text in all figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Have you conducted the in vitro culture experiments following the "cold" conditions?

      We have conducted the experiment in "cold" conditions, but with some modifications to the experimental settings, as the plate reader did not have cooling capacity. Instead, we grew a subset of the isolates (four per clade) in glass vials at constant 20 °C, and measured their OD twice a day. We have included the results in the revised version of the manuscript.

      (2) How many technical and biological replicates were measured for the in vitro culture experiments (Figure 5)? Please add this information to the figure legend and method.

      We measured the growth of four isolates from clade U, nine isolates from clade C and sixteen isolates from clade H. Each isolate was grown three times.

      We have included this information, as requested by the reviewer.

      (3) Making the labels in Figures 2, 3, 5, and 6 bigger would be helpful.

      We have increased the font size of all figures.

      Reviewer #3 (Recommendations for the authors):

      (1) Line 268: "Based on our results in experimentally evolved fruit flies, we propose that within-species competition, thus far largely overlooked, could contribute to ecological adaptation and evolution of the host". Overstatement should be avoided, since the evolution of the host was not directly studied here.

      Our results show that reproductive traits of the host differ upon colonization with each clade. Although we don't test the host's evolution, we speculate that flies differing in their offspring number and developmental time might differ in their overall fitness. Finally, we consider the Discussion section as the right place for speculation and development of hypotheses that can be tested in future work.

      (2) Line 258: "These differences do not explain the clade-specific selection, but reflect the different evolutionary histories of the clades". The temperature factor and its possible role in clade selection would be better discussed at least a little bit.

      In this paragraph we described potential metabolic differences between clades using comparative genomics. We did not find enrichment in a function or group of functions that could explain the different dynamics between clades H and C in the temperature regime.

      In the revised version of the manuscript we highlight that we did not find temperature-specific differences from this analysis.

      (3) Line 252: "...This could explain why clade U, which displayed a high growth rate and carrying capacity in liquid culture". The statement could be further developed with a caution. Even if the isolates that belong to the clade U are outcompeted by H or C, it should be noted that the strain U cannot be used as a true reference for fitness, since it could possess its hidden adaptive properties, not being simply "a loser". Such a hypothesis could explain the maintenance of this strain in the wild.

      We agree with the reviewer in that fitness is relative to the selective environment. Clade U is less fit than H and C in our specific experimental conditions, but it could outcompete them in other conditions, such as wild flies or MRS liquid medium. In the revised version of the manuscript we have rephrased this statement to clarify that we specifically refer to clade U's fitness under the new laboratory conditions.

      (4) In a cold environment, association with the clade C induces developmental delay and produces less progeny, which potentially allows the host to survive in case of harsh conditions and potential food limitation. Could the authors speculate and not exclude that this phenotype could be potentially adaptive? It would be curious to check in further studies whether flies associated with C strains are more stress-resistant, for example.

      We thank the reviewer for this alternative hypothesis. In our manuscript we used the Darwinian definition of fitness; reproductive success of an organism in the focal environment. And thus, both higher progeny per female and shorter developmental time would be beneficial in direct competition with other individuals. It is true that delayed developmental time, or less progeny, could be advantageous in specific cases. This could be the case for D. simulans inoculated with clade C. However, we consider that the fecundity levels observed in D. melanogaster upon inoculation with Clade C (average of 0.06 offspring/female*day in the cold) are too low to sustain a population.

      We have included this hypothesis in the Results section.

      (5) It would be highly recommended to add an experiment to complete the story by measuring the quantity of bacteria in the medium and in the flies. This will resolve the hypothesis (Line 785): "Thus, the ability to exploit this ubiquitous source of carbon and nitrogen could be very advantageous in the fly microbiome context, but would not affect the fitness in liquid culture".

      Following the reviewer's recommendation we included two additional experiments. We measured the bacterial load per fly in the native host, D. simulans, inoculated with the three clades. We also compared the clades' growth speed in solid fly food (without host). In the former experiment, we found similar bacterial loads upon inoculation with clades U and clade H. In contrast, in the latter we found delayed growth of clade U relative to H and C in the food. Thus, chitobiose consumption does not seem to provide an advantage in the fly gut to clade H. We attribute the fitness advantage of H and C to their advantage growing on the laboratory fly food, regardless of the host.

      Both experimental results have been included in the revised version of the manuscript, and the comparative genomics paragraph and discussion have been modified in consequence.

      (6) The chapter "Extended clade-specific differences in KEGG metabolic pathways" could be presented in the main text as it contains important results. These results are mentioned in the chapter "Functional divergence on the genomic level", which looks rather humble when it stands alone as it currently does.

      We appreciate the interest of the reviewer in this supplementary chapter. To keep the length of the manuscript digestible for a broad set of readers, we decided to only include in the main text the functional differences that could play a role in adaptation to the new laboratory environment.

      We consider that a full description of the metabolic differences between the three clades has to be published, as it might be relevant for researchers interested in L. plantarum metabolism. However, it does not fully follow the storyline, as the differences reported in the supplementary, such as nitrate respiration or synthesis of molybdenum cofactors, might not be involved in the clade-specific selection observed in the time series.

      (7) Line 773: "Clades C and H encode a shared genetic repertoire related to sugar/riboflavin metabolism that is lacking in clade U". This indeed allows us to hypothesise that the fixation of these clades in fly populations was due to their improved metabolic capabilities. However, the analysis of fitness shows similarity in flies associated with clades H and U, meaning that sugar/riboflavin metabolism in H does not provide an obvious adaptive trait to flies. Moreover, one could say that sugar metabolism in clade C is maladaptive not only for flies, but also for bacteria in liquid cultures. It is recommended to more clearly state the respective limitations of the study.

      Here we have to make a distinction between bacterial fitness and host fitness. The three clades differ in their (bacterial) relative fitness, as evidenced by the time-series dynamics (Figure 4). In the cited statement we hypothesize that a more versatile sugar metabolism repertoire could increase the bacterial fitness of clades H and C (relative to clade U) in the sugar-rich laboratory diet.

      This is independent of the fitness effect that L. plantarum could have in the host. Finally, as it was discussed in the recommendation 3, fitness is specific to the environment. Clade C is the least fit in liquid MRS in hot conditions, but the fittest in cold experimental conditions.

      (8) The authors should better explain why growth in MRS was not performed in a cold temperature regime to further support or refute the hypothesis that capacity and inflection time could partially explain the higher fitness of bacterial strains from clade U.

      We did not perform this experiment in cold conditions due to technical limitations of the plate reader, that does not have cooling capacity. Nevertheless, following the reviewers' suggestion, we have included in the revised version of the manuscript a new MRS growth experiment in cold-like conditions (constant 20 °C).

      (9) When mentioning that L. plantarum can "increase larval fitness of Drosophila melanogaster relative to germ-free flies" (line 196), the authors should specify in which specific conditions this phenotype was observed, and how relevant the mentioned phenotypes are to the current study.

      Following the reviewer's recommendation, we have modified the paragraph in order to clarify the conditions used in other papers and those used in our work. The references cited in this section (PMID: 21907145, 29290388 and 28062579) report that L. plantarum increases the host fitness in protein-poor diets (12 g/l of dried yeast or less), but not in high-protein diet (50 g/l of yeast or higher). Since our experimental diet contains an intermediate amount of protein (24.3 g/l of dried yeast) we were agnostic of whether L. plantarum would benefit the host or not in our conditions. Regarding the phenotypes, we chose two reproductive traits that are affected by changes in the microbiome according to the literature. Developmental time is directly affected by L. plantarum in the aforementioned papers. Offspring number is another fitness component affected by Drosophila microbiome (PMID: 30510004).

      (10) Provide a reference for line 205: "In axenic D. melanogaster none of the L. plantarum clades provided a fitness advantage to the host relative to germ-free controls, contrary to the effects reported in the literature". If the conditions were different from those in the studies referred to, then it would be of no use to compare the fitness advantage (for example, in Reference 24 another type of diet was used).

      Already covered in recommendation 9.

      (11) Please provide more context to this statement (Line 210): "The high content of dried yeast 24.3 g/l in the fly food used in our experiment likely provided already sufficient amounts of essential amino acids, which negated the growth-promoting effects of L. plantarum". It is not clear why amino acids are taken into account, and what the evidence is for the fact that the amount of essential amino acids was sufficient to abolish growth-promoting effects.

      The whole paragraph was modified in order to clarify the relationship between protein input and nutritional fitness benefit of L. plantarum.

      (12) Please provide measurements of bacterial quantity which would support the statement (Line 215): "The fitness reduction was stronger in the first transfer of flies, likely due to a higher bacterial load".

      Upon request of the reviewer, we have estimated the bacterial load per individual fly in D. simulans. Additionally, we have specified the CFUs inoculated in the vials in transfer 1.

      (13) Correct the typo (line 220): "However, the developmental time was significantly extended after inoculation with clade C at cold temperature (Dunn's test, p < 0.05 05 for all significant comparisons)".

      Done.

      (14) Specify more precisely the temperature conditions referred to in line 226: "In summary, we observed that clade C, which is dominant in the cold-evolved populations, decreases host fitness when axenic flies are inoculated". Does it decrease fitness both in hot and cold environments?

      For the axenic flies, we did find a decrease in fitness in both regimes, yes. We specified it in the revised version of the manuscript.

      (15) Please provide evidence for line 227, or otherwise rephrase it: "The magnitude of this effect varies depending on the environmental temperature, the bacterial load, and the presence of other microbial taxa".

      Novel evidence was provided regarding the role of bacterial load on host fitness.

      (16) Correct the following statement, so that it reproduces the results of the original work (reference 19, line 229): "In a low-protein diet, strains that were not isolated from Drosophila enhanced larval growth relative to germ-free individuals, whereas another Drosophila-associated strain did not have any effect".

      This statement was removed from the revised version. This reference was cited in the discussion to state that: " the nutritional symbiosis in L. plantarum is strain-specific".

      (17) Please provide a rationale for using KEGG Orthologs. Why was this database chosen as an appropriate one, even though it is known to be a non-exhaustive metabolomic resource?

      KEGG is a well-known metabolic database that is widely used in comparative genomics (PMID: 40177264) and built in state-of-the-art software for microbial ecology such as Anvi'o (PMID: 33349678). Other similar gene-to-function databases are less focused on metabolic pathways, such as COG or GO, or limited to specific enzymatic activities, like CAZy. Furthermore, the hierarchical organization of KEGG Orthologs in modules and pathways allowed us to map clade-specific orthologs to the broad metabolic context. For these reasons, we considered KEGG to be the best option for this analysis.

      (18) Line 250: "Therefore, we speculate that the ability to exploit this ubiquitous source of carbon and nitrogen in the lab-maintained fruit flies, could be a strong target of selection in the lab environment". This statement concludes the "Results" section but would be more appropriate for the Discussion section, since the authors do not provide any experimental evidence that could support this statement.

      We have modified this chapter, as covered in recommendation 5.

      (19) Line 276: "However, the intraspecific richness of L.plantarum in our flies was three times higher than that estimated in human gut microbiomes". Note that there are other recent studies which show the presence of several OTUs within L.plantarum isolates (for example PMID: 41484402).

      We thank the reviewer for the reference. We comment on it in the revised manuscript.

      (20) Line 287: "Our finding shows that the well-characterized nutritional symbiosis between Drosophila and L. plantarum depends on the bacterial genotype and cannot be generalized to the entire species". Note that such a conclusion has already been previously stated (for example, PMID: 30008290 and 28993620).

      We thank the reviewer for the references. Indeed, these papers show that some L. plantarum strains are beneficial for the host while others are neutral. Furthermore, as commented by Reviewer #1 in the public review, L. plantarum has been shown to reduce the host's fitness relative to axenic flies (Gould, PNAS, 2018).

      Our observations are novel in two ways. (1) The fecundity observed in D. melanogaster, 0.06 offspring/female/day in average, is lethal (in Gould et al. 2018 fecundity never decreased below 1 offspring/female/day). (2) Clade C outcompetes the other clades in the cold, despite being detrimental for the host.

      We have modified the Discussion to account for the previous work.

      (21) Line 343: "In addition, we obtained L. plantarum genomes from two other experimental evolution studies. Two genomes from the South African experiment and seven genomes from the Portugal experiment". Merge two sentences into one.

      Done.

      (22) Line 360: "At sampling, the age of the flies varied between four and eight days for the hot environment and between nine and 16 days for the cold environment". Please comment on the fact that different age of flies (different physiology) is not the reason for bacterial community differences.

      During maintenance, flies are sampled at different ages because the temperature affects their developmental time. We cannot rule out the hypothesis that age difference drives microbiome differences. Temperature could affect clade competition directly (differences in optimal temperature between clades) or indirectly, by affecting either the host (e.g. changes in Drosophila developmental time alters L. plantarum fitness), the surounding microbiome, or the food (e.g. increased metabolic activity in the hot regime changes nutrients profile). We ruled out the direct effect of temperature with growth experiments in liquid MRS medium and solid fly food, but disentangling the indirect effects is not feasible.

      (23) The majority of figures have low-quality labels that are not legible due to the small size of the font. Please improve.

      Done.

      (24) Figure 1 - Correct the legend: There is no "10" label on the picture. Probably by 10, the authors mean "Generation", while by x10 - number of isogenic replicates.

      Done.

      (25) Figure 2 - No numbers at nodes are indicated, whereas it is announced in the legend that they represent bootstrap support values. In addition, it is recommended to show a reference pangenome in the middle panel to clearly refer to the total size of the possible black bar.

      We added high bootstrap support as coloured nodes in figures 2, 3 and S2.

      We do not understand the reference pangenome request. In the middle panel, each black/white bar corresponds to an orthologous gene that can be either present or absent in each of the genomes. These orthologs were sorted based on hierarchical clustering of the their patterns of abundance (columns present in the same set of genomes, together), not by synteny. Thus, a reference pangenome would be simply a black bar.

      (26) Figure 2: It would be advantageous to add a figure that represents the frequency of each strain in each replicate (at the last time point, for instance). It would explain why some “blue” strains appear to be within the “red” cluster. Otherwise, it is confusing to find cold-evolved bacterial strains in hot-evolved fly populations.

      The frequency of each clade in each replicate is shown in figure 4. We think that it would be more confusing to follow the suggestion of the reviewer, as the isolates were sampled at different time points of the experiment. We would not like to call it a confusion that "blue" strains appear in the "red" cluster, but rather the logical consequence of the color code used in figure 2, which corresponds to the temperature regime in which the isolate was sampled (regardless of its clade). Whereas in the following figures colour represents the clade. It is thus possible to find clade H isolates in the cold temperature regime, as this clade is in low frequency but not completely absent in this regime.

      (27) Figure 3 – Add a label for the X-axis.

      We rotated the tree to be able to increase the genome IDs. We have added the label to the Y-axis.

      (28) Figure 4 - Please indicate how the clade relative abundance was assessed.

      Clade relative abundance was inferred by mapping competitively the short reads against the three clades’ reference sequences. It is specified in the legend now.

      (29) Figure 6 - Total number of F1 flies eclosed normalised by day (during which period?). What do T1 and T2 correspond to?

      During the respective number of days that females were allowed to lay eggs: one day in the hot settings and two days in the cold settings in transfer 1. One day and three days, respectively, in transfer two.

      T1 and T2 correspond to the first and second transfers, as described in the Materials and Methods. In first transfer, flies laid eggs in vials pre-inoculated with a set load of L. plantarum. After egg laying, same adults were then transferred to a sterile set of vials and allowed to lay eggs again (second transfer). Bacterial load in these vials was solely seeded by the parents.

      In order to avoid any confusion, in the revised version of the manuscript we have modified figure 6 to show transfer 1 for both Drosophila species, and moved transfer 2 dataset to supplementary figure S5.

      (30) Figure S2 - label the X-axis.

      We guess the reviewer means Y-axis. Done.

      (31) Figure S3 demonstrates the real data and its variability, so it would be better used instead of Figure 4 (which seems to be just a derivative from Figure S3, not a separate dataset and separate type of analysis).

      As the reviewer suggested, we have replaced figure 4 with figure S3.

      (32) Figure S4: Improve plot title: (e.g., C:H:U = 3:3:3).

      Done.

      (33) Figure S6: It is stated that N = 10; however, some datasets do not have 10 points represented. Please specify why. Also, please specify the meaning of "T1/T2".

      For the inoculation experiment in Drosophila simulans, we had nine replicates per treatment, not ten. This has been corrected in the figure and in the Materials and Methods section.

      T1 and T2 correspond to the first and second transfers, already covered in recommendation 29.

      (34) Table S3: provide legend for values (1 - present in all strains, but 0.04 - what does it mean?).

      It means that 4% of the genomes from this clade harbour the specific gene. We have specified it in the legend of the revised table.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Uphoff et al. propose a structural and mechanistic model in which the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, thereby promoting downstream receptor phosphorylation and signaling. Using AlphaFold-based modeling, the authors predict that the CCP20 domain of SVEP1 binds to TIE1, creating a composite surface that facilitates Angiopoietin association and TIE1 activation. The resulting ternary model (SVEP1-TIE1-ANG) offers a structural rationale for how SVEP1 converts TIE1 into a functional, ligand responsive receptor. Additional models and biological assays suggest roles for other domains of SVEP1, such as CCP5-EGF-L7, although these interactions are predicted with low confidence. The authors interpret these findings as the first structural framework for how SVEP1 enables ANG-TIE1 signaling.

      Strengths:

      (1) The central hypothesis - that SVEP1 enables ANG binding to the orphan receptor TIE1 - is biologically compelling and addresses an important question in vascular biology.

      (2) The AlphaFold-predicted ternary complex (SVEP1-TIE1-ANG) is plausible, high-confidence, and structurally consistent with prior functional data (e.g., poly-Ala scanning from Sato-Nishiuchi et al.).

      (3) The authors' model offers a potential explanation for the previously observed role of SVEP1 in enhancing ANG signaling through TIE1, and may represent the first structural insight into TIE1's transition from orphan to ligand-activated receptor.

      (4) The potential clinical implication - that a combinatorial ligand (ANG+SVEP1) can activate TIE1- could have translational relevance for vascular leak and inflammatory disease.

      Weaknesses:

      (1) Lack of structural validation and mechanistic follow-up: Despite the promising AlphaFold model, there are no figures of the predicted interface, no residue-level interactions shown, no ipTM values reported, and no experimental follow-up to test the interface. PAE plots are incorrectly used as confidence justifications, which is not appropriate for complex predictions.

      We have appended the data showing AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity. We also added ipTM scores and confidence plots for the predicted complexes.

      (2) Biophysical validation is missing: No surface plasmon resonance (SPR), ITC, or biochemical assays are included to confirm ternary complex formation or quantify binding kinetics. Given the manuscript's structural focus, this is a major gap. For instance, an SPR experiment where ANG is immobilized, and TIE1 binding is measured {plus minus} SVEP1, would directly test the model. And allow direct comparison to ANG-TIE2.

      We have addressed this question and performed ELISA assays to measure binding affinities between SVEP1 and TIE1 in presence or absence of ANG1 or ANG2, thus confirming that the affinity is increased in the presence of ANG1 or ANG2.

      (3) Missed opportunity for mutagenesis-driven validation: The manuscript does not include any interface-targeted mutations, despite clear opportunities. For example, mutating T2595 in SVEP1 (to R) or mutating the TIE1-specific residues (residues PL 202-203 to LF) could strongly test the model and potentially reveal dominant-negative behaviors. E.g. A T2595 mutant should block ANG binding but not TIE1 binding.

      We have depicted figures of the interfaces including P202-L203 and included the TIE1 P202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study, for the following reason: Modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2, thus reducing its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both TIE1 and ANG2, the data is not included in the manuscript.

      (4) Overinterpretation of weak models: The additional AlphaFold model involving the CCP5-EGFL7 domains binding TIE1 has extremely low confidence (ipTM < 0.15) when reexamined by this reader and should not be emphasized. There is no biophysical evidence or binding data (SPR) to support this interaction, and its inclusion detracts from the much stronger CCP20 model.

      We agree with this point made by both reviewers and have removed the data on CCP5-EGFL7 from the manuscript.

      (5) Language around modeling is overstated and potentially misleading: Terms like "unequivocal," "high-affinity," or "affirms strong binding" in reference to AlphaFold predictions are inappropriate. These are hypotheses -not confirmations - and must be tested at the biochemical level. This should be clarified throughout the manuscript to ensure non-experts do not misinterpret modeling confidence as binding affinity.

      We agree with the reviewer, and have adjusted the wording.

      (6) Negative stain EM data is not informative due to low resolution and lack of defined interfaces; unless replaced by higher-resolution Cryo-EM, this should be omitted. Better would be co-gel filtration, AUC, or SEC-MALLs with ANG-SVEP1-TIE1.

      We have now appended the data by adding gold-labelled TIE/ANG proteins, thus enhancing clarity.

      (7) Disjointed narrative: The manuscript presents a compelling mechanism involving CCP20-driven ANG binding to TIE1, but then becomes fragmented by introducing the low-confidence CCP5-EGFL7 model and speculative higher-order polymerization models that are not experimentally supported.

      We agree with this point and have have centered the manuscript around CCP20. We removed data concerning CCP5-EGFL7 as suggested by both reviewers.

      Reviewer #2 (Public review):

      Uphoff and colleagues present the results of a study focused on characterizing the binding of SVEP1 to TIE1 along with Angiopoietin-2. Starting with computational prediction of SVEP1 binding to TIE1, the authors identify the region of SVEP1 that serves as a high-affinity ligand for TIE1. Advanced studies identify a weak secondary binding site within SVEP1 that appears to be sufficient but not necessary for its interaction with TIE1 based on in vivo rescue experiments. The most novel contribution of the manuscript seems to be the identification of angiopoietin-1 and -2 as co-factors that seem to enhance the binding of SVEP1 with TIE1 and impact downstream AKT signaling. They propose a complex in which SVEP1 binds to TIE1 and ANG2.

      Although the first set of results is essentially confirmatory, the identification of ANG-2 as a "cofactor" enhancing the binding of SVEP1 to TIE1 and associated downstream signaling (i.e., Figures 3 and 4) is novel and is of interest. However, the manuscript and its conclusions would greatly benefit from some clarifying details and additional experiments to ensure rigor and support specific claims.

      We have addressed the reviewers concerns and significantly appended the manuscript. Most importantly, we provide structural validation of AlphaFold models reporting interfaces, residue-level interactions and ipTM values. We have included new biophysical validation of binding kinetics of SVEP1 and TIE1 in the presence or absence of ANG1 or ANG2. Furthermore, we have removed the data on CCP5-EGFL7 from the manuscript in order to retain focus on the CCP20 domain.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The AlphaFold modeling for the CCP20-based interactions is strong (as determined by this reader rerunning and getting ipTM values and visually inspecting the interactions “because this is not in the manuscript”). As presented, the manuscript stops at the hypothesis-generation stage. Validation is needed to fulfill the paper's title and claims. The structure-function link is not demonstrated, despite an obvious and achievable experimental path (mutagenesis, SPR, kinetics).

      Additional Context and Suggestions

      (1) Show and label AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity.

      We have appended the data in the new supplementary figures 1.1, 1.2, 2.1 and 2.2

      (2) Provide ipTM scores and confidence plots for each predicted complex.

      We have added the values and plots in the new supplementary figures.

      (3) Perform SPR assays with ANG-coated surfaces and measure binding of TIE1 {plus minus} SVEP1. Compare to TIE2 binding for context.

      We performed the proposed experiment using an ELISA assay to measure binding affinities between SVEP-1 and TIE1 in presence or absence of ANG2 and included these data in the manuscript in figure 2.

      (4) Test interface mutants: e.g., T2595R in SVEP1 (should impair ANG recruitment but not TIE1 binding), or PL→LF muta on in TIE1 (should disrupt SVEP1 binding).

      We have depicted figures of the interfaces including P202-L203 and included the TIEP202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study. Our modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2 thus reduce its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both, TIE1 and ANG2, the data is not included in the manuscript.

      (5) Clarify in the Introduction that SVEP1 is a large, multidomain ECM protein to help readers contextualize the domain names early on. Do not use terms like CCP before defining them.

      We added: “Svep1 encodes a 3571 amino acid long extracellular matrix protein containing different domains such as Willebrand factor type A domain (vWF), ephrin-receptor like domains, complex control protein (CCP) domains, and Hyalin repeats at the N-terminus. The C-terminus mainly consists of CCP and EGF domains. Svep1 is expressed in mesenchymal cells, but not in endothelial cells, and functions non-cell-autonomously (Karpanen et al. 2017; Morooka et al. 2017)”.

      (6) Replace or remove negative-stain EM unless higher-resolution cryo-EM data are available.

      We would like to retain the EM data, but have now replaced the negative-stain EM by adding gold-labelled TIE/ANG Proteins to verify that the proteins we show are the ones we expect. The reason we would like to retain the data is that TIE1 has been considered for so many years as an orphan receptor, and thus we consider it appropriate to demonstrate SVEP1/TIE1 binding using multiple methods.

      (7) Reframe claims around modeling to avoid overstatement. For example: "The model suggests a plausible mechanism consistent with prior biochemical data" is more appropriate than "unequivocal".

      We agree with the reviewer, and have adjusted the wording.

      (8) Consider narrowing the focus: the CCP20-TIE1-ANG model is a strong story on its own. The CCP5EGFL7 model and polymerization hypotheses are not essential and may dilute the impact.

      Since this point was made by more than one reviewer, we have removed the data on CCP5EGFL7 from the manuscript.

      (9) Properly define TIE1: Tyrosine kinase with Ig and EGF domains.

      We have corrected the full protein name for TIE1 and added: “TIE1 and Tie2 exhibit a high degree of homology with a globular head domain consisting of three immunoglobulin-like (Ig) domains and three epidermal growth factor-like (EGF) modules and a short stalk formed by three fibronectin type III repeats, while the N-terminal two Ig domains of Tie2 harbor the angiopoietin binding site (Macdonald et al. 2006).” We also added: “D1 and D2 refer to the two N-terminal Ig domains, D3 refers to the three EGF domains and D4 to the third Ig domain of TIE1 or TIE2.”

      (10) Use RTK, not tyrosine kinase receptors TKR.

      Tyrosine Kinase receptor was replaced by receptor tyrosine kinases (RTKs)

      (11) PDBs (.cif and .json files) of the models must be supplied for the readers so they don't need to rerun the AlphaFold jobs.

      We are providing all PDBs with the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) In several locations, the authors state that alphafold detects an "unequivocal" and/or "high affinity" interaction. Experts on computational structure prediction can weigh in, but I am not sure it is accurate to say that alphafold predicts affinity. Quantitative estimates of the prediction confidence or other parameters of the alphafold output are not provided.

      Thank you for the comment, we agree with the reviewer and have adjusted the wording. We have also appended the data showing interfaces, ipTM scores and confidence plots for the predicted complexes.

      (2) In Figure 1C, what concentrations of TIE1 and TIE2 are being used in the SPR experiments shown?

      We have added the concentrations of the proteins to the methods section

      (3) In Figure 1C, what is the affinity constant (KD) of the interaction between SVEP1 and TIE1 and SVEP1 and TIE2?

      We have added the values of affinity constants to the manuscript text.

      (4) In Figure 1C, the authors immobilize a "70kD" fragment of human SVEP1 to determine the interaction between SVEP1 and TIE1. How does the affinity they measure between the SVEP1 fragment and TIE1 compare to the affinity between immobilized full-length SVEP1 with TIE1?

      The largest molecule we used in any assay is not full-length SVEP1, but consisted of a C-terminal SVEP1 protein spanning from the first EGF domain to the C-terminus (approximately 295 kDa) as previous studies have shown that SVEP1 is proteolytically cleaved N-terminal to the first EGF domain, generating a protein of this size. In our ELISA assays, this molecule has a higher affinity to TIE1 than the 70 kD fragment. We have included data for the 70 kDa as well as for the 295 kDa fragment in the manuscript. ELISA assays have used larger SVEP1 fragments (as indicated in the figure legends), the SPR assay was performed with the 70 kDa fragment.

      (5) How does the alphafold structure prediction for the "70kD" fragment of hSVEP1 compare to the prediction of the same 70kD fragment from the full-length protein prediction?

      All modellings using different SVEP1 fragments including the 70 kD and full-length version identify the putative binding site at position CCP20. While AlphaFold3 can predict the correct domain folds in both the 70kDa fragment and full-length SVEP1, the orientation of these domains is highly variable due to flexible linkers between each domain module. This flexibility effects the output confidence metrics, thereby hampering our interpretation of the models. Therefore, we conducted the structural prediction of the complexes with smaller fragments and not the full-length SVEP1.

      (6) What are the amino acids for the 70kD fragment?

      The relevant amino acids are 2261-2890. The accession number and amino acids of each protein have been listed in the key resources table.

      (7) In Figure 1D, what is being depicted by the red stars? This is not explained in the text or figure legend.

      We have replaced figure 1D.

      (8) By itself, Figure 1D is not terribly informative and in my opinion does not support the statement that the authors "were able to directly visulalize the attachment of SVEP1 and TIE1." As a minimum, the authors should repeat the same set of images with SVEP1 and TIE2, but other approaches, such as labeling, could be performed.

      We replaced figure 1D with new data and gold-coated protein enhancing clarity. We think it beneficial to demonstrate SVEP1/TIE1 binding using multiple methods as TIE1 has been considered as an orphan receptor for so many years. We have not performed these experiments with TIE2 proteins as we were not able to show binding of TIE2 to SVEP1 with other assays.

      (9) What is being stained in Figure 1D? Full-length SVEP1/TIE1? Or fragments of these proteins?

      We replaced figure 1D by a new assay with labeled proteins using the 150 kDa version of SVEP1 and the ectodomain of TIE1 as well as ANG1 or ANG2 (new figure 1D and new supplementary figure 2.3). TIE1/ANG proteins were gold-labelled. The protein fragments used in this assay are described in detail in the methods section.

      (10) The authors discover CCP6-EGFL7 as a low-affinity binding region of SVEP1 for TIE1. Is this region in physical proximity to CCP20 (the high-affinity binding region for TIE1) based on alphafold prediction? How would the authors think this is binding TIE1?

      We have removed this data set (see comment to reviewer 1’s request).

      (11) What is the affinity constant (KD) for CCP6-EGFL7 with TIE1?

      We have removed this data set (see comment to reviewer’s 1 request).

      (12) The authors claim that ANG1/ANG2 increase affinity between SVEP1 and TIE1 based on immunoblotting. Immunoblots are semi-quantitative at best. If the claim is higher affinity, I think the authors should measure this by SPR and determine the KD between immobilized SVEP1 with TIE1 in the absence and presence of ANG1 and/or ANG2.

      We conducted ELISA assays (figure 2) showing that the affinity is increased in the presence of ANG1 or ANG2 and agree with this reviewer that this strengthens the data.

      (13) In Figure 3, can the authors explain why ANG1/2 does not pull down with SVEP1/TIE1?

      We noticed that upon transfection of TIE1 into HEK cells, ANG1/2 is almost undetectable anymore in the total lysate. Thus, we believe that after the pull down we are below the detection limit.

      (14) In Figure 3, what is "TL"? I assume total lysate, but this is not specified.

      Thank you, we now specify TL as total lysate.

      (15) In Figure 3 "TL" panel (again, I assume this is total lysate), why are the ANG1/2 immunoblots so weak when co-transfected with TIE1?

      We consider it likely that in the presence of TIE1, ANG1/2 proteins are internalized and digested. Another reason for low signals could be that upon transfection of two plasmids, the amount of ANG1/2 protein is reduced as the cell has limited capacity for transcription and translation.

      (16) In Figure 3B, why is the SVEP1 fragment now 150kD when 70kD fragment was previously used? What domains are contained in this 150kD fragment?

      We now better define the domains of the 150kD SVEP1 fragment. The 150 kD fragment was the one produced first and available in high amounts in our laboratory and thus used for functional assays. The 70 kDa fragment together with ANG2 also induces phosphorylation of AKT, but it was not used in as many conditions/replicates as the amounts we had available were lower.

      (17) In Figure 4A, signaling with SVEP1 by itself and ANG2 by itself should be shown to support the claims being made.

      We added the lines for SVEP1 and ANG2, and also the quantification. SVEP1 itself already affects the phosphorylation of TIE1, most likely because hdLECs produce ANG2 by themselves. We show this with the ANG2 blocking antibody for pAKT.

      (18) For pAKT, what are the concentations of proteins being used and the times of incubation?

      This information is provided in the Materials and Methods section. We added the concentration of the anti-ANG2 antibody, which was missing.

      (19) It seems that p-AKT and AKT are being blotted on different gels. If this is correct, loading controls need to be shown for p-AKT blot.

      We added HSC70 as a loading control for both blots.

      (20) It is interesting that anti-ANG2 antibody inhibits SVEP1-induced p-AKT signaling. As the authors may know, SVEP1 has been identified as a receptor for PEAR1, which also leads to downstream p-AKT signaling, which seems to be independent of ANG2. Do the LECs being used here express PEAR1? If these cells express PEAR1, how do the authors think ANG2 silencing will eliminate SVEP1-associated p-AKT signaling?

      hdLECs express PEAR1. However, we show that p-AKT signaling is attenuated after siRNA KO of TIE1. Thus, the downstream signaling is dependent on TIE1 (Figure3).

      (21) Again, experts on computational structure prediction can weigh in, but I am not sure how to interpret the prediction of the 2:2:2 stoichiometry for the theoretical SVEP1/TIE1/ANG1-2 complex. Are there quantitative estimates of the confidence that can be provided? Did the authors attempt to model this complex with different stoichiometries? It is difficult to know how relevant this model is without any experimental results supporting this result.

      Since 1:1:1 is the smallest possible triple complex, it is our starting point. We can model a 2:2:2 version, but anything larger than this AlphaFold will not run. Furthermore, we now provide quantitative estimates of confidence with the pLDDT, PAE, pTM, ipTM scores for all models including the 2:2:2 complexes.

      Minor comments:

      (1) The authors could consider including a reference to alphafold on line 102.

      We have added a reference for AlphaFold2 and 3

      (2) The authors should refer to surface plasmon resonance (SPR) assays by this term as opposed to using the brand name Biacore.

      We agree with the reviewer and have changed the term Biacore to SPR.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the reviewers for their time and for their valuable inputs throughout the review process. We wish to clarify, one final time, the primary scope and empirical grounding of our work for prospective readers.

      Our study was designed to evaluate whether microsaccades track (in a correlative manner) covert visual-spatial attentional shifting, maintenance, or both. We did so within a single dedicated paradigm, across a large sample (N = 48 human participants). Despite remaining criticisms concerning per-participant event counts and microsaccade classification criteria, the key observation remains a striking dissociation (of the link between microsaccades and covert attention) during the initial shifting and the subsequent maintenance of visual-spatial attention. Moreover, we note how the robust effect observed during shifting (but not maintaining) attention, mitigates residual concerns regarding microsaccade sparsity or signal-to-noise ratio.

      We thus remain confident in the empirical foundation of our work and we invite readers to examine the full paper, supplementary materials, and open-access data to evaluate these findings independently.


      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the editors for their careful evaluation of our article and for their valuable input. Building on these suggestions, we were able to further corroborate our main conclusions, make our article more comprehensive, and thereby substantially strengthen the manuscript.

      We have one additional point of our own: we noticed that in our original submission, we had smoothed the data more than intended. Having caught this, we have now reduced the smoothing employed by 2.5 times compared to the original amount of smoothing (the exact smoothing values have also been added to the methods section). Importantly, however, while this has affected how the results look, this has not affected any of our original conclusions.

      General summary

      We would like to first respond to the major points brought forward by both the editorial summary and the public reviews. As we understand, the two main points that were raised regard: (1) the novelty and theoretical importance of our work and (2) the (in)completeness of our results. We start by providing our response to both of these main points below.

      Novelty and theoretical relevance of the work

      Regarding the novelty of our work, we believe the reviews and, by extension, the editorial summary underappreciated the main theoretical value of the question we addressed. Our work set out to investigate whether microsaccades track covert attentional shifting, attentional maintenance, or both. We fully recognise that there are ample prior studies that investigated and reported a link between microsaccades and covert attention, but also underscore how other studies report seemingly contradicting evidence by reporting that there is no such link. One such example is a recent paper by Willett & Mayo in PNAS (2023). Prompted by the recent hypothesis that this seemingly conflicting evidence may be due to prior work investigating attention ‘in different stages’ (van Ede, PNAS, 2023), we set out to address precisely this using a dedicated task that we designed for this purpose. As acknowledged by the summary and public reviews, this helps to reconcile seemingly opposing views in the literature. In our view, such reconciliation has substantial theoretical value.

      While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’). To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. 

      Finally, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Completeness of results

      Regarding the completeness of our results, the editorial summary points to “the absence of independent measures, single-trial analyses, and neutral-condition controls needed to substantiate the central claims”. While the raised points are valuable, they pertain to issues that are tangential to our primary question and stem from unfortunate misunderstandings of key analytical choices, as we now better clarify. We consider our results complete and comprehensive with regards to the main question our studies set out to answer.

      First, regarding the portrayed “need” for independent measures to define the ‘shift window’ of interest, we clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research. 

      Second, regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the single-trial level in our discussion, and prompted by the valuable feedback, we have expanded on this important contextualisation as part of our revision.

      Finally, regarding the portrayed “need” for a neural-attention control condition, we agree that inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ also at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that challenges our main conclusions. Having clarified this, we embraced this valuable reflection and revised the article to ensure that we do not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation, but refer to ‘the presence of an attentional modulation’ instead.

      Taken together, the explicit design and analysis choices that we made align with the theoretical aims of our study, and the central question we set out to address. The raised points are valuable and we are grateful to have been able to leverage them to improve our article, but we hope to have also clarified how they do not render our findings “incomplete” (as currently portrayed) with regards to the key goal of our article.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining the relationship between microsaccades and covert attention. This question has been widely investigated, with numerous studies showing that during sustained fixation, when subjects covertly attend to a peripheral stimulus, microsaccades tend to be biased toward the attended location. Here, the authors ask whether this microsaccade bias reflects a shift of covert attention or the maintenance of covert attention. They conclude that the bias is primarily driven by attention shifts, a finding that also helps reconcile the seemingly conflicting results of prior research, where the bias was questioned in paradigms that largely involved attention maintenance rather than shifting.

      Strengths:

      The paradigm and conclusions appear sound and supported by the results. A large sample size was used.

      We thank the reviewer for this clear and supportive summary of our work.

      Weaknesses:

      Weaknesses are mostly related to how the authors enforced fixation in the task, and clarifications are needed regarding some methodological details. A more direct comparison of the effect in the two experimental conditions is missing.

      We thank the reviewer for raising these valuable points. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”). Regarding the fixation points, we will address them in our point-by-point replies below.

      Reviewer #2 (Public review):

      Summary:

      This study aims to test the hypothesis that microsaccades are linked to the shifting of spatial attention, rather than the maintenance of attention at the cued location. In two experiments, participants were required to judge an orientation change at either a validly cued location (80% of the time) or an invalidly cued location (20% of the time). This change was presented at varying intervals (ranging from 500 to 3,200 ms) after cue onset. Accuracy and reaction times both showed attentional benefits at the valid versus invalid location across the different cue-target intervals. In contrast, microsaccade biases were time-dependent. The authors report a directional bias primarily observed around 400 ms after the cue, with later intervals (particularly in Experiment 2) exhibiting no biases in microsaccade direction towards the cued location. The authors argue that this finding supports their initial hypothesis that microsaccade biases reflect shifts in attention, but that maintaining attention at the cued location after an attention shift is not correlated with microsaccade direction.

      Strengths:

      The results are straightforward given the chosen experimental design. The manuscript is clearly written, and the presentation of the study and its visualisations are both of a high standard.

      We thank the reviewer for this clear summary of our work.

      Weaknesses:

      The major weakness of this paper is its incremental contribution to a widely studied phenomenon. The link between attention and microsaccades has been the subject of extensive research over the past two decades. This study merely provides a limited overview of the key insights gained from these papers and discussions. In fact, it attempts to summarise previous work by stating that many experiments found a link, while others did not, and provides only a relatively small number of references. To make a significant contribution, I believe the authors should evaluate the field more thoroughly, rather than merely scratching the surface.

      We thank the reviewer for this valuable reflection. For an elaborate response to the perceived novelty, please see our general summary reply above. In addition, we have added a more thorough evaluation of the field to the introduction (page 2, find relevant paragraph below). We hope that this will provide more context for the manuscript and strengthen its contribution to the field.

      Revised paragraph from introduction:

      “This link between microsaccades and covert visual-spatial attention has been demonstrated repeatedly. Early studies linked the direction of microsaccades to the deployment of covert attention [14, 15] and these findings were later replicated and extended. For example, it has been demonstrated in both humans [14–25] and non-human primates [26, 27]; during both externally directed perceptual attention [14, 15, 17, 20, 21, 24–27] and internally directed attention within visual working memory [16, 18, 19, 22, 23]; and in both perception and action tasks following directional cues [28]. Several studies have further linked the directional microsaccade bias to task performance [18, 21, 25, 28, 29]. For example, following spontaneous microsaccades, perception of visual targets presented in the same direction is better [25], and visual discrimination benefits may start already prior to microsaccade execution [21]. Recent evidence further suggests that microsaccades may even play a causal role in shaping the perception of peripheral stimuli [30].”

      The authors then present a potential solution to the conflicting past findings, arguing that attention should be considered a dynamic process that can be broken down into an attention shift and a sustained attention phase. Although the authors present this as a novel concept, I cannot think of anyone in the field who considers spatial attention to be a static entity. Nevertheless, I was curious to see how the authors would attempt to determine the precise timing of the attention shift and manipulate the different stages individually. However, the authors only varied the interval between the onset of the attention cue and the test stimulus, failing to further pinpoint their dynamic attention concept.

      The current version of the experiment, therefore, takes a correlational approach, similar to initial studies by Engbert and Kliegl (2003) and Hafed and Clark (2002). Meanwhile, we have learned a great deal about the link between microsaccades and attention. Below, I will list just a few of these findings to demonstrate how much we already know. It is important to note that, while the present study cites some of these papers, it does not provide a clear overview of how the current study goes beyond previous research.

      (1) Yuval-Greenberg and colleagues (2014) presented stimuli contingent on online-detected microsaccades. A postcue indicated the target for a visual task, and the target could be congruent or incongruent with the microsaccade direction. The authors showed higher visual accuracy in congruent trials. The authors cited that paper, but it is still important to emphasize how this study already tried to go beyond purely correlational links on a single trial level.

      (2) The Desimone lab (Lower et al., 2018) showed that firing rates in monkey V4 and IT were increased when a microsaccade was generated in the direction of the attended target.

      (3) However, attention can modulate responses in the superior colliculus even in the absence of microsaccades (Yu et al., 2022)

      (4) Similarly, Poletti, Rucci & Carrasco (2017) observed attentional modulations in the absence of microsaccades, or comparable attention effects irrespective of whether a microsaccade occurred or not (Roberts & Carrasco, 2019).

      Thus, in light of these insights, I believe the current study only adds incrementally to our understanding of the link between microsaccades and spatial attention.

      We thank the reviewer for this insightful comment, and for pointing out several important studies on this topic. While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’).

      To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. Moreover, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly and appropriately address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the singletrial level in our discussion. Prompted by the reviewer’s valuable feedback, we have expanded on this important contextualisation in our discussion section (page 8: “Therefore, even if microsaccades may more reliably track shifting than maintaining attention, as our current findings show, our findings should not be taken as evidence that microsaccades reliably track attentional shifts at the single-trial level.”). We also incorporated the valuable reference suggestions in our revised manuscript, including in the revised paragraph in our introduction where we provide a more extensive overview of the prior literature, as shown in response to the preceding comment and in the discussion where we discuss the relationship between microsaccades and attention on a single-trial level.

      In general, it is important to have an independent measure of the dynamics of an attention shift. I think a shift of 200-600 ms is quite long, and defining this interval is rather arbitrary. Why consider such a long delay as the shift? Rather than taking a data-driven approach to defining an interval for an attention shift, it would be more convincing to derive an interval of interest based on past research or an independent measure.

      We thank the reviewer for their question. We wish to clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research.

      The present analyses report microsaccade statistics across all trials, but do not directly link single-trial microsaccades to accuracy. Similarly, reaction times and accuracy were analyzed only with respect to valid vs. invalid trials. Here, it would be important to link the findings between microsaccades and performance on a single-trial level. For instance, can the authors report reaction times and accuracy also separately for trials with vs. without microsaccades, and for trials with congruent vs. incongruent microsaccades?

      We thank the reviewer for their sincere interest in our findings and for the great suggestion of an additional analysis. We have now investigated whether trials with a congruent, incongruent or no microsaccade in the shift window (where congruent or incongruent was determined as based on the first microsaccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to stress that our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we have decided not to include these analyses.

      The study would benefit greatly from including a neutral condition to substantiate claims of attentional benefits and costs. It is highly probable that invalid trials would also demonstrate costs in terms of reaction times and accuracy. It would be interesting to observe whether directional biases in microsaccades are also evident when compared to a neutral condition.

      We thank the reviewer for this valuable reflection. We agree that the inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial-numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that in any way challenges our main conclusions.

      Prompted by this reflection, we have ensured to not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation throughout the article, but refer to this only as ‘the presence of an attentional modulation’ instead (such changes were made on pages 3 and 9, and we kept this phrasing consistent in our additions on pages 5 and 22).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The results resemble recent findings by Brandolani et al. (2025), who also showed that the microsaccade bias-in a similar task and using a comparable analysis approach-was restricted to a narrow time window, primarily around the time of the attention shift. The authors should reference this work, discuss similarities and differences, and, given their larger sample size, consider whether they observe a similar correlation between response times and microsaccade rate.

      We thank the reviewer for pointing out this useful reference to us. We have included this article in our introduction, when sketching the current state of the field (page 2) and also mentioned this related article in the discussion (page 8).

      In addition, prompted by this comment, we have investigated whether we observe a similar correlation between response times and microsaccade rate in valid trials. We have investigated this for both possible timeframes: (1) the ‘shift’ timeframe and (2) the ‘maintain’ timeframe. For both experiments, there was no consistent correlation between response times and overall microsaccade rate, as shown in Author response image 1.

      Author response image 1.

      Relationship between the reaction time and the overall saccade rate. This figure shows the relationship between the average reaction time and the average overall saccade rate for both the ‘shift’ period (from 200 to 600 ms after cue onset) and the ‘maintain’ period (from 600 to 1400 ms after cue onset). Each dot represents one participant. Throughout the entire figure, the following significance levels were used: *: p< 0.05, **: p < 0.01, ***: p < 0.001, ****: p < 0.0001.

      (2) I could not find information on the average number of trials per condition and the average number of microsaccades per subject per condition. Ideally, these numbers should be reported (e.g., in a supplementary table). Since the analysis is based on microsaccade direction, knowing how many microsaccades each subject contributed per condition is critical. Microsaccade rates vary substantially across individuals, and subjects with very few events may add noise to the analysis, as proportions of toward/away microsaccades become unreliable.

      We thank the reviewer for pointing out that this relevant information was missing. We have now included these numbers in a supplementary table as suggested (page 17).

      (3) Relatedly, it was unclear whether the time-course analyses were based on collapsing all microsaccade events across subjects or on subject-level averages. In the Methods, the authors state that "the permutation distribution of the largest cluster size was acquired by randomly permuting the trial-average data at the group level 10,000 times," but this is ambiguous. Please clarify.

      We thank the reviewer for pointing out this ambiguity. We have changed the methods section to reflect more clearly that we first obtain time-courses of the microsaccade rate per participant, and subject these time courses to second level statistics using cluster-based permutation analysis (page 13). We have also reworded the sentence you quoted to remove any ambiguity (page 13): “A permutation distribution of the largest cluster size was acquired by randomly permuting the condition labels of each participant’s trial-averaged time course data (i.e. randomly flipping the sign of the difference in rate of toward vs. away saccades) 10,000 times and identifying the size of the largest clusters observed in these randomised data after each permutation.”

      (4) The authors analyze only downward microsaccades, but the cutoff definition is not specified. Presumably, this includes all directions between 180{degree sign} and 360{degree sign}, which may also include nearly horizontal events. This should be clearly stated. In addition, the rate of upward microsaccades should still be shown, divided into up-left and up-right quadrants to parallel the toward/away analysis. This would provide informative context on whether upward microsaccade rates change systematically over time.

      We thank the reviewer for pointing out that this information was missing, and for suggesting this valuable additional analysis. In the methods section, we now explicitly state the angular cutoffs used for the main analysis (page 12). Additionally, we have added a supplementary figure that shows the time course of upwards microsaccades over time (page 21), please see Supplementary Figure S4.

      (5) If the dataset contains enough microsaccadic events per subject, it would be useful to test more conservative angular cutoffs for defining "toward" versus "away."

      We thank the reviewer for this insightful suggestion. We have included an additional analysis, where only microsaccades were included with a direction within a 45° angle around the exacttoward and exact-away directions. This replicated our main finding. The results from this analysis are now included in the supplementary materials (page 24), please see Supplementary Figure S8.

      (6) Figure 2C: It is unclear what the "Center" and "Border" lines represent. The figure is also potentially confusing because it shows microsaccade amplitudes rather than landing positions. Small amplitudes may still bring gaze close to the target; this distinction should be clarified.

      We thank the reviewer for pointing this out. We have changed the “centre” and “border” labels to include more information (pages 6 and 19). We have also added an in-text clarification of the distinction between saccade amplitude and landing position (page 5: “Note that Figure 2C does not show saccade landing positions. While it is theoretically possible for multiple small unidirectional saccades to lead to a larger change in gaze position, a complementary analysis shows that fixation was maintained during the period of peak microsaccade rate in both experiments (see Supplementary Figure S6).”). In addition, in response to the related comment below, we have now also added heatmaps of gaze showing that gaze overall remained close to fixation in our tasks.

      (7) From the Methods, it appears that in Experiment 1, there was no automatic criterion for discarding trials in which gaze deviated from fixation. In Experiment 2, trials were terminated if gaze left a 2{degree sign} window, but given that the target was only 5{degree sign} from fixation, this seems a relatively loose criterion. It would be important to show the distribution of gaze positions during the task to assess whether fixation control was adequate.

      We thank the reviewer for this great suggestion. We have now added a figure to the supplementary materials (page 23) that shows the probability density of gaze position throughout the ‘shift’ period, for left cued trials and right cued trials separately, please see Supplementary Figure S6. We hope that this will further show that even in Experiment 1, fixational control was successful. We also show the difference between left cued and right cued trials, which again shows a gaze bias towards the cued item.

      (8) Did the authors examine whether there was a response time benefit (e.g., RT in congruent microsaccade trials minus RT in incongruent microsaccade trials, as in Brandolani et al., 2025) or an accuracy benefit when microsaccades were directed toward the target?

      We thank the reviewer for their interest in our findings and for the great suggestion of an additional analysis. As we discussed also in response to the general summary from reviewer #2 above, we have now investigated whether trials with a congruent, incongruent or no saccade in the shift window (where congruent or incongruent was determined as based on the first saccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to again stress how our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we decided not to include these analyses. However, please note that we did now include the outcomes of another analysis that more directly targeted the relation between the spatial modulations in microsaccades and task performance, as we turn to below. 

      (9) Was there a relationship between the size of the attentional effect and the magnitude of the microsaccade bias?

      We thank the reviewer also for this insightful question. We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (10) The criteria for minimum microsaccade amplitude and duration are not specified. This should be clarified. I recommend excluding events smaller than ~5 arcmin, as these are likely noise-especially since eye tracking was monocular. Monocular "microsaccades" can be spurious, but this can be determined only with binocular tracking. It is also unclear whether subjects used chin/head rests. A main-sequence plot in the supplementary material would be helpful.

      We thank the reviewer for pointing this out. We have included a main-sequence plot in the supplementary materials (page 23). The main-sequence plot can also be found in Supplementary Figure S7 and suggests that our saccade-detection algorithm worked well with detected saccades following the main sequence. We have also stated more clearly in the methods section that subjects used a chinrest (page 11).

      (11) Please specify the asterisk convention in figure captions (i.e., what * vs. ** vs. *** correspond to in terms of p-values).

      We thank the reviewer for pointing out that these significance levels were not mentioned in every figure caption, so we have added this information to every figure caption where they were missing (page 4, 6, 7 and 19).

      (12) The fact that stricter fixation criteria reduced the size of the effect suggests the possibility that gaze drift toward the target might have conferred an eccentricity advantage in this discrimination task. A direct comparison of the effect in the two experiments would be valuable. The authors should comment on this. It would be informative to plot the average gaze position around the time of peak microsaccade rate in both experiments. Reanalyzing the data post hoc with a stricter trial-selection criterion (e.g., excluding trials where gaze deviated more than 1{degree sign} from fixation) could also be very valuable, as it would systematically test how fixation control influences the observed microsaccade-attention relationship. This would be informative for the community studying this topic.

      We thank the reviewer for these valuable reflections. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

      Regarding the average gaze position around the time of peak microsaccade rate: in response to reviewer #1, under point (7), we have included Supplementary Figure S6 that shows the probability distribution of gaze position throughout the ‘shift’ period (the same figure is found on page 23 in the article), which shows that fixational control was successful in both experiments. This period is also the period of peak microsaccade rate in both experiments.

      We wholeheartedly agree that systematically investigating the effect of fixational control is important for the field as a whole, and this is also precisely why we set out to perform the same experiment in two different experimental settings with regards to fixational control, and why we decided to include the results from both experiment variants side-by-side in our article.

      Reviewer #2 (Recommendations for the authors):

      In addition to my general concerns in the public review, I have the following recommendations.

      (1) Did the authors distinguish between the initial and subsequent microsaccades during their analysis? Is it possible to produce multiple microsaccades when shifting attention, or do the authors only consider the first microsaccade to be linked to an attention shift?

      We thank the reviewer for pointing out this ambiguity. We have now stated more clearly in the methods section that we consider all microsaccades for our analyses (page 12: “Crucially, we did not restrict our analyses to initial saccades; rather, all detected saccades were included. This allowed us to examine oculomotor behaviour during later trial phases, where initial saccades are unlikely to occur.”). We also believe this methodological choice is important, as otherwise it would be conceivable that no microsaccade bias can be found during the ‘sustain’ period, simply because no ‘first’ microsaccades occur anymore.

      (2) Two microsaccades had to be separated by at least 100 ms. This is an unusually long delay.

      Could the authors please specify how many microsaccades were discarded using this criterion?

      This inter-saccade-interval is quite large on purpose, as we want to minimise the probability of counting the same microsaccade twice. We have re-analysed the data with a minimum ISI of 50 ms, and this led to an increase of found saccades of a, respectively, 5.1% and 1.9% increase for Experiments 1 and 2. However, two participants in Experiment 1 led to a much higher increase in saccades than all other participants (these participants had z-scores of 3.9 and 2.4 for the number of additionally found saccades with an ISI of 50 ms; all other z-scores for Experiment 1 were between -0.5 and 0.5). When those two participants were removed, in Experiment 1 only 1.6% more saccades were found.

      (3) If I understand correctly, the authors did not use staircase procedures to eliminate differences in task difficulty between participants. Could the authors demonstrate how task difficulty relates to the link between microsaccades and performance? For example, is the time course of the microsaccade direction bias correlated with performance?

      We thank the reviewer for this suggestion (that overlaps with a comment of Reviewer 1). We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (4) The authors reported using equiluminant stimuli. Could the authors please specify the exact luminance?

      We thank the reviewer for pointing out this missing information. We have now included this information in the methods section (page 11: “, with a luminance of 88.5 cd/m<sup>2</sup>.”). We have also included the luminance of the background (page 11: “luminance: 29.0 cd/m<sup>2</sup>”).

      (5) Could the authors please provide a full polar plot showing all microsaccade directions, and colour-code those included in the analysis?

      We thank the reviewer for this great suggestion on how to present our results even more clearly and comprehensively. We have included a supplementary figure showing the full polar histograms (with 20 radial bins), for all three timeframes of interest: the whole trial, the ‘shift’ period and the ‘maintain’ period (page 20). As requested, the saccades included in the main analyses are colour-coded. See Supplementary Figure S3

      (6) Can the authors please directly compare the main effects between experiment 1 and experiment 2 (Figure 2B)?

      We thank the reviewer for this great suggestion (that was also made by reviewer 1). We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

    1. Author Response:

      We are grateful for the careful and extensive reviews, and are pleased that the reviewers found the work of broad interest to sensory processing. Please find our proposal for revision based on public reviews:

      Reviewer 1

      1) Request for dose-response curve for DL-TBOA and leak current or RMP. We can provide this, at least for the initial phase of the curve relevant to the concentrations used for synaptic experiments. Prolonged exposure to higher concentrations leads to very large cationic currents (through AMPAR) which appear to be damaging to membrane integrity.

      2) We will increase the N for glial vs neuronal block with the blockers we already used; this seems more practical than doing new experiments with different concentrations of TFB-TBOA. 

      3) We can include data to test the effect of blockers or small depolarizations on excitability.

      Reviewer 2

      1) We differentiated experiments with “25-50 uM” from 200 uM DL-TBOA because the higher concentration clearly led to massive AMPAR activation and depolarization block, as shown in Fig 1. We then chose lower concentrations to minimize background current while allowing glutamate build-up during exocytosis.  We felt we were clear on this point. 

      As to reversibility and “off target effects” like synaptic changes, we will provide this information. See also response to Reviewer 1, comment 1.

      2) See response to Reviewer 1, comment 3.

      3) We are certain that increasing stimulus strength increases the number of stimulated fibers, and this is well accepted. The stimulus electrode is placed in the auditory nerve root, well away from recorded cell and synapses, minimizing current spread to synapses. We can compare PPR for weak and strong stimuli in our current dataset to confirm no effects on release probability. As to variations in the intrinsic properties of myelinated auditory nerve fibers and their sensitivity to stimulation, there is no information about this, and do not understand how it would be relevant, particularly in as much as we report a negative result: no difference in blocker effect with small or large numbers of fibers active. The 3 main types of myelinated auditory nerve fiber, Type 1a,b,c, are known to respond to different sound thresholds, but that is a synaptic issue in the inner ear, and apparently not related to the myelinated axon bundle.

      4) We appreciate the reviewer's caution about a role for neuronal transporters and will revise accordingly.  We cited molecular evidence for expression of subtypes in the pre and postsynaptic neurons and in glial cells. Of course, given how ubiquitous such expression is across the brain, we suspect the kinds of experiments we provided offer more direct evidence for function.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Summary:

      This manuscript couples a 32-parameter model with simulation-based inference (SBI) to identify parameter changes that can compensate for three canonical hyperexcitability perturbations (interneuron loss, recurrent-excitatory sprouting, and intrinsic depolarisation). The study demonstrates a careful implementation of SBI and offers a practical ranking of "compensatory levers" that could, in principle, guide therapeutic strategies for epilepsy and related network disorders.

      Strengths:

      (1) By analysing three mechanistically distinct hyper-excitable regimes within the same modelling and inference framework, the work reveals how different perturbations require different compensatory interventions.

      (2) The authors adopt posterior estimation to systematically rank the efficiency of different mechanisms in balancing hyperexcitability.

      (3) Code and data are available.

      We thank the reviewers for their positive comments on our manuscript.

      Weaknesses:

      (1) A highly dense presentation of the simulated models and undefined symbols makes it hard for readers outside the modelling community to follow the biological message. An illustration of the models, accompanied by some explanations and references to the main equations and parameters discussed in this paper, would make the first section much more straightforward.

      Thank you for this feedback. To clarify our methods, we have added Figure 7, which illustrates the dynamics of the point neurons and their synapses. We have also added explanations and definitions of variables right where they appear. These variables were previously defined only in a table on a different page.

      We also moved the methods section to the back of the paper, as is common in many modern manuscripts. We hope that relegating method details to the end makes the manuscript more accessible.

      (2) This methodology appears to be a brute-force approach, requiring millions of simulations to tune 32 parameters in a network of 500-700 cells. It isn't scalable. Moreover, the authors did not use cross-validation, which, with a relatively low increase in computational cost, would provide a quantitative measure as to how well it generalizes; this combination raises doubts about both scalability and reliability.

      Scalability is indeed a key challenge of SBI methods. Amortized neural posterior estimation (NPE) is a brute-force approach in that it samples solely from the prior distribution, which is extremely wide. Many of these samples are therefore not very informative for the biologically plausible dynamics we are interested in, which is a downside of amortized NPE. However, amortized NPE is extremely scalable because once the estimator is trained, it can estimate the parameter distribution of any given output dynamic. We tried to build an amortized NPE for our simulator, but simulation-based calibration (a method to validate posterior estimates using additional simulations) showed that the estimators were unreliable.

      Sequential NPE is not a brute-force approach because it samples from posterior estimates, which are narrower than the prior. Because the amortized NPE failed, we use sequential NPE to create the two estimators for the baseline and the hyperexcitable condition described in the paper. While millions of prior samples are used to generate the initial posterior estimate, which is then sequentially refined, the sequential refinement requires only 80,000 additional simulations. This requires a significant amount of computational resources, which is why we consider the results worth reporting, but we make the simulator, the simulation results, and the trained estimators available, so other researchers can use or train their own estimators without running millions of simulations. We hope our rewrites make the advantages and disadvantages of the approach clearer.

      Regarding reliability and cross-validation, we agree that our initial submission has fallen short. We presented results from only one density estimator per condition, which we considered sufficient given the large number of samples. In the revised version, we present the key results from two additional density estimators trained on partially new training data (Figure 4).

      (3) Several parameters remain so broadly distributed after fitting that the model cannot say with confidence which specific changes matter. Therefore, presenting them as "compensatory levers" is somewhat questionable.

      It is indeed difficult to determine which changes matter because of the simulator’s complexity. Especially the marginal correlation coefficient (Figure 2 C) are small and the pairwise histograms are broad (Figure 2A), because all other parameters are unconstrained. But the conditional correlation coefficients are larger (Figure 2D) and narrower (Figure 2B). We have added the histograms in Figure 2B in the revised version to highlight the difference. We cannot provide a definitive threshold for correlation coefficients to discriminate between important and unimportant mechanisms. Therefore, compensatory mechanisms discovered with SBI should be validated mechanistically, as we do in Figure 5.

      (4) Every conclusion is drawn from simulated data; without testing the predictions on recordings, we have no evidence that the proposed interventions would work in real neural tissue. Because today we cannot diagnose which of the three modelled pathological regimes is actually present in vivo, the paper's recommendations cannot yet be used to guide therapy.

      This is indeed an unfortunate drawback of our current work. We are working to apply this approach to constrain microcircuit simulators with data from epilepsy patients. But that work is currently ongoing and will not fit into the present manuscript.

      Recommendations for the authors:

      Beyond the issues I wrote above, which are methodological, I would like to raise my concern about the way this manuscript is written:

      We highly appreciate this editorial feedback on clarity and style. Such feedback is rare and we have worked to address each point to improve the manuscript.

      (1) Paragraphs - several paragraphs start with: "To quantify/identify/find specific compensatory mechanisms of hyperexcitability with simulation-based inference". It is a good idea to orient the reader with the specific goal of each section, but it is not helpful to repeat the overall message of the paper in every paragraph. Several paragraphs open with "However," or "Additionally,". Please restructure sentences so that connectors appear after a clear topic sentence.

      We have done major rewrites to improve the readability of our manuscript. We start paragraphs with more specific context sentences, rather than the broad research goal, and also made paragraphs much shorter with clearer main messages.

      (2) Section 1: The way you present NPE, it would seem like it's specific to neuroscience (and it's not). The paragraph starting at line 26 is not clear. Please revise it. Line 29 - missing a "." before the next sentence begins. Avoid phrases like "for the longest time".

      We now stress that NPE, like SBI, is used across scientific domains.

      (3) Section 2 was tough to read. Please present each equation in its own numbered display, followed immediately by a plain-language explanation of every symbol and parameter. Provide an illustrative diagram: a small schematic of the AdEx neuron, synaptic connections, and the three perturbations. Even a simple block figure will orient nonexperts. Keep critical methodological decisions (priors, summary statistics, simulation length) in the main text, but move voluminous tables of parameter bounds, learning rates, and hardware specs to the supplement. Remove mentions of which Python functions you used. Readers care about algorithmic choices, not function names. Please reserve specific code references for the GitHub README.

      We have added schematic panels at the beginning of Figures 3 & 4 and added Figure 7, which illustrates the neuron and synapse models of the simulator. We also made major rewrites to the methods section to remove programmatic implementation details and define variables where they appear.

      In general, I think it would be a good idea to have an editor to polish syntax, verb tense consistency, and punctuation. A thorough language edit will improve the readability and impact of this manuscript.

      We have attempted to improve the points raised by the reviewer. In particular, we have carefully rewritten verb tense and punctuation throughout the revised manuscript.