10,000 Matching Annotations
  1. Last 7 days
    1. eLife Assessment

      This Review Article provides a comprehensive overview of whole-brain activity changes induced by brain stimulation and effectively summarizes the current state of the field. However, the integrative framework spanning spatial and mechanistic scales, which is presented in the discussion, should be introduced earlier to guide the reader. A more cohesive conceptual framework throughout the manuscript would improve the synthesis of the literature and enhance accessibility.

    2. Reviewer #1 (Public review):

      Summary:

      This paper is a comprehensive review of perturbation studies, and the state-dependence of the brain's response to perturbation at the circuit, mesoscale, and macroscale level.

      Strengths:

      The strengths of the paper are the thorough description of many perturbation studies at different levels of organization, and the integration of both experimental and modeling studies. The review clearly communicates the need to consider 1) brain or local-population state, and 2) multiple levels of organization, in order to understand perturbation responses. Another major strength is the ability for the reader to reproduce figures using the EBRAINS platform.

      Weaknesses:

      The major weakness is that the review does not include a significant integration across scales, and as a result reads like three separate (though comprehensive) reviews. Currently, the only integration across the scales is in a brief conclusion paragraph. I would recommend adding an additional section, in which the overarching picture is discussed. (i.e. a unifying view of state dependence, and what is learned by considering across scales), and more prefacing in the introduction of the overarching message and framework to the review.

    3. Reviewer #2 (Public review):

      Summary:

      In this review article, the authors discuss the whole brain activity changes induced by brain stimulation. They review the literature on how these activity changes depend on the cognitive state of the brain and divide the results by the scale of the change being induced, from microscale changes across small groups of neurons, up to macroscale changes across the entire brain. Finally, they describe attempts to model these changes using computational models.

      Strengths:

      The review provides an overview of the results within this sub-field of neuroscience, and the authors are able to discuss a lot of prior results. The framing of the changes in neuronal activity in terms of computational changes is also a helpful approach.

      We thank the authors for the updates that they have made in response to our original comments. Their attempts to address many of the comments that we raised have greatly improved the paper. We believe that there are two major points that still require some additional changes:

      (1) We raised the concern that the results within each of the three spatial scales did not join together into a cohesive single framework. The authors responded by updating the conclusion section to provide a more conceptual picture linking the different spatial scales. This is much appreciated. However, by placing this framework at the end of the paper, it prevents the reader from using this understanding to building a conceptual model as they progress through the paper. We would ask that the authors intersperse this conceptual picture within the main text, and to then re-emphasize it in the conclusions. This would frame each section in terms of the findings that led directly to it and, therefore, allow the reader to build a conceptual understanding within each section. As one example, we note that the authors have made no changes to the mesoscale processing section. Therefore, when reading that section, it is completely unclear how any of the results seen in the microscale may relate to the changes observed at the mesoscale.

      (2) The authors have greatly improved their explanation of the complexity metrics. However, the paper still lacks a conceptual understanding for why "perturbation-based complexity metrics" are a reasonable way to study the state-dependent dynamics? What does studying perturbations provide that studying the spontaneous activity in different states alone, would not provide? Why is this the preferred way to study such dynamical systems? Such a justification would strongly support the analyses reviewed in the paper and would increase the reader's understanding of the methodology.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This Review Article provides a thorough overview of whole-brain activity changes induced by brain stimulation and summarizes the current state of the field. However, it lacks integration across spatial and mechanistic scales, which limits the reader's ability to understand how the different findings relate to one another. In addition, several key concepts are not explained in sufficient depth for non-expert readers. The manuscript would benefit from the development of a cohesive conceptual framework to more clearly synthesize the existing literature.

      Thank you for the positive assessment. We fully agree, and as suggested we have added a new conclusion paragraph that outlines a synthesis of the paper and suggests a conceptual framework :

      “In this paper, we have reviewed aspects of neuronal responsiveness, from the microscale level of neurons and circuits, the mesoscale level of single brain areas, and the macroscale level of the whole brain. At the microscale, it is apparent that the circuit operating in an asynchronous mode displays the highest responsiveness, as seen in brain slices (D’Andola et al., 2018). The underlying mechanism is that the high levels of synaptic « noise » in asynchronous states set neurons in a high responsive mode, as seen in models of single neurons (Ho & Destexhe, 2000). This higher responsiveness is confirmed at mesoscale, and can be seen for example with Utah-array recordings comparing wake and anesthesia (Dwarakanath et al., 2025). Similarly, propagating waves occur in the asynchronous state in awake monkey (Muller et al., 2014), and sensory inputs evoke more propagating patterns (and higher PCI) in wakefulness with asynchronous states compared to slow-wave states of anesthesia in mice (Montagni et al., 2024). At the whole-brain scale, experiments also find that evoked responses are more complex and propagating compared to slow-wave states (Massimini et al., 2005), a situation which models can reproduce (Goldman et al., 2023; Sacha et al., 2025). Other measures, such as fluidity (Breyton et al., 2024) and reversibility (Camassa et al 2024) also point to the same conclusion. Collectively, these results show that asynchronous and irregular activity states set neurons in a high responsive mode, which in turn impacts network behavior and favors the propagation of activity as mesoscale propagating waves, or macroscale activity patterns that propagate across brain regions. It is therefore not surprising that the best correlate of conscious states is the asynchronous activity (Koch et al., 2016).”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper is a comprehensive review of perturbation studies and the state-dependence of the brain's response to perturbation at the circuit, mesoscale, and macroscale levels.

      Strengths:

      The strengths of the paper are the thorough description of many perturbation studies at different levels of organization, and the integration of both experimental and modeling studies. The review clearly communicates the need to consider (1) brain or local-population state, and (2) multiple levels of organization, in order to understand perturbation responses. Another major strength is the ability for the reader to reproduce figures using the EBRAINS platform.

      Weaknesses:

      Two major points of improvement should be resolved with the review, in order to make it useful for a broad audience.

      The first is that the review does not include a significant integration across scales, and as a result, reads like three separate (though comprehensive) reviews. Currently, the only integration across the scales is in the brief conclusion paragraph. I would recommend adding an additional section, in which the overarching picture is discussed. (i.e. a unifying view of state dependence, and what is learned by considering across scales). This need not be too long, but it should be longer than a single conclusion paragraph.

      Thank you for the positive assessment. We fully agree with the excellent suggestion of adding a concluding paragraph where the conceptual framework and overarching picture are presented. Please see the new conclusion paragraph that we added to the paper (copied above in the reply to Editors).

      The second major weakness is that there is a lack of clarity on many points throughout, which is needed for the reader to fully understand the results described.

      See our answer to the specific comments below for the list of unclear points.

      Reviewer #2 (Public review):

      Summary:

      In this review article, the authors discuss the whole-brain activity changes induced by brain stimulation. They review the literature on how these activity changes depend on the cognitive state of the brain and divide the results by the scale of the change being induced, from microscale changes across small groups of neurons, up to macroscale changes across the entire brain. Finally, they describe attempts to model these changes using computational models.

      Strengths:

      The review provides an overview of the results within this subfield of neuroscience, and the authors are able to discuss a lot of prior results. The framing of the changes in neuronal activity in terms of computational changes is also a helpful approach.

      Weaknesses:

      However, the authors are not able to contextualize these results within a single framework, i.e. explaining from first principles how different aspects of stimulus-induced changes interact to generate functional changes in the brain, and how different changes - at distinct spatiotemporal scales - combine to form larger effects. This is a significant weakness in generating a review of the literature, since the authors do not provide a cohesive conceptual framework on which to frame the results. Similarly, the authors do not explain how their different computational models fit together, and how one can get a singular computational understanding of the distinct mechanisms of brain activity changes due to stimulation under different brain states, by combining the results derived from each separate model.

      Thank you for the positive assessment. This is an excellent suggestion, actually also requested by Reviewer 1. We have now added a new conclusion paragraph where the conceptual framework is explained (copied above in the reply to Editors).

      Major Comments:

      (1) The authors have written this review as if it were intended for an audience who is already familiar with the topics. For example, they introduce concepts like complexity, spiral vs planar waves, without much explanation.

      Thank you for this helpful comment. We agree that the Introduction should be more accessible to readers who are less familiar with these concepts, and we have therefore revised the text to provide a clearer definition of complexity and a more explicit explanation of propagating wave patterns.

      Specifically, we now clarify that complexity can be understood as the richness of the set of accessible states of a system, which in our context can be related to the diversity of slowwave propagation modes and to the high-entropy, desynchronized activity of the awake brain. We also expanded the description of propagating slow waves to distinguish planar from spiral waves and to explain how their relative prevalence changes with anesthesia depth.

      In the Introduction we replaced the sentence “New methods … at various scales” with “New methods for characterizing the complexity of network dynamics and their response patterns have emerged, particularly recently (Krohn et al., 2023; Wolf et al., 2018), and are presented here at various scales. Here, complexity is associated with the set of accessible states of a system (Parisi, 2006). In the present context, this notion can be linked to the diversity of slow-wave propagation modes and to the richness (i.e., the entropy) of perturbation-evoked responses in brain activity.”

      While in Results (p. 13) the sentence “Spontaneous slow waves … administered (Huang et al., 2010).” has been expanded in “Spontaneous slow waves can also display propagating patterns, as shown in anesthetized mice (Huang et al., 2010; Mohajerani et al., 2010; Pazienti et al., 2022; Stroh et al., 2013). These patterns may take the form of planar waves, which travel across the cortex along a relatively regular front, or spiral waves, which rotate around a central core and therefore produce a more complex spatiotemporal organization. Under relatively deep anesthesia, spiral waves occur more frequently than planar waves, whereas the opposite imbalance is observed as anesthesia is lightened (Huang et al., 2010).”

      (2) Regarding complexity, the authors present a quantification termed PCI. However, in the associated box, they state that PCI could be implemented in a number of different ways, using analogous metrics (which are, nonetheless, not identical). Yet the authors simply claim that all these metrics are sufficiently similar to be grouped together as "PCI". The authors do not provide much intuition about this, and they also don't present any other potential quantifications. This makes any interpretation of their results strongly dependent on your understanding of the concept of PCI. It would be helpful to present some other, analogous metric to demonstrate that the results that the authors are focusing on are not somehow tied to the specific computational structure of the PCI metric.

      Thank you for pointing out to this inconsistency. We agree that the rationale for focusing on perturbational complexity was not sufficiently introduced in the original version of the manuscript.

      Broadly speaking, complexity measures used in consciousness research can be divided into two major classes. The first includes observational measures, which are computed from spontaneous ongoing activity and quantify statistical dependencies within neural time series. The second includes perturbational measures, which quantify the deterministic causal interactions revealed by a controlled perturbation of the system and their spatiotemporal propagation across the network (see Sarasso et al., 2021).

      The primary aim of the present Review was to discuss how complexity changes across spatial and temporal scales in response to perturbations. For this reason, we focused on perturbational complexity measures and, in particular, on the Perturbational Complexity Index (PCI), which remains one of the most widely adopted and validated approaches in this category.

      As described in Box 2, different implementations of PCI have been proposed. The two most established versions are PCI based on Lempel–Ziv complexity (PCI^LZ) and PCI based on state transitions in principal component space (PCI^ST). Although these implementations differ algorithmically, they were developed to operationalize the same theoretical construct and have been shown to correlate strongly when applied to the same datasets (Comolatti et al., 2019). For this reason, throughout the Review we use the term “PCI” as an umbrella label encompassing these related perturbational complexity measures.

      Importantly, all complexity measures discussed in the studies reviewed here belong to this broader class of perturbational approaches. While adaptations of the original algorithms are often required when dealing with different recording modalities and spatial scales, these modifications mainly concern preprocessing and signal representation rather than the underlying theoretical construct being quantified.

      To clarify this point, we have revised the Introduction to explicitly motivate our focus on perturbational complexity, to distinguish perturbational from observational complexity measures, and to explain why different PCI implementations can be discussed within a common conceptual framework. We believe that these additions make the rationale of the Review substantially clearer and reduce the impression that the conclusions depend on a specific implementation of PCI.

      (3) The authors divide the review into sections organized by the spatial extent of the effects that they are exploring (e.g. from microscale to macroscale). However, they don't bring together these insights into a cohesive structure - for example, by providing potential explanations of the macroscale effects by using the microscale changes.

      We agree – and this is now the focus of the newly-added conceptual-framework conclusion paragraph.

      (4) The authors completely ignore any aspect of cell-type specificity in their review, despite the known importance of specific cell types at the microcircuit scale. This makes it difficult to map their results onto the true biological system.

      We agree that cell-type specificity could be made more explicit. The revised manuscript now clarifies that several models already include cell-type specificity. For example, the AdEx-based models distinguish excitatory regular-spiking or pyramidal populations with adaptation from inhibitory fast-spiking populations without adaptation. This differentiation is not made with other models like leaky or quadratic integrate-and-fire models. At the mesoscale, mean-field approaches can be derived for different structures, such as cortex, thalamus, hippocampus, striatum, or cerebellum, and can incorporate the experimentally experimentally observed firing properties of relevant cell classes.

      (5) The authors introduce several different computational models, such as the Hopf model, the AdEx model, and the MPR model. However, they do not provide the reader with a conceptual understanding of the structure of each of these models (except through potentially more complex terminology, e.g. the Hopf model is a "phenomenological StuartLandau nonlinear oscillator"). Additionally, though they present the results of each simulation, they don't provide the reader with intuition about how these models compare against each other, and how best to interpret results derived from each model.

      Very good question, and the answer is not easy. If the goal is to capture large-scale phenomena with models as simple as possible, then Stuart-Landau, Hopf, or Jahnsen-Rit may be appropriate. This approach is rather top-down. But if the goal is to assess how microscopic changes (synaptic receptors for example) affect large-scale brain activity, then we need a bottom-up approach, where mean-field models are derived. We can better explain this.

      We agree that while the technical definitions of the whole-brain models (Hopf, AdEx, MPR) were provided, a clear conceptual framework comparing their underlying structures, specific trade-offs, and interpretation guidelines was missing. We have substantially revised the "Macroscale" section on Page 22 to provide immediate intuition regarding what each model represents structurally (e.g., macroscopic phenomenology vs. microscopic biological realism). We emphasized the structural assumptions of each of them, as well as the explicit utility in interpreting brain responsiveness. This ensures readers understand exactly why a researcher would choose one model over another depending on the mechanistic question at hand.

      (6) In several cases, the authors make statements that they appear to believe to be completely straightforward (and require no justification), but that do not appear so to the reader. For example, they mention: "In wakefulness and REM sleep, ..., the membrane potential is depolarized and close to the spike threshold, which explains why neurons respond more reliably and with less response variability compared with slow-wave sleep". However, this statement is not obvious to the reader and requires explanation (for example, in a system that is close to balance, bringing cells closer to the firing threshold can result in increased response jitter).

      We agree that the original statement was an over-simplification of a complex situation. We have revised it to avoid suggesting that depolarization alone monotonically increases reliability. The relevant mechanism is the combination of depolarization, desynchronized high-conductance synaptic input, balanced fluctuations, and reduced tendency to enter long silent Down states. In this regime, weak inputs are more likely to be converted into spikes and propagate through the network. However, too high conductance, excessive noise, can shunt inputs, enhance jitter, or saturate the network. We are now more explanatory.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As stated in the public review, there is a lack of clarity on many points throughout, which is needed for the reader to fully understand the results described.

      Points needing clarification:

      (1) sPCI (slice Perturbational complexity index) - is this different from other PCIs in box 2? Regardless, the metric and its interpretation should be briefly explained in the main text.

      We use sPCI to refer to the PCI measure adapted for application to cortical brain slices (D’Andola et al., 2017; see Box 2). It relies on the same core algorithm as PCI, namely the Lempel–Ziv complexity of the spatiotemporal pattern of significant responses (Casali et al., 2013) but differs in the preprocessing steps required for slice recordings. We now added this clarification to the text.

      (2) Page 9 "by decreasing fast inhibition but also enhancing it" - What does that mean? More info about the model is needed.

      Thanks for raising this point, since this sentence was indeed confusing. We have revised it now.

      (3) Page 9 "balance between segregation and integration, a crucial ingredient on which sPCI relies" - How is this balance seen in the figure? All I see is sPCI and blockage of GABA.

      The comment is correct, and this mention of segregation and integration has now been eliminated.

      (4) Figure 3D, Page 11 "two different desynchronized (AI) states in a network of AdEx neurons"- What are the two different states? Why is the response different?

      The different AI states correspond to different synaptic strength parameters, we added this precision in the text.

      (5) Figure 3B - "Bifurcation diagram showing the different activity regimes displayed by spiking neuron network." Which model? Multiple are cited. This is a general issue throughout where multiple models are mentioned in the text, and it's unclear which is shown in the figure.

      We agree with the Reviewer's helpful remark. We have revised the manuscript to explicitly state the types of models depicted in the different panels of Figure 3. Corresponding details have also been incorporated into the relevant text in the Results section (previously pages 10–12)."

      A few editorial issues:

      (1) The text in many of the figure panels was too small to read. This is a significant issue that must be addressed.

      We will fix this at the next round, can you please let us know which figures are not visible?

      (2) I recommend reading through for writing flow. E.g. In the first paragraph of the introduction, there are two sentences that start with "importantly, ..." in a row.

      Thanks for noting this – this is now fixed.

      (3) Figure 1E - How does the color on the left relate to the right? What is the y-axis?

      The colour code corresponds to the latency of activation (light blue, 0 ms; red, 300 ms). The Y-axes is the global mean field power (voltage). It has now been included in the figure caption.

      (4) Figure 1C - What is the stimulus?

      The triangle corresponds to the electrical stimulation of the homotopic area 18 of the contralateral hemisphere. This information is now included in the figure legend.

    1. eLife Assessment

      This study presents an important finding regarding the role of oxytocin neurons in thermogenesis and behavioral thermoregulation. The use of numerous converging methods, including behavior, fiber photometry, optogenetics, thermal recordings, metabolic analyses, and more, produces a multi-dimensional dataset delivering findings that provide solid support for the conclusions. The conclusions could be further strengthen by more extensive analyses of behavior and determining whether it is the release of oxytocin (rather than co-release of glutamate) from the PVN that is critical for the transition between behavioral states, nevertheless, the manuscript had many strengths, the findings are novel, and this work opens new doors for understanding the role of the PVT in thermoregulation. This work will be of strong interest to the thermoregulation, social behavior, and oxytocin signaling communities.

    2. Reviewer #1 (Public review):

      Summary:

      The authors identify and investigate a specific population of PVNOT neurons (oxytocin neurons of the paraventricular hypothalamus) that seem to be involved in both behavioral and autonomic thermoregulation. These cells are activated by social thermoregulatory behaviors, but can influence thermoregulation in both social and social contexts, specifically during transitions and when mice are at low core body temperature (Tb).

      Comments on revised version.

      The authors have addressed my concerns with clear and reasonable explanations and altered the text accordingly. This has improved the paper, but it still feels in some parts like a patchwork of nice work and discoveries stitched together. Further changes to format, analysis, and some experimental work could hugely improve the manuscript. I see that will surely come from future work, and this is the authors' choice.

      Regarding the lack of behavioral analysis, I think it's fair for them to keep it for future studies.

      I am happy to see they take and expand the opto inhibition suggestion. Again, that experiment would be nice for this paper, but not crucial.

      Regarding discussing Raam et al 2026. It is good that they detail the practical decision of using females. What I meant was that, given that both papers study calcium dynamics around the time when mice engage in social thermoregulatory behaviour, they could have speculated on potential dmPFC-PVN functional connectivity, for example. Or the fact that Raam found that females showed fewer huddling behaviour than males at 5{degree sign}C (however, Vandendoren tested 15{degree sign}C, not 5{degree sign}C). Discussion of these features would be welcome, but maybe all of the current scope.

      Overall, this is a very strong paper.

    3. Reviewer #2 (Public review):

      This is a very interesting study from Vandendoren and colleagues examining the role of PVN oxytocin neurons during thermoregulatory behaviors, in particular during thermoregulatory huddling. The findings are important and have implications for the thermoregulation field as well as the social/naturalistic behavior field. The findings are compelling and use a combination of state-of-the-art tools (photometry, optogenetics, automated behavior tracking, thermal imaging, and core body temperature measurement), often in combination with each other, to produce a rigorous and high-dimensional dataset.

      Comments on revised version.

      I appreciate the effort the authors have put into addressing all of my questions, and I have no remaining concerns.

    4. Reviewer #3 (Public review):

      Summary:

      This study investigates how the activity of hypothalamic paraventricular oxytocin (PVNOT) neurons relates to physiological states in female mice, with a particular focus on behavioral states and thermogenic sympathetic activity. To address this question, the authors combined automated video-based behavioral classification with calcium imaging of PVNOT neuron activity. Sympathetic thermogenesis was inferred from surface temperature changes measured by infrared thermography, and the authors have made their custom analysis scripts available. The authors report that strong, pulsatile activation of PVNOT neurons was "occasionally" observed immediately before transitions from resting to active states. This observation suggests that PVNOT neuronal activity may facilitate the transition from rest to activity. This phenomenon was observed in both pair-housed and individually housed animals. Taken together, these findings raise the possibility that the oxytocinergic system contributes to naturalistic behavior transitions even in the absence of social interactions. However, concerns regarding the selectivity of GCaMP expression in oxytocin-expressing neurons call into question the validity of the recorded PVNOT neuronal activity. The revised manuscript improves the presentation and interpretation of the data. Nevertheless, because the authors have not provided additional experiments or analyses addressing the major methodological concerns, the evidence supporting the central conclusions remains essentially unchanged.

      Strengths:

      The oxytocinergic neural system is believed to subserve a wide range of physiological functions. Elucidating these roles requires monitoring PVNOT neuronal activity under diverse behavioral contexts, as well as manipulating this activity to establish causal relationships. In this study, the authors present a technically sound experimental framework that integrates behavioral tracking in both individually and group-housed mice with the monitoring and manipulation of PVNOT neuron activity. This setup represents a valuable methodological resource for researchers investigating the physiological functions of oxytocin.

      Weaknesses:

      (1) Immunohistochemical validation of selective GCaMP expression in oxytocin-expressing neurons showed that only 24-51% of GCaMP-positive neurons expressed oxytocin. As an alternative approach, the authors argue that the similarity between calcium dynamics recorded in virgin and lactating animals supports the identity of the recorded neurons as oxytocin neurons. While this physiological comparison is interesting, it does not constitute direct evidence for cell-type specificity of GCaMP expression. The revised manuscript now acknowledges that in situ hybridization targeting oxytocin mRNA would provide a more reliable validation, but such validation has not been performed. Therefore, uncertainty regarding the identity of the recorded neurons remains, limiting confidence in the interpretation of the calcium imaging data.

      (2) Although the authors' interpretation is generally consistent with the data presented, their main conclusions rely heavily on observational findings. Moreover, optogenetic stimulation of PVNOT neurons failed to robustly recapitulate behavioral state transitions (Figs. 6D and S5B). Further interventional experiments remain necessary to rigorously test the authors' interpretation and establish a causal relationship between PVNOT activity and rest-to-active transitions. In particular, loss-of-function approaches targeting the PVNOT system, such as OXTR antagonism, inhibitory optogenetics, or cell-type-specific ablation, remain essential to determine whether perturbation of this system alters behavioral state transitions. Although the authors expanded the Discussion to acknowledge this limitation, the revised manuscript provides no additional experimental evidence addressing it.

      Comments on revised version.

      I appreciate the authors' efforts to clarify the manuscript and to discuss the limitations more explicitly. Nevertheless, because my major concerns have been addressed primarily through revised interpretation rather than new evidence, my overall assessment of the scientific support for the principal conclusions remains unchanged.

    5. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version.

      As discussed before, the authors employ a wide range of techniques (FOS IHC, FP for fine scale PVN OXT population dynamics, behavioural analysis, core and surface temperature tracking, physiological recordings to assess AAV specificity, optogenetic activation of PVN OXT neurons, and projection tracing) to address a clear question. The outcomes of these techniques seem to drive the same conclusion that PVN OXT neurons signal transitions from rest to arousal (behavioural and thermogenic) in a state-dependent manner:

      - FOS data identifies PVN OXT population activity following behavioural onset

      - Ca activity in these cells peaks at behavioural and thermogenic state transitions

      - Rump temperature and BAT activity increase at state transition points

      - Optogenetic stimulation of these cells recapitulates the thermogenic effects seen during physiological state transitions (in low body temperature animals) with a trending increase in physical activity

      Despite the inconclusive IHC results when validating the specificity of their AAV, the virgin female/ lactation experiment is convincing that they are specifically targeting PVN OXT neurons. The rationale for this experiment is clearer in the revised manuscript.

      Generally, in terms of the revised manuscript, the authors give strong responses to reviewer comments, either incorporating feedback, or giving clear explanations for the choices they made in the original manuscript. The revised manuscript is clearer about the question the authors aim to address, the reasons for their choice of experiments, and the limitations of the techniques used.

      We thank the reviewer for the close attention to the manuscript, the response to reviewers, and the revision, all of which have improved the manuscript.

      Criticisms:

      I appreciate and agree with the authors' point that this manuscript is more fundamental than simply social basis oxytocin neuron function. This is point is well made by their data, and in the revised text. However, I still believe more behavioural analysis would be welcome to any reader.

      They partly justify the lack of behavioural analysis in Figure 6 with the problem of "animal merging" on the SGBS images. However, in Figure 6C, they confirm that, in solo conditions, the SGBS readings are consistent with core body temperature readings. So why not stick to core body temperature, opto stimulate and analyse the social behaviour with DLC (with normal video recordings)?

      This is a good suggestion. Because we find that quiescent huddling (paired) bouts were associated with stronger body temperature regulation compared to solo quiescence and other behavioral states, and because PVNOT peak probability and frequency were higher in the paired compared to solo context, these experiments are warranted. We made the following edits to the discussion:

      “Future experiments should attempt to disentangle the effects of PVNOT light stimulation on social vs. non-social aspects of these behavioral state transitions; of particular interest would be to examine how light stimulation affects the duration and thermoregulatory control of social huddling.”

      The lactation validation still seems out of place in manuscript order. It is a very valuable validation, but it feels more like supplementary data for Figure 1. I feel the authors wanted it as a main figure because of how much work it must have been. In that case, it still makes more sense to include it in Figure 1.

      The purpose of the lactation experiment arose from the inadequacy of using histology to test whether AAV-transfected cells were oxytocin-immunoreactive. Because we observed intense oxytocin immunoreactivity in the fibres lining the ventricle, and less reactivity in the cell bodies than what we would have predicted from the Oxytocin-Cre-dependent AAV, we turned to the known physiological relationship between oxytocin-positive neurons and lactation. As such, this study is not associated with Figure 1, which demonstrates our initial, coarse-grained findings relating FOS activity in the PVN and in oxytocin-positive neurons during social thermoregulation.

      To your point, it typically does make sense to have the cellular validation “up front” as supporting or background information that enables the downstream experiments. However, what gives this data credibility as a standalone figure is the novel finding that PVNOT neurons display burst-like patterns of activity outside the context of lactation. Previous discussions with experts in the field, along with a review of the literature, unexpectedly led us to the observation that the burst-like patterns we observed during the transition from rest to wake and thermogenesis in virgin females represents a new aspect of oxytocin neuron physiology. Because we wanted to directly compare the new virgin female activity pattern (i.e., Figure 2) with the known lactation activity pattern, we decided it made the most sense to combine the validation aspect with the novel aspect into a standalone figure.

      Though their lactation experiment validates that they are targeting PVN OXT neurons, their optogenetic stimulation protocol may not be specifically inducing OXT release from these cells. PVN OXT neurons co-release glutamate but can also release glutamate independently of OXT following lower frequency tonic stimulation. OXT release from PVN neurons requires pulsatile stimulation at a higher frequency (Leithead et al., 2021; Piñol et al., 2014; Lincoln & Wakerley, 1975). In this paper, the authors use a low stimulation frequency (10Hz) and continuous pulse train (20s) to optogenetically manipulate the target PVN population which may bias the cells towards glutamate release over OXT. Therefore, though they find evidence that PVN OXT neurons are involved in driving the transition between states in their other experiments, their optogenetic stimulation may not necessarily involve OXT release/signalling. It may be valuable to separate this out to identify the signalling molecule underlying this behavioural/ thermogenic transition. This could be done by using an opto protocol that recapitulates physiological OXT release.

      The authors do however mention that isolating the specific contribution of OXT signalling compared to other co-transmitted molecules was not the aim of this study, so this is not an essential question for this manuscript.

      Thank you for this thoughtful point. We agree our optogenetic stimulation experiment should be interpreted as activation of PVNOT neurons rather than as selective evidence for oxytocin release or oxytocin signaling. PVNOT neurons can co-release glutamate (an idea we had also briefly touched upon in the Limitations and caveats section), and the stimulation pattern/frequency may influence the relative engagement of fast glutamatergic transmission versus peptide release. We agree the lactation literature, including Lincoln et al., highlights the importance of high-frequency pulsatile activity for oxytocin release, and that Piñol et al. provide evidence that PVNOT-linked glutamatergic transmission can interact with oxytocin-receptor-dependent modulation of downstream synapses–so thanks for pointing these out.

      We made revisions to support our protocol and now acknowledge this important aspect of the neuronal physiology. In Results, we now explain why we selected 10Hz: this frequency was grounded in the study by Fukushima et al. (2022), where 10Hz stimulation of PVNOT terminals in the rMR elicit thermogenic responses and 10Hz stimulation of PVNOT somata produce thermogenesis that’s dependent on oxytocin receptors in rMR.

      In the Limitations section, we now cite these three references to include broader context around stimulation frequency and differential release. We emphasize that our optogenetic data demonstrate sufficiency of PVNOT neuron activation, but do not establish whether the downstream thermogenic and behavioral effects are mediated by oxytocin, glutamate, or both. We note that resolving this issue will require future experiments using stimulation-pattern comparisons together with receptor-targeted pharmacology or genetic loss-of-function approaches.

      References

      Leithead, A. B., Tasker, J. G., & Harony-Nicolas, H. (2021). The interplay between glutamatergic circuits and oxytocin neurons in the hypothalamus and its relevance to neurodevelopmental disorders. Journal of neuroendocrinology, 33(12), e13061. https://doi.org/10.1111/jne.13061

      Lincoln, D. W., & Wakerley, J. B. (1975). Factors governing the periodic activation of supraoptic and paraventricular neurosecretory cells during suckling in the rat. The Journal of physiology, 250(2), 443-461. https://doi.org/10.1113/jphysiol.1975.sp011064

      Piñol, R. A., Jameson, H., Popratiloff, A., Lee, N. H., & Mendelowitz, D. (2014). Visualization of oxytocin release that mediates paired pulse facilitation in hypothalamic pathways to brainstem autonomic neurons. PloS one, 9(11), e112138. https://doi.org/10.1371/journal.pone.0112138

      A loss of function experiment to test for sufficiency would be a nice addition to further confirm their claims, but the authors mention that there were technical limitations to their attempts at inhibiting PVN OXT neurons. I appreciate the authors declaring that the DREADDs attempt suffered from unfortunate confounds. But for optogenetic attempts, I don't think they need a closed-loop system to get some useful results. They still can shine the light at "random" moments (that will correspond to random body temperatures) and then separate the data per body temperature.

      We thank the reviewer for this constructive suggestion. Such an experiment would strengthen our claims and complement the optogenetic activation (Fig. 6). Reviewer 3 brought up a similar concern.

      Building directly on the reviewer’s proposal, we now describe a loss-of-function experiment as an important next step. Optogenetic inhibition of PVNOT neurons can be delivered at pseudo-random times across light and rest phase. Because animals spend extended periods at rest during this phase, a substantial fraction will fall within established rest bouts, which can then be analyzed and stratified by body temperature, as the reviewer notes. The prediction is that silencing PVNOT neurons during rest should prolong the average duration of rest bouts and delay the onset of activity and thermogenesis, relative to matched unstimulated bouts.This provides a direct test of whether PVNOT activity is necessary for the transition from rest to activity. We have revised the Limitations and caveats section to describe this experiment.

      “Third, although we show that PVNOT neurons are sufficient to drive thermogenic and behavioral transitions (Fig. 6), we did not perform acute loss-of-function experiments. Such experiments are warranted because decreases in baseline PVNOT calcium activity were associated with transitions toward the onset of quiescence (Fig. 3I-L), suggesting this system may bidirectionally regulate thermo-behavioural state. A tractable next step would be to optogenetically inhibit PVNOT neurons during established rest bouts, delivered at pseudo-random times across the light and rest phase and analyzed post hoc by behavioral state and body temperature; we predict that silencing during rest would prolong the average duration of rest bouts and delay the onset of activity and thermogenesis. Pairing the inhibition with selective oxytocin antagonist (such as L-368,899), would further test whether the thermogenic and autonomic components of these transitions are oxytocin receptor dependent rather than driven by glutamate released by the same neurons.”

      Lastly, the mention of Raam et al. 2026 is insufficient. The authors just mention it regarding the potential differences with males, to be explored in future experiments. Even if not using males in the current study doesn't affect the stated conclusions, the fact that they chose females because "their thermo-behavioural states were readily discernible" is a considerable bias. Testing males in this very study might be out of scope, but more discussion is warranted.

      We thank the reviewer for this point. We agree that our decision to study females deserves fuller treatment, and we have expanded the Limitations and caveats section accordingly.

      We want to be clear about the rationale, because it was methodological rather than an assumption of sex specificity. Our previous study on behavioral thermoregulation in mice (Landen et al., 2024) showed that, during the light/rest phase, females–but not males–display clearly rhythmic episodes of rest and activity that align with transitions between thermoregulatory states, and are therefore well suited to the analyses that form the core of this study. This choice does constrain the generality of our findings to females, but it does not affect the validity of the conclusions we draw, all of which concern PVNOT neurons in females.

      At the same time, we agree that whether these mechanisms extend to males is a substantive open question and we now say so explicitly. A direct comparison in males, while beyond the scope of the present study, is an important next step, and the recently defined neural basis of collective thermoregulatory huddling (Raam et al. 2026) offers a useful framework for that work. We have modified the Discussion/Limitations and caveats as follows:

      “We focused on females for a practical reason: during the light and rest phase, females show clear, rhythmic bouts of rest and activity, which makes transitions between thermoregulatory states readily discernible and well suited to the analyses around each state transition used here (Landen et al., 2024). This choice constrains the generality of our conclusions, which pertain specifically to females. Because oxytocin signaling can differ between sexes (https://doi.org/10.1016/j.yfrne.2015.04.003), and because the neural control of thermoregulatory behavior may not be identical in males, whether the PVNOT dynamics we describe operate similarly in males remains an open question. Testing males directly was beyond the scope of the present study, but it is an important next step, particularly as the neural basis of collective thermoregulatory huddling has recently begun to be defined (Raam et al. 2026).”

      Reviewer #2 (Public review):

      Summary:

      This is a very interesting study from Vandendoren and colleagues examining the role of PVN oxytocin neurons during thermoregulatory behaviors, in particular during thermoregulatory huddling. The findings are important and have implications for the thermoregulation field as well as the social/naturalistic behavior field. The findings are compelling and use a combination of state-of-the-art tools (photometry, optogenetics, automated behavior tracking, thermal imaging, and core body temperature measurement), often in combination with each other, to produce a rigorous and high-dimensional dataset.

      Comments on revised version.

      I appreciate the effort the authors have put into addressing all of my questions, and I have no remaining concerns.

      Thanks for the comments; they have greatly improved the manuscript.

      Reviewer #3 (Public review):

      Summary:

      This study investigates how the activity of hypothalamic paraventricular oxytocin (PVNOT) neurons relates to physiological states in female mice, with a particular focus on behavioral states and thermogenic sympathetic activity. To address this question, the authors combined automated video-based behavioral classification with calcium imaging of PVNOT neuron activity. Sympathetic thermogenesis was inferred from surface temperature changes measured by infrared thermography, and the authors have made their custom analysis scripts available. The authors report that strong, pulsatile activation of PVNOT neurons was "occasionally" observed immediately before transitions from resting to active states. This observation suggests that PVNOT neuronal activity may facilitate the transition from rest to activity. This phenomenon was observed in both pair-housed and individually housed animals. Taken together, these findings raise the possibility that the oxytocinergic system contributes to naturalistic behavior transitions even in the absence of social interactions. However, concerns regarding the selectivity of GCaMP expression in oxytocin-expressing neurons call into question the validity of the recorded PVNOT neuronal activity.

      Strengths:

      The oxytocinergic neural system is believed to subserve a wide range of physiological functions. Elucidating these roles requires monitoring PVNOT neuronal activity under diverse behavioral contexts, as well as manipulating this activity to establish causal relationships. In this study, the authors present a technically sound experimental framework that integrates behavioral tracking in both individually and group-housed mice with the monitoring and manipulation of PVNOT neuron activity. This setup represents a valuable methodological resource for researchers investigating the physiological functions of oxytocin.

      Thanks for the comments. We are encouraged to hear this framework will open new doors in understanding how the oxytocin system regulates behavior and energy homeostasis.

      Weaknesses:

      (1) Immunohistochemical validation of selective GCaMP expression in oxytocin-expressing neurons showed that only 24-51% of GCaMP-positive neurons expressed oxytocin. As an alternative approach, the authors demonstrate that GCaMP-expressing PVN neurons in virgin females exhibit calcium peaks during rest-wake transitions with kinetics similar to those observed in PVNOT neurons during early lactation. However, this comparison is based solely on population-level peak profiles and does not provide direct evidence for cell-type specificity of GCaMP expression in oxytocin neurons. This limitation substantially undermines the validity of the optical calcium imaging data. In situ hybridization targeting oxytocin mRNA, rather than immunohistochemistry, may provide a more reliable assessment of expression specificity.

      We view our data as showing strong evidence that the recorded neurons include, but may not be limited to, PVNOT neurons for the following two reasons: (1) as the reviewer notes, our longitudinal experiment shows conservation in the physiological and biophysical profile of these neurons in females that went from virgins to parturition and lactation, and (2) as described in Discussion/PVNOT neurons in context of arousal and peptidergic PVN cell-types, non-OT cell-types in the PVN do not show this pulsatile busting profile.

      In the “Discussion/Thermal tracking and validation of PVNOT recording specificity” section we had stated “We note that the animals were perfused at ~ZT4–8, before we were aware that somatic OT immunoreactivity in PVN neurons reaches a daily low during the early light phase [56]”. We now add to this the idea, suggested by the reviewer, that “In situ hybridization targeting oxytocin mRNA, rather than immunohistochemistry, may provide a more reliable assessment of expression specificity.”

      (2) Although the authors' interpretation is generally consistent with the data presented, their main conclusions rely heavily on observational findings. Moreover, optogenetic stimulation of PVNOT neurons failed to robustly recapitulate behavioral state transitions (Figs. 6D and S5B). Further interventional experiments will be necessary to more rigorously test the authors' interpretation and to establish mechanistic insight into the causal relationship between PVNOT activity and rest-to-active transitions. In particular, loss-of-function approaches targeting the PVNOT system, such as OXTR antagonism, inhibitory DREADDs, or cell-type-specific ablation, will be essential to determine whether perturbation of this system alters behavioral state transitions These points should be addressed in future studies.

      Reviewer 1 brought up a similar concern. We have added to the Discussion/Limitations and caveats to address this.

      “Third, although we show that PVNOT neurons are sufficient to drive thermogenic and behavioral transitions (Fig. 6), we did not perform acute loss-of-function experiments. Such experiments are warranted because decreases in baseline PVNOT calcium activity were associated with transitions toward the onset of quiescence (Fig. 3I-L), suggesting this system may bidirectionally regulate thermo-behavioural state. A tractable next step would be to optogenetically inhibit PVNOT neurons during established rest bouts, delivered at pseudo-random times across the light and rest phase and analyzed post hoc by behavioral state and body temperature; we predict that silencing during rest would prolong the average duration of rest bouts and delay the onset of activity and thermogenesis. Pairing the inhibition with selective oxytocin antagonist (such as L-368,899), would further test whether the thermogenic and autonomic components of these transitions are oxytocin receptor dependent rather than driven by glutamate released by the same neurons.”

      Note: as described in the previous response to reviewers, we have tried inhibitory DREADDs in this system and have concluded that it is of little value because delivering DREADD ligand requires handing the animals for an IP injection—a procedure that disrupts sleep/rest and induces stress hyperthermia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors have answered our criticisms and can proceed as they chose. This is an important paper, and it is the author's choice whether to develop their research here or in a subsequent paper.

      Thank you.

      Reviewer #2 (Recommendations for the authors):

      I thank the authors for citing my pre-print, as suggested by Reviewer 1. The paper has now been published and the authors may like to cite the published version (doi.org/10.1038/s41593-026-02224-0).

      Thank you.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors now interpret their results as indicating that PVNOT activity biases the system toward state transition (from rest to active), rather than acting as a deterministic trigger. This interpretation is reasonable. However, the wording "PVNOT peaks (or neurons) predict transitions to behavioral arousal and thermogenesis" may be misleading. If arousal and thermogenesis occur in more than 80% of cases following PVNOT peaks, then such peaks could reasonably be described as "being predicted". Otherwise, the terminology should be revised for clarity.

      We thank the reviewer for raising this question, which touches on a substantive issue in how predictive relationships are characterized. We agree that "predicts" can misleadingly imply a high positive predictive value: i.e., that a large fraction of peaks are followed by transitions.

      This is not the claim we intend, nor is it the appropriate statistical criterion. A variable is predictive when it shifts the conditional probability (or, here, the conditional distribution) of the outcome relative to its base rate — the criterion underlying likelihood ratios, relative risk, and signal-detection measures — rather than when it exceeds an absolute occurrence threshold such as 80%. By this standard, a peak can be informative even if transitions do not follow the majority of peaks, provided transitions are substantially more likely (or thermogenically warmer) when a peak precedes them than when one does not.

      Our data support precisely this. The logistic regression shows peaks are much more probable immediately before rest offset than at other transitions or at baseline, and our new analysis shows that transitions preceded by peaks carry significantly larger post-offset Tb increases than those without. We are not claiming peaks act as a deterministic trigger, and we agree with the reviewer that they are not present before every transition.

      To keep our language aligned with these results, we have revised the wording to avoid "predict" where it could imply high hit-rate determinism, replacing it with comparative phrasing. Accordingly, we have revised the terminology throughout the manuscript: where a claim concerns timing, we now state that peaks “precede” transitions. We have removed “predict”/”predictive” from the section heading, figure legend, introduction and results as follows.

      “Then, we discovered that PVNOT calcium dynamics during huddling were associated with increased likelihood of transitions to body warming and arousal.”

      “PVNOT neuronal activity precedes transitions towards thermogenesis and behavioral arousal in social and non-social contexts.”

      Fig. 3 legend title: “PVNOT peaks are associated with increased likelihood of thermogenic rest-to-active transitions.”

      “Thus, PVNOT peaks are at least five-fold more likely to occur near the offset of quiescence/quiescent compared to onset, and signal an increase in physical activity—a correlate of behavioral arousal 53 and a means of increasing metabolic rate and Tb [26]”

      “Thus, for nesting and active huddling, PVNOT peaks are two- to three- fold more likely to occur at bout onset than offset.” Dropping flagged word here lol.

      “Together these results suggest that elevated PVNOT activity dynamics precede the offset of two rest states (quiescence and quiescent huddling) by approximately 100 seconds, and the onset of two post-quiescence active states (nesting and active huddling) by around 20 seconds, in solo and paired mice respectively.”

      “Moreover, PVNOT peaks aligned with the low point of a U-shaped body temperature profile: on average, Tb decreased before, and increased after, the time of the calcium peak in both solo and paired conditions (Fig. 3O,R). Together, these results suggest that PVN<sup>OT</sup> peaks occur during a low Tb trough and mark a subsequent rise in Tb.”

      (2) Regarding the 400-sec latency of BAT surface temperature increases following optogenetic stimulation, the authors now attribute this delay to slow peptidergic transmission. However, the authors should consider prior findings showing that BAT temperature increased immediately following optogenetic stimulation of PVN→rMR oxytocin neurons in anesthetized rats (Fukushima et al., 2022).

      My hunch is that doing this in anesthetized rats gives a stronger signal to noise… not sure if I can back that up though.

      At the least we can add a sentence that says “rMR oxytocin neurons immediately increases BAT temperature, while infusion of OXT or NMDA in the rMR results in BAT temperature increases after approximately one minute…” (see Fig. 3,4,5).

      We thank the reviewer for redirecting us to Fukushima et al. (2022). We note, however, that in that study the fast-responding variable was BAT sympathetic nerve activity, whereas the BAT temperature itself rose over several minutes following both optogenetic stimulation (their Fig. 4F, quantified at 5 and 10 minutes) and focal rMR infusion of oxytocin or NDMA (their Fig. 5, multiminute traces). This thermal timescale is comparable to the one we observe.

      The remaining difference could reflect methodological differences: we stimulated PVNOT somata rather than rMR terminals, measured intrascapular surface rather than BAT temperature directly, and recorded in awake, freely behaving animals (rather than anesthetized animals) in which competing thermoeffector and behavioral processes are active. Consistent with a methodological basis for the delay, focal infusion of oxytocin or NMDA into the rMR in that study increased BAT temperature over roughly a minute (Fukushima et al., 2022). Slow, diffuse peptidergic neuromodulation may further contribute, oxytocin is released from large dense-core vesicles and can act over extended time scales (Ludwig and Leng, 2006; Parmaksiz and Kim, 2025; Qian et al., 2023), although our data cannot isolate this mechanism from the factors above or from fast glutamatergic co-transmission that likely accompanies PVNOT activation (Hrabovszky and Liposits, 2008).

      (3) In the previous review, clarification was requested regarding the rationale and histological basis for intravenous FluoroGold injection. While the authors have now added methodological details, they should also incorporate the following explanatory text (previously provided in their rebuttal) into the manuscript for readers unfamiliar with PVN histological analyses:

      "Intravenous injection of FluoroGold (FG) was used to histologically differentiate between magnocellular and parvicellular oxytocin neurons in the PVN. Because the posterior pituitary is located outside the blood-brain barrier, i.v. FG is selectively taken up by terminals of magnocellular neurons and retrogradely transported to their cell bodies. This allows us to infer the neuroanatomical identity (magno- vs. parvicellular) of the PVNOT neurons of interest."

      We thank the reviewer for this suggestion. We have added the explanatory text to the results subsection, “PVN<sup>OT</sup> cellular projections to the rMR”. The text now reads: “rMR cell types in mice, we used FluoroGold (FG to disambiguate magno- vs. parvocellular PVN<sup>OT</sup> projections [67] (Fig. S6A-C). Because the posterior pituitary is located outside the blood-brain barrier, intravenous FG is selectively taken up by terminals of magnocellular neurons and retrogradely transported to their cell bodies. This allows us to infer the neuroanatomical identity (magno- vs. parvicellular) of the PVN<sup>OT</sup> neurons of interest.”

    1. eLife Assessment

      The manuscript presents a primary important finding, namely that microbial riboflavin-derived MR1 ligands are pharmacological activators of human MAIT cells and that MR1 ligand stimulation enhances MAIT-mediated tumor killing across multiple tumor models. While the evidence presented is generally solid, there are some limitations of the in vitro and in vivo models used, as well as some overstatements about the broader applicability of the research, which should be revised.

    2. Reviewer #1 (Public review):

      The manuscript from Zhu et al. identifies microbial riboflavin-derived MR1 ligands as potent pharmacological activators of human MAIT cells and provides evidence that MR1 ligand stimulation can enhance MAIT-mediated tumor killing across multiple solid tumor models. The study is conceptually interesting and supported by a broad combination of human primary samples, tumor cell lines, 3D models, SC transcriptomics, and xenograft experiments. Overall, the data largely support the central conclusion that MR1 ligand stimulation can strongly activate human MAIT cells and enhance anti-tumor cytotoxicity. However, the broader conclusions concerning endogenous MAIT mobilization, tumor specificity, and translational potential are not yet fully supported by the current data and should either be moderated or addressed with additional experiments.

      Comments:

      (1) The authors use one-way ANOVA throughout the manuscript, but this may not be appropriate for some analyses, particularly when multiple experimental factors are present and their interaction effects need to be considered. For example, Figure 3f appears to involve multiple factors, for which a two-way ANOVA may be more appropriate. Similar issues may apply to other panels.

      (2) In Figure 3f, the authors show data from patients #1 and #2 and state that the experiment is representative of three experiments. What does the reported "n=4" represent in this figure?

      (3) There appears to be a discrepancy between Figure 3f and Supplementary Figure 3b. The two panels appear to use the same treatment conditions and the same label, and both appear to use patient #1 samples, yet the reported values are different. Please clarify the experimental design and explain the reason for this discrepancy.

      In addition, the gating strategy used to define live tumor cells should be clearly described in the figure legend and/or Methods. The authors define "live tumor cells" as MR1/5-OP-RU tetramer-CD45- cells. However, in primary liver tumor samples, the CD45-/tetramer- population may contain other non-hematopoietic cells, such as fibroblasts, and therefore may not exclusively represent tumor cells. The authors should clarify whether additional tumor-specific markers or other criteria were used. The gating strategies for the relevant flow cytometry experiments should be provided in the Supplementary figures.

      (4) I have some concerns regarding the claims of "selective activation of anti-tumor inflammatory pathways rather than generalized cytokine release" and "avoiding induction of tumor-supportive mediators." The authors show that MAIT cells stimulated with 5-OP-RU can substantially reduce tumor cell viability. Therefore, the cellular composition of the co-culture is likely to change considerably during the assay, which may affect the absolute levels of cytokines and other soluble mediators detected. For example, reduced tumor cell numbers could lead to lower production of tumor-derived factors such as VEGF, potentially confounding the interpretation that these mediators are not induced by MAIT activation. The authors should consider whether cytokine measurements have been normalized to viable cell numbers or otherwise account for differences in tumor cell abundance.

      (5) The in vivo tumor models may show substantial variability between independent experiments. Rather than presenting a single representative experiment, the authors should consider showing pooled data from all independent experiments, with the total number of mice clearly indicated.

      (6) Why did the authors use an MR1-overexpressing tumor cell line for the in vivo studies rather than the parental cells with endogenous MR1 expression, together with MR1-KO cells as a negative control? The authors demonstrate that MR1 is detectable across multiple tumor cell lines and that endogenous MR1 expression is sufficient to support MAIT-mediated killing in vitro. Moreover, MR1 overexpression substantially enhances tumor cell susceptibility to MAIT-mediated killing. Therefore, it is unclear whether the strong therapeutic efficacy observed in vivo reflects physiologically relevant MR1 expression or is driven by artificially elevated MR1 expression. An in vivo comparison using parental and MR1-KO tumor cells would substantially strengthen the translational relevance and establish whether the therapeutic effect can be achieved at endogenous levels of MR1.

      (7) How is tumor specificity of MAIT achieved ? The authors propose that MAIT-cell activation by MR1 ligands provides an antigen-independent approach for tumor targeting. However, MR1 is broadly expressed and is not tumor specific. While the relative sparing of T and B cells in Figure 7B provides some evidence of cell-type selectivity, this does not establish tumor versus normal tissue specificity. It remains unclear whether activated MAIT cells can discriminate tumor cells from other normal MR1-expressing cells and tissues. This raises an important question regarding the potential systemic toxicity of MAIT cells activated by systemic administration of 5-OP-RU. In particular, could other MR1-expressing cells be targeted when a large number of MAIT cells are simultaneously activated? The authors should consider assessing systemic toxicity in vivo, for example by examining serum ALT/AST levels and tissue pathology, and/or by evaluating the effects of MAIT + 5-OP-RU in tumor-free animals. At least, the potential specificity and safety limitations of systemic MR1 agonism should be discussed.

    3. Reviewer #2 (Public review):

      The manuscript by Zhu et al. describes MAIT cell activation by riboflavin metabolites presented by MR1. The authors provide solid evidence for this activation and anti-cancer functional consequence using an array of selected cell lines, primary ex vivo and engineered xenograft models. Broadly, the results are thorough and well controlled, and provide a highly informative insight into the metabolite-MAIT-cancer cell interactions. However, the majority of this work is undertaken using models that preferentially express key targets, and whilst still useful, the (current) broader implications of this research are overstated. Additionally, the suggested MAIT modulation of the tumor microenvironment requires clarification.

      Major Comments:

      (1) In Figures 2b-d, the authors suggest microbial metabolite stimulation of PBMC cultures increased MAIT cell frequency up to 60%. Whilst their flow data is compelling, the frequency of one population can be influenced by changes in other populations. A form of absolute or relative-to-total count should be used.

      (2) The statements regarding cytokine induction in Figure 4e are too strong; many of those inflammatory cytokines are not automatically and consistently tumour-suppressive. The line 299 '...were not induced' may just reflect death of tumor cells. It would be useful to include tumour cell-only controls in Figure 4.

      (3) Figure 7 is interesting, but the authors' conclusion that MAIT+5-OP-RU controls the tumor microenvironment is not robustly supported by their evidence.

      a) It is not clear how CD14+ cells established a sustained suppressive environment.

      b) It is not clear how the peritoneal addition of microbial metabolites 'significantly enhanced MAIT-mediated tumor control'. The authors show that the addition of 5-OP-RU reduced the number of GFP-expressing tumour cells present in peritoneal lavage fluid. There is limited evidence to suggest this occurs through MAIT cells or MR1 in this figure.

      c) It is difficult to draw conclusions from peritoneal lavage flow when some experimental groups received cells IP, but then all groups were equally assessed for key populations, and all data are presented as frequencies. The authors should use absolute counts (or similar) to appropriately show changes in cell populations to account for varying total/live/cd45+ cell compartments.

      d) It would be necessary at a minimum to include 5-OP-RU-only controls, and ideally include MR1 blocking or the cancer line with MR1 removed. Alongside this, the authors should substantially reduce the strength of their statements on microbial metabolite-MAIT suppression of the tumor microenvironment.

    4. Author response:

      Reviewer #1 (Public review):

      The manuscript from Zhu et al. identifies microbial riboflavin-derived MR1 ligands as potent pharmacological activators of human MAIT cells and provides evidence that MR1 ligand stimulation can enhance MAIT-mediated tumor killing across multiple solid tumor models. The study is conceptually interesting and supported by a broad combination of human primary samples, tumor cell lines, 3D models, SC transcriptomics, and xenograft experiments. Overall, the data largely support the central conclusion that MR1 ligand stimulation can strongly activate human MAIT cells and enhance anti-tumor cytotoxicity. However, the broader conclusions concerning endogenous MAIT mobilization, tumor specificity, and translational potential are not yet fully supported by the current data and should either be moderated or addressed with additional experiments.

      We thank the reviewer for the positive feedback. We will address all comments and suggestions point by point.

      Comments:

      (1) The authors use one-way ANOVA throughout the manuscript, but this may not be appropriate for some analyses, particularly when multiple experimental factors are present and their interaction effects need to be considered. For example, Figure 3f appears to involve multiple factors, for which a two-way ANOVA may be more appropriate. Similar issues may apply to other panels.

      We thank the reviewer for this valuable comment. We will carefully review the statistical analyses and revise the tests as appropriate, including the use of two-way ANOVA where multiple experimental factors are present.

      (2) In Figure 3f, the authors show data from patients #1 and #2 and state that the experiment is representative of three experiments. What does the reported "n=4" represent in this figure?

      We thank the reviewer for this valuable comment. We will revise the figure legend to clearly define what the reported n = 4 represents.

      (3) There appears to be a discrepancy between Figure 3f and Supplementary Figure 3b. The two panels appear to use the same treatment conditions and the same label, and both appear to use patient #1 samples, yet the reported values are different. Please clarify the experimental design and explain the reason for this discrepancy.

      In addition, the gating strategy used to define live tumor cells should be clearly described in the figure legend and/or Methods. The authors define "live tumor cells" as MR1/5-OP-RU tetramer-CD45- cells. However, in primary liver tumor samples, the CD45-/tetramer- population may contain other non-hematopoietic cells, such as fibroblasts, and therefore may not exclusively represent tumor cells. The authors should clarify whether additional tumor-specific markers or other criteria were used. The gating strategies for the relevant flow cytometry experiments should be provided in the Supplementary figures.

      We thank the reviewer for this valuable comment. Figure 3f (patient #2) and Supplementary Figure 3b (patient #1) were generated using samples from different patients. We will clarify this in the revised manuscript and provide the relevant gating strategies in the Supplementary Information.

      (4) I have some concerns regarding the claims of "selective activation of anti-tumor inflammatory pathways rather than generalized cytokine release" and "avoiding induction of tumor-supportive mediators." The authors show that MAIT cells stimulated with 5-OP-RU can substantially reduce tumor cell viability. Therefore, the cellular composition of the co-culture is likely to change considerably during the assay, which may affect the absolute levels of cytokines and other soluble mediators detected. For example, reduced tumor cell numbers could lead to lower production of tumor-derived factors such as VEGF, potentially confounding the interpretation that these mediators are not induced by MAIT activation. The authors should consider whether cytokine measurements have been normalized to viable cell numbers or otherwise account for differences in tumor cell abundance.

      We thank the reviewer for this important comment. We agree that differences in tumor cell abundance may affect cytokine measurements. We will moderate our claims accordingly and acknowledge this limitation in the revised manuscript.

      (5) The in vivo tumor models may show substantial variability between independent experiments. Rather than presenting a single representative experiment, the authors should consider showing pooled data from all independent experiments, with the total number of mice clearly indicated.

      We thank the reviewer for this valuable comment. We will provide pooled data from all independent in vivo experiments and clearly indicate the total number of mice.

      (6) Why did the authors use an MR1-overexpressing tumor cell line for the in vivo studies rather than the parental cells with endogenous MR1 expression, together with MR1-KO cells as a negative control? The authors demonstrate that MR1 is detectable across multiple tumor cell lines and that endogenous MR1 expression is sufficient to support MAIT-mediated killing in vitro. Moreover, MR1 overexpression substantially enhances tumor cell susceptibility to MAIT-mediated killing. Therefore, it is unclear whether the strong therapeutic efficacy observed in vivo reflects physiologically relevant MR1 expression or is driven by artificially elevated MR1 expression. An in vivo comparison using parental and MR1-KO tumor cells would substantially strengthen the translational relevance and establish whether the therapeutic effect can be achieved at endogenous levels of MR1.

      We thank the reviewer for this important comment. We agree that comparison with endogenous MR1 expression would strengthen the translational relevance of our findings. We will include new in vivo experiment comparing parental tumor cells.

      (7) How is tumor specificity of MAIT achieved? The authors propose that MAIT-cell activation by MR1 ligands provides an antigen-independent approach for tumor targeting. However, MR1 is broadly expressed and is not tumor specific. While the relative sparing of T and B cells in Figure 7B provides some evidence of cell-type selectivity, this does not establish tumor versus normal tissue specificity. It remains unclear whether activated MAIT cells can discriminate tumor cells from other normal MR1-expressing cells and tissues. This raises an important question regarding the potential systemic toxicity of MAIT cells activated by systemic administration of 5-OP-RU. In particular, could other MR1-expressing cells be targeted when a large number of MAIT cells are simultaneously activated? The authors should consider assessing systemic toxicity in vivo, for example by examining serum ALT/AST levels and tissue pathology, and/or by evaluating the effects of MAIT + 5-OP-RU in tumor-free animals. At least, the potential specificity and safety limitations of systemic MR1 agonism should be discussed.

      We thank the reviewer for this important comment. To further evaluate the potential safety concerns associated with systemic MR1 ligand stimulation, we will include a new experiment assessing the effects of MAIT cells plus 5-OP-RU in tumor-free animals. We will also discuss the potential specificity and safety limitations of systemic MR1 agonism in the revised manuscript.

      Reviewer #2 (Public review):

      The manuscript by Zhu et al. describes MAIT cell activation by riboflavin metabolites presented by MR1. The authors provide solid evidence for this activation and anti-cancer functional consequence using an array of selected cell lines, primary ex vivo and engineered xenograft models. Broadly, the results are thorough and well controlled, and provide a highly informative insight into the metabolite-MAIT-cancer cell interactions. However, the majority of this work is undertaken using models that preferentially express key targets, and whilst still useful, the (current) broader implications of this research are overstated. Additionally, the suggested MAIT modulation of the tumor microenvironment requires clarification.

      We thank the reviewer for the positive feedback. We will address all comments and suggestions point by point.

      Major Comments:

      (1) In Figures 2b-d, the authors suggest microbial metabolite stimulation of PBMC cultures increased MAIT cell frequency up to 60%. Whilst their flow data is compelling, the frequency of one population can be influenced by changes in other populations. A form of absolute or relative-to-total count should be used.

      We thank the reviewer for this valuable comment. We will provide absolute cell counts and/or normalized data to more accurately assess changes in MAIT cell frequency.

      (2) The statements regarding cytokine induction in Figure 4e are too strong; many of those inflammatory cytokines are not automatically and consistently tumour-suppressive. The line 299 '...were not induced' may just reflect death of tumor cells. It would be useful to include tumour cell-only controls in Figure 4.

      We thank the reviewer for this valuable comment. We agree that the statements regarding cytokine induction should be interpreted more cautiously. We will revise the relevant claims.

      (3) Figure 7 is interesting, but the authors' conclusion that MAIT+5-OP-RU controls the tumor microenvironment is not robustly supported by their evidence.

      (a) It is not clear how CD14+ cells established a sustained suppressive environment.

      We thank the reviewer for this valuable comment. We will include additional experiments to further characterize the contribution of CD14+ cells to the observed suppressive environment.

      (b) It is not clear how the peritoneal addition of microbial metabolites 'significantly enhanced MAIT-mediated tumor control'. The authors show that the addition of 5-OP-RU reduced the number of GFP-expressing tumour cells present in peritoneal lavage fluid. There is limited evidence to suggest this occurs through MAIT cells or MR1 in this figure.

      We thank the reviewer for this valuable comment. We will include additional T-cell and T-cell + 5-OP-RU control groups to further determine the contribution of MAIT cells to the observed tumor control.

      (c) It is difficult to draw conclusions from peritoneal lavage flow when some experimental groups received cells IP, but then all groups were equally assessed for key populations, and all data are presented as frequencies. The authors should use absolute counts (or similar) to appropriately show changes in cell populations to account for varying total/live/cd45+ cell compartments.

      We thank the reviewer for this valuable comment. We will provide absolute cell counts, in addition to frequencies, to account for differences in total and viable CD45+ cell numbers.

      (d) It would be necessary at a minimum to include 5-OP-RU-only controls, and ideally include MR1 blocking or the cancer line with MR1 removed. Alongside this, the authors should substantially reduce the strength of their statements on microbial metabolite-MAIT suppression of the tumor microenvironment.

      We thank the reviewer for this valuable comment. We will include additional T-cell and T-cell + 5-OP-RU control groups and will substantially moderate our statements regarding microbial metabolite-mediated modulation of the tumor microenvironment.

    1. eLife Assessment

      This study shows that partial cone photoreceptor loss induces pathway-specific circuit remodeling in the mouse retina, with alpha OFF-sustained and OFF-transient retinal ganglion cells adapting differently through changes in their pre- and postsynaptic circuits. The results are valuable because they provide a key understanding of the diversity of circuit remodeling in retinal degeneration. The data are convincing, although clearer reporting of the numbers of independent animals and retinas, a more rigorous distinction between compensation and circuit change, and discussion of the functional consequences would strengthen the mechanistic insights.

    2. Reviewer #1 (Public review):

      Summary:

      Lee et al. investigate how parallel retinal pathways respond to a common loss of photoreceptor input. The authors induce partial cone loss in adult mice and compare the functional responses of sustained OFF alpha (sOFFa) and transient OFF alpha (tOFFa) ganglion cells, together with changes in their presynaptic circuits. Using targeted patch-clamp recordings, linear-nonlinear analyses, pharmacological dissection of inhibitory inputs, and quantitative synaptic imaging, they show that the two pathways do not respond uniformly to cone loss. tOFFa ganglion cells exhibit more extensive changes in spatiotemporal receptive fields than sOFFa ganglion cells, with contributions from excitatory transmission, presynaptic glycinergic inhibition, direct GABAergic and glycinergic inhibition, and intrinsic properties. At the same time, transformations between synaptic input and spike output partially preserve ganglion cell signaling despite the loss of cones.

      Strengths:

      This is a technically careful and high-quality study. The comparison of two well-defined ganglion cell types and their dominant bipolar-cell pathways provides an unusually detailed view of where circuit modifications arise following a shared perturbation. The combination of recordings at successive stages of signal processing, pharmacological manipulations, and synaptic imaging is a particular strength. The use of partial stimulation in control retina also helps distinguish the immediate consequence of reduced input from subsequent circuit changes. The resulting conclusion that common photoreceptor loss produces pathway-specific forms of remodeling rather than a uniform retinal response is interesting and well supported. The work adds to our understanding of the diversity and circuit specificity of responses to retinal degeneration.

      Weaknesses:

      The principal limitations concern the precision of some mechanistic interpretations rather than the central observation of pathway-specific remodeling. First, the framework used to classify effects as compensation or circuit change sometimes treats the absence of a statistically significant difference as evidence that two conditions are equivalent. Second, the numbers of animals and retinas contributing to the main physiological and anatomical comparisons are not consistently reported, making it difficult to evaluate the independence of measurements obtained from multiple cells, images, or synaptic puncta. Finally, the consequences of the observed remodeling for the visual signals carried by these pathways remain unclear. This is particularly relevant for tOFFa ganglion cells, which have been implicated in responses to looming or approaching dark objects. The altered temporal filtering, center-surround organization, and input-output transformation could preserve, degrade, or otherwise transform such signals. These issues qualify the mechanistic and functional interpretation but do not substantially weaken the main conclusion that the two pathways respond differently to partial cone loss.

    3. Reviewer #2 (Public review):

      Summary:

      This is an elegant, rigorous, and thought-provoking study that examines how different neural circuits are altered in response to loss of a common sensory input. To study this question, the authors use the mouse retina as a model system to investigate how downstream retinal circuits undergo modifications following a well-controlled partial loss of cone photoreceptors.

      Strengths:

      The experiments were conducted with a high degree of rigor, and the authors carefully considered and implemented appropriate controls throughout the study. Multiple parameters were tested, including pharmacological approaches to assess responses from different ganglion cell types. In addition, the authors complemented their functional data with confocal imaging to further support their findings. Overall, this is a well-written paper that provides a thorough analysis demonstrating how two similar ganglion cell types undergo distinct adaptations (i.e., compensation versus remodeling) in response to the loss of the same sensory input.

      Weaknesses:

      No additional experiments are needed. However, the authors may wish to consider the following points:

      (1) Do the differences in compensation versus remodeling observed in ganglion cells reflect changes in the OPL? Different bipolar types may remodel their dendrites and form aberrant contacts with rods in the absence of cones. However, this would be challenging to test because there are currently no good markers for different bipolar types.

      (2) It would be interesting to determine whether these functional changes can be detected at the transcriptomic level or whether they are mediated primarily through post-translational modifications.

    1. eLife Assessment

      This study presents a useful finding on using diverse experimental systems to understand how neuromodulatory signals shape glial inflammatory signaling; however, the strength of evidence is inadequate. The astrocyte-enrichment method used may permit contamination by microglia, oligodendrocyte-lineage cells, or neurons, complicating attribution of TNF expression specifically to astrocytes. This concern is compounded by the strong microglial response to Gi manipulation and the lack of quantitative validation of chemogenetic cell-type specificity in vivo. These weaknesses have hindered further evaluation of the claims.

    2. Reviewer #1 (Public review):

      In this manuscript, the authors explore whether GPCR signaling in astrocytes affects the production of TNF by astrocytes and, to a lesser extent, microglia. Unfortunately, the method used by the authors to acquire astrocyte-enriched cultures is known to result in meaningful rates of contamination by myeloid cells (microglia and others), oligodendrocyte-lineage cells, and neurons. Alternative methods of generating highly enriched astrocyte cultures, as well as purifying astrocytes with little to no neuronal or myeloid contamination across age and brain regions, have shown no evidence of TNF expression by astrocytes (Zhang et al., J Neurosci, 2014; Zhang et al., Neuron, 2016; Clarke et al., PNAS, 2018). In fact, the paper cited by the authors as demonstrating differences between human and rodent astrocytes found no evidence of TNF expression in immature or mature human astrocytes (Zhang et al., Neuron, 2016). The idea that the majority of the observed TNF transcriptomic signal, at least in culture, comes from myeloid or neuronal contamination also aligns with the authors' observation that myeloid-enriched cultures act identically to astrocyte-enriched cultures.

      The authors also use a GFAP virus to drive GPCR signaling in astrocytes and neuronal progenitor cells in their cultures, but, given that these cultures are known to have meaningful contamination by other cell types, such signaling could be due to astrocyte → microglia/neuron signaling or other multicellular pathways that cannot be excluded. Similar concerns mean that we cannot assume the effect of DREADD activation of astrocytes in vivo (Figure 6) reflects a bulk change in TNF expression driven by astrocyte-specific changes rather than by multicellular signaling.

      The most compelling evidence for their claim of astrocyte TNF expression comes from the human-induced astrocytes. However, their antibody staining is not sufficient to claim these cells are truly astrocyte-like. Antibody staining is highly prone to non-specificity, as highlighted by the fact that their ALDH1L1 antibody staining appears perfectly nuclear despite ALDH1L1 being a cytoplasmic protein.

      To address both the purity concerns of the astrocyte-enriched cultures and the concerns about the astrocyte identity of the induced astrocytes, the authors should perform RNA sequencing. By profiling gene expression in these cultures at the genome-wide level, readers can truly assess the degree of contamination and thus the likelihood of the proposed mechanism (i.e., astrocyte-specific TNF production). Importantly, previous studies have suggested that very little neuronal and myeloid contamination is required to dramatically change cellular responses (Foo et al., Neuron, 2011; Liddelow et al., Nature, 2017).

    3. Reviewer #2 (Public review):

      Summary:

      Abbasi et al. examine how signaling through the major G-protein pathways (Gs, Gq, and Gi) influences tumor necrosis factor expression in astrocytes and microglia. Using a combination of pharmacological receptor activation, chemogenetic manipulation, primary rodent glial cultures, human induced pluripotent stem cell-derived astrocytes, and an in vivo astrocyte-targeted Gi manipulation, the authors report a broadly consistent pattern in which Gs- and Gq-associated signaling reduces tumor necrosis factor expression, whereas Gi signaling increases it. The study's cross-species and cross-preparation design, spanning astrocytes and microglia as well as in vitro and in vivo systems, provides a potentially valuable framework for understanding how neuromodulatory pathways may regulate glial inflammatory signaling.

      Strengths:

      A major strength of the study is the breadth of experimental systems used, which includes primary rat glia, human induced pluripotent stem cell-derived astrocytes, and an in vivo manipulation, allowing for comparison across species and levels of biological complexity. The use of chemogenetic receptors in astrocytes provides relatively direct control over Gq and Gi signaling, and these experiments yield consistent effects on both tumor necrosis factor messenger RNA and protein, strengthening the internal validity of the astrocyte findings. The observation that similar directional effects are seen in human-derived astrocytes and in microglial cultures further supports the idea that aspects of this regulatory relationship may be conserved across glial cell types. More broadly, the study addresses an important and timely question about how neuromodulatory signaling pathways interface with glial inflammatory outputs, and it generates a coherent set of observations that could serve as a foundation for more mechanistic work.

      Weaknesses:

      The central claim that Gs, Gq, and Gi signaling broadly and directly constitute a general regulatory code for tumor necrosis factor expression is more expansive than the current evidence fully supports. In particular, the evidence for Gs-dependent effects is indirect, relying on beta-adrenergic receptor activation and forskolin-mediated adenylyl cyclase stimulation rather than direct manipulation of Gs itself, leaving uncertainty about pathway specificity. More generally, the use of different endogenous receptors to represent each G-protein class in microglia complicates interpretation, since individual receptors may engage additional signaling pathways beyond their canonical G-protein coupling, limiting the extent to which the results can be attributed to G-protein class alone.

      The in vivo experiment also does not definitively establish the cellular source of the observed increase in tumor necrosis factor, as measurements are taken from bulk cortical tissue following astrocyte-targeted Gi activation. This leaves open the possibility that the observed changes arise indirectly from other cell types, particularly microglia, which are shown elsewhere in the study to be strongly responsive to Gi-related manipulations. In addition, the specificity of chemogenetic expression in vivo is not quantitatively demonstrated, further limiting cell-type attribution.

      There are also important issues related to experimental design and statistical interpretation. Across several experiments, it is unclear whether reported sample sizes reflect independent biological replicates, technical replicates, or imaging fields, which is especially consequential for the human induced pluripotent stem cell-derived astrocyte experiments where donor-level independence is not clearly established. The in vivo design also appears to treat hemispheres as independent observations despite their paired nature, which may inflate statistical independence given the small sample size.

      Finally, several conclusions would benefit from more cautious framing. The data support differential regulation of tumor necrosis factor relative to interleukin-1 rather than strict cytokine specificity, and measurements based solely on messenger RNA should not be interpreted as direct evidence of cytokine production. The comparison between glial signaling effects and neuronal excitation or inhibition also juxtaposes fundamentally different biological readouts and should not be interpreted as a direct functional opposition. Overall, while the study provides interesting and potentially important observations, the broader pathway-level and cell-type-specific conclusions are not yet fully established by the current experimental evidence.

    4. Author response:

      Reviewer 1 is concerned that our astrocyte enriched cultures have significant contamination of microglia or other myeloid cells, OPCs (and related cells) and neurons. Further, they assert that purified astrocytes do not express TNF.

      That astrocytes can’t make TNF directly contradicts our previous paper (Heir, et al., JNeurosci, 2024) showing that the TNF driving homeostatic plasticity is generated by astrocytes. The reviewer seems to want to dispute that paper, which is not really the topic of the current paper (which covers the regulation of TNF production, not whether particular cell types make TNF). The Nedergaard group also saw TNF release from human astrocytes (Wang, et al., 2006). The papers cited by the reviewer (all from the same group) rely on RNAseq data, which has limited depth and cannot distinguish if something is not expressed or simply below threshold. Further, as these datasets were generated from astrocytes isolated from brain (which has normal levels of activity), the astrocytic TNF expression would be very low. Plenty of data supports that astrocytes can express TNF when stimulated (by LPS or other activators), and our previous paper shows that this is also true when neuronal activity is blocked (or absent). But at baseline, astrocyte TNF is quite low and likely undetectable as assayed in those papers.

      Here we are using highly purified astrocyte cultures. The reviewer is perhaps unfamiliar with the type of cultures we are using. Given that we use mechanical disruption to remove neurons, followed after 1-2 weeks by shaking to remove microglia, and finally cell passaging, all before experiments, it is surprising that the reviewer thinks there could be neuronal contamination. Neurons cannot survive that procedure, and we do not observe them by morphology or immunostaining, nor see neuronal markers by qPCR.  The microglial contamination is also minimal, as noted in the manuscript, with qPCR for microglia markers is at noise levels (Iba1 Ct value of 34.6), while GFAP shows robust expression (Ct of 16.8; >100,00 fold more than Iba1). But it is possible, if unlikely, that some small number of microglia are making a lot of TNF. However, treating our cultures with the microglia toxin LME (used in Heir, et al., 2024) did not alter our results, further suggesting microglia are not contributing here. Other contaminating cell types (in the OPC lineage, for example) are also possible. However, the majority of cells in our astrocyte-enriched culture are positive for TNF by immunostaining (done while blocking protein export, to prevent any release of TNF). This makes it highly probably that astrocytes are producing TNF (and this production is regulated by g-protein signaling). To verify this, we will show TNF protein in cells co-labeled with astrocyte markers in our upcoming revision of the paper. This will definitively identify astrocytes as producing TNF in these rat cultures. With the human iPSC-derived astrocytes, microglial contamination is not possible (this requires a completely different differentiation protocol). We agree the ALDH1L1 labeling is not as expected, but it is unclear if this is an antibody issue or mis-localized protein. However, the cells also label with S100beta and GFAP, making the astrocyte identity the most likely option by far. We have additional qPCR data showing expression of ALDH1L1 by these cells, in addition to the other astrocyte markers (which will also be added to the revision). The in vivo situation is more complex, and we can’t exclude that astrocyte-DREADD signaling here indirectly alters TNF production in other cells. However, given the direct regulation of astrocyte TNF production in culture, the simplest explanation is that the same is occurring in vivo.

      Reviewer 2 was concerned that the Gs data was indirect and the use of pharmacological approaches with microglia. As for Gs signaling, it is a bit unclear what the reviewer is suggesting as an alternative hypothesis. We activate the Gs-coupled beta-adrenergic receptor to reduce TNF levels and get the same effect by activating adenylyl cyclase, the canonical downstream pathway from Gs-coupled receptors. While it is possible that beta-adrenergic receptors could have alternate coupling or that Gs activation acts on additional pathways, it seems odd to argue that Gs would not be working through adenylyl cyclase activation when activating the cyclase yields the same response. Certainly the most parsimonious explanation is that Gs-couple receptors act through adenylyl cyclase to reduce TNF production.

      As for the use of pharmacology with microglia, this was the more expedient solution to the difficulty of using AAV virus on microglia. Gathering the necessary Cre and conditional DREADD lines was an impractical solution in terms of time and resources. However, the pharmacology of these receptors is well characterized, as is the g-protein coupling. Given that the results are identical to the results from more specific manipulations in astrocytes, it seems reasonable to conclude that there is a common pattern of GPCR regulation of TNF production. The criticism that non-canonical pathways can be activated by these receptors seems equally true for the DREADDs, as these are just GPCRs with mutated binding sites. If anything, the forskolin experiment is the most specific, yet the reviewer dislikes this approach. The overall consistency of the responses, whether due to DREADD activation, native receptors or direct activation of adenylyl cyclase, is the strongest argument.

      This reviewer was also concerned about the limits of in vivo experiments. As noted above, we agree that the in vivo situation is less controlled and indirect effects are possible. However, since the direct action on astrocytes in a defined culture system is identical to what we observe in vivo, the most likely explanation is that the GPCR is having the same effect on TNF production, rather than leading to an unknown secondary signaling which then alters TNF production in microglia (or other cell types).

      The remaining concerns about sample size, statistics, etc will be fully addressed in an upcoming revision. All reported n’s are biological replicates. The iPSCs were generated from 3 distinct unrelated individuals.

    1. eLife Assessment

      This important study shows that neural responses to visual input during active vision are more closely aligned with self-generated eye movements than with fixation onset. The evidence for neural responses being more aligned to saccade-related events than fixation-related events is convincing, but the evidence for the more novel finding that neural responses are most closely aligned to peak saccade curvature is currently incomplete. This work will be of interest to visual perception and sensory processing researchers.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining MEG responses to participants free-viewing natural visual images. The vast majority of our knowledge of visual processing in the brain comes from studies where visual input is presented during fixation and the neural response is measured relative to stimulus onset. Even studies that include eye movements tend to either analyze the data relative to the start of each new fixation, or ignore saccades as noise. The current study simultaneously measures MEG and eye-tracking during active vision, and conducts a variety of analyses testing which of the saccade-related events produce the best alignment to the neural data. Five human participants viewed thousands of complex natural scene images while freely moving their eyes. MEG data were then binned as a function of saccade duration and aligned to different fixation and saccade events. M100 responses were better aligned with the preceding saccade onset than the current fixation onset. An additional analysis showed that when MEG signals were decomposed into independent components, the majority of the components showed more alignment and variance explained from saccade-related events (saccade onset, peak velocity, peak visual motion energy, and peak saccade curvature) compared to fixation-onset-defined events; the strongest performing of these factors was the time of peak saccade curvature. A final analysis compared the similarity of MEG topographies measured from stimulus onset (as would be standard in a static design) to those linked to peak saccade curvature and fixation onset, showing that stimulus onset responses were quite dissimilar to the active vision aligned events.

      Strengths:

      Overall, I think this is a fundamentally important research question, taking a novel and interesting approach. I very much like the idea behind this study. My enthusiasm is somewhat tempered by the weaknesses described below. However, at the very least I think this study would be valuable as a key launching point for future explorations, and for pushing the field into a much-needed new direction.

      Weaknesses:

      In its current state, the manuscript seems preliminary/incomplete in terms of both data analysis and engagement with the prior literature.

      (1) In terms of the theoretical contribution, there are several potential contributions, some supported more by the data than others, and some more novel than others. In my rough assessment, from most general to most specific:<br /> a. Static vision is not the same as active vision. Supported somewhat by the analyses. Not novel (there are several studies both recent and older making this point, aside from the vaguely referenced sink-source sentence in the discussion), but this is still an understudied/underappreciated area.<br /> b. Neural responses are better aligned to saccade-related events than fixation-related events. Supported pretty compellingly by the analyses, and pretty novel. An important theoretical contribution.<br /> c. Peak saccade curvature is the saccade-related event explaining most variance. An extremely novel finding, but not well supported by the current data. At best, this seems a preliminary, exploratory hint of something to investigate further. It's intriguing but lacking in both empirical support (e.g. is this even consistent across subjects?) and theoretical discussion (what would it mean / what would be the mechanisms of such a link?).

      (2) There is a small number of subjects, and for several main analyses, the data are pooled across them. Small N's can be reasonable in cases where there is large data for each subject. But it is standard to show the subjects individually to confirm reliability. Figure 1 does this nicely, but then for the main analyses examining the ICs and variance explained by the different saccade-related events (Figures 2C-F), the data were pooled across subjects. Strong conclusions are being drawn from the pooled data (e.g. highest proportion of explained variance from the peak curvature event), but it's unclear if this is consistent across subjects or potentially dominated by 1 or 2 subjects. Indeed, when the "best" score is presented for each participant (Fig 2E), only 2 of the 5 subjects showed peak saccade curvature as the best. And these results look strikingly different across subjects (P5 doesn't even look anything like an M100 response).

      (3) Several parts of the results and methods are hard to follow. I had to read the paper several times to understand it. In many cases, the methods text doesn't even link with the results (e.g. the term "M100" is not anywhere in the methods).

      (4) Several parts of the results felt under-explored:<br /> a) The analysis in Figure 1E is very interesting, but it's not reported in enough detail. There are no quantitative results here, just a visual of a distribution and a description of it being broad. I would be particularly interested in seeing the mean alpha reported for the best sensor for each participant (i.e. linking with the rest of that figure).<br /> b) How consistent is the timepoint of peak saccade curvature? It appears to increase with saccade duration, but is it a fixed / consistent percentage of saccade duration? If not, what factors cause it to vary? How similar is this timepoint to the optimal alpha from the analysis in Figure 1E? Would binning the data based on peak saccade curvature instead of saccade duration produce even better alignments for Figure 1D?<br /> c) For the Figure 3 analysis comparing static scene-onset responses to the saccade- and fixation-related responses: I am wondering how much of the difference is actual saccade-related activity vs a true difference in visual processing. It seems the interpretation is that "visual processing", when measured in static contexts, is very different from when measured in active contexts. But what's being compared is not visual processing specifically, but the entire whole-brain MEG response. I think in order to make this conclusion more compelling, there needs to be some way of filtering out these influences. E.g., a study that presents a simulated saccade condition, where a participant keeps their eyes fixated but views snapshots of the visual scene mimicking the exact saccade sequence of another subject.

      (5) The discussion felt too thin. See some specific points below. In general, combined with the fact that the results were often hard to follow and sparse, I was left with the impression that this report was being forced into a shorter format than necessary.

      (6) How do microsaccades and other types of eye movements fit into this story?

    3. Reviewer #2 (Public review):

      Summary:

      Although our visual system is continuously analyzing the current visual scene, its processing proceeds in discrete episodes separated by brief eye movements (saccades). It has generally been assumed that the analysis of the next visual snapshot begins in earnest when the eyes land on a new fixated location just after a saccade, but there have been various studies indicating that at least some amount of processing occurs earlier, as the system anticipates the impending eye movement. Here, the authors use magentoencephalography (MEG) measurements to record visually-driven responses and determine at what point exactly the processing of a new visual snapshot begins.

      Strengths:

      (1) The work is concise and to the point, and the techniques used are a good way to answer the underlying question about visual processing, since they reflect widespread activity in the brain (rather than activity at a particular location or structure).

      (2) The use of natural images and extensive data collection from 5 participants is a nice feature of the experimental design which permits characterization of the common effects and of variance across individuals.

      (3) The data are analyzed rigorously, but the results are also understood intuitively; for instance, by visual comparison of responses aligned on fixation onset versus saccade onset.

      (4) The results provide a clean characterization of when visual analysis begins relative to saccade onset under natural viewing conditions.

      Weaknesses:

      (1) There were questions about how the scene-onset condition was established, and how data were selected for it.

      (2) The significance of the results is slightly overstated; the text would benefit if some of the claims were phrased with a bit more carefully.

      (3) In particular, the issue of how motor-related processes (versus stimulus-related content) may determine the processing of the next visual snapshot should be discussed with a bit more nuance.

      These are minor weaknesses. Overall, I found the work to be novel and instructive, as it bridges neurophysiological and psychophysical findings in a satisfactory way.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript addresses a fundamental question in cognitive neuroscience: which event should serve as the temporal reference for neural processing during natural vision? While fixation onset has traditionally been treated as the analogue of stimulus onset in free-viewing experiments, the authors convincingly demonstrate that this assumption is incomplete.

      The study utilizes a remarkable natural-viewing dataset consisting of simultaneous MEG and eye-tracking recordings collected during the exploration of thousands of natural scenes. The authors compare several candidate eye-movement events and evaluate which event best explains the timing of the early M100 response. Across several complementary analyses, saccade-related events consistently outperform fixation onset, with peak saccade curvature emerging as the event that best predicts neural response timing.

      Strengths:

      A particular strength of the work is that the conclusions do not rely on a single analytical approach. Instead, multiple independent analyses converge on the same interpretation, increasing confidence that the observed timing relationships are robust rather than analysis-specific. The comparison between natural-viewing responses and classical stimulus-onset responses is especially compelling and highlights qualitative differences in their spatiotemporal organization. Of particular conceptual importance, the findings support the broader perspective that perception is intrinsically linked to action and internally generated sensorimotor processes. This aligns well with growing evidence that oculomotor action and active sampling play central roles in perception. The work contributes to an important ongoing shift in how natural vision should be studied experimentally and interpreted theoretically.

      Weaknesses:

      I identified no major weaknesses in the study. The main limitation is the relatively small number of participants, despite the exceptionally rich dataset. Future work in larger cohorts and across complementary electrophysiological recording modalities will help establish the generalizability of the reported temporal relationships.

    5. Author response:

      We would like to thank the editor and reviewers for their constructive and thoughtful feedback. We appreciate the reviewers' assessment that our work addresses a fundamentally important research question through a novel approach. We are also glad the reviewers found our data to be rigorously analysed, and that they valued our focus on the whole cortex rather than localised regions or electrodes. We are encouraged by the overall assessment of our work and welcome the suggestions for improving the manuscript. Below, we summarise how we plan to address the reviewers' comments in our revision:

      Analyses

      - We will include quantitative results for Fig. 1E that describe the distribution of the optimal alpha values across sensors and participants.

      - For Figures 2 and 3 we will include analyses of individual participants in the Appendix.

      - We will provide more detailed descriptions on how the timing of peak saccade curvature relates to saccade onset and to the identified optimal alpha value.

      Presentation of Methods and Results

      - We will phrase our claims and conclusions more carefully and nuanced throughout, ensuring direct coverage by the data and analyses.

      - We will revise the currently complex sections of the Methods and Results to improve clarity and readability.

      - We will be more explicit about how the data for scene onset were selected.

      Revision of the Discussion

      - We will extend the Discussion section to address possible mechanisms linking the timing of peak saccade curvature and ERF initiation. We will also provide a more thorough discussion of existing and more recent literature on the topic.

      - We will emphasise the main takeaway of the study: the observation that saccade-related processes are more important to the M100 than previously thought, and, reversely, that this component may be less directly related to fixation-locked responses. We will also present our observation of peak saccade curvature as a starting point for future research, as it was not intended as conclusive mechanistic insight into how and why this process relates to early cortical responses.

    1. eLife Assessment

      This is a valuable manuscript that leverages information already being collected in mosquito surveillance, but that is currently discarded, building on ideas developed during the SARS-CoV-2 pandemic. The work is rigorous, involving ample data from real-world surveillance, independent laboratory testing, and simulation modeling to validate the inference method. The evidence is solid for demonstrating feasibility and biological plausibility.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors were responsive to the previous comments and, where needed, edited the manuscript to improve clarity around assumptions and to highlight specific sensitivity analyses.]

      Summary:

      This manuscript seeks to make use of information about Ct values from PCR testing of mosquito pools for West Nile virus infection to make inferences about mosquito prevalence and West Nile risk. It does so through analysis of empirical data and simulated data with a realistic agent-based model.

      Strengths:

      This work is conceptually innovative for mosquito-borne viruses, building on ideas developed primarily during work on SARS-CoV-2. Exploring this topic is worthwhile regardless of the outcome. The use of data, testing in multiple labs, and complementarity of modeling and empirical data analysis are all strengths of the approach.

      Weaknesses:

      Some of the primary weaknesses include a dependence of the results on relatively narrow model assumptions, and lack of compelling improvement over existing methods. None of these are fatal flaws but are instead modest weaknesses that limit the potential of or excitement about the method.

    3. Reviewer #2 (Public review):

      Summary:

      The authors extend their previous population-based Ct-value framework for inferring community epidemic trajectories from human infections to vector infections, using mosquitoes as vectors for West Nile virus. They use agent-based modelling to distinguish virus-positive detections arising from non-active infection states from those reflecting active infections, and then apply this framework to mosquito surveillance data from Colorado and Texas.

      Overall, this is a well-designed and carefully evaluated study. The manuscript proposes a feasible and potentially valuable framework for vector infection surveillance. The findings are supported by both mechanistic agent-based simulations and applications to real-world mosquito surveillance data, which strengthens the biological plausibility and practical relevance of the proposed approach.

      Strengths:

      A major strength of the study is its clear methodological extension from human infection surveillance to vector infection surveillance. The agent-based modelling framework provides a useful basis for distinguishing active infections from virus-positive detections that may reflect non-active infection states. The application to surveillance data from two different geographic settings further supports the feasibility of the framework. Overall, the study is carefully designed, and the model schematic and main analyses are generally clear.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript seeks to make use of information about Ct values from PCR testing of mosquito pools for West Nile virus infection to make inferences about mosquito prevalence and West Nile risk. It does so through analysis of empirical data and simulated data with a realistic agent-based model.

      Strengths:

      This work is conceptually innovative for mosquito-borne viruses, building on ideas developed primarily during work on SARS-CoV-2. Exploring this topic is worthwhile regardless of the outcome. The use of data, testing in multiple labs, and the complementarity of modeling and empirical data analysis are all strengths of the approach.

      Weaknesses:

      Some of the primary weaknesses include a dependence of the results on relatively narrow model assumptions, and a lack of compelling improvement over existing methods. None of these weaknesses are fatal flaws; they are modest weaknesses that limit the potential of or excitement about the method.

      Thank you for the comment

      Reviewer #2 (Public review):

      Summary:

      The authors extend their previous population-based Ct-value framework for inferring community epidemic trajectories from human infections to vector infections, using mosquitoes as vectors for West Nile virus. They use agent-based modelling to distinguish virus-positive detections arising from non-active infection states from those reflecting active infections, and then apply this framework to mosquito surveillance data from Colorado and Texas.

      Overall, this is a well-designed and carefully evaluated study. The manuscript proposes a feasible and potentially valuable framework for vector infection surveillance. The findings are supported by both mechanistic agent-based simulations and applications to real-world mosquito surveillance data, which strengthens the biological plausibility and practical relevance of the proposed approach.

      Strengths:

      A major strength of the study is its clear methodological extension from human infection surveillance to vector infection surveillance. The agent-based modelling framework provides a useful basis for distinguishing active infections from virus-positive detections that may reflect non-active infection states. The application to surveillance data from two different geographic settings further supports the feasibility of the framework. Overall, the study is carefully designed, and the model schematic and main analyses are generally clear.

      Weaknesses:

      (1) It would be helpful if the authors could provide plots showing variation across locations and over time. This would further support the claim made in the paragraph at lines 101-107.

      Thank you for the comment. Our supplementary Material Figures S5 and S6 already included these visualisations. However, we note that these were not referenced in the manuscript. We have now referenced these within the lines:

      “First, the variation we observe is consistent across five trapping seasons and two states (Figures S5 and S6).”

      (2) Figure 2: The model schematic is clear in terms of workflow, but it would benefit from more information on model parameterization. In particular, it would be helpful to clarify which parameters or migration rates were estimated from the data and which were assumed based on prior literature.

      Thank you for the comment. All the parameters are used from the literature and recorded in the Supplementary Material. However, we have now added a note in the caption of Figure 2, referencing the Supplementary Material as below:

      “Overall structure of the agent-based model (parameters were derived from the literature; see Supplementary Material S2, S3 and S4)”

      (3) Figure 4: I wonder whether the authors examined how changes in the proportion of mosquitoes with static viral-kinetics trajectories would affect the observed bimodal distribution. Relatedly, it would be useful to know whether there is a threshold proportion at which the method becomes less able to distinguish active from static viral-kinetics patterns.

      Thank you for the comment. We have conducted this in analysis and have already included the relevant figures in the Supplementary Material, In particular, Figures S14 (in Section S7) and S22. We have referenced Figure S14 where we discuss the proportion of mosquitoes with static viral-kinetics trajectories that would affect the observed bimodal distribution. However, we had not included a reference to Figure S22, where we illustrate the proportions at which the method becomes less able to distinguish active from static viral-kinetics patterns. We have now included this reference in the same line.

      “We found that the simulated pooled Ct values aligned well with the observed data when the percentage viral load inherited from birds was 100% and the probability of a productive or non-productive infection in the mosquitoes was 0.5, capturing the bimodal distribution of low Ct values (from productively infected mosquitoes) and high Ct values (from non-productively infected mosquitoes) (Figure 4 (B), (C) & (D); Supplementary Material S7.1, Figure S14 for Ct distributions of viral inheritance probability vs. model change probability and Figure S22 for the accuracy and confidence-interval coverage across different productive infection proportions).”

      Conclusion:

      Overall, the evidence is reasonably strong for demonstrating the feasibility and biological plausibility of the proposed framework. Some conclusions would be further strengthened by additional sensitivity analyses on key assumptions, especially the proportion of static viral-kinetics trajectories and spatial-temporal heterogeneity across surveillance sites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for th e authors):

      (1) ll 70-73 - There are a lot of ideas in this sentence. It would be useful to support this with a schematic figure or something like that, which illustrates the conceptual predictions made here. Such a step is necessary given the novelty of what is being explored here.

      This portion of the introduction has been simplified to better introduce the core observations that formed our hypothesis:

      “The substantial variation in viral quantities observed during cross-sectional entomological surveillance suggests more complex vector/virus interactions, and the precedence set by population SARS-CoV-2 testing in humans suggests that Ct value data for WNV in mosquitoes could inform metrics of disease risk to humans. However, there are substantial differences in the epidemiology and biology of WNV infection in mosquitoes with SARS-CoV-2 in humans.”

      We have chosen not to include an additional schematic figure, as the core idea (population viral loads reflect the convolution of infection incidence and within-host viral kinetics) is illustrated in the referenced literature, and the main message of the current manuscript is the explore how this phenomenon is observed in arbovirus vector surveillance.

      (2) ll 134-137 - Mosquito species is another factor that could result in wide variation in Ct values due to differences in vector competence and infection kinetics. Given that 2-4 mosquito species are present in these pools with unknown frequencies, this seems like a potentially major source of unexplained variation.

      Importantly, Culex pipiens and Culex tarsalis mosquitoes are separated prior to testing for WNV. We have clarified in the legend for Figure 1 that Ct values are presented from pools of either Culex pipiens/restuans/salinarus or Culex tarsalis.

      Additionally, we have added the following text to the Materials and Methods section:

      “For identification purposes, Cx. Pipiens species mosquitoes are not separated from the Cx. Salinarius or Cx. Restuans, which are nearly identical morphological. However, Cx. Pipiens is far more abundant than either Cx. Salinarius or Cx. Restuans in Nebraska.”

      We also show in Figure S22, S23, and S24 that we observe similar variation in Ct values across mosquito species, location, and epi week, and thus we do not think that differences between species play a major impact in our findings.

      (3) ll 173-175 - I believe that this is a consequence of the trapping method. Could this please be spelt out a bit more?

      Indeed, all of the data in this manuscript were derived from mosquitoes collected in CDC Light Traps that attract host-seeking mosquitoes (i.e mosquitoes looking for a bloodmeal). We are not considering vertical transmission in our model as it has been reported to occur infrequently in laboratory studies. Therefore, WNV-positive mosquitoes collected in CDC Light Traps have been exposed to WNV through a previous blood meal from a bird. We have clarified the text to include this explanation:

      “The mosquito pool Ct value data in this study come from specimens collected using CDC Light Traps that are baited with CO2, specifically targeting host-seeking mosquitoes. Vertical transmission is not factored into our model, thus, for a WNV-positive mosquito to be captured in the pool, it must have already obtained one blood meal from an infected bird and be seeking its next blood meal, which introduces a delay between infection and being captured.”

      (4) ll 177-178 - Doesn't the temporal trend in Ct values primarily reflect temporal changes in mosquito infection prevalence?

      Thank you for the comment. We agree that the temporal changes in mosquito infection prevalence is the main factor influencing the distribution of the Ct values in pools, as the time-since-infection distribution of trapped mosquitoes does not vary sufficiently to lead to trapping mosquitoes at very different points in their viral kinetics trajectory. Our intended point from this sentence was, given the prevalence and pool size, the variation in the viral load of the infected mosquitoes does not vary in time as all infected mosquito are trapped after they have reached a constant high-viral load level. We have now revised this sentence to reflect this.

      “As the infected mosquitoes progress from increasing viral load to a high set-point viral load, temporal trends in pooled Ct values primarily reflect time-varying infection prevalence and the number of infected mosquitoes in each pool. The remaining non-temporal variation in pooled Ct values reflect individual-level variation in mosquito set-point viral loads.”

      (5) Section 2.2 - It would seem that the assumed viral kinetics in birds would be important to this line of reasoning, given that that determines initial viral load ingested by mosquitoes. I am unclear on what was assumed in the model regarding viral kinetics in birds.

      Thank you for the comment. We have discussed the viral kinetics of the birds in detail in Section 5.3 and Supplementary Material S3. However, we agree that we have not explicitly mentioned this in Section 2.2. Therefore, we have added a reference to these sections in the following paragraph:

      “The model assumes that the mosquito's initial viral load is proportional to the infector bird's viral load (see Section 5.3 and Supplementary Material S3 for further details on the bird viral kinetics model).”

      (6) ll 194-202 - Whilst you have shown that this hypothesis leads to predictions that are consistent with the data, this is a relatively narrow hypothesis, and others are neither discussed nor refuted.

      There are two features of the data which we discuss. First, the substantial variation in Ct values across pools. This is described in detail in Section 2.1. The second observation is the bimodal pattern, which L192-202 refers to. While we agree that we have not modelled alternative hypotheses, our point is that the distribution of pooled Ct values is bimodal, and capturing some mosquitoes with very low viral loads is the most plausible explanation for the mode at high Ct values. However, we contend that this is actually a fairly broad hypothesis, as there are many plausible mechanisms generating mosquito infections with low viral loads, which we already discuss (discussion section beginning “This could be explained by a variety of factors…”). No changes have been made to the manuscript.

      (7) Section 2.3, first paragraph - The problem with this approach is that these simulations depend on a number of assumptions and parameter settings that are not estimated as part of the model fitting process. Thus, the model is very narrow and contingent on these narrow and not compellingly justified assumptions.

      Thank you for the comment. While we agree with the reviewer that this is a potential limitation of our study, we have discussed this in detail in the discussion. As mentioned in the manuscript “the main objective of this study was not to formally fit the multi-scale agent-based model to the data, but rather to understand how individual-level viral kinetics in mosquitoes are reflected in pooled surveillance data, and to demonstrate the use of pooled Ct values in estimating WNV infection prevalence”, we believe the assumptions and model are sufficient to address the research objectives. Furthermore, the fact that simulated Ct value distributions from the ABM can be used directly with the prevalence estimation method to give similar estimates to the existing PooledInfRate package supports the validity of our assumptions, though we agree that this does not necessarily mean all of our assumptions are correct, nor that our model is generalisable to other settings. No changes have been made to the manuscript.

      (8) Section 2.3, second paragraph - So the newly proposed method using Ct values does no better than the existing method using binary data?

      Thank you for the comment. We agree with the reviewer that our method and the existing PooledInfRate package perform similarly at the estimated prevalence levels of WNV. However, the Ct-based method, as we have discussed and shown, is robust at all prevalence levels where the binary-only method fails, and our method can distinguish the productive and non-productive prevalence.. Thus, while the prevalence estimates are similar under both methods for the current dataset, the novelty lies in the ability to reconstruct prevalence using the data in an entirely different way, and the proof-of-concept for how Ct values may harbour more biological information than treating pools as positive/negative. We believe that these points are sufficiently discussed throughout the manuscript. We have not made any changes to the manuscript.

      (9) ll 252-254 - This may only be true because the simulation model and the inference model are identical. If the inference model were misspecified (due, for example, to incorrect assumptions about kinetics, etc), this result would likely weaken.

      Thank you for the comment. The difference in robustness between the binary-only method and Ct-based method is not a feature of the method, but rather of how the data is used. At higher prevalence, all pools are likely to have at least one positive mosquito in them, and thus all pools will be positive, removing all information to discriminate between different prevalence levels. In contrast, the Ct-based method is able to still discriminate between prevalence levels even when all of the pools are positive, as there is still information based on whether the positive pools have low or high Ct values. No changes have been made to the manuscript.

      (10) ll 272-273 - Again, this is highly dependent on built-in model assumptions.

      Thank you for the comment. We agree that the performance of a model can depend on the underlying assumptions and the structure of the model. This is true in general for any model-based inference technique (see White, 1982, for example). Therefore, our simulations, results and interpretations are intended to be evaluated under the model structures and underlying assumptions we have used throughout the manuscript. However, to be explicit, we have now added this line at the end of the paragraph that included the sentence.

      “These results are based on the model structure and the underlying assumptions we used and they may be affected by model misspecification, including incorrect assumptions (see White, 1982, for example).”

      (11) ll 273-281 - Can this be done with pooled data only, or does it require individual mosquito Ct values? The latter would seem to be less practical to obtain in real-world applications.

      Thank you for the comment. As we have cited the related work for SARS-CoV-2, in theory, these methods are applicable when individual Ct values are present. Both pooled data and individual-level data will work, but using pooled data requires the pooling and dilution process to be modelled explicitly. However, as the reviewer mentions, for mosquito surveillance, this is not a practical approach as mosquitoes are always pooled prior to testing to reduce effort and costs. No changes have been made to the manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) The supplementary figures do not appear to be ordered according to their first mention in the manuscript, which makes them somewhat harder to follow.

      Thank you for the helpful comment. We have now made sufficient changes to the Supplementary Material and updated the references in the manuscript. Where possible, the supplementary figures and sections are now numbered and presented in the order of their first mention in the manuscript.

      (2) Lines 71-73: This sentence is somewhat vague, and I was not fully clear on the intended message. The authors may wish to revise it for clarity.

      Thank you for the comment. Similar comments have been made by Reviewer #1. We have revised this sentence for clarity.

    1. eLife Assessment

      This work makes an important contribution to understanding the role of calreticulin and the Del52 variant in calcium homeostasis in cells based on convincing evidence. There was some question as to whether the work in vitro fully recapitulated the physiological setting in which the homozygous Del52 variant has been reported to have defects in calcium homeostasis. Nonetheless, the studies are well performed using appropriate methods to address these questions and make an important contribution to the understanding of calreticulin, and the Del52 variant, function in health and disease.

    2. Reviewer #1 (Public review):

      The authors attempted to compare calcium binding properties of wildtype calreticulin with calreticulin deletion mutant (CRTDel52) associated with myeloproliferative neoplasms.

      The researchers conducted their study using advanced techniques They found almost no difference in calcium binding between the two proteins and observed no impact on calcium signaling, specifically store-operated calcium entry (SOCE). The study also noted an increase in ER luminal calcium-binding chaperone proteins. Surprisingly, the authors selected flow cytometry as a technique for measurements of ER luminal calcium. Considering limitations of this approach it would be better to use alternative approaches. This is particularly important as previous reports, using cells from MPN patients, indicate reduced ER luminal calcium and effects on SOCE (Blood, 2020). This issue matters because earlier research with MPN patient cells reported reduced ER luminal calcium levels and altered SOCE (Blood, 2020). How do the authors explain the difference between their results and previous findings about lower ER luminal calcium and changed SOCE in MPN patient cells expressing CRTDel52? Other studies have found that unfolded protein responses are activated in MPN cells with CRTDel52 calreticulin (see Blood, 2021), and increased UPR could account for higher levels of some ER resident calcium-binding proteins observed here. Overall, it remains unclear how this work improves our understanding of MPN or clarifies calreticulin's role in MPN pathophysiology.

      Comments on revised version.

      The authors have addressed the points raised in the original review. However, given the absence of significant differences between the wild-type and mutant proteins, the relevance of this work to MPN pathology remains unclear. The novelty of the study is limited, as calcium has generally not been considered a significant factor in MPN pathology associated with mutant calreticulin.

    3. Reviewer #2 (Public review):

      Summary:

      Tagoe and colleagues present a thorough analysis of the calcium (Ca2+) binding capacity of calreticulin (CRT), an endoplasmic reticulum (ER) Ca2+-buffer protein, using a mutant version (CRT del52) found in myeloproliferative neoplasms (MPNs). The authors use purified human CRT protein variants, CRT-KO cell lines, and an MPN cell line to elucidate the differing Ca2+ dynamics, both on the level of the protein and on cell-wide Ca2+-governed processes. In sum, the authors provide new insights into CRT that can be applied to both normal and malignant cell biology.

      First the authors purify CRT protein and perform isothermal titration calorimetry to quantify the Ca2+ binding capacity of CRT. They use full-length human CRT, CRT del52, and two truncations of CRT (1-339 and 1-351, the former of which should lead to the entire loss of low affinity Ca2+ binding). While CRT del52 has previously been shown to lead to a decrease in Ca2+ binding affinity in other models, the ITC data shows that this is retained in CRT del52.

      Next, the authors utilize a CRT-KO cell line with subsequent addition of CRT protein variants to validate these findings with flow cytometric analysis. Cells were transfected with a ratiometric ER Ca2+ probe, and fluorescence indicates that CRT del52 is unable to restore basal ER Ca2+ levels to the same extent as CRT wild-type. To translate these findings to MPNs, the authors perform CRT-KO in a megakaryocytic cell line, where reconstitution with either CRT variant did not cause a difference in cytosolic calcium levels. The authors further test store-operated calcium entry (SOCE), an important process to maintaining ER Ca2+ levels, in these cells, and find that CRT-KO cells have lower SOCE activity, and that this can be slightly recovered with CRT addition.

      Finally, the authors ask whether other effects of CRT-KO/reconstitution can affect cellular Ca2+ signaling pathway and levels. RNASeq analysis revealed showed that CRT-KO lead to an increase in various chaperone protein expressions, and that reconstitution with CRT del52 is unable to reduce expression to the same extent as reconstitution with CRT wildtype.

      Comments on revised version.

      The authors have sufficiently addressed my concerns from the first review.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study investigates low-affinity Ca2+ binding by WT calreticulin and mutant calreticulin associated with type I myeloproliferative neoplasms, as well as the impact on Ca2+ fluxes in suspension cultures of megakaryocyte-like cells in vitro in response to ER Ca2+ ATPase inhibitors that deplete endoplasmic reticulum (ER) Ca2+ store and open plasma membrane Ca2+ channels through STIM1-Orai interactions. The results are important in that they show that Ca2+ binding by calreticulin and store-operated Ca2+ entry are not fundamentally impacted by the type I deletion mutation in calreticulin, which rules out a direct effect of the calreticulin mutation on its own low-affinity Ca2+ binding and any broad impact on ER Ca2+ regulation. The strength of the data and methods used ranges from solid to convincing, although the use of suspension-based flow cytometric assays to investigate ER Ca2+ levels and Ca2+ entry can be challenged. High-affinity Ca2+ binding sites could be further considered, and possible confounding effects of Abl kinase activity in the megakaryocyte-like cell lines could be offset.

      The authors thank the editors and the reviewers for the summary, comments and many helpful suggestions. In the revised manuscript, we have used fluorimetry for more precise ER calcium measurements (new Figure 5), clarified the high-affinity calcium binding site concern based on our previous work, and addressed possible effects of BCR-ABL translocation kinase activity upon cellular calcium signaling using the drug Imatinib (new supplements to Figure 6 and 7).

      Public Reviews:

      Reviewer #1 (Public review):

      The researchers conducted their study using advanced techniques. They found almost no difference in calcium binding between the two proteins and observed no impact on calcium signaling, specifically store-operated calcium entry (SOCE). The study also noted an increase in ER luminal calcium-binding chaperone proteins. Surprisingly, the authors selected flow cytometry as a technique for measurements of ER luminal calcium. Considering the limitations of this approach, it would be better to use alternative approaches.

      Thank you for this suggestion. We have undertaken fluorimetry-based ER calcium measurements, which are shown in a new Figure 5. These also indicate similar ER calcium levels in CRT-KO HEK293T cells, compared to those reconstituted with wild-type CRT and CRT<sub>Del52</sub>.

      This is particularly important as previous reports, using cells from MPN patients, indicate reduced ER luminal calcium and effects on SOCE (Blood, 2020). This issue matters because earlier research with MPN patient cells reported reduced ER luminal calcium levels and altered SOCE (Blood, 2020). How do the authors explain the difference between their results and previous findings about lower ER luminal calcium and changed SOCE in MPN patient cells expressing CRTDel52?

      We thank the reviewer for asking for these clarifications. We have revised the discussion to address some of these points and also clarify the findings of the referenced study (Di Buduo et al., 2020) which did not directly measure ER calcium levels. We also discuss findings from a related study with cultured megakaryocytes from patients that indicated different effects of type I vs type II mutations (Pietra et al., 2016). In the absence of engineered controls, sample-to-sample heterogeneities in primary cells make it difficult to attribute any measured differences as direct effects of CRT mutations. Different from these experiments, by using purified proteins and ITC, our studies show that the Del52 mutant has calcium-binding characteristics resembling that of the wild-type protein. Additionally, through genetic manipulations in cell lines, our studies directly address the effects of calreticulin KO and its Del52 mutation upon ER luminal and cytosolic calcium levels, and cellular SOCE signals. We did not measure significant differences in any of these parameters between the KO cells and those reconstituted with wild-type calreticulin or the Del52 mutant. As noted by the editors, these results show that Ca2+ binding by calreticulin and SOCE in a cell are not fundamentally impacted by the type I deletion mutation.

      Other studies have found that unfolded protein responses are activated in MPN cells with CRTDel52 calreticulin (see Blood, 2021), and increased UPR could account for higher levels of some ER-resident calcium-binding proteins observed here.

      These points are addressed in the discussion. Either protein misfolding in cells with wild-type calreticulin deficiency or the sensing of cellular calcium perturbations could induce the expression of ER calcium-binding proteins in calreticulin-deficient cells, although we favor the latter model for the reason specified in the discussion. Regardless of the precise mechanisms underlying the expression changes in calcium-binding proteins, the upregulated factors are predicted to compensate for calreticulin deficiency and contribute to the maintenance of the overall cellular calcium homeostasis.

      Overall, it remains unclear how this work improves our understanding of MPN or clarifies calreticulin's role in MPN pathophysiology.

      Multiple studies referenced in the manuscript have suggested links between altered calcium signaling/binding by CRT mutants and MPN pathogenesis. Our studies indicate that ER and cytosolic calcium levels and SOCE are not directly impacted by the MPN type I CALR mutation, points noted in the abstract and discussion. Thus, calcium signaling may not play a specific role in MPN CALR mutant pathology via suggested mechanisms. We are confident that readers will find these results important for better understanding the role of calreticulin type I mutations in MPN.

      Reviewer #1 (Recommendations for the authors):

      This study aimed to express, purify, and evaluate low-affinity calcium binding by a calreticulin deletion mutant (CRTdel52) that is linked to myeloproliferative neoplasms (MPN). The researchers performed cell imaging, flow cytometry, and isothermal titration calorimetry to compare calcium binding between wild-type calreticulin and CRTDel52. They assessed cytosolic calcium levels and store-operated calcium entry (SOCE) in HET293T cells (CRT knocked-out background) and in megakaryoblastic MEG-1 cells. Additionally, they examined changes in the abundance of endoplasmic reticulum (ER) resident calcium-binding proteins in cells expressing either wild-type or mutant CRT.

      This study is well executed but lacks clear relevance to MPN, and it is not clear how this work advances our knowledge of calreticulin biology. No differences were found in calcium binding between wild-type calreticulin and CRTDel52, nor was SOCE impacted. They noticed, however, a compensatory increase in the abundance of some ER resident calcium-binding proteins. The lack of any significant changes in calcium behavior between wild type and CRTDel52 is not surprising based on the known amino acid sequence of calreticulin and calreticulin mutant and based on our knowledge about CRT calcium binding in general. Consequently, it is not clear how this work advances our understanding of the pathophysiology of MPN. Calcium may not play a critical role in the MPN pathology; instead, CRTDel52 secretion and receptor signaling appear more central. Further research should address how these findings relate specifically to MPN and cell biology, in general.

      Our current studies demonstrate increased expression of other calcium-binding proteins in the context of heterozygous MPN type I CALR mutations (Figure 8C) or conditions resembling homozygous MPN type I CALR mutations (Figures 8D-8G). These results, together with findings of maintained ER and cytosolic calcium levels and SOCE signals (Figures 4-7 and Figure 6, supplemental Figure 2 and Figure 7, supplemental Figure 1), indicate that altered calcium binding/signaling by Del52 does not directly contribute to MPN pathology.

      What is the biological or pathophysiological relevance of the CRTDel52-KDEL construct?

      The KDEL sequence is important for the ER retention of CRT (Sonnichsen et al., 1994), and its addition was expected to at least partially remedy the ER retention defect of CRT<sub>Del52</sub>. This point is clarified in the revised results section.

      The rise in ER calcium-binding proteins is noteworthy but anticipated, given likely genetic changes from UPR pathway activation in these cells. Is this relevant to MPN?

      We suggest that increased expression of other calcium-binding proteins in the context of heterozygous MPN type I CALR mutations (Figure 8C) or conditions resembling homozygous MPN type I CALR mutations (Figures 8D-8G) would contribute to the maintenance of the cell’s calcium signaling capacity.

      How do the authors explain the difference between their results and previous findings about lower ER luminal calcium and changed SOCE in MPN patient cells expressing CRTDel52?

      The findings related to SOCE are addressed in the points discussed above and in the revised discussion. Related to ER luminal calcium, the study by Ibarra et al. (Ibarra et al., 2022) reported that CRT<sub>Del52</sub> overexpressed in U2OS cells (expressing endogenous CRT) had reduced ER calcium levels compared to the same cells expressing WT CRT or CRT<sub>Ins5</sub>. Those measurements did not use a ratiometric ER calcium probe, and additionally it is possible that the expression of compensatory calcium-binding proteins is more muted in cells expressing endogenous wild-type CRT.

      Other studies have found that unfolded protein responses are activated in MPN cells with CRTDel52 calreticulin (see Blood, 2021), and increased UPR could account for higher levels of some ER-resident calcium-binding proteins observed here.

      We agree that increased UPR could account for higher levels of some ER-resident calcium-binding proteins. As noted in the revised discussion, regardless of the precise mechanisms underlying the expression changes in calcium-binding proteins, the upregulated factors are predicted to compensate for calreticulin deficiency and contribute to the maintenance of the overall cellular calcium homeostasis.

      It is not clear why flow cytometry was a choice of technique for measurements of ER calcium. Pacific Blue was detected at 405 nm excitation and 452-455 nm emission, while unbound probe signals appeared in the AmCyan channel (405 nm excitation/498 nm emission). It is unnecessary to mention fluorochrome labels (like Pacific Blue or AmCyan) for channels that are not in use. Only the channels actually utilized need to be specified: The calcium-bound GEM-CEPIA1er probe's signal was collected using the Pacific Blue channel (excitation at 405 nm, emission at 452-455 nm), while the unbound probe's signal was detected using the AmCyan channel (excitation at 405 nm, emission at 498 nm). Since it's not possible to monitor emission only at a specific wavelength with a Fortessa, could this be a different channel that is being recorded?

      Additionally, due to the similar spectra and potential for bleed-through between Pacific blue and Amcyan, compensation is likely necessary. Therefore, you should include single-stained control cells containing only one probe for proper reporting. Additionally, only the "GEM-CEPIA1er probe" is displayed, while the second probe is referred to solely as "unbound".

      A single genetically encoded GEM-CEPIA1er probe (Suzuki et al., 2014) was used for measuring both the bound and unbound signals. The GEM-CEPIA1er probe was excited with the 405 nm violet laser. In the methods section of the revised manuscript, the GEM-CEPIA1er probe wording is included for describing both the bound and unbound signal collections. Additionally, we have undertaken new spectrofluorimetric experiments (new Figure 5), which allow for the distinct emission peaks to be recorded corresponding to the Ca<sup>2+</sup>-bound and Ca<sup>2+</sup>-unbound signals. Similar results were obtained as reported for the flow cytometry-based experiments.

      Increased expression of wild-type calreticulin compared to parental cells should impact on ER calcium content and dynamics in back-transfected HEK293T-KO or MEG-1 cells. Direct ER calcium measurements in HEK293 cells with various calreticulin constructs would significantly strengthen this presentation.

      Our experiments were structured to compare calcium signaling in cells expressing only wild-type CRT or CRT<sub>Del52</sub> (resembling homozygous type I MPN CALR mutations) compared to CRT-KO cells. The over-expression of CRT in the reconstituted cells compared to endogenous expression level is a limitation of our study which we have acknowledged in the revised manuscript discussion. Understanding the effects of over-expression of wild-type CRT vs the CRT<sub>Del52</sub> mutant upon ER and cytosolic calcium signals and SOCE is interesting, but beyond the scope of the present study.

      The authors should examine the immunolocalization of CRTDel52 and wild-type protein in HEK293 cells.

      Previous published studies from another lab showed that CRT<sub>Del52</sub> is secreted from HEK cells and that the addition of a KDEL sequence to CRT<sub>Del52</sub> reduces secretion and induces its increased intracellular accumulation (Arshad and Cresswell, 2018). This point is noted in the revised results section and the reference is cited. This appears to be the general theme in primary cells and cell lines. Previous studies and our own prior published studies have shown that CRT<sub>Del52</sub> (but not wild-type CRT) is detectable in the media of cell lines and patient serum as well on the cell surface of primary cells and cell lines (Kaur et al., 2024, Venkatesan et al., 2021, Pecquet et al., 2023).

      SDS-PAGE of purified proteins is overloaded, and chromatograms show extra peaks or shoulders, making protein quality assessment uncertain.

      Representative chromatograms, peaks corresponding to protein monomers used for ITC analyses and the relevant gels are clarified in the revised manuscript. In new analyses since the original submission, intact protein mass spectrometry was undertaken for CRT<sub>Del52</sub>. The results indicate a 35-42 amino acid truncation in different preparations. The truncated proteins would still include acidic residues (between 340–351) previously implicated in low-affinity calcium binding by murine CRT that are shared between wild-type and CRT<sub>Del52</sub>. This new information is now included in the revised results section.

      Analysis of SOCE in calreticulin-deficient cells and cells reconstituted with calreticulin or overexpressing the protein has already been reported (PMID12324449).

      The indicated reference and additional related papers examining effects of CRT deficiency and overexpression on cellular calcium signaling (Arnaudeau et al., 2002, Bastianutto et al., 1995, Mery et al., 1996, Nakamura et al., 2001) are cited in the revised manuscript.

      Reviewer #2 (Public review):

      Tagoe and colleagues present a thorough analysis of the calcium (Ca2+) binding capacity of calreticulin (CRT), an endoplasmic reticulum (ER) Ca2+-buffer protein, using a mutant version (CRT del52) found in myeloproliferative neoplasms (MPNs). The authors use purified human CRT protein variants, CRT-KO cell lines, and an MPN cell line to elucidate the differing Ca2+ dynamics, both on the level of the protein and on cell-wide Ca2+-governed processes. In sum, the authors provide new insights into CRT that can be applied to both normal and malignant cell biology.

      First, the authors purify CRT protein and perform isothermal titration calorimetry to quantify the Ca2+ binding capacity of CRT. They use full-length human CRT, CRT del52, and two truncations of CRT (1-339 and 1-351, the former of which should lead to the entire loss of low-affinity Ca2+ binding). While CRT del52 has previously been shown to lead to a decrease in Ca2+ binding affinity in other models, the ITC data show that this is retained in CRT del52.

      Next, the authors utilize a CRT-KO cell line with subsequent addition of CRT protein variants to validate these findings with flow cytometric analysis. Cells were transfected with a ratiometric ER Ca2+ probe, and fluorescence indicates that CRT del52 is unable to restore basal ER Ca2+ levels to the same extent as CRT wild-type. To translate these findings to MPNs, the authors perform CRT-KO in a megakaryocytic cell line, where reconstitution with either CRT variant did not cause a difference in cytosolic calcium levels. The authors further test store-operated calcium entry (SOCE), an important process for maintaining ER Ca2+ levels, in these cells, and find that CRT-KO cells have lower SOCE activity, and that this can be slightly recovered with CRT addition.

      Finally, the authors ask whether other effects of CRT-KO/reconstitution can affect the cellular Ca2+ signaling pathway and levels. RNASeq analysis revealed that CRT-KO leads to an increase in various chaperone protein expressions, and that reconstitution with CRT del52 is unable to reduce expression to the same extent as reconstitution with CRT wildtype.

      Strengths:

      The authors provide new insights into CRT that can be applied to both normal and malignant cell biology.

      We thank the reviewer for the recognition that this study is important for our understanding of both normal and malignant cell biology.

      Weaknesses:

      (1) The authors should consider discussing the high-affinity Ca2+ binding site more in the introduction. Can they show a proof-of-concept experiment that validates that incubation of recombinant CRT reduces the function of that high-affinity Ca2+ binding site?

      In a previous study (Wijeyesakere et al., 2011), we showed that at a starting calcium concentration of 0 mM and with CaCl<sub>2</sub> injections to a final concentration of 70-80 mM the measured K<sub>D</sub> value was 16.6 mM for calcium binding to wild type murine calreticulin, (which has ~95% sequence identity with human calreticulin), corresponding to the high-affinity site. On the other hand, at a starting calcium concentration of 50-100 mM and CaCl<sub>2</sub> injections to a final concentration of 700-850 mM, the measured K<sub>D</sub> value for calcium binding to wild-type murine calreticulin was 590 mM (corresponding to the low-affinity sites). We did not observe the high-affinity sites when the starting calcium concentration was 50 mM and calcium injections were at 33 mM each; similar conditions are used in the present study. These points are clarified in the revised manuscript in the results section.

      (2) For Figure 2B, do you have an explanation for why the purified proteins run higher than predicted (48-52kDa) - are these proteins still tagged with pGB1?

      Yes, the purified proteins shown in Figure 2B retained a GB1 tag. This point is clarified in the revised methods.

      (3) The MEG-01 cell line has the BCR:ABL1 translocation, while CRT mutations are strictly found in BCR:ABL1 negative MPNs. Could these experiments be repeated in these cells treated with imatinib to decrease these effects, or see if basal MEG-01 Ca2+ levels/activity are changed with or without imatinib?

      Thank you for this important point. We have assessed cytosolic calcium levels in MEG-01 cells that were treated or not treated with imatinib in new Figure 6, supplemental Figures 1 and 2 and Figure 7, supplemental Figure 1) and show that the prior results hold in imatinib-treated cells.

      References

      ARNAUDEAU, S., FRIEDEN, M., NAKAMURA, K., CASTELBOU, C., MICHALAK, M. & DEMAUREX, N. 2002. Calreticulin differentially modulates calcium uptake and release in the endoplasmic reticulum and mitochondria. J Biol Chem, 277, 46696-705.

      ARSHAD, N. & CRESSWELL, P. 2018. Tumor-associated calreticulin variants functionally compromise the peptide loading complex and impair its recruitment of MHC-I. J Biol Chem, 293, 9555-9569.

      BASTIANUTTO, C., CLEMENTI, E., CODAZZI, F., PODINI, P., DE GIORGI, F., RIZZUTO, R., MELDOLESI, J. & POZZAN, T. 1995. Overexpression of calreticulin increases the Ca2+ capacity of rapidly exchanging Ca2+ stores and reveals aspects of their lumenal microenvironment and function. J Cell Biol, 130, 847-55.

      DI BUDUO, C. A., ABBONANTE, V., MARTY, C., MOCCIA, F., RUMI, E., PIETRA, D., SOPRANO, P. M., LIM, D., CATTANEO, D., IURLO, A., GIANELLI, U., BAROSI, G., ROSTI, V., PLO, I., CAZZOLA, M. & BALDUINI, A. 2020. Defective interaction of mutant calreticulin and SOCE in megakaryocytes from patients with myeloproliferative neoplasms. Blood, 135, 133-144.

      IBARRA, J., ELBANNA, Y. A., KURYLOWICZ, K., CIBODDO, M., GREENBAUM, H. S., ARELLANO, N. S., RODRIGUEZ, D., EVERS, M., BOCK-HUGHES, A., LIU, C., SMITH, Q., LUTZE, J., BAUMEISTER, J., KALMER, M., OLSCHOK, K., NICHOLSON, B., SILVA, D., MAXWELL, L., DOWGIELEWICZ, J., RUMI, E., PIETRA, D., CASETTI, I. C., CATRICALA, S., KOSCHMIEDER, S., GURBUXANI, S., SCHNEIDER, R. K., OAKES, S. A. & ELF, S. E. 2022. Type I but Not Type II Calreticulin Mutations Activate the IRE1alpha/XBP1 Pathway of the Unfolded Protein Response to Drive Myeloproliferative Neoplasms. Blood Cancer Discov, 3, 298-315.

      KAUR, A., VENKATESAN, A., KANDARPA, M., TALPAZ, M. & RAGHAVAN, M. 2024. Lysosomal degradation targets mutant calreticulin and the thrombopoietin receptor in myeloproliferative neoplasms. Blood Adv, 8, 3372-3387.

      MERY, L., MESAELI, N., MICHALAK, M., OPAS, M., LEW, D. P. & KRAUSE, K. H. 1996. Overexpression of calreticulin increases intracellular Ca2+ storage and decreases store-operated Ca2+ influx. J Biol Chem, 271, 9332-9.

      NAKAMURA, K., ZUPPINI, A., ARNAUDEAU, S., LYNCH, J., AHSAN, I., KRAUSE, R., PAPP, S., DE SMEDT, H., PARYS, J. B., MULLER-ESTERL, W., LEW, D. P., KRAUSE, K. H., DEMAUREX, N., OPAS, M. & MICHALAK, M. 2001. Functional specialization of calreticulin domains. J Cell Biol, 154, 961-72.

      PECQUET, C., PAPADOPOULOS, N., BALLIGAND, T., CHACHOUA, I., TISSERAND, A., VERTENOEIL, G., NEDELEC, A., VERTOMMEN, D., ROY, A., MARTY, C., NIVARTHI, H., DEFOUR, J. P., EL-KHOURY, M., HUG, E., MAJOROS, A., XU, E., ZAGRIJTSCHUK, O., FERTIG, T. E., MARTA, D. S., GISSLINGER, H., GISSLINGER, B., SCHALLING, M., CASETTI, I., RUMI, E., PIETRA, D., CAVALLONI, C., ARCAINI, L., CAZZOLA, M., KOMATSU, N., KIHARA, Y., SUNAMI, Y., EDAHIRO, Y., ARAKI, M., LESYK, R., BUXHOFER-AUSCH, V., HEIBL, S., PASQUIER, F., HAVELANGE, V., PLO, I., VAINCHENKER, W., KRALOVICS, R. & CONSTANTINESCU, S. N. 2023. Secreted mutant calreticulins as rogue cytokines in myeloproliferative neoplasms. Blood, 141, 917-929.

      PIETRA, D., RUMI, E., FERRETTI, V. V., DI BUDUO, C. A., MILANESI, C., CAVALLONI, C., SANT'ANTONIO, E., ABBONANTE, V., MOCCIA, F., CASETTI, I. C., BELLINI, M., RENNA, M. C., RONCORONI, E., FUGAZZA, E., ASTORI, C., BOVERI, E., ROSTI, V., BAROSI, G., BALDUINI, A. & CAZZOLA, M. 2016. Differential clinical effects of different mutation subtypes in CALR-mutant myeloproliferative neoplasms. Leukemia, 30, 431-8.

      SONNICHSEN, B., FULLEKRUG, J., NGUYEN VAN, P., DIEKMANN, W., ROBINSON, D. G. & MIESKES, G. 1994. Retention and retrieval: both mechanisms cooperate to maintain calreticulin in the endoplasmic reticulum. J Cell Sci, 107 (Pt 10), 2705-17.

      SUZUKI, J., KANEMARU, K., ISHII, K., OHKURA, M., OKUBO, Y. & IINO, M. 2014. Imaging intraorganellar Ca2+ at subcellular resolution using CEPIA. Nat Commun, 5, 4153.

      VENKATESAN, A., GENG, J., KANDARPA, M., WIJEYESAKERE, S. J., BHIDE, A., TALPAZ, M., POGOZHEVA, I. D. & RAGHAVAN, M. 2021. Mechanism of mutant calreticulin-mediated activation of the thrombopoietin receptor in cancers. J Cell Biol, 220, e202009179.

      WIJEYESAKERE, S. J., GAFNI, A. A. & RAGHAVAN, M. 2011. Calreticulin is a thermostable protein with distinct structural responses to different divalent cation environments. J Biol Chem, 286, 8771-85.

    1. eLife Assessment

      The work presented represents important new evidence linking the cascade of neural processes triggered by memory-based prediction errors. The study uses an impressive collection of approaches and methods to characterize and measure cognitive control, arousal, and memory changes as a function of memory-based violations. The analyses are technically sophisticated and rigorous and, taken together, provide solid evidence that there are multiple processes accompanying prediction errors, and that they differentially relate to successful encoding.

    2. Reviewer #1 (Public review):

      This manuscript describes a multi-modal study of associative learning and memory in humans that combines scalp EEG, pupillometry and behavioral analysis to explore the construct of mnemonic prediction errors (MPEs), in terms of their relationship to attention and cognitive control. Across two pooled studies, participants performed associative memory tasks in which they learned the relationship between a cue word (action verb) and subsequent picture (animate or inanimate) with a strong vs. weak (4 or 1 repetitions) encoding manipulation. At test, participants were encouraged to generate a prediction following the cue word to determine whether the subsequently presented picture was a match or mismatch. The timecourse of pupillary responses during match decisions were decomposed using temporal principal components analysis, which identified 6 distinct and overlapping processes. Some of the components (PC3/PC4) exhibited sensitivity to both the strength and mismatch conditions, as well as behavior (both RT and accuracy) and retrieval success on the subsequent trial. Furthermore, relationships were also observed between pupillary responses (specifically for PC4) and both frontal theta and posterior alpha power measures obtained from scalp EEG in Experiment 2, as well as for frontal theta and subsequent learning from mismatch stimuli (assessed using subsequent memory findings from a surprise recognition test). The authors suggest the findings indicate that MPEs elicit changes in attention, arousal and cognitive control which impact subsequent learning.

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPEs.\

    3. Reviewer #2 (Public review):

      Summary:

      The authors studied cognitive control and attention in response to mnemonic prediction errors (MPEs): situations in which the external reality violates internal memory-based predictions. The behavioral task first established strong versus weak predictions, and then either confirmed or violated these predictions. The authors examined markers of cognitive control (frontal theta) and attention (posterior alpha suppression, pupil response) while strong and weak predictions were confirmed or violated. They found increased cognitive control (frontal theta) for strong MPEs, which correlated with subsequent memory. Markers of attention (alpha suppression, pupil response) also accompanied strong MPEs but did not correlate with subsequent memory. Pupil response was investigated using an interesting approach that decomposes the response into different components, finding that different components respond earlier or later and show different correlations with MPEs and their strength. The authors also investigated how EEG, reaction time, and pupil responses correlated with one another, providing further insight into the mechanism underlying the response to MPEs. Together, the study points toward multiple control and attention mechanisms involved in MPE response and memory.

      Strengths:

      The study has a clear behavioral paradigm with multiple measures - behavioral, EEG, and pupillometry that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a multi-modal study of associative learning and memory in humans, that combines scalp EEG, pupillometry and behavioral analysis to explore the construct of mnemonic prediction errors (MPEs), in terms of their relationship to attention and cognitive control. Across two pooled studies, participants performed associative memory tasks in which they learned the relationship between a cue word (action verb) and subsequent picture (animate or inanimate) with a strong vs. weak (4 or 1 repetitions) encoding manipulation. At test, participants were encouraged to generate a prediction following the cue word to determine whether the subsequently presented picture was a match or mismatch.

      The timecourse of pupillary responses during match decisions were decomposed using temporal principal components analysis, which identified 6 distinct and overlapping processes. Some of the components (PC3/PC4) exhibited sensitivity to both the strength and mismatch conditions, as well as behavior (both RT and accuracy) and retrieval success on the subsequent trial. Furthermore, relationships were also observed between pupillary responses (specifically for PC4) and both frontal theta and posterior alpha power measures obtained from scalp EEG in Experiment 2, as well as for frontal theta and subsequent learning from mismatch stimuli (assessed using subsequent memory findings from a surprise recognition test). The authors suggest the findings indicate that MPEs elicit changes in attention, arousal and cognitive control which impact subsequent learning.

      Strengths:

      This manuscript has many strengths, including a clever study design, thoughtful integration of multiple neurocognitive measures, and a set of rigorous and technically sophisticated analyses, which reveal a large set of relationships among the measures and behavior. The findings demonstrating brain/physiology-behavior relationships are particularly important, in that they point to potential functional consequences of MPEs.

      Weaknesses:

      The technical proficiency and complexity of the study and analysis also presents a clear limitation and challenge for interpretation. It is likely that readers, even those that are quite knowledgeable about the methods, constructs, and questions being addressed will often struggle (as this reviewer did) to keep the large set of findings in mind and gain understanding of how they all fit together.

      Indeed, it seems like there many threads running together in the paper which make it challenging to find the through-line of the key findings. The authors do address some of the key questions motivating the paper in the Introduction, but the results are somewhat ambiguous with regard to the primary question of the study as to whether the detection of MPEs leads to interaction among cognitive control, attention, and arousal. To their credit, the authors tackle this question through both cross-correlation and formal mediation analyses, and summarize these in diagrammatic figures (Figure 3, Figure 6). Yet it is not resolved whether the results represent a clear answer pointing to independence, or rather a lack of statistical power, or ill-resolved formulation of the mediational relationship. In particular, the cross-correlation suggests that posterior alpha suppression in response to MPEs does precede frontal theta, yet this indirect relationship does not explain the variation in trial-by-trial RTs on mismatches. This suggests a potential model misspecification.

      In addition to the primary interaction issue mentioned above (between cognitive control, attention & arousal), the Introduction lays out a number of claims:

      (1) That pupil size will be more sensitive to strong than weak MPEs.

      (2) That MPE-linked increases in attention (indexed with posterior alpha suppression) and arousal (indexed with pupil size) will be linked to learning.

      (3) That MPE learning will vary as a function of prediction strength.

      Given the focus on learning, it is somewhat surprising that learning is not included in the mediation models. As the authors indicate in the Discussion, the use of trial-by-trial RT variation to drive the mediation model might be problematic, given that the RTs are sensitive to a range of factors beyond mnemonic prediction strength and also are under competing pressures (longer for mismatches than matches, due to surprise-linked slowing, but also faster following stronger rather than weaker mnemonic predictions). Thus, an alternative possibility might be to use trial-by-trial recognition of mismatches as the outcome variable in mediation models rather than trial-by-trial RT as the independent variable.

      A large component of the results (Sections 2 and 3) is devoted to analyses of cue-linked pupil and EEG processes that putatively reflect mnemonic predictions (i.e., occurring before picture probes are presented and match/mismatch detection, i.e., MPEs occur). Yet these Results and the subsequent pupillary PCA components (PC1 and PC5) that are elicited are not well-integrated with the primary themes of the paper or the causal hypotheses. One finding that does seem to figure prominently (in that it is mentioned in Abstract, Introduction & Discussion) relates to the amount of attention allocated to the mnemonic prediction generation. Yet this finding is not well emphasized in the Results themselves. Possibly it refers to the negative relationship between posterior alpha during memory retrieval and the magnitude of pupillary PC3 component, described in Section 3. But it was quite challenging to identify amongst the wealth of results described in this Section as well as the others. More generally, the large amount of findings described across all four lengthy Results sections makes it challenging for readers to discern what are the key ones that the authors would like to highlight.

      It is recommended that the authors do another pass through the paper to better highlight the most critical findings that they want to emphasize or which are most interpretable from a mechanistic and causal flow perspective and then de-emphasize or move other findings to the Supplemental Materials. Although the authors are to be commended for such a rigorous and comprehensive set of analyses, there are so many of them and findings, that the key points get buried and the reader needs to struggle potentially unnecessarily to identify the key take-away points.

      We thank Reviewer 1 for the helpful feedback on how the manuscript can be further strengthened. We recognize that the rich set of findings can overwhelm the reader, resulting in difficulty discerning the main take aways about the effects of mnemonic prediction errors. We particularly appreciate Reviewer 1’s encouragement to restructure the manuscript so as to focus on the findings reported in Sections 1 and 4 of the original revision; the current revision now focuses on these key observations.

      As part of this restructuring, Reviewer 1 also proposed moving the content from Sections 2 and 3 of the original revision to the Supplement. We agree with the Reviewer that the questions addressed in these sections on retrieval-related processes are not the main focus of the paper, but that they are informative in their own right. To avoid their getting lost in the Supplement, we decided that these results would be better served in a separate manuscript and thus we have removed them entirely. 

      We acknowledge in the revised manuscript that the mediation and cross-correlation analyses were exploratory and that these specific analyses may not be well powered in the current experiments. With respect to Reviewer 1’s concerns about the specification of the mediation models, we were motivated to test whether MPEs trigger an increase in cognitive control that in turn, triggers an increase in attention and/or arousal (Fig. 4a); as such, we designed the model to assess whether, on strong MPE trials, frontal theta mediates the relationship between prediction strength and attention/arousal. As noted in the manuscript and raised by Reviewer 1, mismatch RT here is an imperfect measure of trial-level prediction strength. Future experiments that selectively elicit strong MPEs and have a more controlled measure of trial-level prediction strength may be better equipped to address these questions about interactions between control, attention, and arousal. We agree with Reviewer 1 that models assessing subsequent memory as an outcome would be desirable. However, given that (a) we did not find strong evidence for interactions at the time of a strong MPE and (b) we only observed a relationship between frontal theta and subsequent memory (but not posterior alpha or pupil), subsequent memory mediation models do not appear to be well justified. Altogether, these findings illuminate open avenues for future research.

      Reviewer #2 (Public Review):

      Summary:

      The authors studied cognitive control and attention in response to mnemonic prediction errors (MPEs): situations in which the external reality violates internal memory-based predictions. The behavioral task first established strong versus weak predictions, and then either confirmed or violated these predictions. The authors examined markers of cognitive control (frontal theta) and attention (posterior alpha suppression, pupil response) while strong and weak predictions were confirmed or violated. They found increased cognitive control (frontal theta) for strong MPEs, which correlated with subsequent memory. Markers of attention (alpha suppression, pupil response) also accompanied strong MPEs but did not correlate with subsequent memory.

      Pupil response was investigated using an interesting approach that decomposes the response into different components, finding that different components respond earlier or later and show different correlations with MPEs and their strength. The authors also investigated how EEG, reaction time, and pupil responses correlated with one another, providing further insight into the mechanism underlying the response to MPEs. Together, the study points toward multiple control and attention mechanisms involved in MPE response and memory.

      Strengths:

      The study has a clear behavioral paradigm with multiple measures — behavioral, EEG, and pupillometry — that offer an investigation into different aspects of MPE response and memory.

      The study is also very comprehensive in looking at multiple phases in processing MPEs: the prediction phase (prior to the violation), the response to MPEs, and subsequent memory of MPEs, all within one study. Specifically, the link between neural mechanisms and subsequent memory is a major advancement, as most prior studies did not include this component. Mechanisms underlying subsequent memory of MPEs are theoretically important, as a primary function of MPEs is to promote learning and memory. As the authors mention, the different neural and pupillary signals are not robustly correlated, suggesting multiple mechanisms underlying MPE detections, which is interesting, offers avenues for future research, and can facilitate a better theory of how MPEs are processed in the brain. Finally, the decomposition of pupil response into different components and their correlation with behavior (RT during match/MPE detection) is interesting.

      Weaknesses:

      The methods are rigorous, and the data support the claims. The weaknesses are minor and are offered here as avenues for future research.

      (4) The relationships the authors find between brain measures and pupil components were largely not specific to mismatches/matches. Thus, the specificity of this relationship is untested.

      (5) The results with subsequent memory are important and address a major gap in the field that largely did not relate neural effects of MPE to subsequent memory. However, one major limitation of the study is that the authors did not test memory for matches. I understand the logic of avoiding testing matches. Because matches were repeated more times in the study, it’s not a fair comparison and could change participants’ overall criterion for old/new decisions. Future research could address this, e.g., by testing weak matches or potentially using a between-subject design.

      We appreciate Reviewer 2’s helpful feedback during the review process and encouraging comments on the strengths of the manuscript. We agree and note in the revision that it would be illuminating for future studies to contrast memory for events that violate and confirm mnemonic predictions.

      Comments on revised version

      The authors addressed all my concerns. I appreciate the authors’ thoughtful and detailed response.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It was challenging to read the paper with key findings happening at the time of the MPE presented first, and then to go “backwards in time” to examine process that occurred at preceding time periods (i.e., prior to probe presentation), and then again forward in time to examine learning related processes.

      (1) Restructuring the Results

      In this regard two distinct recommendations are made:

      - Move Sections 2 and 3 to Supplemental Materials, to maintain the focus on the key findings related to detection of MPEs and their effects on subsequent learning.

      - An alternative structure would be to present the cue-locked findings first, which relate to mnemonic predictions, then to those that occur when those predictions get violated, and finally to the learning process that occur following MPEs and which can be detected with subsequent recognition tests.

      We thank the Reviewer for this encouragement to restructure the manuscript. To address concerns regarding the density of the paper and the cohesiveness of the findings, we removed the content that was in Sections 2 and 3 of the original revision. We will publish those results in a separate paper with additional analyses to more directly address previously raised questions regarding the mechanisms indexed by frontal theta and posterior alpha at the time of retrieval.  

      (2) Integration of the summary figures

      At the minimum, a recommendation would be to better integrate Figure 6, and maybe various versions of Figure 3, earlier into the text, preferably even in the Introduction, and then repeatedly refer to them throughout the results. For example, in Figure 6, linking the leftmost panel of the figure to Sections 2 and 3 is critical, and the righthand panel to Sections 1 with the rightmost part related to learning explicitly linked to Section 4.

      We moved the summary figure up to now be Figure 1; we reference this figure in the Introduction and throughout the manuscript; and we additionally make reference in the figure to the association between frontal theta and PC3 at the time of a strong MPE as well as the cross-correlation outcome. We hope these modifications further aid the reader in identifying the main findings of the manuscript.

      (3) Specification of the mediation models

      Additionally, for the mediation models it is quite unclear why mismatch RT is treated as the index of MPE magnitude, as this seems to be where the problem may lie in model fitting. Why not think of this as an outcome variable (since would seem to be a causal outcome of the underlying functional processes elicited by mismatch detection)?

      We thank the reviewer for these thoughtful comments. Our goal in designing the mediation models in Fig. 4a-c was to test our hypothesis that strong MPEs trigger an increase in cognitive control that, in turn, triggers an increase in attention and/or arousal; thus, attention/arousal should be the outcome and cognitive control should be the mediator. Given that increases in attention and cognitive control were selectively observed for strong and not weak MPEs, the models were restricted to strong MPEs. We therefore needed a trial-level measure of MPE magnitude to determine whether stronger MPEs in the strong condition elicit greater attention/arousal by engaging more cognitive control. As such, we decided to use mismatch RT as a proxy measure of MPE magnitude; we acknowledge in the text that this measure is imperfect. While a model with mismatch RT as the outcome variable is possible, the relative timing of the attention/arousal effects (which largely occur following responses) would render interpretation to be more challenging. Had stronger evidence of indirect effects emerged in Figs. 4b-c, we could have conducted model comparison with RT as an outcome. We acknowledge in the text that these analyses were exploratory and characterization of potential indirect effects will require more data and a more precise measure of MPE magnitude.

      Alternatively, examining trial-by-trial recognition memory of mismatches as the relevant outcome variable would also seem to capture the functional process of interest. In this regard, have the authors examined whether trial-by-trial RT on mismatches predicts subsequent recognition of these items? If this direct relationship holds, it could be a target for mediation analyses in itself.

      We agree with the reviewer that in theory, a mediation model predicting subsequent memory would be a desirable test of an integrated model of the mechanisms underlying MPE-driven learning. However, the mediation analyses conducted to address the functional relationships at the time of a prediction error (Fig. 4) are not well-powered to begin with; this limitation is raised in the Results and the Discussion. A mediation model predicting subsequent memory would be similarly underpowered; given that there is not strong evidence for indirect effects at the time of a strong MPE, and that only frontal theta – and not posterior alpha nor pupil – predicts subsequent memory, we think that such a mediation model is not well justified to include. Such a model should be more directly tested by well-powered designs in future studies.

      (4) Hippocampal theta in the Introduction

      The Introduction discusses hippocampal theta as well as frontal theta yet also makes clear that the former is not really well-detected or analyzed using scalp EEG. Consequently, a recommendation would be to remove this paragraph from the Introduction, since it can be misleading and a “red herring” for the reader, and instead only bring up this point in the Discussion section, as a pointer to the need for future research using methods that may be more sensitive to hippocampal interactions with PFC regions.

      We appreciate this point and moved discussion of the hippocampus from the Introduction to the Discussion.

    1. eLife Assessment

      This manuscript presents an important and timely contribution by incorporating desolvation barriers into coarse-grained models of biomolecular condensates. The findings are convincing, supported by a clear physical model and systematic simulations showing effects on phase behavior, packing, and dynamics. Some clarification and broader context would improve the manuscript, but it provides a foundation that will be of use for developing more realistic coarse-grained interaction schemes.

    2. Reviewer #1 (Public review):

      This manuscript is very interesting and timely. By introducing the critical effects of desolvation barriers and solvent (water)-separated minima into the implicit-solvent potentials (of mean force, PMFs) for coarse-grained molecular dynamics simulations of biomolecular liquid-liquid phase separation (LLPS), this work fills a gap that should be apparent to researchers of protein folding in the past couple of decades but has so far escaped deserved attention such that these basic features of aqueous solvation have seldom, though not never, been invoked in recent studies of biomolecular condensates. Although the present paper deals almost exclusively with homopolymers, this work can be a foundation for the future development of a new, more physical coarse-grained interaction schemes for simulating amino acid sequence-dependent effects, which I presume is the authors' ongoing or next endeavor. The results presented in this manuscript are highly valuable.

      However, there is room for improvement in the authors' description of (i) the broader impact of effects of desolvation barrier and solvent-separated minimum in the thermodynamics of biomolecular condensates, especially with regard to the ramifications on hydrostatic pressure-dependent effects; (ii) the physical implication of using a 20-parameter hydropathy scale rather than a 210-parameter pairwise amino acid interaction scheme; and (iii) temperature-dependent effects, including the authors' discussion of "enthalpic" and "entropic" contributions. In all these aspects, the authors' discussion should be put in a more comprehensive context of the existing literature. At a few other places, description of the methods and results should be clarified as well.

      Comments on revised version

      The authors have thoroughly and adequately addressed all my previous concerns and suggestions. The manuscript is now significantly improved in terms of clarity and proper placement in the context of prior works on desolvation effects in protein conformations.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in the molecular simulation of biomolecular condensates. Most residue-level coarse-grained models used for IDP phase separation employ implicit solvent and represent effective interactions through relatively simple pairwise potentials. While these models have been very useful, they usually do not explicitly distinguish direct contacts from solvent-separated interactions, nor do they include an energetic barrier associated with water removal. This manuscript attempts to address that limitation by introducing desolvation-inspired terms into coarse-grained models and examining their consequences for phase behavior, chain conformations, dense-phase packing, and dynamics.

      The central idea is physically well motivated. Using a simple homopolymer model, the authors show that increasing the desolvation barrier suppresses phase separation, whereas stabilizing solvent-separated contacts enhances phase separation. They further show that solvent-separated interactions can reduce dense-phase over-compaction, which is a meaningful result given the known challenges in obtaining both accurate single-chain dimensions and realistic dense-phase properties from the same coarse-grained model. The finding that desolvation-like terms can reshape dense-phase packing without simply rescaling the overall interaction strength is interesting and could be useful for future model development. I also found the attempt to connect conformational changes across dilute and dense phases with thermal distance from the critical point to be intriguing. The dynamic analysis, including the FRAP-like simulations and the discussion of kinetic arrest during coarsening, adds another useful dimension to the work.

      Overall, I think this is a useful and potentially important contribution.

      Comments on revised version.

      The authors have addressed my earlier comment regarding conformational changes between the dilute and condensed phases. One small additional suggestion is that they may find two related studies useful in this context: Devarajan et al., Nature Communications (2024), on relationships between dilute-phase conformations and condensate material properties, and Wang et al., Chemical Science (2024), which examines sequence-dependent conformational changes during condensation for both model polyampholyte sequences and naturally occurring IDPs. These studies may provide some complementary context for the discussion. This is simply a literature suggestion and does not affect my overall assessment of the revised manuscript.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This manuscript is very interesting and timely. By introducing the critical effects of desolvation barriers and solvent (water)-separated minima into the implicit-solvent potentials (of mean force, PMFs) for coarse-grained molecular dynamics simulations of biomolecular liquid-liquid phase separation (LLPS), this work fills a gap that should be apparent to researchers of protein folding in the past couple of decades but has so far escaped deserved attention such that these basic features of aqueous solvation have seldom, though not never, been invoked in recent studies of biomolecular condensates. Although the present paper deals almost exclusively with homopolymers, this work can be a foundation for the future development of a new, more physical coarse-grained interaction scheme for simulating amino acid sequence-dependent effects, which I presume is the authors' ongoing or next endeavor. The results presented in this manuscript are highly valuable.

      We thank the reviewer for all these positive comments.

      However, there is room for improvement in the authors' description of (i) the broader impact of effects of desolvation barrier and solvent-separated minimum in the thermodynamics of biomolecular condensates, especially with regard to the ramifications on hydrostatic pressure-dependent effects; (ii) the physical implication of using a 20-parameter hydropathy scale rather than a 210-parameter pairwise amino acid interaction scheme; and (iii) temperature-dependent effects, including the authors' discussion of "enthalpic" and "entropic" contributions. In all these aspects, the authors' discussion should be put in a more comprehensive context of the existing literature. At a few other places, the description of the methods and results should be clarified as well. Accordingly, the authors should revise the manuscript to address the following items thoroughly within the revised manuscript (not merely in the response letter) with the additional references mentioned below included in the revised discussion:

      (1) In several places, e.g., on line 77 (p.2), the authors appear to suggest that "implicit-solvent representation" is the origin of the deficiency in commonly utilized coarse-grained potentials that this study is aiming to rectify. But desolvation barriers and solvent-separated minima are also features of implicit-solvent representations; they are just features that should be incorporated in more accurate implicit-solvent potentials. This point is stated quite clearly and accurately in the Abstract (p.1) but not consistently in the rest of the text. The authors should check the entire text carefully to ensure that a coherent, accurate perspective is presented.

      We thank the reviewer for pointing out this important issue. We agree that implicit-solvent representation itself is not the origin of the deficiency. Our intention is to incorporate desolvation-inspired effective terms within an implicit-solvent coarse-grained framework, because many commonly used implicit-solvent potentials do not directly account for the desolvation features, such as the desolvation barrier and the solvent-separated potential well.

      We have revised the Abstract, Introduction, Results, and Discussion to make this distinction consistent throughout the manuscript. The revised text now emphasizes that the model remains an implicit-solvent CG model, but contains additional effective terms inspired by desolvation features observed in all-atom PMFs.

      Corresponding changes:

      (1) (page 1, lines 17–20) The Abstract identifies the model as an implicit-solvent CG model with added desolvation terms:

      "Here, guided by all-atom simulations and experimental measurements, we develop a desolvation-aware implicit-solvent CG model by incorporating residue-level desolvation terms directly into the pairwise energy function and apply it to investigate LLPS of intrinsically disordered proteins."

      (2) (page 2, lines 80–81) The Introduction retains the implicit-solvent description of existing residue-level CG models:

      "Despite these advances, most residue-level CG models rely on implicit solvent representations, in which individual water molecules are not explicitly represented."

      (3) (page 2, lines 87–89) The specific limitation is identified as the absence of a direct account of the multi-step desolvation process:

      "More importantly, conventional implicit-solvent CG models used for LLPS usually do not directly account for the multi-step desolvation process that accompanies the transition from a dilute solution to a dense condensate."

      (4) (page 4, lines 172–174) The added terms are described as part of a desolvation-inspired effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (5) (page 15, lines 521–522) The Discussion restates the implicit-solvent model framework:

      "To address this challenge, we developed a desolvation-aware implicit-solvent CG framework that incorporates desolvation barrier and solvent-separated terms into the pairwise potential."

      (2) In the discussion of the importance of desolvation barriers and solvent-separated minima in the Introduction (pp.1-3), connections should be drawn to recent works that utilize these PMF features to rationalize hydrostatic pressure (P)-modulated effects on biomolecular LLPS, including the P-dependent reentrant phase separation of alpha elastin; see Cinar et al. (2019) Chem Eur J 25:13049 (https://chemistryeurope.onlinelibrary.wiley.com/doi/full/10.1002/chem.201902210) and references therein, especially discussions around Figures 10, 11 & 13 in this reference.

      We thank the reviewer for bringing this literature to our attention. We agree that pressure-modulated LLPS provides important context for the physical relevance of desolvation barriers and solvent-separated minima. We have therefore expanded the Introduction and Discussion to connect our model to prior work on hydrostaticpressure-dependent condensate behavior, including pressure-dependent reentrant phase separation of alpha-elastin.

      Corresponding changes:

      (1) (page 3, lines 93–95) The Introduction connects the PMF features to hydrostaticpressure-dependent LLPS:

      "Related studies on hydrostatic pressure effects have further suggested that desolvation barriers and solvent-separated minima can help rationalise pressure-modulated LLPS behaviors, including the pressure-dependent reentrant phase separation of α-elastin Cinar et al. (2019, 2018)."

      (2) (page 15, lines 533–537) The Discussion states the pressure-dependent implication conservatively:

      "These findings may also provide a useful physical basis for future studies of pressure-dependent condensate behavior, as pressure-induced changes in hydration, solvent-separated states, and desolvation barriers have been proposed to contribute to pressure-modulated and reentrant LLPS Dias and Chan, 2014); Cinar et al. (2019, 2018)."

      (3) In the lower panels of Figures 2D, E (p.5), what do the differently colored small circles in the double-minimum free energy profiles represent? Does the color shading have the same meaning as that in the upper panels? If so, what do the positions of the circles on the free energy profile represent? The authors should clarify this.

      We thank the reviewer for identifying this ambiguity. The small circles in the lower panels of Figures 2D and 2E are qualitative schematic representations of residue-pair configurations along the effective pair-potential profile. Their blue and green colors distinguish the barrier-variation and solvent-separated-well cases, respectively; they are not a quantitative scale and do not encode temperature or population magnitude. The positions of the circles indicate the direct-contact, barrierregion, or solvent-separated regions, while the density of circles schematically represents the population of configurations.

      We have clarified this interpretation in the Figure 2 caption and aligned the Results text with the redistribution among direct-contact, barrier-region, and solvent-separated states.

      Corresponding changes:

      (1) (page 6, Figure 2D and E lower panels) The schematics distinguish low and high ε_b or ε_ss and use the density and position of the circles to depict populations in the direct-contact, barrier-region, and solvent-separated regions; the blue and green colors distinguish the two parameter families and are not a quantitative scale.

      (2) (page 6, Figure 2 caption) The caption defines the population encoding used in the lower panels:

      "The lower panels schematically illustrate how changes in ε<sub>b</sub> and ε<sub>ss</sub> alter the distribution of residue-pair configurations. The small circles indicate schematic populations of residue-pair configurations along the potential profile, with denser circles representing a higher population."

      (3) (page 5, lines 210–213) The Results text connects the schematics to redistribution among the three residue-pair states:

      "These opposing effects suggest that the desolvation potential regulates macroscopic phase behavior by redistributing residue-pair configurations between direct-contact, barrier-region, and solvent-separated states (lower panels of Figure 2D and E)."

      (4) The discussion regarding entropy and enthalpy around Figure 2 is quite confusing as it stands. What do the authors mean exactly by the association of entropy or enthalpy with the desolvation barrier of the solvent-separated minimum? Are they referring to conformational entropy?

      We thank the reviewer for pointing out this ambiguity. We agree that our original wording around entropy and enthalpy could be misleading, because it might imply a rigorous thermodynamic decomposition of the PMF. In the revised manuscript, we have therefore clarified that the effect of the desolvation barrier refers to an entropyrelated configurational restriction of residue-pair configurations near the barrier region, rather than the overall conformational entropy of the entire chain. We also replaced the previous "enthalpic stabilization" wording with "effective free-energy stabilization" to avoid implying that the solvent-separated minimum is treated as a purely enthalpic contribution.

      Corresponding changes:

      (1) (page 5, lines 199–201) The barrier effect is described in terms of the sampled residue-pair population:

      "Analysis of residue-residue radial distribution functions showed that higher ε<sub>b</sub> suppresses the population of configurations near the barrier region (Figure 2—figure Supplement 1C)."

      (2) (page 5, lines 201–202) The entropy-related statement is restricted to configurational sampling near the barrier:

      "This reduction in the statistical weight of barrier-region configurations can be interpreted as an entropy-related configurational restriction and thus disfavors phase separation."

      (3) (page 5, lines 209–210) The solvent-separated minimum is described as an effective free-energy contribution:

      "This solvent-separated minimum provides effective free-energy stabilization for water-mediated configurations and thereby promotes phase separation."

      (4) (page 6, Figure 2D and E lower panels) The schematic headings are "Barrier-mediated Restriction" and "Solvent-separated Stabilization", avoiding a strict entropy-enthalpy decomposition.

      (5) Do the authors assume that the PMF (effective implicit-solvent potential) is a purely enthalpic term? It appears to be the authors' assumption. If so, the assumption has to be stated clearly in their discussion of "entropy" vs "enthalpy" around Figure 2.

      We thank the reviewer for raising this important point. We do not assume that the PMF obtained from all-atom simulations is a purely enthalpic term. The PMF is a free-energy profile that contains enthalpic and entropic contributions. The current manuscript defines the PMF as −k<sub>B</sub> T lnP(r), uses the all-atom PMFs to motivate a nonbonded effective coarse-grained potential, and describes the solvent-separated minimum as providing effective free-energy stabilization. We do not perform a rigorous enthalpy-entropy decomposition, and the revised wording avoids implying such a decomposition.

      Corresponding changes:

      (1) (page 16, lines 594–595) The Methods define the PMF as a free-energy profile obtained from the radial probability density:

      "The potential of mean force (PMF) was computed as PMF(r) = −k<sub>B</sub>T ln P(r), where P(r) is the radial probability density obtained from the production trajectory."

      (2) (page 5, Figure 1 caption) The CG interaction is labeled as an effective potential rather than as an enthalpic PMF decomposition:

      "Pairwise effective potential incorporating desolvation-inspired terms. Different curves correspond to different desolvation parameters."

      (3) (page 4, lines 172–174) The parameters are described as shaping an effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (4) (page 5, lines 209–210) The solvent-separated contribution is described using free-energy language:

      "This solvent-separated minimum provides effective free-energy stabilization for water-mediated configurations and thereby promotes phase separation."

      (6) Closely related to points 3-5 above, it should be stated clearly that the "temperature" used in the authors' simulations does not represent experimental temperature if the authors are using purely enthalpic effective potentials because PMFs are in fact temperature-dependent. This clarification is necessary to avoid misunderstanding. In this regard, it should be noted that temperature-dependent effective interactions have been used for modeling biomolecular condensates in analytical theory (Lin, Song, Forman-Kay & Chan, J Mol Liq 2017, already in the citation list) as well as in coarse-grained molecular dynamics simulations [Dignon et al. (2019) ACS Cent Sci 5:821-830 (https://pubs.acs.org/doi/10.1021/acscentsci.9b00102); Chakravarti & Joseph (2025) Protein Sci 34:e70284 (https://onlinelibrary.wiley.com/doi/10.1002/pro.70284)]. The latter two studies, not cited currently, are particularly relevant and thus should be cited because the authors may wish to incorporate temperature-dependent features in their ongoing or future effort in constructing a more comprehensive coarse-grained interaction scheme for biomolecular LLPS simulation.

      We agree with the reviewer that the simulation temperature should be interpreted carefully. In the present simulations, the effective potential is temperature-independent within each chosen parameter set. Therefore, the reduced temperature primarily serves as a model temperature controlling the relative strength of thermal fluctuations, rather than as a direct experimental temperature. We have clarified this point in the revised manuscript and added relevant references on temperature-dependent effective interactions, which represent an important direction for future model development.

      Corresponding changes:

      (1) (page 4, lines 177–179) The manuscript states that the absolute simulation temperature is not an experimental temperature:

      "Because the effective interaction parameters used in the model are temperature-independent, the absolute simulation temperature should not be directly interpreted as an experimental temperature."

      (2) (page 4, lines 182–184) The reduced temperature is identified as a model temperature:

      "Accordingly, T<sup>*</sup> should be interpreted primarily as a model temperature that controls the relative strength of thermal fluctuations, rather than as having a direct quantitative correspondence with experimental temperature."

      (3) (page 14, lines 481–485) The FUS LC temperature comparison is framed cautiously:

      "It is worth noting that the residue-level coarse-grained models used here employ temperature-independent effective interaction parameters. As a result, the temperature values reported here cannot be interpreted as quantitatively equivalent to experimental temperatures, particularly when they deviate substantially from room-temperature conditions."

      (4) (page 15, lines 563–567; continues on page 16, lines 568–569) The Discussion identifies temperature-dependent effective interactions and a corresponding future extension:

      "In addition, effective interactions themselves can be temperature-dependent, as demonstrated in analytical theories and coarse-grained simulations of biomolecular condensates Lin et al. (2017); Dignon et al. (2019); Chakravarti and Joseph (2025). Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      (7) In tackling "entropy" vs "enthalpy", it should be noted that the temperature dependence of the effective interactions entails an entropic contribution (which is itself temperature dependent) in addition to conformational entropy. As for the effective potential with desolvation barrier and solvent-separated minimum, it should be noted that the decomposition into entropic and enthalpic contributions at the direct contact, desolvation barrier, and solvent-separated minimum can be dramatically different, see, e.g., MaCallum et al. (2007) PNAS 104:6206-6210 (https://www.pnas.org/doi/full/10.1073/pnas.0605859104) and references therein.

      We thank the reviewer for this important clarification. We agree that temperature-dependent effective interactions can contain entropic contributions beyond conformational entropy and that the balance of enthalpic and entropic contributions may differ among the direct-contact minimum, desolvation barrier, and solvent-separated minimum. The present model does not decompose the PMF into temperature-dependent enthalpic and entropic components; accordingly, we have avoided assigning those components to individual PMF features. The Discussion cites explicit-solvent PMF analyses when noting residue-pair and temperature dependence and identifies temperature-dependent desolvation parameters as an important future extension.

      Corresponding changes:

      (1) (page 15, lines 561–563) The Discussion cites explicit-solvent PMF work when noting residue-pair and temperature dependence:

      "Explicit-solvent PMF analyses have shown that desolvation barrier heights and solvent-separated minima can differ substantially among residue pairs and may also exhibit temperature dependence Cinar et al. (2019); Debiec et al. (2014); MacCallum et al. (2007)."

      (2) (page 15, lines 563–566) The manuscript states that effective interactions can themselves depend on temperature:

      "In addition, effective interactions themselves can be temperature dependent, as demonstrated in analytical theories and coarse-grained simulations of biomolecular condensates Lin et al. (2017); Dignon et al. (2019); Chakravarti and Joseph (2025)."

      (3) (page 15, lines 566–567; continues on page 16, lines 568–569) Temperature-dependent desolvation parameters are identified as a future model extension:

      "Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      (8) P.7, line 340: The proportionality relation follows directly from the standard FloryHuggins result T_c = T chi(T)/chi_c, thus the proportionality constant is exactly 1/chi_c. Is this the standard relation that the authors are invoking here? The authors should clarify this.

      We thank the reviewer for pointing out the missing intermediate steps. Yes, the relation we invoked is based on the standard Flory-Huggins critical condition. In the revised manuscript, we have expanded the derivation to explicitly show how the critical condition chi(T_c) = chi_c leads to the relation between chi(T_sim) - chi_c and the normalized thermal distance (T_c - T_sim)/T_sim.

      We also revised the wording to avoid presenting this as a universal law. The relation is now presented as a simulation-supported trend within the present model, rationalized by a simplified linear-response assumption between Delta R_g and the excess interaction strength.

      Corresponding changes:

      (1) (page 7, lines 263–264) The critical-condition substitution is now shown explicitly:

      "At the critical point, χ(T<sub>c</sub>) = χ<sub>c</sub>, which gives ε<sub>eff</sub> = k<sub>B</sub>T<sub>c</sub>χ<sub>c</sub>. Substituting this relation into the expression for χ(T<sub>sim</sub>) yields χ(T<sub>sim</sub>) = χ<sub>c</sub>T<sub>c</sub>/T<sub>sim</sub>."

      (2) (page 7, line 265) The resulting relation is written as Equation (2):

      "χ(T<sub>sim</sub>) − χ<sub>c</sub> = χ<sub>c</sub> (T<sub>c</sub> − T<sub>sim</sub>)/T<sub>sim</sub>."

      (3) (page 7, lines 266–267) The fixed-chain-length assumption is stated explicitly:

      "For systems with the same chain length, χ<sub>c</sub> is a fixed constant. Thus, the deviation from the critical interaction parameter is directly related to the rescaled thermal distance (T<sub>c</sub> − T<sub>sim</sub>)/T<sub>sim</sub>."

      (9) The study on dynamic consequences on pp.8-11 is interesting, but clarifications are necessary:

      (i) The vertical schematic in Figure 4A should be explained in detail in its entirety. As it stands, no explanation is provided either in the figure caption or in the text. In particular, what does "elasticity driven" refer to?

      (ii) The top snapshot in Figure 4A is labeled t_sim = 0 ns. Does it mean that the snapshot shown is the only chain configuration that the authors used to start the simulation, and that the snapshot does NOT represent the result of any time evolution, no matter how short the duration is? However, if that is the case, why is this snapshot identified with spinodal decomposition if it is not the product of a time evolution from a more homogeneous configuration?

      (iii) Related to (ii) - do the rectangular boxes shown represent the entire simulation box or just part of the box containing the polymer chains? One would imagine that if the top snapshot represents spinodal decomposition, the simulation would have been started at a more uniform distribution a short time prior? Why is this not the case?

      (iv) What precisely do the small yellow beads and black-colored springs in the zoomin image of Figure 4E represent?

      We thank the reviewer for all these inspiring comments and questions. We agree that the original Figure 4 schematic did not sufficiently explain the sequence of dynamical events and the meaning of several graphical elements. We have therefore revised both the Figure 4 caption and the Results text to make the schematic self-contained and to clarify how it relates to the quantitative analyses in Figure 4F and G.

      First, we replaced the phrase "elasticity driven" with a more precise description of "viscoelastic resistance". In the revised text, interfacial tension is described as favoring domain fusion thermodynamically, whereas transient inter-chain network connectivity within dense domains generates viscoelastic resistance to the deformation required for coalescence. This revision avoids implying that elasticity is the driving force and instead identifies it as a resistance that delays domain fusion kinetically during the plateau regime.

      Second, we clarified the meaning of t_sim = 0 ns and its relation to spinodal decomposition. The system was equilibrated at a supercritical temperature to obtain a homogeneous one-phase state and was then instantaneously quenched to the target temperature. The label t_sim = 0 ns denotes the first snapshot immediately after the quench. It is not the only initial configuration used in all simulations; the reported kinetic metrics were averaged over six independent slab simulation replicas.

      The t_sim = 0 ns snapshot is therefore described as a homogeneous but thermodynamically unstable post-quench state. Spinodal decomposition refers to the subsequent amplification of the initial density fluctuations after the quench, including the development of interconnected density fluctuations within 1-2 ns, rather than to a preceding evolution represented by the t_sim = 0 ns snapshot.

      Third, we clarified that the rectangular snapshots in Figure 4A show the entire simulation box viewed along the z-axis. The subsequent snapshots show how post-quench density fluctuations grow and reorganize into dense domains during spinodal decomposition.

      Finally, we clarified the symbols in the zoom-in schematic of Figure 4E. Yellow beads now denote residues involved in transient inter-chain contacts, and black springs denote schematic network connections formed by these contacts. These elements are meant to illustrate transient network connectivity and are not additional simulated particles or force-field terms.

      Corresponding changes:

      (1) (page 10, Figure 4A) The vertical schematic now labels the progression as "Thermodynamic instability", "Kinetic arrest (viscoelastic resistance)", "Domain coarsening (interfacial-tension dominated)", and "Dynamic equilibrium (chain self-diffusion)".

      (2) (page 10, Figure 4E) The plateau schematic labels the competing effects as "Interfacial Tension" and "Transient network resistance".

      (3) (page 10, Figure 4 caption) The caption defines the snapshots and the vertical schematic:

      "Upper snapshots show the simulation box along the z-axis at t<sub>sim</sub> = 0, 10, and 500 ns. The vertical schematic summarizes the dynamical progression described in the main text, from the post-quench spinodal instability to kinetic arrest, domain coarsening, and dynamic equilibrium."

      (4) (page 11, lines 368–370) The first recorded time point after the quench is defined explicitly:

      "Here, t<sub>sim</sub> = 0 ns denotes the first snapshot immediately after the temperature quench, corresponding to a homogeneous but thermodynamically unstable nonequilibrium state."

      (5) (page 10, Figure 4 caption) The yellow beads and black springs are defined:

      "In the zoom-in view, yellow beads denote residues involved in transient inter-chain contacts, and black springs denote schematic network connections formed by these contacts."

      (6) (page 12, lines 401–404) The Results explain the physical meaning of the transient network:

      "In the zoom-in schematic in Figure 4E, this transient network is represented by connections between residues involved in inter-chain contacts, illustrating how multivalent interactions can resist domain deformation during the plateau regime."

      (10) In discussing dynamic effects, it is useful to draw connections to related works on the effect of chain flexibility on "aging" of condensate [Biswas & Potoyan (2024) PRX 45:9222-9245 (https://journals.aps.org/prxlife/abstract/10.1103/PRXLife.2.023011)] and characterization of viscoelasticity in simulations of biomolecular condensates [Tejedor et al. (2023) J Phys Chem B 127:4441-4459 (https://pubs.acs.org/doi/10.1021/acs.jpcb.3c01292)], as the effects of desolvation can be explored further based on these prior works.

      We thank the reviewer for these important references. We have added connections to simulation studies of condensate viscoelasticity and aging. The revised manuscript now places our dynamic results in the context of transient network connectivity, chain flexibility, sticker lifetime, desolvation-associated rigidification, and viscoelastic or aging-like material behavior.

      We present these connections conservatively as relevant context and as future directions for extending the current model, rather than claiming a new universal dynamic mechanism.

      Corresponding changes:

      (1) (page 12, lines 397–399) The dynamics section cites simulation-based rheological analyses of condensate viscoelasticity:

      "Similar viscoelastic effects have recently been quantified in molecular simulations of biomolecular condensates using rheological analyses of time-dependent material properties Tejedor et al. (2023)."

      (2) (page 12, lines 399–401) The manuscript connects condensate aging to chain flexibility, sticker lifetime, and desolvation-associated rigidification:

      "Molecular simulations of condensate aging have further highlighted the roles of chain flexibility, sticker lifetime, and desolvation-associated rigidification in promoting more solid-like states Biswas and Potoyan (2024)."

      (3) (page 12, lines 423–427) The kinetic interpretation is connected to viscoelastic andaging-like behavior:

      "The sensitivity of kinetic arrest and coarsening dynamics to desolvation parameters underscores the importance of incorporating desolvation features into coarse-grained potentials for more physically plausible molecular simulations of LLPS, especially when connecting microscopic interaction lifetimes to emergent viscoelastic or ageing-like material behavior."

      (4) (page 16, lines 569–572) The Discussion identifies simulation-based rheological analysis as a future direction:

      "An additional direction would be to combine these potentials with simulation-based rheological analyses to quantify how desolvation reshapes condensate viscoelasticity, aging-like maturation, and long-time material relaxation Tejedor et al. (2023); Biswas and Potoyan (2024)."

      (11) Much of the present study is based on the original HPS formulation of Dignon et al. (2018). In this regard and also in anticipation of future development of improved interaction schemes, several issues should be stated and discussed, even if briefly:

      (i) The original HPS model has a basic shortcoming in accounting for the relative interaction strengths of, among others, arginine vs lysine residues [Das et al. (2020) PNAS 117:28795-28805 (https://www.pnas.org/doi/10.1073/pnas.2008122117)].

      (ii) Compared to 210-parameter pairwise interaction schemes, such as KH in Dignon et al. (2018) and Joseph et al. (2021), the 20-parameter interaction scheme is likely too restrictive to account for pairwise amino acid residue interactions [Wessén et al. (2022) J Phys Chem B 45:9222-9245 (https://pubs.acs.org/doi/10.1021/acs.jpcb.2c06181)].

      (iii) The height of the desolvation barrier may vary significantly for different amino acid residue pairs, see, e.g., Figure 11 of Cinar et al. (2019) mentioned above (and references therein). The authors should discuss these nuances in the revised version. They may also wish to take them into consideration in future investigations.

      We thank the reviewer for the suggestion to clarify these limitations. We have revised the Discussion to acknowledge explicitly the limitations of the 20-parameter hydropathy-scale representation relative to more flexible 210-parameter pairwise interaction schemes for describing amino-acid-pair interactions. We have also added discussion emphasizing that future desolvation-aware models should incorporate residue-pair-specific parameters for the desolvation barrier and solvent-separated potential well.

      Corresponding changes:

      (1) (page 13, lines 445–446) The scope of the averaged baseline parameterization is stated explicitly:

      "This uniform parameterization captures the generic desolvation features of the PMFs but does not resolve residue-pair-specific variations in desolvation energetics."

      (2) (page 15, lines 551–554) The Discussion identifies the limitation of the 20-parameter HPS representation, including Arg/Lys interactions:

      "In particular, HPS-type models use a 20-parameter hydropathy-scale representation, which is useful for capturing generic IDP phase behavior but is not flexible enough to resolve residue-pair-specific chemical effects, such as the distinct interaction patterns of arginine and lysine residues Das et al. (2020)."

      (3) (page 15, lines 554–557) The greater flexibility of 210-parameter pairwise schemes is described:

      "More general 210-parameter pairwise interaction schemes, such as KH-type and related residue-pair-specific models, provide greater flexibility for encoding amino acid-pair preferences and capturing sequence-specific interaction heterogeneity Dignon et al. (2018b); Joseph et al. (2021); Wessén et al. (2022)."

      (4) (page 15, lines 557–561) The limitation of using one averaged desolvation parameter set is stated:

      "Second, the present desolvation model employs a single set of averaged parameters (α<sub>b</sub>, α<sub>ss</sub>) for all residue pairs. While this simplification is effective for isolating the generic physical consequences of desolvation, it has limitations in describing the pair-specific variations in the desolvation barrier and the solvent-separated minimum."

      (5) (page 15, lines 566–567; continues on page 16, lines 568–569) Residue-specific and temperature-dependent parameters are identified as a future extension:

      "Future extensions of the model could therefore incorporate residue-specific and temperature-dependent desolvation parameters derived from bottom-up parameterization or expanded experimental datasets, thereby enhancing predictive accuracy for sequence-dependent LLPS."

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in the molecular simulation of biomolecular condensates. Most residue-level coarse-grained models used for IDP phase separation employ implicit solvent and represent effective interactions through relatively simple pairwise potentials. While these models have been very useful, they usually do not explicitly distinguish direct contacts from solvent-separated interactions, nor do they include an energetic barrier associated with water removal. This manuscript attempts to address that limitation by introducing desolvation-inspired terms into coarse-grained models and examining their consequences for phase behavior, chain conformations, dense-phase packing, and dynamics. Strengths:

      The central idea is physically well motivated. Using a simple homopolymer model, the authors show that increasing the desolvation barrier suppresses phase separation, whereas stabilizing solvent-separated contacts enhances phase separation. They further show that solvent-separated interactions can reduce densephase over-compaction, which is a meaningful result given the known challenges in obtaining both accurate single-chain dimensions and realistic dense-phase properties from the same coarse-grained model. The finding that desolvation-like terms can reshape dense-phase packing without simply rescaling the overall interaction strength is interesting and could be useful for future model development. I also found the attempt to connect conformational changes across dilute and dense phases with thermal distance from the critical point to be intriguing. The dynamic analysis, including the FRAP-like simulations and the discussion of kinetic arrest during coarsening, adds another useful dimension to the work.

      Weaknesses:

      At the same time, there are several places where the manuscript would benefit from more careful framing. First, the desolvation terms are still effective coarse-grained parameters rather than a direct representation of water molecules. The language sometimes gives the impression that desolvation is being treated explicitly, whereas the model introduces desolvation-inspired effective interactions into an implicitsolvent framework.

      We thank the reviewer for the positive assessment and constructive suggestions. We agree that the desolvation terms should be described as effective coarse-grained parameters rather than explicit water molecules. We have revised the manuscript to describe the model as a desolvation-aware implicit-solvent coarse-grained framework with desolvation-inspired effective interaction terms.

      Corresponding changes:

      (1) (page 1, lines 17–20) The Abstract identifies the model as an implicit-solvent CG model:

      "Here, guided by all-atom simulations and experimental measurements, we develop a desolvation-aware implicit-solvent CG model by incorporating residue-level desolvation terms directly into the pairwise energy function and apply it to investigate LLPS of intrinsically disordered proteins."

      (2) (page 4, lines 151–153) The Results describe the added contributions as desolvation-related effective terms:

      "These observations underscore the importance of incorporating desolvation-related effective terms and exploring the effects of different desolvation strengths on the thermodynamics and kinetics of protein LLPS."

      (3) (page 4, lines 172–174) The pair interaction is described as a desolvation-inspired effective potential:

      "Together, these parameters shape the desolvation-inspired effective potential and modulate the statistical balance between direct-contact and solvent-separated configurations."

      (4) (page 15, lines 541–543) The Discussion emphasizes that the framework retains water-mediated features within an implicit-solvent representation:

      "By retaining key water-mediated features while preserving the computational efficiency of implicit-solvent representations, this framework provides a mechanistic means to decouple overall phase-separation propensity from condensed-phase packing."

      Second, the conformational analysis is interesting, but the broader context of prior work on dilute-to-dense phase conformational reorganization of IDPs could be more clearly discussed. This would help clarify what is new in the present work, whether it is the conformational change itself, its dependence on desolvation terms, or the proposed scaling with distance from the critical point.

      We thank the reviewer for this suggestion. We agree that the conformational change itself should be placed in the context of prior work. The contribution of the present analysis is not simply the observation that IDP conformations can reorganize upon condensation. Rather, we examine how desolvation-inspired effective terms modulate dilute- and dense-phase conformations and how the conformational change correlates with thermal distance from the critical point within the present model.

      We have revised the Results section discussing Figure 3 to cite prior work and to state the interpretation of ΔR_g more clearly.

      Corresponding changes:

      (1) (page 6, lines 230–232) Prior work on conformational reorganization upon condensation is cited:

      "Previous studies have shown that IDP condensation can reorganize chain conformations by redistributing the balance between intra-chain and inter-chain interactions Wei et al. (2017); Hazra and Levy (2021); Tesei et al. (2021); von Bülow et al. (2025)."

      (2) (page 7, lines 241–242) The phase-dependent conformational response to the desolvation parameters is introduced:

      "In addition to the difference between dilute- and dense-phase conformations, varying the desolvation parameters further reveals a phase-dependent conformational response (Figure 3A–C)."

      (3) (page 7, lines 243–244) The stronger response in the dilute phase is stated directly:

      "Increasing ε<sub>b</sub> or decreasing ε<sub>ss</sub> shifts the dilute-phase R<sub>g</sub> distributions toward larger values, whereas the dense-phase R<sub>g</sub> remains comparatively insensitive to these parameter changes."

      (4) (page 7, lines 247–249) The source of the desolvation-dependent variation in ΔR_g is identified:

      "As a result, the desolvation-dependent variation in ΔR<sub>g</sub> = R<sub>g</sub><sup>dense</sup> − R<sub>g</sub><sup>dilute</sup> arises predominantly from the conformational changes of isolated chains in the dilute phase."

      (5) (page 7, lines 254–256) The observed relationship is presented as an approximate trend in the simulated systems:

      "Notably, data from the simulated systems approximately follow a common trend, revealing a strong correlation between the magnitude of conformational change and the thermal distance to the phase transition point (R<sup>2</sup> = 0.942, Figure 3D)."

      Third, the dynamic results are potentially useful, but the manuscript should more clearly articulate what is nontrivial beyond the expected slowing of local rearrangements by an added barrier in the potential.

      Overall, I think this is a useful and potentially important contribution.

      We thank the reviewer for this constructive comment and the positive overall assessment. We have revised the dynamics section to clarify that the nontrivial result lies in the competition between two effects: although the desolvation barrier directly slows local rearrangements, its reduction of dense-phase packing can reverse the net mobility trend at fixed temperature. At matched thermodynamic quench depth, the intrinsic slowing associated with energy-landscape roughness becomes evident. We also clarified that desolvation modulates transient kinetic arrest and domain-scale coarsening, not only local rearrangements.

      Corresponding changes:

      (1) (page 11, lines 353–357) The fixed-temperature and matched-quench-depth analyses are summarized as opposing contributions:

      "Together, the fixed-temperature and renormalized analyses in Figure 4C and D reveal two distinct and opposing contributions of desolvation to condensate dynamics. At fixed temperature, increasing ε<sub>b</sub> loosens dense-phase packing and thereby increases the measured diffusion coefficient, whereas at matched thermodynamic quench depth, the same parameter change suppresses chain mobility by roughening the microscopic energy landscape and slowing local rearrangements."

      (2) (page 11, lines 358–361) The multiscale interpretation is stated explicitly:

      "Condensate dynamics therefore emerge from a balance between density-regulated mobility and energy-landscape-regulated mobility, with macroscopic packing determining the dominant trend and microscopic barrier roughness imposing an additional kinetic modulation. This interplay highlights how desolvation reshapes condensate dynamics across multiple physical scales."

      (3) (page 12, lines 421–423) The dynamics section distinguishes the result from simple local slowing:

      "This picture shows that desolvation does more than slow down local chain rearrangements through an added barrier. It also regulates the balance between fluctuation growth, transient arrest, and domain coarsening, thereby shaping the evolution of phase-separated domains."

      Reviewer #2 (Recommendations for the authors):

      (1) The model is physically motivated and useful, but I would encourage the authors to be more precise in describing the added terms as desolvation-inspired effective interactions rather than explicit desolvation.

      We thank the reviewer for the comment and suggestion. We have revised the manuscript accordingly and describe the added terms as desolvation-inspired effective interactions within an implicit-solvent CG framework throughout the Abstract, Results, and Discussion. More detailed changes are provided in our response to the first point raised in Reviewer #2's Public Review.

      (2) The desolvation barrier is introduced as part of the equilibrium pair potential, and therefore it is expected to affect not only kinetics but also the phase boundary through changes in the configurational partition function. The manuscript would benefit from clarifying this point, since the term "barrier" may otherwise suggest a primarily kinetic role. In particular, the authors should explain whether the observed shift in T_c reflects a change in the effective pair attraction, for example, through the integrated Boltzmann weight or second virial coefficient, rather than only an entropic penalty associated with restricted configurations.

      We thank the reviewer for this important point. We agree that the desolvation barrier is part of the equilibrium pair potential and therefore affects the phase boundary through the Boltzmann-weighted sampling of residue-pair configurations, not only through kinetic slowing.

      Following this recommendation, we added a bead-level second virial coefficient analysis based on the effective pair potential. This analysis provides a pair-potentiallevel measure of the integrated effective attraction and clarifies why increasing the barrier lowers T_c, whereas stabilizing the solvent-separated minimum raises T_c.

      Corresponding changes:

      (1) (page 5, lines 203–205) The equilibrium role of the barrier is stated explicitly:

      "At the pair-potential level, the desolvation barrier modifies the equilibrium Boltzmann weight and thereby alters the integrated effective attraction, as quantified by the bead-level second virial coefficient B<sub>2</sub>."

      (2) (page 5, lines 205–207) The barrier-dependent second virial coefficient is connected to the shift in critical temperature:

      "Specifically, increasing ε<sub>b</sub> makes B<sub>2</sub>/σ<sup>3</sup> larger (Figure 2—figure Supplement 1G), indicating a weaker integrated effective attraction and providing a thermodynamic basis for the lower T<sub>c</sub><sup>*</sup>."

      (3) (page 5, lines 207–209) The solvent-separated-well trend is linked to a smaller second virial coefficient:

      "By contrast, deepening the solvent-separated well ε<sub>ss</sub> elevates T<sub>c</sub><sup>*</sup> (Figure 2E), which is associated with the enhanced population of solvent-separated configurations and a smaller B<sub>2</sub>/σ<sup>3</sup> (Figure 2—figure Supplement 1D, H)."

      (4) (page 18, lines 662–665) The Methods specify the integration range and the quantity used to compare integrated effective attraction:

      "The upper limit of integration r<sub>c</sub> is set as 3σ, which is sufficiently large to capture the full range of interactions while ensuring numerical convergence. The reduced value B<sub>2</sub>/σ<sup>3</sup> was used to compare the integrated effective attraction under different desolvation parameters."

      (5) (Figure 2—figure supplement 1G, H) The new panels report the integrated effective attraction:

      "(G, H) Bead-level second virial coefficient (B<sub>2</sub>/σ<sup>3</sup>) calculated from the effective pair potential under varying ε<sub>b</sub> at fixed ε<sub>ss</sub> = 0.02 kcal/mol (G) and varying ε<sub>ss</sub> at fixed ε<sub>b</sub> = 3.12 cal/mol (H)."

      (3) The conformational analysis in Figure 3 is interesting and potentially important. It would help to better place this result in the context of prior work showing dilute-todense phase conformational reorganization of IDPs, and to clarify what is new here beyond that broader observation.

      We thank the reviewer for the comment and suggestion. We have revised the Results section discussing Figure 3 to place dilute-to-dense conformational reorganization of IDPs in the context of previous studies and then to emphasize the specific contribution of the present work.

      The revised text clarifies that, within the present model, desolvation-inspired interactions mainly regulate chain conformations in the dilute phase, whereas dense-phase conformations remain comparatively insensitive. Detailed changes are provided in our response to the second point raised in Reviewer #2's Public Review.

      (4) The proposed scaling between ΔR_g and distance from the critical point is intriguing, but the argument relies on simplifying assumptions. I would present this more as an empirical scaling supported by a plausible theoretical argument rather than a general result.

      We thank the reviewer for this helpful suggestion. We agree that the correlation between Delta R_g and the distance from the critical point relies on simplifying assumptions and should not be presented as a general law. In the revised manuscript, we have softened the interpretation and now present this relationship as an empirical correlation supported by a simplified Flory-Huggins-based theoretical argument.

      Corresponding changes:

      (1) (page 7, lines 254–256) The relationship is described as an approximate trend in the simulated systems:

      "Notably, data from the simulated systems approximately follow a common trend, revealing a strong correlation between the magnitude of conformational change and the thermal distance to the phase transition point (R<sup>2</sup> = 0.942, Figure 3D)."

      (2) (page 7, lines 256–258) The interpretation is limited to an association with thermal distance from the critical point:

      "This result suggests that the conformational response to phase separation is closely associated with how far the system resides thermally from the critical point."

      (3) (page 7, lines 263–264) The critical-condition derivation is shown explicitly:

      "At the critical point, χ(T<sub>c</sub>) = χ<sub>c</sub>, which gives ε<sub>eff</sub> = k<sub>B</sub>T<sub>c</sub>χ<sub>c</sub>. Substituting this relation into the expression for χ(T<sub>sim</sub>) yields χ(T<sub>sim</sub>) = χ<sub>c</sub>T<sub>c</sub>/T<sub>sim</sub>."

      (4) (page 7, lines 268–272) The structural relation is explicitly introduced as a first-order linear-response approximation:

      "The thermodynamic driving force χ(T<sub>sim</sub>) – χ<sub>c</sub> can then be related to the structural observable ΔR<sub>g</sub>. Since ΔR<sub>g</sub> captures the structural transition from an intrachain-interaction-dominated state in the dilute phase to an interchain-interaction-dominated state in the dense phase, we assume, as a first-order approximation, that this conformational shift responds approximately linearly to the excess interaction strength, expressed as ΔR<sub>g</sub> ∝ [χ(T<sub>sim</sub>) − χ<sub>c</sub>]."

      (5) (page 8, lines 285–287) The unscaled relation is labeled as an empirical scaling approximation:

      "Although the complete relation in Equation (3) contains an additional T<sub>sim</sub> factor, the unscaled quantities remain strongly correlated over the simulated range. We therefore use T<sub>c</sub> − T<sub>sim</sub> ∝ ΔR<sub>g</sub> as an empirical scaling approximation."

      (5) The dynamics section would benefit from a statement of what is nontrivial, since a desolvation barrier is expected to slow local rearrangements.

      We thank the reviewer for the comment and suggestions. As described above, we have revised the dynamics section to clarify what is nontrivial beyond the expected slowing of local rearrangements by an added barrier. The revised text emphasizes that desolvation affects condensate dynamics through competing effects of macroscopic packing and microscopic energy-landscape roughness, and that it also regulates transient kinetic arrest and domain-scale coarsening. More detailed changes are provided in our response to the third point raised in Reviewer #2's Public Review.

    1. eLife Assessment

      This study presents a valuable perspective on platelet-mediated fibrin compaction, proposing that fibrin fibers undergo "winding" or coiling, a concept with potential relevance for thrombosis and clot mechanics. While the revised manuscript has improved substantially, with clearer presentation, appropriately softened language, and high-quality experimental data, direct evidence for a causal link between cytoskeletal swirling and fibrin winding/compaction is still incomplete; in particular, the actomyosin dependence and rotational fiber movements are consistent with the proposed model but do not exclude alternative mechanisms. The winding/swirl mechanism should therefore be viewed as a hypothesis supported by the observations rather than a mechanism directly demonstrated by the experiments. With this qualification, the evidence is solid and support the main conclusions.

    2. Reviewer #1 (Public review):

      This paper reports a previously unrecognized mechanism by which platelets compact fibrin fibers during clot retraction. Rather than simply pulling on fibers, the authors propose that platelets generate swirling motions that wind and loop fibrin into dense structures.

      While the results are intriguing, the underlying physical mechanism remains unexplained. In particular, it is unclear how platelets generate swirling motion capable of inducing fibrin coiling, especially when suspended in 3d fibrin mesh. This raises concerns about the conclusions. Also, does fibrin have inherent chirality or structural asymmetry that could promote coiling independently of platelet activity? Furthermore, platelet retraction typically involves platelet aggregation rather than isolated cells, and it is unclear how fibrin coiling would proceed in clustered platelets.

      Comments on revised version.

      The authors have significantly improved the manuscript and enhanced the presentation of the results. In my opinion, the physical mechanism responsible for the compaction of fibers into the coiled structures caging platelets remains somewhat elusive. Nevertheless, I find the results convincing, and I believe the study will make a valuable contribution to the field.

    3. Reviewer #2 (Public review):

      Summary:

      Grichine et al. investigate platelet-mediated fibrin compaction using human donor platelets and propose a novel mechanistic model in which platelets generate contractile forces and wind fibrin fibres into compact, coiled structures. Using a combination of 2D spreading assays, 3D clot imaging via expansion microscopy, live-cell imaging, and computational modelling, the authors present evidence of cage-like fibrin architectures, coiled fibre morphologies, and platelet-centred "rosette" structures that are present during fibre compaction. They suggest the involvement of actomyosin in fibre compaction and, overall, the study addresses an important and longstanding question in thrombosis and haemostasis while offering a conceptually novel perspective on clot compaction.

      Strengths:

      The integration of multiple imaging modalities is a notable strength. In particular, the 2D fibre-retraction assay provides a useful model for understanding the spatiotemporal dynamics of platelet-mediated fibrin compaction, which could be applied to other systems and may yield detailed mechanistic insights into biological processes. The live-imaging approaches are particularly well executed and provide valuable dynamic insights into fibre accumulation and compaction.

      Weaknesses:

      The primary weakness of the paper is the absence of direct evidence demonstrating the mechanism of fibre compaction via cytoskeletal swirling. Consequently, the relationship between platelet dynamics and fibrin organisation, including coordinated measurements of platelet motion and fibre rearrangement, is not directly assessed (perhaps due to technical barriers). However, the paper does provide solid evidence through myosin inhibition and computational modelling, demonstrating how platelets might mediate fibre compaction.

      Comments on revised version.

      Overall, the study addresses an important question in thrombosis and haemostasis and introduces a potentially impactful conceptual framework for understanding clot compaction. The imaging approaches and datasets presented will be valuable to the community, particularly to researchers interested in platelet mechanics and fibrin organisation. The possibility that fibres can be compacted extracellularly through cytoskeletal swirling represents a compelling and relatively unexplored mechanism. Therefore, this paper does a good job of establishing a thought-provoking mechanism with solid supporting evidence, although a direct demonstration of the underlying molecular mechanism requires further investigation.

    4. Reviewer #3 (Public review):

      Summary:

      This work aims to understand the mechanisms that platelets use to interact with and compact fibrin fibers during clot formation. This is an important process during wound healing and recent work has demonstrated that platelets play a critical role in generating the force required to drive accumulation of fibrin. The authors argue that current models are insufficient to account for the observed reduction in clot volume and propose that platelets actively 'wind-up' these fibers by undergoing myosin-dependent rotation. While interesting, the experiments performed by the authors do not directly test this mechanism and further evidence is required to support their claims.

      Weaknesses:

      (1) The motivation to switch from the system used in Figure 1 and 2 to the '2D fiber-retraction assay' is not clear. While the authors state that this system has 'reduced complexity' the differences between these assays appears to disrupt the 'cage-like' organization of fibrin around platelets shown in Figure 1 and 2 (compare images in Figure 2 with those in Figure 4). An in-depth comparison of two methods is needed to support the conclusions from the 2D system. Furthermore, the change in plasma volume (Figure 2 vs Figure 7) should also be tested - the authors state that this increases fibrin fiber formation, but this is not quantified or demonstrated in the figures. Notably, this appears to change the morphology of the fibrin fibers shown (comparing Figure 2 and Figure 7).

      (2) It is unclear how the classification of platelets as 'fiber-winding' versus 'fiber compaction' differs in Figure 2. The criteria used for these classifications should be stated. Further, it seems premature to characterize fibers as wound without having established this earlier in the manuscript.

      (3) Is the 'gearwheel' different from the 'cage' of fibrin fibers? They appear similar, but it is difficult to distinguish between these with only qualitative descriptions of these phenotypes.

      (4) The quantification of platelet extensions in Figure 9 is confusing. While the those in 9A are clear, those in 9B are not. For instance, what is the difference between #7 and #8 in the middle panel of 9B? It does not seem like #8 is labeling an extension.

      (5) It is unclear what the modeling accomplishes as there is no comparison between the results of these simulations and their experiments.

      (6) The data presented in Figure 12 provides the most direct support for their mechanism, but falls short of directly testing their claims. These experiments should be repeated to include blebbistatin to test the contribution of myosin and include quantitative rather than qualitative comparisons of these experiments.

      Comments on revised version:

      The manuscript is substantially improved in clarity and organization. The authors have adequately addressed most of my concerns regarding presentation and interpretation through revisions to the text and figures. However, my primary mechanistic concerns remain unresolved. Although the proposed model is now presented more cautiously, the revised manuscript still does not directly test the central mechanism, and the conclusions therefore remain insufficiently supported.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the editor and the reviewers for their time and efforts to evaluate our manuscript. We have taken into account all the comments and revised the manuscript accordingly which has considerably strengthened the message.

      Several parts of the manuscript, have been extensively rewritten to add explanations and clarify our hypothesis and claims, this also led us to add four new references.

      In addition, we have added the following new figures:

      - New part of figure 2. Figure 2E shows a platelet in a constrained clot after 4h of retraction with the fibrin cage around the platelet center still present. The actin staining of the platelet shows that radial actin fibers are present in each bulb extending to the platelet center. This observation supports our hypothesis that in each bulb an individual cytoskeletal swirling could take place resulting in the accumulation of fibrin fibers at the base of each bulb.

      - Modification of figure 3, to include the criteria used to define four categories of platelets and associated fibers in the 2D fiber retraction assay (new Fig. 3C).

      - New figure 13, illustrating the quantification of fibrin fibre compaction mediated by platelets in the 2D fiber retraction assay and the rotational movement of a fiber mass (video 9).

      - New supplementary figure 1, showing the result of a new model simulation in the absence of cytoskeleton swirling. Under this condition the fibrin fiber does not loop around the platelet bulb.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper reports a previously unrecognized mechanism by which platelets compact fibrin fibers during clot retraction. Rather than simply pulling on fibers, the authors propose that platelets generate swirling motions that wind and loop fibrin into dense structures.

      While the results are intriguing, the underlying physical mechanism remains unexplained. In particular, it is unclear how platelets generate swirling motion capable of inducing fibrin coiling, especially when suspended in 3d fibrin mesh. This raises concerns about the conclusions.

      The reviewer is right, it is difficult to imagine how platelets in a 3D fibrin mesh can accumulate fibers at the base of their extensions to form a cage-like fiber organisation around the center of the platelets. We therefore developed the 2D fibre-retraction assay, which we believe provides important insight for the coiled fiber accumulations above spread platelets in the 2D situation but also provides a framework for interpreting similar processes that may occur within a 3D clot. In response, we have placed greater emphasis on clarifying and strengthening the comparison between the potential mechanistic aspects in the 2D and 3D assays, in order to better support our proposed model (see Results, section: "Platelets, spread on a 2D surface, organize fibers above them", last paragraph). In addition, the Ideas and Speculations section of the discussion has been extensively rewritten to provide more detailed explanations about the potential mechanism leading to fibrin fiber accumulations around platelet bulbs in a 3D fibrin mesh.

      Also, does fibrin have inherent chirality or structural asymmetry that could promote coiling independently of platelet activity?

      Yes, double-stranded fibrin protofibrils have a helical twist [1]. Furthermore, a clot formed in the absence of platelets and other cellular components shows intrinsic tensile forces [2]. However, we show that inhibition of actomyosin actions prevents fibrin fiber accumulation in the 2D fibre-retraction assay providing evidence that platelet actions are necessary to observe the coiled fibers above spread platelets. This has been accentuated in the revised version and three references have been added.

      Furthermore, platelet retraction typically involves platelet aggregation rather than isolated cells, and it is unclear how fibrin coiling would proceed in clustered platelets.

      Under the in vitro fiber retraction conditions used in our study (constrained or unconstrained clots or even in the 2D assay) individual platelets are homogenously distributed within the forming clot or on the coverslip. Therefore, there are no big platelet aggregates or clusters of platelets under our experimental conditions and the results can only demonstrate how individual platelets act on fibrin fibers. This point has been emphasized in the revised version (Discussion, third paragraph).

      Reviewer #2 (Public review):

      Summary:

      Grichine et al. investigate platelet-mediated fibrin compaction using human donor platelets and propose a novel mechanistic model in which platelets generate contractile forces and wind fibrin fibers into compact coiled structures. Using a combination of 2D spread assays, 3D clot imaging via expansion microscopy, live-cell imaging, and computational modelling, the authors present evidence of cage-like fibrin architectures, coiled-fibre morphologies, and platelet centred "rosette" structures present during fibre compaction. They further suggest that actomyosin-driven cytoskeletal dynamics, potentially involving rotational or swirling motion, underlie this proposed winding mechanism, analogous to DNA looping and compaction. The study addresses an important and longstanding question in thrombosis and hemostasis and offers a conceptually novel perspective on clot compaction.

      Strengths:

      The integration of multiple imaging modalities is a notable strength of this paper. In particular, the 2D fiber-retraction assay provides a useful model for understanding the spatio-temporal dynamics of platelet-mediated fibrin compaction, which can be applied to other systems and may yield detailed mechanistic insights into biological processes. The live-imaging approaches are particularly well executed and offer valuable dynamic insight.

      Weaknesses:

      The primary weakness of this paper lies in its descriptive nature and its reliance on correlative rather than causal evidence. Several interpretations are not uniquely supported by the data presented. For example, the categorisation of fibrin accumulation in 2D assays as "fiber winding" and "fibre compaction" remains descriptive without establishing winding as a mechanism.

      When introducing the 2D fiber-retraction assay (figure 3) in the revised version, we now only mention the terms fiber accumulation and compaction to better align with the level of evidence, since wound-up fibers cannot be distinguished in this figure. The criteria to establish the four categories of platelets and associated fibers in the 2D fiber retraction assay have now been included in figure 3C.

      Nevertheless, coiled fibers above spread platelets are clearly visible in figure 4 and 8 and dynamic fiber rotations or winding-up are observed in figure 12 and video 9. These observations have been presented more cautiously, as indicative rather than definitive evidence of a winding mechanism.

      Alternative mechanisms, such as circular bundling, stacked fibers under tension, or fibrin crosslinking-induced aggregation, are neither excluded nor investigated.

      For fibrin fiber bundling, staggered or crosslinked protofilaments no platelet actions are necessary as described previously [2,3]. Since we observed a clear difference between +/- blebbistatin conditions in the 2D fiber-retraction assay, the fiber compaction we observe depends on platelet actions. Consequently, we consider these alternative mechanisms unlikely based on our data. This has been stated explicitly in the results section and discussion and three references have been added.

      Although the authors present compelling live imaging, establishing winding as a dynamic phenotype would require quantitative analyses, such as measuring angular velocities and coiling rates.

      We have incorporated quantitative measurements (new figure 13) about platelet mediated fibrin fiber compaction and angular rotation velocities to complement the observations obtained from live imaging. It is important to note, however, that angular velocities and coiling rates are likely influenced by the number of fiber–fiber contacts present at the time coiling occurs. Specifically, an increased number of contacts is expected to elevate tension within the network, thereby modulating the forces generated by platelets and, consequently, affecting both velocity and coiling dynamics.

      The use of a second fluorophore-labelled fibrin population could further strengthen evidence for rotational dynamics.

      These live videos are quite difficult to acquire because of the following reasons:

      - Small platelet size

      - Heterogeneity of platelets within the population (10 d half-life, old platelets may not be able to compact fibers efficiently).

      - The speed of the process and the time needed to adjust parameters for image acquisition, necessitates an arbitrary choice of the acquisition window and only one acquisition (90 min) per sample preparation is possible.

      - Furthermore, the laser-induced illumination can perturb the observed processes. We therefore use high-spatial-resolution 3D confocal time-lapse imaging, performed in photon-counting mode with very low laser excitation.

      For these reasons, the use of additional markers would be technically challenging and could perturb the delicate equilibrium and dynamics of the process under investigation.

      Similarly, the inference of rotational contractility or actomyosin "swirling", based on chiral actin organisation and blebbistatin treatment, is not sufficiently supported to conclude that platelets actively wind or loop fibrin fibers.

      Importantly, in the 2D fiber-retraction assay, we do not propose that the rotational actomyosin activity leads to a contractility of the platelets which would allow fiber retraction. Rather, we suggest that cytoskeletal actomyosin swirling (as demonstrated for nucleated cells by Bershadsky's team) can induce rotational dragging of extracellular bound fibrin fibers around the pseudonucleus of spread platelets thereby promoting accumulation of fibrin fibers (shown in figure 12C, video 9, third panel). Consistent with this interpretation, inhibition of myosin by blebbistatin prevents the accumulation of fibrin fibers above spread platelets in the 2D fibre retraction assay (Fig. 3).

      The mathematical model, while complementary and well-constructed, relies on multiple assumptions and lacks predictive validation.

      We thank the reviewer for this insightful comment and acknowledge that the proposed model relies on several important assumptions. In our view, the most significant assumption is that integrin molecules undergo rotational downstream motion as a consequence of their coupling to the swirling cytoskeleton. To assess the necessity and impact of this assumption, we provide an additional simulation performed in absence of the cytoskeletal swirling. Under this condition the fibrin fibers are not looped around the platelet bulb (this result has been added as supplementary figure 1). This analyses also provides further validation of the proposed model and underlying mechanism. At the same time, it is important to emphasize that the primary purpose of the model was to examine whether the hypothetical swirling dynamics of the cytoskeleton, together with the associated receptors, could in principle reproduce the experimentally observed fibrin organization.

      Appraisal:

      While the authors successfully document intriguing fibrin architectures and provide a compelling descriptive framework, they do not fully demonstrate a mechanistic model of active fibrin winding by platelets. The conclusions regarding platelet-driven winding and rotational dynamics are not sufficiently supported by direct or quantitative evidence. To substantiate these claims, the study would benefit from experiments that directly link platelet dynamics to fibrin organisation, including coordinated measurements of platelet motion and fibre rearrangement. As it stands, the results are suggestive but do not definitively support the proposed mechanism.

      Discussion and Impact:

      Despite these limitations, the study addresses an important question in thrombosis and hemostasis and introduces a potentially impactful conceptual framework for understanding clot compaction. The imaging approaches and datasets presented will be valuable to the community, particularly for researchers interested in platelet mechanics and fibrin organisation. However, the overall impact will depend on whether the proposed mechanism can be more rigorously validated. In its current form, the study presents an interesting and thought-provoking model, but would benefit from either stronger experimental support for the proposed mechanisms or a more cautious interpretation of the findings.

      We agree that the proposed mechanism requires further validation. In the revised version we have added a new result (figure 2E) showing that radial actin filaments are present in each bulb of a platelet in a constrained clot, supporting the possibility that rotational cytoskeletal movements could take place in individual bulbs. In a new figure 13, we have also quantified fiber compaction and the angular velocity of a rotating fibrin mass observed in video 9. Furthermore, in the revised manuscript, we present a more cautious and explicitly hypothesis-driven interpretation of the mechanism. We hope that the publication of our observations will be of interest to researchers in the field of thrombosis and clot mechanics who possess the specialized tools and expertise necessary to rigorously evaluate and either substantiate or refute the proposed mechanistic model.

      Reviewer #3 (Public review):

      Summary:

      This work aims to understand the mechanisms that platelets use to interact with and compact fibrin fibers during clot formation. This is an important process during wound healing, and recent work has demonstrated that platelets play a critical role in generating the force required to drive the accumulation of fibrin. The authors argue that current models are insufficient to account for the observed reduction in clot volume and propose that platelets actively 'wind up' these fibers by undergoing myosin-dependent rotation. While interesting, the experiments performed by the authors do not directly test this mechanism, and further evidence is required to support their claims.

      We do not "propose that platelets actively 'wind up' these fibers by undergoing myosin-independent rotation" of the whole platelet, but rather of the cytoskeleton winding-up extracellular fibrin fibers attached to integrin receptors.

      Weaknesses:

      (1) The motivation to switch from the system used in Figures 1 and 2 to the '2D fiber-retraction assay' is not clear. While the authors state that this system has 'reduced complexity', the differences between these assays appear to disrupt the 'cage-like' organization of fibrin around platelets shown in Figures 1 and 2 (compare images in Figure 2 with those in Figure 4). An indepth comparison of two methods is needed to support the conclusions from the 2D system.

      We agree that the cage-like fibrin organization around platelets is disrupted in the 2D fibre-retraction assay when platelets are completely spread on the coverslip before they have encountered fibrin fibers (Fig. 4). This has been explicitly stated in the revised version. However, some platelets in the 2D fiber-retraction assay form the same number of extensions as platelets in a 3D clot (Fig. 9 A, B) and are not completely spread on the glass surface. For these platelets a cage-like fibrin organisation is retained under the 2D conditions (Fig. 5 and 6). Nevertheless, the fiber density at the base of the bulbs is higher in the 2D assay than under the constrained 3D clot retraction conditions (Fig. 1C and Fig. 2), probably because in the 2D condition the fibers are less constrained and readily available for compaction.

      Furthermore, the change in plasma volume (Figure 2 vs Figure 7) should also be tested - the authors state that this increases fibrin fiber formation, but this is not quantified or demonstrated in the figures. Notably, this appears to change the morphology of the fibrin fibers shown (comparing Figure 2 and Figure 7).

      We thank the reviewer for raising this point. We would like to clarify that Figure 2 and Figure 7 correspond to two distinct experimental setups: the constrained clot retraction assay (Figure 2) and the 2D fiber-retraction assay (Figure 7). As such, they are not directly comparable. We understand, however, that the reviewer is likely referring to the apparent differences between Figures 3–6 (lower plasma volume, higher fiber density) and Figures 7–8 (higher plasma volume, lower apparent fiber density).

      The reduced number of visible fibers in the latter condition is not solely a consequence of plasma volume per se, but rather results from the formation of a labile fibrin gel at higher plasma concentrations, which is lost during the fixation and aspiration steps. This effect was initially observed across samples from two donors with differing plasma fibrinogen levels. In one case, an unusually low fibrinogen concentration allowed the addition of higher plasma volumes without inducing gel formation. In contrast, in the other sample, a more typical fibrinogen level resulted in gel formation under the same conditions.

      Importantly, we performed all experiments using matched donor plasma and platelets. As a result, the precise fibrinogen concentration could not be determined prior to experimentation. Nonetheless, post hoc measurements confirmed that fibrinogen levels in most donor samples fell within the normal physiological range, which allowed us to always use the same plasma volumes for low and high plasma concentrations (4ul/ml PBS and 7 ul/ml PBS, respectively) except for one donor as mentioned above.

      (2) It is unclear how the classification of platelets as 'fiber-winding' versus 'fiber compaction' differs in Figure 2. The criteria used for these classifications should be stated. Further, it seems premature to characterize fibers as wound without having established this earlier in the manuscript.

      The reviewer probably refers to figure 3 and he is right; it is premature to mention fiber winding at this stage of the results section (see our response to reviewer #2). In the revised version, we have modified figure 3 to include the criteria used to classify the platelets into four different categories (Fig. 3C).

      (3) Is the 'gearwheel' different from the 'cage' of fibrin fibers? They appear similar, but it is difficult to distinguish between them with only qualitative descriptions of these phenotypes.

      The "gearwheel" is observed for completely spread platelets in the 2D fiber-retraction assay and a figure illustrating our hypothetical speculations to compare the 2D gearwheel with the 3D clot situation is presented in the discussion under the "Ideas and Speculations" paragraph (now Fig. 14). We have given a more comprehensive explanation of the proposed mechanism in the revised version.

      (4) The quantification of platelet extensions in Figure 9 is confusing. While those in 9A are clear, those in 9B are not. For instance, what is the difference between #7 and #8 in the middle panel of 9B? It does not seem like #8 is labeling an extension.

      For the platelet shown in the middle panel of Figure 9B, the extensions cannot be clearly distinguished in the MIP (Maximum Intensity Projection) image because extension #8 is positioned above extension #7 and is therefore superimposed in the projection. However, the two extensions can be differentiated when examining the 3D image stack (Video 4, upper panel). As indicated in the figure legend, the number of extensions was determined manually by scrolling through the z-stack image sequence. In the revised version, we will also define the abbreviation “MIP” as Maximum Intensity Projection.

      (5) It is unclear what the modeling accomplishes, as there is no comparison between the results of these simulations and their experiments.

      We thank the reviewer for this valuable concern. We chose not to combine the experimental fibrin organization and the modeling results within the same figure panel, as the resulting image would be too complex and difficult to interpret. We have, however, added a supplementary figure 3 showing the results of a new simulation in the absence of cytoskeletal swirling. Under these conditions no winding of the fibrin fiber around the platelet bulb can be observed. It is also important to emphasize that the comparison between the model and the experimental data was intended to be primarily qualitative rather than quantitative.

      (6) The data presented in Figure 12 provides the most direct support for their mechanism, but falls short of directly testing their claims. These experiments should be repeated to include blebbistatin to test the contribution of myosin and include quantitative rather than qualitative comparisons of these experiments.

      As mentioned already above, these live videos are quite tricky to acquire because of the following reasons: - small platelet size

      - Heterogeneity of platelets within the population (10 d half-life, old platelets may not be able to compact fibers efficiently).

      - The speed of the process and the time required to optimize imaging parameters, necessitate the selection of an arbitrary acquisition window. Consequently, only a single acquisition of approximately 90 min can be performed per sample preparation, with no guarantee that relevant platelet-fibrin interactions can be acquired in the acquisition window.

      - Furthermore, after blood donation, the first sample is usually ready to be acquired around 3 pm, acquisition time 90 min. At least 10 successful acquisitions per condition would be required to ensure statistical robustness, but maximal 4 can be acquired per donor, because platelet samples start to deteriorate within twelve hours after blood donation.

      Taken together, the intrinsic heterogeneity of the platelet population, the low likelihood of capturing informative events, and the limited availability of suitable imaging resources at our institute render a robust and quantitative comparison between conditions with and without blebbistatin extremely challenging, if not impractical, within a reasonable timeframe.

      In accordance with the reviewer's request, we have added a new figure 13 to the revised version, presenting quantitative data on the platelet-mediated fibre compactions and the speed of angular fibrin rotations observed in video 9.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Throughout the manuscript, it is difficult to map the data presented in the figures to the text in the results section. Often, many subpanels are referred to collectively (for example, 'Fig 4 AE and animation, Video 3' on line 150), and the reader is left to piece together how this data fits into the statements in the results section. More guidance from the authors would help to understand the connection between these data and their conclusions.

      In the revised version, we have provided clearer explanations to make it easier to understand the conclusions drawn from the data. Concerning the indication "Fig 4 A-E and animation, Video 3" just means that platelets shown in panels A-E of figure 4 can also be visualized in the animation video 3. We have also put an effort to clearly indicate which figure part is presented in the associated video.

      There are also many figures that contain redundant information. The authors should consider revising these figures and including some of these repeated images as supplemental figures.

      As noted by the reviewers, our study provides predominantly qualitative observations essentially because it is not obvious to choose parameters which would be pertinent and could be quantified accurately using expansion microscopy. A quantitative analysis would allow to show the quantification and a representative image to describe the phenotypes of platelet-mediated fibre organisations. Without a quantitative analysis, we consider it more appropriate to provide multiple examples, enabling the reader to assess the consistency as well as the variability across repeated observations.

      Additional References

      (1) Jansen KA, Zhmurov A, Vos BE, et al. Molecular packing structure of fibrin fibers resolved by X-ray scattering and molecular modeling. Soft Matter. 2020;16(35):8272-8283.

      (2) Spiewak R, Gosselin A, Merinov D, et al. Biomechanical origins of inherent tension in fibrin networks. J Mech Behav Biomed Mater. 2022;133:105328.

      (3) Ramanujam RK, Lavi Y, Poole LG, Bassani JL, Tutwiler V. Understanding blood clot mechanical stability: the role of factor XIIIa-mediated fibrin crosslinking in rupture resistance. Res Pract Thromb Haemost. 2025;9(4):102871.

      (4) Gaertner F, Ahmad Z, Rosenberger G, et al. Migrating Platelets Are Mechano-scavengers that Collect and Bundle Bacteria. Cell. 2017;171(6):1368-1382 e1323.

    1. eLife Assessment

      This valuable study examines how the rodent prelimbic cortex represents learned and generalized threat over time and identifies distinct stable and dynamic neuronal populations that contribute to these representations. The evidence is convincing, supported by longitudinal calcium imaging, appropriate control groups, and sophisticated analyses showing that the relevant neural signals cannot be explained simply by freezing behavior. The work provides a conceptual framework for understanding how stable threat-related representations supporting memory generalization and discrimination can be maintained despite ongoing changes in neuronal ensemble composition.

    2. Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.

      To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Comments on revised version.

      The authors have addressed my previous concerns well, and the revised manuscript is substantially improved. In particular, the additional analyses strengthen the conclusion that prelimbic cortical activity reflects learned threat value rather than simply freezing behavior, while the revised framing and additional controls clarify the interpretation of the longitudinal neural dynamics. This paper represents an important contribution to our understanding of the neural mechanisms supporting aversive learning, memory, and generalization.

    3. Reviewer #2 (Public review):

      The authors have substantially revised the paper in response to the original review, which is greatly appreciated. It is clear that it will eventually make a nice contribution to the literature. This being said, the following points are somewhere between major and minor in term of their implications for interpretation of the study results. If they were to be addressed, the paper would be again improved.

      There are a few remnants of the past language that are not helpful re interpretation of the study results: 1) "Specifically, the observed population gradients could emerge either from the pooled activity of frequency-selective neurons that respond to individual tones or from neuronal subpopulations that integrate information across tones to encode their learned threat-value."; and 2) "Together, these findings suggest that the PL integrates sensory similarity with learned threat value to generate stable representations that support adaptive generalization and discrimination." Neither of these statement follows what has been shown in the study, even with inclusion of the results from the GLM analysis (see point 4 below).

      (1) This paragraph in the Discussion is difficult to follow: "Generalization has traditionally been explained by perceptual similarity (Shepard, 1987), whereby stimuli resembling a conditioned cue recruit overlapping sensory representations and evoke similar behavioral responses (Corches et al., 2019; Grosso et al., 2018). Although perceptual similarity clearly influences the extent of generalization, accumulating evidence indicates that it cannot fully account for generalized responding (Verra et al., 2026). More recent frameworks propose that associative learning assigns learned value to novel stimuli by integrating their sensory similarity with previous experience, allowing behavior to scale according to predicted biological significance (Verra et al., 2026; Zaman et al., 2023). Our findings provide a neural framework consistent with these ideas. Sensory similarity promoted consistent neuronal population responses across tones, whereas associative learning organized these responses into graded representations that tracked learned threat value across the stimulus continuum. Thus, sensory similarity appears to define the neuronal substrate upon which associative learning constructs value-based representations that support graded behavioral generalization."

      While the revisions have removed the many unnecessary references to inference and integration, this paragraph seems like it is adhering to the original idea of how the authors wished to present their work. If the authors wished to talk about something more than perceptual similarity in the context of generalization, they should have used a task that lends itself to a more-than-perceptual-similarity explanation. Again, the inclusion of the GLM analysis is suggestive for some of what the authors wish to say, but doesn't justify the statements that: "Sensory similarity promoted consistent neuronal population responses across tones, whereas associative learning organized these responses into graded representations that tracked learned threat value across the stimulus continuum." In short, the analysis does not substitute for the design that could have and should have been used to assess learned threat value independently of sensory similarity.

      (2) The next paragraph in the Discussion is also confusing. "Such reorganization has been proposed to provide flexibility by allowing new information to be incorporated into existing cortical representations while preserving stable behavioral performance (Mau et al., 2020; Zaki & Cai, 2024). Several mechanisms could contribute to this turnover, including systems consolidation, retrieval-induced reconsolidation or memory updating, and repeated nonreinforced stimulus exposure (Lacagnina et al., 2019; Mau et al., 2020; Sangha, 2015; Zaki & Cai, 2024). Although our experiments cannot distinguish between the first two possibilities, the behavioral data argue against extinction as the primary explanation. Extinction is generally associated with the formation of new CS+-safety associations (Bouton et al., 2021), whereas discrimination ratios increased across retrieval sessions, indicating that animals progressively improved their discrimination between threat-associated and safe stimuli rather than acquiring generalized safety responses. This pattern is consistent with previous work showing that discrimination learning sharpens stimulus representations and narrows behavioral generalization gradients (Dunsmoor & LaBar, 2013; Herzog et al., 2021; Jenkins & Harrison, 1960; Lommen et al., 2017). Importantly, turnover was not uniform across the population. Graded neurons retained remarkably consistent response profiles across retrieval sessions, and their activity remained more strongly associated with learned threat value than with freezing behavior. These observations indicate that stable components of the population code can coexist with extensive reorganization of surrounding neuronal ensembles."

      The issue with repeated testing is *not* caused by extinction per se. The issue is that non-reinforcement across the repeated testing should differentially affect the CS+ and CS-. Specifically, it should extinguish responding to the CS- stimulus at a rate that matches its distance from the CS+, thereby sharpening the CS+ versus CS- discrimination in precisely the ways that have been observed. Ergo, the repeated testing *is* a problem for inferences that might be drawn about the way that generalization gradients change with time; and *is* a problem for statements regarding "dynamic reorganization of cortical activity patterns over time." There is nothing in the study that allows one to comment on the reorganization of cortical activity patterns over time. The reorganization can and should be attributed to the repeated testing, which is confounded with time. Nonetheless, the reorganization must be due to the repeated testing and NOT time as the present findings are inconsistent with the well-documented broadening of generalization gradients with time.

      (3) In the next paragraph, the authors state: "At the same time, narrower generalization gradients and improved discrimination across retrieval sessions suggests ongoing memory updating. These observations are consistent with contemporary theories proposing that systems consolidation and retrieval-dependent updating are complementary processes through which memories continue to evolve after learning (Mau et al., 2020; Tome et al., 2024; Zaki & Cai, 2024)."

      In general, I'm not sure why one would invoke systems consolidation or retrieval-induced reconsolidation as an explanation for any of the present findings: they are not explanations of much at all. In this specific text, the authors seem to be implying an updating process that occurs independently of what is learned across the repeated sessions of testing. Why? The changes that occur in the behaviour and neuronal representations are perfectly explicable in terms of additional learning that occurs - of the sort that I hope to have made clear in my previous comment. Why invoke more than what is needed to explain the observed pattern of results?

      (4) Re the GLM analysis - The authors write that: "the fact that the GLM analysis indicates that these neurons reflect learned threat value more than freezing behavior, suggests that they encode an abstract property of the learned stimulus rather than simply mirroring behavioral output."

      This is fine if freezing fully indexes the state of conditioned fear and there are no other behaviours in which animals express their fear. If, however, fear is expressed in a range of other behaviours that are likely coordinated by the PL (e.g., startle, vigilance, scanning, orienting to source of danger), this interpretation of the GLM analysis is unwarranted. This is an important point and would be worth noting somewhere in the paragraph where the statement appears.

    4. Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under examined in the field.

      Comments on revised version.

      The authors have convincingly and thoroughly addressed my concerns. I have no further issues regarding this study.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public review:

      Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure. To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to neural changes across sessions. Repeated retrieval can induce memory updating (reconsolidation) or extinction, the latter involving the formation of a new association between the CS+ and safety. Although memory updating may have contributed to the ensemble reorganization observed here, several aspects of our findings argue against extinction as the primary explanation for the observed neural changes.

      First, we observed substantial neuronal ensemble turnover beginning with the first retrieval session. This early turnover is consistent with previous observations in the prefrontal cortex [1, 2] and with growing evidence that cortical memory representations remain dynamic throughout systems consolidation [3, 4]. Longitudinal studies have shown that neurons are continuously recruited into and removed from cortical memory ensembles while memory expression remains stable [1-4].

      Second, we calculated discrimination ratios to quantify discrimination of each tone relative to the CS+ across retrieval sessions (Figure S1). These analyses showed that discrimination increased, rather than decreased, over successive retrieval sessions, a pattern inconsistent with the behavioral profile expected if repeated testing had induced extinction.

      Finally, one of the most novel findings of our study is that ensemble turnover does not affect all neuronal populations equally. The graded neurons identified by our clustering analysis maintained their identity and functional organization across retrieval sessions, and their activity was better explained by tone threat value than by freezing behavior (Figure 8). This selective stability indicates that ensemble reorganization is not a uniform process but instead preferentially affects specific neuronal subpopulations while preserving a stable threat-value generalization gradient. Thus, although repeated retrieval may contribute to ongoing ensemble reorganization, our results demonstrate that this process is selective and largely spares the neuronal subpopulations that encode graded threat-value representations.

      Accordingly, we have revised the Discussion to explicitly acknowledge these points as follows:

      “The ensemble turnover observed here is consistent with previous studies demonstrating dynamic reorganization of cortical activity patterns over time [1-3, 5]. Such reorganization has been proposed to provide flexibility by allowing new information to be incorporated into existing cortical representations while preserving stable behavioral performance [4, 6]. Several mechanisms could contribute to this turnover, including systems consolidation, retrieval-induced reconsolidation or memory updating, and repeated nonreinforced stimulus exposure [4, 6-8]. Although our experiments cannot distinguish between the first two possibilities, the behavioral data argue against extinction as the primary explanation. Extinction is generally associated with the formation of new CS+-safety associations [9], whereas discrimination ratios increased across retrieval sessions, indicating that animals progressively improved their discrimination between threat-associated and safe stimuli rather than acquiring generalized safety responses. This pattern is consistent with previous work showing that discrimination learning sharpens stimulus representations and narrows behavioral generalization gradients [10-13]. Importantly, turnover was not uniform across the population. Graded neurons retained remarkably consistent response profiles across retrieval sessions, and their activity remained more strongly associated with learned threat value than with freezing behavior. These observations indicate that stable components of the population code can coexist with extensive reorganization of surrounding neuronal ensembles.” Pg. 19

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      This is an important point, which was also highlighted by the other reviewers. To directly address this concern, we implemented the generalized linear model (GLM) analysis suggested by Reviewer 3. We modeled the neuronal activity time series using both tone identity and freezing behavior as simultaneous predictors. Because tone identity was fixed across trials whereas freezing varied from trial to trial, the GLM allowed us to dissociate their independent contributions to neuronal activity.

      As described in the original submission, freezing was estimated from the miniscope's onboard inertial measurement unit (IMU), which measures body acceleration along three axes. Rather than classifying freezing using a fixed threshold, we estimated the continuous probability of freezing from the accelerometer signal using a Gaussian mixture model. This probabilistic estimate was incorporated directly into the GLM together with tone identity, providing a conservative test of whether neuronal activity was better explained by freezing behavior or by the auditory stimulus.

      We applied the GLM both to all sound-responsive neurons contributing to the population response curves (Figure 4) and to the graded and frequency-selective neuronal subpopulations identified by our clustering analysis (Figure 8). Across both experimental groups and all analyses, the median regression coefficients (β) associated with tone identity were consistently larger than those associated with freezing, indicating that tone identity contributed more strongly to neuronal activity. Moreover, tone coefficients exhibited graded monotonic profiles that closely tracked the learned threat value of each tone, with graded neurons showing the strongest gradients (Figures 4a, 8a, and 8e). Consistent with previous reports [14, 15] freezing accounted for a modest but significant component of PL activity. However, only 6–8% of graded neurons were classified as freezing-dominant, indicating that for the vast majority of these neurons, tone identity was the stronger predictor. Together, these findings demonstrate that the graded representation of learned threat value persists after accounting for freezing behavior, supporting our conclusion that PL activity reflects learned threat value rather than merely the behavioral expression of fear.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      We thank the reviewer for this suggestion. We now include measures of registration quality in the resubmission. Specifically, we calculated shifts in centroid distances, proportion of ROIs retained across all sessions, and representative examples of matched imaging fields over time (Fig, S3).

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      We corrected correspondence between text and Figure.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      We clarified the labelling of the Figure 2a and call the graphs “activity-plots”.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      For Figure 3, we maintained the previous one-way ANOVAs assessing changes in AUC per day to be able to note significance on the Figure panels. However, we added a three-way mixed-effects analysis, using group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30) as variables, with frequency and day of testing as repeated measures. The results were described as follows (statistical details Table S1):

      “To determine how AUC varied across groups over time, we performed a three-way mixed-effects ANOVA with group (CS+15, CS+3, and no shock), frequency (3, 7, 11, and 15 kHz), and time (test days 1, 15, and 30) as factors, with repeated measures on frequency and time. For positive responder neurons, the analysis revealed significant main effects of group (p < 0.001) and time (p < 0.05), as well as a significant group × frequency interaction (p < 0.001), whereas the time × frequency and group × time × frequency interactions were not significant (p > 0.05; Table S2a). Tukey-corrected post hoc comparisons showed that, in the CS+15 group, AUC differed between all frequency pairs except 11 and 15 kHz (p < 0.05). In the CS+3 group, the AUC at 3 kHz differed from those at 7, 11, and 15 kHz (p < 0.05), whereas no significant frequency differences were observed in the no-shock controls (p > 0.05). For negative responder neurons, the only significant effect was a time × frequency interaction (p < 0.01). Tukey-corrected simple-effects analyses revealed that, on day 30, the AUC at 15 kHz differed from those at 3, 7, and 11 kHz (p < 0.05; Table S2b). Because this pattern was observed across all experimental groups, including the no-shock controls, it is unlikely to reflect associative learning. These results indicate that although the AUC exhibited modest changes over time, these changes were not group-specific and therefore do not support learning-dependent alterations in neuronal responses. Together, these results show that despite substantial neuronal turnover, PL population responses encode generalization gradients, closely matching behavioral expression.” Pg. 9

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term inference may overstate the cognitive processes engaged by the current task. Accordingly, we revised the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli. The new GLM analyses further support this interpretation by demonstrating that, in the conditioned groups, neuronal activity at both the population and single-neuron levels is explained substantially better by tone identity than by freezing behavior (Figures 4 and 8). We therefore retained the term threat value, as our results indicate that PL activity primarily reflects learned threat value rather than simply the expression of freezing behavior, but removed inference.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

      We corrected the language and replaced valence for “threat value”

      Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      We thank the reviewer for this thoughtful comment. We agree that our original wording may have implied that turnover was a spontaneous process. Repeated retrieval provides opportunities for updating the learned contingencies associated with both the conditioned and generalization stimuli, and therefore changes in ensemble composition across sessions need not arise independently of experience. We also agree that the stability of graded neuronal representations may be related to the animal's certainty about the learned contingencies. However, in our data the graded neuronal population remained remarkably stable across retrieval sessions, whereas changes occurred primarily within the dynamic, frequency-selective neuronal populations. This suggests that stable ensembles preserve representations of learned threat value while updating is concentrated in a distinct neuronal subpopulation. We have now incorporated these ideas into the Discussion.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      We thank the reviewer for these specific criticisms. We revised the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.

      (a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      (a) We hypothesized that the PL generates representations of learned threat value that support threat generalization and discrimination, and that these representations emerge from the coordinated activity of stable and dynamic neuronal subnetworks, preserving consistent relationships among stimuli despite ongoing cellular turnover.

      (b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      (b) The summary statement was rewritten: " Together, these findings provide a neural framework for understanding how the PL supports adaptive threat generalization and discrimination.” pg. 4

      (c) 'In CS<sup>+</sup> 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS<sup>+</sup>15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      We adopted the reviewer's suggested rewording: " In CS<sup>+</sup> 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting learned contingency value across testing days" pg. 9

      We will systematically review the entire manuscript to ensure consistency with this revised framing.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we followed Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using the time series activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis allowed us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we concluded that tone identity was a stronger predictor of neuronal activity than freezing. These results are summarized in Figures 4 for all cells contributing to population responses and Figure 8 for the main neuron types identified in the clustering analysis (frequency-selective and graded neurons). All details of this extensive new analysis are shown in red in the revised resubmission.

      In addition, we revised the text to remove the terminology of “learned and inferred valence” throughout the manuscript.

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      We thank the reviewer for this helpful comment. We agree that our use of the term valence in describing the no-shock controls was imprecise. Because none of the tones was associated with reinforcement in this group, there was no learned valence that could modulate neuronal activity. Our intention was simply to convey that, although both positive and negative sound-responsive neurons were present, the population responses did not vary systematically across tone frequencies. We have revised this section accordingly.

      We also agree that our original wording overstated the interpretation of the graded population responses. Our data do not demonstrate that associative learning is required for sound responsiveness itself; rather, they show that associative learning is required for the emergence of graded population responses that distinguish tones according to their learned threat value. We have revised the text to make this distinction explicit.

      Finally, we agree that our previous references to "learning and inference" were not justified by the behavioral paradigm. We have removed this language throughout the manuscript and now describe the findings more directly as graded representations of learned threat value that closely parallel the observed behavioral generalization gradients.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      We thank the reviewer for this thoughtful comment. We agree that repeated retrieval is an inherent limitation of longitudinal memory studies and that repeated non-reinforced presentations of the tones provide opportunities for memory updating. Accordingly, we have revised the Discussion to explicitly acknowledge that repeated retrieval may contribute to the ensemble reorganization observed across sessions through memory updating or reconsolidation processes (Discussion, pg. 19).

      We also agree that the progressive sharpening of the behavioral generalization gradients across retrieval sessions is consistent with memory updating. Both the behavioral data (increased discrimination ratios) and the neuronal data (progressively sharper population generalization gradients among neurons active after conditioning) indicate that the memory representation became more precise over time. We now discuss this possibility explicitly in the revised Discussion. We also agree that fear generalization often broadens with time; however, this is not universal. Under discriminative conditioning paradigms, repeated retrieval can instead produce progressively narrower generalization gradients [11]. We have revised the Discussion to clarify this distinction and added the appropriate references (pg. 19).

      While the reviewer's interpretation is therefore plausible, we do not believe it fully accounts for our observations. If repeated non-reinforced presentations were the sole driver of the observed neuronal changes, one might expect a more uniform reorganization across the neuronal populations engaged by the task. Instead, the reorganization was highly selective. Neurons encoding graded threat value remained remarkably stable across retrieval sessions, whereas neuronal turnover occurred primarily within the frequency-selective subpopulations. Thus, although repeated retrieval may update the memory representation, the neuronal substrate supporting graded threat-value coding is largely preserved while refinement occurs within a distinct neuronal subpopulation.

      Moreover, we observed substantial neuronal turnover beginning with the first retrieval session, consistent with previous longitudinal studies showing that cortical memory ensembles remain dynamic despite stable memory [1-4]. This early emergence of turnover suggests that repeated testing alone is unlikely to account for the continuous population dynamics observed throughout the experiment.

      Rather than viewing these findings as evidence exclusively for either memory updating or systems consolidation, we believe they are more consistent with current models proposing that these processes occur in parallel. Several influential frameworks argue that memories are continuously modified through retrieval while simultaneously undergoing systems-level reorganization [4, 6, 16, 17].We have therefore revised the Discussion to interpret the longitudinal changes more conservatively as reflecting the combined influence of retrieval-dependent memory updating and systems-level reorganization.

      In summary, we have revised the manuscript to better acknowledge the contribution of repeated retrieval while emphasizing what we believe is the principal finding of our study: despite substantial turnover within the overall ensemble, the neuronal population encoding graded threat value remained remarkably stable, whereas refinement occurred primarily within dynamic frequency-selective neuronal populations.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS<sup>+</sup>. In CS<sup>+</sup>15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      We agree that, because population similarity is highest for the 3/3, 15/15, and 15/11 tone pairs and freezing is also greatest for these same stimuli, the neural data could, in principle, reflect a correlate of behavioral expression rather than an independent representation of learned threat value.

      To directly address this possibility, we implemented a generalized linear model (GLM) to dissociate the contributions of tone identity and freezing behavior to neuronal activity. Across all analyses, tone identity consistently explained substantially more variance in neuronal activity than freezing behavior. Importantly, this finding held not only for the full population of sound-responsive neurons used to generate the population similarity analyses (Figure 4), but also for both the stable graded neurons and the dynamic tone-selective neuronal populations identified by our clustering analysis (Figure 8). Thus, although freezing behavior contributes modestly to PL activity, it cannot account for the enhanced similarity of population vectors across stimulus presentations or the graded population responses that form the basis of our conclusions.

      In addition, the temporal dynamics of the population vector similarity analysis are not entirely consistent with the interpretation that PL activity simply reflects the expression of freezing behavior. Population vector similarity peaked during the first 5 seconds following tone onset, whereas freezing occurred intermittently throughout the tone presentations. Although this temporal relationship does not establish causality, it is consistent with the interpretation that PL activity reflects the learned threat value associated with each tone rather than merely tracking the magnitude of freezing.

      Finally, we have revised the manuscript to more clearly acknowledge the correlational nature of these analyses. Specifically, we now state that population vector similarity is associated with, rather than determines, the degree of threat generalization.

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      We made this correction. (“These findings indicate that population-level similarity at stimulus onset scales with behavioral threat generalization”. pg. 13)

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      We agree that the phrase "integrated learned valence" is unnecessarily opaque and we replaced it with more precise language “Our previous analyses demonstrated that threat-value generalization gradients are represented at the population level. However, these findings do not reveal how these representations arise. Specifically, the observed population gradients could emerge either from the pooled activity of frequency-selective neurons that respond to individual tones or from neuronal subnetworks that integrate information across tones to encode their learned threat-value.” (Pg. 13)

      (8) Another example of what has been a common theme in this review:

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      We thank the reviewer for pointing out that this section was unclear. We agree that our original wording was imprecise and could be interpreted as implying cognitive processes that were not directly tested in the present study. Accordingly, we have revised the terminology throughout the manuscript. We no longer refer to "inferred emotional valence" or "core memory content" and instead describe these neurons more specifically as exhibiting graded representations of learned threat value.

      This is not the interpretation we intended. To determine whether these neurons primarily reflected defensive behavior rather than learned stimulus value, we implemented a Generalized Linear Model (GLM) that dissociates the contributions of tone identity and freezing behavior to neuronal activity. Across the entire neuronal population, as well as within the stable graded and dynamic tone-selective neuronal subpopulations, tone identity consistently explained substantially more variance than freezing behavior (Figures 4 and 8). Furthermore, after accounting for freezing, the regression coefficients of the graded neurons continued to follow the learned threat value of the tones, exhibiting opposite monotonic gradients in the CS+15 and CS+3 groups. If these neurons simply tracked defensive behavior irrespective of the stimulus presented, this relationship would not be expected to persist after accounting for freezing. We therefore conclude that the activity of this stable neuronal subpopulation is better explained by graded representations of learned threat value than by defensive behavior alone, and we have revised the manuscript accordingly.

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      We thank the reviewer for highlighting that this section was unclear. We agree that the original phrasing was insufficiently precise. Our intention was to convey that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. This observation led us to hypothesize that neurons recruited into the active population after conditioning— likely dynamic, frequency-selective neurons—also contribute to these graded population responses through modulation of their firing rates.

      The reviewer correctly notes that the phrase "shaped by associative processes" was too vague. By this we meant that the firing properties of these neurons are modified by the animal's associative history, including both the original conditioning experience and any retrieval-dependent updating that may occur during subsequent test sessions. We have revised the manuscript to make this interpretation explicit rather than leaving it to the reader to infer.

      To test this hypothesis, the GLM analysis we implemented dissociated the contributions of tone identity and freezing behavior to neuronal activity. After accounting for freezing, tone identity (i.e., learned threat value) remained a significant predictor of neuronal responses. Importantly, this was also true for the dynamic, frequency-selective neurons (Fig. 8e–f), indicating that these neurons contribute to population-level representations of learned threat value through firing-rate modulation rather than simply reflecting defensive behavior.

      To clarify our interpretation, we have rewritten the relevant section as follows:

      "Graded clusters encode generalization gradients but constitute only a subset of the active neuronal population. Nevertheless, population-level representations, which incorporate all active neurons, remain robust and accurately preserve these gradients. This observation led us to hypothesize that neurons recruited over time (e.g., dynamic, frequency-selective cells) also contribute to threat-value representations. Consistent with findings in the hippocampus showing that neurons can encode task contingencies through firing-rate modulation despite responding selectively to a single location (Gagliardi et al., 2024; Huxter et al., 2003; Sanders et al., 2019), we tested whether dynamic, frequency-selective clusters exhibited firing-rate differences proportional to learned threat value." (page 15)

      Regarding the reviewer's suggestion that the characteristics of the newly recruited neurons may reflect learning during repeated non-reinforced test sessions, we agree that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population, and we now explicitly acknowledge this possibility in the Discussion. However, we do not believe that our findings are fully explained by repeated non-reinforced retrieval alone. First, no-shock control animals underwent the same repeated testing but failed to develop graded neuronal representations, indicating that repeated exposure in the absence of associative learning is insufficient to account for the observed changes. Second, both behavioral discrimination and the corresponding population-level neural gradients became progressively sharper over time, consistent with refinement of learned threat representations rather than an effect of repeated testing alone, which must lead to extinction.

      In summary, we thank the reviewer for highlighting both the ambiguity of our original wording and an important alternative interpretation. In response, we have clarified the text to explicitly define what we mean by associative processes, added a GLM analysis demonstrating that the newly recruited neurons encode learned threat value beyond freezing behavior, and revised the Discussion to acknowledge that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population.

      (10) The following points all relate to the Discussion and reiterate many of the points above.

      (a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      We modified the language as stated in the prior points.

      (b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      We incorporated the new GLM analysis to address this point and conclusions.

      (c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      We agree that the term "graded representational axis" was insufficiently defined and could be interpreted in multiple ways. Because this terminology was not essential to our conclusions, we have removed it and instead describe the observed phenomenon as a graded population representation of learned threat value across tone frequencies. We also removed the statement suggesting that recurrent connectivity provides a stable scaffold for these representations, as this mechanistic interpretation is not directly supported by our data.

      We also agree that some sections of the manuscript overstated the scope of our conclusions and have revised the wording accordingly. Our study uses neuronal activity recorded during memory retrieval after learning, an approach widely used in studies of systems consolidation to infer how learned information is represented within neural populations. Accordingly, we have revised the manuscript to explicitly state that our findings pertain to neural representations of learned threat value during memory retrieval rather than the broader content of emotional memories.

      Finally, we agree that our data are correlational and do not establish the causal role of the neuronal representations we identify. Throughout the manuscript, we now refer more precisely to population- and single-neuron correlates of learned threat value during memory retrieval following auditory fear conditioning.

      (d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

      We thank the reviewer for this thoughtful comment. Our interpretation that these neurons reflect preexisting sensory-driven properties of PL cortex is based on two observations. First, tone-selective neuronal clusters were present in both conditioned and no-shock control animals, consistent with previous reports of sensory responsiveness in PL cortex [18, 19]. Second, these responses were already present during the first retrieval session, when the intermediate frequencies were presented for the first time. Thus, they cannot be explained by repeated exposure to those tones across subsequent test sessions.

      We therefore interpret the frequency-selective response properties as pre-existing features of PL circuitry that are present independently of conditioning. In contrast, associative learning modifies the firing activity of these neurons, allowing them to contribute to graded representations of learned threat value. This interpretation is supported by our GLM analysis, which showed that, after accounting for freezing, tone identity significantly predicted the activity of frequency-selective neurons in conditioned animals but not in no-shock controls. Thus, while the frequency-selective response properties are present independently of learning, associative learning modifies how these neurons encode learned threat value. We have revised the manuscript to clarify this distinction.

      Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded neuronal responses may partially reflect freezing-related activity rather than representations of learned threat value. In the revised manuscript, we acknowledge that previous studies have identified PL neurons whose activity tracks freezing independently of stimulus identity or associative content. To directly address this possibility, we implemented the reviewer's suggestion by fitting a generalized linear model (GLM) to the neuronal activity time series derived from the Ca<sup>2+</sup> signals, using tone identity and freezing behavior as predictors. Because tone identity is fixed across trials, whereas freezing varies both during tone presentation and across trials (see below our answer to the Recommendations to Authors), this approach allowed us to dissociate their respective contributions to neuronal activity. We are grateful for this excellent suggestion, which has substantially strengthened both the manuscript and the conclusions that can be drawn from our data. The new analyses are summarized in Figures 4 and 8.

      In the points below we summarize the new findings.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

      We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such as combinations of optogenetic and two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.

      We examined inter-individual variability in freezing relative to the proportion of graded cells but the number of mice used in this study (CS+3= 5 and CS+15=7) did not give us enough power to reach significance.

      Therefore, we modified the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the correlational nature of our conclusions.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Many sections of the paper should be rewritten along the lines that I have suggested in my public review and below.

      Minor Comments:

      (1) INTRO - 'This broad accessibility reduces spatial specificity and increases learning variability...'.

      Broad accessibility of what, exactly? And how does the 'broad accessibility' reduce spatial specificity and increase learning variability? That is, I do not understand what the terms 'reduced spatial specificity' and 'increased learning variability' refer to at this point in the first paragraph...

      We rewrote the introduction and discussion to address the points raised by the reviewer.

      (1) INTRO - 'The prelimbic cortex (PL) contributes to the expression (Burgos-Robles et al., 2009; SierraMercado et al., 2011; Sotres-Bayon & Quirk, 2010) and the proper discrimination and generalization of threat memories (Rosas-Vidal et al., 2025; Stujenske et al., 2022).'

      What is achieved by calling it 'proper' discrimination and generalization? Can't one simply say that the PL contributes to the expression, discrimination, and generalization of threat memories?

      This was corrected.

      (3) INTRO - '... and that population-level firing rate similarity across stimulus presentations determines threat generalization'.

      Or, alternatively, that generalization of conditioned freezing responses from the tone CS to variants along the dimension of Hz values is reflected in systematic changes in the firing rate of PL neuronal ensembles; when the test stimulus is similar to the conditioned stimulus, the two elicit similar behavioural responses and evoke a similar population-level firing rate in the respective PL neuronal ensembles.

      We have revised the Introduction as stated above. However, as discussed in our detailed responses, freezing behavior cannot fully account for the observed patterns of PL activity.

      (4) METHODS - 'Memory retrieval was tested on days 1, 15, and 30 after conditioning to probe early, long-term, and remote memory (Bontempi et al., 1996). During retrieval, mice were tested in a novel context with the CS<sup>+</sup>, CS1<sup>-</sup>, and two intermediate frequencies (7 and 11 kHz), presented in semirandom order, with each tone repeated three times (Fig. 1a).'

      Why was testing conducted in a different context than that of conditioning? This is likely to result in an underestimation of generalization to the different tones...

      In tone fear conditioning, it is always customary to test in a different context to dissociate conditioning to the context vs conditioning to the tones, which usually take place simultaneously in the same context [20]. Therefore, testing generalization in a novel context gives the correct estimate of generalization to the tones in the absence of contextual conditioning confounds. Please note that while overall freezing levels may be lower in a novel context due to the absence of contextual conditioning, the relative generalization gradient across tones — which is what your study measures — is unlikely to be systematically distorted by context change.

      (5) RESULTS - 'No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that freezing reflected associative learning.'

      The inference doesn't follow from the result described. Was there more freezing among animals in the shocked groups compared to those in the no-shock group? I presume so - my point is that this comparison is the one that most directly speaks to the presence or absence of associative learning.

      Experimental animals exhibited not only higher overall freezing but also graded freezing responses across tone frequencies. It is important to note that no-shock controls did not display this pattern, ruling out the possibility that the different frequencies themselves elicited graded behavioral responses. To clarify this point, we revised the sentence as follows: "No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that the graded freezing patterns resulted from associative learning rather than the acoustic properties of the tones." (Pg. 6)

      (6) 'Across animals and sessions, we identified distinct neuronal populations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound (Fig. 2b)...'

      To be clear, do you mean to say that there were distinct neuronal populations that consistently [i.e., across all three sessions] increased their responses to the tones [positive modulation], decreased their responses to the tones [negative modulation], showed variable responses to the tones [mixed responses], and did not respond to tones [not modulated]?

      The sentence refers to neuronal populations identified within each recording session based on their responses to the tones, not to neurons that maintained the same response profile across all three sessions. We have revised the text to make this distinction explicit. The only stable patterns across sessions were observed in graded neurons that were stable across retrieval.

      “Across animals, we identified distinct neuronal subpopulations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound in each session (Fig. 2b)” Pg. 7

      (7) What does 'active' mean in relation to Figure 2? Does this refer to neurons that displayed either positive responses, negative responses, and/or mixed responses? In the text, it is stated that 'Sound responder neurons were classified using a test that detected modulation based on magnitude relative to baseline variability, allowing reliable identification of both transient and sustained responses while remaining robust to noise...'

      I can't work out if this is the same classification criteria used for the determination of positive modulation, negative modulation, and mixed responding.

      We thank the reviewer for pointing out this ambiguity. In Figure 2, the term "active" referred to neurons that exhibited significant sound-evoked modulation and were subsequently classified as showing positive, negative, or mixed responses. Thus, active and sound-responsive refer to the same population of neurons. We removed the word active to avoid confusion.

      The reference to transient and sustained responses describes the temporal profile of the calcium signals rather than separate response categories. Some neurons exhibited brief calcium transients that rose and decayed rapidly, whereas others displayed sustained activity throughout the tone presentation. The sound-response detection algorithm was designed to reliably identify both temporal response profiles. We have revised the manuscript to make these definitions explicit. We modified the sentence as follows: “Sound-responsive neurons were identified using a statistical test that detected activity modulation relative to baseline variability, allowing reliable identification of responses while remaining robust to noise. This approach was effective for neurons exhibiting either brief calcium transients that rose and decayed rapidly or sustained activity throughout the tone presentation.” Pg. 7-8

      (8) 'A moderate proportion of neurons was present across all retrieval sessions, with no differences between groups (p > 0.05).'

      Do you mean to say that 'A moderate proportion of neurons was ACTIVE across all retrieval sessions, with no differences between groups (p > 0.05)'?

      We replaced the word present and replaced it with “active”. Pg. 8

      (9) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'If neurons encoding graded responses carry core mnemonic information, they should exhibit enhanced stability over time. To test this hypothesis, we quantified the proportion of registered neurons that retained their cluster identity across at least two retrieval sessions and compared these values to a shuffled null distribution (10,000 iterations), with multiple comparisons controlled using the BenjaminiHochberg procedure.'

      What does the comparison to the shuffled null distribution tell us exactly? I accept that some neurons were stable positive responders across at least two sessions. The comparison to the shuffled null distribution creates a false impression about the robustness of this stability or the 'enhanced stability over time'.

      Our intention in comparing the observed stability to a shuffled null distribution was to evaluate whether the proportion of neurons retaining cluster identity exceeded chance levels expected from random assignment. The shuffled distribution therefore provides a statistical baseline against which the observed degree of stability can be evaluated. We agree, however, that the wording “enhanced stability over time” may be confusing regarding this finding. We rephrased this paragraph to clarify that a subset of neurons retained cluster identity across all retrieval sessions at levels greater than expected by chance as follows:

      “These data demonstrate that graded clusters remain consistently active at levels exceeding chance, preserving their cellular identity and providing a stable representation of learned contingencies and generalization gradients.” Pg. 15.

      (10) ABSTRACT. The abstract states that, 'Stimulus-evoked population similarity scaled precisely with behavioral generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing, indicating that population-level similarity predicts inferred threat.'

      I believe that the sentence could be rewritten as, 'Stimulus-evoked population similarity reflected the degree of generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing.'

      We revised the text according the reviewer’s suggestion; however, we had to shorten the sentence due to word limits. “Population similarity tracked behavioral generalization, whereas consistent population states emerged only for shock-associated or highly generalized tones.” Pg. 2

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The ensembles with graded activation in proportion to stimulus valence are described at various points in the manuscript as "maintaining the emotional 'gist'", "preserving core components of the memory trace", and "preserving core components of the memory trace". This conclusion is premature because there is an alternative interpretation. The graded response ensembles would also be consistent with coding for the freezing behavior itself, irrespective of the specific memory or stimulus association that drives it. An ensemble that encodes a behavior in this way would not be considered mnemonic, just as motor neurons in the spinal cord are not, even if they may fire during a conditioned response. Indeed, previous work has identified neurons in the prelimbic cortex that encode freezing independently from the stimuli that signal an aversive outcome (e.g., Kyriazi, Headley, and Pare 2020; Casanova, Pouget, ..., Vetere 2024).

      There are two ways the authors can address this point.

      (a) Fit a generalized linear model to the time series of inferred spiking activity from the Ca2+ signal and include stimuli and freezing as predictors. Since freezing behavior is inconsistent across trials, while stimulus presence is fixed, they can be disassociated. If, after accounting for freezing, responsiveness neurons still show a graded coding of stimuli that agrees with inferred aversiveness, this would strengthen their claim that they have identified an ensemble that corresponds with mnemonic or salience aspects of the stimuli.

      (b) Conduct no further analysis but cover the issue in the discussion as a limitation to their study and to dampen some of the language throughout the manuscript that implies that a memory trace has been identified.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an important alternative interpretation is that graded-response ensembles could reflect freezing-related activity rather than representations of learned threat value. To directly address this possibility, we implemented a Generalized Linear Model (GLM) analysis, as suggested by the reviewer. The GLM was fitted to the activity of every sound-responsive neuron included in the population analyses and simultaneously incorporated tone identity and continuous freezing probability (derived probabilistically from miniscope acceleration) as predictors, allowing us to quantify their independent contributions to neuronal activity.

      We want to note that freezing was quantified from the miniscope's inertial measurement unit (IMU) using a two-component Gaussian mixture model applied to the log-transformed body-acceleration signal. Rather than classifying freezing with a binary threshold, we used the posterior probability of the low-movement state as a continuous freezing regressor. This approach captures graded variations in immobility and provides a more conservative test of tone encoding, because it accounts for more behaviour-related variance than a binary classifier, making it more difficult to detect an independent contribution of tone identity.

      We applied the GLM both to all sound-responsive neurons contributing to the population response curves and separately to the identified frequency-selective and graded neuronal subpopulations. Across all analyses, tone identity consistently explained neuronal activity better than freezing. Furthermore, the freezing-corrected tone β coefficients scaled with learned threat value, with graded neurons exhibiting the strongest monotonic gradients, indicating that they provide the most robust representation of learned threat value. These findings demonstrate that the graded coding of learned threat value persists after accounting for freezing behavior and therefore cannot be explained simply by the behavioral expression of fear. The new analyses are presented in Figures 4 and 8. Notably, although freezing-dominant neurons were present in both the tone-selective and graded populations, they represented only a small fraction of each group and were least prevalent among graded neurons (6–8%), further supporting the conclusion that graded neurons primarily encode learned threat value.

      In addition, we revised the manuscript to avoid language implying that these neuronal populations constitute a mnemonic trace. Instead, we consistently describe them as encoding learned threat value, a more accurate interpretation that is directly supported by the new GLM analyses.

      (2) The title makes a seemingly causal claim by using the term 'arise', "Learned and inferred valence arise from interactions between stable and dynamic subnetworks". While it is true that the authors show that both stable and dynamic ensembles encode valence, they do not demonstrate that the behavioral expression of valence depends on these codes, nor their interaction. Experimentally testing this is beyond the scope of this study (holographic two-photon stimulation of transient and stable ensembles?), but they may be able to get closer to it by examining inter-individual variability. The authors could measure the proportion of neurons in each subject that participate in the stable (graded responding) and dynamic (stimulus-specific) ensembles, and see if they predict individual differences in the expression of freezing behavior or its generalization. Indeed, this correlation may change across testing days.

      We agree that the term “arise” in the title may imply a stronger causal relationship than is directly supported by the present data. We modified the title in the resubmission as follows: “Complementary stable and dynamic prelimbic ensembles encode learned threat value underlying generalization and discrimination”

      The new LGM analysis confirms that a large proportion of neural activity can be predicted by tone threat value; therefore, we think this title fully captures our findings.

      We also appreciate the reviewer's suggestion to examine inter-individual variability. In the revised manuscript, we tested whether the proportion of graded neurons correlated with freezing behavior. However, the limited number of experimental animals in each experimental group provided insufficient statistical power to reliably assess this relationship. Accordingly, we revised the manuscript to clarify that our conclusions are based on correlational observations rather than causal inferences.

      In summary, we revised the title and related language throughout the manuscript to avoid implying causal mechanisms beyond the scope of the current experiments.

      Minor points:

      (1) I was surprised by the absence of an ensemble in the No-shock group that responded uniformly to all stimuli. Can the authors confirm this?

      Yes, we confirm this finding. It was unexpected to us as well. We would like to clarify, however, that some control neurons may have responded to more than one frequency, but these responses were too infrequent or too weak to be classified as a distinct graded neuronal population by our clustering algorithm. Thus, while broadly responsive neurons may have been present in the control group, they did not form a robust, identifiable ensemble comparable to that observed after fear conditioning.

      (2) Several different approaches were used to analyze the same Ca2+ responses to stimuli across testing days. These were the "Sound responder classification", "Average stimulus-aligned trace procedure", "Population similarity over time across tone pairs", and the construction of "Stimulus response vectors". These feature differing alignment/binning/interpolation, normalization, and response quantification procedures, and it is unclear why they cannot all be in agreement, at least when it comes to alignment and normalization.

      We thank the reviewer for this careful reading of our Methods. All analyses were performed on the same underlying calcium imaging dataset, but they were designed to address different aspects of the data and therefore required different preprocessing steps. The analyses share a common initial pipeline leading to the calcium traces (all z-scored across the session). Differences in subsequent processing (e.g., use of ΔF/F versus CASCADE-deconvolved activity, normalization, baseline correction, temporal binning, and interpolation) were introduced only when required by the specific analysis method.

      To make this clearer, we have substantially revised the Methods. We added a new overview of preprocessing section that summarizes the common preprocessing pipeline and explicitly distinguishes the shared steps from those that are analysis-specific. We also included a summary table describing the input signal (ΔF/F or CASCADE-deconvolved activity), normalization procedure, and temporal processing used for each analysis. Finally, the individual Methods sections were revised to eliminate redundancies and more clearly describe the steps to avoid confusion. We hope these revisions make the rationale for the different preprocessing procedures and the overall analytical workflow more transparent (Pg. 22-23)

      (3) In the methods section "Window-wise response quantification" the Ca2+ signal was baselinesubtracted and divided by the standard deviation in the baseline across trials, but that data was already presumably z-normalized to the baseline of each trial ("Data alignment and normalization"). This second step of normalization seems excessive. Why is it not sufficient to just take the average peri-stimulus response across the z-normalized trials from the "Data alignment and normalization" section? This is simpler and would capture the effect size of the response relative to baseline.

      We thank the reviewer for this careful observation.

      The two operations are also not the same normalization applied twice; they standardize different sources of variability. During the alignment step, each trial is z-scored relative to its own baseline by dividing by the standard deviation of that trial's baseline across time. This places all trials on a common within-trial scale before averaging. In the window-wise step, the trial-averaged, baseline-subtracted response is expressed relative to the standard deviation of the per-trial baseline levels across trials—a distinct quantity that reflects trial-to-trial baseline stability rather than within-trial fluctuations. The purpose of this second term was to down-weight windows in cells with unstable baselines across trials, and it entered the analysis only as a significance criterion; the magnitude threshold defining a sound responder was applied to the trial-averaged baseline-relative response itself. We have revised the Methods to clarify the distinct roles of these two normalization steps.

      To further address this concern, we re-ran the sound-responder classification after removing the second (between-trial) normalization step, so that responder detection depended only on the per-trial baselinerelative response magnitude and its temporal persistence. Across all cells, tones, and sessions (n = 89,504 cell–tone–session classifications), the two procedures agreed on 95.3% of labels. The small fraction of cells whose labels changed were almost exclusively those lying immediately at the detection threshold: 86.6% of changes involved cells moving into or out of the "modulated" category, whereas direct reversals between excitatory and inhibitory classification occurred in only 8 of 89,504 cases (0.009%). Consistent with the between-trial standard deviation being a less stable quantity when few trials are available, label changes were approximately twice as frequent in the three-trial retrieval sessions (5.2%) as in the ten-trial conditioning sessions (2.1%). Overall responder proportions changed only minimally (positive responders +2.2%, negative responders +4.3%), and all population-level findings—including the graded threat-value gradient across tones, its absence in no-shock controls, and its persistence after controlling for freezing in the , as analysis—were unaffected. These analyses demonstrate that our conclusions are robust to this methodological choice.

      (4) It would increase confidence in the tracking of neurons across days if the authors showed some example images of neurons tracked across days.

      We added an example in the Supplement. Additionally, we now provide measures of registration quality (Fig. S3)

      (5) Table S2 is a bit confusing. I take it that Graded A/B were only for CS15, and Graded C/D/E were only for CS3. Also, Common B1-4 were the cells with positive responses to individual stimuli, and Common C1-4 were the cells with negative responses to stimuli. If this is the case, it should be explained in the figure legend (or even better, clusters should be named and numbered consistently in all figures.

      Thank you for pointing this out, we corrected the Table to indicate which test corresponds to which figure and cluster, specifying which ones were positive or negative modulated. Please note old Table 2 is now Table 5

      (6) The term network and subnetworks implies some connectivity between neurons, but in this study, it is used to refer to the ensembles of cells activated in a similar manner. Since connectivity is never assessed, it would be better if the authors stuck to the terms ensembles or populations.

      We changed the wording and now use ensembles or populations

      (7) The legend for Figure S3 has the text 'eded', which seems to be a typo.

      We corrected this typo.

      References

      (1) Kitamura, T., et al., Engrams and circuits crucial for systems consolidation of a memory. Science, 2017. 356(6333): p. 73–78.

      (2) DeNardo, L.A., et al., Temporal evolution of cortical ensembles promoting remote memory retrieval. Nat Neurosci, 2019. 22(3): p. 460–469.

      (3) Tome, D.F., et al., Dynamic and selective engrams emerge with memory consolidation. Nat Neurosci, 2024. 27(3): p. 561–572.

      (4) Mau, W., M.E. Hasselmo, and D.J. Cai, The brain in motion: How ensemble fluidity drives memory-updating and flexibility. Elife, 2020. 9.

      (5) Gallego, J.A., et al., Long-term stability of cortical population dynamics underlying consistent behavior. Nat Neurosci, 2020. 23(2): p. 260–270.

      (6) Zaki, Y. and D.J. Cai, Memory engram stability and flexibility. Neuropsychopharmacology, 2024. 50(1): p. 285–293.

      (7) Lacagnina, A.F., et al., Distinct hippocampal engrams control extinction and relapse of fear memory. Nat Neurosci, 2019. 22(5): p. 753–761.

      (8) Sangha, S., Plasticity of Fear and Safety Neurons of the Amygdala in Response to Fear Extinction. Front Behav Neurosci, 2015. 9: p. 354.

      (9) Bouton, M.E., S. Maren, and G.P. McNally, Behavioral and Neurobiological Mechanisms of Pavlovian and Instrumental Extinction Learning. Physiol Rev, 2021. 101(2): p. 611–681.

      (10) Jenkins, H.M. and R.H. Harrison, Effect of discrimination training on auditory generalization. J Exp Psychol, 1960. 59: p. 246–53.

      (11) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.

      (12) Herzog, K., et al., Reducing Generalization of Conditioned Fear: Beneficial Impact of Fear Relevance and Feedback in Discrimination Training. Front Psychol, 2021. 12: p. 665711.

      (13) Lommen, M.J.J., et al., Training discrimination diminishes maladaptive avoidance of innocuous stimuli in a fear conditioning paradigm. PLoS One, 2017. 12(10): p. e0184485.

      (14) Casanova, J.P., et al., Threat-dependent scaling of prelimbic dynamics to enhance fear representation. Neuron, 2024. 112(14): p. 2304–2314 e6.

      (15) Kyriazi, P., D.B. Headley, and D. Pare, Different Multidimensional Representations across the Amygdalo-Prefrontal Network during an Approach-Avoidance Task. Neuron, 2020. 107(4): p. 717–730 e5.

      (16) McKenzie, S. and H. Eichenbaum, Consolidation and reconsolidation: two lives of memories? Neuron, 2011. 71(2): p. 224–33.

      (17) Winocur, G. and M. Moscovitch, Memory transformation and systems consolidation. J Int Neuropsychol Soc, 2011. 17(5): p. 766–80.

      (18) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.

      (19) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.

      (20) Phillips, R.G. and J.E. LeDoux, Differential contribution of amygdala and hippocampus to cued and contextual fear conditioning. Behav Neurosci, 1992. 106(2): p. 274–85.

    1. eLife Assessment

      This is a valuable study on the electrophysiological and computational underpinnings of the accumulation of intermittent glimpses of sensory evidence for decisions. The authors present convincing EEG and behavioural evidence to support their claims. The work will be of interest to cognitive and systems neuroscientists working on decision-making.

    2. Reviewer #1 (Public review):

      Summary:

      This paper characterises the physiological and computational underpinnings of the accumulation of intermittent glimpses of sensory evidence, with a focus on the centroparietal positivity and motor beta lateralization. The main finding is that the centroparietal positivity builds up during evidence accumulation but falls back to baseline during gaps, while motor beta lateralization maintains a continuous a sustained representation throughout the gap and until response.

      Strengths:

      - Elegant combination of electroencephalography and computational modelling.<br /> - Innovative task design, including parametric manipulation of gap duration.<br /> - The authors describe results of two separate experiments, with very similar results, in effect providing an internal replication.

      Weaknesses:

      - In their response to the reviewers, the authors now include a figure illustrating the relationship between the centroparietal positivity and motor beta lateralisation. However, in the absence of statistical analyses, it remains difficult to draw firm conclusions about this relationship.

      - The paper does not provide an exhaustive characterisation across sensors and frequency bands. However, as the data are publicly available, these questions could be addressed in future work.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript examines decision-making in a context where the information for the decision is not continuous, but separated by a short temporal gap. The authors use a standard motion direction discrimination task over two discrete dot motion pulses (but unlike previous experiments, fill the gaps in evidence with 0-coherence random dot motion of differently coloured dots). Previous studies using this task (Kiani et al., 2013; Tohidi-Moghaddam et al., 2019; Azizi et al., 2021; 2023) or other discrete sample stimuli (Cheadle et al., 2014; Wyart et al., 2015; Golmohamadian et al., 2025) have shown decision-makers to integrate evidence from multiple samples (although with some flexible weighting on each sample). In this experiment, decision-makers tended not to use the second motion pulse for their decision. This allows the separation of neural signatures of momentary decision-evidence samples from the accumulated decision-evidence. In this context, classic electroencephalography signatures of accumulated decision-evidence (central-parietal positivity) are shown to reflect the momentary decision-evidence samples.

      Strengths:

      The authors present an excellent analysis of the data in support of their findings. In terms of proportion correct, participants show poorer performance than predicted if assuming both evidence samples were integrated perfectly. A regression analysis suggested a weaker weight on the second pulse, and in line with this, the authors show an effect of the order of pulse strength that is reversed compared to previous studies: A stronger second pulse resulted in worse performance than a stronger first pulse (this is in line with the visual condition reported in Golmohamadian et al., 2025). The authors also show smaller changes in electrophysiological signatures of decision-making (central parietal positivity, and lateralised motor beta power) in response to the second pulse. The authors describe these findings with a computational model which allows for early decision-commitment, meaning the second pulse is ignored on the majority of trials. The model-predicted electrophysiological components describe the data well. Some flexible weighting of the second pulse also described the data well (in line with previous studies), but this explanation suffers from additional model complexity. In particular, this analysis of model-predicted electrophysiology is impressive in providing simple and clear predictions for understanding the data.

      Weaknesses:

      Behaviour in this experiment is different from previous experiments which use very similar designs (Kiani et al., 2013; Tohidi-Moghaddam et al., 2019; Azizi et al., 2021; 2023). The authors provide some possible explanations for this in the discussion. Overall performance in this experiment was much worse than previous experiments: Participants achieved ~85% correct following 400 ms of 33 - 45% coherent motion. In previous work, performance was ~90% correct following 240ms of 12.8% coherent motion. A second weakness is that, while bounded model can describe the data in this manuscript, it cannot explain the data from previous experiments showing a stronger weight on the second pulse.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      We have addressed the outstanding points made by the reviewers and provide a detailed description of the additional analyses performed & key results below. We have also updated the manuscript to reflect these additional results, and to contain a more detailed consideration of alternative plausible models.

      We also note that we have corrected one figure panel (Fig.2 panel B, Exp. 2 only), where we identified a small bug in the visualisation code whereby the data of either one or two participants was not correctly plotted in some conditions. This makes no difference to the reported effects.

      Please note that the reviewers acknowledged that your introduction now more broadly refers to the previous work from various groups on motor beta lateralisation (MBL).

      (1) Evaluating the correlation between CPP and MBL, which is key for supporting the claim that CPP is feeding MBL. If, as you are alluding to in your rebuttal, single-trial estimates of CPP are too noisy, trials could be binned based on CPP.

      As requested, we now provide additional analyses binning the data by CPP amplitudes, for the high-low coherence conditions at P1. Full details are provided below. In both experiments we find that, for a given coherence, greater CPP amplitudes at P1 correlate with stronger motor beta lateralisation.

      (2) Examining the possibility of down-weighting (in line with previous studies) compared to your current bounded integration description. Specifically, does the your model predict a bi-modal CPP-P2 distribution that is not evident in the data?

      We have now fit 3 additional models investigating alternative mechanisms that might account for the behavioural results. In particular, we have explored 4 different ways in which flexible weighting of the second pulse might account for both the behavioural and neural data. A full account of the results is provided below, and has been included in the manuscript. In sum, we find that a model which directly and uniformly downweighs evidence from the second pulse (as opposed to indirectly through little or no distance remaining to bound, as in our model) can account for the behavioural data well, but it cannot recapitulate CPP-P2 results unless an accumulation-terminating bound is also included in the model, and the additional complexity of a model with these two free parameters is not supported by model comparison. An alternative model where P2 is downweighed as an inverse function of P1 strength (i.e., stronger downweighing for P1-high coherence pulses) could recapitulate both the behavioural and neural data, but again, this was not favoured by model comparison metrics that account for complexity in the current dataset. We have added a piece on this in the discussion, noting how previous studies mentioned by the reviewer such as Cheadle et al., 2014 and Glickman et al., 2022 find a consistency bias where later evidence is boosted when it agrees with the earlier evidence, opposite to the dampening suggested by the model here, but that a key distinction in our task is that P2 always agreed with P1, so that a dampening might be plausible if subjects tend to withdraw some of their attention from the confirmatory P2 based on the strength of P1.

      Regarding the CPP-P2 distribution, our original bounded model does indeed predict a bimodal CPP-P2 distribution with a peak at 0 arising from the early termination trials, which does not appear in our data (see Fig. S19 and related reply below). However, that EEG noise precludes the detection of any such bimodality in single trial amplitude distributions is demonstrated by the fact a bimodal distribution is strongly predicted for CPP-P1 amplitudes due to the two coherences, most strongly in fact for the unbounded model since there would be nothing to cap the higher-coherence, yet no trace of such bimodality is evident there either, due to EEG noise.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper characterises the physiological and computational underpinnings of the accumulation of intermittent glimpses of sensory evidence, with a focus on the centroparietal positivity and motor beta lateralization. The main finding is that the centroparietal positivity builds up during evidence accumulation but falls back to baseline during gaps, while motor beta lateralization maintains a continuous a sustained representation throughout the gap and until response.

      Strengths:

      - Elegant combination of electroencephalography and computational modelling.

      - Innovative task design, including parametric manipulation of gap duration.

      - The authors describe results of two separate experiments, with very similar results, in effect providing an internal replication.

      Weaknesses:

      - A direct characterization of how the centroparietal positivity and motor beta lateralization interact is missing, which limits the novelty. In their reply to reviewers, the authors argue that the signal-to-noise ratio of EEG signals is insufficient for such analyses at the single-trial level. If so, a binned or trial-averaged approach could still be attempted.

      As requested, we have now performed an additional analysis binning trials according to single-trial CPP-P1 amplitudes. To this aim, we sorted trials according to P1 coherence, and median-split them within condition according to the CPP-P1 amplitudes integrated in a time window around the grand-averaged peak [0.4 to 0.6s] after pulse onset, on the same subset of electrodes as in the manuscript. We then plotted motor beta lateralisation (MBL) as the difference in [Contra - Ipsi] hemispheres. Stronger negativities thus indicate stronger lateralisation towards the correct response. In all 4 cases, (both experiments and both coherence levels), higher CPP amplitudes were associated with stronger lateralisation from 0.5s post-pulse onwards (Author response image 1).

      Author response image 1.

      MBL (bottom) traces aligned to P1 onset (time = 0), median split by CPP amplitude [0.4-0.6s] post pulse onset, within P1 coherence condition. Trials with stronger CPPP1 potentials were linked to stronger MBL lateralisation toward the correct response.

      - An exhaustive characterisation of sensors and frequency bands is also missing. In their reply to reviewers, the authors suggest that this would detract from their hypothesis-driven focus. I disagree: the main hypothesis and figures could remain centred on the centroparietal positivity and motor beta lateralization, with a more comprehensive mapping of sensors and frequencies placed in supplementary material. Since the purpose of the paper is to examine EEG-based decision signals in a novel behavioural context, a broader characterisation of the underlying EEG landscape would seem appropriate.

      To broaden our characterisation, we have now included an additional supplementary figure that describes another distinct, relevant EEG signal. Fig. S12 shows the lateralised readiness potential (LRP), a lateralised motor preparation signal that has long been used as an index of relative motor preparation with high temporal resolution (Eimer, 1998; Kelly & O’Connell, 2013; Vidal et al., 2015).The LRP is typically computed as the difference in voltage between [IpsiContra] lateral motor electrodes with respect to eventual response, and it captures the fact that the contralateral motor cortex exhibits more pronounced negative ramps than the ipsilateral one immediately preceding action execution. The EEG landscape characterised in our paper thus comprises four distinct signals that are all functionally relevant to the task, including occipital alpha power, relevant for attention & temporal expectation encoding, which was included both in the main manuscript (Fig. 2) and the supplement (Figs. S9, S11). Given our already extensive supplementary material (18 figures) focused on our main research questions, we feel that a full, hypothesis-free exploration across the dimensions of frequency, space (sensors) and time, considering that there are 60 experimental conditions among which differences may be tested for (2 directions x 3 gaps x 4 coherence pairings in exp 1, plus 2 directions x 4 gaps x 4 coherence pairings in exp 2, plus single-pulse trials), would render the supplemental materials excessive in volume. Again, the data will be shared publicly for future exploration of these many dimensions.

      Reviewer #2 (Public review):

      Summary:

      This manuscript examines decision-making in a context where the information for the decision is not continuous, but separated by a short temporal gap. The authors use a standard motion direction discrimination task over two discrete dot motion pulses (but unlike previous experiments, fill the gaps in evidence with 0-coherence random dot motion of differently coloured dots). Previous studies using this task (Kiani et al., 2013; Tohidi-Moghaddam et al., 2019; Azizi et al., 2021; 2023) or other discrete sample stimuli (Cheadle et al., 2014; Wyart et al., 2015; Golmohamadian et al., 2025) have shown decision-makers to integrate evidence from multiple samples (although with some flexible weighting on each sample). In this experiment, decision-makers tended not to use the second motion pulse for their decision. This allows the separation of neural signatures of momentary decision-evidence samples from the accumulated decision-evidence. In this context, classic electroencephalography signatures of accumulated decision-evidence (central-parietal positivity) are shown to reflect the momentary decision-evidence samples.

      Strengths:

      The authors present an excellent analysis of the data in support of their findings. In terms of proportion correct, participants show poorer performance than predicted if assuming both evidence samples were integrated perfectly. A regression analysis suggested a weaker weight on the second pulse, and in line with this, the authors show an effect of the order of pulse strength that is reversed compared to previous studies: A stronger second pulse resulted in worse performance than a stronger first pulse (this is in line with the visual condition reported in Golmohamadian et al., 2025). The authors also show smaller changes in electrophysiological signatures of decision-making (central parietal positivity, and lateralised motor beta power) in response to the second pulse. The authors describe these findings with a computational model which allows for early decision-commitment, meaning the second pulse is ignored on the majority of trials. The model-predicted electrophysiological components describe the data well. In particular, this analysis of model-predicted electrophysiology is impressive in providing simple and clear predictions for understanding the data.

      Weaknesses:

      Some readers may be left questioning why behaviour in this experiment is so different from previous experiments which use almost exactly the same design (Kiani et al., 2013; TohidiMoghaddam et al., 2019; Azizi et al., 2021; 2023). Overall performance in this experiment was much worse than previous experiments: Participants achieved ~85% correct following 400 ms of 33 - 45% coherent motion. In previous work, performance was ~90% correct following 240ms of 12.8% coherent motion. A second weakness is that, while the authors present a model which describes the data based on pre-mature decision-commitment, they do not examine explanations from the existing literature, that evidence is flexibly weighted, and do not provide any analyses which could be used to compare these descriptions. While their model can describe the data in this manuscript, it cannot explain the data from previous experiments showing a stronger weight on the second pulse.

      The revised version of the manuscript includes a detailed discussion about possible reasons why our stimulus characteristics, task design & experimental protocol may have led to the observed behavioural results (lines 605 onwards). Furthermore, we have now included an extended model comparison as a supplementary note which examines alternative models that could account for the observed data and an additional discussion section that links it to the existing literature.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have responded to each of the comments in the previous review. The manuscript introduction and discussion have been substantially improved, and now more adequately address the previous literature. Limited improvements were made to the analysis, although the authors acknowledged why the suggested improvements from the reviewers were unlikely to be successful, but did not attempt to address the comments using other methods.

      One common theme to both reviews was that, although the model broadly describes the data, it is not fully tested, and alternative descriptions are not fully considered.

      We thank the reviewer for their careful consideration of our data & their reply. We agree it is important to formally test alternative descriptions, most particularly those involving downweighting of the processing of P2, and we have now done so. If participants were simply downweighting P2 by implementing generally smaller drift rates regardless of P1, we would expect the CPP-P2 to exhibit the same classical pattern as in P1, with higher amplitudes following high-coherence P2. The interaction pattern we observe, whereby CPP-P2 amplitudes are systematically lower following high-coherence P1, within each P2 coherence, can only be explained by some dependency between the processes occurring at P1, and those following at P2. What if, as the reviewer suggested originally, “additional evidence from the second pulse was down-weighted according to certainty following the first pulse?” We thus consider this also.

      To formally test this, we have now fit four further models and compare both the fit quality to accuracy data and the predicted EEG results to our original implementation. All models were fit to the grand-averaged accuracy data for all conditions, using 10K simulations, and for both experiments separately (as was done for the bounded model presented in the manuscript). For EEG simulations in unbounded models, we assumed that the CPP signal fell back down to zero upon the dots turning blue, as for the simulations in the main manuscript. Fig. S15 illustrates the fit quality as measured by means of Bayesian Information Criterion (BIC), and we go into more detail on each model in turn below.

      “(1) Unbounded model + P1-independent P2 drift rate reweighting (DR<sub>P2</sub>)

      First, we fit a model with no bound in which the drift rate for P2 was estimated by reweighting (positive or negative) the P1 drift rate via an additional free scaling parameter w. This model thus had the same complexity as our original one (k = 3 free parameters): two drift rate parameters for high and low coherence of P1 (d<sub>high</sub>, d<sub>low</sub>), plus a scaling parameter, w, dictating the strength of P2 relative to same-coherence P1. The drift rate for P2 (d<sub>P2</sub>) was computed as the d*w, where d equals d<sub>high</sub> or d<sub>low</sub> depending on P2 coherence, multiplied by the scaling parameter w, so that if w < 1, P2 drift rates would decrease compared to the same coherence in P1. This model implements the reviewers’ suggestion that systematic downweighting of the second pulse might also account for the results (while also allowing for upweighting (w > 1) for the sake of flexibility).

      This first model (DR<sub>P21</sub>) yielded a similar fit quality to the behavioural data compared to our original bounded model (Bnd), as measured by BIC (Fig. S15). The new model broadly recapitulated the key behavioural results, including order effects, (Fig. S16,A) and could also recapitulate the generally lower CPP-P2 amplitudes (Fig. S16,B). However, this model failed to capture the key coherence-based pattern observed in the CPP-P2 data. Namely, while our EEG results showed that CPP-P2 in trials following P1-low coherence pulses reached overall higher amplitudes than that in trials following P1-high coherence pulses (see manuscript Fig. 3), this model’s simulations predicted that CPP-P2 should scale only with P2 coherence, showing higher amplitudes and steeper build-up rates for P2-high trials, regardless of P1 coherence (Fig. S16C). This is at odds with our empirical results.”

      “(2) Unbounded model + P1-dependent P2 Drift rate reweighting (invDR<sub>P2</sub>)

      Next, we tested a model in which P2 downweighting could depend on P1 strength. That is, we made the P2 drift rate scaling parameter w inversely proportional to P1 coherence so that P2 drift rate d<sub>P2</sub> = d*w/d<sub>P1</sub>, where d equals d<sub>high</sub> or d<sub>low</sub> depending on P2 coherence, and d<sub>P1</sub> indicates the preceding P1 coherence. This implements a kind of certainty weighting, whereby evidence following a strong P1 is more strongly dampened than evidence following a weak P1. This model could recapitulate the key behavioural findings (Fig. S17A), and also qualitatively captured the CPP-P2 effects (i.e. P1-low trials reaching overall higher amplitudes than P1-high trials, Fig. S17C), although the magnitude of this effect was substantially smaller than predicted by the original simple bounded model with no drift rate modulations. However, BICs indicated that this model provided an overall worse fit to the accuracy data compared to the original bounded model (Fig. S15).”

      (3) Bounded models + P2 drift rate reweighting (Bnd + DR<sub>P2</sub>, Bnd+ invDR<sub>P2</sub>)

      Finally, we investigated how well models with both a bound and either of the two P2 drift rate scaling methods we investigated above (uniform downweighting, P1-dependent downweighting) could capture the data, thus effectively testing two extensions of our original implementation that allowed for flexible reweighting of P2.

      The systematic reweighting model with a bound (Bnd + DR<sub>P2</sub>) could recapitulate all key behavioural and EEG findings (Fig. S18), but the additional complexity of the model was not supported by BIC (Fig. S15). Crucially, the reason that this model could recapitulate the CPPP2 results was still the presence of a bound, although we note that the proportion of trials that were predicted to terminate early was reduced in this model compared to the original bounded model presented in the manuscript (c.f. Fig. 4). Yet, this relatively small fraction of trials where accumulation ended early meant that 1) in some trials no accumulation was allowed to occur at all during P2, and 2) where it occurred, the DV was closer to the bound following P1-high coherence trials, thus needing to accumulate less further evidence before reaching a bound and yielding smaller CPP-P2 amplitudes overall in those trials. The inclusion of a bound was thus key to allow a model with systematic P2 downweighting to account for both behavioural and EEG data.

      The model with P1-based scaling of P2 drift rates with a bound could also recapitulate all key behavioural and EEG findings (Fig. S18), but again, the additional complexity of the model was not supported by BICs (Fig. S15).

      “Conclusion

      This extended modelling exercise suggests that 1) a simple bounded model is favoured by model comparison, 2) an alternative model of equal complexity which includes a P1dependent systematic downweighting of P2 rather than a bound can produce qualitatively similar results, the common feature of both viable models being the push-pull relationship between P1 and P2, and 3) more complex models including both a bound and P2 modulations can also account for both the behavioural and EEG results, but the additional complexity is not supported by the current data. While it is possible that some flexible weight modulations occur, these are not sufficiently influential to justify its inclusion in the model. Future work would nevertheless be warranted to explore this possibility in more detail using tailored task paradigms. For the scope of this paper, we have maintained the bounded account in the main manuscript as it is the one supported by the model comparison in the current dataset, but we have also included the alternative P1-based downweighting account as a supplementary figure, along with some additional discussion.”

      In response to my previous comment 3, the authors show their model predicts that there should be no CPP-P2 if the bound is reached before P2, otherwise CPP-P2 is similar to CPPP1 (Figure R2). The argument in the manuscript is that the lower CPP-P2 is because of this bound. The distribution of CPP-P2 amplitudes should therefore have higher variance than CPP-P1 amplitudes, and one might even predict a second mode in the distribution, around 0 amplitude (those trials that terminated before P2). The authors do not show this.

      In Figure R3, it looks like the data have been normalised independently for CPP-P1 and CPPP2 (since the means are approximately the same); normalisation also prevents a comparison of the variance. However, it is apparent that there is no bimodality in the CPP-P2 distribution - were there substantially more trials with 0 CPP-P2 amplitude than CPP-P1? Is the EEG data actually more consistent with a model that systematically downweights P2?

      In the previous Figure R3, data were normalised across CPP-P1 and CPP-P2, not separately. We plot the non-normalised values here (Author response image 2), for comparison, along with median, variance and skewness values for CPP-P1 and CPP-P2. Additionally, we attach the single-participant plots at the bottom of this document (Author response image 3)

      We reanalysed the non-normalised data, excluding outliers (defined as values exceeding the mean +/- 3 times the standard deviation, computed for each pulse & for each participant separately). We found, in both experiments, lower median amplitudes (Exp. 1: t(21) = 2.68, p = 0.013; Exp. 2: t(20) = 4.37, p < 0.001), higher variances (Exp. 1: t(21) = 0.59, p = 0.55; Exp. 2: t(20) = 2.33, p = 0.03) and more positive skewness (Exp. 1: t(21) = 1.79, p = 0.08; skewness: 0.027 vs. 0.69; Exp. 2: t(20) = 3.85, p <0.001; skewness: 0.027 vs. 0.214) in CPP-P2 compared to CPP-P1, although variance and skewness effects were only significant in Exp. 2.

      The reviewer argued in the previous review as well as here that increased variance would be predicted by the model – this is correct, but we believe that the mere presence of higher CPPP2 variances in our empirical data does not, on its own, necessarily support the model. That is because this increased variance could be explained by other factors, such as the increased EEG signal complexity of data at P2 compared to P1. This increased complexity naturally arises from overlapping potentials from CPP-P1 (which can be corrected for, but will increase data noise and thus variability nonetheless), as well as the various gap durations across conditions, which would also affect pre-P2 dynamics. Thus, while our data (partially - in Exp. 2 only) support the reviewer’s interpretation, we would be cautious in using the observation of increased variability as evidence for or against our model given the considerations above.

      Author response image 2.

      A. Empirical CPP–P1 and CPP-P2 amplitude [450-550ms post-pulse] distributions, pooled across coherences. Data were not normalised within-participant. Data were baselined 100ms before pulse onset prior to CPP-P2 amplitude extraction. In both experiments, CPPP2 amplitudes had a lower median (vertical line) amplitude and higher variance than CPP-P1. B. CPP amplitude simulations based on the original bounded model. A high number of trials where CPP-P2 amplitude should equal 0 due to early terminations. C. CPP amplitude simulations based on the inverse weighting model (invDR<sub>P2</sub>). The model did not predict any trials with zero CPP-P2 amplitude because early terminations were not allowed. Rather, the mean distribution shifted towards lower predicted amplitudes, because P2 was downweighted proportionally to P1 coherence.

      The reviewer asks: “Were there substantially more trials with 0 CPP-P2 amplitude than CPPP1?”. Given the noisiness of single-trial EEG data due to high-frequency artifacts, and/or spurious signal drifts (as illustrated by the raw CPP value distributions in Author response image 2A above), it is not be possible to directly detect trials on which CPP = 0. Instead, this must be inferred by other means. If the CPP is in reality at 0 on a larger proportion of trials, then on average, this should manifest as a higher fraction of trials with lower amplitudes, resulting in a more positively skewed distribution of CPP-P2 amplitudes compared to CPP-P1. As reported above, in both experiments we find that skewness is higher in CPP-P2 than in CPP-P1, in line with this hypothesis.

      Regarding the question “Is the EEG data actually more consistent with a model that systematically downweights P2?”, we point to the additional modelling we conducted in response to the comment above. To recapitulate, we find that a model that systematically downweights P2 could account for behavioural findings, but could only recapitulate the key EEG CPP-P2 patterns if an accumulation-ending bound was also included in the model. Instead, a model where P2 is downweighted as a function of P1 strength could qualitatively capture both behavioural and EEG findings, but was not strongly supported by goodness of fit measures in this dataset. We note, however, that the latter model does not predict a bimodal CPP-P2 distribution, but rather a shifted mean and more positive skewness for CPP-P2 trials (Author response image 2C; skewness: CPP-P1 = 1.08; CPP-P2 = 1.24). Thus, in that respect, it does appear to provide a better qualitative recapitulation of the single-trial CPP-P2 data. However, given the extent of EEG noise, the bimodal underlying distribution of our bounded model would also translate to a unimodal, skewed distribution as observed, so this does not provide a strong basis for adjudication. Underscoring this, it is noteworthy that an unbounded model in fact predicts a more separated bimodal distribution for P1 than a bounded model, yet, again, with EEG noise, we are not able to identify any such bimodality in the empirical P1 amplitude distribution.

      Minor:

      In the discussion, the authors write "We also used a narrower range of coherences than the previous studies, which possibly lends itself to calibrating a bound to achieve acceptable accuracy while saving cognitive effort." (Page 22). Perhaps this should be reworded. The range of coherence in this study was ~26-44% in Exp 1 (a difference of 18%) in previous experiments the coherence was 3.2-12.8% (a difference of ~10%). The ranges of performance were similar.

      We agree that the phrasing could be improved. We meant to say that we used only two coherences (high-low) with less than a twofold difference between them, instead of multiple levels of evidence strength in previous studies (e.g. 0,3.2,6.4,12.8) – we have clarified this.

      The authors mention in their rebuttal "However, in contrast to previous studies, we did not include any feedback on a trial-by- trial basis, instead only providing feedback at the end of each block indicating the average accuracy." Actually, Kiani et al., 2013 also only gave feedback at the end of each block. This is also implied in the discussion. I suggest this be removed as the common feedback in Kiani et al., 2013 suggests this cannot explain the difference.

      The methods in Kiani et al. 2013 state “At the end of motion stimulus, a 400–1000 ms delay period (truncated exponential) was imposed before the Go signal, disappearance of the fixation point, was presented. The subject was required to report the net direction of motion within 1 s after the Go signal by pressing a left or right key. Distinctive auditory feedback was delivered for correct and error responses. On trials with 0% coherence, the type of feedback was chosen randomly.“ We understand this means feedback was provided after every trial. If this interpretation is wrong, we would like to kindly ask the reviewer to point us to the relevant methods section so that we can correct the manuscript.

      Author response image 3.

      Individual CPP-P1 (blue) and CPP-P2 (orange), for both experiments (non-z-scored). Vertical lines indicate median CPP amplitudes for each pulse, respectively.

    1. eLife Assessment

      This important study measures the development of hearing in the zebra finch and shows that two-day-old hatchlings have no detectable electrophysiological response to 95 dB SPL clicks. Because this stimulus carries energy across a broad range of sound frequencies and is some 60 dB louder than faint, high-pitched parental calls, the findings challenge prior proposals that such calls serve as a channel for heat-dependent prenatal developmental programming. Two reviewers found the evidence on the absence of auditory responses convincing, supported by auditory brainstem recordings using validated methods with appropriate controls and by a separate laser vibrometry experiment ruling out detection through egg vibration. A third reviewer disputed these findings and questioned whether absent physiological markers of hearing definitively establish deafness. On the evidence presented, embryonic hearing is unlikely to mediate the developmental effects reported for heat whistle playback. Whether those effects reflect some other pathway or require re-interrogation of the behavioral phenomenon is now the open question, alongside a need for auditory measurements more sensitive than the auditory brainstem response.

    2. Reviewer #1 (Public review):

      This work by Antonnen et al. was triggered by claims of auditory-mediated effects on altricial avian embryos which were published without any direct evidence that the relevant parental vocalizations were actually heard. I agree with Anttonen et al. that, based on the available evidence about avian auditory development, those claims are highly speculative and therefore necessitate more direct experimental verification.

      Attonen et al. have embarked on a comprehensive series of experiments to

      (1) Better characterize acoustically the relevant parental vocalizations (heat whistles; in a separate preprint, not reviewed here)

      (2) Characterize the auditory sensitivity of zebra finches at various stages of their posthatching development. Despite the long-standing importance of the zebra finch as a songbird model in neuroethology of learned vocalizations, the auditory development of the species had not been studied so far.

      (3) Explore an alternative hypothesis of how the parental vocalizations might be perceived.

      The principal method used here is the non-invasive recording of ABR (auditory brainstem response), a standard neurophysiological method in auditory research. The click-evoked ABR provides a quick and objective assessment of basic hearing sensitivity that does not require animal training. Weaknesses of the technique include its limited frequency specificity and low signal-to-noise ratio. The authors are experienced with ABR measurements and well aware of those issues. ABR responses in zebra finches are shown to gradually appear during the first week posthatching and to mature in subsequent weeks, consistent with the auditory development in other altricial bird species studied previously. When matching the acoustic properties of parental heat whistles and auditory sensitivities, hearing of the parental heat whistles by zebra finch hatchlings was convincingly excluded. Although not directly measured, this also convincingly extrapolates to zebra finch embryos. Finally, the authors tested the hypothesis that parental heat whistles could induce perceptible vibrations of the egg and thus stimulate the embryo via a different modality. The method used here was laser doppler vibrometry, an appropriate, state-of-the-art technique that the authors also have proven experience with. The induced vibrations were shown to be several orders of magnitude below known vibrotactile sensitivities in mammals and birds. Thus, although zebra finch vibrotactile thresholds were not obtained directly, the hypothesis of vibrotactile perception of parental heat whistles by zebra finch embryos could also be rejected convincingly.

      In summary, even when considering some weaknesses of the techniques (which the authors are aware of), the conclusions of the paper are well supported: Auditory and/or vibration perception of parental heat whistles can be excluded as an explanation for previous reports of developmental programming for high ambient temperatures. As a constructive suggestion towards resolving the apparent paradox, the authors recommend to repeat some of the crucial, previous playback experiments at lower sound levels that better match the natural parental vocalizations.

    3. Reviewer #2 (Public review):

      This study by Anttonen, Christensen-Dalsgaard and Elemans describes the development of hearing thresholds in an altricial songbird species, the zebra finch. The results are very clear and along what might have been expected for altricial birds: at hatch (2 days post-hatch), the chicks are functionally deaf. Auditory evoked activity in the form of auditory brainstem responses (ABR) can start to be detected at 4 days post-hatch but only at very loud sound levels. The study also shows that ABR response matures rapidly and reaches adult like properties around 25 days post-hatch. The functional development of the auditory system is also frequency dependent with a low to high frequency time course. All experiments are very well performed. The careful study throughout development and with the use of multiple time-points early in development is important to further ensure that the negative results found right after hatching are not the result of the experimental manipulation. The results themselves could be classified as somewhat descriptive but, as the authors point out, they are particularly relevant and timely. Since 2016, there has been a series of studies published in high profile journals that have presumably showed the importance of prenatal acoustic communication in altricial birds, mostly in zebra finches. This early acoustic communication would serve various adaptive functions. Although acoustic communication between embryos in the egg and parents has been shown in precocial birds (and crocodiles), finding an important function for pre-natal communication in altricial birds came as a surprise. Unfortunately, none of those studies performed a careful assessment of the chicks' hearing abilities. This is done here, and the results are clear: zebra finches at 2 and 6 days post hatch are functionally deaf. Since, it is highly improbably that the hearing in the egg is more developed than at birth, one can only conclude that zebra finches in the egg (or at birth) cannot hear the heat whistles. The paper also ruled out the detection on egg vibrations as an alternative path. The prior literature will have to be corrected, or further studies conducted to solve the discrepancies. For this purpose, the "companion" paper on bioRxiv that studies the bioacoustical properties of heat calls from the same group will be particularly useful. Researchers from different groups will be able to precisely compare their stimuli.

      Beyond the quality of the experiments, I also found that the paper was very well written. The introduction was particularly clear and complete (yet concise).

      Weaknesses:

      My only minor criticism is that you don't discuss potential differences between behavioral audiograms and ABRs. Optimally, one would need to repeat the work of Okanoya and Dooling with your set up and using the same calibration. The ~20dB difference might be real or it might be due to SPL measured with different instruments, at different distances, etc. Either way, you could add a sentence in the discussion that states that even with the 20 dB difference in audiogram heat whistles would not be detected during the early days post-hatch, but that adding a (novel) behavioral assay in young birds could further resolve the issue.

      More Minor Points.

      (1) As mentioned in the main text, the duration of pips (form pips to bursts) affects the effective bandwidth of the stimulus. I believe that you could give an estimate of this effective bandwidth given what is know from bird auditory filters. I think that this estimate could be useful to compare to the effective bandwidth of the heat-call which you can now also estimate.

      (2) Fig 5b. label the green and pink areas as song and heat-call spectrum. Also note that in the legend you say: "Green and red areas display the frequency windows related to the best hearing sensitivity of zebra finches and to heat calls, respectively". I don't think this is what you meant. I agree that 1-4 kHz is the best frequency sensitivity of zebra finches but you probably meant green == "song frequency spectrum" and pink == "heat call spectrum". In either case the figure and the legend need clarification.

      (3) Fig 5c. Here also, I would change the song and heat-call labels to "song spectrum", "heat call spectrum". You don't want readers to think you used song and heat calls in these experiments (maybe next time?). For the same reason, maybe in 5a you could add a cartoon of the oscillogram of a frequency sweep next to your speaker.

      (4) Methods. In your description of the stimulus, you describe "5ms long tone bursts" but these are the tone pips in the main part of the manuscript. Use the same terms.

      Comments on revisions:

      In the latest version of this manuscript, Anntonen et al have diligently addressed all the issues raised by the reviewers. As I mentioned in our discussions among reviewers, it is impossible to "prove" a null hypothesis and both methods and experimental design could always be improved. At this stage however, they provide convincing evidence that the sound intensity of heat calls are below the hearing thresholds of zebra finch chicks.

    4. Reviewer #3 (Public review):

      Summary

      This study aims to contest recent findings that prenatal exposure to natural sounds and anthropogenic noise before hatching affects development and fitness in an altricial songbird. To this aim, it attempts to estimate hearing capacities of zebra finch nestlings. It first uses responses to clicks in nestlings and adults to estimate differences in hearing thresholds but uses experimental parameters that systematically lower nestling responses. It then measures responses to tones but restrict nestling data to a protocol (long tones) that failed in adults; response to tones with a correct protocol (short tones) is repeated in adults only. Thirdly, it includes data on loud airborne sound making eggs vibrate, even though this has no relevance to embryonic vibration perception by direct contact with the incubating parent or in incubators. Lastly, it includes a lengthy discussion on how these results, even though inaccurate or incorrect, would show that zebra finch nestlings and embryos are "functionally deaf". It fails to note that even if the experiment had been performed correctly and still failed to detect an auditory response in 2 day hatchlings - which is unlikely given the above and findings in other songbirds - this would not somehow eliminate the developmental effects of prenatal sounds that have been empirically demonstrated.

      Strength:

      The study is not performed adequately to bring reliable answers, but it addresses the important topic of hearing development in altricial songbirds and the long-held (but untested) assumption that altricial avian embryos cannot hear. More broadly, there is a need to reassess avian auditory perception - and the methodological approaches to measure it - given the accumulating evidence that some bird species respond to high frequency biological sounds beyond their known hearing range.

      Weaknesses:

      The study presents many experimental flaws that specifically compromise response detection in immature animals in the first experiment, and data from the remaining two experiments are invalid. The revision did not fix any of these issues.

      i) Response to clicks, Fig 1: Unlike what the revision claims, the deviations from validated protocols (too rapid stimulation, too few measurements, low temperature, reliance on clicks only) do lead to a potentially large underestimation of nestling hearing sensitivity. Calculation and new data trying to show that these deviations would not matter are wrong or insufficient. Even without strict conventions on ABR methodology, it is striking that every single experimental choice made here i) greatly differs from other avian studies and ii) reduces - rather than maximises - response detection, specifically for low amplitude signals (as in nestlings or near thresholds).

      ii) Response to frequency tones, Fig 2: The experiment on frequency sensitivity (long tone burst) failed in adults (positive control) with 0 to only about half of the adults responding (Fig S1), and sensitivity underestimated by 30dB (Fig2). Measures with a failed positive control are invalid, even if the overall pattern of the few data points obtained in nestlings vaguely resemble (but does not match) expectations. That the experiment was repeated in adults with a correct protocol (short tone pips) - but not in nestlings - is highly misleading.

      iii) Attempts to validate the protocols above by comparing to published adult values are meaningless (response A.3.1; Fig 5). The one experiment being compared (short tones) was only performed in adults in this study. This cannot validate results obtained by different methodologies in nestlings, for responses to clicks or long tones.

      iv) Vibrations: No previous results on the effect of zebra finch heat calls on development rely on the hypothesis or assumption that airborne sound from heat calls makes eggs vibrate. This idea is solely attributable to the authors of the preprint, and is not biologically or experimentally realistic. All speculations from this experiment (Fig3) on vibration perception by embryos or heat call effects are meaningless.

      v) Writing and presentation: The text throughout is highly misinformative for non-specialists or any reader not carefully inspecting figures (including in supplementary material) and methods. The confusion is greatly aggravated in the (6-page long) discussion which i) fails to recognise and account for the study shortcomings, and instead ii) greatly overstates the results and what they mean, and iii) misrepresents current knowledge by excluding highly relevant studies showing evidence of early sound perception in embryos and hatchlings, and introducing many errors in the presentation of published papers. This creates an illusion of a strong mismatch between heat-calls and zebra finch sensory capacities, for which the study actually does not provide any evidence for, and which is extremely unlikely to exist.

      vi) The revision did not address any of the major flaws of the study outlined above (see detailed assessment below). In particular,<br /> ** For point i):<br /> - estimations for the effect of having done too few (400) measurements are wrong - the effect on hearing thresholds cannot be calculated, but would be much greater than 4dB with the expected 37% noise reduction with the standard 1000 sweeps;

      - the new data provided does not measure the impact of high stimulus rate, and measures on adults largely underestimate effects on nestling response;

      - body temperature, now provided in the revision, is at the lowest extreme for the species, which may increase hearing thresholds;

      - tones do elicit a stronger and earlier response than clicks. Whether this is related to stimulus duration does not change the fact that clicks underestimate hearing thresholds, and delay hearing onset by several days.<br /> ** For points ii) and iv):<br /> No new data or any valid explanation was provided in the revision. It is still the case that nestling frequency responses were obtained with an experiment where the positive control failed; and the vibrating egg experiment is irrelevant to vibration perception in embryos and any observed effects of heat calls on development.<br /> ** The revision greatly lengthens the speculations about heat call perception, based on inaccurate or totally incorrect data.

      Conclusion and impact:<br /> i) Overall, the study fails to provide any reliable estimate of zebra finch nestling hearing capacities: it underestimates hearing onsets and thresholds (clicks) and gives no information on nestling frequency sensitivity. It is extremely likely that a better designed ABR experiment, or a more sensitive methodology (e.g. electrophysiology), would have detected a response in 2 day-old nestlings, as in other songbirds. The conclusion of "deafness" in hatchlings (or embryos, which were not tested) is clearly unsupported.<br /> If zebra finch hatchlings were clearly deaf, a well-designed study would have shown this a lot more convincingly.

      ii) The data on adults in not new (3 prior studies) - although this study generally underestimates zebra finch high frequency sensitivity. The presentation in relation to heat-calls is flawed: true values of heat-call frequency range and sound levels (>6kHz at 45dB) fall within the adult hearing range.

      iii) Without any new evidence, this study does not progress our understanding of heat-call and noise impact on development. Exactly the same issue remains: that heat-call and noise effects on development contradict the general VIEW on hearing ontogeny in altricial birds. But we have learnt nothing from this study on hearing ontogeny or frequency sensitivity. The study provides no reliable neuroscience data to advance the debate.

      iv) The largest section of the preprint is a highly speculative discussion, based on erroneous data and wrong interpretations, as well as a misrepresentation of what is known. That these issues would not be recognised is - in my view - of serious concern.

      Detailed assessment:

      The summary by the authors of my major concerns (R3.A0) is incorrect and misreport many of my statements, without giving any meaningful answers. I will not engage in such discussion.

      My actual three major concerns do remain:

      (1) Results on Fig 1 overestimate hearing onset and thresholds: All experimental parameters chosen to test responses to clicks are known to lower detection. From their cumulative impact, there is absolutely no doubt that the data in Fig 1 underestimate hearing capacities, and disproportionately so in nestlings compared to adults. Therefore, no quantitative estimate of nestling hearing threshold, absolute or relative (i.e. the estimated "54dB difference"), can be taken from this study. The age of hearing onset based on clicks (4 day old) is also wrong.

      (2) All results on Fig 2 are false: 40 to 100% of adults (positive control) failed to respond within their normal hearing range (Fig S1). Data on nestling frequency sensitivity using this methodology (Fig 2) are clearly invalid.

      (3) All results on Fig 3 are irrelevant: they assume zebra finch parents are hoovering in front of their nest while heat-calling rather than incubating their eggs, which is nonsense. Any speculation on the role of bone-conduction or vibrotactile stimulation for embryonic heat-call detection is simply unfounded.

      The authors' responses to these comments are largely mistaken (see below), and do not change the 3 facts stated above. Overall, the data produced are unreliable, and so are the interpretations and conclusions.<br /> The conclusions that i) heat-calls fall outside of adult hearing range and ii) young nestlings (and embryos) are deaf, are both incorrect. The study provides no actual estimate of nestling hearing and how far off their hearing range heat-calls fall.

      (1) Responses to clicks underestimate hearing abilities, specifically in nestlings.

      (i) DISPROPORTIONATE EFFECT OF PROTOCOL ON NESTLINGS: As reported in other avian studies (e.g. Brittan Powell et al 2004), because of their immature neural system, nestlings show low amplitude waveforms compared to adults. This occurs even with sound much louder than their hearing threshold. For example here, 8 day old nestlings show waveforms of only low amplitude at 95dB, even though they respond to sound 40dB softer, at 55dB (Fig 1H).

      This characteristic of nestling waveforms means that:<br /> - nestling responses are harder to detect,<br /> - nestling responses are more easily attenuated below the detection criteria (>2 S/N),<br /> - a low amplitude response in nestlings at a given sound level does not predict that softer sounds will not be perceived by the animal.

      Any deviation in protocol that attenuates waveform amplitude will therefore disproportionately affect nestling thresholds (and artificially lead to the "54dB" estimate).<br /> Even without universal standards (resp R3.A1.1), knowing this, the protocol should be adjusted to maximise response detection. This study does exactly the opposite.

      (ii) INSUFFICIENT MEASUREMENTS: using only 400 sweeps, rather than the typical 1000 sweeps, reduces response detection, especially in nestlings.

      - The claim that using 1000 sweeps would only decrease the hearing threshold by 4dB (response A3.2; ms L530) is false:<br /> - Based on the square root relationship mentioned by the authors, using 1000 averages instead of 400 would improve the signal-to-noise ratio by 37%, which would allow detecting many small amplitude waves which are currently hidden in the abnormally high noise.<br /> - Lowering the noise floor level by 4dB does not mean that the hearing threshold would only decrease by 4dB. The improvement in hearing threshold would be much greater, especially in nestlings.

      (iii) UNSUITABLY HIGH STIMULATION RATE: the new data on pairs of clicks greatly underestimate the attenuation by high stimulation rates, and so do adult measurements compared to nestlings'.<br /> - the click rate used in this study (25 clicks per second) is 5 to 25 times faster than most previous studies in young birds (e.g. Saunders et al 1973, 1974: 1 stim/sec or less; Katayama 1985: 3.3 clicks/sec; Brittan-Powell et al 2004: 4 stim/sec), and 6 times faster than zebra finch heat-calls.<br /> - responses to pairs of clicks (new data in Fig S3; L 543-552, response A3.3) does not measure the response dampening caused by high stimulation rates. Response attenuation (adaptation) after a single click is much weaker than after a train of 400 consecutive clicks (or even just 30). This paired-click measure ignores the cumulative attenuation observed with many consecutive stimuli. Paired-click paradigm is used to measure immediate refractoriness for other purposes (e.g. temporal resolution, diagnosis tool) but does not replicate high stimulation rates.<br /> It is unclear why the authors chose to use this weak approximation rather than simply replicating the measurements at the same slow rate as in other avian studies. This would have given a straightforward answer, comparable to previous studies that have demonstrated the detrimental impact of fast rate on response strength by directly comparing responses to different stimulation rates (e.g. Saunders et al 1973; Brittan-Powell et al 2004).<br /> - Adults are less sensitive to high stimulation rates than immature individuals (e.g. Saunders et al 1973; Khayutin 1985; Brittan-Powell et al 2004 and 6 references cited therein). Effects of fast rate in adults (new data in Fig S3) therefore largely underestimate effects on nestlings (not measured), and the high stimulation rate used in this study increases the relative difference between nestlings and adults.<br /> - Given the 2 points above, it is incorrect to conclude from this new data that the high stimulation rate used had "no effect on ABR amplitude or thresholds" (responses A3.3 L551). Instead, given that 40ms (i.e. interval for 25 clicks/ sec) is at the limit of what adults can handle after a single click, this new data does confirm that the stimulation rate is indeed too high and underestimates hearing in nestlings (and adults to a lesser extent).

      (iv) LOW BODY TEMPERATURE can reduce ABR response.<br /> - The average body temperature (39.5C) now provided in the revised manuscript (response A2) is at the lowest extreme of the range for zebra finches. The normal average body temperature for zebra finches is 41C at low ambient temperature, rising to 43-44C at high ambient temperatures (when heat-calls would be produced). This value of 41C is consistent across many studies in wild and domestic zebra finches (e.g. Wojciechowski et al 2020 [avr 41C at 23C]; Udino and Mariette 2022 [avg 41, min=40C at 32C]; Bech and Midtgard 1981 [avr ~41C]; Pessato et al 2022 avg=41C at 27C]). The only study reporting a body temperature average as low as 39C (Cooper et al 2020) was indeed under hypothermic conditions, obtained in adult zebra finches under extreme fasting conditions (17hrs of food deprivation), at ambient temperature below thermoneutrality, during the night (i.e. during "nocturnal hypothermia", when bird body temperature is normally lower).<br /> Therefore 39.5C does corresponds to hypothermic conditions for most individuals. While this body temperature remains considerably higher than that used experimentally to supress hearing response, using a below-normal body temperature, added to other factors, may lower wave amplitude and therefore increase hearing threshold estimates.

      (v) CLICKS UNDERESTIMATE HEARING ONSETS<br /> - My statement that, compared to tones, clicks elicit a smaller response, at a later age, is correct (Saunders et al 1973; Brittan-Powell et al 2004).<br /> - It does not matter whether this is due to differences in stimulation duration (response R3.A2.1), since clicks are always much shorter than tones. What matters is that hearing onset based on responses to clicks has been found to underestimate the earliest age at which an auditory response can be obtained by several days (Saunders et al 1973).<br /> - Without any valid nestling data for tones (see below), this study, based on clicks only, does not provide the true age of hearing onset.<br /> - The suggestion of using 170dB clicks (response R3.A2.1) is very odd. The correct approach would have been to use a correct protocol for tones in nestlings, rather than solely correcting it for adults (see below).

      (vi) CONFUSING INTERPRETATION OF WAVE AMPLITUDE AND LATENCY<br /> - the text implies throughout that nestlings lacking an adult-like amplitude and latency is a sign of poor hearing (e.g. L122 "ABR wave I gradually reaches maturity 25 days after hatching").<br /> - It should be clarified that this is a characteristic of immature systems but does mean these nestlings do not hear. For example, this characteristic persists even in 10d old nestlings which have similar hearing thresholds to adults.

      (2). Results on nestling frequency sensitivity are invalid. Using long (25ms) tones with subdermal electrodes is incorrect, and not used by anyone other than the authors. The failure to repeat the experiment with a correct protocol (5ms tones) in nestlings is misleading:

      (i) When 60% of 4-day old chicks respond to clicks (Fig 1), they are described as "functionally deaf" (L78, L261). By contrast, when 0% to 60% of adults respond to tones in the core of their hearing range (Fig S1 for data shown in Fig 2), this is merely presented as a methodological limitation (text added L186-197), and the authors still proceed to presenting results on nestlings with a method that failed in adults (positive control).

      (ii) When the positive control fails, the experiment is failed. This principle applies in Neuroscience as in any scientific field.

      (iii) That the experiment repeated in adults with the standard and correct short tone protocol (5ms) gave normal results, is not a validation. Instead, it proves the point that 25ms is unsuitable (regardless of whether that is due to rising time, L399). This correction does NOTHING to fix results in nestlings, which were ONLY done with the unsuitable 25ms tones.

      (iv) The mixture of adult data with 5 and 25ms tones, instead, creates confusion, because differences in protocols between nestlings and adults are blurred in the text. The abnormally small proportion of adults responding (0 to 60%) with 25ms tones is only shown in supplementary material, and attenuated in the main text (L186: "about half").

      (v) The claim (response R3.A1.2) that long 25ms tones is commonly used with the ABR methodology used here is false:<br /> ** ALL (but 1) avian studies using tones of 20ms or longer (cited by myself or by the authors in their reply R3.A1.2) used a different ABR methodology, with electrodes implanted through the skull. The only exception, using 25ms tones with subdermal electrodes (as here), is a study by the authors themselves.<br /> ** All other avian studies with subdermal electrodes used short tones, of 5ms or less (e.g. Brittan-Powell et al 2002, 2004; Henry & Lucas 2008), as in the corrected adult experiment.<br /> ** My initial comment already specified this difference in electrode placement.<br /> ** The results here (failing in adults) unquestionably show that the method of long tones with subdermal electrodes does not work, and the literature shows that no one else uses it.

      (vi) When the aim of the study is to contest effects of high frequency sounds (L39-45), it is odd to choose a protocol specifically directed at testing responses to very low frequencies (responses A3.1, R3A.0; L398-403) and that compromises all results.

      (vii) That the shape of the few data points obtained in nestlings with long tones would broadly resemble expectations (responses A3.1, R3A.0) is not a validation: See ii.<br /> - The shape is not even correct:<br /> ** Why aren't 4 and 6 day old nestlings responding to any frequency when they were responding to clicks (in spite of poor detection conditions for clicks)?<br /> *** Why are they not responding, when 2 day old flycatchers respond to tones from 1 to 4kHz at 45dB, and 0.5 to 5kHz at 60dB (Aleksandrov and Dmitrieva 1992, Korneeva et al 2006)?

      (viii) That only relative measures between nestlings and adults matter (resp R3.A4; L404) is incorrect. Measures that are qualitatively (see vii) and quantitatively (see i) wrong cannot be compared.<br /> If these data with long tones were indeed acceptable, why repeat the experiment in adults with a correct protocol, even before this flaw was pointed to during the review?

      (ix) The unusually low number of measurements (400 instead of 1000; L530) will have reduced detectability of low amplitude responses, in nestlings and near thresholds, here as in the click experiment (see above). The claim that this would only increase threshold by 4dB (responses A3.2, L530) is wrong (see above).

      (x) There is no explanation as to why only half of the 6-day old nestlings (n=5-6) were tested with tones. Would having 10 in this group shown a response at 6-day old?

      (xi) Even the correct 5ms tone protocol in adults overestimated thresholds at higher frequencies (>3kHz) by up to 30dB compared to other published estimates (Fig 5). If inter-population differences among domestic zebra finches (L 393; Fig 5 legend; resp R3.A4) were enough to cause 30dB differences, results on domestic zebra finches here should not be extrapolated to those on wild-derived Australian zebra finches documenting developmental effects of heat calls.

      The authors give no other elements than the above in their responses that could demonstrate the validity of their nestling frequency data.

      Based on other studies, we can expect zebra finch hearing to develop sensitivity to high frequencies after that to middle frequencies. But pretending that this study provides any evidence towards this is wrong. There is no valid data.

      (3) The experiment using loud sounds to make eggs vibrate is biologically and experimentally meaningless. The claim that these measurements would rule out vibration perception in embryos is totally unfounded. The preprint is creating the illusion of having ruled-out a mechanism that they have not tested.

      (i) zebra finches are not hovering in front of their nest when heat-calling. They are in physical contact with the eggs while incubating, with vibrations expected to travel directly through solids from adults to eggs, without attenuation.

      (ii) the claim that previous studies on heat-call developmental impact rely on the assumption of airborne sound making egg vibrate (response A3.5, R3.A8) is incorrect. This is conceptually wrong and there is no indication of this in any paper.

      (iii) vibration perception in embryos in birds and other taxa (e.g. amphibian, reptiles and insects) rely on direct contact with the source, not on loud airborne sounds shaking eggs. Beyond any consideration on heat-calls, claiming that this experiment could give any indication of vibration perception in embryos is nonsense.

      The authors give no information in their responses that could demonstrate the validity of this experiment.

      (4) Interpretation and discussion

      All other songbird studies, individually and collectively, show much greater hearing sensitivity in nestlings that what this study is trying to suggest (e.g Khayutin 1985; Korneeva et al 2006; Aleksandrov and Dmitrieva 1992; Rivera et al 2018; Platzen and Magrath 2004; Haff & Magrath 2012). The only way the authors can reach their conclusion is by:<br /> - presenting data that greatly underestimate hearing sensitivity in zebra finches (see sections 1 and 2) and overstating them, and;<br /> - misrepresenting current knowledge by excluding relevant papers and being unclear about what the literature shows.<br /> This study tries to impose the idea that zebra finch do not detect any sound before day 4-6 post-hatch, and have extremely rudimental hearing until day 8-10 post-hatch. By contrast, other studies show very young hatchlings respond to sound 1 to 3 days after hatching (first age tested) and show quite sophisticated, and totally functional, responses to relevant sounds at 5 days old.

      This is not to say that hearing does not continue to improve post-hatch in birds, or that sensitivity to mid frequencies post-hatch would not precede that to high frequencies. But statements throughout this study are so exaggerated and/or wrong that nothing can be learnt. It is undeniable that this study fails to provide the useful and balanced assessment needed to establish were true heat-calls actually sit relative to zebra finch adult, nestling and embryonic hearing range.

      It is literally impossible to correct every wrong statement in the discussion and responses to reviewers. I focus here on some examples related the claims of "nestling deafness" or of heat-calls being outside of adult zebra finch hearing range, as well as inaccuracies leading to an apparent match with current literature.

      (i) MISALIGNMENT WITH OTHER STUDIES AND MISREPORTING

      *** Neurological evidence

      - Highly relevant evidence on response to sound in zebra finch embryos (Rivera et al<br /> 2018) is totally excluded.<br /> Excluding this study on the basis that it used a different methodology that does not directly quantifying auditory sensitivity (resp R3.A6a) is a poor justification. Evidence, even indirect or imperfect, should be brought to the attention of the readers.

      - This applies to many other studies cited in my first review and arbitrarily excluded here. If one wants to conclude on "deafness", all evidence, even indirect, should be considered. Excluding non-ABR studies means the conclusion cannot be extended beyond flat ABR traces.

      - Other studies show much greater sensitivity that what the text describes:

      * Flycatcher hatchlings at 2-3d post hatch (first age tested), respond across a wide range of frequencies (0.3 to 5kHz), at low to moderate sound levels (45-65dB)<br /> (Aleksandrov and Dmitrieva 1992, Korneeva et al 2006).

      * Stating that these studies in flycatchers "likely yield lower thresholds" (L364) is an astonishing understatement. Thresholds in Aleksandrov and Dmitrieva (1992) at 2-3 day old were 35 dB lower than those here at 4 day old with clicks, and 60 to 80dB lower than those with tones of 1-2kHz at 8 day old.

      * Claims that "sensitivity improves rapidly postnatally, with ~40 dB threshold decreases (L365)" is also misreporting these studies' findings. Thresholds decreased by 25dB consistently at 9 out of 11 frequencies tested from 0.3 to 8kHz (Aleksandrov and Dmitrieva, 1992). A difference of 40dB was only found at 5-6kHz (Aleksandrov and Dmitrieva, 1992). Likewise, improvement also varied from 25 to 40dB in Korneeva 2006. These do not average to "~40 dB".

      * Even birds developing 4 times slower than songbirds (budgerigars) show a response at 5 day old. My statement that Brittan-Powell 2004 shows an auditory response at 5d old is correct. Re-response R3.A2.2: Fig 1B shows one example for ONE individual. All other figures based on multiple individuals shows at least some individuals responded at 5-6 days old at frequencies less than 4Khz (Fig 1C, 4 and 5).

      - Many inaccuracies on avian hearing remain uncorrected in the revision.<br /> * e.g. L 329: extrapolating high frequency hearing (>6khZ) from precocial species is incorrect because even adult chicken and ducks are not sensitive to high frequencies.<br /> The authors imply elsewhere that species of songbirds cannot be compared (resp A4, R3.A6), but make extrapolations from species that are far more remote, phylogenetically, developmentally and ecologically than other songbirds.

      *** Behavioural evidence

      - contradicts this study findings:

      * When correctly cited and described, the literature, does not support the authors' statement that nestling behavioural response "typically emerges between ~5-10 days post-hatch" (L358). It emerges earlier, at an unknown age, including potentially from hatch (present at 1.5day) for innate responses (see below).

      * Results in other songbirds are not consistent with this study finding that a "response to loud click stimuli is first detectable at 4-8 days post hatch" (L230). Instead they show nestling hearing capacities described here are abnormally poor.

      * Therefore, the conclusion that "The timeline of behavioural studies closely matches the onset and maturation of ABR responses observed here in zebra finches" (L360) is wrong.

      - shows a very early response (1-2day post hatch), with no known onset:

      * songbird behavioural response to sound, with parental alarm call suppressing begging, has been demonstrated in nestlings as young as 1.5 or 3 days old (Khayutin 1985, Korneeva et al 2006, Aleksandrov and Dmitrieva 1992). None of these results are mentioned in the revision when discussing behavioural evidence (L357-360).

      * Instead, they exclusively mention studies that started testing nestlings at 5-6 days old, or much later (e.g. 17 day old: Suzuki 2011; 14 day old: Barati & McDonald 2017).

      * ALL of these studies (cited or not) demonstrated a significant response of nestlings to calls on the first age tested. NONE tested the onset of this response.

      * the one study looking at progression across 3 ages, at 5, 8 and 11 day old shows parental alarm calls suppress nestling calling at day 5 as much as later on, with no effect of age (Platzen and Magrath 2004). The authors failed to acknowledge this in their response (resp A4) or revision (L359).

      * the claim that Haff & Magrath (2012) showed nestlings did not respond at 5-6 day old but did at 10-11 days (resp A4) is wrong. At 5 day old, they responded to their own species alarm calls, as well as to another similar sounding species and to the sound of predators themselves (Table2 in Haff & Magrath 2012; as in Platzen and Magrath 2004).

      - learning, not just hearing, improves nestling response with age:

      * Haff & Magrath (2012) showed that by 10-11 day old, nestlings had learnt to also respond to heterospecifics, demonstrating that learning improves nestling response with age.

      * the intensity of the response to low frequency sound improved more with age than that to high frequency calls. If improvement were related to hearing limitations, the opposite would be expected (Haff & Magrath (2012).

      - nestlings discriminate complex calls, including at high frequency

      * By 5-6 day old, nestlings can already discriminate several different sounds indicative of danger, among the complex natural acoustic background (Haff & Magrath 2012).

      * nestling do not respond indiscriminately to any calls, but only to relevant sounds that specifically present a threat to them (Haff and Magrath 2012; Magrath et al 2006).

      * the idea that nestlings only distinguish "low frequency broadband cues" (L362) is inaccurate (resp R3.A6). Nestlings respond to scrubwren chip calls and fairywren alarm calls that are narrowband calls with the fundamental at 8 and 10 kHz respectively (Platzen and Magrath 2004; Haff and Magrath 2012).

      * These studies indeed "do not imply mature auditory sensitivity" (L362), they show that auditory maturity is not needed to show a perfectly functional response to biologically meaningful sounds in a natural context.

      - Overall, every other songbird species tested shows greater hearing sensitivity than that proclaimed here for zebra finches. It is very unlikely that zebra finches would be such an outlier.

      (ii) INCORRECT CONCLUSION ON DEAFNESS

      Deafness of young zebra finch nestlings cannot be demonstrated because:

      - the study has no valid response to tones in nestlings to establish the age of hearing onset.

      - responses to clicks are greatly underestimated (see 1). Had correct, more sensitive, protocol parameters been used, it is very likely that:<br /> * Most 2 day old nestlings would have responded to clicks, instead of 60% of 4 day old nestlings.<br /> * Thresholds would be lower, especially in nestlings.<br /> The >54dB difference between nestlings and adults based on an assumed 95dB threshold in 2d old nestlings is wrong.

      - based on data on other songbird nestlings, a difference of 25dB would be more realistic (at frequencies of 1-2 kHz, comparable to clicks, Aleksandrov and Dmitrieva 1992, Korneeva et al 2006). This greatly contrasts with the >54dB estimate here. While the authors qualify their estimate as "conservative" (legend Fig4, L293), it is actually greatly overestimated.

      - Even assuming the data in this study were correct, the interpretation is erroneous. Concluding "deafness of young nestlings" is incorrect when 60% of 4 day old nestlings respond to clicks at 80dB. This is especially wrong given that the ABR method systematically overestimates thresholds by 20-40dB.

      - That ABR is used in humans to diagnose deafness (L263) does not make ABR the most suitable method for birds, when evidence shows that other methods (e.g. electrophysiology or behaviour) are more accurate. Hearing screening in humans can only use non-invasive methods, and has additional criteria than accuracy (price, ease, etc).

      - The discussion fails to acknowledge the implications of the limitations of the study (detailed above). For example, talking in broad terms of differences in protocols (L392-395), does not tell readers that this study, because of the parameters chosen, led to an underestimation of zebra finch hearing capacities.<br /> The discussion instead greatly overstates what the study shows (e.g. L229-232, 261, 277-278, 288, 332, 340, etc).

      (iii) INCORRECT CONCLUSION THAT HEAT-CALL FALL OUTSIDE THE ADULT HEARING RANGE.

      - heat-calls fall outside the adult hearing range solely because of the authors' decision to:<br /> * restrict heat-call frequency range to the authors' own estimation (7-10kHz) on 4 birds instead of using published values that caused developmental impact (e.g. Katsis et al: 6-10kHz), or were produced in vitro (5.9kHz; Anttonnen et al 2025);

      * lowering heat call sound level (34dB at 10 cm) to measurements on 4 isolated birds under unknown temperature conditions for an unknown amount of time, when the same paper gives values of 43.4 dB at 1 m (range: 30.7-53.2) in standard in vitro conditions.

      * As soon as EITHER of these two values is corrected, the statement that heat calls are outside zebra finch adult hearing range is false.

      * In addition, assuming constant heat-call sound level across contexts is unreasonable. That zebra finch can produce inspiratory syllable very similar to heat calls and during inspiration at 65dB during song (Goller & Dalley 2001) argues against the assumption that heat-calls are always soft.

      - Other species also show perception, not just production (resp R3.A10a), of calls above their known hearing range (10-20kHz; e.g. Duque et al 2020). The zebra finch is not the only case in birds of mismatch between signals perceived and known hearing range.

    5. Reviewer #4 (Public review):

      Reviewing Editor:

      There is a virtuous circle between behavior and neuroscience. Sometimes neuroscience identifies a signal that no behavioral study alone could discern: place cells in rats, replayed song sequences in sleeping birds, compass-like signals in the central complex of the fruit fly. Sometimes behavioral observation is what inspires neuroscience: precise measurements of short and long latency reflexes implicate different neural pathways, and directed escape responses in fish are so fast that they require specialized and lateralized escape circuits. And sometimes a behavioral claim requires an animal's sensory system to possess a capacity nobody had documented. A prime example is bat echolocation: Spallanzani inferred in 1793 that bats navigate by hearing, but the claim was dismissed for over a century because the signal he proposed could not be detected. It was vindicated only when Griffin and Galambos measured the ultrasonic emissions directly in 1940, and the loop closed fully when Suga and colleagues found cortical neurons in the mustached bat tuned to precisely the echo delays and Doppler shifts the behavior required.

      In this manuscript we are faced with a fascinating and contemporary instance of this important dialogue between behavior and neurophysiology. Several high profile papers have reported that incubating zebra finch parents produce high frequency "heat calls" when ambient temperatures rise, and that playback of these calls to eggs during late incubation alters offspring growth, begging behavior, thermal preference, and reproductive success in adulthood. That function requires that the embryo be able to sense, presumably to hear, the heat calls. This is textbook ecology and often cited as the prime example of adult behavior affecting the development of their unborn (or unhatched) offspring.

      Yet the auditory capacity of very young zebra finches had never been carefully examined. This paper does exactly that, using auditory brainstem responses, a method that detects synchronous volleys of afferent input through low levels of the auditory system. ABRs are not perfect and may miss very subtle or sparsely represented signals, but they are a time-tested way to assess hearing capacity and the development of a system.

      Here, the authors show, convincingly in my view and in the views of Reviewers 1 and 2, that 2 DPH hatchlings have no detectable ABR to a 95 dB SPL broadband click, and that sensitivity then rises progressively across the first two postnatal weeks. This is the most consequential result, since it bears directly on whether an auditory route to heat call perception is feasible at all. A second finding is that ABR wave I amplitude continues to mature until 20 to 25 days post hatch, coinciding with the onset of sensory song learning, just as young finches need to form a template of an adult male song to copy.

      The key question relevant to mediating this review process is: If a two-day-old hatchling shows no detectable auditory brainstem response, over 400 averaged sweeps, to a 95 dB SPL broadband click, in 14/14 animals tested, then how could that same animal, two days earlier and inside an egg, have been sensitive to a far weaker stimulus of roughly 33 dB SPL at 6.8 kHz?

      Below I set out the considerations that influenced my judgment as I handled these reviews.

      Strengths:

      The developmental series is the strength of the design. Seven ages, with multiple early time points, means the negative result at 2 DPH is not a lone flat trace but sits at one end of a graded and internally consistent trajectory. Body temperature was monitored and held within {plus minus}0.5{degree sign}C, and differed by no more than 0.6{degree sign}C across age groups. Every stimulus parameter was applied identically at every age, so the developmental comparison is a within-method one. The vibrometry experiment addresses the most obvious alternative modality with an appropriate and state-of-the-art technique.

      The key findings related to the development of the ABR response and, presumably, the hearing capacity. At two days post hatch, the authors presented broadband clicks of 20 µs duration at 95 dB SPL peak equivalent, a stimulus carrying energy across the full spectrum including the 6.8 kHz of the heat whistle. Clicks were delivered at 25 Hz with a 40 ms inter-stimulus interval, 400 presentations per condition with alternating polarity, and responses were scored both by an automated signal-to-noise criterion and by independent visual inspection. No response was observed in any of the 14 animals tested. Two days later, under the same protocol, eight of thirteen animals responded, with a mean threshold of 80.6 {plus minus} 1.6 dB SPL and wave I latencies of roughly 4.0 to 4.7 ms against 1.8 to 2.5 ms in adults. These responses were small, broad, and slow, the signature of an immature auditory system. By 6 DPH nine of ten animals responded, and from 8 DPH onward every animal did at every age tested. Adult thresholds settle at 40.6 {plus minus} 4.0 dB SPL, more than 54 dB below the 2 DPH value. The 4 DPH data show the preparation resolves weak, desynchronized responses under exactly the parameters used at 2 DPH, and that sensitivity emerges on a trajectory consistent with auditory development in other birds and in mammals.

      On Reviewer 3's methodological objections:

      Reviewer 3 raised a series of objections to the recording parameters: the stimulus rate is faster than in comparable avian developmental studies, 400 sweeps is fewer than the conventional 1000, body temperature sits at the low end of the reported range for the species, and clicks lag tones in developmental onset. Each of these describes a mechanism that would reduce the amplitude of an evoked response relative to some ideal stimulus. But in aggregate my take was that none of these would completely abolish a response that is detectable two days later. That distinction is the crux of my reading and others may disagree. A response reduced by 70 percent is still a response, and the question at 2 DPH is not why the trace was small but why there was no trace at all over 400 averages in all 14 animals tested.

      Reviewer 3's concerns are difficult to translate into threshold shifts, and on the narrow point that amplitude reductions do not map cleanly onto decibels, the difficulty that the reviewer faced in converting stimulus concerns to impact on dB threshold is well taken. But the figures supplied are themselves bounded: a 37 percent noise reduction from additional sweeps, a 70 percent amplitude reduction from stimulus rate in the youngest birds, a four day shift in apparent onset from using clicks rather than tones. These are large effects. They are not unbounded ones. And because the decibel is a logarithmic unit, in which every 20 dB corresponds to a tenfold change in sound pressure, the gap they are being asked to explain, exceeding 54 dB, amounts to a difference of more than 500-fold.

      The 4 DPH data bear directly on this, because every parameter at issue was applied identically at that age. The same 25 Hz rate, the same 400 sweeps, the same temperature, the same click stimulus resolved small, dispersed, long latency responses in 8 of 13 animals two days later. These are precisely the weak and desynchronized responses an immature auditory system is expected to produce, and precisely what the objections predict should have been lost. Whatever theoretically possible deficits in stimulus design, they did not prevent detection of a marginal response in animals two days older, and no mechanism has been proposed by which their cost would fall from total suppression at 2 DPH to negligible at 4 DPH.

      A 95 dB SPL broadband click carries energy at every frequency including 6.8 kHz and exceeds the level at which heat whistles arrive at 10 cm by more than 60 dB, before any attenuation by shell or by an incubating parent. Even a substantial underestimate of hatchling sensitivity leaves that gap open, and at 6.8 kHz it is a broadband stimulus being compared against a narrowband signal in the frequency region that matures last, in this dataset and in every developmental series reported.

      Sensory systems as matched filters:

      A second consideration: sensory systems as matched filters (survival-critical sensory responses are typically associated with expanded sensory sensitivity for them). Failure to detect an auditory response is not proof of functional deafness. An evoked potential measures synchronized population activity, so sparse discharge below the threshold for ABR sensitivity (a population response) can never be excluded. This is the crux of Reviewer 3's concerns.

      But this second consideration makes the difficulty of finding any evidence of hearing or vibrotactile responsiveness in these animals more compelling. Rüdiger Wehner, working on desert ant navigation, noted that sensory systems are matched filters. They are tuned to the narrow slice of the world that matters for survival and reproduction and discard the rest. From the tuning of an ant's polarization channel you can infer what it navigates by, and from the work of Capranica we know that the tuning of a frog's inner ear relates to the sound of conspecific frogs. If acoustic or vibrotactile processing of heat calls were genuinely critical to the embryo, natural selection would have built the outsized sensitivity to receive it. Instead there is a greater than 54 dB deficit and a frequency range where nothing is detectable at all. I interpret this as evidence that evolution did not build auditory sensitivity into two-day-old hatchlings, which suggests functional deafness in the embryo.

      One might invert this argument and object that the filter in the embryo is precisely for the heat whistle itself, so that neither a broadband click nor a 25 ms tone burst, the only two stimuli presented to 2 DPH animals, is the right probe. Feature detectors that respond to a specific signal while ignoring stimuli of far greater energy are common: cricket AN2 neurons fire to bat-like pulse intervals and stay quiet to loud broadband noise, and anuran midbrain neurons are interval-tuned and silent to spectrally matched noise exceeding the call in level. Embryonic sensory systems also carry transient specializations that later disappear, and the heat whistle is unusually specifiable, narrowband near 6.8 kHz and delivered in rhythmic trains. A detector tuned to that rhythm would not have been engaged by any stimulus used here, since the actual call was never played to any animal at any age.

      But feature selectivity is typically computed centrally, downstream of peripheral transduction, and ABR wave I reflects auditory nerve output, below any circuit that could implement pattern selectivity. A central detector still requires afferent input to operate on, and it is that input which seems to be absent here. It's conceivable that the cochlea itself exhibits some very special tuning to the heat whistle and implements some kind of yet-to-be discovered surround suppression at the periphery, such that a click would not be able to activate the hair cells in the heat whistle's frequency range. But 2 DPH animals were also tested directly with narrowband 6 and 8 kHz stimuli at 95 dB SPL, bracketing the heat whistle frequency, also without response. The surviving version of the objection therefore requires a pathway with too few fibers to generate a far-field response, tuned to a temporal structure nobody tested, in an animal whose auditory nerve shows no signal at some 60 dB above the natural signal level.

      Caveats:

      The paper comes with some caveats, which the authors note in their discussion. First, no embryo was tested. The embryonic claim is an extrapolation from hatchlings. I regard it as a reasonable one, since sound must additionally traverse the shell and since auditory systems gain rather than lose function as development proceeds, but it is an extrapolation nonetheless. Second, frequency-specific thresholds in nestlings were not obtained with a full stimulus set. The low to high developmental sequence therefore rests on a conservative stimulus. Pip data in nestlings would be a valuable addition to the record. Finally, the study offers no pre-neural measure, such as cochlear microphonics or otoacoustic emissions. Such a measure would distinguish a cochlea that is not transducing from one that transduces without synchronized output, and would address the strongest form of the objection raised in review.

      Summary:

      Absence of evidence is not evidence of absence, and some future study could in principle identify single neuron responses to heat calls in an embryo. The evidence in this paper makes that outcome unlikely in my assessment. It is worth noting what the precedent actually supports. Prenatal hearing is real and well documented in precocial birds: mallard embryos deprived of exposure to their own calls fail to recognize the maternal assembly call after hatching, and chickens show evoked responses well before hatching. But those animals hatch at a developmental stage altricial songbirds do not reach until days later, and the calls involved carry their energy below 3 kHz, the frequency region that comes online first. The claim at issue here requires sensitivity at 6.8 kHz, the region that matures last, in an animal at a far earlier stage. Something may still be happening to the embryo during heat calls, but whatever the mechanism, it is probably not the embryo listening. Behavioral claims that lack strong precedent, such as hearing through an eggshell, need to sit in a virtuous circle with neurophysiological investigation, and this study is what that looks like from the physiological side.

    6. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work by Antonnen et al. was triggered by claims of auditory-mediated effects on altricial avian embryos, which were published without any direct evidence that the relevant parental vocalizations were actually heard. I agree with Anttonen et al. that, based on the available evidence about avian auditory development, those claims are highly speculative and therefore necessitate more direct experimental verification.

      Attonen et al. have embarked on a comprehensive series of experiments to:

      (1) Better characterize acoustically the relevant parental vocalizations (heat whistles; in a separate preprint, not reviewed here)

      (2) Characterize the auditory sensitivity of zebra finches at various stages of their posthatching development. Despite the long-standing importance of the zebra finch as a songbird model in neuroethology of learned vocalizations, the auditory development of the species has not been studied so far.

      (3) Explore an alternative hypothesis of how the parental vocalizations might be perceived.

      The principal method used here is the non-invasive recording of ABR (auditory brainstem response), a standard neurophysiological method in auditory research. The click-evoked ABR provides a quick and objective assessment of basic hearing sensitivity that does not require animal training. Weaknesses of the technique include its limited frequency specificity and low signal-to-noise ratio. The authors are experienced with ABR measurements and well aware of those issues. ABR responses in zebra finches are shown to gradually appear during the first week posthatching and to mature in subsequent weeks, consistent with the auditory development in other altricial bird species studied previously. When matching the acoustic properties of parental heat whistles and auditory sensitivities, hearing of the parental heat whistles by zebra finch hatchlings was convincingly excluded. Although not directly measured, this also convincingly extrapolates to zebra finch embryos. Finally, the authors tested the hypothesis that parental heat whistles could induce perceptible vibrations of the egg and thus stimulate the embryo via a different modality. The method used here was laser doppler vibrometry, an appropriate, state-of-the-art technique that the authors also have proven experience with. The induced vibrations were shown to be several orders of magnitude below known vibrotactile sensitivities in mammals and birds. Thus, although zebra finch vibrotactile thresholds were not obtained directly, the hypothesis of vibrotactile perception of parental heat whistles by zebra finch embryos could also be rejected convincingly.

      In summary, even when considering some weaknesses of the techniques (which the authors are aware of), the conclusions of the paper are well supported: Auditory and/or vibration perception of parental heat whistles can be excluded as an explanation for previous reports of developmental programming for high ambient temperatures. As a constructive suggestion towards resolving the apparent paradox, the authors recommend repeating some of the crucial, previous playback experiments at lower sound levels that better match the natural parental vocalizations.

      (R1. A1) We thank the reviewer for their time and effort to thoroughly review our paper and for the positive comments on our manuscript. In the revised manuscript we have addressed the concerns that you have raised in the Joint recommendations above (Pages 1-4).

      Reviewer #2 (Public review):

      This study by Anttonen, Christensen-Dalsgaard, and Elemans describes the development of hearing thresholds in an altricial songbird species, the zebra finch. The results are very clear and along what might have been expected for altricial birds: at hatch (2 days post-hatch), the chicks are functionally deaf. Auditory evoked activity in the form of auditory brainstem responses (ABR) can start to be detected at 4 days post-hatch, but only at very loud sound levels. The study also shows that ABR response matures rapidly and reaches adult-like properties around 25 days post-hatch. The functional development of the auditory system is also frequency dependent, with a low-to-high frequency time course. All experiments are very well performed. The careful study throughout development and with the use of multiple time-points early in development is important to further ensure that the negative results found right after hatching are not the result of the experimental manipulation. The results themselves could be classified as somewhat descriptive, but, as the authors point out, they are particularly relevant and timely. Since 2016, there have been a series of studies published in high-profile journals that have presumably shown the importance of prenatal acoustic communication in altricial birds, mostly in zebra finches. This early acoustic communication would serve various adaptive functions. Although acoustic communication between embryos in the egg and parents has been shown in precocial birds (and crocodiles), finding an important function for prenatal communication in altricial birds came as a surprise. Unfortunately, none of those studies performed a careful assessment of the chicks' hearing abilities. This is done here, and the results are clear: zebra finches at 2 and 6 days post-hatch are functionally deaf. Since it is highly improbable that the hearing in the egg is more developed than at birth, one can only conclude that zebra finches in the egg (or at birth) cannot hear the heat whistles. The paper also ruled out the detection on egg vibrations as an alternative path. The prior literature will have to be corrected, or further studies conducted to solve the discrepancies. For this purpose, the "companion" paper on bioRxiv that studies the bioacoustical properties of heat calls from the same group will be particularly useful. Researchers from different groups will be able to precisely compare their stimuli.

      Beyond the quality of the experiments, I also found that the paper was very well written. The introduction was particularly clear and complete (yet concise).

      Weaknesses:

      My only minor criticism is that the authors do not discuss potential differences between behavioral audiograms and ABRs. Optimally, one would need to repeat the work of Okanoya and Dooling with your setup and using the same calibration. The ~20dB difference might be real, or it might be due to SPL measured with different instruments, at different distances, etc. Either way, you could add a sentence in the discussion that states that even with the 20 dB difference in audiogram heat whistles would not be detected during the early days post-hatch. But adding a (novel) behavioral assay in young birds could further resolve the issue.

      (R2. A0) We thank the reviewer for their time and effort to thoroughly review our paper, and for the positive comments on our manuscript.

      In our revision, we have added a new figure (Fig 5) and three new paragraphs (Lines 387-422) in the discussion to compare all published ABR and behavioral audiograms and the differences between and among these datasets. Our adult data is consistent with the reported findings from four other labs despite differences in stimulus design, setups and genetic background of the animals, providing strong support for our findings.

      Furthermore we have more clearly presented our argument why we think juvenile and embryos cannot detect heat whistles. For clarification, we have added a new figure (Fig 4) and four new paragraphs (Lines 271-355) in the discussion.

      We agree with the reviewer that more data is needed on the development of hearing in songbirds and zebra finches especially; both anatomical data and functional, such as innervation. We emphasize this need in our discussion (Lines 352-355 and Lines 405-411).

      More Minor Points:

      (1) As mentioned in the main text, the duration of pips (from pips to bursts) affects the effective bandwidth of the stimulus. I believe that the authors could give an estimate of this effective bandwidth, given what is known from bird auditory filters. I think that this estimate could be useful to compare to the effective bandwidth of the heat-call, which can now also be estimated.

      (R2. A1) Please see answer A3.4 under Joint Recommendations.

      (2) Figure 5b. Label the green and pink areas as song and heat-call spectrum. Also note that in the legend the authors say: "Green and red areas display the frequency windows related to the best hearing sensitivity of zebra finches and to heat calls, respectively". I don't think this is what they meant. I agree that 1-4 kHz is the best frequency sensitivity of zebra finches, but they probably meant green == "song frequency spectrum" and pink == "heat call spectrum". In either case, the figure and the legend need clarification.

      (R2. A2) We thank the reviewer for pointing out these issues. We have changed the figure and legend accordingly. In the meantime, we published a paper measuring the in vivo source levels of the heat whistles (Anttonen et al., Current Biology 2025), and we have carefully gone through this manuscript to incorporate those findings and adjust the text accordingly. We have therefore adjusted the analysis in Fig 3B to correct for the narrower frequency distribution of the heat whistles.

      (3) Figure 5c. Here also, I would change the song and heat-call labels to "song spectrum", "heat call spectrum". The authors would not want readers to think that they used song and heat calls in these experiments (maybe next time?). For the same reason, maybe in 5a you could add a cartoon of the oscillogram of a frequency sweep next to your speaker.

      (R2. A3) We thank the reviewer for pointing out these issues We mention the frequency sweep in the legend for panel A, but decided against including this in the figure to prevent too much clutter. We have changed the figure = legend to (new text underlined):

      Legend Fig 3. “A Setup used to measure sound-induced vibrations of eggs. A 94 dB, 0.25-to-10 kHz frequency sweep was played at the eggs to determine the vibration transfer function.”

      (4) Methods. In the description of the stimulus, the authors describe "5ms long tone bursts", but these are the tone pips in the main part of the manuscript. Use the same terms.

      (R2. A4) Thank you for catching this, we have changed this into “5ms long tone pips".

      Reviewer #3 (Public review):

      Summary

      Following recent findings that exposure to natural sounds and anthropogenic noise before hatching affects development and fitness in an altricial songbird, this study attempts to estimate the hearing capacities of zebra finch nestlings and the perception of high frequencies in that species. It also tries to estimate whether airborne sound can make zebra finch eggs vibrate, although this is not relevant to the question.

      Strength

      That prenatal sounds can affect the development of altricial birds clearly challenges the long-held assumption that altricial avian embryos cannot hear. However, there is currently no data to support that expectation. Investigating the development of hearing in songbirds is therefore important, even though technically challenging. More broadly, there is accumulating evidence that some bird species use sounds beyond their known hearing range (especially towards high frequencies), which also calls for a reassessment of avian auditory perception.

      Weaknesses

      Rather than following validated protocols, the study presents many experimental flaws and two major methodological mistakes (see below), which invalidate all results on responses to frequencyspecific tones in nestlings and those on vibration transmission to eggs, as well as largely underestimating hearing sensitivity. Accordingly, the study fails to detect a response in the majority of individuals tested with tones, including adults, and the results are overall inconsistent with previous studies in songbirds. The text throughout the preprint is also highly inaccurate, often presenting only part of the evidence or misrepresenting previous findings (both qualitatively and quantitatively; some examples are given below), which alters the conclusions.

      Conclusion and impact

      The conclusion from this study is not supported by the evidence. Even if the experiment had been performed correctly, there are well-recognised limitations and challenges of the method that likely explain the lack of response. The preprint fails to acknowledge that the method is well-known for largely underestimating hearing threshold (by 20-40dB in animals) and that it may not be suitable for a 1-gram hatchling. Unlike what is claimed throughout, including in the title, the failure to detect hearing sensitivity in this study does not invalidate all previous findings documenting the impacts of prenatal sound and noise on songbird development. The limitations of the approach and of this study are a much more parsimonious explanation. The incorrect results and interpretations, and the flawed representation of current knowledge, mean that this preprint regrettably creates more confusion than it advances the field.

      (R3. A0) We thank the reviewer for their detailed and critical assessment. We agree that establishing auditory sensitivity in very young altricial birds is technically challenging and that careful interpretation of ABR data is essential. We also appreciate the reviewer’s recognition of the importance of obtaining direct physiological data on auditory development.

      However, we respectfully disagree with the reviewer’s central claim that our methodology is flawed or that our conclusions are unsupported. Many of the concerns raised reflect misunderstandings of ABR methodology, selective interpretation of the literature, or assumptions that are not supported by empirical evidence. Below, we address the main points in turn.

      As also stated in our Provisional response, the reviewer’s critique can be distilled into four main arguments:

      (1) ABR cannot be reliably measured in very small animals.

      (2) Our stimulus design (especially 25 ms tone bursts) invalidates frequency-specific results.

      (3) ABR thresholds should be corrected to behavioral thresholds, which would alter conclusions.

      (4) Our findings are inconsistent with prior studies in songbirds.

      We address each of these below before responding point-by-point.

      (1) Suitability of ABR in small animals.

      Reviewer claim: ABR may not be suitable for very small hatchlings.

      This claim is not supported by existing evidence. ABR measures summed neural activity, and signal amplitude depends in part on the distance between neural tissue and recording electrodes. In smaller animals, this distance is reduced, which can increase signal amplitude and improve signal-to-noise ratio.

      Consistent with this, ABR has been successfully recorded in animals substantially smaller than zebra finch hatchlings, including zebrafish (Jørgensen et al., 2012), 10 mm froglets (Goutte et al., 2017) and 5 mm salamanders (Capshaw et al., 2020). It is in fact much more surprising the technique still provides robust signals even in extremely large animals such as Minke whales, where the distance between electrodes and brain is on the decimeter scale (Houser et al., 2024). We have extensive experience of recording ABRs in such small systems.

      Thus, there is no principled reason why ABR would be an invalid method to study auditory sensitivity in zebra finch hatchlings.

      (2) Stimulus design and tone duration

      Reviewer claim: Use of 25 ms tone bursts invalidates frequency-specific results.

      We agree that stimulus duration affects frequency specificity and ABR detectability. However, the reviewer’s assertion that there is a single “correct protocol” (≤5 ms) is inaccurate. In avian ABR studies, stimulus duration varies depending on experimental goals.

      Our choice of 25 ms tone bursts was intentional and necessary to accurately represent low frequencies (down to 250 Hz), ensuring sufficient cycles per stimulus in the plateau segment of 15 ms and minimizing spectral splatter (see auditory brainstem response design considerations discussed in Lauridsen et al., 2021) and our responses below.

      Key clarifications:

      a) Click-evoked ABRs form the basis of our conclusions about onset of hearing, not tone bursts.

      b) Tone bursts were used primarily to assess frequency-dependent maturation, not detect earliest sensitivity.

      c) We explicitly demonstrate that:

      - 25 ms bursts yield higher thresholds (lower sensitivity)

      - 5 ms pips yield lower thresholds and align with published ABR audiograms

      We have now:

      - Further clarified the rationale of stimulus design in the methods (Line 520-530) and added a section in the discussion (Lines 397-411).

      - Included additional comparison between burst and pip datasets (Lines 387-395).

      - Clarified that conclusions about early hearing do not depend on tone-burst data (Fig 4 and Lines 271-294).

      (3) ABR vs behavioral thresholds

      Reviewer claim: Failure to correct ABR thresholds (20–40 dB) invalidates conclusions.

      We agree that ABR thresholds typically overestimate behavioral thresholds. However, we disagree that this invalidates our conclusions.

      Importantly:

      a) We do not replace measured ABR data with corrected values, as this would be methodologically inappropriate.

      b) Instead, we:

      - Present measured ABR thresholds transparently (Fig 1-3)

      - Compare them directly to published behavioral audiograms (Fig 5)

      - Explicitly discuss the expected offset (Lines 413-422)

      In the revised manuscript we:

      a) Add a new figure (Fig. 5) compiling all published ABR and behavioral audiograms

      b) Show that:

      - ABR and behavioral audiograms have similar shapes (Fig 5, new discussion Lines 387-422)

      - Offsets are typically ~20 dB (Line 413-422)

      Crucially, even under conservative corrections:

      - Early hatchlings remain far less sensitive than adults (>54 dB SPL) to clicks.

      - Heat whistle levels remain at or below detection limits even in adults (new figure Fig 4)

      - The developmental gap (>50 dB between adults and 2 DPH hatchlings) remains decisive.

      Thus, incorporating ABR–behavioral differences does not change the central conclusion.

      (4) Consistency with prior literature in developing songbirds.

      Reviewer claim: Results contradict previous studies in developing songbirds.

      We respectfully disagree. The cited studies fall into three categories:

      (1) Behavioral studies (e.g., alarm-call responses)

      (2) Gene expression studies (e.g., ZENK activation)

      (3) Different species with different developmental trajectories

      None of these directly measure auditory sensitivity thresholds in zebra finch embryos or hatchlings.

      We emphasize:

      - Behavioral responses do not provide threshold measurements

      - Neural activation (e.g., ZENK) does not demonstrate functional perception thresholds

      - Cross-species comparisons must consider differences in developmental timing.

      We have now expanded the Discussion to explicitly address these studies and clarify how they relate to our findings (Lines 357-367). Our data are consistent with what is known about the physiology of auditory development in all birds studied so far.

      (5) Final statement

      We have revised the manuscript extensively to:

      - Clarify methodology and experimental design

      - Expand discussion of ABR limitations

      - Incorporate additional literature and comparisons

      - Correct inconsistencies in reporting

      We maintain that our central conclusions—that early zebra finch hatchlings lack detectable auditory brainstem responses and are unlikely to perceive parental heat calls at natural levels—is robust and supported by the data. We go into more detail in our point-by-point rebuttal below.

      Detailed assessment

      For brevity, only some references are included below as examples, using, when possible, those cited in the preprint (DOI is provided otherwise). A full review of all the studies supporting the points below is beyond the scope of this assessment.

      (A) Hearing experiment

      The study uses the Auditory Brainstem Response (ABR), which measures minute electrical signals transmitted to the surface of the skull from the auditory nerve and nuclei in the brainstem. ABR is widely used, especially in humans, because it is non-invasive. However, ABR is also a lot less sensitive than other methods, and requires very specific experimental precautions to reliably detect a response, especially in extremely small animals and with high-frequency sounds, as here.

      (1) Results on nestling frequency sensitivity are invalid, for failing to follow correct protocols:

      (R3. A1.1) We disagree that our protocol is invalid. There is no universal ABR protocol standard in birds. Our approach is consistent with established principles of stimulus design and is validated by:

      - Robust click-evoked responses

      - Consistent developmental trajectories

      - Agreement between pip-based ABR and behavioral audiograms

      We now clarify this explicitly. See response R3. A0 (Stimulus design and tone duration).

      The results on frequency testing in nestlings are invalid, since what might serve as a positive control did not work: in adults, no response was detected in a majority of individuals, at the core of their hearing range, with loud 95dB sounds (Figure S1), when testing frequency sensitivity with "tone burst".

      This is mostly because the study used a stimulation duration 5 times larger than the norm. It used 25ms tone bursts, when all published avian studies (in altricial or precocial birds) used stimulation of 5ms or less (when using subdermal electrodes as here; e.g., cited: Brittan-Powell et al 2004; not cited: Brittan-Powell et al 2002 (doi: 10.1121/1.1494807), Henry & Lucas 2008 (doi: 10.1016/j.anbehav.2008.08.003)). Long stimulations do not make sense and are indeed known to interfere with the detection of an ABR response, especially at high frequencies, as, for example, explicitly tested and stated in Lauridsen et al 2021 (cited).

      (R3. A1.2) ABR with long-duration stimuli were shown previously to work perfectly well in birds and do not interfere with the detection of an ABR response. Longer stimuli have been used, e.g. in the following bird papers:

      - Amin et al., J Neurophysiol 2007 (cited) on zebra finch ABR: 20 ms tone bursts

      - Korneeva et al. (2006), evoked responses from field L in flycatcher (cited by the reviewer): 20 ms

      - Saunders et al. (1973) (cited): 60 ms tone bursts.

      - Larsen ON, Wahlberg M, Christensen-Dalsgaard (2020) Amphibious hearing in a diving bird, the great cormorant (Phalacrocorax carbo sinensis), J Exp Biol, doi:10.1242/jeb.217265: 25 ms tone bursts

      Human ABR has been measured even with long-duration speech signals (duration 40 ms and longer). See for example Binkhamis et al., Ear and Hearing 40: 659-670, 2019.

      Furthermore, the reviewer unfortunately misunderstood some aspects in Lauridsen et al., which tested a specifical method exploiting neural phase locking, and showed that this method has a low-frequency bias because neural phase locking decreases at high frequencies.

      Also, Lauridsen et al clearly show the reason for using 25 ms bursts: because we aimed to measure frequencies down to 250 Hz, we need to have a sufficient number of cycles (3) in the plateau segment to represent the frequency adequately (Lauridsen et al. fig 1). A 15 ms plateau contains 3 cycles, plus 5 ms rise/fall time (1 cycle) equals 25 ms. We have clarified this in our methods (L520-525) and added a section in the discussion (L397-411).

      Thus, long-duration stimulations make sense and have been used before successfully. See also response R3.A0 (Stimulus design and tone duration).

      Adult response was then re-tested with a correct 5ms tone duration ("tone-pip"), which showed that, for the few individuals that responded to 25ms tones, thresholds were abnormally high (c.a. by 30dB; Figure 2C).

      Yet, no nestlings were retested with a correct protocol. There is therefore no valid data to support any conclusion on nestling frequency hearing. Under these circumstances, the fact that some nestlings showed a response to 25ms tones from day 8 would argue against them having very low sensitivity to sound.

      (R3. A1.3) Please see answer A3.1 under Joint Recommendations and R3.A0 (Stimulus design and tone duration).

      (2) Responses to clicks underestimate hearing onset by several days:

      Without any valid nestling responses to tones (see # 1), establishing the onset of hearing is not possible based on responses to clicks only, since responses to clicks occur at least 4 days after responses to tones during development (Saunders et al, 1973). Here, 60% of 4-day-old individuals responding to clicks means most would have responded to tones at and before 2 days post-hatch, had the experiment been done correctly.

      (R3. A2.1) We disagree that clicks necessarily underestimate onset.

      Clicks are broadband stimuli that:

      - Recruit large neural populations

      - Are commonly used to detect early auditory responses

      The cited delay between tone and click responses reflects stimulus energy differences, not an inherent limitation of clicks. We have clarified this in the revision and softened language to refer to “no detectable ABR response” rather than absolute deafness.

      The report that Saunders could only see responses to clicks later than to tones only reflects that he used click amplitudes that were insufficiently high. The ABR responses reported were to extremely intense tones (110 dB SPL) of long duration (60 ms).

      If Saunders had used clicks (duration 60 µs) with comparable sound energy, they would have had been very difficult to produce. He should have used clicks with an amplitude 1000 times (60 dB) higher than the tones to produce the same sound energy. This would have been clicks at 170 dB SPL, equivalent to the sound at the mouth of a medium-sized military cannon. Applying this pressure would not be a recommendable method for hearing assessment, but instead lead to irreversible hearing damage.

      In budgerigars, hearing onset occurs before 5 days post hatch, since responses to both clicks and tones were detectable at the first age tested at 5dph (Brittan-Powell et al, 2004).

      (R3. A2.2) This is not how we interpret the cited paper. They state that ‘Responses were first obtained from 1-week-old at high stimulation, and their click responses (Fig. 1) show no wave 1 peak at 6 days post-hatch. Also, their conclusion (p 3101) states that ‘budgerigars probably cannot hear at hatching’.

      (3) Experimental parameters chosen lower ABR detectability, specifically in younger birds: Very fast stimulus repetition rate inhibits the ABR response, especially in young:

      (a) The stimulus presentation rate (25 stim/ sec) is 6 times faster than zebra finch heat-calls, and 5 to 25 times faster than most previous studies in young birds (e.g., cited: Saunders et al 1973, 1974: 1 stim/sec or less; Katayama 1985: 3.3 clicks/sec; Brittan-Powell et al 2004: 4 stim/sec).

      Faster rates saturate the neurons and accordingly are known to decrease ABR amplitude and increase ABR latency, especially in younger animals with an immature nervous system.

      In birds, this occurs especially in the range from 5 to 30 stim/sec (e.g., cited: Saunder et al 1973, Brittan-Powell et al 2004). Values here with 25 rather than 1-4 stim/min are therefore underestimating true sensitivity.

      (R3. A3a) Please see answer A3.3 under Joint Recommendations.

      (b) Averaging over only 400 measures is insufficient to reliably detect weak ABR signals: The study uses 2 to 3 times fewer measures per stimulation type than the recommended value of 1,000 (e.g., Brittan-Powell et al 2002, 2024; Henry & Lucas 2008). This specifically affects the detection of weak signals, as in small hatchlings with tiny brains (adult zebra finches are 12-14g).

      (R3. A3b) Please see answer A3.2 under Joint Recommendations.

      (c) Body temperature is not specified and strongly affects the ABR:

      Controlling the body temperature of hatchlings of 1-4 grams (with a temperature probe under a 5mm-wide wing) would be very challenging. Low body temperature entirely eliminates the ABR, and even slight deviance from optimal temperature strongly increases wave latency and decreases wave amplitude (e.g., cited: Katayama 1985).

      (R3. A3c) Please see answer A2 under Joint Recommendations.

      (d) Other essential information is missing on parameters known to affect the ABR: This includes i) the weight of the animals,

      (R3. A3d-i) These important details have now been added in Table S6.

      (ii) whether and how the response signal was amplified and filtered,

      (R3. A3d-ii) Signal was amplified 500 times (74 dB). We have included these important details to the methods (Line 508).

      (iii) how the automatised S/N>2 criteria compared to visual assessment for wave detection,

      (R3. A3d-iii) There is no universally accepted/fitting protocol for performing ABR recordings in various animals, we decided to perform both visual and automated criteria detection of thresholds. The automated criterion in our experience is more strict approach than visual detection of thresholds, because using an automated criteria for threshold detection will remove potential experimenter bias from the results. We added this in our methods (Lines 583-585).

      (iv) what measures were taken to allow the correct placement of electrodes on hatchlings less than 5 grams.

      (R3. A3d-iv) We have placed electrodes in much smaller animals than 5 grams, and the common landmarks (ear opening, midline of skull) could easily be identified in the hatchlings.

      (4) Results in adults largely underestimate sensitivity at high frequencies, and are not the correct reference point:

      (a) Thresholds measured here at high frequencies for adults (using the correct stimulus duration, only done on adults) are 10-30dB higher than in all 3 other published ABR studies in adult zebra finches (cited: Zevin et al 2004; Amin et al 2007; not cited: Noirot et al 2011 (10.1121/1.3578452)), for both 4 and 6 kHz tone pips.

      (b) The underlying assumption used throughout the preprint that hearing must be adult-like to be functional in nestlings does not make sense. Slower and smaller neural responses are characteristic of immature systems, but it does not mean signals are not being perceived.

      (R3. A4) We acknowledge variation across studies and now include a comprehensive comparison of all zebra finch audiograms (new Fig. 5 and discussion Lines 387-422).

      Importantly:

      - Our pip-based audiogram aligns with previous ABR studies

      - Differences in high-frequency sensitivity likely reflect methodological variation or population differences.

      Our conclusions rely on relative developmental changes, not absolute thresholds.

      (5) Failure to account for ABR underestimation leads to false conclusions:

      (a) Whether the ABR method is suitable to assess hearing in very small hatchlings is unknown. No previous avian study has used ABR before 5 days post-hatch, and all have used larger bird species than the zebra finch.

      (R3. A5a) As stated above (R3.A0), small animals should give better signals, and we have been able to measure ABR in much smaller animals previously.

      (b) Even when performed correctly on large enough animals, the ABR systematically underestimates actual auditory sensitivity by 20-40 dB, especially at high frequencies, compared to behavioural responses (e.g., none cited: Brittan-Powell et al 2002, Henry & Lucas 2008, Noirot et al 2011). Against common practice, the preprint fails to account for this, leading to wrong interpretations.

      (R3. A5a) See our answer to R3.A0 above.

      For example, in Figure 1G (comparing to heat call levels), actual hearing thresholds would be 3040dB below those displayed. In addition, the "heat whistle" level displayed here (from the same authors) is 15dB lower than their second measure that they do not mention, and than measures obtained by others (unpublished data). When these two corrections are made - or even just the first one - the conclusion that heat-call sound levels are below the zebra finch hearing threshold does not hold.

      (R3. A5a) Our conclusion that heat whistles are unlikely to be perceived does not rely on a single dataset or method, but on the convergence of three independent constraints: (i) signal amplitude, (ii) adult auditory sensitivity, and (iii) developmental immaturity of the auditory system.

      First, heat whistles are low-amplitude signals. Our in vivo measurements show levels of ~33 dB re 20 µPa at 10 cm and ~14 dB at 1 m (Anttonen et al., 2025, Curr Biol). Even allowing for uncertainty in near-field estimation, a conservative upper bound at very close range (<5 cm) is ~40 dB SPL.

      Second, the most sensitive available measure—behavioral audiograms—places adult zebra finch thresholds at ~40 dB SPL at ~6 kHz (Okanoya and Dooling, 1987, J Comp Psychol), increasing steeply toward higher frequencies. Thus, even under optimal conditions, heat whistles fall at or below the detection threshold of adults, and only potentially at very close range.

      Third, auditory sensitivity in early development is substantially reduced. Our ABR data show a ≥40–60 dB decrease in sensitivity in hatchlings relative to adults for click stimuli, which provide the most favorable conditions for eliciting responses. Because frequency-specific sensitivity develops later, thresholds at 6–8 kHz are expected to be even higher in hatchlings and embryos.

      Taken together, these constraints define a narrow and unfavorable detection window: a low-amplitude, high-frequency signal positioned at the edge of adult sensitivity, combined with a large developmental decrease in auditory sensitivity. Under these conditions, it is unlikely that heat whistles are detectable by hatchlings or embryos.

      Importantly, this conclusion does not depend on precise correction factors between ABR and behavioral thresholds. Even when considering the most sensitive behavioral data and conservative estimates of sound level, the signal remains at or below the limits of detection in adults, and far below expected sensitivity in early developmental stages.

      We have included this argument more clearly in our discussion (Line 271-367), illustrated by new Fig. 5.

      (c) Rather than making appropriate corrections, the preprint uses a reference in humans (L180), where ABR is measured using a much more powerful method (multi-array EEG) than in animals, and from a larger brain. The shift of "10-20dB" obtained in humans is not applicable to animals.

      (R3. A5c) Again our conclusions rely on relative developmental changes, not absolute thresholds. The clinical practice in humans to measure ABR is with 4 electrodes, not a multi-electrode EEG array.

      Animal studies where ABR audiograms have been compared directly to psychophysical audiograms show differences of around 20 dB. For example in Brittain powell et al 2002, audiogram comparisons between behavioral and ABR in budgerigars were made within the same lab, same animal population and by the same people, leading to 20 dB difference. Our discussion includes a new paragraph on this topic (Lines 413-422).

      (6) Results are inconsistent with previous findings in developing songbirds:

      (R3. A6) We now explicitly discuss all cited studies. Key points:

      - Early behavioral responses do not imply high-frequency sensitivity

      - Studies in other species do not directly translate to zebra finches

      - None of the cited work provides direct measures of auditory thresholds in embryos

      As expected from all of the above, results and conclusions in the preprint are inconsistent with findings in other songbirds, which, using other methods, show for example, auditory sensitivity in: a) zebra finch embryos, in response to song vs silence (not cited: Rivera et al 2018, doi: 10.1097/WNR.0000000000001187)

      (R3. A6a) We thank the reviewer for pointing out this study. We agree that the question of auditory responsiveness in embryos is important, and that a range of approaches have been used to address it. However, the study cited (Rivera et al., 2018) does not directly measure auditory sensitivity, but instead infers auditory processing from differences between treatment groups exposed to different acoustic conditions. As such, it is not directly comparable to physiological measures of hearing sensitivity, such as ABR or behavioral thresholds.

      In addition, interpretation of these results is complicated by limited characterization of the acoustic environment and differences in experimental handling between groups, which may introduce confounding factors unrelated to auditory perception. Given these considerations, and because our study focuses specifically on quantifying auditory sensitivity using established physiological methods, we have chosen not to include a detailed discussion of this work.

      (b) flycatcher hatchlings at 2-3d post hatch (first age tested), across a wide range of frequencies (0.3 to 5kHz), at low to moderate sound levels (45-65dB) (cited: Aleksandrov and Dmitrieva 1992, not cited: Korneeva et al 2006 (10.1134/S0022093006060056)).

      (R3. A6b) Korneeva et al. (2006) and Aleksandrov and Dmitrieva (1992) report evoked responses in very young flycatcher hatchlings across a broad frequency range. Notably, these measurements were obtained using more invasive recording approaches (e.g., implanted electrodes in Field L) in unanesthetized birds, which are known to yield lower thresholds compared to far-field ABR recordings under anesthesia. These methodological differences likely account for part of the higher sensitivity reported.

      Importantly, even in flycatchers, auditory sensitivity shows substantial postnatal improvement: thresholds decrease by up to ~40 dB over the first days after hatching, and the upper frequency limit expands from ~4 to ~7 kHz. Thus, while absolute sensitivity may differ across species and methods, the overall developmental trajectory—gradual improvement in sensitivity and progressive extension toward higher frequencies—is consistent with our findings and with broader patterns reported in songbirds.

      We have added this paper in our discussion (Lines 362-365).

      (c) songbird nestlings at 2-6d post hatch, which discriminate and behaviourally respond to relevant parental calls or even complex songs. This level of discrimination requires good hearing across frequencies (e.g., not cited: Korneeva et al 2006; Schroeder & Podos 2023 (doi: 10.1016/j.anbehav.2023.06.015)).

      (R3. A6c) The species mentioned are different species from our study species. In the Pied flycatchers (Korneeva et al. 2006) experimental conditions were different: recordings were made from unanesthetized nestlings with implanted electrodes directly in the brain (field L), so likely with better SNR. The audiograms show a 10 dB SPL threshold after day 11, so the species may be considerably more sensitive than the zebra finch. The swamp sparrows in the Schroeder and Podos (2023) behavioral study were exposed for 4 days starting at 4-7 days post-hatch, so the study does not address embryonal hearing.

      (d) zebra finch nestlings at 13d post-hatch, which show adult-like processing of songs in the auditory cortex (CNM) (Schroeder & Remage-Healey 2021, doi: 10.1002/dneu.22802).

      (R3. A6d) This study does not conflict with our data. Even though sensitivity is lower at 10 days than in adults, cortical processing could still be ‘adult-like’.

      (e) zebra finch juveniles, which are able to perceive and learn song syllables at 5-7kHz (fundamental frequency) with very similar acoustic properties to heat calls, and also produced during inspiration (Goller & Daley 2001, doi: 10.1098/rspb.2001.1805).

      (R3. A6e) This result is not in conflict with our data. First, the onset of song learning occurs earliest at 20 DPH as discussed in the paper and our work demonstrates that click-evoked ABR thresholds are adult-like at 20 DPH. In the cited paper, the tutoring experiments were initiated at 35 DPH so the auditory system of the studied juveniles is mature.

      Second, even though Goller & Daley 2001 do not report the source level of the specific syllable or the playback sound pressure levels, the source level of the inspiratory notes is comparable to other syllables, and thus around ~67 dB and ~34 dB louder than heat whistles.

      NONE of these results - which contradict results and claims in the preprint - are mentioned.

      Instead, the preprint focuses on very slow-developing species (parrots and owls), which take 2-4 times longer than songbirds to fledge (cited: Brittan-Powell et al 2004; Köppl & Nickel 2007; Kraemer et al 2017).

      (R3. A6f) We have included papers in our discussion that reflect the known data (to our best knowledge) on the developmental neurophysiology and neuroanatomy of the auditory system and not proxies thereof.

      (7) Results in figures are misreported in the text, and conclusions in the abstract and headers are not supported by the data:

      For example:

      (a) The data on Figure 1E shows that at 4 days old, 8 out of 13 nestlings (60%) responded to clicks, but the text says only 5/13 responded (L89).

      (R3. A7a1) We apologize for this typo. Corrected.

      When 60% (4dph) and 90% (6dph) of individuals responded, the correct term would be that "most animals", rather than "some animals" responded (L89).

      (R3. A7a2) We have rephrased this sentence into “observable in most animals during” as suggested.

      Saying that ABR to loud sound appeared "in the majority only after one week" (L93) is also incorrect, given the data.

      (R3. A7a3) We have rephrased this sentence into: “Thus, sounds at loud, yet physiologically relevant SPLs do not evoke ABRs in the first days after hatching, but do so in all animals at 8 DPH.” (Lines 97-99).

      It follows that the title of the paragraph is also erroneous.

      (R3. A7a4) The paragraph title supports our conclusions and we will keep it.

      (b) The hearing threshold is underestimated by 40dB at 6 and 8Kz on Fig 2C, not by "10-20dB" as reported in the text (L178).

      (R3. A7b) We have changed the title of this section and moved the last sentence to the discussion to remove the focus on heat whistles. We added a paragraph in the discussion to specifically address the difference between ABR and behaviorally measured audiograms (Lines 413-422).

      (B) Egg vibration experiment

      (8) Using airborne sound to vibrate eggs is biologically irrelevant:

      (R3. A8.1) We agree that parental contact could influence vibration transmission.

      However, (1) prior studies assume airborne sound transmission, and (2) our experiment tests this assumption directly. We now clarify this scope (Lines 328-342 and Figure 4) and discuss contact-based transmission as a potential future direction.

      The measurement of airborne sound levels to vibrate eggs misunderstands bone conduction hearing and is not biologically meaningful: zebra finch parents are in direct contact with the eggs when producing heat calls during incubation, not hovering in front of the nest. This misunderstanding affects all extrapolations from this study to findings in studies on prenatal communication.

      (R3. A8.2) The definition of bone conduction is the response to sound that is not mediated by a functional middle ear, but through the skull. In the earlier study, the eggs were stimulated by sound from a headphone, so that is the reason for using the same stimulation here. See also joint response A3.5 above.

      (C) Misrepresentation of current knowledge

      (9) Values from published papers are misreported, which reverses the conclusions:

      (R3. A9) We thank the reviewer for identifying inconsistencies and have:

      - Corrected heat whistle frequency ranges consistently through our paper

      - Added a comprehensive comparison figure gathering all available audiograms (Fig 5)

      - Expanded discussion of high-frequency hearing.

      These revisions do not alter our conclusions.

      Most critical examples:

      (a) Preprint: "Zebra finch most sensitive hearing range of 1-to-4 kHz (Amin et al., 2007; Okanoya and Dooling, 1987; Yeh et al., 2023)" (L173).

      Actual values in the studies cited are:

      1-to-7kHz, in Amin et al 2007 (threshold [=50dB with ABR] is the same at 7kHz and 1KHz).

      1-to-6 kHz, in Okanoya and Dooling (the threshold [=30dB with behaviour] is actually lower at 6kHz than at 1KHz).

      1-to-7kHz, in Yeh et al (threshold [=35-38dB with behaviour] is the same at 7kHz and 1KHz).

      (R3. A9a.1) In this sentence presenting our results (“sensitive hearing range of 1-to-4 kHz”) we originally wrote that “these are consistent with the following papers (Amin et al., 2007; Okanoya and Dooling, 1987; Yeh et al., 2023)". This latter part was left out during the writing process. This explains the different numbers. We apologize for this mistake.

      To avoid confusion, in our revision we have placed all ABR curves together into new Fig 5 and have included a new paragraph to discuss the differences (Lines 387-422).

      Note that zebra finch nestlings' begging calls peaking at 6kHz (Elie & Theunissen 2015, doi: 10.1007/s10071-015-0933-6), would fall 2kHz above the parents' best hearing range if it were only up to 4kHz.

      (R3. A9a.2) Of course that is possible. However this representation is incorrect because begging calls are harmonic sounds with a fundamental frequency around 500 Hz and formant at 6 kHz. Begging calls thus contain lots of energy at frequencies below 6 kHz, while the heat whistles do not. The peak frequency of heat whistles is also their lowest frequency component.

      (b) The preprint incorrectly states throughout (e.g., L139, L163, L248) that heat-calls are 7-10kHz, when the actual value is 6-10kHz in the paper cited (Katsis et al, 2018).

      (R3. A9b) The authors in Katsis et al. 2018 provided a range of 6-10 kHz estimated from the spectrogram without any further specification of methods. In another manuscript, we have quantified the heat whistle frequency (Anttonen et al Curr Biol https://doi.org/10.1016/j.cub.2025.08.054) to be 6.8 ± 0.6 kHz. We have changed this accordingly throughout our manuscript.

      (c) Using the correct values from these studies, and heat-calls at 45 dB SPL (as measured by others (unpublished data), or as measured by the authors themselves, but which is not reported here (Anttonen et al 2025), the correct conclusion is that heat calls fall within the known zebra finch hearing range.

      (R3. A9c) Please see our answer R3.A5a. We have included this argument more clearly in our discussion (Lines 271-355), illustrated by new Fig. 4.

      (10) Published evidence towards high-frequency hearing, including in early development, is systematically omitted:

      (a) Other studies showing birds use high frequencies above the known avian hearing range are ignored. This includes oilbirds (7-23kHz; Brinklov et al 2017; by 1 of the preprint authors, doi: 10.1098/rsos.170255) and hummingbirds (10-20kHz; Duque et al 2020, doi: 10.1126/sciadv.abb9393), and in a lesser extreme, zebra finches' inspiratory song syllables at 57kHz (Goller & Dalley, 2001).

      (R3. A10a) We agree that some bird species produce or use acoustic signals extending into high frequencies. However, signal production is not evidence of perceptual sensitivity. Many animals, including birds and mammals, produce signals that contain harmonic or broadband components extending beyond their most sensitive hearing range without implying functional detection at those frequencies.

      The cited examples (oilbirds, hummingbirds, inspiratory song syllables in zebra finches) concern signal production or ecological specializations in different species, not measured auditory sensitivity in zebra finches, and particularly not during early development. As such, they do not provide evidence that zebra finches—adults or embryos—can detect low-amplitude, narrowband signals in the 6–7 kHz range.

      Our study explicitly addresses auditory sensitivity using physiological measurements, which is the relevant metric for evaluating detectability.

      (b) The discussion of anatomical development (L228-241) completely omits the well-known fact that the avian basilar papilla develops from high to low frequencies (i.e., base to apex), which - as many have pointed out - is opposite to the low-to-high development of sensitivity (e.g., cited: Cohen & Fermin 1978; Caus Capdevila et al 2021).

      (R3. A10b) We agree that the avian basilar papilla develops from base to apex (high to low frequency). We have now added a sentence in the Discussion to acknowledge this (Lines 406411).

      Importantly, morphological development does not directly translate to functional sensitivity. Functional hearing depends critically on factors such as hair cell innervation, synaptic maturation, and central auditory processing, which are known to develop over time.

      Our data show a low-to-high frequency progression in functional sensitivity, consistent with previous physiological studies. This apparent mismatch between anatomical gradients and functional onset has been noted in other systems and likely reflects the later maturation of neural encoding rather than hair cell differentiation per se. We now clarify this distinction in the revised manuscript (Lines 406-411).

      (c) High frequency hearing in songbirds at hatching is several orders of magnitude better than in chickens and ducks at the same age, even though songbirds are altricial (e.g., at 4kHz, flycatcher: 47dB, chicken-duck: 90dB; at 5kHz, flycatcher: 65dB, chicken-duck: 115dB; Korneeva et al 2006, Saunders et al 1974). That is because Galliformes are low-frequency specialists, according to both anatomical and ecological evidence, with calls peaking at 0.8 to 1.2kHz rather than 2-6kHz in songbirds. It is incorrect to conclude that altricial embryos cannot perceive high frequencies because low-frequency specialist precocial birds do not (L250;261).

      (R3. A10c) We agree that species differ in their auditory ecology and frequency specialization, and we do not claim that all altricial birds share identical developmental trajectories.

      However, the cited comparisons involve different species, methodologies, and developmental timelines, which limits their direct comparability. In particular:

      Developmental staging is not directly comparable across species using days post-hatch alone.

      - Different methods (e.g., invasive recordings vs. ABR vs behavioural assays) yield systematically different thresholds.

      - Ecological specialization (e.g., low-frequency vs. broadband species) influences adult audiograms and likely developmental trajectories.

      We have revised the Discussion to explicitly acknowledge these limitations and to avoid overgeneralization across species. Importantly, our conclusions are based on within-species comparisons (adult vs. hatchling zebra finches) combined with measured signal levels of heat whistles. These constraints are sufficient to evaluate detectability without relying on cross-species extrapolation.

      (11) Incorrect statements do not reflect findings from the references cited For example:

      (a) "in altricial bird species hearing typically starts after hatching" (L12, in abstract), "with little to no functional hearing during embryonic stages (Woolley, 2017)." (L33).

      There is no evidence, in any species, to support these statements. This is only a - commonly repeated - assumption, not actually based on any data. On the contrary, the extremely limited evidence to date shows the opposite, with zebra finch embryos showing ZENK activation in the auditory cortex in response to song playback (Rivera et al, 2018, not cited).

      The book chapter cited (Woolley 2017) acknowledges this lack of evidence, and, in the context of song learning, provides as only references (prior to 2018), 2 studies showing that songbirds do not develop a normal song if the song tutor is removed before 10d post-hatch. That nestlings cannot memorise (to later reproduce) complex signals heard before d10 does not mean that they are deaf to any sound before day 10.

      Studies showing hearing in young songbird nestlings (see point 6 above) also contradict these statements.

      (R3. A11a) We agree that the precise onset of hearing in altricial embryos is not well established. We have therefore revised the wording in the Abstract and Introduction to avoid categorical statements and instead reflect the limited available evidence (Lines 13-16 and 33-37).

      Our data provide direct physiological measurements showing extremely low sensitivity immediately after hatching, which constrains the likelihood of functional hearing in earlier embryonic stages.

      Regarding the cited ZENK study, we note that immediate early gene expression indicates neural activation but does not provide a measure of auditory sensitivity or detection thresholds. As such, it cannot be directly compared to physiological or behavioral measures of hearing.

      (b) "Zebra finch embryos supposedly are epigenetically guided to adapt to high temperatures by their parents high-frequency "heat calls" " (L36 and L135).

      This is an extremely vague and meaningless description of these results, which cannot be assessed by readers, even though these results are presented as a major justification for the present study. Rather than giving an interpretation of what "supposedly" may occur, it would be appropriate to simply synthesize the empirical evidence provided in these papers. They showed that embryonic exposure to heat-calls, as opposed to control contact calls, alters a suite of physiological and behavioural traits in nestlings, including how growth and cellular physiology respond to high temperatures. This also leads to carry-over effects on song learning and reproductive fitness in adulthood.

      (R3. A11b) We thank the reviewer for raising this point. In the revised manuscript, we have replaced the previous phrasing with a more precise and neutral summary of what these studies report, namely that embryonic exposure to heat-call playbacks has been associated with differences in physiological and behavioral traits.

      Our study, however, addresses a distinct question—whether such acoustic signals are detectable by embryos given known constraints on signal amplitude and auditory sensitivity. The cited studies do not directly quantify auditory perception or the physical sound environment experienced by embryos. As a result, they do not provide a direct test of the sensory mechanism required for acoustic communication. A detailed evaluation of experimental design and interpretation in those studies is beyond the scope of the present manuscript, and we therefore limit our discussion to assessing the biophysical and physiological plausibility of the proposed mechanism.

      (c) "The acoustic communication in precocial mallard ducks depends specifically on the lowfrequency auditory sensitivity of the embryo (Gottlieb, 1975)" (L253)

      The study cited (Gottlieb, 1975) demonstrates exactly the opposite of this statement: it shows that duckling embryos, not only perceive high frequency sounds (relative to the species frequency range), but also NEED this exposure to display normal audition and behaviour post-hatch. Specifically, it shows that duckling embryos deprived of exposure to their own high-frequency calls (at 2 kHz), failed to identify maternal calls post-hatch because of their abnormal insensitivity to higher frequencies, which was later confirmed by directly testing their auditory perception of tones (Dimitrieva & Gottlieb, 1994).

      (R3. A11c) We thank the reviewer for this clarification and have revised the relevant text. Our intention was to highlight that embryonic auditory experience can shape postnatal behavior, not to imply strict low-frequency limitation. Therefore we already included the actual frequency in the original sentence. We have removed the non-descriptive term “low-frequency” (Lines 330-332).

      (12) Considering all of the mistakes and distortions highlighted above, it would be very premature to conclude, based on these results and statements, that altricial avian embryos are not sensitive to sound. This study provides no actual scientific ground to support this conclusion.

      (R3. A12) We respectfully disagree with the reviewer’s conclusion.

      Our study does not make a general claim that altricial embryos are incapable of perceiving sound. Rather, we evaluate a specific hypothesis: whether zebra finch embryos and hatchlings can detect sound and parental heat whistles.

      Our conclusions are based on the convergence of:

      (1) Measured low sound pressure levels of heat whistles,

      (2) Established adult auditory thresholds (behavioral data),

      (3) A large developmental decrease in auditory sensitivity demonstrated by our ABR measurements.

      Even under conservative assumptions, these constraints place heat whistles at or below adult detection thresholds and far below expected sensitivity in hatchlings and embryos.

      Thus, our conclusion is not based on absence of evidence, but on quantitative constraints that make detection unlikely under biologically realistic conditions.

      Recommendations for the authors:

      Joint recommendations:

      In response to the joint recommendations, we have:

      - Expanded methodological transparency (temperature, electrode setup, stimulus parameters),

      - Added new data (Fig S3) and figures (Fig 4 and 5),

      - Clarified ABR limitations and interpretation,

      - Strengthened the separation between measured results and interpretation,

      - Reframed conclusions to avoid overstatement.

      These revisions leave the central two conclusions unchanged: 1) zebra finch hatchlings and embryos are functionally deaf, and 2) under biologically realistic conditions, heat whistles are unlikely to be detectable by zebra finch hatchlings or embryos.

      (A) Reviewers 1 and 2:

      Much of the reviewer discourse revolved around providing clarifications of methodology for measuring the ABR and caveats for interpretation. There was near consensus with reviewers 1 and 2 on issues related to the ABR, which should be addressed.

      We appreciate the reviewers’ consensus that the main conclusions are supported, while requesting clarification of methodological details and interpretation of ABR measurements.

      (1) Please address all of the issues raised by reviewers 1 and 2 above.

      (A1) All points raised by Reviewers 1 and 2 have been addressed in detail in our point-by-point rebuttal below. In addition, we have revised the manuscript to improve clarity, added new figures (Fig. 4, 5), and substantially expanded the Discussion with eight new paragraphs to better contextualize our findings.

      (2) Please also

      - clarify all aspects of experimental details of the ABR that were missing, including temperature control (estimate body and ambient temperatures during ABR recordings,

      - please address the possibility of hypothermia of hatchlings that could have reduced ABR responses,

      - and potential local head cooling due to surgical exposure and its likely effect on highfrequency response depression).

      (A2) In our revision, we have expanded the Methods section (L479-485) and added the following new data:

      Body and ambient temperature/hypothermia

      We have now included the body temperatures during ABR recordings in new table S6. These data show that:

      (1) Body temperature was stable throughout recordings,

      (2) Temperatures were within the physiological range,

      (3) Conditions were consistent across all age groups.

      Importantly, even the youngest hatchlings maintained stable temperatures and showed no indication of hypothermia. Therefore, differences in ABR responses cannot be attributed to temperature effects.

      Potential cooling due to surgical exposure

      This concern does not apply to our experiments. We used subdermal needle electrodes, which do not require surgical exposure. Therefore, no local cooling of the head occurred, and no tissue exposure could affect high-frequency sensitivity. We have added a clarifying sentence in the Methods section to explicitly state this (Line 501-503).

      (B) Reviewer 3 also had additional requests for clarification that should also be addressed:

      (3.1) Stimulus duration too long: The study used 25 ms tone bursts instead of the standard {less than or equal to} 5 ms "pips." Could this prevent reliable ABR detection, especially at high frequencies?

      (A3.1) We agree that stimulus duration affects ABR characteristics and now clarify our rationale in the manuscript.

      - The 25 ms tone bursts were deliberately chosen to ensure sufficient cycle representation at low frequencies (down to 250 Hz) and to avoid frequency splatter.

      - Using a constant duration across frequencies ensures comparable stimulus energy.

      Importantly:

      - The 25 ms data yield audiogram shapes consistent with both click responses (Fig 2C) and published behavioral data (new Fig 5).

      - To address potential high-frequency limitations, we included a dataset using 5 ms tone pips, which produced thresholds consistent with published ABR studies (new Fig 5).

      Thus, both stimulus types support the same conclusion: a gradual maturation of hearing sensitivity from low to high frequencies. We have expanded the Discussion with three paragraphs to clarify these methodological trade-offs (Lines 387-422).

      (3.2) Were 400 sweeps enough averaging? Might a signal appear at 1000 or more?

      (A3.2) Signal-to-noise ratio improves with the square root of the number of averages. Increasing from 400 to 1000 sweeps would therefore reduce thresholds by at most ~4 dB. This magnitude is small relative to the >54 dB developmental differences observed, and the large gap between signal levels and detection thresholds. Thus, increasing sweep number would not alter the conclusions. We now clarify this explicitly in the Methods (L530).

      (3.3) Was the repetition rate too high? How does the stimulus presentation affect the ABR? Might a signal have emerged with 1-4 per second?

      (A3.3) We have clarified stimulus presentation rates in the revised manuscript:

      - Clicks were presented at 25 Hz. Control measurements (now included as Supplementary Fig. S3) show no effect of this rate on ABR amplitude or threshold.

      - Tone bursts and pips were presented at ~3 Hz, consistent with commonly used rates that avoid neural adaptation. We apologize for leaving this out in our original submission.

      We now explicitly describe these parameters and their rationale in the Methods (Lines 543-552).

      (3.4) If possible, provide an estimate of the effective bandwidth of the tone pips and compare it with the bandwidth of the parental heat-whistles.

      (A3.4) We agree that stimulus bandwidth differs between tone pips and heat whistles, and that broader signals may stimulate multiple auditory filters. Shorter stimuli (e.g., 5 ms pips) have broader bandwidth and may stimulate multiple filters—particularly at low frequencies—potentially lowering thresholds, whereas longer stimuli (25 ms bursts) are more frequency-specific and may yield higher thresholds. At higher frequencies (including the heat whistle range), this effect is expected to be smaller.

      However, quantitative correction is currently not possible due to a lack of species-specific data on auditory tuning curves in zebra finches. The only available avian data (budgerigar; Saunders et al., 1979) suggest auditory filter bandwidths (Q10 dB, i.e., the bandwidth 10 dB below the peak divided by peak frequency) of ~1.4 at low frequencies and ~1000 Hz at higher frequencies, but how multifilter stimulation affects thresholds is unknown and likely species- and frequency-dependent.

      Given these uncertainties, direct comparison between tone stimuli and heat whistles requires strong assumptions. We therefore suggest that future studies should measure responses to natural heat whistles directly.

      (3.5) Egg Vibration Experiment. Address the possibility that if a parent were physically lying on top of an egg and generated a heat call, parental body vibration could significantly communicate some perceptual vibrotactile signal to the egg. Reviewer 3 raised the possibility that the experiments in this paper tested the extent to which an auditory input can vibrate the egg - what if a vocalizing bird was on the egg?

      (A3.5) We agree that embryos may receive multiple types of sensory input from parents, including direct mechanical cues.

      However, our experiment specifically tests the hypothesis proposed in prior work: that airborne sound (heat whistles) induces egg vibrations sufficient for perception. Our findings show that airborne sound-induced vibrations are orders of magnitude below known vibrotactile sensitivity thresholds.

      Regarding parental contact:

      - Heat whistles are produced by an aerodynamic whistle mechanism, not tissue vibration (Anttonen et al., Curr Biol 2025), meaning most respiratory energy is radiated as sound rather than dissipated as heat/vibration in the body.

      - A parent sitting on the egg would attenuate airborne sound transmission, not amplify it.

      We now clarify in the Discussion that other cues (e.g., respiration, direct contact, temperature) may exist and need to be included in new experiments (Lines 349-352). Even so, these are distinct mechanisms and were not the hypothesis tested in prior playback studies.

      In our paper, we will not add a detailed discussion of these prior papers as this is outside the scope of this paper. Instead, we added a paragraph what would be a constructive way forward (Lines 349-355). Hopefully somebody in the community will have the good fortune to secure research funding to continue this benchmarking work.

      (4) Finally, all reviewers agreed that some more context on the ABR and its relationship to functional hearing could be provided, with less direct focus on the heat-call experiments.

      (A4) We agree and in the revision discussion have compiled all published ABR and behavioral audiograms (Fig. 5) and added new paragraphs on functional hearing (Lines 261-269), and ABR vs behavioural audiograms (Lines 413-422).

      Furthermore, to remove focus on the heat whistles, we have moved all heat-whistle-specific interpretation out of the Results into a single, focused Discussion section (Line 271-355).

      The behavioral studies mentioned below show auditory responses (e.g., begging suppression). However, these behaviors are tested between ~5–10 days post-hatch (consistent with our findings), in different species, and do not provide quantitative sensitivity thresholds, nor do they address detectability of low-amplitude, high-frequency signals like heat whistles.

      In our revision we added a new discussion paragraph including these behavioral studies (Lines 357-367).

      For example, there are ample cases in the literature of altricial birds exhibiting behavioral evidence of auditory sensitivity by reducing begging calls in response to parental alarm calls:

      Platzen & Magrath (2004) - Playback of parental alarm calls nearly abolished nestling non-begging calls and reduced begging in scrubwrens. Proc. R. Soc. B 271:1271-1276.

      Different species: scrubwrens. Playback age: 5-, 8- and 11-DPH nestlings.

      Magrath, Haff, Horn & Leonard (2006) - Review and experiments on the developmental shift to silence/freeze after aerial alarm calls as chicks become fledglings; documents nestling quieting to alarms. Proc. R. Soc. B 273:2335-2341.

      Different species: scrubwrens. Playback age: 7-9 DPH nestlings, and 2- 4 days after fledging.

      Magrath, Pitcher & Dalziell (2007) - Nestlings respond to the sound of a predator's footsteps and parental food/alarm calls; includes begging suppression following predator sounds. Anim. Behav. 74:1117-1129.

      Different species: scrubwrens. Playback age: 8 DPH nestlings.

      Haff & Magrath (2012) - Nestlings suppress calling after heterospecific alarm calls (when acoustically similar to conspecific alarms), indicating generalized auditory danger recognition. Anim. Behav. 84:e.g., 495-505 (article).

      Different species: scrubwrens. Playback age: 5-6 and 10-11DPH nestlings. They show that 10-11 days old suppress calling while 5-6 days old do not.

      Barati & McDonald (2017) - Noisy miner nestlings suppress begging after conspecific alarm calls and some heterospecific cues; stronger/longer suppression for terrestrial-predator alarms. Sci. Rep. 7:9563.

      Different species: Noisy miner (Manorina melanocephala ). Playback age: 14 DPH. Nestlings started to vocalise at 5 DPH.

      Suzuki (2011) - In Paridae, parental alarm calls encode predator type; prior work (cited within) shows young of altricial species suppress vocalizations to alarms. Curr. Biol. 21:15-20.

      Different species: great tits. Playback age: 17 DPH.

      Can you please contextualize the present results about the timing of auditory development with the above body of work with respect to the timing of alarm call-induced begging call suppression?

      In our revision we have added a new paragraph in the discussion on these papers (Lines 357-367), and highlight the need for comparative work on hearing development in different species (Lines 352-355 and Lines 405-411).

      (5) Strictly speaking, a flat ABR does not equal deafness - at the extreme, an average of 10,000 trials may pull out a minuscule signal. Thus, the more rigorous path would be, in the results section, to ensure that statements summarize the data as they are, representing an absence of a neural signal.

      Save the interpretation of what this may mean for the discussion, and provide alongside this interpretation the necessary caveats related to temperature, rendition rate, averaging, etc.

      Clarify the conditions where a flat ABR demonstrates or fails to demonstrate immature deafness.

      Expand clarification for how the known 20-40 dB difference between ABR and behavioral thresholds can exist if a flat ABR can be interpreted as deafness.

      Consider refraining from concluding deafness from a flat ABR. Discuss that behavioral, single-unit, or alternative physiological assays might detect responses below the ABR threshold. If such cases exist, cite.

      (A5) We thought about this considerably before starting our measurements. What constitutes the absence of a signal? Even with intracellular recordings of all but one of the auditory neurons, the last one could still contain a signal and theoretically transmit information to the nervous system. We agree with the reviewers that absence of an ABR response should not be equated with absolute deafness. We have revised the manuscript accordingly and removed all statements implying “deafness” from the Results. The Results now strictly report presence or absence of detectable ABR responses.

      However, in both clinical and comparative contexts, absence of ABR responses at high SPLs (e.g., 90–95 dB) is widely interpreted as functionally non-responsive hearing. The developmental shift we observe (>54 dB) is far larger than typical ABR–behavioral offsets (20–40 dB). In the Discussion, we have added a new paragraph arguing that we think that the term functional deafness is reasonable here (Lines 261-269).

    1. eLife Assessment

      This manuscript presents important work on macrostructure formation in a freshwater filamentous cyanobacterium, focusing on its ability to aggregate and buckle. The authors employ a wide range of experiments and approaches, including time-lapse imaging and 3D theoretical modelling, to establish the physical dynamics of filaments, such as gliding motility and buckling. They demonstrate that, in addition to gliding motility, filament length and flexibility are essential for the ability of cyanobacteria to capture particles and form aggregates. Overall, the study provides convincing evidence and advances our understanding of the role of microbial motility in the environment. It will be of broad interest to biophysicists and environmental microbiologists.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors investigate the mechanisms underlying macrostructure formation in a freshwater filamentous cyanobacterium strain, F. draycotensis, focusing on how its ability to aggregate and form these structures depends on the physical properties of the filaments. Using experimental observations, they demonstrate that the cyanobacterium actively captures and surrounds particles, a process driven primarily by gliding motility.

      To explain these physical dynamics, the authors present a 3D model indicating that particle collection relies on filament length, as well as a specific mechanical response, namely, filament buckling and the subsequent formation of loops of bundles of filaments. While the authors have previously documented the buckling and looping characteristics of this strain, this study provides new insight by demonstrating that these physical phenomena are essential for particle capture and collection.

      Strengths:

      This manuscript benefits from a rigorous and detailed quantitative analysis of video recordings, which clearly documents the motility, buckling behaviour, and particle collection dynamics of the filaments.

      The authors effectively validate their hypothesis by using a naturally shorter filamentous strain, which fails to collect particles, suggesting that filament length is indeed a critical parameter.

      To further confirm the length dependency within the same species, the authors experimentally generated shorter filaments of F. draycotensis. The fact that these shortened filaments also lose the capacity to collect particles provides strong evidence supporting their proposed mechanism.

      Weaknesses:

      There is a conceptual concern. The authors linked the specific physical properties of this strain to evolutionary data, highlighting that the studied lineages diverged approximately two billion years ago. This creates a misleading impression that particle collection via flexible looping filaments is a recent evolutionary adaptation. However, particle collection has been observed in other cyanobacteria, such as Trichodesmium, which features short, rigid filaments. Therefore, the term "emerging" does not seem appropriate for the title and text. The capacity to collect particles in the studied strain F. draycotensis appears to be primarily a function of physical characteristics (filament length and flexibility) rather than evolutionary age. Any cyanobacterial strain possessing similar physical properties is likely to exhibit comparable behaviour, rendering the evolutionary timeframe largely irrelevant to the core mechanism. In addition, the phylogenetic tree presented in Figure S5 does not reflect the current consensus on cyanobacterial evolution and systematics and does not align with modern phylogenomic frameworks (see, for example, Strunecky et al., 2023 https://doi.org/10.1111/jpy.13304). There is also no such order Cyanobacteriales, which has been mentioned in a few older publications but is clearly outdated.

      Another concern is that the authors nearly completely ignore the role of type IV pili in the gliding motility of cyanobacteria, including filamentous strains. For a long time, there was a misconception that the gliding motility of cyanobacteria was due to slime protrusion. Slime plays a role in this process. However, several studies have shown that filamentous strains also use type IV pili to glide on surfaces. The authors should discuss this and include it in their model. In addition, the authors concluded that gliding motility is responsible for particle collection by Fluctiforma draycotensis. Although I believe that their conclusion is correct, there might be several limitations to the experiments which allow for other reasons to be considered. Their conclusions were based on the use of a non-motile strain and an unspecified community without the motile Fluctiforma draycotensis strain. The problem I see here is that it is not clear why this strain is not motile; it could be because of the lack of type IV pili, mutations which alter their functionality, defects in slime secretion, any other mutation (e.g. in chemoreceptors), cellular structure, metabolism, or combinations of these. Furthermore, it is possible that the community changes its composition and behaviour when it lives without the cyanobacterium with a rich carbon source (glucose) or with a non-motile cyanobacterium which may not secrete slime or, for example, a signalling component which controls behaviour of the bacteria in the community. For that reason, the authors should be more cautious with their conclusion that solely motility behaviour of Fluctiforma draycotensis is responsible for particle collection. Additional factors might be responsible for these effects.

    3. Reviewer #2 (Public review):

      Summary:

      The authors studied aggregation, buckling, and particle collection by the filamentous cyanobacterium Fluctiforma draycotensis, as well as by the filamentous Pseudanabaena sp. (order Pseudoanabenales). They performed a range of experiments, from imaging individual gliding filaments to multiple-day experiments showing the formation of large aggregates around a particle formed from a precipitate. They also developed a model of buckling filaments to argue that the ability of elastic filaments to collect particles and form macrostructures is confined to a part of the filament phase space in terms of length and flexibility, meaning that gliding combined with certain filament length and flexibility naturally reproduces the observations.

      Strengths:

      This is an impressive study that uses multiple tools to connect macrostructure formation with filaments' gliding motility and buckling. It adds an important perspective on the biological and physical factors at play in the emergence of aggregates.

      Weaknesses:

      The authors ignore the possibility that filament behavior plays an important role in the emergence of the observed patterns. Cyanobacteria have been shown to control their gliding motility (Pfreundt et al Science 2023; Kurjahn et al Nature Comm 2024), and their molecular motors are known to be regulated by chemotaxis-like signaling pathways (Risser ARM 2025). As far as I know, how the coordination between the pulling agents along an individual filament works is actively debated, but there seems to be little doubt that it exists. 

      To illustrate this point better, note that the aggregation observed by the authors is consistent with the length-dependent ability of filaments to coordinate gliding (I'm not saying this is how it works in Fluctiforma draycotensis; I'm saying it's consistent). Suppose the coordination requires sufficiently long filaments, which could be the case when signaling molecules travel along the filament, propagating information about when individual pulling agents should reverse. In such a model, short filaments act randomly because they fail to coordinate gliding by the time they glide off nascent aggregates, whereas longer filaments can perform informed reversals because they have more time for coordination. Such behavior then explains the lack of aggregation in Pseudanabaena sp. (via behavior, not lack of stiffness). Note that Trichodesmium is stiff; its filaments do not buckle, yet Trichodesmium forms organized aggregates via tightly controlled motility. Note also that, as the authors report, since Pseudanabaena sp. is both shorter and faster, its filaments have relatively (to the time needed to glide the filaments' length) little time to coordinate reversals. In my opinion, whether the observed patterns passively emerge from gliding and buckling or result from active behavior remains an open question.

      I also have a small suggestion regarding this statement on model novelty:

      The essential novelty of this model is that the filament itself is active and out of equilibrium, and additionally, the forces and torques are applied locally along its centreline, and not at its extremities as in previous steady-state mechanical studies of elastic, twistable filaments such as DNA [31-33] (see Methods and SI).

      This statement needs to be revised as it ignores a substantial body of work on self-organization of active filaments: (R. E. Isele-Holder, J. Elgeti, G. Gompper, Soft Matter 2015; Pfreudnt et al, Science 2023; Faluweki et al PRL 2023; Kurjahn et al Nature Comm 2024).

      Last point: the authors often say that their observations are reproducible ('...reproducibly forms macroscopic granules...'). What is meant? Different experiments on different days, different aliquots?

    4. Reviewer #3 (Public review):

      Summary:

      The authors report and characterize the formation of aggregate microstructures by the motile filamentous cyanobacterium Fluctiforma draycotensis, which exhibits gliding motility accompanied by rotation along the long axis while excreting EPS. In experiments with motile F. draycotensis cultures, they observed the formation of granular structures composed of cyanobacteria and other material (iron, polystyrene beads, etc.), with macrostructures on the scale of 1mm within 24 hours. The structures were motile at speeds comparable to that of the cyanobacteria filaments, resulting in their growth through coalescence over time. Notably, such macrostructures were absent in nonmotile F. draycotensis, pointing to the role of filament motility in their formation. Through experiments examining the micro-scale dynamics, inert material such as small polystyrene beads was found to be transported by the gliding, buckling, and plectoneme dynamics of the filaments, pointing to the underlying mechanism by which particles are collected into larger-scale microgranule structures.

      To interrogate the properties that drive the cyanobacteria filament buckling, plectoneme formation, and entanglement, the authors develop a mechanical model for filaments as nearly inextensible, slender bodies with resistance to twisting and bending under active gliding forces and torques and responding to fluid flows and surface adhesion. They derive expressions for the thresholds for buckling and twisting instabilities, which are additionally demonstrated and interrogated through simulation via the Immersed Boundary Method. Most importantly, bending and plectoneme formation only occur with sufficiently long filaments, and the threshold is shorter for bending than for plectoneme formation. Experimental observations with wild-type filaments agree with the model-predicted thresholds. The authors perform additional experiments with shorter filaments below both thresholds, including the filamentous bacterium Pseudanabaena, which fail to collect particles (though can in principle form macrostructures).

      Strengths:

      This work appears to be novel (notably, the discovery and characterization of the particle collection behavior of a filamentous cyanobacterium) and has interesting implications for both naturally observed cyanobacterial macrostructures as well as the controllable parameters in engineering them. The experimental and modeling work is well motivated, contributing to the broader understanding of macrostructure formation and material aggregation through active filament dynamics (not exclusive to cyanobacteria), as well as the underlying physical properties governing important filamentous cyanobacterium dynamics. As such, I would expect the results of this paper to be of broad interest to both biophysicists and microbiologists. Generally, the manuscript is well written with clear, compelling figures that illustrate the important conclusions of this study.

      Weaknesses:

      In the section on "Shorter gliding filaments cannot collect particles nor form granule macrostructures", the filamentous cyanobacteria considered *all* fall below the predicted thresholds for bending and twisting. The "long" F. draycotensis are 60 microns in length, notably less than the 120 and 320 micron thresholds derived in the previous section as well as the lengths of filaments considered in Figure 3D, yet these "long" 60 micron filaments form macrostructures. How can this be understood in the context of the model predictions? Is the nature of the macrostructures in Figure 4B, the microscale parameters, or the collection of particles somehow different than those with filaments an order of magnitude longer in earlier parts of the paper? The paper would be stronger if these sorts of questions were addressed in the text and/or with supplementary figures.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary.

      In this manuscript, the authors investigate the mechanisms underlying macrostructure formation in a freshwater filamentous cyanobacterium strain, F. draycotensis, focusing on how its ability to aggregate and form these structures depends on the physical properties of the filaments. Using experimental observations, they demonstrate that the cyanobacterium actively captures and surrounds particles, a process driven primarily by gliding motility.

      To explain these physical dynamics, the authors present a 3D model indicating that particle collection relies on filament length, as well as a specific mechanical response, namely, filament buckling and the subsequent formation of loops of bundles of filaments. While the authors have previously documented the buckling and looping characteristics of this strain, this study provides new insight by demonstrating that these physical phenomena are essential for particle capture and collection.

      Strengths:

      This manuscript benefits from a rigorous and detailed quantitative analysis of video recordings, which clearly documents the motility, buckling behaviour, and particle collection dynamics of the filaments.

      The authors effectively validate their hypothesis by using a naturally shorter filamentous strain, which fails to collect particles, suggesting that filament length is indeed a critical parameter.

      To further confirm the length dependency within the same species, the authors experimentally generated shorter filaments of F. draycotensis. The fact that these shortened filaments also lose the capacity to collect particles provides strong evidence supporting their proposed mechanism.

      We thank the reviewer for the accurate summary of our work and for their identified strengths of the study.

      Weaknesses:

      There is a conceptual concern. The authors linked the specific physical properties of this strain to evolutionary data, highlighting that the studied lineages diverged approximately two billion years ago. This creates a misleading impression that particle collection via flexible looping filaments is a recent evolutionary adaptation. However, particle collection has been observed in other cyanobacteria, such as Trichodesmium, which features short, rigid filaments. Therefore, the term ”emerging” does not seem appropriate for the title and text. The capacity to collect particles in the studied strain F. draycotensis appears to be primarily a function of physical characteristics (filament length and flexibility) rather than evolutionary age. Any cyanobacterial strain possessing similar physical properties is likely to exhibit comparable behaviour, rendering the evolutionary timeframe largely irrelevant to the core mechanism.

      We would like to first clarify our use of the term “emergent”. It seems that the referee took this in an evolutionary context, whereas we are using this term in the context of its use in systems dynamics, and referring to: “a complex entity displaying behaviors that its components do not have on their own, and emerge only when they interact in a wider whole”. Here, particle collection and dynamic aggregate formation “emerges” from the buckling and interaction of many filaments.

      With regards to the evolution of particle collection behavior, our comment on the evolutionary distance between F. draycotensis and Pseudoanabena sp. was meant to highlight the point that particle collection seems to be a function of physical characteristics and motility: Despite a large evolutionary distance, and possibly many biological differences, a physics-based argument is capturing the difference between the particle collection ability of these two organisms. Thus, we are in agreement here with the reviewer. We did not intend to make any arguments about “evolutionary age” of the particle collection behavior.

      We see that the short, evolutionary comment in the Introduction has confused the reviewer and potentially is confusing to other readers too. We will therefore remove this evolutionary comment from the Introduction section of the revised manuscript and make the point in more detail in the Discussion section.

      In addition, the phylogenetic tree presented in Figure S5 does not reflect the current consensus on cyanobacterial evolution and systematics and does not align with modern phylogenomic frameworks (see, for example, Strunecky et al., 2023 https://doi.org/10.1111/jpy.13304). There is also no such order Cyanobacteriales, which has been mentioned in a few older publications but is clearly outdated.

      We thank the reviewer for this comment, as it has made us realise that we never explained our choice of taxonomic framework in the manuscript, and perhaps this is the source of the confusion.

      The order Cyanobacteriales does exist: it is the order-level name applied in the Genome Taxonomy Database (GTDB) [5, 6], currently the most comprehensive and actively curated genome-based taxonomy of prokaryotes. GTDB classifies taxonomic groups algorithmically, as monophyletic groups in a concatenated marker-protein phylogeny with ranks normalised by relative evolutionary divergence. This has produced a number of re-groupings and new names relative to the older, morphology-derived classifications; many of these have since been formally proposed under the International Code of Nomenclature of Prokaryotes and the SeqCode [2], and are progressively being adopted by the NCBI. The placement of Cyanobacteriales, and of the other orders shown in Figure S5, can be inspected directly on the GTDB “Taxonomy Tree” (see here for the orders within the class Cyanobacteriia).

      We would also like to note that we do not see our tree and the framework of Strunecky et al. as being in conflict. Strunecky et al. constructed their phylogenomic backbone using GTDB-Tk and the same 120-marker concatenated alignment that the GTDB itself uses. What differs between the two schemes is therefore not the underlying phylogeny but the nomenclature applied to the resulting clades: Strunecky et al. work within the botanical tradition and combine the phylogenomic tree with phenotypic characteristics, thereby proposing ten new orders and fifteen new families, whereas GTDB assigns rank boundaries purely by evolutionary divergence and so draws broader order limits. In practice, the GTDB order Cyanobacteriales spans several of the families (e.g. Oscillatoriales and Coleofasciculales) and orders (e.g. Chroococcales and Nostocales), that are proposed within the Strunecky et al. work. Our reason for adopting the GTDB nomenclature is for practical reasons specific to this study. F. draycotensis is a recently described organism [3] that is not included in Strunecky et al. and has no placement in their tree. In GTDB it falls within a family-level lineage (placeholder name JAAUUE01) inside the Cyanobacteriales, with the sequenced members of the Coleofasciculaceae as its closest relatives. We could not have assigned it to one of the Strunecky orders without inventing a placement. The same applies to some of the other, recent metagenomically described cyanobacteria [10], which similarly have no assigned names in the literature. GTDB, by contrast, provides a reproducible, algorithmic assignment for all of these genomes, and is now widely used for this reason in genome- and metagenome-based studies of cyanobacteria (e.g. [1]). We therefore used it consistently throughout.

      Finally, with regards to the reviewer’s point about the tree itself, we would like to note that Figure S5 was intended only to convey the evolutionary distance between F. draycotensis and Pseudanabaena sp., and it was built from a modest set of six concatenated ribosomal protein markers using an approximate maximum-likelihood method with SH-like local support values. This is considerably less rigorous than the 120-marker RAxML and Bayesian analysis of Strunecky et al., and we agree that a stronger tree may be preferable. For the revised manuscript we are recomputing the tree from a substantially larger set of concatenated single-copy marker genes, using IQTREE with model selection and non-parametric bootstrap support. We would note, however, that the specific conclusion drawn from this figure — that the two strains we use for our experimental work, namely F. draycotensis and Pseudoanabena sp. belong to deeply divergent cyanobacterial lineages — is supported by the deep backbone of the cyanobacterial tree, which is stable across marker sets and inference methods, and is equally supported by the tree of Strunecky et al.

      We will make these points clearer in the Methods and Discussion sections of the revised manuscript, as well as the Figure S5 legend.

      Another concern is that the authors nearly completely ignore the role of type IV pili in the gliding motility of cyanobacteria, including filamentous strains. For a long time, there was a misconception that the gliding motility of cyanobacteria was due to slime protrusion. Slime plays a role in this process. However, several studies have shown that filamentous strains also use type IV pili to glide on surfaces. The authors should discuss this and include it in their model.

      The reviewer is correct that we did not include molecular details of gliding motility in our biophysical model. They are also correct to point out that pili and slime biosynthesis genes are shown to be involved in gliding motility [8]. It is, however, still unclear how these factors interact to produce mechanical gliding forces that can result in filament rotation (observed only in some filamentous cyanobacteria), filament reversal, as well as decoordination during such reversals, which we have previously shown in F. draycotensis [9]. Therefore, we have chosen to keep the biophysical model at a coarse-grained, phenomenological level. Instead of explicitly modelling the detailed molecular mechanisms behind force generation, we model only the minimum necessary physical forces and torques needed to reproduce the observed rotation and translation of the filament during gliding under de-coordinated conditions. This model is able to reproduce the experimentally observed buckling and twisting of filaments, and is therefore sufficient and useful to achieve a coarse-grained understanding of mechanical forces and their relationship to buckling, twisting and entanglement, which are the main processes we focus on here. As molecular details behind force generation in rotating, filamentous cyanobacteria become available, more detailed physical models can be constructed. We also note, in this context, that the two filamentous cyanobacteria we compare both encode the type IV pilus machinery, so the presence of a pilus motor does not by itself distinguish a particle-collecting from a non-collecting strain (see our response to the reviewer’s next point).

      We will make these points clearer in the Methods and Discussion sections of the revised manuscript.

      In addition, the authors concluded that gliding motility is responsible for particle collection by Fluctiforma draycotensis. Although I believe that their conclusion is correct, there might be several limitations to the experiments which allow for other reasons to be considered. Their conclusions were based on the use of a non-motile strain and an unspecified community without the motile Fluctiforma draycotensis strain. The problem I see here is that it is not clear why this strain is not motile; it could be because of the lack of type IV pili, mutations which alter their functionality, defects in slime secretion, any other mutation (e.g. in chemoreceptors), cellular structure, metabolism, or combinations of these. Furthermore, it is possible that the community changes its composition and behaviour when it lives without the cyanobacterium with a rich carbon source (glucose) or with a non-motile cyanobacterium which may not secrete slime or, for example, a signalling component which controls behaviour of the bacteria in the community. For that reason, the authors should be more cautious with their conclusion that solely motility behaviour of Fluctiforma draycotensis is responsible for particle collection. Additional factors might be responsible for these effects.

      Our conclusion that gliding motility is the main factor underpinning particle collection is based on several observations.

      Firstly, on the macroscopic scale we present several control experiments where we did not observe particle collection: (i) in the community featuring a non-motile F. draycotensis, and with mostly the same other bacterial species as the community featuring the motile F. draycotensis, (ii) in a bacterial community derived from the original F. draycotensis community but lacking any cyanobacteria, (iii) in the original community with physically shortened F. draycotensis, and (iv) in another cyanobacterial community featuring different bacteria and a naturally shorter, filamentous gliding cyanobacteria Pseudanabaena sp. A straightforward, parsimonious explanation that satisfies all these observations is that particle collection is underpinned by physical characteristics of gliding filamentous cyanobacteria.

      Secondly and more directly, in time-lapse microscopy imaging we repeatedly observe clusters of beads being moved by gliding filaments, and thereby being collected into larger clusters. Thus, whilst factors such as slime secretion also contribute, the primary mechanism driving the observed particle motion seems to be that particles stick to filaments and are carried around with them as they glide. We cannot rule out a contribution of pili to bead attachment and transport. We note, however, that both cyanobacteria compared here encode the type IV pilus machinery. In a homology survey of the two genomes, Pseudanabaena sp. and F. draycotensis both carry orthologues of the core T4P components — the assembly ATPase PilB, the retraction ATPase PilT, the inner-membrane platform protein PilC, the prepilin peptidase PilD, and the alignment-complex proteins PilM and PilF — together with the hormogonium-associated hmpD, hmpF and hmpG. Pseudanabaena sp. is therefore not pilus-deficient, and it does glide, yet it does not collect particles. The difference between the two organisms consequently cannot be attributed to the presence or absence of the pilus motor, which we would argue supports the physical argument we make here. Consistent with this, we have not identified mutations in pilus-related genes in the mutant, non-motile F. draycotensis.

      We are currently in the process of preparing another manuscript describing the mutations that led to motility loss in the non-motile F. draycotensis, as well as the proteins that are differentially expressed in the motile and non-motile F. draycotensis. These analyses will shed more light on the molecular mechanisms abolishing motility and how they might be influencing particle collection.

      In the revised manuscript, we will make these points clearer in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      The authors studied aggregation, buckling, and particle collection by the filamentous cyanobacterium Fluctiforma draycotensis, as well as by the filamentous Pseudanabaena sp. (order Pseudoanabenales). They performed a range of experiments, from imaging individual gliding filaments to multiple-day experiments showing the formation of large aggregates around a particle formed from a precipitate. They also developed a model of buckling filaments to argue that the ability of elastic filaments to collect particles and form macrostructures is confined to a part of the filament phase space in terms of length and flexibility, meaning that gliding combined with certain filament length and flexibility naturally reproduces the observations.

      Strengths:

      This is an impressive study that uses multiple tools to connect macrostructure formation with filaments’ gliding motility and buckling. It adds an important perspective on the biological and physical factors at play in the emergence of aggregates.

      We thank the reviewer for the accurate summary of our work and highlighting the strengths of the study.

      Weaknesses:

      The authors ignore the possibility that filament behavior plays an important role in the emergence of the observed patterns. Cyanobacteria have been shown to control their gliding motility (Pfreundt et al Science 2023; Kurjahn et al Nature Comm 2024), and their molecular motors are known to be regulated by chemotaxislike signaling pathways (Risser ARM 2025). As far as I know, how the coordination between the pulling agents along an individual filament works is actively debated, but there seems to be little doubt that it exists.

      To illustrate this point better, note that the aggregation observed by the authors is consistent with the length-dependent ability of filaments to coordinate gliding (I’m not saying this is how it works in Fluctiforma draycotensis; I’m saying it’s consistent). Suppose the coordination requires sufficiently long filaments, which could be the case when signaling molecules travel along the filament, propagating information about when individual pulling agents should reverse. In such a model, short filaments act randomly because they fail to coordinate gliding by the time they glide off nascent aggregates, whereas longer filaments can perform informed reversals because they have more time for coordination. Such behavior then explains the lack of aggregation in Pseudanabaena sp. (via behavior, not lack of stiffness). Note that Trichodesmium is stiff; its filaments do not buckle, yet Trichodesmium forms organized aggregates via tightly controlled motility. Note also that, as the authors report, since Pseudanabaena sp. is both shorter and faster, its filaments have relatively (to the time needed to glide the filaments’ length) little time to coordinate reversals. In my opinion, whether the observed patterns passively emerge from gliding and buckling or result from active behavior remains an open question.

      We appreciate the comment by the reviewer. We certainly agree that behavioral responses exist in filamentous cyanobacteria and will interplay with the physical aspects to produce exciting, complex dynamics. Besides the exemplar ideas that the reviewer provides, there can be many other scenarios involving behavioral responses, such as responses to light and to quorum sensing molecules or photosynthesis-generated radicals. For example, in F. draycotensis we have observed photo-responses at the aggregate level, which we are are currently studying. Photoresponses are also observed in Trichodesmium aggregates [7]. In general, a full understanding of the interaction of the biological (i.e. behavioral) and the physical aspects will require several future studies.

      In the current study, however, we focus on characterising the physical aspects of gliding motility alone, combined with experimental observations. We believe that this approach is important to establish a form of “null expectation” from the physics of gliding, elastic filaments alone. Currently, the molecular mechanisms responsible for coordinating the reversal behaviour of multiple filaments are still unclear, so it is difficult to experimentally demonstrate behavioural contributions to aggregate formation, e.g. via experiments where such behaviour is switched off. In the meantime, simulations such as those presented here allow us to test more precisely the potential role of activity, coordinated reversals and the elastic properties of the filament. In future it will be interesting to scale up the presented model to include multiple interacting filaments, and to systematically test the respective roles of active coordination behaviour for one individual filament (reversals) and for multiple interacting filaments (where contacts modulate activity), as well as the physical properties (length and flexibility). Such modelling studies can then identify if a ‘purely physical’ model can or cannot generate realistic aggregates, and pinpoint whether additional coordination mechanisms are needed to regulate aggregation. By testing the combination of different physical and biological coordination mechanisms, it would then help to indicate how much of a role is played by various potential active coordination behaviours.

      We will bring out this point more clearly in the Discussion section of the revised manuscript.

      I also have a small suggestion regarding this statement on model novelty:

      The essential novelty of this model is that the filament itself is active and out of equilibrium, and additionally, the forces and torques are applied locally along its centreline, and not at its extremities as in previous steady-state mechanical studies of elastic, twistable filaments such as DNA [31-33] (see Methods and SI).

      This statement needs to be revised as it ignores a substantial body of work on self-organization of active filaments: (R. E. Isele-Holder, J. Elgeti, G. Gompper, Soft Matter 2015; Pfreudnt et al, Science 2023; Faluweki et al PRL 2023; Kurjahn et al Nature Comm 2024).

      We agree with the reviewer that there is a significant literature on active filaments, some of which we have already cited and will now discuss in more details, as well as adding and discussing the suggested additional references. Our statement on “model novelty” refers to the analysis of buckling instabilities of biological filaments, and in particular DNA, due to a combination of forces and torques. To our knowledge, this has only be studied explicitely by [4], and only in the local (resistive force theory) limit. The elastohydrodynamic simulations coupled to local active forces and torques, as we implemented here, are therefore novel and will expand the analysis of both microbial filaments and other biological polymers. We will clarify these points in the Methods and Discussion sections of the revised manuscript.

      Last point: the authors often say that their observations are reproducible (’...reproducibly forms macroscopic granules...’). What is meant? Different experiments on different days, different aliquots?

      The “replicability” statement was in reference to different experiments started on different days using cultures obtained from serial transfer experiments, as well as cultures re-initiated from cyrostocks. This point will be made clear in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The authors report and characterize the formation of aggregate microstructures by the motile filamentous cyanobacterium Fluctiforma draycotensis, which exhibits gliding motility accompanied by rotation along the long axis while excreting EPS. In experiments with motile F. draycotensis cultures, they observed the formation of granular structures composed of cyanobacteria and other material (iron, polystyrene beads, etc.), with macrostructures on the scale of 1mm within 24 hours. The structures were motile at speeds comparable to that of the cyanobacteria filaments, resulting in their growth through coalescence over time. Notably, such macrostructures were absent in nonmotile F. draycotensis, pointing to the role of filament motility in their formation. Through experiments examining the micro-scale dynamics, inert material such as small polystyrene beads was found to be transported by the gliding, buckling, and plectoneme dynamics of the filaments, pointing to the underlying mechanism by which particles are collected into larger-scale microgranule structures.

      To interrogate the properties that drive the cyanobacteria filament buckling, plectoneme formation, and entanglement, the authors develop a mechanical model for filaments as nearly inextensible, slender bodies with resistance to twisting and bending under active gliding forces and torques and responding to fluid flows and surface adhesion. They derive expressions for the thresholds for buckling and twisting instabilities, which are additionally demonstrated and interrogated through simulation via the Immersed Boundary Method. Most importantly, bending and plectoneme formation only occur with sufficiently long filaments, and the threshold is shorter for bending than for plectoneme formation. Experimental observations with wild-type filaments agree with the model-predicted thresholds. The authors perform additional experiments with shorter filaments below both thresholds, including the filamentous bacterium Pseudanabaena, which fail to collect particles (though can in principle form macrostructures).

      Strengths:

      This work appears to be novel (notably, the discovery and characterization of the particle collection behavior of a filamentous cyanobacterium) and has interesting implications for both naturally observed cyanobacterial macrostructures as well as the controllable parameters in engineering them. The experimental and modeling work is well motivated, contributing to the broader understanding of macrostructure formation and material aggregation through active filament dynamics (not exclusive to cyanobacteria), as well as the underlying physical properties governing important filamentous cyanobacterium dynamics. As such, I would expect the results of this paper to be of broad interest to both biophysicists and microbiologists. Generally, the manuscript is well written with clear, compelling figures that illustrate the important conclusions of this study.

      We thank the reviewer for the accurate summary of our work and recognising the broad relevance of the study.

      Weaknesses:

      In the section on “Shorter gliding filaments cannot collect particles nor form granule macrostructures”, the filamentous cyanobacteria considered “all” fall below the predicted thresholds for bending and twisting. The “long” F. draycotensis are 60 microns in length, notably less than the 120 and 320 micron thresholds derived in the previous section as well as the lengths of filaments considered in Figure 3D, yet these “long” 60 micron filaments form macrostructures. How can this be understood in the context of the model predictions? Is the nature of the macrostructures in Figure 4B, the microscale parameters, or the collection of particles somehow different than those with filaments an order of magnitude longer in earlier parts of the paper? The paper would be stronger if these sorts of questions were addressed in the text and/or with supplementary figures.

      We thank the reviewer for this point. Indeed as we mention in the text, the ‘long’ population has a mean length of 60 micron. However, as shown in the length distribution plot in Fig 4A, the maximum filament lengths observed in these populations (within the samples used for microscopy) are 560 microns for the long filaments, versus 240 microns for the short filaments. Thus, we expect the long population to contain multiple filaments that can buckle and a few that can form plectonemes, whilst the short population might have some buckling filaments and none that form plectonemes. We stress that Fig 4A only shows the length distribution for what we believe to be a representative sample taken from the long and short populations, not the full data from the entire population.

      We will revise the main text to include the maximum filament lengths of the two populations as well as the mean values. We will also add lines to Fig 4A to indicate the buckling and plectoneme threshold lengths from the analytical estimate for the F. draycotensis filaments (same values as in Fig 3), to make it clear that the long population contains more buckling/plectoneming filaments than the short population.

      References:

      (1) E. S. Cameron et al. “Diversity and specificity of molecular functions in cyanobacterial symbionts”. In: Sci Rep 14.1 (2024), p. 18658. issn: 2045-2322 (Electronic) 2045-2322 (Linking). doi: 10.1038/s41598-024-69215-8. url: https: //www.ncbi.nlm.nih.gov/pubmed/39134591.

      (2) M. Chuvochina et al. “Proposal of names for 329 higher rank taxa defined in the Genome Taxonomy Database under two prokaryotic codes”. In: FEMS Microbiol Lett 370 (2023). issn: 1574-6968 (Electronic) 0378-1097 (Print) 0378-1097 (Linking). doi: 10.1093/femsle/fnad071. url: https://www.ncbi.nlm.nih.gov/ pubmed/37480240.

      (3) S. J. N. Duxbury et al. “Niche formation and metabolic interactions contribute to stable diversity in a spatially structured cyanobacterial community”. In: ISME J (2025). issn: 1751-7370 (Electronic) 1751-7362 (Linking). doi: 10.1093/ismejo/ wraf126. url: https://www.ncbi.nlm.nih.gov/pubmed/40577531.

      (4) Raymond E. Goldstein, Thomas R. Powers, and Chris H. Wiggins. “Viscous Nonlinear Dynamics of Twist and Writhe”. In: Physical Review Letters 80.23 (June 1998), pp. 5232–5235. issn: 1079-7114. doi: 10.1103/physrevlett.80.5232.

      (5) D. H. Parks et al. “A standardized bacterial taxonomy based on genome phylogeny substantially revises the tree of life”. In: Nat Biotechnol 36.10 (2018), pp. 996–1004. issn: 1546-1696 (Electronic) 1087-0156 (Linking). doi: 10.1038/ nbt.4229. url: https://www.ncbi.nlm.nih.gov/pubmed/30148503.

      (6) D. H. Parks et al. “GTDB release 10: a complete and systematic taxonomy for 715 230 bacterial and 17 245 archaeal genomes”. In: Nucleic Acids Res 54.D1 (2026), pp. D743–D754. issn: 1362-4962 (Electronic) 0305-1048 (Print) 03051048 (Linking). doi: 10.1093/nar/gkaf1040. url: https://www.ncbi.nlm.nih. gov/pubmed/41123020.

      (7) U. Pfreundt et al. “Controlled motility in the cyanobacterium Trichodesmium regulates aggregate architecture”. In: Science 380.6647 (2023), pp. 830–835. issn: 1095-9203 (Electronic) 0036-8075 (Linking). doi: 10.1126/science.adf2753.

      (8) Douglas D Risser. “Motility in Filamentous Cyanobacteria”. In: Annual Review of Microbiology 79 (2025).

      (9) Jerko Rosko et al. “Cellular coordination underpins rapid reversals in gliding filamentous cyanobacteria and its loss results in plectonemes”. In: eLife 13 (2025), RP100768.

      (10) A. Scarampi et al. “Enrichment of convergent metabolic functions in microbial communities through imposed and emergent environmental niches”. In: bioRxiv (2026). doi: 10.64898/2026.02.11.705344.

    1. eLife Assessment

      This study presents valuable findings regarding the role of the P2X7 receptor in retinoic acid signaling-mediated retinal remodeling following photoreceptor degeneration. However, the evidence supporting the main conclusions is currently incomplete, as key comparisons are confounded by differences in genetic background, and the cellular source of receptor expression is not sufficiently resolved. While the work offers mechanistic insights of interest to researchers studying retinal degeneration, additional experimental controls are needed to fully support the authors' claims.

    2. Reviewer #1 (Public review):

      In this study, Telias et al. identify the P2X7 receptor as a key component of retinoic acid signaling-mediated remodeling in the degenerating rd1 retina. The authors report increased P2X7R expression in the inner retina following photoreceptor loss and link P2X7R signaling to retinal ganglion cell hyperactivity, membrane hyperpermeability, and altered calcium homeostasis. Genetic deletion of p2rx7 abolishes ganglion cell hyperpermeability and reduces several features of pathological remodeling in the rd1 retina. The study provides valuable mechanistic insight into retinal remodeling following photoreceptor degeneration. However, several issues currently limit the strength of the conclusions. In particular, key comparisons are confounded by differences in genetic background, the cellular source of P2X7R expression is not sufficiently resolved, and several experiments require additional controls and more cautious interpretation.

      Major comments:

      (1) The Methods state that C57BL/6J mice were used as wild-type controls, whereas rd1 mice were maintained on a C3H/HeJ background, and the rd1-p2rx7 knockout line is on a mixed background. Direct comparisons among these groups are therefore potentially confounded by strain-specific differences. Authors should use littermate rd1 het mice as healthy controls in all their experiments.

      (2) The authors do not demonstrate P2X7R expression in RGCs. The P2X7R signal in the rd1 retina shown in Figure 1B appears saturated and is therefore difficult to compare directly with the WT image in Figure 1A. Furthermore, Figure 1E indicates that overall P2X7R fluorescence in the GCL is not significantly different between WT and rd1 retinas, whereas the representative images appear to suggest a marked increase. The authors should provide images acquired and displayed under identical settings and consider including retinal whole-mount staining with an RGC-specific marker. P2X7R abundance should then be quantified specifically within identified RGCs in both healthy and rd1 retinas. As mentioned above, het rd1 mice should be included. As an additional control, the authors also should include the staining of rd1 p2x7r KO retinas.

      (3) Does the increase in P2X7R abundance correlate with photoreceptor loss? Please include staining of younger rd1 mice along with het rd1 littermates.

      (4) In Figure 1G-H, the description of the reporter is internally inconsistent: the Results refer to an artificial mini-Pax6 promoter, whereas the Figure 1 legend describes the Ple344 neuronal mini-promoter derived from Tubb3. Please clarify which promoter is used and what cell population the ECFP signal labels. Figure 1G should explain the function of each reporter element and how RAR activity is inferred. Figure 1H should include an RGC marker such as RBPMS and provide quantitative analysis of RBPMS-positive, reporter-positive, and Yo-Pro-positive cells. Additional controls are needed to exclude effects of viral transduction or retinal inflammation on Yo-Pro uptake. They should include rd1 retinas without AAV, rd1 het retinas with and without the reporter AAV.

      (5) In addition, lines 143-144 state that two experiments were performed, but only one is described in that paragraph; the text should be reorganized or clarified.

      (6) Yo-Pro-1-positive cell density in rd1 mice between Figure 1H and Figure 2D is different. Why?

      (7) Constitutive deletion of p2rx7 may cause developmental or compensatory changes that could contribute to the observed phenotype in Figures 2 and 3. Additional controls are therefore needed to distinguish acute effects of p2rx7 loss from developmental consequences. The authors should assess whether p2rx7 deletion alters retinal cell-type composition, including RGC density, or affects the timing or extent of photoreceptor degeneration. Perform a rescue experiment to determine whether overexpression of p2rx7 in the knockout background restores the phenotype. For all experiments, het rd1 control mice should be included.

      (8) Figure 3D and E experiments should include control AAV expression such as GFP.

      (9) Figure 5 experiments should include control rd1 het mice.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors used genetic, transcriptomic, imaging, and electrophysiological approaches to investigate the role of P2X7R in retinal remodeling in rd1 mice. The authors propose that RA signaling upregulates P2X7R, leading to altered Ca²⁺ signaling, HCN1 expression, and spontaneous RGC hyperactivity.

      Strengths:

      Multiple lines of experiments were conducted, focusing on an important question in the field.

      Weaknesses:

      However, some aspects of the experimental design, statistical analysis, and interpretation require clarification. In particular, the genetic controls and causal evidence should be strengthened before the proposed RA-P2X7R-Ca²⁺-HCN1 pathway can be fully supported.

      Some major issues:

      (1) The Abstract states that P2X7R deletion prevents the upregulation of RA-responsive genes, whereas the Results state that RAR-dependent genes were not consistently changed by P2rx7 deletion. This central statement should be corrected and clarified.

      (2) The P2rx7 knockout model requires further discussion. JAX strain 005576 targets exon 13 and has previously been reported to retain truncated P2X7 transcripts with residual activity (Masin et al., 2012). The authors should avoid describing this allele as complete loss of the entire P2rx7 gene unless additional isoform-specific validation is provided.

      (3) The evidence that RAR directly regulates P2rx7 transcription remains incomplete. Predicted promoter motifs and reduced P2X7R protein after BMS-493 treatment are supportive, but they do not establish direct transcriptional regulation. Measurement of P2rx7 mRNA, RAR promoter occupancy, or promoter mutagenesis would strengthen this conclusion.

      (4) Several statistical results require verification. In Figure 2H, four paired eyes are analyzed using an unpaired Mann-Whitney test, and the reported p<0.001 is difficult to reconcile with n=4. Similarly, the significance levels in Figure 5A-D are not compatible with a two-sided Wilcoxon rank-sum test using n=3 mice per group. The statistical unit, exact P values, and number of biological replicates should be rechecked.

      (5) The overall causal pathway remains partially inferential. The study does not directly show that increased Ca²⁺ causes HCN1 upregulation or that HCN1 is required for RGC hyperactivity. These steps should either be experimentally tested or presented as a proposed model rather than an established mechanism.

    1. eLife Assessment

      This important work considerably advances our understanding of the nature of ongoing thought across individuals and across task contexts. Using a combination of fMRI brain data and experience sampling, they find that activation of the MDN leads to greater stability of deliberate task-focused thoughts during task performance. The evidence supporting the claims is convincing and both analytically strong and creative, offering a novel approach to capture thought stability and its relationship to independent brain maps. The findings will be of broad interest to cognitive scientists and neuroimagers.

    2. Reviewer #1 (Public review):

      Summary:

      This paper suggests an alternative model for the function of the multiple demand network. Specifically, its role is not necessarily to sustain cognitive control and maintain task sets, but rather to "stabilize task-appropriate modes of thought". Evidence for this would be that the MD network is responsible for maintaining a particular thought state during a task. To investigate this, they use a combination of fMRI brain data during a set of 14 tasks and experience sampling in a different set of participants performing the same tasks. Using dimensionality reduction, they reduced the space of task features (and brain systems) to a smaller, more tractable set of dimensions and examined whether stability in specific thought components was related to recruitment of specific brain systems during the task.

      Strengths:

      Overall, this is an interesting and creative study with strong analytic methods that do a good job accounting for confounds or alternative explanations (save one I mention below).

      Weaknesses:

      I have mostly minor comments and one major one.

      Major:

      The principal finding is that tasks that evoke brain activity patterns that resemble the MDN also had more stable "deliberate task focus" features. While all the analysis and controls are impressive, I'm still left with the sense that this is reifying something we already know or that alternative explanations are more parsimonious than the MDN induces stability in "thought".

      I thought an example might be easiest to understand my point: If I gave participants a series of working-memory-related tasks. Some of these are the crème de la crème, and others are sloppy and poorly designed. Then suppose I assess them on measures related to deliberate thought; I'd likely find the "good" tasks elicit more consistent/reliable deliberate task focus. I also would bet money that these same tasks would evoke canonical WM and MDN activity patterns more than the sloppy tasks. This isn't evidence of MDN stabilizing patterns, but rather that both stable thought patterns and activity in the MDN share a common cause. Thus, I would predict that with my thought experiment, your analysis would find the same result. So it strikes me as a strong alternative possibility for these results is that tasks that reliably evoke deliberate task focus are also those that more strongly and consistently evoke working-memory demand (i.e., Figure 2 shows that they are primarily driven by the executive/WM tasks).

      Minor:

      (1) The PCA was reviewed previously, and I don't want to relitigate a prior method, but I had one minor concern. It would be useful to know how the principal results are based merely on the "deliberate" or "focus" items specifically. Is the thought space necessary, or do the individual items that likely drive the "deliberate task focus" PC essentially replicate the main result?

      (2) I struggled with the motivation for projecting the task data onto a resting state FC analysis that focuses on "gradients". I understand that with 14 tasks activation maps, data reduction is a good thing. But as someone who isn't as enmeshed in this work, I didn't follow why this specific "atlas" was chosen over any other (parcellations, meta-analytic maps of canonical networks, etc.). Maybe a brief sentence saying why this and not that would help readers who find themselves in my shoes.

    3. Reviewer #2 (Public review):

      Summary:

      The study's aim was to establish whether stability in thought patterns relates to the particular thought pattern, the task context, or their interaction. And further, whether stable thought patterns could be linked with distinct brain patterns.

      Strengths:

      The core reliability framing is novel, and the trait/state/interaction decomposition is a fruitful way to pose the question, leading to the finding that stability is neither a pure trait nor a pure task property, but emerges from their interaction, which is a solid contribution.

      The aim to characterise aspects of stability across individuals and across tasks was achieved and is well supported.

      Weaknesses:

      While the paper makes excellent use of existing data sets, the independent samples, i.e., one study sample for the cognitive/thought-sampling data, and several different study samples to generate the brain maps, do limit the brain-behaviour conclusions that can be drawn.

    1. eLife Assessment

      This important study presents compelling evidence for dopamine-based representation learning in the mouse olfactory tubercle. While this work suggests that gradient descent in deep networks serves as a helpful framework for understanding neural representations, the empirical evidence provided remains incomplete. Nonetheless, this work will be of significant interest to systems and computational neuroscientists, cognitive scientists, and machine learning theorists.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript reports on simultaneous neural recordings in the olfactory tubercle (OTu) and ventral tegmental area in head-fixed mice performing a simple go-nogo odor-guided reversal task. In the task, there were 3 odors that predicted lick-spout water at 0%, 50%, and 100% probability, with 0% and 100% odors reversing at some point each session. The authors fitted this neural data to a value function approximator in which reward prediction errors fed back onto state representations, allowing the optimal set of representations to be learned. They found that model predictions of adjustments to state representations were correlated with trial-by-trial changes in OTu neural activity, from which the authors conclude that such a system, with dopaminergic errors feeding back onto OTu representations of states, which then generate reward predictions, is biologically plausible.

      Strengths:

      This is a novel and creative modeling approach that has important implications. It seems to be showing the biological plausibility of a model that can learn state representations rather than relying on a fixed set of states that are programmed into the model. This makes tremendous sense, because the real world is much less well-defined than the kinds of tasks conventionally used by neuroscientists to probe reinforcement learning. As such, it is an important demonstration.

      Weaknesses:

      The task used in this study is quite simple in its state space, and in particular in how it maps sensory stimuli (odors) onto states, such that the task would seem not to require a system that can learn state representations, or at least that it would not be ideal for testing such a model. This mismatch raises some questions about why this model would perform as well as it appears to be doing here.

      The model has two updating functions, both using dopaminergic RPE's. One of these maps raw stimuli to state representations using a parameter termed theta; the second maps state representations to a value prediction, using a parameter termed w. The interaction of these two updating functions seems to be giving the model its interesting characteristics. But it is critical to test what the first of these updates is doing in this model, given that odor stimuli appear to map straightforwardly onto states. For example, one might test the effect of ablating this part of the model, leaving only updates of what the authors term w. Relatedly, one might test the extent to which OTu neurons show simple odor selectivity before and after reversals, asking whether these neurons reflect state representations in this model merely by being selective for a particular odor, or if they develop a more complex kind of responsiveness.<br /> The authors compare the full model, which uses a gradient descent update, to a series of alternatives. The fact that the full model performs significantly better than any of the alternatives leads to the conclusion that the brain is using something like this model in this task. But these alternative models are all reduced or simplified versions of the primary model. This suggests that the full model is the best version within the basic framework posed by the authors. But to draw the conclusion that this model is capturing what is occurring in the brain, one would want to test how this model would perform compared to a different class of model, in particular one that assumes a fixed set of states.<br /> A second weakness of the paper is that the authors do not show the behavioral or raw neural data, which would be important to summarize for the sake of transparency and to help readers get an intuitive sense of what is going on in the task and why the model performs as well as it does. One essential issue is: how much training do mice receive before neural data used in the analysis are collected? Do mice get pre-exposure to contingency reversals before analyzed neural data is collected? Relatedly, how quickly (i.e., in how many trials) do mice show behavioral evidence of having learned initial contingencies and then reversed contingencies? What is the behavioral criterion? With regard to the number of trials mice take to learn the reversals, this can change enormously over training, and such changes could have a big effect on how the model performs. Regarding the neural data, one would want to show some measure of odor selectivity of SPN's and DAN's and how each population responds to delivery and omission of reward in different conditions.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper, the authors use electrophysiological recordings from the olfactory tubercle (OTu) and ventral tegmental area (VTA) of mice learning an olfactory Pavlovian conditioning task to demonstrate the consistency of changes in OTu odour responses with gradient descent updates minimising reward prediction error (RPE). The paper is clearly written and the work well motivated. The authors address a gap in the literature on animal reinforcement learning by providing neural evidence for gradient-based representation learning, something that had been proposed but not yet tested. The results are convincing and the limitations comprehensively addressed. Of particular interest is the proposal that OTu SPNs could solve the weight transport problem through knowledge of the sign of their downstream connections from their expression of either D1/D2 receptors. This makes a concrete experimental prediction that future research could test.

      Strengths:

      The paper provides one of the first demonstrations backed by neural recordings that representation learning in the brain is consistent with gradient descent. It shows how, although weight transport may be biologically implausible, the brain appears to find other ways to compute a gradient in multilayered networks for efficient learning. The study builds nicely on recent work in systems neuroscience and provides evidence for a concrete implementation in the OTu-VTA circuitry of mice.

      The paper demonstrates that changes in OTu striatal projection neuron (SPN) activity over trials are proportional not only to the RPE relayed by VTA dopamine neurons, but also account for the influence of each particular SPN on the RPE. If an SPN decreases the RPE when active, its activity will increase on the next trial after a positive RPE. The activity will instead decrease for an SPN that increases the RPE. This relation is encapsulated in the update rule of Equation 5.

      Weaknesses:

      The main weakness of the paper is that it provides only indirect evidence for the update mechanism by inferring synaptic weights based on the (justified) assumption of VTA dopamine neurons encoding RPE. This is still a substantial contribution to understanding representation learning in the brain, though I do think that the authors could provide some additional evidence to further convince readers. One idea could be to analyse the distribution of inferred weights and validate whether it agrees with known statistics of connectivity between OTu SPNs and their downstream projections (e.g. fraction of D1/D2 SPNs).

      Further, as dopamine neurons are known to have asymmetrical responses for positive and negative RPEs (with the dips in activity related to negative RPEs being generally smaller) I'd expect an improvement in the correlation of the learning updates particularly after the reversal if the authors account for this in the model.

    4. Reviewer #3 (Public review):

      Most models of reinforcement learning in the brain treat the question of how the external world is represented as an afterthought. In "Error driven representation learning in the mesolimbic system", the authors begin by calling attention to this limitation of previous work, then proceed to show that neural activity in part of the ventral striatum evolves in a manner consistent with dopaminergic RPE-driven representation learning. Overall, the question is interesting, the modeling is well done, and the claims are bold. While I am not convinced that olfactory tubercle outputs _mainly_ reflect state features acquired through error-driven learning (see main points below), I am now more willing to believe that they might. I am confident that this work will spark discussion.

      Strengths:

      The latent weight trajectory inference approach is a nice application of Kalman filtering, the model validation and comparison steps are well executed and thorough, the discussion is nicely written and includes an appropriate caveat about the weight transport problem, and the general idea of striatal output being value-like but also having state-like aspects that evolve over time is thought-provoking.

      Weaknesses:

      In my view, the main limitation of this work is that the authors do not clearly rule out (1) non-representation learning and (2) non-error-driven learning. This fits with the narrow research question stated at the end of the introduction (L71-72), which is confirmatory in nature and does not claim to exclude alternatives, but other aspects of the framing are less consistent.

      Main points of criticism

      (1) Representation vs. value learning

      The first two pages of the manuscript gave me the impression that the authors wish to draw a clear distinction between representation and value learning, and that they would squarely position this paper as a study of representation learning. The substance of the work does not seem consistent with this positioning, a mismatch that could be addressed either by incorporating new analysis and discussion or by changing the framing.

      Examples of emphasis on representation:

      - Abstract L14-17 defines value and representation learning and states why they are different.

      - First three paragraphs of introduction explain the power of learned representations.

      - Results L103-108 attribute state representations to OTu and value to downstream regions.

      For this level of emphasis, it would be good to see a convincing argument that OTu MSN activity is better understood as a state representation than as the output of a value function.

      Options for establishing OTu as state and not value:

      - Explicitly claim in the introduction that previous work has established this, and explain why evidence of stimulus valence being encoded in this area (citations on L68-70) does not favour the value interpretation. Discussing how OTu differs from other parts of ventral striatum that are canonically seen as value coding would also help.

      - Directly compare the extent of state vs. value encoding in the present data. The $w=1$ control is a good step in this direction, but its connection to value coding is mentioned only in passing.

      (2) Error driven learning vs. other types of learning

      Similar to my previous point, the authors seem to claim that the representational changes they study are specifically error driven. While the authors include a good number of controls and ablations, it was not obvious to me that any of them correspond to a form of non-error driven learning that could plausibly generate useful representations. Adding Hebbian learning or a sparsifying learning rule would strengthen this aspect of the work.

    1. eLife Assessment

      This revised submission remains a valuable contribution to influenza surveillance. After further addressing a previously highlighted methodological issue, the claims of providing an early warning tool aligns more with the reported study results and is a solid addition to the evidence base.

    2. Reviewer #2 (Public review):

      [Editors' note: The Reviewing Editor has assessed the revised article without further input from the original reviewers. The Reviewing Editor noted the authors further addressed a methodological concern, and eLife's Assessment remains unchanged from the previous review.]

      Summary:

      The study aimed to assess the associations between meteorological drivers and influenza is important although not new. The authors used 6 years of surveillance data and deep learning models, combining distributed lag non-linear models (DLNM) with Bayesian-optimized LSTM neural networks for predictive modeling. The key interest in this area is to explore the subtropical locations, where influenza is less common and circulates year-round. The authors further claimed that such an association could be able to provide an early warning in the community.

      Strengths:

      Study design based on a prospective cohort to analyse the data for retrospective outcomes.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      The study aimed to assess the associations between meteorological drivers and influenza is important although not new. The authors used 6 years of surveillance data and deep learning models, combining distributed lag non-linear models (DLNM) with Bayesian-optimized LSTM neural networks for predictive modeling. The key interest in this area is to explore the subtropical locations, where influenza is less common and circulates year-round. The authors further claimed that such an association could be able to provide an early warning in the community.

      Strengths:

      Study design based on a prospective cohort to analyse the data for retrospective outcomes.

      We would like to express our sincere and heartfelt gratitude to all of you for your exceptionally thorough, constructive, and intellectually rigorous evaluation of our manuscript. The breadth and depth of the feedback we have received reflect a high standard of scientific scrutiny that we deeply respect and appreciate.

    1. eLife Assessment

      This important study compares how different classes of drugs act on the SARS-CoV-2 main protease, a key antiviral target, and shows that many of them work by controlling whether the enzyme assembles into its active dimeric form. The evidence, based on a range of complementary biophysical methods, is convincing and points to the interface between the two protein protomers, including a newly found binding site, as a promising target for broad-spectrum antiviral drugs. This work will be of interest to biochemists and virologists working on treatments for coronaviruses.

    2. Reviewer #2 (Public review):

      Summary:

      This manuscript presents a sophisticated investigation into the mechanisms by which different inhibitor classes affect the SARS-CoV-2 main protease (Mpro), a pivotal antiviral drug target. This study reveals that effective inhibition can be achieved by modulating the stabilization of the essential dimeric state. It also indicates the dimer interface could be a druggable allosteric site, which may offer a strategy for developing broad-spectrum anticoronaviral agents.

      Strengths:

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      Comments on latest version:

      All of my concerns have been adequately addressed.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Reviewer #1 (Public review):

      Summary:

      Since dimerization is essential for SARS-CoV-2 Mpro enzymatic activity, the authors investigated how different classes of inhibitors, including peptidomimetic inhibitors (PF-07321332, PF-00835231, GC376, boceprevir), non-peptidomimetic inhibitors (carmofur, ebselen, and its analog MR6-31-2), and allosteric inhibitors (AT7519 and pelitinib), influence the Mpro monomer-dimer equilibrium using native mass spectrometry. Further analyses with isotope labeling, HDX-MS, and MD simulations examined subunit exchange and conformational dynamics. Distinct inhibitory mechanisms were identified: peptidomimetic inhibitors stabilized dimerization and suppressed subunit exchange and structural flexibility, whereas ebselen covalently bound to a newly identified site at C300, disrupting dimerization and increasing conformational dynamics. This study provides detailed mechanistic evidence of how Mpro inhibitors modulate dimerization and structural dynamics. The newly identified covalently binding site C300 represents novelty as a druggable allosteric hotspot.

      Strengths:

      This manuscript investigates how different classes of inhibitors modulate SARS-CoV-2 main protease dimerization and structural dynamics, and identifies a newly observed covalent binding site for ebselen.

      Weaknesses:

      None. The requested mutagenesis data have been provided in the revised manuscript, and all of my previous concerns have been satisfactorily addressed.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      None. The overall quality of the manuscript has been substantially improved. The authors have added supportive mutagenesis data in the revised manuscript to validate the proposed role of C300. All of my concerns have been adequately addressed.

      We appreciate the reviewer’s recognition of the improvements made in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents a sophisticated investigation into the mechanisms by which different inhibitor classes affect the SARS-CoV-2 main protease (Mpro), a pivotal antiviral drug target. This study reveals that effective inhibition can be achieved by modulating the stabilization of the essential dimeric state. It also indicates the dimer interface could be a druggable allosteric site, which may offer a strategy for developing broad-spectrum anticoronaviral agents.

      Strengths:

      The identification of dimer interface stabilization/destabilization as distinct inhibitory mechanisms and the discovery of C300 as a potential allosteric site for ebselen are important contributions to the field. The experimental approach is modern, multi-faceted, and generally well-executed.

      Comments on revised version:

      The authors have very nicely addressed most of the previous comments raised. But one comment remains to be clarified relating to original point 5 and the authors' response:

      "We agree with the reviewer about the need for quantitative rigor in reporting HDX changes. We have calculated the fractional deuterium uptake difference for each peptide fragment discussed in the text between the inhibitor-bound and unbound states. These values, along with their statistical significance (p-values from a two-tailed t-test), have been provided in the revised manuscript (Legends for Figures 3 and 4). Although the HDX change of residues 296-306 is relatively small (<5%), this region showed a reproducible difference with low experimental variability and statistical significance (p < 0.05). Given its location within the C-terminal dimerization interface and its consistency with native MS, we interpret this change as a subtle local conformational perturbation."

      Two questions remain for the statements in line 376-380. First, while it is stated "residues 296-304 in the C-terminal region of Mpro were more flexible upon ebselen binding", the segment of 296-306 is shown Figure 4c. Second, the HDX change for this segment upon ebselen binding is very subtle in the figure (in contrast to the significant HDX change of the same segment in the protein upon PF-07321332 binding), thus making the strong conclusion that "This suggests that ebselen targeting C300 may induce structural changes in the C-terminal helical segment, weakening key hydrogen bonds at the dimer interface and ultimately inhibiting activity" not convincing. The reviewer would suggest the authors either delete this conclusion or largely tone it down.

      We thank the reviewer for the recognition of our efforts and agree with the reviewer’s suggestion. We have corrected “residues 296–304” to “residues 296–306” in Line 377 and removed the statement “This suggests that ebselen targeting C300 may induce structural changes in the C-terminal helical segment, weakening key hydrogen bonds at the dimer interface and ultimately inhibiting activity.”, as suggested.

    1. eLife Assessment

      This study presents important findings that genome-wide DNA double-strand breaks trigger large-scale genomic amplification (DIGA), which does not require origin re-licensing, but is dependent on some proteins that are involved in break-induced replication (BIR). While the findings and model are compelling, the evidence supporting the claims is largely incomplete, as it depends primarily on a single assay to examine amplification. This work will be of interest to scientists in the DNA repair, replication, and cell cycle fields, as well as cancer biologists.

    2. Reviewer #1 (Public review):

      Summary:

      In this report, the authors investigate the mechanisms underlying large-scale genomic amplification (DIGA) induced by genome-wide DNA double-strand breaks (DSBs), such as those generated by ionizing radiation (IR).

      Strengths:

      The authors demonstrate that DSB-induced DIGA does not require origin re-licensing but is dependent on proteins involved in break-induced replication (BIR). This finding represents a major strength of the study, as it reveals a previously unrecognized mechanism of DSB-induced genomic amplification. Additional strengths include the demonstration that DIGA is promoted by DNA end resection and suppressed by the 53BP1-RIF1-shieldin pathway. The authors also show that SET8 and SUV4-20H1 have opposing effects on CDT1 overexpression-induced and IR-induced re-replication. Furthermore, the finding that the extent of DIGA in cancer cells correlates with sensitivity to IR has potential implications for cancer treatment.

      Weaknesses:

      However, more comprehensive studies are needed to strengthen the conclusions. Although IR- or DSB-induced DIGA was observed in multiple cancer cell lines, the overall mechanisms underlying why DIGA is more pronounced in certain cell lines but not others remain unclear. Beyond p53 status, additional factors that determine DIGA susceptibility should be investigated. In addition, the evidence that DIGA-associated DNA synthesis occurs within the same cell cycle is not yet sufficient. For instance, a time-course experiment (EdU versus DAPI) should be performed after IR to determine when DIGA initiates relative to normal S-phase DNA replication. A parallel analysis of re-replication at different time points following MLN4924 treatment would provide a useful comparison. Cyclin B and phospho-histone H3 (Ser10) levels should be monitored to define cell cycle stage. Additional evidence supporting a BIR-dependent mechanism, such as demonstrating conservative DNA synthesis and mapping DIGA sites following AsiSI-induced DSBs, would be needed.

      Overall, this study provides new insights into the mechanisms driving genome-wide amplification following DSB formation. Additional mechanistic studies will further strengthen the conclusions and broaden the impact of this work.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Benamar et al. investigate whether DNA double-strand breaks induce extensive abnormal DNA synthesis in cancer cells and whether this response contributes to the cytotoxic effects of ionizing radiation and other DNA-damaging treatments. The authors refer to this process as double-strand-break-induced genomic amplification (DIGA). Mechanistically, they propose that insufficient protection of broken DNA ends permits excessive resection, followed by RAD51-dependent strand invasion and RAD52-POLD3-POLD4-dependent DNA synthesis resembling break-induced replication.

      Overall, the study addresses an interesting and potentially important question. Break-induced replication was originally defined as a pathway that repairs one-ended DNA breaks through extended synthesis from an invaded homologous template. Studies in yeast established that this process involves a migrating DNA-synthesis bubble, conservative inheritance of newly synthesized DNA, frequent template switching, and high mutagenicity. Related forms of DNA synthesis have subsequently been described in mammalian cells at collapsed replication forks, telomeres, under-replicated mitotic regions, and some transcription-associated DNA breaks.

      The important advance of this study is the proposal that double-strand-break-associated DNA synthesis, with several features of break-induced replication, can become sufficiently extensive to cause a measurable increase in total cellular DNA content and that this synthesis correlates with radiation-induced cell death. However, the physical structure, genomic distribution, and extent of the additional DNA have not yet been directly established.

      Strengths:

      Radiation-induced breaks are generally considered in relation to DNA repair, chromosome rearrangements, checkpoint activation, mitotic failure, senescence, and cell death. The possibility that broken DNA ends can also initiate extensive DNA synthesis provides a potentially important additional link between defective break repair, genome amplification, and treatment-induced cytotoxicity.

      The authors examine this phenomenon using several complementary approaches and comparisons across multiple cancer cell lines. The finding that several double-strand-break-inducing treatments produce a similar response suggests that the phenotype is not restricted to ionizing radiation or to a single cellular background.

      The authors also make efforts to distinguish DIGA from canonical origin-dependent re-replication, such as that caused by CDT1 stabilization and inappropriate origin relicensing. They find that, following irradiation, CDT1 is degraded rather than stabilized and that CDT1 depletion does not suppress DIGA. Perturbation of ORC1 or ORC2 also has little effect on the phenotype. In addition, depletion of SET8 or loss of SUV4-20H1 increases, rather than decreases, DIGA.

      The authors carry out a systematic analysis of DNA-end protection and processing. Loss of ATM, RNF8, RNF168, 53BP1, RIF1, or Shieldin components enhances DIGA, whereas depletion or inhibition of CtIP, MRE11, EXO1, RAD51, RAD52, POLD3, or POLD4 suppresses it. These experiments support a model in which insufficient end protection permits excessive DNA-end resection, followed by strand invasion and recombination-associated DNA synthesis involving factors linked to break-induced replication.

      Overall, the findings place DIGA within a growing body of evidence that recombination-associated DNA synthesis can extend well beyond short repair patches. The study attempts to connect the processing of double-strand breaks with abnormal DNA synthesis, increased cellular DNA content, and the cytotoxic response to radiation.

      Major weaknesses and suggested experiments:

      (1) The main evidence for DIGA is the appearance of cells with greater-than-G2/M DNA content by flow cytometry. BrdU incorporation shows that active DNA synthesis occurs within this population, and the density-gradient experiments provide further evidence for newly synthesized DNA. However, these methods do not reveal the physical nature of the proposed genomic amplification-for example, which genomic regions are copied or how long the synthesis tracts are.

      This is important for the central conclusion of the study because mammalian break-induced replication can produce genomic duplications, but the extent of synthesis depends strongly on the type of DNA lesion. Previous work has shown that repair of damaged replication forks can produce segmental genomic duplications (Costantino et al., 2014; PMID: 24310611). By contrast, recent measurements at defined two-ended double-strand breaks suggest that mammalian break-induced-replication tracts can be relatively restricted (Li et al., 2021; PMID: 33470420; Shah et al., 2024; PMID: 39368985).

      Thus, the large increase in total DNA content detected by flow cytometry in this study would require many simultaneous synthesis events, very long synthesis tracts, or an alternative process such as whole-genome duplication. The authors should therefore characterize the additional DNA directly. Whole-genome sequencing of sorted cells with greater-than-G2/M DNA content could be particularly informative.

      (2) The genetic dependencies are consistent with synthesis initiated from resected DNA ends, but they do not directly demonstrate that the new DNA synthesis begins at double-strand breaks. DNA damage could indirectly alter replication-origin activity, cell-cycle progression, or DNA synthesis at genomic regions distant from the original lesions. Increased DNA content alone therefore does not establish amplification initiated directly at DNA breaks.

      (3) MLN4924-induced re-replication is used throughout the study, but it is unclear whether the authors have combined MLN4924 with ionizing radiation and measured the resulting DNA synthesis. This experiment could clarify whether conventional re-replication and DIGA are independent, overlapping, or mechanistically connected.

      (4) The results involving canonical non-homologous end joining are complex. Loss or inhibition of DNA-PKcs increases DIGA, whereas loss of XRCC4 or XLF strongly suppresses it, and loss or inhibition of LIG4 has a weaker suppressive effect. The authors suggest that XRCC4 and XLF stabilize broken DNA ends and thereby permit the synthesis reaction independently of final ligation. This is an interesting model, but the current evidence does not yet establish it. XRCC4 and XLF can affect end synapsis, break persistence, chromosome fusion, resection, and cell-cycle progression. Their loss could therefore reduce DIGA through several indirect mechanisms. The authors should measure DNA-end resection, RAD51 loading, persistence of double-strand breaks, and recruitment of RAD52 or POLD3 in XRCC4-, XLF-, and LIG4-deficient cells. Complementation with separation-of-function mutants that differentially affect XRCC4-XLF end bridging and LIG4 recruitment could test whether end stabilization, rather than ligation, is important. Without this, the opposing effects of upstream and downstream non-homologous end-joining components remain difficult to understand.

      (5) The correlation between the amount of DIGA and radiation sensitivity across cancer cell lines is interesting. However, cell lines differ in many factors that influence radiation responses, including p53 status, apoptosis, checkpoint activity, ploidy, homologous recombination capacity, and proliferation rate. Matched models would provide stronger evidence. Conversely, restoring DNA-end protection in a DIGA-prone cancer cell line could test whether this reduces susceptibility. These experiments would help determine whether DIGA is a general property of cancer cells or a feature of particular repair-defective genetic backgrounds.

    4. Reviewer #3 (Public review):

      Summary:

      In this study, the authors show that exposure of cancer cells to DSBs induced enzymatically or by ionizing radiation (IR) results in large-scale genomic amplifications (DIGA: DSB-induced genomic amplifications) that are detectable by FACS. This phenomenon was observed in diverse cancer cell lines and was shown to be limited by p53 in paired cell lines. In principle, DIGA could result from re-initiation of DNA synthesis within the same cell cycle, mitotic segregation errors, or repair-associated DNA synthesis. The authors conduct experiments to distinguish between these possibilities and conclude that DIGA is the result of extensive RAD51-dependent DNA synthesis.

      Strengths:

      The observation of genome amplification in response to DSBs specifically in cancer cells is striking, particularly at high IR doses or by AsiSI endonuclease induction. At 9 Gy, almost 50% of cells have a >4N DNA content. The correlation between high levels of DIGA and reduced clonogenic survival of melanoma cells suggests that DIGA contributes to IR-induced cytotoxicity. The authors convincingly show that DIGA occurs in a single cell cycle and is not due to chromosome mis-segregation at mitosis. Because DIGA is partially dependent on resection nucleases, POLD3, POLD4, RAD52 and RAD51, the authors conclude that DIGA results from repair synthesis by a BIR-like process.

      There are significant weaknesses in the study:

      (1) It is unclear to this reviewer how BIR-like synthesis could increase genomic DNA copy number by 50% or more, as indicated by the FACS plots. In the case of AsiSI, there are ~150 sites that are efficiently cleaved in U2OS cells. Each of these sites would need to prime extensive synthesis by an inefficient migrating D-loop mechanism. Studies of BIR-like synthesis at DSBs in U2OS cells have shown that repair tracts are fairly short, around 3-10 kb in length. Thus, it is hard to rationalize how BIR at a few hundred DSBs could initiate such large changes in genome size.

      (2) The BrdU-FACS plot shown in Figure S4 shows that most of the cells have an 8N content 48 and 72 h after AsiSI induction and do not have much BrdU incorporated. This finding would suggest that genomic DNA doubling is mostly independent of de novo synthesis. Also, there does not seem to be a continuum from 4N up to 8N in these plots as one might expect from variable length DNA synthesis tracts.

      (3) The requirement for XRCC4, XLF and LIG4 to promote DIGA is puzzling if synthesis occurs by a BIR-like mechanism. Although there is evidence that DNA synthesis at DSBs can be initiated by homology-directed strand invasion and terminated by NHEJ, this mechanism would limit the extent of synthesis to a few kb at each site. One might have expected an increase in DIGA in the absence of core NHEJ factors because of the increased number of DSBs available for BIR.

      (4) Although the authors rule out re-replication as a contributor to IR-induced DIGA, the FACS profiles of MLN4924 and 9 Gy treatment look remarkably similar, and both would be inhibited by aphidicolin treatment. A previous study showed that DSBs induce endoreduplication in Arabidopsis (Adachi et al. PNAS 2011).

    5. Author response:

      We thank the reviewers for their careful and constructive evaluation of our study. We are encouraged that the reviewers recognized the significance of DSB-induced genomic amplification (DIGA) and the evidence implicating DNA-end processing and recombination-associated DNA synthesis in this response.

      To our knowledge, this study provides the first description of DIGA as a large-scale increase in genomic DNA content following the induction of DSBs in cancer cells and represents an initial effort to define factors that regulate this phenomenon. The present work shows that DIGA can be induced by several sources of DSBs, involves de novo DNA synthesis, is genetically distinguishable from canonical CDT1-dependent origin re-licensing, is regulated by pathways controlling DNA-end protection and resection, and requires RAD51, RAD52, POLD3, and POLD4.

      At the same time, we agree that many important questions remain regarding the physical organization and genomic distribution of the additional DNA, the sites from which synthesis originates, the length and number of synthesis tracts, and the full determinants that render some cancer cells more susceptible to DIGA than others. We view these as important questions that arise from the initial characterization of this previously unrecognized phenotype and that will require substantial additional investigation.

      Reviewer #1:

      We agree that p53 status alone does not explain the considerable variation in DIGA observed among the cancer cell lines examined. Our experiments using isogenic HCT116 cells identify p53 as one factor capable of limiting DIGA, while the genetic studies implicate DNA-end protection and resection pathways as additional determinants. The present data, however, do not establish which of these or other pathways account for the differences among individual cancer cell lines. Defining the molecular basis for this variability will require systematic comparison of DIGA-prone and DIGA-resistant cells.

      The reviewer also raises an important question regarding the temporal relationship between DIGA and normal S-phase DNA replication. Our conclusion that the increase in DNA content involves de novo synthesis within the same cell-cycle interval is supported by BrdU incorporation in cells with >4N DNA content, the persistence of DIGA when progression through mitosis is blocked by nocodazole, and the detection of newly synthesized DNA in synchronized irradiated cells. These experiments do not, however, define precisely when DIGA-associated synthesis begins relative to normal S-phase replication. More detailed time-resolved analysis will be required to establish this relationship.

      We also agree that the present findings support a BIR-like mechanism rather than providing a complete physical reconstruction of classical BIR. The dependence of DIGA on DNA-end resection, RAD51, RAD52, POLD3, and POLD4 provides genetic evidence for recombination-associated DNA synthesis with features of BIR. Direct determination of synthesis-tract architecture, template usage, and genomic distribution will be required to define the underlying synthesis mechanism more completely.

      Reviewer #2:

      We agree that direct characterization of the additional DNA represents an important next step in understanding DIGA. The current study demonstrates a substantial increase in cellular DNA content, de novo DNA synthesis within the >4N population, and dependence on factors involved in DNA-end resection, strand invasion, and BIR-associated synthesis. These experiments do not determine which genomic regions are amplified or the length of individual synthesis tracts. Genomic analysis of cells undergoing DIGA should help determine whether the observed increase in DNA content reflects numerous amplification events, extensive synthesis from a subset of sites, or a different organization of the additional DNA.

      We appreciate the reviewer highlighting the study by Costantino et al. (2014), which provided important evidence that BIR-associated repair of damaged replication forks can generate segmental genomic duplications in human cells. The genomic alterations characterized in that study, however, arose under a substantially different experimental setting. Costantino et al. induced replication stress through cyclin E overexpression and analyzed copy-number alterations accumulated over a three-week period in clonally derived cells. Among these alterations, amplifications smaller than 200 kb were reduced following depletion of POLD3 or POLD4, leading the authors to propose that this subset of segmental duplications may represent BIR events, whereas larger amplifications and deletions could involve other repair mechanisms.

      DIGA, however, differs from these previously described alterations in several readily observable respects. DIGA develops over approximately one to three days following acute induction of DSBs by IR or AsiSI and produces increases in total cellular DNA content sufficiently large to be detected directly by flow cytometry. Thus, the two phenomena differ in their mode of induction, kinetics, and scale. At the same time, the involvement of POLD3 and other recombination-associated factors in both settings raises the possibility that they share aspects of the underlying DNA-synthesis machinery. The present data do not establish whether the additional DNA in DIGA consists of numerous segmental duplications, substantially longer synthesis products, or another genomic configuration.

      We also agree that our experiments do not directly demonstrate that DIGA-associated DNA synthesis initiates precisely at individual DSB sites. The ability of AsiSI-generated DSBs to induce DIGA, together with its dependence on DNA-end resection, RAD51, RAD52, POLD3, and POLD4, links the phenomenon closely to DSB processing. Direct mapping of newly synthesized DNA relative to defined DSBs will ultimately be required to determine where DIGA-associated synthesis originates.

      The reviewer asks whether MLN4924-induced re-replication and DIGA have been examined simultaneously. We have not examined this combination. Our distinction between these processes instead rests on their different genetic requirements. In particular, depletion of CDT1 strongly suppresses MLN4924-induced re-replication but does not suppress IR-induced DIGA, and DIGA is stimulated while rereplication is inhibited by the depletion of SET8. These observations argue against canonical CDT1-dependent origin re-licensing as the mechanism underlying DIGA, although they do not exclude more complex interactions between replication and DSB-associated DNA synthesis.

      We agree that the effects of XRCC4, XLF, and LIG4 are mechanistically intriguing and not yet fully understood. The present experiments establish that loss of XRCC4 or XLF, and to a lesser extent LIG4, suppresses DIGA, whereas loss or inhibition of DNA-PKcs enhances it. Stabilization of broken DNA ends by XRCC4/XLF is one possible interpretation, but effects on end resection, DSB persistence, repair-pathway choice, or other functions of these proteins could also contribute. The opposing effects of different components of the NHEJ machinery therefore identify an important mechanistic question that remains to be resolved.

      Finally, we agree that the correlation between DIGA and radiation sensitivity across the melanoma cell-line panel does not by itself demonstrate causality. The data establish an association between the propensity to undergo DIGA and sensitivity to IR. Because these cell lines differ in multiple additional properties that may influence the radiation response, matched models in which DIGA can be selectively altered will be important for determining the extent to which DIGA itself contributes to radiation-induced loss of proliferative capacity.

      Reviewer #3:

      We agree that the magnitude of the increase in DNA content is one of the most interesting unresolved features of DIGA. Previous analyses of BIR-associated synthesis at defined mammalian lesions have generally described synthesis events considerably smaller than the total increase in DNA content observed here. Our experiments do not establish the length or number of individual synthesis events responsible for DIGA. We therefore use the term BIR-like to describe the genetic requirements of the process rather than to imply that each DSB gives rise to a single exceptionally long BIR tract. Determining how many genomic sites participate and how much DNA is synthesized at individual sites will be important for understanding how the large increase in total DNA content is generated.

      Regarding the AsiSI BrdU experiments, it is important to note that BrdU was provided as a one-hour pulse immediately before harvesting. BrdU signal at 48, 72, or 96 hours therefore reports DNA synthesis occurring during that particular one-hour interval and does not measure the cumulative DNA synthesis that preceded the measurement. Consequently, relatively modest BrdU incorporation in cells that have already accumulated high DNA content does not indicate that the preceding increase occurred independently of DNA synthesis. Conversely, these experiments alone do not define the physical mechanism by which the additional DNA accumulated.

      We agree that the requirement for XRCC4, XLF, and LIG4 is unexpected under a simple model of BIR. As noted above, the present experiments establish this genetic relationship but do not define its molecular basis. The differential effects of DNA-PKcs and downstream NHEJ factors suggest that individual NHEJ components may influence DIGA through functions that are not adequately represented by viewing the pathway simply as a linear ligation reaction.

      Finally, we agree that the similarity between the flow-cytometric profiles produced by IR and MLN4924, as well as their common sensitivity to aphidicolin, does not by itself distinguish the underlying mechanisms. The distinction in the current study instead derives from their different genetic requirements, particularly the dependence of MLN4924-induced re-replication on CDT1 compared with the lack of such a requirement for DIGA, together with the differential effects of factors involved in DNA-end protection, resection, and DSB repair. These observations argue that DIGA is not simply the consequence of canonical origin re-licensing (i.e., rereplication or endoreduplication), while leaving open the possibility that additional replication-associated mechanisms contribute to the phenotype.

      In summary, we appreciate the reviewers highlighting several important mechanistic questions raised by our findings. We view these questions as natural extensions of the initial discovery and characterization of DIGA. The present study identifies a large-scale DSB-associated increase in genomic DNA content and establishes important roles for DNA-end protection, resection, strand invasion, and BIR-associated factors in regulating this response. Determining the genomic architecture of the additional DNA, the sites and molecular intermediates from which synthesis originates, and the cellular determinants of DIGA susceptibility will be important goals for future studies.

    1. eLife Assessment

      This valuable paper by Dong et al. describes an analysis of mutant phenotypes of the Rab GTPases Rab5, Rab7 and Rab11 in Drosophila second order olfactory neuron development. The revised version presents a convincing characterization and comparison of the different rab mutants on projection neuron development, with clear differences for the three rabs and by inference for the early, late and recycling endosomal functions executed by each.

    2. Reviewer #1 (Public review):

      Summary:

      Dong et al. present an in-depth analysis of mutant phenotypes of the Rab GTPases Rab5, Rab7, and Rab11 in Drosophila second-order olfactory neuron development. These three Rab GTPases are amongst the best-characterized Rab GTPases in eukaryotes and have been associated with major roles in early endosomes, late endosomes, and recycling endosomes, respectively. All three have been investigated in Drosophila neurons before; however, this study provides the most detailed characterization and comparison of mutant phenotypes for axonal and dendritic development of fly projection neurons to date. In addition, the authors provide excellent high-resolution data on the distribution of each of the three Rabs in developmental analyses.

      Strengths:

      The strength of the work lies in the detailed characterization and comparison of the different Rab mutants on projection neuron development, with clear differences for the three Rabs and by inference for the early, late, and recycling endosomal functions executed by each.

      Comments on revised version.

      The authors conducted extensive revision experiments, especially to characterize developmental defects. Efforts to identify cargoes were not successful. The evidence is now convincing.

    3. Reviewer #2 (Public review):

      Summary:

      This study by Dong et al characterizes the roles of highly-expressed Rab GTPases Rab5, Rab7 and Rab11 in the development and wiring of olfactory projection neurons in Drosophila. This convincing descriptive study provides complementary approaches of Rab expression and localization profiling, conventional dominant negative mutants, and clonal loss of function mutants to address the roles of different endosomal trafficking pathways across circuit development. They show distinct distributions and phenotypes for different Rabs. Overall, the study sets the stage for future mechanistic studies in this well-defined central neuron.

      Strengths:

      Beautiful imaging in central neurons demonstrates differential roles of 3 key Rab proteins in neuronal morphogenesis as well as interesting patterns of subcellular endosome distribution. These descriptions will be critical for future mechanistic studies. The manuscript is well-written and explanatory, very accessible to a wide audience without sacrificing technical accuracy.

      Comments on revised version.

      The paper is greatly improved with new data and better-nuanced interpretation. The clarity of explanations for a non-Drosophila reader have been redone particularly well.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aimed at a comprehensive phenotypic characterization of the roles of all Rab proteins expressed in PN neurons in developing Drosophila olfactory system. Important data are shown for a number of these Rabs with small/no phenotypes (in the Supplements) as well as the main endosomal Rabs, Rab5, 7, and 11 in the main figures.

      Strengths:

      The mosaic analysis is a great strength allowing visualization of small clones or single neuron morphologies. This also allows some assessment of cell autonomy of the observed phenotypes. The impact of the work lies in the comprehensiveness of the experiments. The rescue experiments are a strength. The added developmental data strengthen the impact of the paper.

      Weaknesses:

      The main weakness is that the experiments do not address the mechanisms that are affected by the loss of these Rab proteins, especially in terms of the most significant cargos. The insights thus do not extend far beyond what is already known from other work in many systems.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dong et al. present an in-depth analysis of mutant phenotypes of the Rab GTPases Rab5, Rab7, and Rab11 in Drosophila second-order olfactory neuron development. These three Rab GTPases are amongst the bestcharacterized Rab GTPases in eukaryotes and have been associated with major roles in early endosomes, late endosomes, and recycling endosomes, respectively. All three have been investigated in Drosophila neurons before; however, this study provides the most detailed characterization and comparison of mutant phenotypes for axonal and dendritic development of fly projection neurons to date. In addition, the authors provide excellent high-resolution data on the distribution of each of the three Rabs in developmental analyses.

      Strengths:

      The strength of the work lies in the detailed characterization and comparison of the different Rab mutants on projection neuron development, with clear differences for the three Rabs and by inference for the early, late, and recycling endosomal functions executed by each.

      Weaknesses:

      Some weakness derives from the fact that Rab5, Rab7, and Rab11 are, as acknowledged by the authors, somewhat pleiotropic, and their actual roles in projection neuron development are not addressed beyond the characterization of (mostly adult) mutant phenotypes and developmental expression.

      We would like to thank Reviewer #1 for their appreciation of our characterization of distinct Rab mutants.

      Reviewer #2 (Public review):

      Summary:

      This study by Dong et al. characterizes the roles of highly-expressed Rab GTPases Rab5, Rab7, and Rab11 in the development and wiring of olfactory projection neurons in Drosophila. This convincing descriptive study provides complementary approaches to Rab expression and localization profiling, conventional dominantnegative mutants, and clonal loss-of-function mutants to address the roles of different endosomal trafficking pathways across circuit development. They show distinct distributions and phenotypes for different Rabs. Overall, the study sets the stage for future mechanistic studies in this well-defined central neuron.

      Strengths:

      Beautiful imaging in central neurons demonstrates differential roles of 3 key Rab proteins in neuronal morphogenesis, as well as interesting patterns of subcellular endosome distribution. These descriptions will be critical for future mechanistic studies. The cell biology is well-written and explanatory, very accessible to a wide audience without sacrificing technical accuracy.

      Weaknesses:

      The Drosophila manipulations require more explanation in the main text to reach a wide audience.

      We appreciate Reviewer #2’s analysis of our work and thank them for their suggestions to improve the clarity of our manuscript.

      Reviewer #3 (Public review):

      Summary:

      The authors aimed at a comprehensive phenotypic characterization of the roles of all Rab proteins expressed in PN neurons in the developing Drosophila olfactory system. Important data are shown for a number of these Rabs with small/no phenotypes (in the Supplements) as well as the main endosomal Rabs, Rab5, 7, and 11 in the main figures.

      Strengths:

      The mosaic analysis is a great strength, allowing visualization of small clones or single neuron morphologies. This also allows some assessment of the cell autonomy of the observed phenotypes. The impact of the work lies in the comprehensiveness of the experiments. The rescue experiments are a strength.

      Weaknesses:

      The main weakness is that the experiments do not address the mechanisms that are affected by the loss of these Rab proteins, especially in terms of the most significant cargos. The insights thus do not extend far beyond what is already known from other work in many systems.

      We thank this reviewer for their feedback and appreciation of our genetic manipulations.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Consensus suggestions after discussion of all three reviewers:

      All three reviewers agree that the morphological and phenotypic analysis of the fly olfactory neurons is a strength of the manuscript. The shared perceived weakness is that the experiments do not address the mechanisms that are affected by the loss of these Rab proteins, especially in terms of the most significant cargos; the findings are in line with a large body of literature.

      The three reviewers feel that the manuscript could be strengthened greatly by adding data on an actual cargo (cell surface proteins?) and a more detailed analysis of the actual developmental origin (what happens when during axon and dendrite development) with respect to sorting of such cargo in the neurons they analyzed.

      We appreciate the time and effort of all three of our reviewers and share their interest in both identifying Rab-regulated cargos as well as determining the developmental origins of the Rab phenotypes. We have added three additional main figures (new Figure 4, Figure 8, and Figure 9), two supplemental figures (Figure 1—figure supplement 1 and 2), two additional supplemental tables (Table S2 and S3), and five additional panels (in Figure 3 H–L) of mutant developmental phenotype analysis.

      Regarding cargos, we also share the reviewers’ desire to identify cargos regulated by each Rab and made attempts to do so but were ultimately unable to achieve this goal. The main obstacles to this were: (1) it is not known which cell-surface proteins are most robustly endocytosed in PNs; without this knowledge it is difficult to identify candidates whose localization would predominantly reflect endosomal rather than plasma membrane distribution, making it challenging to detect changes in compartment-specific localization upon Rab perturbation; (2) reagents to evaluate cell-surface proteins in PNs are not cell-type-specific making it difficult to evaluate changes in their distribution in PNs; (3) tagged overexpressed proteins are either unavailable or expressed at levels too high to sensitively detect changes in their distribution. We have elaborated on each of these points below and feel that cargo identification, while an important future direction, is beyond the scope of the present study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      There are a number of experiments and ideas that the authors might consider to further improve on this work.

      (1) The idea, introduced by the authors in the introduction, that Rab-mediated recycling of cell surface proteins back to the membrane versus degradation is, of course, excellent and interesting. It is less clear how this applies to the present study. The functions of Rab5, Rab7, and Rab11 are so widespread, potentially affecting many signaling roles resulting in primary or secondary effects on membrane and even cytoskeletal regulation, that it remains unclear whether the mutant phenotypes are related to the recycling or degradation of cell surface receptors. To link the idea to experimentation, the authors have an excellent opportunity in their system to look at endogenously tagged cell surface proteins (or at least one example), many of which the Luo lab has characterized in these neurons, to minimally correlate cell surface protein defects to the observed developmental defects.

      We understand this critique and share this reviewer’s interest in identifying the specific cargos regulated by each Rab during development. We attempted to use antibodies to evaluate changes in cell-surface protein localization in response to disrupting individual Rabs but were unable to reliably distinguish shifts in association with specific endosomal compartment as many available antibodies label cell-surface proteins expressed in antennal lobe cells beyond projection neurons (such as olfactory receptor neurons, glia, or local interneurons) which complicates analyses.

      Additionally, although we have, in other work, generated multiple 'flp-on' tags for PN cell-surface proteins, these cannot be used in combination with the MARCM system, as it relies on a heat-shock-inducible flp to label singlePN clones. Heat shock would simultaneously induce tag expression in other cells expressing the tagged gene, preventing PN-type-specific detection. This incompatibility thus prevents us from simultaneously perturbing individual Rabs and tracking corresponding changes in surface-protein localization with single-cell resolution.

      Moreover, for proteins that are not highly endocytosed, it is difficult to separate plasma-membrane from endosomal localization, and we currently do not know which cell-surface proteins are most robustly endocytosed in PNs. Thus, while we share the reviewer’s interest in identifying candidate cargos, technological limitations make it difficult to achieve this goal within the scope of the current study.

      (2) The mutant phenotypes are mostly characterized based on adult outcomes. Maybe a little more can be learned about when and how Rab5, Rab7, or Rab11 function is locally required by characterizing the developmental processes that lead to, e.g., aberrant dendritic development in Rab5 and Rab11.

      We also feel that charting the developmental origins of Rab mutant phenotypes is important. Prior to mid-pupal stage (around 48 hours after puparium formation), glomeruli in the antennal lobe have not yet assumed their stereotyped positions, which complicates analyses and interpretation; thus, many of our analyses are conducted at the adult stage. For Rab11 mutants we did perform many developmental analyses to evaluate the origins of the axonal development (Figure 6—figure supplement 1) and dendrite elaboration phenotypes (Figure 5 J–L) we observed at the adult stage. We realize that the developing axonal analyses were in supplemental material where they could be missed. We have moved these data to the main figures (Figures 8 and 9) and emphasized these analyses. Further, we extended our Rab5 mutant analyses to evaluate developmental phenotypes (Figure 3H– L and Figure 4). We believe that these new analyses have strengthened the manuscript.

      (3) Regarding the subcellular localization analyses: a collection of endogenously tagged Rabs in Drosophila has been generated by Dunst et al. (2015), which is surprisingly not cited. Maybe the authors could consider looking at the endogenous localization of Rabs using the tagged version in parallel to their overexpressed tagged versions.

      We have now cited and discussed this paper (line 82) and thank the reviewer for pointing out this omission. We previously attempted to evaluate these endogenously tagged Rab proteins in PNs; however, since PN dendrites project into the antennal lobe, a dense neuropil region containing PN dendrites, ORN axons, glial processes, and neurites of local interneurons, we are unable to resolve individual Rab puncta from cytosolic (non-vesicle associated Rabs) or evaluate Rab localization in a cell-type-specific manner. For this reason, we focused on evaluating the localization of tagged Rab proteins from UAS-transgenes using a MARCM rescue strategy. We directly addressed this in the text (starting on line 83).

      (4) It is maybe not entirely surprising that Rab5 and Rab11 have the strongest phenotypes, as these have been implicated in early and recycling endosomal processes in basically all eukaryotic cells with major implications for signaling throughout development and function, often causing cell death (and in the case of Rab5, tumorigenic phenotypes in flies). By contrast, the Drosophila brain can develop in the absence of Rab7 (Cherry et al., 2013; also the reference for the Rab7 null mutant, not Chan et al., 2011). A key concern in any developing fly cell rendered mutant using clonal analysis is the perdurance of RNA or protein (ultimately even maternal contribution), which could be addressed by discussion or experimentally.

      We thank the reviewer for pointing out our citation error, which we have now corrected.

      However, we note that Cherry et al. (2013) found that loss of Rab7 causes pupal lethality at stages prior to completion of 50–80% of development and can also cause embryonic lethality when maternal Rab7 contribution is blocked. This indicates that the whole organism cannot fully develop in the absence of Rab7. And while Cherry and colleagues did evaluate overall brain morphology in Rab7 mutant pupae, they did not look at the development of individual cell types. So, it is still unclear how loss of this GTPase affects the development of individual central nervous system neurons.

      Thus, to understand whether Rab7 has functions in PN development, we used the QMARCM system to perform Rab7 LOF analyses in PN clones. While we did not observe any phenotypes in single-cell MARCM clones (Figure 6), we did see mild defects in neuroblast clones (Figure 6—figure supplement 1A–C). Since single-cell MARCM clones are more susceptible to RNA/protein perdurance, we further evaluated Rab7 function by expressing a Rab7 dominant-negative transgene in DL1-PNs using a DL1-specific GAL4 driver (Figure 6—figure supplement 2), which circumvents potential perdurance issues mentioned by this reviewer. Importantly, this same transgene produces dendrite targeting defects when expressed in all PNs (Figure 1I), confirming its efficacy. However, no phenotypes were observed when expression was restricted to DL1-PNs, suggesting that Rab7 may not be required in DL1-PNs for their dendrite targeting. Given that both Rab7 mutant neuroblast clones and pan-PN expression of Rab7 dominant negative causes PN dendrite targeting defects we conclude that Rab7 is nonautonomously required for dendrite targeting of DL1-PNs.

      We have softened our language with regards to the Rab7 analysis and have emphasized, and strengthened, our previous discussion of these points beginning on line 268 in the results section and on line 427 of the discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 1B, it would be useful to show the circuit over multiple developmental stages, rather than just in its final form.

      We have added this to Figure 1; it is now panel C. Thank you for this suggestion.

      (2) Expression analysis of Rabs in Figure 1C - how do these levels and ratios compare to the whole brain? Whole body?

      Unfortunately, we are unable to evaluate how Rab expression in PNs compares to all other cells in the brain as there is no sequencing data available for this organ at this time point. We did compare the expression of endosomal Rabs between PNs and their presynaptic targets, ORNs. We found that many Rabs displayed similar expression patterns between these two cell types during development. We have added a new paragraph on this, beginning on line 100 and we added two additional supplemental figures (Figure 1—figure supplement 1 and 2).

      (3) The authors should include at least a few sentences comparing the current approach and results to previous comprehensive Rab protein expression analysis, for example, in PMID 17409086, 22000105, 22844416, and 33666175.

      Thank you for pointing out this omission, we have amended it beginning on line 82.

      (4) For the non-Drosophila reader (for example, a cell biologist working on endosomal traffic in cultured neurons), the paper is less accessible. Some examples:

      We thank this reviewer for their suggestions for ways to clarify our work for the non-Drosophila reader. We have addressed each of their points.

      (a) The severity difference between Rab5, Rab11, Rab7 and Rab4, Rab 21, Rab35 isn't immediately obvious from the images to someone who doesn't work with this system - does the brightness of the ectopic growths indicate the number of ectopically grown processes? It might help to have half a sentence to make this difference more accessible for the readers who aren't familiar with this system.

      We have clarified this beginning on line 112.

      (b) Figures 1D-E require more extensive description of the experimental setup with orthogonal expression systems than is provided briefly in the cartoon, figure legend, and supplement. For example, it should be noted what white vs blue represents in the marked glomeruli.

      We have clarified this point beginning on line 114.

      (c) There should be at least one sentence introducing what is marked and what it means when MARCM clones are first shown in Figure 2B, in addition to the supplemental figure.

      We have added a detailed explanation of MARCM on line 155.

      (d) It's not clear to a non-expert what the meaning is of no innervation of non-adPN glomeruli in wild-type in Figure 2E. This requires a sentence of explanation.

      We have added additional details about this on line 159 and 169.

      (e) Can a control image be shown for the experiment in 3B?

      We have added an additional set of control images in Figure 3B on the left.

      (5) The experiment measuring axonal projection to the lateral horn in Rab5 clones in Figure 3 J-L is underpowered (n=3 for mutant). While this may be due to the frequency of an overall projection defect as shown in Figure 3B, it makes it difficult to assess the robustness of the terminal phenotype. Further, for clarity, similar measurements (e.g., bouton width) should be aligned vertically between E-G and J-L.

      We have performed additional analyses on Rab5 axons in the lateral horn and added them to a new Figure 4. The n’s are now n=10 for controls and n=7 for Rab5 mutants. Additionally, we have aligned similar measurements in the figure panels and standardized the axes of each graph so that it is easier to compare between developmental stages.

      (6) The argument that cell-type-specific phenotypes are due to distinct cargoes is weak. The same cargo could have different functions or signaling properties in different cell types (e.g., "Taken together, the distinct branching phenotypes observed in the mushroom body versus lateral horn suggest that Rab5 may regulate the trafficking of a distinct set of cargos in each axonal compartment"). Similarly, this argument is just one of many possibilities, as the effect could be quite indirect (eg via mis-regulated signal transduction): "Yet, the terminal boutons in Rab5 mutants were nearly 2-fold larger than those of controls (Figure 3G), suggesting Rab5 regulates the trafficking of cell-surface proteins that normally restrain bouton growth " and "Page 12 "Rab7mediated degradation does not have a major role in regulating axon or dendrite development" - should be softened since Rab7 may easily play an important but redundant role.

      We have made all of these changes and removed references to trafficking of specific CSPs.

      (7) Statistics need to be added to: Figure 3B, Figure 5D-E, Figure 2 - figure supplement 1C, Figure 4 - figure supplement 1E, F, Figure 5 - figure supplement 1A, E, Figure 6 - figure supplement 1K.

      We thank this reviewer for pointing out this omission, we have added statistical measurements to our graphs.

      As to not visually overwhelm readers with statistical measurements on already dense graphs (such as Figure 7B), we have added two supplemental tables (Table S2 and S3) that display all the results of all of the comparisons performed in the statistical tests.

      We have cited this table in the figure legends and in-text figure references.

      In addition to the methods, we have also added the exact statistical test and post-hoc tests (when applicable) to the figure legends.

      (8) The BSDC identifier for the UAS-Rab11-mCherry stock may be incorrect.

      It appears that some of the values in the ‘Identifiers’ column of the Key Resources table shifted downward. We have fixed this and appreciate that the reviewer pointed this out.

    1. eLife Assessment

      This fundamental study delivers a population reference panel for long-read sequencing-based structural variants and demonstrates its utility for disease association analyses in the UK Biobank, expanding genetic discovery beyond conventional SNV-based approaches. The authors provide convincing evidence through extensive benchmarking and systematic analyses that the panel improves structural variant detection and can support fine-mapping of trait associations. The broader impact will depend on timely access to the imputed UK Biobank resource, or on provision of a practical, reproducible pipeline and associated summary statistics.

    2. Reviewer #1 (Public review):

      Summary:

      The authors sequenced 888 individuals from the 1000 Genomes Project using the Oxford Nanopore long-read sequencing method to achieve highly sensitive, genome-wide detection of structural variants (SVs) at the population level. They conducted solid benchmarking of SV calling and systematically characterized the identified SVs. While short-read sequencing methods, including those used in the 1000 Genomes Project, have been widely applied, they exhibit high accuracy in detecting single nucleotide variants (SNVs) and small insertions and deletions but have limited sensitivity for SV detection. This study significantly enhances SV detection capabilities, establishing it as a valuable resource for human genetic research. Furthermore, the authors constructed an SV imputation panel using the generated data and imputed SVs in 488,130 individuals from the UK Biobank. They then conducted a proof-of-principle genome-wide association study (GWAS) analysis based on the imputed SVs and selected traits within the UK Biobank. Their findings demonstrate that incorporating SV-GWAS analysis provides additional insights beyond conventional GWAS frameworks focusing on SNVs, particularly in improving fine-mapping.

      Strengths:

      The authors constructed a high-sensitivity reference panel of genome-wide SVs at the population level, addressing a critical gap in the field of human genetics. This resource is expected to significantly advance research in human genetics. They demonstrated the imputation of SVs in individuals from the UK Biobank using this panel and conducted a proof-of-concept SV-based GWAS. Their findings highlight a novel and effective strategy for integrating SVs into GWAS, which will facilitate the analysis of human genetic data from the UK Biobank and other datasets. Their conclusions are supported by comprehensive analyses.

      Weaknesses:

      The authors have addressed many of my previous comments, and I appreciate their efforts. However, I still have two related concerns.

      (1) Shortly after reviewing this manuscript last year, my laboratory obtained access to the UK Biobank (UKB) Tier 3 dataset for an unrelated project. In August 2025, I searched the UKB Research Analysis Platform (UKB-RAP) for the imputed structural variant (SV) dataset described in this manuscript but was unable to locate it. After contacting UKB, I was informed that they were developing the system for releasing the data. To the best of my knowledge, the dataset remains unavailable. A major contribution of this work is the generation of an imputed SV resource for approximately 500,000 UKB participants with extensive phenotypic information. If this resource is not accessible to the research community, even to authorized UKB users, the practical impact and utility of the study are substantially diminished.

      (2) Given that the imputed SV dataset is currently unavailable, it becomes even more important for the authors to provide a detailed, ready-to-run SV imputation pipeline for UKB-RAP, even if the "data processing simply consisted of running standard bioinformatics tools with the parameters exactly as described in the manuscript". In particular, the pipeline should include practical information such as computational requirements (e.g., memory and storage), expected running time, and estimated cost. Anyone with experience using UKB-RAP will agree that reproducing large-scale analyses on the platform can be both technically complex and financially expensive. Such pipeline would greatly improve the reproducibility and accessibility of this work.

      Because my initial assessment of the manuscript was generally positive, I do not wish to change my overall evaluation, summary, or assessment of its strengths. However, I would view the work even more favorably if either (i) the imputed SV dataset became publicly available to authorized UKB users, or (ii) the authors extended their SV-GWAS analyses to the full range of UKB phenotypes and released the resulting summary statistics, analogous to the Pan UKBB ("https://pan.ukbb.broadinstitute.org/") resource. Although this would require considerable additional effort and computational resources, it would substantially enhance the long-term value and impact of the study.

      Finally, I would like to emphasize that these comments are not intended to create unnecessary difficulties for the authors or the editors. Rather, I believe this highlights a broader issue in the use of this kind of large public datasets: reviewers cannot independently verify key results, and readers cannot readily build upon the work if the underlying resources are inaccessible, even after obtaining authorized access to the original dataset. I hope the authors, together with the eLife editors and UK Biobank where appropriate, can help facilitate the timely release of this valuable resource.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Our revision includes:

      (1) The generated data from long-read whole-genome sequencing of 1000 Genomes Project samples, including FASTQ files, SV calls, and the imputation panel, are now openly available via ENA and OpnMe. The imputed structural variant data have been submitted to UK Biobank for release through the UK Biobank Research Analysis Platform, subject to UK Biobank release procedures. SV-WAS summary statistics have been made available via OpnMe.

      (2) Clarification of analyses and methods, addition of two new Supplementary Figures, and correction of minor issues throughout the manuscript.

      (3) A significantly expanded Discussion to address the reviewers’ comments and better contextualise our methods and results.

      eLife Assessment

      This fundamental work significantly enhances our understanding of how structural variants influence human phenotypes. The conclusion is convincingly supported by rigorous analyses of long-read sequencing data. If the raw data are made publicly available, these high-quality datasets and findings will further advance our knowledge of genetic variation in the human population.

      We thank the editors for this positive assessment of our work. The raw long-read sequencing data (FASTQ files) can now be accessed through the European Nucleotide Archive (ENA) under accession number PRJEB89727, as part of a larger collection of 1019 sequenced probands from the 1000 Genomes Project (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The generated imputation panel and the structural variant calls, based on the 888 probands used in the present manuscript, remain freely available for download at https://opnme.com/genomiclens. We have now added the summary statistics of 32 SV-wide association studies to the same resource. In addition, we have submitted the imputed SV genotypes for UK Biobank participants to the UK Biobank; once processed by UK Biobank, these genotypes will be released via the UK Biobank Research Analysis Platform (RAP).

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors sequenced 888 individuals from the 1000 Genomes Project using the Oxford Nanopore long-read sequencing method to achieve highly sensitive, genome-wide detection of structural variants (SVs) at the population level. They conducted solid benchmarking of SV calling and systematically characterized the identified SVs. While short-read sequencing methods, including those used in the 1000 Genomes Project, have been widely applied, they exhibit high accuracy in detecting single nucleotide variants (SNVs) and small insertions and deletions but have limited sensitivity for SV detection. This study significantly enhances SV detection capabilities, establishing it as a valuable resource for human genetic research. Furthermore, the authors constructed an SV imputation panel using the generated data and imputed SVs in 488,130 individuals from the UK Biobank. They then conducted a proof-of-principle genome-wide association study (GWAS) analysis based on the imputed SVs and selected traits within the UK Biobank. Their findings demonstrate that incorporating SV-GWAS analysis provides additional insights beyond conventional GWAS frameworks focusing on SNVs, particularly in improving fine mapping.

      The authors constructed a high-sensitivity reference panel of genome-wide SVs at the population level, addressing a critical gap in the field of human genetics. This resource is expected to significantly advance research in human genetics. They demonstrated the imputation of SVs in individuals from the UK Biobank using this panel and conducted a proof-of-concept SV-based GWAS. Their findings highlight a novel and effective strategy for integrating SVs into GWAS, which will facilitate the analysis of human genetic data from the UK Biobank and other datasets. Their conclusions are supported by comprehensive analyses.

      We thank the reviewer for highlighting the value of our SV imputation reference panel.

      Weaknesses:

      (1) Although the authors employ state-of-the-art analytical approaches for the identification of SVs, the overall accuracy remains suboptimal, as indicated by an F1 score of 74.0%, particularly in tandem repeat regions. To enhance accuracy, it would be beneficial to explore alternative SV detection methods or develop novel approaches. Given the value of the reference panel and the fact that improved SV accuracy would lead to more precise SV imputation and GWAS results, investing effort in methodological refinement is highly encouraged.

      Accurate SV calling remains an active area of research and is beyond the scope of the present study. Tandem repeat regions are particularly challenging for standardised SV detection. We believe that achieving a benchmark for NA12878 of F1 = 74% on a genome-wide level and, notably, F1 = 91% when excluding longer tandem repeats, represents strong performance. This result is especially convincing when considering that our benchmarking compared the SV calls to data generated using a different sequencing technology and processed using different bioinformatics pipelines.

      (2) From the Methods section, it appears that the authors employed Beagle for both imputation and the UK Biobank imputation.

      (a) It would be better to explicitly clarify this in the Results section and provide a detailed description of the corresponding procedures and parameters in the Methods section for both analyses, as this represents a key aspect of the study.

      We thank the reviewer for these suggestions. Accordingly, we added the clarification to the Methods section that the leave-one-out imputation used exactly the same pipeline and settings as the UK Biobank imputation (page 14, section “Leave-one-out imputation performance”):

      “We excluded one individual from the panel and imputed SVs for this individual using the panel of the remaining 887 samples, applying exactly the same pipeline and settings as those later used for SV imputation into UK Biobank (see below).”

      (b) Additionally, Beagle is not specifically designed for SV imputation, the imputation quality of SVs is generally lower than that of SNVs. Exploring strategies to improve SV imputation, such as developing a novel method with reference panel data, may enhance performance.

      As stated in our manuscript (page 4), we believe that, in our study, the imputation quality of SVs is lower than of that of SNVs primarily because of a) the greater difficulty of SV calling compared to SNV genotype calling and b) the heterogeneity in SV representation across samples. Once SVs are encoded as bi-allelic markers in the reference panel, they can be imputed using the same LD/haplotype-based framework as any other variants. Accordingly, improving SV imputation is likely to benefit most from more accurate upstream SV calling and genotyping (e.g., through more robust multi-sample calling and harmonised variant representations) and not so much from improved or SV-specific imputation methods. While improved imputation is an important research direction, it is beyond the scope of the present manuscript.

      (c) It is also important to assess how this reduced imputation quality may influence GWAS results. For instance, it would be useful to examine whether associated SVs exhibit higher imputation quality and whether SVs with lower quality are less likely to achieve significant association signals. In addition, the lower imputation quality observed for INV, DUP, and BND variants (Figure 3) may be due to their greater lengths (Figure 2). It is better to investigate the relationship between SV length and imputation quality.

      We agree that imputation quality can influence GWAS results. For example, for the FEV1/FVC phenotype, SVs with INFO > 0.9 are almost twice as likely to reach genome-wide significance (p < 5e-8) compared with SVs with 0.7 < INFO < 0.9 (odds ratio 1.95; Fisher’s exact test p-value 2.5e-5). This is consistent with the intuitive notion (applicable to any variants, not only to SVs) that greater uncertainty in the imputed genotypes dilutes association signals and therefore reduces power. For a detailed discussion of the relationship between allele frequency, imputation accuracy, and GWAS association results, see Zhang et al., Human Molecular Genetics 31(1):146–155 (2022), https://doi.org/10.1093/hmg/ddab203

      We have now investigated the relationship between SV length and imputation quality (the new Supplementary Figure 6). The results suggest that the observed association between imputation quality and SV size is primarily driven by the SV-size–dependent minor allele frequency in the imputation panel.

      (3) All examples presented in the manuscript focus on SVs that overlap with genes. It may also be valuable to investigate SVs that do not overlap with genes but intersect with enhancer regions. SVs can contribute to disease by altering regulatory elements, such as enhancers, which play a crucial role in gene expression. Including such analyses would further demonstrate the utility of SV-GWAS and provide deeper insights into the functional impact of SVs.

      We agree with the reviewer that examining SVs intersecting with enhancer regions could be an interesting direction for future studies, as it would provide additional insights into regulatory mechanisms and disease associations. However, in the present proof-of-principle study, we prefer focusing on SVs overlapping with genes and have now highlighted this in additional detail in the revised manuscript (Discussion, page 7):

      “In the present proof-of-principle study, we focused on SVs overlapping with the coding sequence of genes. In future applications of our SV imputation panel, more refined gene mapping approaches could be employed, e.g., including SVs overlapping enhancer regions or epigenetic marks. Such an enhanced mapping would increase the number of identified associated genes and thus provide additional insights into regulatory mechanisms and disease biology.”

      (4) The data availability link currently provides only a VCF file ("sniffles2_joint_sv_calls.vcf.gz") containing the identified SVs.

      (a) It would be beneficial for the authors to make all raw sequencing data (FASTQ files) and key processed datasets (such as alignment results and merged SV and SNV files) available. Providing these resources would enable other researchers to develop improved SV detection and imputation methods or conduct further genetic analyses.

      Thank you for emphasising the importance of data sharing, which we agree with.

      The Data Availability section of the manuscript already includes a link to https://opnme.com/genomiclens, where we made both the SV calls and the full and reduced SV imputation panel files freely available. We have now added SV summary statistics from 32 SVwide association studies to the same resource. Based on the reviewer’s request, we now also reference the ENA repository project PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727), where the raw FASTQ files are available for download, in the manuscript.

      We have appended the Data Availability statement on page 22 of the revised manuscript as follows:

      “Raw SV calls, the long-read sequencing-based SV imputation panel, and the SV summary statistics from 32 SV-wide association studies are available through the OpnMe initiative of Boehringer Ingelheim GmbH (https://opnme.com/genomiclens). The raw long-read sequencing data (FASTQ files) for the 1000 Genomes Project samples included in this study are accessible via the European Nucleotide Archive under accession number PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The dataset analysed here constitutes a subset of this broader collection.”

      (b) Furthermore, establishing a dedicated website for data access, along with a genome browser for SV visualization, could significantly enhance the impact and accessibility of the study. Additionally, all code, particularly the SV imputation pipeline accompanied by a detailed tutorial, should be deposited in a public repository such as GitHub. This would support researchers in imputing SVs and conducting SV-GWAS on their own datasets.

      The Methods section provides a full and detailed description of the imputation pipeline and parameters in the section “Preprocessing and imputation of SVs into UK Biobank” on page 14 of the revised manuscript. Our data processing simply consisted of running standard bioinformatics tools with the parameters exactly as described in the manuscript.

      Reviewer #1 (Recommendations for the authors):

      (1) In the Results section, Figure 3b is mentioned before 3a, and it is better to switch them in the Figure.

      Thank you for highlighting this fact. We acknowledge that typically the sequence of sections matches exactly between text and figures. However, in this specific case, we would prefer to deviate from the norm: In our opinion, Figure 3 is easier to interpret in its current sequence. At the same time, the text flows more logically in its current sequence, describing 3b before 3a. We would therefore prefer to stick to the current order, even if it means that Fig. 3b is described before 3a in the text.

      (2) Page 10, "Figure 1e" -> "Figure 2e".

      Thank you, we corrected this issue.

      (3) Page 14, "Leave-one out" -> "Leave-one-out".

      Thank you, we corrected this mistake.

      (4) It is better not to use abbreviations in the subheadings, especially "UKB" (page 3).

      Thank you, we changed the acronym ‘UKB’ to ‘UK Biobank’ in all subheadings.

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to develop a novel and efficient method for SV detection, utilizing data from the 1000 Genomes Project (1KGP) for modeling and calibration. This method was subsequently validated using UK population data and applied to identify structural variants associated with specific disease phenotypes.

      Strengths:

      Third-generation single-molecule sequencing data offers several advantages over traditional high-throughput sequencing methods, particularly due to its long-read lengths, which provide valuable insights into significant forms of genomic variation. The authors have developed an efficient method for detecting structural variations and optimizing the utilization of genomic data. We hope that this method will continue to be refined, enabling researchers to more effectively leverage long-read data, high-throughput data, or even a synergistic combination of both.

      Weaknesses:

      Although this research contributes to our ability to more effectively utilize long-length and high-throughput data, there are some key issues that need to be addressed in terms of analyzing the specific results as well as writing the article.

      Reviewer #2 (Recommendations for the authors):

      (1) How to discuss the lower detection rate of structural variations (SVs) in East Asian populations, it is worth considering whether the authors' training dataset, which may have been based on raw data with insufficient representation of East Asian individuals, could have introduced a bias favoring other populations. This potential bias might arise from the relatively limited data available for Asian ancestry. Alternatively, the observed differences could also be influenced by the role of natural selection, which may have shaped the genomic landscape of East Asian populations in distinct ways. Further investigation is needed to clarify these possibilities.

      Thank you for raising this important point. Although an interesting research direction, a detailed investigation of the factors affecting SV detection rates is beyond the scope of the present study. However, we do not think that the lower detection rate in East Asians is due to an underrepresentation of Asian ancestry in our dataset. To explain this to all readers, we have added the following explanation to page 7 of the Discussion:

      “In this context, we observed that the number of SVs detected per individual differed between superpopulations. We identified the highest average number of SVs in individuals of African descent and a slightly lower average in East Asians, compared to the other superpopulations. While we included a higher number of African ancestry individuals, the number of East Asian individuals included in our reference panel was comparable to the number of individuals from other, non-African ancestries. In fact, it was even larger than the number of European ancestry individuals (AFR n=241, SAS n=171; EAS n=168; EUR n=164; AMR n=144). Therefore, we do not expect a major bias from underrepresentation of any superpopulation in the training dataset. It is well established that African ancestry is more diverse than is the case for other superpopulations [32, 33] and previous studies indicate that East Asian populations tend to exhibit slightly lower genetic diversity compared to European populations [34], which is consistent with the lower observed SV counts per genome.”

      (2) The authors did not present the results of the detection of CNV.

      Copy number variations (CNVs) are considered a subclass of structural variants. In our analysis, we detected deletions and duplications, which represent the most common forms of CNVs. However, we did not specifically investigate high copy-number SVs, as these are often larger than what can be reliably detected using long-read sequencing. Large-scale CNVs are typically identified in biobank studies through analysis of intensity data from genotyping microarrays using tools like PennCNV, and there is extensive literature supporting the use of this microarray approach in UK Biobank and other genotyped cohorts, see for example Aguirre et al.: Phenomewide Burden of Copy-Number Variation in the UK Biobank. Am J Hum Genet. 2019, Aug 1;105(2):373-383. doi:10.1016/j.ajhg.2019.07.001.

      (3) Multiple testing correction is essential for ensuring the validity of large-scale structural variation (SV) association analyses. It is strongly recommended that the statistical methods and correction strategies employed, such as Bonferroni correction or false discovery rate (FDR) control, be explicitly detailed to enhance the transparency and reliability of the findings.

      For genome-wide SV association analyses, we applied the commonly used genome-wide significance threshold of 5e-8, which is standard in genome-wide studies. Given that these were exploratory proof-of-principle analyses illustrating use cases for SV analyses, we decided not to correct on top of that for multiple testing for the number of traits (32) tested. For the pQTL analyses, we further adjusted this threshold using a Bonferroni-type correction based on the number of proteins tested (1,463), to account for the increased number of multiple comparisons.

      We have now added a more detailed description of this multiple testing procedure to the Methods subsection “SV-wide association studies in UK Biobank” on page 16 of the revised manuscript:

      “In the exploratory SV-WAS, we used the standard threshold for genome-wide significance of p < 5×10<sup>-8</sup>. For the pQTL analyses, we applied Bonferroni correction for multiple testing on top of that genome-wide threshold, correcting for the number of tested protein levels (n=1463): p < 5×10<sup>-8</sup>/ 1463 = 3.4×10<sup>-11</sup>.”

      (4) The study primarily relied on data from the 1000 Genomes Project (1KGP) and the UK Biobank; however, the UK Biobank cohort is predominantly composed of individuals of European ancestry, which may restrict the generalizability of the research findings to other populations.

      Our reference panel was constructed to cover multiple ancestries, enabling imputation for diverse populations. Thus, our imputation panel can be applied to biobanks around the world and is freely available for this purpose. As a proof of principle, we have demonstrated the feasibility and performance of SV imputation in UK Biobank as an example of a broadly accessible cohort. We are looking forward to biobanks from diverse ancestries downloading our imputation panel and applying it to their populations.

      (5) Although the study employed long-read sequencing technology, the validation of structural variation (SV) detection accuracy predominantly relied on internal data, such as 'leave-one-out' validation. To further strengthen the reliability of the SV detection methods, it is recommended to incorporate additional external independent datasets for validation.

      The leave-one-out procedure in our study was used to validate the imputation performance, not the accuracy of SV detection. To assess SV calling accuracy, we performed extensive benchmarking against external SV call datasets derived from PacBio long-read sequencing and Illumina short-read sequencing. These details are provided under the subheading ‘Structural variant calling and benchmarking’ in the Results section on page 2 of the manuscript.

      (6) Some of the SVs mentioned in the study overlap with disease association loci in the GWAS Catalog, but functional annotation and exploration of the biological mechanisms of these SVs are more limited. It is suggested that LD can be added to analyse whether there are SNP that are highly linked to them to further explore their functions.

      We thank the reviewer for this suggestion. We have actually conducted an analysis addressing exactly this question: We performed conditional association analyses of the SV signals with nearby short variants (SNPs and InDels) at the SV locus. Such a conditional analysis addresses whether the observed SV association is influenced by LD-correlated SNPs or not. The results of this analysis are reported in Supplementary Tables 16 and 17. These tables include both the conditional analysis results and the LD between each SV and the variant at the locus with the second-highest evidence for an association.

      Researchers interested in exploring the biological significance of the SV-WAS results in more detail can now download the full SV-WAS summary statistics from https://opnme.com/genomiclens.

      (7) The discussion section could be further expanded to explore the role of SV in complex diseases and its potential application in precision medicine. For example, it could discuss how SV information can be integrated into existing GWAS frameworks to enhance the accuracy of disease risk prediction.

      Thank you for the suggestion, we have now added the following sentences to the discussion (page 7/8):

      “Structural variants can influence complex disease biology through either the disruption of coding sequence or an altered regulation of gene expression. Such effects may not be well captured by short variants alone. Incorporating SVs into GWAS and follow-up analyses would thus provide more accurate disease risk prediction, uncover underlying pathomechanisms by highlighting actionable pathways and targets, and support precision medicine by providing biomarkers for patient stratification.”

      (8) The geographic labeling of certain samples in Figure 2 appears to contain inaccuracies. For instance, the CDX sample, which represents the Dai population from Xishuangbanna in China's Yunnan Province, is currently mislabeled as originating from China's Inner Mongolia. This discrepancy should be corrected to ensure the accuracy of the data representation.

      We apologise for the misunderstanding. The geographic map in Figure 2a serves as an illustrative mapping of the samples to countries. It is intended to provide readers with an overview of population coverage, rather than to indicate the precise geographic origins of individual populations. The populations CDX, CHB, and CHS are displayed within the outline of China in alphabetical order, without any intention to indicate their exact geographic origin. We changed the respective figure caption to make this clear (page 19 of the revised manuscript):

      “Map of the 888 samples from the 1000 Genomes project, mapping the samples to countries and not indicating detailed geographical origins of populations.”

      Reviewer #3 (Public review):

      This study successfully identified genetic loci associated with various traits by generating large-scale long-read sequencing data from a diverse set of samples. This study is significant because it not only produces large-scale long-read genome sequencing data but also demonstrates its application in actual genetics research. Given its potential utility in various fields, this study is expected to make a valuable contribution to the academic community and to this journal. However, there are several critical aspects that could be improved. Below are specific comments for consideration.

      Strengths:

      Producing high-quality, large-scale variant datasets and imputation datasets

      Weaknesses:

      (1) Data availability

      Currently, it appears that only the Genomic Lens SV Panel is available on the webpage described in the Data Availability section. It is unclear whether the authors intend to release the raw sequencing data. Since the study utilized samples from the 1000 Genomes Project, there should be no restriction on making the data publicly accessible. Given this, would the authors consider making the raw sequencing reads publicly available? If so, NCBI SRA or EBI ENA would be the most appropriate repositories for data deposition. I strongly encourage the authors to consider public data release. Additionally, accessing the Genomic Lens SV Panel data does not seem straightforward. The manuscript should provide a more detailed description of how researchers can access and utilize these data. In my opinion, the best approach would be to upload the variant data (VCF files) to a public database such as the European Variation Archive (EVA) hosted by EBI.

      I strongly request that the authors publicly deposit the variant data. At a minimum:

      (a) The joint genotype data for all 888 samples from the 1000 Genomes Project must be publicly available.

      Thank you for emphasising the importance of data sharing, which we agree with.

      The Data Availability section of the manuscript already includes a link to https://opnme.com/genomiclens, where we make both the SV calls and the full and reduced SV imputation panels (provided as multi-sample VCF files) freely available. Based on the reviewer’s request, we now also reference the ENA repository project PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727), where the raw FASTQ files are available for download.

      We have appended the Data Availability statement on page 22 of the revised manuscript as follows:

      “Raw SV calls, the long-read sequencing-based SV imputation panel, and the SV summary statistics from 32 SV-wide association studies are available through the OpnMe initiative of Boehringer Ingelheim GmbH (https://opnme.com/genomiclens). The raw long-read sequencing data (FASTQ files) for the 1000 Genomes Project samples included in this study are accessible via the European Nucleotide Archive under accession number PRJEB89727 (https://www.ebi.ac.uk/ena/browser/view/PRJEB89727). The dataset analysed here constitutes a subset of this broader collection.”

      (b) For the UK Biobank samples, at least allele frequency data should be disclosed.

      Supplementary Table 5 includes the allele frequencies of the SVs imputed into UK Biobank.

      (c) Since eLife has a well-established data-sharing policy, compliance with these guidelines is essential for publication in this journal.

      By sharing the FASTQ files, the SV calls, the SV imputation panels, the SV summary statistics, and (once processed by UK Biobank) the genotypes of SVs imputed into UK Biobank, we are providing all SV data generated in our study.

      (2) Long-read sequencing data quality

      While the manuscript presents N50 read length and mean or median read base quality for each sample in a table, it would be highly beneficial to visualize these data in figures as well. A violin plot or similar visualization summarizing these distributions would significantly improve data presentation.

      Notably, the base quality of ONT long-read sequencing data appears lower than expected. This may be attributed to the use of pore version 9.4.1, but the unexpectedly low base quality still warrants attention. It would be helpful to include a small figure within Figure 2 to illustrate this point. A visual representation of read length distribution and base quality distribution would strengthen the manuscript.

      We thank the reviewer for this suggestion. We have now included two violin plots (the new Supplementary Figure 1) to the revised manuscript, summarising a) the N50 read length per sequencing run and b) the median read quality per sequencing run. These plots provide a clearer visualisation of the underlying distributions. We do not consider the ONT base quality to be low. Importantly, structural variant detection is generally robust to modest variations of per-base quality. Therefore, we do not expect the observed base quality levels to significantly affect SV calling in this study.

      (3) Variant detection precision, recall, and F1 score

      This study focuses on insertions and deletions (indels) {greater than or equal to}50 bp, but it remains unclear how well variants <50 bp are detected. I am particularly interested in the precision, recall, and F1 score for variants between 5-49 bp.

      While ONT base quality is relatively low, single-base variants are challenging to analyze, but variants {greater than or equal to}5 bp should still be detectable as their read accuracy is still approximately 90%, making analysis feasible. Given that Sniffles supports the detection of variants as small as 1 bp, I strongly encourage the authors to conduct an additional analysis.

      A simple two-category classification (e.g., 5-49 bp and {greater than or equal to}50 bp) should suffice. Additionally, a comparative analysis with HiFi and short-read sequencing data would be highly valuable. If possible, I strongly recommend that all detected variants {greater than or equal to}5 bp be made publicly available as VCF files.

      Because short InDels are available from high-coverage Illumina sequencing data generated for the same individuals (i.e., the data referred to as the NYGC dataset in our manuscript), we decided against calling such short variants from our lower coverage Oxford Nanopore data and thus concentrated our efforts on reliably calling longer variants covering at least 50 bp, consistent with the conventional definition of structural variants.

      (4) Assembly-based methods

      Given the low read accuracy and low sequencing depth in this dataset, it is understandable that genome assembly is challenging. However, the latest high-quality human genome datasets-such as those produced by the Human Pangenome Reference Consortium (HPRC)demonstrate that assembly-based approaches provide significant advantages, particularly for resolving complex and long structural variants.

      Since HPRC data also utilize 1000 Genomes Project samples, it would be highly informative to compare the accuracy of ONT sequencing in this study with HPRC's assembly-based genome data. The recent publication on 47 HPRC samples provides a valuable reference for such a comparison. Given its relevance, the authors should consider providing a comparative analysis with HPRC data.

      The aim of the present study was to generate an SV reference panel that enables SV imputation for large biobanks. Detailed assessments of ONT sequencing quality in general and comparisons to other sequencing efforts and technologies are out of scope for the present manuscript. We invite the scientific community to use the FASTQ files provided at ENA for conducting such detailed assessments in follow-up studies.

    1. eLife Assessment

      Muenker and colleagues use an optical tweezer setup to apply oscillatory forces to endocytosed/phagocytosed glass beads over a wide frequency range (from ~1 to 1000 Hz) and probe cytoplasmic material properties at multiple time scales in six different cell types. Using statistical methods and principal component analysis, they find that the active and passive mechanical properties of cells can be described by 6 parameters (from power law fits) that allow characterizing the viscous and elastic nature of the cytoplasmic material as well as an effective active energy driven by cellular metabolism. Overall, this is a very well done and important work, using compelling and state-of-the-art methods.

    2. Reviewer #1 (Public review):

      Summary:

      In this MS, Muenker and colleagues, explore the intracellular mechanics of a range of animal adherent cells. The study is based on the use of an optical tweezer set up, which allows to apply oscillatory forces on endocytosed/phagocytosed glass beads with a large frequency range (from ~1 to 1000 Hz) , allowing to probe cytoplasm material properties at multiple time scales. By switching off the laser trap, the authors also record the positional fluctuations of beads, to extract passive rheological signatures. The combination of both methods allow to fit 6 parameters (from power law fits) that allow to characterize the viscous and elastic nature of the cytoplasm material as well as an effective active energy driven by cellular metabolism. Using these methodologies, the authors first establish/confirm, using HeLa cells, that the cytoplasm is more solid like at short frequencies, and more fluid like at higher frequencies, and that these material states depend on both microtubules and actin cytoskeleton. The manuscript then goes on to explore how these parameters evolve in other 6 cell types including muscles, highly migratory and epithelial cells. These results show for instance that muscle cells are much stiffer, while migratory cells are more fluid like with an increased active energy. Finally using statistical methods and principal component analysis , the authors establish some mechanical fingerprints (activity, fluidity and resistance) that allow to distinguish cell's mechanical state and relate it to their particular functions.

      Strengths:

      Overall, this is a very well executed work, which provides a large body of rigorous numbers and data to understand the regulation of cytoplasm mechanics and its relation to cell state/function. This work opens up on the possibility to systematically link cellular phenotype and cytoskeleton organization to intracellular mechanical signatures among many cell types and contexts.

    3. Reviewer #2 (Public review):

      Summary:

      By analyzing cells' frequency-dependent viscoelastic properties and intracellular activity through microrheology, Münker et al simplify the complex active mechanical state into six key parameters that constitute the mechanical fingerprint. They apply this concept to cells treated with cytoskeleton-inhibiting drugs. Additionally, a comprehensive statistical analysis across various cell types shows how cells coordinate their mechanical properties within a defined phase-space marked by activity, mechanical resistance, and fluidity.

      Strengths:

      (1) The distribution of the six parameters: they have been well characterized based on established theories, and they can be used to understand cell-type-specific biomechanical differences. The examples of muscle cells and immune cells were profound and informative.<br /> (2) Efforts to perform dimension reduction of parameter space into activity (E), fluidity (C1) and resistance (A) are insightful and will be helpful for future characterization of cell mechanics.

      Comments on revised version.

      In the original submission, cytochalasin B alone showed little effect on viscoelastic and active energy parameters, and it was unclear whether this reflected a true absence of actin's role or an artifact of the perturbation method used. In the revised manuscript, the authors addressed this by repeating the cytochalasin B measurements with larger sample sizes and adding latrunculin A, a mechanistically distinct and more potent actin-depolymerizing drug, together with immunostaining to confirm cytoskeletal disruption. This convincingly shows that actin depolymerization does affect the solid-like prefactor and fluidity, resolving the original concern.

      Nocodazole-induced microtubule depolymerization previously did not appear to reduce the solid-like property A, which was unexplained. The revised manuscript removes the speculative compensation-mechanism explanation, adds a discussion comparing the results to prior AFM literature (explaining the discrepancy as reflecting different mechanical compartments probed - cortex vs. intracellular), and the new data now show a significant reduction of A with nocodazole treatment as well. This weakness is resolved.

    4. Reviewer #3 (Public review):

      Summary:

      Cells and tissues are viscoelastic materials. However, metabolic processes that underly survival, growth and migration render the cell as an active matter at non-equilibrium. These two facts contribute to the difficulty of probing mechanical properties especially with sub-cellular resolution. However, the concept that the mechanical phenotype can be indicative of normal physiology necessitates approaches of defining the cellular phenotype. Here, Muenker et al evokes a powerful argument for mapping intracellular mechanics using optical tweezer- active microrheology. They present a suite of parameters towards a definition of a mechanical fingerprint. This is a compelling idea. There are some concerns as detailed below

      Strengths:

      These are technically challenging experiments and the authors provide systematic approaches to probe a system at non-equilibrium.

      Weaknesses:

      The importance of the mechanical fingerprint is diluted due to some missing controls needed for biological relevance. As it reads, sinusoidal waves are applied sequentially from 1- 1024Hz.<br /> Please clarify if amplitude is the same for each frequency, also how many frequencies are used?

      On this point, due to perturbations due to alterations in pre-stress, are the orders of frequencies randomized?

      How many beads are probed in a given cell.

      Is the graph in 1 c G', G" per cell or average of many cells?

      Figure 1e is quite nice, however is there an equivalent performed in a non-linear ECM such as collagen for comparison, in a similar vein can the equivalent be calculated for cells with/without treatment with low doses of cycloheximide to reduce protein synthesis? Yes, cytoskeletal elements are important for cell mechanics, but cytoplasm crowding is often an overlooked factor.

      The biggest issue is the interpretation of the different factors as each of these cells have different energetic needs.<br /> The comparison between cancer cells with different aggressiveness, immune and epithelial cells.<br /> For example, some types of cancer cells will be dominated by glycolysis vs oxphos, which will influence both the cytoplasmic and nuclear mechanics?

      It would be useful to carefully assess factors not restricted to<br /> a) Cytoskeleton<br /> b) Protein synthesis<br /> c) Metabolic state

      For similar lines and/ or cells where there are lineages that are either more metastatic in cancer, normal counterpart or drug resistant in an effort to link the fingerprint to a biological output. Specifically, is migration, proliferation, survival correlated with the measurements.

      The reviewer is sensitive to the technical difficulties of the experiments. However, the interpretation and importance of the mechanical fingerprinting requires additional work as mentioned above.

    5. Author response:

      The following is the authors’ response to the original reviews.

      General comments:

      You will see that many of the reviewers’ comments overlap. From our discussion with them, we agree that several of these comments should be addressed in this study, particularly comments related to the interpretation of the effect of drugs acting on the cellular cytoskeleton (reviewers #1 and #2). We also agree that the comparison of isogenic cell lines such as the mcf10a series or the 4T1 series should address some of the concerns regarding the interpretation of the mechanical fingerprint (reviewer #3). Also, certain methodological aspects should be easily clarified (reviewers #1 and #3).

      We also agreed that other comments may be more difficult to address in the context of this study. This is the case for comments related to establishing a link between different mechanical signatures and different cellular functions/outcomes (Reviewers #1 and #3). One could test whether migration or proliferation is altered by changing the mechanical fingerprint, or you could simply discuss these aspects by carefully reviewing the literature to corroborate mechanical signatures with known cellular phenotypes (e.g. migration speed, adhesion, cell size...). This is also the case for comments on the influence of other cellular parameters such as molecular crowding and energy metabolism, which could be left for future work or where you could use a low dose of cycloheximide (below the level of deleterious effects) to address the effect of cytoplasmic proteins (reviewer #3).

      We thank the editor for providing this helpful overview of the requested revisions. We have carefully addressed these points throughout the revised manuscript. The only difficulty was to establish the isogenic cell lines as requested. It took us over 18 months to find a source of these cells in Europe, and since then we are trying hard, but not successful to get these cells stably growing in the condition necessary for the optical tweezers experiments. As we have now spent more than 2 years on this without success, we decided to resubmit the paper without this part to not further delay this manuscript. The additional experiments and revisions have substantially strengthened the manuscript. Especially, the addition of Latrunculin A as suggested was an excellent request, as now the results regarding actin depolymerization and mechanical properties are in excellent agreement with the expected effects, as Latrunculin A is much more efficient in depolymerizing actin than cytochalasin B. The major changes are summarized below, followed by a detailed point-by-point response to all reviewer comments.

      General changes

      (1) Repeated measurements on wild-type HeLa cells.

      (2) Repeated all Cytochalasin B and Nocodazole experiments and increased the number of analyzed cells to approximately 60 per condition.

      (3) Performed additional experiments using Latrunculin A and combined Latrunculin A + Nocodazole treatment.

      (4) Performed immunostainings for all cytoskeletal perturbation conditions (WT, Cytochalasin B, Latrunculin A, Nocodazole, Cytochalasin B + Nocodazole, and Latrunculin A + Nocodazole).

      (5) Refined the rheological analysis procedure and expanded the methodological description.

      (6) Revised the manuscript text throughout and expanded the discussion of limitations and biological interpretation.

      Public Reviews:

      Reviewer #1 (Public Review):

      A limit of the paper is that the biological mechanisms by which intracellular mechanics is modulated (e.g. among cell types) remains unexplored and only briefly discussed. Yet this limit is greatly offset by the rigor of the approach.

      We thank the reviewer for this positive assessment and agree that a more extensive discussion of the biological mechanisms underlying the observed mechanical fingerprints strengthens the manuscript. We have substantially expanded the Discussion and Conclusion sections to address potential contributions of cytoskeletal organization, intracellular transport, molecular crowding, and metabolic state. In addition, we now discuss the relationship between the identified mechanical phase space and known cellular phenotypes where appropriate, while explicitly outlining the limitations of the current study and the need for future investigations linking intracellular mechanics to cellular function.

      Reviewer #2 (Public Review):

      The most difficult part of the method is the part with actin polymerization inhibition with cytochalasin B. The data shows that viscoelastic parameters as well as active energy parameters are unaffected by cytochalasin B. It is reasonable to expect that elasticity will reduce and fluidity will increase upon application of such a drug. The stiffness-reducing effect was observed only when CB was used with nocodazole most likely because of phagocytosis of the bead, which is governed by microtubule. The use of other actin-depolymerizing drugs such as latrunculin A would be needed to test actin’s role in mechanical fingerprints. If actin’s role is only explained by accompanying microtubule inhibition, it is not a convenient system to directly test the mechano-adaptation process.

      We thank the reviewer for this important suggestion. To strengthen the interpretation of the actin perturbation experiments, we repeated the Cytochalasin B measurements with an increased number of cells and performed additional experiments using Latrunculin A, a mechanistically distinct and more potent actin-depolymerizing compound. Together with complementary immunostaining experiments, these additional data reveal distinct contributions of the two major cytoskeletal systems to the intracellular mechanical fingerprint. Whereas actin depolymerization primarily affects intracellular stiffness and fluidity, microtubule depolymerization has the strongest effect on intracellular activity while also contributing to cellular softening. These additional experiments provide a substantially clearer interpretation of the respective roles of actin filaments and microtubules in shaping the intracellular mechanical fingerprint.

      Depolymerization of MT with nocodazole did not reduce the solid-like property A. Adding discussion and comparison with other papers in the literature using nocodazole will be helpful in understanding why.

      We thank the reviewer for this suggestion. We have expanded the discussion and now compare our observations with previous AFM studies investigating Nocodazole treatment. While AFM measurements of cortical mechanics often report little change or even increased stiffness after microtubule depolymerization, our intracellular measurements reveal pronounced softening and strongly reduced intracellular activity. We now discuss that this difference likely reflects the distinct intracellular mechanical compartment probed by intracellular microrheology compared with cortical AFM measurements.

      Overall, the usefulness of the concept of mechanical fingerprints and comparisons with other cell mechanics studies (from other groups) will make this manuscript stronger.

      We thank the reviewer for this suggestion. Throughout the revised manuscript we have strengthened the comparison of the mechanical fingerprint with previous literature. In particular, we now discuss the cytoskeletal perturbation experiments in the context of published AFM studies, compare the observed mechanical differences between cell types with previous measurements where available, and expand the discussion of the biological interpretation and limitations of the proposed mechanical fingerprint.

      Reviewer #3 (Public Review):

      The importance of the mechanical fingerprint is diluted due to some missing controls needed for biological relevance.

      We thank the reviewer for raising this important point. To strengthen the biological interpretation of the mechanical fingerprint, we performed substantial additional experiments, including repeated cytoskeletal perturbation measurements with increased sample sizes, additional Latrunculin A experiments, and complementary immunostaining analyses. We also expanded the discussion to address the influence of factors beyond the cytoskeleton, including molecular crowding and metabolic state, and explored possible relationships between the proposed mechanical phase space and cellular phenotypes. While we agree that future studies using well-controlled isogenic model systems will be required to establish direct links between intracellular mechanics and biological function, we believe that the additional experiments and expanded discussion substantially strengthen the biological relevance of the present study.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      A caveat of the general methodology, which is partially acknowledged in the MS is that beads are endocytosed and likely end up in specific lysosomal compartments. Therefore, it is not clear whether the mechanical fingerprint fully represent the material properties of bulk cytoplasm, and not something more specific to lysosomal organelles. For instance, lysosome motion may be largely driven by motors moving along MT cytoskeletal track, and the extracted effective energy may as such not fully represent the crowding and effective active temperature of the cytoplasm. This limit certainly affect the interpretation of the results in other cell types, in which membrane trafficking and cytoskeletal organization may vary largely. I believe it would be very important to outline this limitation of the work and discuss it in light of the results obtained throughout.

      We thank the reviewer for pointing out this important limitation, which was not sufficiently addressed in the original manuscript. We have now acknowledged this issue throughout the manuscript and added a limitation section to the conclusion to clarify that our findings specifically relate to internalized objects surrounded by a membrane and therefore primarily reflect the properties of membrane-bound organelles in the 1 µm size regime, rather than the bulk cytoplasm as a whole.

      We consider this focus on membrane-enclosed intracellular objects to be biologically relevant and interesting in its own right. Alternative approaches for introducing tracer particles, such as microinjection or particle guns, are generally more invasive and less reproducible. We therefore deliberately focused on phagocytosed beads as a minimally perturbative and robust experimental system in this study.

      The evolution of the mechanics in Hela Cells using cytoskeletal drugs in interesting, but I was confused by the fact that authors interpret the effect of cytochalasin solely on the cortex. As they are probing intracellular rheology, variations (or lack thereof) may rather reflect bulk F-actin networks? Also the compensation mechanism is interested, but it would need to be strengthened by immunostaining for instance, to support the claim, that microtubule depolymerization enhances F-actin networks.

      We thank the reviewer for this important comment. To elaborate on the effect of cytoskeletal filaments, we extended our analysis by repeating the experiments, increasing the number of samples, and investigating the effect of an additional drug, Latrunculin A. Additionally, we conducted immunostaining with subsequent confocal imaging to deepen our understanding of the effect of the respective drugs. The additional experiments reveal that actin and microtubules contribute differently to the fingerprint. Actin depolymerization primarily affects intracellular stiffness and fluidity, whereas microtubule depolymerization has the strongest effect on both mechanics and intracellular activity. Combined perturbation produces the largest overall effect. Based on these additional data, we no longer invoke the compensation mechanism proposed in the original manuscript. While interactions between the actin and microtubule cytoskeleton have been reported previously, our immunostaining experiments do not provide evidence for a compensatory increase in actin organization following microtubule depolymerization. We have therefore removed this interpretation from the revised manuscript and replaced it with a discussion based on the newly acquired perturbation and imaging data.

      The final figure using principal component analysis is very interesting, but it would be important to link this to phenotypic signatures of the different cells. Could the authors try to link resistance, fluidity and activity to the different functions/behavior of cells? For instance, some of these cells are migratory but some may move much faster than others, and it would be very interesting to correlate the degree of activity or fluidity with speed of migration, or cell shape/size/contractile state for example.

      Indeed, this is an important point. Establishing direct links between the mechanical fingerprint and functional cellular properties such as migration, contractility, proliferation, or morphology would substantially strengthen the biological interpretation of the identified phase space. We carefully considered this suggestion and explored several approaches to relate the measured mechanical parameters to cellular phenotype. However, obtaining directly comparable quantitative functional data across all investigated cell types proved challenging. Parameters such as migration speed, adhesion, and contractility depend strongly on experimental conditions, including substrate properties, assay design, and culture conditions, making literature values difficult to compare across studies. To address the reviewer’s concern, we expanded the discussion and incorporated comparisons to available literature where appropriate. For example, previous studies have reported higher migration rates for HeLa cells compared with MCF7 cells, which is qualitatively consistent with the higher intracellular activity observed in HeLa cells. However, due to the limited comparability and availability of quantitative functional data across the investigated cell types, we refrained from performing a formal correlation analysis. In addition, we grouped the investigated cell lines according to several broad phenotypic classifications, including epithelial/mesenchymal character, cancer status, metastatic potential, and migratory potential, and examined their distribution within the proposed phase space. While this exploratory analysis provides additional biological context, it did not reveal robust relationships that could support definitive conclusions regarding structure–function relationships. We therefore agree with the reviewer that establishing direct links between intracellular mechanical fingerprints and cellular function represents an important next step. To this end, future studies will combine intracellular rheological measurements with independently quantified functional assays, ideally in well-controlled isogenic model systems.

      Reviewer #2 (Recommendations For The Authors):

      The study needs more thorough validation against known technology (such as AFM) or literature, e.g., rheological change upon the same drugs used in the current study.

      We thank the reviewer for this suggestion. We have expanded the discussion of the cytoskeletal perturbation experiments and now compare our observations to previous AFM studies and related literature on cytoskeletal mechanics. Consistent with AFM measurements of cortical mechanics, actin depolymerization using Cytochalasin B or Latrunculin A resulted in a reduction of cellular stiffness. In contrast, microtubule depolymerization produced effects that differ from many AFM studies, which report either no change or an increase in cortical stiffness following Nocodazole treatment. We now explicitly discuss that this discrepancy likely reflects the different mechanical compartments probed by the two techniques. AFM predominantly measures the actin-rich cell cortex, whereas our intracellular microrheology measurements probe the mechanical environment experienced by membrane-bound intracellular particles. We therefore interpret the differing response to microtubule depolymerization as evidence that intracellular active mechanics and cortical mechanics can be influenced by distinct physical mechanisms. These comparisons have been incorporated into the Results and Discussion sections of the revised manuscript.

      Page 8: Citation to Fig. 3a is missing before mentioning Fig. 3b.

      We revised the manuscript to ensure that all references are given in an appropriate order.

      Proper uses of hyphens are recommended to avoid confusion. For example, ’a yet not understood change’ can be written as ’ a yet-not-understood change’.

      We thank the reviewer for this suggestion. We carefully revised the manuscript to improve the use of hyphenation and compound modifiers throughout the text. The specific example highlighted by the reviewer, as well as similar constructions, have been corrected to improve readability and avoid ambiguity.

      Reviewer #3 (Recommendations For The Authors):

      As it reads, sinusoidal waves are applied sequentially from 1- 1024Hz. Please clarify if amplitude is the same for each frequency, also how many frequencies are used? On this point, due to perturbations due to alterations in pre-stress, are the orders of frequencies randomized?

      We thank the reviewer for pointing out this ambiguity. We have revised the manuscript to provide a more detailed description of the active microrheology protocol. Specifically, we now state that all measurements were performed using a constant trapping-laser oscillation amplitude of 200 nm and that the applied frequencies were 1, 2, 4, 8, 16, 32, 64, 128, 256, 512, and 1024 Hz. The frequencies were applied sequentially in increasing order and were not randomized. This information has now been added to the manuscript.

      How many beads are probed in a given cell?

      We thank the reviewer for this question. We have clarified this point in the Methods section and now explicitly state that only a single phagocytosed probe particle was analyzed per cell. Of course, many different cells, and hence beads, have been analyzed per cell type.

      Is the graph in 1 c G’, G” per cell or average of many cells?

      We thank the reviewer for pointing out this ambiguity. In the original version of the manuscript, Figure 1b showed data from a representative cell, whereas Figure 1c displayed an average over multiple cells. To avoid confusion, we revised Figure 1 and now show representative data from a single measurement throughout the analysis workflow (Figure 1c,e,f).

      Figure 1e is quite nice, however, is there an equivalent performed in a nonlinear ECM such as collagen for comparison, in a similar vein can the equivalent be calculated for cells with/without treatment with low doses of cycloheximide to reduce protein synthesis? Yes, cytoskeletal elements are important for cell mechanics, but cytoplasm crowding is often an overlooked factor.

      We thank the reviewer for this important suggestion, and we are glad that the reviewer likes figure 1e. Regarding non-linear ECM, we have not done such experiments using optical tweezers. Collagen is a highly heterogeneous material and using the small deformations that we can obtain using the optical tweezers, our access to the non-linear contributions is rather limited.

      However, we agree that factors beyond the cytoskeleton, including molecular crowding and protein content, can make important contributions to intracellular mechanics. While investigating these effects experimentally, for example through cycloheximide treatment, would be highly interesting, such studies were beyond the scope of the present work.

      The primary focus of this study was to establish and validate a mechanical fingerprint for intracellular active microrheology and to investigate how this fingerprint responds to perturbations of the cytoskeleton. The additional experiments performed during revision therefore concentrated on strengthening the interpretation of the cytoskeletal contributions.

      At the same time, we agree that molecular crowding represents an important alternative mechanism influencing intracellular mechanics. We have therefore expanded the Discussion and Conclusion sections to explicitly acknowledge this limitation and now cite recent studies demonstrating strong effects of molecular crowding on intracellular rheology (Umeda et al,. 2023, Ebata et al., 2023). We further discuss that, besides cytoskeletal organization, metabolic state, intracellular transport, and molecular crowding are likely contributors to the observed mechanical fingerprint.

      The biggest issue is the interpretation of the different factors as each of these cells have different energetic needs. The comparison between cancer cells with different aggressiveness, immune and epithelial cells. For example, some types of cancer cells will be dominated by glycolysis vs oxphos, which will influence both the cytoplasmic and nuclear mechanics? It would be useful to carefully assess factors not restricted to

      (a) Cytoskeleton

      (b) Protein synthesis

      (c) Metabolic state

      For similar lines and/or cells where there are lineages that are either more metastatic in cancer, normal counterpart or drug resistant in an effort to link the fingerprint to a biological output. Specifically, is migration, proliferation, survival correlated with the measurements. The reviewer is sensitive to the technical difficulties of the experiments. However, the interpretation and importance of the mechanical fingerprinting requires additional work as mentioned above.

      We thank the reviewer for this thoughtful comment. We agree that intracellular mechanics is likely influenced by a broad range of biological factors beyond the cytoskeleton, including metabolic state, molecular crowding, intracellular transport processes, and protein synthesis. We also agree that the biological significance of the mechanical fingerprint would be strengthened by establishing direct links to functional cellular outputs such as migration, proliferation, or survival. To address the first point, we have expanded the Discussion and Conclusion sections of the manuscript to explicitly acknowledge that the observed fingerprint is unlikely to be determined solely by cytoskeletal organization. In particular, we now discuss the potential contributions of metabolic state, intracellular transport, and molecular crowding, and cite recent studies demonstrating the importance of these factors for intracellular mechanics. To address the second point, we explored several strategies to relate the measured mechanical fingerprints to cellular phenotype. We expanded the discussion of available literature, including examples where mechanical properties and migratory behavior appear qualitatively consistent. In addition, we grouped the investigated cell lines according to broad biological characteristics, including epithelial/mesenchymal character, cancer status, metastatic potential, and migratory potential, and examined their distribution within the proposed phase space. While this exploratory analysis provides additional biological context, it did not reveal robust relationships that would support definitive conclusions regarding structure–function relationships. We therefore agree that establishing direct links between intracellular mechanics and cellular function represents an important next step. Such studies will require quantitative functional assays performed under controlled and directly comparable conditions, ideally using well-defined isogenic model systems. We now discuss these limitations and future directions explicitly in the revised manuscript.

    1. eLife Assessment

      This valuable study reports a series of artificial-selection experiments for microbiomes associated with improved drought performance in rice. A major strength is the solid experimental design using multiple starting soil communities, which can guide others in designing related experiments. While interpretation of the results is constrained by inadvertent microbial dispersal between samples, the limited effectiveness of the sterile controls and the absence of healthy well-watered plants, the work nevertheless provides a helpful proof of concept for host-mediated microbiome selection, identifying candidate taxa, functions and simplified communities for downstream study. As a first step towards microbiome engineering in this area, it will be of particular interest to colleagues working in plant-microbiome interactions, microbial community selection and microbiome design.

    2. Reviewer #1 (Public Review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers.]

      Summary:

      The study claims to explore plant microbiome engineering using host-mediated selection as a strategy to enhance rice growth and drought tolerance.

      Strengths:

      The authors have derived and identified simplified microbiomes from wild microbial communities of rice fields, deserts, and serpentine seep soils by selecting microbiomes from plants with desired phenotypes across generations. Metagenome-assembled genomes revealed enriched functions, such as glycerol-3-phosphate and iron transport, known to mediate plant-microbe interactions during drought.

    3. Reviewer #2 (Public Review):

      Summary:

      In this study, Styer et al. impose artificial selection on root-associated microbiomes to increase drought tolerance in rice plants using different soils as starting microbiomes. Using NDVI and biomass as a proxy for plant health, they find that iterative passaging of the microbiomes of the best-performing plants increased plant resilience to drought stress in a soil-dependent manner. The study makes use of numerous controls. The authors survey the microbiota of the plants across generations, using an array of interesting analyses to characterize their observations. Firstly, the authors find that the acquired microbiomes are divergent towards the beginning of the selection experiment, but nearly converge later suggesting that the selected communities become more similar over time. One reason is that the diversity of the microbiomes severely decreases after only one or two generations of selection AND that microbes from each inoculation source appear to easily disperse across the experiment, leading to microbiome homogeneity. The authors then present an analysis to correlate ASVs with the NDVI and Biomass over the course of the experiment (using the rice soil selection lines) to develop hypotheses about which ASVs may impact plant traits.

      Strengths:

      The authors set out to refine the understanding of microbiome artificial selection, a topic of recent interest to the plant microbiome field. The authors use an established approach (Mueller et al), expanding upon it by including multiple starting soil inocula to ask whether the strength of selection varies by input microbiome. This is an important and novel question. Using drought resilience as measured by NDVI and plant biomass to select upon was a wise choice for this type of study, given their relative ease and quickness to assess. The inclusion of several types of controls, multiple selection lines, and several starting soil inocula showed a thoughtful experimental design. The analyses were diverse, non-standard, and attempted to address microbiome dynamics on multiple fronts. I am not necessarily convinced by some of the conclusions (see below), however, I think this study examines an important and exciting topic in the area of plant microbiomes. I predict the findings of the experiments will inform a wide audience of researchers attempting similar studies and be helpful in their designs.

    4. Reviewer #3 (Public Review):

      Summary:

      In this work, Styer et al. explore host selection as a means for recruiting microbes that may aid their host under stressful conditions, in this case under drought stress, as an alternative to target-SynCom design. They do so by subjecting rice plants to several generations of soil transplantation, and by using the most successful rice plants as donors for the next generation. By using several NGS approaches and very thorough bioinformatics analysis, the authors identify potential microbial taxa and the associated functions enriched in the conditions of interest.

      Strengths:

      In general, I think this approach was very much needed in the field as an alternative to SynComs, which are still not readily usable in croplands. This work sets the grounds for future similar approaches, using different stresses and different host plants.

      In this work, the experimental setup is well thought-through and well-replicated. In addition, an exhaustive set of preliminary experiments was performed before deciding on the final panel of soils to use and scoring methodology. The figures are clear and well-explained.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Reviewing Editor, the Senior Editor, and the three reviewers for their careful and constructive assessment of our manuscript. We were encouraged that the reviewers found the question timely and novel, the experimental design thoughtful and well-replicated, and the analyses diverse and informative. The reviewers also raised a number of valuable concerns, which clustered around three themes: (i) the framing of host-mediated selection as microbiome “engineering” versus a proof of concept; (ii) the interpretive challenges introduced by microbial dispersal and the resulting limits on the sterile-inoculated controls; and (iii) requests for clearer methodological detail and additional context from the recent literature. We have revised the manuscript to address these points through clearer framing, expanded discussion, and fuller methodological detail. Consistent with the nature of this long-term experiment, our revisions strengthen the interpretation and presentation of the existing dataset rather than adding new experiments.

      eLife Assessment

      The study has also shortcomings in that the rescuing effect is not benchmarked against healthy well-watered plants, the sterilized controls do not add much information, and the dispersal between inocula confounds the interpretation of the results… the presentation would overall benefit from more extensive consideration of recent developments in the field.

      We appreciate this balanced summary and have revised the manuscript accordingly. We have reframed the abstract and Introduction to present the study explicitly as a proof of concept rather than a completed engineering effort (ll. 27–31; ll. 96–101); we now address the well-watered benchmarking limitation and the limits of the sterile-inoculated controls directly in the Discussion (ll. 543–551); we discuss dispersal and its confounding effect on interpretation head-on, including an alternative hypothesis (ll. 546–551); and we have incorporated the recent studies suggested by Reviewer 3 (ll. 206–209, 442, 476–478). Each change is detailed in the point-by-point responses below.

      Reviewer #1 (Public Review):

      Weaknesses:

      The findings demonstrate the efficacy of host-mediated microbiome selection, but the engineering part for enhancing rice performance under drought-stress conditions has not been provided. The proposed mechanisms rely on correlations but not direct experimental proofs.

      We agree, and we have adopted this framing throughout. Our study demonstrates host-mediated selection as a discovery framework rather than a completed engineering pipeline, and we now say so explicitly: the abstract has been reframed (ll. 27–31) and a statement added at the end of the Introduction (ll. 96–101) clarifying that the work reproducibly enriches beneficial taxa and functions and yields simplified candidate communities, but does not yet benchmark those communities against single-isolate inoculants or test them in the field or against a resident native microbiome. We likewise agree that the functional inferences from our metagenome-assembled genomes (MAGs) are correlative; we now state this explicitly in the Methods and Discussion (ll. 776–778) and note that establishing causal roles for individual taxa or genes will require targeted isolation and genetic manipulation.

      Reviewer #1 (Recommendations For The Authors):

      The experimental design… could benefit from more detailed explanations. For instance, what are the criteria for choosing these soils and how are they relevant to rice growth phenotype? Also, the word ‘generation’ is misleading as it implies the use of seed-to-seed experiments… It would also be good to explain why the authors chose 6 generations for rice fields and 4 generations for deserts and serpentine seep. Importantly, the contribution of the rice seed microbiome… has not been considered and is also missing from… the discussion.

      We have addressed each part of this comment. Soil selection criteria: the Results section “Source inocula bacterial diversity” describes our rationale — we screened nine field soils in a pilot experiment, then selected the three that both supported rice growth and had negligible taxonomic overlap (providing three distinct starting points), with a stated per-soil expectation (rice-adapted, drought-adapted, and high-diversity). We are happy to expand this further if the reviewer feels additional detail is needed. “Generation”: we now define this term as a single 40-day selection cycle rather than a seed-to-seed generation (l. 109). Six vs. four generations: we explain in the Results (l. 309) that, having observed convergence of microbiome composition across soil treatments by the fourth selection generation, we concentrated resources on Rice Field and carried it through two additional cycles. Seed microbiome: we have added a note to the Discussion (ll. 444–446) that, although seeds were surface-sterilized before each generation, a residual contribution of seed-borne endophytes common to all treatments cannot be excluded.

      The authors stated that microbiomes were not selected for propagation into future generations in control lines. In this case, have the authors tested if the control LI microbiome in SG1 through SG6 did or did not significantly change in all the soil types?

      We have clarified the role of the live-inoculated (LI) lines in the text. LI lines were well-watered controls that were re-inoculated each generation with unsterilized selection-line material; they were included to identify drought-enriched taxa (by contrast with the droughted selection lines) and to test whether drought-optimized microbiomes were deleterious under well-watered conditions — not as an independently propagated selection line. Because LI communities were re-derived from selection-line inocula each generation, their composition necessarily tracked the changes occurring in the selection lines; this is the basis of the SL-versus-LI differential-abundance analysis (Figure 7B, Supplemental Figure 9). We note that comprehensive, temporally resolved 16S sequencing was performed for Rice Field, so we are appropriately cautious about extending LI comparisons across every soil type, and we have tempered our conclusions from the control lines accordingly (ll. 543–551).

      In Figure 3B (rice field), the tolerance in terms of AUC NDVI contrastingly increases to the biomass values in SG5 and SG6… the [NDVI] does not seem to be a good measure… It would be interesting to analyze these data sets under normal conditions… include representative pictures of all the ‘generations’… It will also be important to include the LI control data in Figures 3B and 3C.

      We appreciate these suggestions and respond to each. Our metric is biomass-adjusted AUC NDVI, which we use precisely to separate drought performance from plant size; NDVI itself was validated against shoot water content in preliminary experiments (R = 0.98; Supplemental Figure 2E), so we are confident it is an appropriate, validated proxy for drought status. We have substantially expanded the Methods to explain this adjustment and why the metric can diverge from raw biomass (ll. 666–669). Regarding the specific additions requested: analyzing the well-watered plants as a phenotypic dataset, adding representative images for every generation, and plotting LI data in Figure 3B/C would each require new analyses or figures that are outside the scope of this revision; moreover, LI plants were never droughted and therefore have no drought-response score comparable to the SL and SI lines, so they cannot be placed on the same axes. Representative images contrasting the first and last selection generations are already provided in Figure 3A. We have, however, added an explicit acknowledgement that our design does not quantify the absolute magnitude of drought rescue relative to well-watered performance (ll. 551–552).

      It is less clear how sterile soils acquired environmental taxa over time. Was this a seepage of microbes from inoculated samples to the calcinated clay, possibly via the water irrigation system? In this regard, four Venn diagrams representing all the generations… would be relevant.

      Each plant was grown in an individual container with its own separate water reservoir (Supplemental Figure 4), so shared irrigation was not a route of transfer; the most likely routes are airborne movement and handling within the growth chamber, together with within-treatment shuffling of plants. Our dispersal analysis (Figure 5) already traces the origins of taxa in each treatment, and Supplemental Figure 6 quantifies the ASVs shared among treatments over generations; we have added explicit criteria for these origin assignments (ll. 248–253). We therefore prefer to retain the existing Figure 5 / Supplemental Figure 6 presentation rather than add four separate Venn diagrams, which would convey the same information less quantitatively, but we are glad to reconsider if the editor feels a Venn representation would help readers.

      What is the logic behind the so-called ‘immigrating taxa’ in this study?

      “Immigrating” (dispersed) taxa are those that appear in a treatment despite not being attributable to that treatment’s own starting material — i.e., ASVs not detected in that treatment’s field soil or enrichment-generation inoculum, which must therefore have arrived by dispersal from other treatments or from the growth-chamber environment. We have made this definition explicit in the text (ll. 248–253).

      The decrease in alpha-diversity in subsequent generations… should be thoroughly discussed. Have authors tried to culture these few remaining taxa? If yes… tested for their individual drought tolerance supported by physiological assays… If no, is the microbiome of SG6 (and associated functions) ideal or sufficient to create drought tolerance in field conditions?

      We have expanded the discussion of the diversity decline. In addition to niche filtering along the soil-to-root gradient and dilution-to-extinction (already discussed), we now note that DNA-based profiling cannot distinguish metabolically active cells from relic DNA or dormant/non-viable cells, so part of the apparent collapse in diversity may reflect enrichment for the taxa that were active in the original inoculum (ll. 206–209). We agree that culturing the remaining taxa and characterizing them with physiological assays (e.g., water potential, water-use efficiency, stomatal conductance) is a valuable next step; these experiments are outside the scope of the present study, which we have now framed explicitly as a proof of concept, and we identify field validation of selected communities as a key open question (ll. 96–101).

      The result that Ideonella was identified as the dominant taxa in all selection conditions is highly interesting… This… should have been followed up for isolating the strains and performing direct tests to test their importance for conveying drought stress.

      We agree that isolating and directly testing dominant taxa such as Ideonella is the logical next step, and we now emphasize that a central value of host-mediated selection is that it yields simplified communities from which such taxa can be more readily isolated (ll. 27–31). These isolation and functional-validation experiments are beyond the scope of the current study and we have framed them as future directions rather than undertaking them here.

      The MAGs shown in Figure 8 have apparently ‘been assigned to ASVs…’. These data are not shown anywhere… the MAG data only give correlations but not direct genetic proofs of the biological functions of the identified genes.

      We have expanded the Methods to describe how each MAG was matched to an ASV (by closest taxonomic assignment and by concordance of relative abundance across samples), and we now state explicitly that these assignments are approximate and that the functional inferences drawn from them are correlative rather than definitive (ll. 776–778). We would be glad to add a supplemental table listing the MAG-to-ASV assignments if the reviewer or editor would find it useful; because it reports assignments already in hand, it requires no new analysis.

      Reviewer #2 (Public Review):

      Strengths:

      I think this study examines an important and exciting topic in the area of plant microbiomes. I predict the findings of the experiments will inform a wide audience of researchers attempting similar studies and be helpful in their designs.

      We thank the reviewer for recognizing the novelty of this complex experiment as well as the effort we put into designing it. Like the reviewer, we hope that this manuscript can serve a wide audience and help inform subsequent experiments in this new topic area.

      Weaknesses:

      Although the controls were well designed, the dispersal of the microbiomes erased the utility of the sterile inoculated (SI) controls… the SI lines acquired microbes from the experiment and never appeared to significantly deviate from the SL plants. The dispersal of the microbes… also minimizes any conclusions that can be made about the different starting inocula and how prone to selection they may be.

      We agree that microbial dispersal confounded our ability to use the sterile-inoculated (SI) plants to account for batch variation between generations. By maintaining each plant as a spatially discrete unit (individual pots and watering reservoirs), we had originally intended SI plants simply to acquire a similar consortium of environmental microbiota each generation. Truly axenic SI plants would have been better suited to this purpose, but would have severely limited the number of replicates and replicate selection lines we could include. We have now addressed this limitation directly in the Discussion (ll. 543–551): we state that the SI lines cannot be treated as static, microbe-free baselines, that the batch-to-batch variation they were meant to capture is only partially controlled, and that dispersal limits the strength of the conclusions we can draw about differences between starting inocula. This shared trajectory of selection and intended control lines has been observed in other host-mediated selection studies but rarely discussed in detail, and we now foreground it as a lesson for experimental design.

      Reviewer #2 (Recommendations For The Authors):

      My first concern is the framing of the approach… the authors never show that this approach has better efficacy than single-isolate inoculates… The phase of the research is still proof of concept, understandably, but these caveats should be mentioned/addressed head-on in the Introduction and Discussion.

      We agree and have made these caveats explicit rather than implicit. The abstract now frames the work as identifying candidate taxa and communities rather than delivering a finished engineering solution (ll. 27–31); the Introduction now states plainly that this is a proof of concept that does not benchmark the passaged communities against single-isolate inoculants or evaluate them in the field or against a resident native microbiome (ll. 96–101); and the Discussion reiterates these limitations (ll. 543–551).

      I disagree with the authors that the selected microbiota better approximate field conditions (line 55) - because… the diversity of the microbiome is drastically reduced… it is likely that exclusion of taxa is just as important as the passaging of bacterial members to see the desired effect.

      We take this point and have revised the sentence at (former) line 55 accordingly (now l. 57): we now say that community-level screening more closely approximates field complexity than single-isolate screens only at the outset, and we no longer imply that the selected (diversity-reduced) communities better approximate the field. We agree that taxon exclusion may be as important as enrichment; this is consistent with our balance analyses, in which the denominator groups comprise taxa negatively associated with phenotype (Figure 7), and with the diminishing returns we observe as diversity collapses. We have also added an explicit sentence to the Discussion (l. 488) stating that the exclusion of detrimental taxa may be as important as the enrichment of beneficial ones, and that a microbiome’s finite membership may contribute to the diminishing returns of selection we observe over generations.

      (1) It is unclear what the reason (or methodology) for correcting NDVI by biomass. Much of the findings hinge on corrected NDVI values, so a more thorough explanation of the correction method… would benefit the reader.

      We have substantially expanded this explanation in the Methods (ll. 666–669). We now state that biomass and AUC NDVI were anti-correlated (Supplemental Figure 12) and that we adjusted for plant size by taking the residuals of a linear regression of AUC NDVI on shoot dry-weight biomass, using these biomass-adjusted values as our measure of drought performance so that selection would reflect drought tolerance rather than plant size alone.

      (2) Are data for panels B and C of Figure 3 scaled?… how can one have a negative area under the curve if all the NDVI values are positive? For panel B, the representative plant images are much larger than 0.8 grams.

      This is a helpful catch, and the confusion stems from our terse original description. The values plotted are the biomass-adjusted AUC NDVI (regression residuals), which are centered on zero by construction; negative values therefore indicate poorer-than-expected drought performance for a plant of a given size and do not reflect negative raw NDVI or a negative raw area under the curve. We now explain this explicitly (ll. 666–669). In panel C, shoot biomass is plotted as dry weight in grams; the representative plant images in panel A are qualitative illustrations and are not scaled to the biomass axis. We will make the axis labels and legend state the units and the residual nature of the adjusted metric explicitly (noted in our accompanying figure-revision guide).

      (3) The dispersal analysis… What are the criteria for classifying ASVs as specific to an input source? Was it that they were observed in all samples of field soil, i.e. was a prevalence threshold implemented? Could they be observed in any other soil at a smaller threshold?

      We have added the criteria explicitly (ll. 248–253). An ASV was attributed to a given soil treatment if it was detected (present/absent) in that treatment’s field-soil or enrichment-generation inoculum samples; ASVs detected in none of the field soils or source inocula were designated environmental in origin (“unk/env”), and ASVs meeting the criterion for more than one treatment were assigned to each. Assignments were thus based on detection in the source samples rather than on an abundance-prevalence threshold within later generations.

      This reviewer finds the results around [inoculum source] inconclusive… the serpentine seep microbiome appears to provide more benefit from the first round of selection than any other soil… The slope of improvement… is different between soils, but mainly because the serpentine microbiomes start out conveying greater benefits than the other soils.

      We agree the Serpentine Seep result is not clear-cut. The Discussion already presents inoculum provenance as one of several factors shaping the outcome rather than a decisive one, and we have now added an explicit acknowledgement that Serpentine Seep conferred comparatively large benefits in the earliest cycles before plateauing, so its weaker response to continued selection may reflect an early approach to a performance ceiling rather than an inherently poorer substrate for selection (l. 423). We have tempered our “source matters” language accordingly.

      Have the authors assessed the biomass and ndvi of the well-watered plants?… showing this data would allow the reader to assess the degree to which the microbiomes are rescuing the plant… and… whether tradeoffs exist… under fully watered conditions.

      We have added an explicit statement that our design does not pair each droughted line with a well-watered readout of the same phenotype, so we refrain from estimating the absolute magnitude of drought rescue (ll. 551–552). We note, however, that shoot biomass increased in parallel with drought performance across selection generations (Figure 3C), which provides no evidence that selection for drought tolerance came at a cost to growth under our conditions. Collecting matched well-watered phenotypes to quantify effect size and trade-offs is a worthwhile aim for future work but would constitute a new analysis beyond this revision.

      How can the authors exclude the possibility that environmental microbes pre-existing in the growth chamber taxonomically overlap with the field soil-specific microbes?… the alternative hypothesis should be mentioned… A clearer representation of the ASVs categorized as source soil-specific in Figure 5… would be useful and how many of these ASVs make up the bar plots.

      We now state this alternative hypothesis explicitly: because dispersed taxa came to dominate all treatments, we cannot fully exclude that taxa shared across treatments were recruited from a common growth-chamber pool rather than dispersing directly between soils (ll. 546–549). We note that the two processes are difficult to distinguish retrospectively, but that the bias of each SI line toward its own treatment’s native diversity (Figure 5) is more consistent with genuine cross-treatment dispersal. Regarding the figure, the number of ASVs underlying each origin category is available in Supplemental Figure 6; we describe in the accompanying figure-revision guide how the Figure 5 legend can be clarified to state the assignment criteria and the ASV counts.

      The sterile inoculated plants were a nice control in theory, but I question their utility… A contrast that should be made is the microbiomes of only SI plants. It is striking that sterilized controls assemble and retain more microbes from the unsterilized starting inoculum. I would expect everything to be acquired from dispersal.

      We agree, and we have foregrounded this in the Discussion (ll. 543–551). As the reviewer notes, SI communities were biased toward their own treatment’s native diversity rather than being assembled entirely from dispersal (Figure 5) — an informative observation, but one that also demonstrates why the SI lines cannot serve as the clean, microbe-free baseline we had intended. We now treat this as a key design lesson and note that a fully isolated (e.g., gnotobiotic) control would be required to separate these effects in future experiments (l. 560).

      Reviewer #3 (Public Review):

      Weaknesses:

      Sterile/non-inoculated calcined clay also tends to enrich similar microbes… In a future experiment, the work would benefit from including a truly sterile control… the reader may get to wonder whether these efforts are necessary at all… This is discussed across the paper but not directly addressed and I think the manuscript would benefit from a clear argument for or against this idea.

      We thank the reviewer for this insightful point and have made our argument explicit rather than leaving it implicit. First, we agree a fully isolated, truly sterile control would strengthen future iterations of this design; the manuscript notes that gnotobiotic plants would be the ideal (if costly) means of achieving this (l. 560). Second, on whether selection is necessary if plants recruit beneficial microbes from the environment: the phenotypic gains seen even in the sterile-inoculated lines do not indicate that selection was superfluous, but rather that those plants recruited from a metacommunity that was itself being optimized by selection in the neighboring selection lines each generation. In other words, environmental acquisition propagated the benefits of selection across the shared growth-chamber environment rather than replacing it. We have clarified this reasoning in the Discussion (ll. 543–551).

      Reviewer #3 (Recommendations For The Authors):

      It is mentioned multiple times… that host genotype is the driver of the microbiota selection… However, this is not the case [multiple lines] and therefore I don’t find that surprising that there is a convergence of the microbiota across soils and selection rounds.

      We agree and have added text making this explicit: all plants were a single, near-isogenic rice genotype, and because host genotype is itself a strong filter on microbiome composition, the use of one genotype — together with shared environmental conditions and selection criteria — makes convergence across lines an expected rather than a surprising outcome (ll. 438–441). We have softened language that could be read as attributing selection to host-genotype variation.

      Another possibility… is that those microbes that are found in the later generations are actually the ones that were active/alive in the initial inoculum. It is not possible to rule out that most of the sequenced microbes in the input were not actually dead. Similar observations were made… in Duran et al. 2022. New Phytol.

      We have added this possibility to the manuscript, noting that DNA-based profiling cannot distinguish metabolically active cells from relic DNA or dormant/non-viable cells, so part of the apparent diversity decline may reflect enrichment for the subset of taxa that were active in the original inoculum, with reference to the transplantation work the reviewer cites (Durán et al. 2022; ll. 206–209).

      In the shotgun data, was there any observation of other microbes present (fungi, virus)? Did they follow the same trends as the bacterial communities?… I think addressing this will be very interesting and very novel.

      We agree this is an interesting question. Our shotgun workflow was designed and assembled specifically to recover high-quality bacterial and archaeal MAGs, and a rigorous cross-kingdom analysis (fungi, viruses) would require dedicated, eukaryote- and virus-specific assembly, binning, and reference databases — a substantial new analysis that lies outside the scope of this revision. We therefore flag cross-kingdom community dynamics as a promising direction for future work rather than presenting a new analysis here.

      Any interesting overlap with the results found in Karasov et al. 2022 (biorxiv)?

      We have added a comparison to drought-driven selection on host-associated microbiomes in Arabidopsis (Karasov et al. 2022) at the relevant point in the Discussion (l. 442).

      In Liu et al., 2024 Nat. Comms, the authors found Devosia as an interesting candidate for disease suppression (to add to the discussion?).

      Added — we now note that Devosia, one of the lesser-known genera enriched in our experiment, has recently been highlighted as a candidate mediator of disease suppression in the rhizosphere (Liu et al. 2024; l. 476).

      Lipids as a signal for host-microbe interaction: Rich et al., 2021 Science.

      Added — in the functional-enrichment discussion we now cite lipids as increasingly recognized central signaling molecules in host–microbe symbioses (Rich et al. 2021; l. 478).

    1. eLife Assessment

      This valuable study combines experiments and theory to investigate the role of spontaneous correlated activity in establishing aligned topographic maps of neural activity in higher-order sensory areas and will be of interest to researchers studying multisensory integration and brain development. The revised work presents solid evidence that spontaneous activity is correlated and spatially organized across the relevant cortical areas and that, in a computational model, an intermediate level of such correlation can refine a coarse initial connectivity scaffold into aligned maps containing neurons responsive to one or both sensory modalities.

    2. Reviewer #1 (Public review):

      Dwulet et al. combined experimental and modeling approaches to investigate how correlated spontaneous activity in the mouse's primary visual (V1) and primary somatosensory (S1) areas drives the development of multisensory integration in area RL. Notably, they focused on early developmental stages, before sensory experience occurs. Consistent with previous experimental findings, the authors first demonstrated that spontaneous activity becomes more sparse across development in all three areas, as measured by event amplitude, event duration, and participation ratio. Using a linear mixed model analysis to compare the maturation of this spontaneous activity, they found evidence that S1 matured the fastest. The authors then presented experimental evidence suggesting that these spontaneous events were moderately correlated both spatially and temporally.

      They hypothesized that activity-dependent mechanisms use these correlations to establish connectivity across these regions. To test this hypothesis, the authors modeled a feedforward network with connections from S1 to RL and from V1 to RL, where the strength of connections depended on a Hebbian term for potentiation and a heterosynaptic term for depression. By investigating different levels of V1-S1 correlations, they found that moderate levels of correlation led to the significant development of topographically organized connectivity while maintaining a mix of bimodal and unimodal cells in RL. Additionally, when simulating a network with a more mature S1, they observed that topographical maps improved not only between S1 and RL but also between V1 and RL. Finally, the authors use linear regression to suggest that the mixture of bimodal and unimodal cells in RL is optimal for encoding the maximum amount of information from both V1 and S1.

      Comments on revised version:

      The revision closes most of the data-model gaps raised in my original review. The authors have clarified the experimental measures and statistical comparisons, improved the spatial correlation-map analysis, added a temporal-lag analysis that argues against stereotyped traveling waves, expanded the model description, and performed additional simulations examining the effects of differences in spontaneous activity. Taken together, these changes provide solid support for the paper's principal conclusion: structured and moderately correlated activity can, within the proposed model, guide the refinement of an initially coarse connectivity scaffold into aligned multisensory representations.

      The remaining limitations primarily concern the more specific claim that the somatosensory pathway matures first and guides refinement of the visual pathway. The experiments support the conclusion that spontaneous activity in the somatosensory cortex matures earlier. However, the proposed consequence of this difference is carried in the model by an assumed stronger initial somatosensory-to-higher-order connectivity bias, motivated by pilot anatomical observations that are not included quantitatively in the manuscript. The new supplementary simulations suggest that differences in activity amplitude and frequency alone are insufficient, but the parameters are varied over ranges considerably smaller than the differences measured experimentally, and event duration is not varied. These simulations therefore do not strongly establish that the measured activity differences are insufficient to produce the effect. In addition, the Figure 4 caption and portions of the Discussion continue to imply that more mature somatosensory activity itself instructs map alignment, whereas the revised Results present the more qualified conclusion that an additional connectivity difference is required.

      A few internal inconsistencies also remain. The abstract still describes activity in the three areas as being recorded simultaneously, although the cellular-resolution recordings were acquired sequentially; only the wide-field data were collected simultaneously across areas. The revised model also assigns spontaneous events durations and intervals in milliseconds, while the measured calcium events last several seconds and occur only a few times per minute.

    3. Reviewer #2 (Public Review):

      The revised manuscript has substantially improved, and the authors have satisfactorily addressed most of the concerns raised in my original review. Overall, the experimental evidence and its relationship to the computational model are now presented more clearly and rigorously, substantially strengthening the manuscript.

      My main reservation in the original review concerned the role of the initial topographic connectivity bias in the computational model. The revised manuscript provides a clearer interpretation of this aspect. The initial bias represents a coarse activity-independent scaffold, while the final organization of the maps depends on its interaction with the structure and degree of correlated spontaneous activity. Importantly, the simulations show that the presence of the initial bias alone does not determine the final connectivity pattern. I therefore consider the computational results substantially better supported and interpreted in the revised manuscript. Nevertheless, as the model starts from a predefined coarse topographic organization of the projections from the primary sensory cortices to RL, in my opinion, the results demonstrate how structured spontaneous activity can refine and align an initially organized connectivity scaffold, rather than showing that spontaneous activity itself establishes this topographic organization.

      This distinction is relevant when interpreting the central mechanistic conclusion of the study. The work provides convincing support for the idea that correlated spontaneous activity can contribute to the refinement and alignment of multisensory cortical maps, conditional on the existence of an initial coarse topographic organization. Establishing experimentally how this initial connectivity is organized during the relevant developmental period, and how spontaneous activity modifies it, remains an important question for future work.

      Overall, I consider the revised manuscript considerably stronger than the original submission. Most of my previous concerns have been adequately resolved, and the study provides valuable experimental and computational insight into how spontaneous activity may contribute to the development of aligned multisensory representations.

    4. Reviewer #3 (Public review):

      Summary:

      The study by Dwulet et al. explores how the development of spontaneous neural activity in primary sensory cortices influences the co-alignment of multiple sensory modalities in higher-order brain areas (HOAs). To address this question, they focus on connectivity between the primary visual (V1) and somatosensory (S1) cortices and an associative cortical area (RL) in mice. The authors combine experimental (wide-field and two-photon calcium imaging) and computational approaches to show that spontaneous activity matures at a different pace across these brain regions. Their data indicate that S1 develops more rapidly than V1, which is possibly beneficial for RL's integration of visual and somatosensory inputs through correlated spontaneous activity. Using a computational model, they demonstrate that a moderate correlation between V1 and S1 activity can optimally guide the formation of bimodal neurons in RL, which are crucial for maximizing the decodability of multisensory stimuli. This finding highlights the role of correlated spontaneous activity in primary sensory cortices in establishing co-aligned topographic multimodal sensory representations in downstream circuits.

      Strengths:

      The manuscript is well written and it provides strong enough evidence to support the main claim of the authors. The insights on the role of correlated activity on instructing co-aligned multisensory maps in HOAs are not trivial and are an important advancement for the field.

      Weaknesses:

      In the opinion of this reviewer, the study has no major weaknesses. A drawback of the work is that none of the predictions of the computational modeling have been corroborated through mechanistic experimental manipulations of early brain activity.

      Comments on revised version:

      The authors have addressed all my previous concerns. I have no further comments.

    5. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Dwulet et al. combined experimental and modeling approaches to investigate how correlated spontaneous activity in the mouse's primary visual (V1) and primary somatosensory (S1) areas drives the development of multisensory integration in area RL. Notably, they focused on early developmental stages, before sensory experience occurs. Consistent with previous experimental findings, the authors first demonstrated that spontaneous activity becomes more sparse across development in all three areas, as measured by event amplitude, event duration, and participation ratio. Using a linear mixed model analysis to compare the maturation of this spontaneous activity, they found evidence that S1 matured the fastest. The authors then presented experimental evidence suggesting that these spontaneous events were moderately correlated both spatially and temporally.

      They hypothesized that activity-dependent mechanisms use these correlations to establish connectivity across these regions. To test this hypothesis, the authors modeled a feedforward network with connections from S1 to RL and from V1 to RL, where the strength of connections depended on a Hebbian term for potentiation and a heterosynaptic term for depression. By investigating different levels of V1-S1 correlations, they found that moderate levels of correlation led to the significant development of topographically organized connectivity while maintaining a mix of bimodal and unimodal cells in RL. Additionally, when simulating a network with a more mature S1, they observed that topographical maps improved not only between S1 and RL but also between V1 and RL. Finally, the authors use linear regression to suggest that the mixture of bimodal and unimodal cells in RL is optimal for encoding the maximum amount of information from both V1 and S1.

      However, there are significant gaps between the experimental data and the modeling setup, which weaken the paper's conclusions. Additionally, some key details are omitted, making it difficult to fully assess their analysis and interpret some of their figures.

      (1) Some of the statistical measures and techniques in Figure 1 could benefit from clearer definitions. While the thresholds for activation (peak with at least 5% dF/F0) and events (20% of recorded cells activated simultaneously) are provided, event duration and participation rate are not clearly defined. Based on this definition of event alone, it is unclear why the minimum participation rate in Figure 1F is not 20%. Additionally, the conclusion that S1 matures earlier than RL and V1 could be strengthened by including a direct comparison between S1 and RL, as the current analysis only compares these areas to V1.

      We thank the reviewer for this comment. We have now updated the Methods to include the event duration as time above half max, participation rate as % of cells out of total in that region active during an event. Also, the threshold of 20% recorded cells to identify an event was incorrectly stated, in fact the threshold was 5% consistent with what the reviewer observed in Figure 1F. This error has been corrected throughout the Methods. We chose 5% because spontaneous activity significantly sparsifies over development, with events involving far fewer cells, as previously shown by multiple studies (Golshani et al., 2009; Rochefort et al., 2009; Gribizis et al., 2019; Leighton et al., 2021; Murakami et al., 2022; Chini et al., 2022; reviewed in Lakhera et al., 2024).

      For the linear mixed model (LMM) analysis, we used V1 as a reference just for convenience, but this has no influence on the results. We now added a direct comparison using each area as reference in the LMMs. Several Supplementary Tables (S1-3) now show these results with coefficient estimates and stars showing statistical significance and are mentioned in the legend of Figure 1 and the main text.

      (2) The wide-field experiments in Figure 2 could be expanded to support the feedforward modeling assumptions. Currently, the spatial and temporal correlations presented leave open the possibility that these spontaneous events are traveling waves propagating from V1 to RL to S1 (or vice versa). This scenario would suggest a different connectivity scheme for the model. Clarifying this point with additional data analysis, specifically including temporal correlations involving RL, could provide stronger support for the model's assumptions.

      We agree with the reviewer that the correlation analyses shown in Figure 2 do not differentiate between two possibilities: activity that travels smoothly from one cortical area to another, thereby correlating correlations between these areas, versus activity that is spatially confined to individual areas but occurs near-synchronously across those areas. To address this point, we have revised Figure 2 in two ways.

      First, we added examples of spontaneous activity showing near-synchronous but spatially distinct activation of sub-areas in V1, RL and S1 (new Figure 2D). These examples show that localized activity can remain confined to individual sensory cortical areas and RL, while occurring at similar times across areas. Thus, the observed correlations are not simply due to single large events spreading continuously across the entire imaged field.

      Second, we added a lagged cross-correlation analysis between V1 and S1 activity (new Figure 2G). This analysis shows that the correlation between V1 and S1 peaks close to zero lag and decays for both positive and negative lags. This argues against a stereotyped travelling-wave-like propagation from V1 to S1 or from S1 to V1 with a fixed delay. The cross-correlation curves show a mild asymmetry, with somewhat higher correlations when S1 precedes V1. However, because the dominant peak is centered near zero lag, we interpret the data primarily as evidence for near-synchronous, spatially structured coactivity across sensory areas, rather than fixed directional propagation.

      Together, these two analyses support the modeling abstraction that V1 and S1 provide temporally correlated, spatially structured inputs to RL. We have added the new activity examples and the lagged cross-correlation analysis to Figure 2 and revised the Results accordingly. Although these analyses do not exclude all forms of propagating activity, they argue against the specific concern that the correlations are dominated by stereotyped traveling waves passing sequentially through V1, RL, and S1.

      (3) The functional correlation map in Figure 2D appears contradictory to the authors' modeling assumption that inputs are correlated spatially in V1 and S1. While V1 seed points align topographically with RL, this organization breaks down when extended into S1. In contrast, and in support of the modeling assumption, Figure 2E shows clearer topography across all three regions. A discussion of this discrepancy would be helpful, as it's a key conclusion of the figure. Additionally, it is unclear when this data was collected during development. Clarifying the developmental stage and analyzing how this map changes over time could strengthen the results.

      We thank the reviewer for pointing out this ambiguity. In the original version, the functional correlation maps were generated using separate seed locations in V1 and S1, and the interpretation relied heavily on thresholded RGB maps in which each pixel was assigned to the color channel with the strongest correlation. This representation made it difficult to directly compare the V1- and S1-seeded maps and may have given the impression that topographic organization was preserved in one direction but not the other.

      We have therefore revised the analysis and presentation of Figure 2. Instead of using separate seeds in V1 and S1, we now use common seed locations in RL and compute the correlations of these RL seeds with activity across the imaged cortical field. This allows us to ask directly whether different RL locations are associated with spatially distinct regions in both V1 and S1. We now show both the raw correlation maps, in which the RGB channels reflect the correlation values for the three RL seeds (new Figure 2E), and the thresholded/maximum-channel representation, in which each pixel is assigned to the strongest of the three color channels (new Figure 2F). The raw correlation maps make the correlation structure visible without relying solely on thresholding, whereas the thresholded representation highlights the spatial ordering of the strongest correlations.

      With this revised analysis, the topographic relationship across V1, RL, and S1 is clearer and no longer depends on comparing separate V1- and S1-seeded maps. We also clarified in the figure legend how the RGB maps are computed and how thresholded pixels are represented.

      The reviewer also asked about the developmental stage and progression of this phenomenon. The example shown in Figure 2 was recorded at PN9, and we now state this explicitly. In addition, we added examples from PN9–PN13 in Supplementary Figure S1, showing that similar functional correlation-map structure is present across the developmental period analyzed here. This is consistent with previous work showing that retinotopy-like patterns in higher visual areas can be recovered from functional-connectivity analysis of spontaneous activity before eye opening (Murakami et al., 2022), and with recent work showing that retinotopy-like and somatotopy-like patterns of ongoing activity, together with their rough topographic correspondence in RL, are already present before eye opening at PN10–11 (Matsumoto, Murakami & Ohki, 2025).

      (4) The modeling of spontaneous events with fixed amplitude and duration seems inconsistent with the experimental data in Figure 1, which shows variability in these parameters. This is particularly confusing in Figure 4, where S1 maturation is modeled as a stronger topographical alignment with RL, but the experimental data defines maturation based on amplitude, duration, and event rates. Justifying these modeling choices or adapting the model to reflect experimental variability would create a better connection between the theory and data.

      We agree with the reviewer that the original presentation did not sufficiently distinguish between the experimentally measured maturation of spontaneous activity and the way S1 maturation was implemented in the model. In the experiments (Figure 1), earlier maturation of S1 was reflected by lower event amplitudes, shorter durations, and higher event rates. In contrast, the original model explored the effect of a stronger or more spatially refined S1-to-RL projection (Figure 4). This modeling choice was motivated by pilot anatomical data suggesting that projections from S1 to RL become more elaborate earlier than projections from V1 to RL at comparable developmental ages. We include examples of these pilot data (Author response image 1), but we have not included them in the manuscript because the dataset is preliminary and does not yet allow for a sufficiently complete quantitative analysis.

      Author response image 1.

      Projections from V1 and S1 to RL at different developmental ages. Pilot anatomical data suggest that the S1 projection to RL becomes more elaborate and mature earlier than the V1 projection.

      To address the reviewer’s concern more directly, we have now extended the model to incorporate differences in the spontaneous activity patterns of V1 and S1, including the lower amplitude and higher frequency of S1 events. We then examined how these activity differences interact with different levels of initial connectivity bias between the primary sensory cortices and RL (Supplementary Figure S2). We also quantified the resulting topography, map alignment, and fraction of bimodal RL neurons as a function of the S1 bias and included these additional plots in Figure 4 (panels C-E).

      This analysis shows that incorporating the more mature S1-like activity patterns alone was not sufficient to generate the appropriate topographic and aligned maps. Rather, the model still required an initial connectivity bias, together with an appropriate level and structure of correlated activity. This is consistent with the results shown in Figure 3B,E,G–I and discussed in our response to Reviewer 2, point 3, where we show that the initial bias does not by itself determine the final map structure, but instead interacts with the level of V1–S1 correlation. We have added the new analysis to Supplementary Figure S2 and revised the text to clarify the interpretation. Rather than presenting the stronger S1 bias as a direct consequence of the more mature S1 activity dynamics revealed through the differences in amplitude, duration, and event rate, we now frame it as a model prediction: earlier S1 maturation may need to be accompanied by, or act through, a more advanced anatomical or functional S1-to-RL projection, whose refinement still depends on the temporal and spatial structure of spontaneous activity.

      The results suggest that differences in spontaneous activity dynamics and differences in projection maturity may act together during the emergence of topographically aligned multisensory maps, with neither component alone being sufficient to determine the final organization. Future experiments will be needed to establish whether such an S1-to-RL connectivity bias is present systematically, to quantify its developmental progression, and to disentangle the relative contributions of more mature spontaneous activity dynamics and more mature connectivity.

      (5) Several important details of the mathematical model are missing or unclear, partly due to typos. The Results section mentions the general framework of the input correlation matrix (e.g., "S1 and V1 neurons were driven by a combination of events, independent and shared in each V1 and S1" and "each independent event activated a randomly chosen, contiguous set of neurons"), but the specifics are not fully explained. Additionally, the caption of Figure 5 refers to a non-linear transfer function (a sigmoid), but these details are not provided in the Methods section, which instead suggests a linear model was used. A careful review of the main text and Methods section would help ensure that all the necessary details are included and that the story is both complete and accurate.

      We thank the reviewer for pointing out these missing details and inconsistencies. We have carefully revised the Results, figure captions, and Methods to make the model description more complete and internally consistent.

      First, we clarified how spontaneous input events were generated. Specifically, V1 and S1 activity was constructed from independent events in each area and shared events across the two areas. These event streams were generated using Poisson processes, with the rates chosen such that the total event rate was matched across simulations while varying the fraction of shared versus independent events. We also clarified that each event activated a spatially contiguous group of neurons, thereby implementing local spatial correlations within each primary sensory area, while shared events activated corresponding topographic locations in V1 and S1.

      Second, in the Methods we clarified the use of the nonlinear transfer function in the decoding analysis shown in Figure 5. The simulated RL activity was transformed with a sigmoid nonlinearity before performing the regression analysis, and we have now added the corresponding equation (15) to the Methods.

      Third, we clarified the distinction between the numerical decoding analysis and the analytical calculation of the optimal weight matrix. The decoding analysis uses the nonlinear transformation described above, whereas the analytical calculation uses a linearized version of the model to obtain a tractable closed-form solution. We now state this explicitly in the Methods to avoid the impression that two inconsistent models were used.

      Finally, we corrected several typographical errors and checked that the Results, Methods, and figure captions use consistent terminology for the input generation, correlation structure, and decoding analysis.

      (6) While Figure 5 supports the paper's conclusion that a mixture of unimodal and bimodal neurons in RL optimizes information encoding, the authors missed an opportunity to strengthen the connection between the model and experimental data. Specifically, they could apply this reconstruction method to the experimental data and examine how RL's ability to reconstruct V1/S1 activity changes across development. Their model predicts that this performance would improve over time, and if this trend is observed in the experimental data, it would provide strong validation that these feedforward connections are developing in line with the model's predictions.

      We agree with the reviewer that applying the reconstruction analysis directly to the experimental data would provide an important additional test of the model. However, the current experimental datasets are not well suited for this analysis. The two-photon recordings used to characterize spontaneous activity in V1, S1, and RL were acquired sequentially rather than simultaneously, and therefore cannot be used to reconstruct V1/S1 activity from RL activity. In principle, a related analysis could be attempted using the wide-field recordings, which are simultaneous across cortical areas. However, these data have lower spatial resolution, include movement-related variability, and do not provide cellular-resolution measurements of RL activity. We explored this possibility, but the resulting reconstructions were not sufficiently reliable or interpretable to include in the manuscript.

      We now state this explicitly as a limitation in the Discussion and identify simultaneous multiarea recordings at cellular resolution as an important future test of the model. Such experiments would make it possible to determine whether the ability of RL activity to reconstruct V1/S1 activity improves across development, as predicted by the model.

      Reviewer #2 (Public review):

      The authors aim to investigate the role of spontaneous activity in shaping the development of multisensory integration in the brain, specifically focusing on the connections between primary visual and somatosensory sensory areas (V1 and S1) and a higher-order cortical area rostrolateral to V1 (RL). They seek to understand how spontaneous activity guides the formation of aligned topographic maps and the emergence of bimodal neurons in RL.

      First, the authors found that spontaneous activity in all three areas sparsifies over time, but S1 exhibits more mature patterns earlier than V1 and RL. They claimed that correlated activity among neighboring regions of these areas during development carries topographic information. These data were used to implement a computational model that employed Hebbian rules of synaptic plasticity. The model indicated that correlated spontaneous activity can generate topographic connectivity between S1/V1 and RL and bimodal neurons in RL. The model suggested that the more mature spontaneous activity in S1 can guide map alignment between V1 and RL. In addition, the model also suggested that a mixture of bimodal and unimodal neurons in RL is optimal for decoding information from V1 and S1.

      While the data presented in the manuscript is promising and provides preliminary insights into the role of spontaneous activity in multisensory integration, it would be beneficial to strengthen the experimental foundation regarding the correlation between V1, S1, and RL. Incorporating more rigorous spatio-temporal analyses of spontaneous activity could enhance the robustness of these findings.

      Here are some important concerns:

      (1) The analysis of how spatial topography influences activity correlations in Figure 2 has several issues.

      (1a) While squares in V1 and S1 covered a small area of these sensory areas, the correlated territories in RL covered the entire area of RL. The topographic map in V1 continues caudally, so where is the rest of the map in RL? Something similar applies to the relationship between S1 and RL.

      We thank the reviewer for pointing out this ambiguity. In the original version, the functional correlation maps were generated using separate seed locations in V1 and S1, and the interpretation relied heavily on thresholded RGB maps in which each pixel was assigned to the color channel with the strongest correlation. This made it difficult to directly compare the V1- and S1-seeded maps and could give the impression that the correlation structure extended differently across RL depending on the chosen seed area.

      We have therefore revised the analysis and presentation of Figure 2. Instead of using separate seeds in V1 and S1, we now use common seed locations in RL and compute the correlation of each RL seed with activity across the imaged cortical field. This allows us to ask more directly whether different locations in RL are associated with spatially distinct regions in both V1 and S1. We now show both the raw correlation maps, in which the RGB channels reflect the correlation values for the three RL seeds (new Figure 2E), and the thresholded/maximum channel representation, in which each pixel is assigned to the strongest of the three color channels (new Figure 2F). The raw correlation maps make the correlation structure visible without relying solely on thresholding, whereas the maximum-channel representation highlights the spatial ordering of the strongest correlations.

      With this revised analysis, the topographic relationship across V1, RL, and S1 is clearer and no longer depends on comparing separate V1- and S1-seeded maps. We also clarified in the figure legend and Methods how the RGB maps are computed, how the maximum-channel maps are generated, and how thresholded pixels are represented. In addition, we added Supplementary Figure S1 to show further functional-correlation-map examples across PN9, PN10, and PN13 recordings, with seed locations in V1, S1, or RL as indicated in each panel.

      (1b) It is essential to know how areas were drawn. High precision is required.

      Consistent delineation of cortical areas is absolutely essential for interpreting the functional correlation maps. We have therefore expanded the Methods to describe how cortical areas were delineated from the wide-field recordings. Briefly, recordings were acquired in a field of view defined relative to lambda and the midline, and cortical-area outlines were assigned using published reference maps together with the spatial organization of spontaneous activity patterns and functional correlation maps. This approach follows the procedure we previously validated for developmental wide-field recordings (Leighton et al., 2021).

      To make this transparent, we added Supplementary Figure S3, which illustrates how the reference-map-based outlines were overlaid on the imaging field of view and how functional correlation maps and individual network events helped identify the boundaries of V1 and neighboring areas. We also clarified this in the Methods.

      (1c) It is not clear if correlated activity means different events in sync or large events that cover 2 or all 3 cortical areas of interest. The figure points to the second option, which contradicts the size of events at these stages, mainly in the oldest mice analyzed here.

      The reviewer asks whether the correlations reflect spatially confined events occurring near-synchronously in different cortical areas, or instead large events spanning V1, RL, and S1. To clarify this point, we revised Figure 2 to show representative activity traces and individual frames from the wide-field recordings (new Figure 2B–D). These examples show that activity can be localized to distinct subregions within V1, RL, and S1 while occurring at similar times across areas. Thus, the observed correlations are not well explained by single large events spreading continuously across the entire imaged field.

      We have revised the Results and Figure 2 to make this clearer. In addition, the lagged cross-correlation analysis in Figure 2G shows that V1–S1 correlations peak near zero lag and decay for both positive and negative lags, arguing against a stereotyped travelling-wave-like propagation between the two primary sensory cortices as the dominant explanation for the observed correlations.

      (1d) It is fundamental to know in detail and provide examples of how the detection of events was performed. For instance, could the dispersion of light from an event in V1 close to RL cause the detection of activity in RL?

      The reviewer asks how events were detected in the wide-field recordings and whether light dispersion could lead to false-positive correlations between neighboring areas. We have clarified this point in the Methods. For the functional correlation analyses shown in Figure 2, we did not perform event detection. Instead, the correlation maps were computed from the continuous fluorescence time courses by calculating Pearson correlations between seed region activity and the activity of every pixel in the field of view. Thus, the functional correlation maps do not depend on detecting or assigning individual events.

      To address the concern about whether correlations could reflect light spread from large events rather than genuine co-activity across areas, we revised Figure 2 to include representative activity traces and individual frames from the wide-field recordings. These examples show that activity can be spatially confined to distinct subregions in V1, RL, and S1 while occurring at similar times across areas. This argues against the interpretation that the correlations are simply caused by a single event spreading continuously across the imaged field or by light dispersion from one area into another. We have also described the area delineation procedure in more detail in the Methods and added Supplementary Figure S3 to illustrate how activity patterns and functional correlation maps were used to assign outlines of distinct cortical areas.

      Although wide-field imaging cannot completely exclude minor contributions from light scattering near area borders, the spatially localized activation patterns and the topographically ordered correlation maps support the interpretation that the correlations reflect genuine nearsynchronous co-activity across V1, RL, and S1.

      (2) For the correlations among V1, S1, and RL, it is crucial to have a consistent method to delineate the borders of cortical areas. The authors mention in one sentence that areas were drawn according to a reference map. More details are needed to convince the reader that the borders are accurate, especially because their shape and position change with age.

      As described in our response to point 1b, we have expanded the Methods to clarify how cortical-area borders were delineated in the wide-field recordings. Briefly, recordings were acquired in a field of view defined relative to lambda and the midline, and cortical area outlines were assigned using published reference maps together with the spatial organization of spontaneous activity patterns and functional correlation maps. We also added Supplementary Figure S3, which illustrates how the outlines based on reference maps were overlaid on the imaging field of view and how functional correlation maps and individual network events helped identify the boundaries of V1 and neighboring areas. This makes the delineation procedure more transparent across animals and developmental ages.

      (3) The results from the model seem to be based on the initial bias in connectivity between neighboring cells from the different areas. Then, it seems straightforward that implementing correlated activity with Hebbian and synaptic depression rules will force the strengthening of connections between spatially close cells. Despite this apparent predisposition of the model towards a defined outcome, the flaws in the experimental data used prevent a rigorous interpretation of the computational model.

      We understand the reviewer’s concern that the initial topographic bias could predispose the model toward the emergence of topographic maps. However, the model results show that this bias is not by itself sufficient to determine the final organization (Figure 3B,E,G– I). When V1–S1 correlations are weak, many RL neurons decouple from the primary sensory inputs, resulting in poor topography and few bimodal neurons (Figure 3E,G–I). Conversely, when V1–S1 correlations are very strong, the two input maps become highly aligned, but topography is degraded because many RL neurons receive similar visual and somatosensory inputs at the same topographic location, thereby overriding the initial topographic bias (Figure 3E,G,H). Thus, the initial bias does not simply determine the final map structure. Rather, appropriate topography, map alignment, and the emergence of a mixture of unimodal and bimodal neurons require an intermediate level of correlated activity.

      We have revised the manuscript to make this interpretation clearer. We also strengthened the experimental basis for the activity structure used in the model by revising Figure 2 and the corresponding Results and Methods. The revised analyses now show near-synchronous but spatially distinct activation of V1, RL, and S1, a lagged cross-correlation analysis arguing against stereotyped travelling-wave-like propagation between V1 and S1, and functional correlation maps computed from common RL seed locations. Together, these additions clarify the spatial and temporal structure of the spontaneous activity used to motivate the model.

      Finally, as described in our response to Reviewer 1, point 4, we have extended the model to test the role of the initial bias more directly in combination with experimentally measured differences in V1 and S1 activity patterns. In this analysis, we incorporated these activity differences and examined how they interact with different levels of initial connectivity bias (Supplementary Figure S2). These simulations show that more mature S1-like activity patterns alone are not sufficient to generate the appropriate topographic and aligned maps, and that an initial connectivity bias is required. At the same time, consistent with Figure 3, this bias does not by itself determine the final organization; the outcome also depends on the temporal correlation structure of V1 and S1 activity.

      We agree that the initial topographic bias remains an important modeling assumption, consistent with the idea that coarse activity-independent mechanisms provide an initial scaffold for later activity-dependent refinement. We now present the model accordingly: not as showing that correlated activity alone creates topography from an entirely unstructured circuit, but as showing how structured spontaneous activity can refine an initially coarse topographic scaffold to produce aligned multisensory maps and a mixture of unimodal and bimodal RL neurons.

      (4) In the Introduction, the authors nicely and briefly explain the role of primary and higher order sensory cortices in information processing. They also explain how spontaneous activity during development helps to build these circuits by refining connections or establishing hierarchies. They continue explaining the relevance of aligning different topographic maps to allow multisensory integration. Then they provide some examples of sites of multisensory integration. This provides a general context for the data presented in the Results section; however, and importantly, there is no specific introduction of why they are interested in RL and its interaction with V1 and S1. The authors should introduce the RL area and explain why it is an interesting site for multisensory processing.

      We thank the reviewer for pointing this out. We have revised the Introduction to make the rationale for focusing on RL more explicit. Specifically, we now introduce RL as a higher-order cortical area located between V1 and S1 that receives topographically organized input from both primary sensory cortices and contains overlapping visual and tactile representations. We also clarify that RL is a particularly relevant area for studying multisensory map alignment because corresponding locations in visual and whisker space can converge onto the same RL neurons, including bimodal neurons. Finally, we expanded the Introduction to explain that RL has been implicated in visually guided tactile behaviors and cross-modal generalization, making it an appropriate model system for studying how aligned multisensory representations emerge during development.

      (5) The results shown in Figure 1 corroborate published data from Golshani et al, Rochefort et al, Murakami et al. While the reproduction of data is more than welcome, the authors should specify which part of the data is completely new and acknowledge clearly the rest as corroboration of previous data. The sentence "As described in previous experiments ..." partially acknowledges this fact but is not clear enough. In addition, the transition between this part of the manuscript and the next data is not smooth. Data seems to be used to feed the model so perhaps the organization of the manuscript leaves room for improvement.

      We thank the reviewer for pointing this out. We have therefore revised the Results to clarify that the developmental sparsification of spontaneous activity in V1 is consistent with previous work, including Portera-Cailliau, Konnerth, Hanganu-Opatz, Crair and Ohki labs as well as our own (Siegel et al. 2021) and that similar developmental trends in S1 and RL corroborate and extend these observations across the sensory and higher-order cortical areas analyzed here.

      We also clarified what is new in the present analysis. Specifically, our contribution is not simply to reproduce previously described developmental sparsification, but to compare V1, S1, and RL within the same experimental and statistical framework, revealing that S1 exhibits more mature activity features earlier than V1 and RL. We also revised the transition to the next section to make clearer how these measurements motivate the subsequent analysis of temporally and spatially correlated spontaneous activity between V1, S1, and RL.

      Reviewer #3 (Public review):

      Summary:

      The study by Dwulet et al. explores how the development of spontaneous neural activity in primary sensory cortices influences the co-alignment of multiple sensory modalities in higher order brain areas (HOAs). To address this question, they focus on connectivity between the primary visual (V1) and somatosensory (S1) cortices and an associative cortical area (RL) in mice. The authors combine experimental (wide-field and two-photon calcium imaging) and computational approaches to show that spontaneous activity matures at a different pace across these brain regions. Their data indicate that S1 develops more rapidly than V1, which is possibly beneficial for RL's integration of visual and somatosensory inputs through correlated spontaneous activity. Using a computational model, they demonstrate that a moderate correlation between V1 and S1 activity can optimally guide the formation of bimodal neurons in RL, which are crucial for maximizing the decodability of multisensory stimuli. This finding highlights the role of correlated spontaneous activity in primary sensory cortices in establishing co-aligned topographic multimodal sensory representations in downstream circuits.

      Strengths:

      The manuscript is well written and it provides strong enough evidence to support the main claim of the authors. The insights on the role of correlated activity on instructing co-aligned multisensory maps in HOAs are not trivial and are an important advancement for the field.

      Weaknesses:

      In the opinion of this reviewer, the study has no major weaknesses. A drawback of the work is that none of the predictions of the computational modeling have been corroborated through mechanistic experimental manipulations of early brain activity.

      We thank the reviewer for their positive assessment of the manuscript and for highlighting the importance of the model predictions. We agree that a direct mechanistic perturbation of early spontaneous activity would provide an important future test of the model. Such experiments could, for example, perturb the temporal correlation structure between V1 and S1 during the relevant developmental window and then test whether this affects the alignment of V1/S1 maps in RL and the emergence of bimodal RL neurons.

      In the present study, we focused on identifying candidate features of spontaneous activity that could instruct multisensory map alignment and testing their sufficiency in a computational model. We now explicitly acknowledge in the Discussion that causal perturbations of early spontaneous activity will be needed to validate the model predictions experimentally. We believe this provides an important direction for future work while preserving the main conclusion of the current study: that structured, moderately correlated spontaneous activity provides a plausible developmental mechanism for refining aligned multisensory representations in higher-order cortex.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Additional comments/suggestions for the figures:

      (1) In Figure 1D-G, some of the dots lie almost directly on top of each other, essentially "hiding" certain data points. Using different shapes for each of the three regions might help alleviate this issue and make the data more visually distinct.

      We thank the reviewer for this suggestion. We have revised Figure 1D-G so that the three cortical regions are shown with different marker shapes. This should make overlapping data points easier to distinguish and clarify that each point corresponds to the average value for one animal and cortical region at the indicated postnatal age.

      (2) In Figure 2D-E, RGB color values are used to represent the highest correlation coefficient across the three seeded areas. It would be more informative if these also depicted the magnitude of the correlations, possibly through a color gradient. Additionally, the black regions in these panels are not currently defined and should be clarified.

      We have revised the functional correlation-map analysis and its presentation in Figure 2, as suggested by the reviewer. In the revised figure, the main correlation-map panels now use three seed locations in RL and show the resulting correlations across the imaged cortical field. We present the maps in two complementary ways. First, the raw RGB correlation map shows the correlation values for all three seed locations, with the intensity of each color channel reflecting the magnitude of the corresponding Pearson correlation coefficient (new Fig. 2E). Second, the maximum-channel representation assigns each pixel to the seed location with the strongest correlation, while still preserving correlation strength through pixel intensity (new Fig. 2F).

      We have also added color scales to relate pixel intensity to correlation magnitude and clarified that black pixels in the maximum-channel representation correspond to pixels below the correlation threshold used for visualization. The figure legend and Methods now describe how the RGB maps and maximum-channel maps were computed. Finally, we added Supplementary Figure S1 with additional examples from PN9, PN10, and PN13 recordings, with seed locations in V1, S1, or RL as indicated in each panel. This illustrates that they are quite similar across the ages investigated here.

      (3) I found Figure 3F a bit difficult to interpret without referring to the Methods section for the definitions of Topography and Alignment. Since these definitions are relatively short and essential for understanding all the modeling figures, I suggest moving them into the main text where they are first introduced.

      The definitions of Topography and Alignment have been added to the text where they are introduced.

      (4) In Figures 3-5, it is unclear what causes the variability in the model’s responses, as there are two potential sources of randomness: the initial random connectivity matrix and the correlated inputs driving the system. Are either of these fixed? For example, is the distribution of dots along the y-axis in Figure 3G-H, which corresponds to zero correlation between V1 and S1, driven by variability in the initial connectivity matrix, the random timing of input events, or a combination of both? If it’s a combination, it would be interesting to tease this effect apart by fixing one form of randomness and recreating these plots.

      In the original simulations in Figures 3–5, neither source of randomness was fixed across runs: each point corresponds to an independent developmental realization with a newly sampled initial connectivity matrix and a newly sampled sequence of spontaneous input events. The initial connectivity was random but weakly biased toward matched topographic location, while spontaneous activity consisted of stochastic independent and shared events activating randomly chosen contiguous groups of neurons (as explained in the main text and Methods). Thus, for example, the spread of points at zero V1–S1 correlation in Figures 3G–H reflects a combination of variability in the initial connectivity and variability in the independent V1 and S1 event histories. At zero correlation, no shared V1–S1 events are present, so this spread does not reflect variability in correlated shared events, but rather run-to-run differences in the two independently refined maps.

      We have clarified this point in the text and figure legend. We agree that fixing one source of randomness while varying the other would be an interesting additional analysis to decompose the relative contribution of initial wiring versus input history. However, the goal of the present simulations was to characterize the ensemble of possible developmental outcomes when both initial connectivity and spontaneous activity vary, as expected biologically.

      This interpretation is also consistent with the earlier two-layer model from developmental refinements from retina/thalamus to V1 (Wosniack et al., eLife 2021) on which our model builds, where final receptive fields emerge from the interaction between weak biased initial connectivity and stochastic structured spontaneous activity. In the current three-layer extension, the same principle applies to two converging projections, from V1 to RL and from S1 to RL. The initial topographic bias constrains the possible map structure, while the spatiotemporal statistics of V1 and S1 activity determine whether the two maps remain separate, align, or collapse into overly bimodal representations.

      (5) The specific parameter values used to create the panels in the modeling figures (Figures 3 and 4) should be made clearer, at least in the figure captions. For example, in Figure 3E, the exact values for the “weak,” “medium,” and “strong” correlations should be provided. Additionally, Figure 4 does not mention the strength of the correlated input considered, which should be specified as well.

      The values for the weak, medium and strong correlations have been added to the figure caption of Figure 3. The input correlation for Figure 4 is also now specified in the figure caption.

      (6) There is an odd vertical line in Figure 3I that doesn’t appear to be discussed or defined. Its purpose should be clarified, or the line should be removed if it is unintentional.

      This line has been removed.

      (7) There is a typo in the caption for Figure 3. Panel 'K' should be panel 'J'.

      This typo has been corrected.

      (8) In the text, the authors write "With these connectivity refinements, the generated activity in RL became sparser in terms of amplitude and participation rate (Figure 3J)." While this appears to be the case for this single example, it is difficult to confirm without zooming in on the panel. These quantities should be computed across multiple instances, and a summary plot should be provided to support this statement.

      The experimentally measured developmental sparsification of RL activity is quantified (independent of the model) in Figure 1D–F.

      We see how the original wording placed too much weight on the illustrative example in Figure 3J. We have revised the text to clarify that Figure 3J shows a representative simulation illustrating how RL activity changes as V1/S1-to-RL connectivity refines, rather than a separate population-level quantification across model instances.

      At the same time, this example is not meant to introduce a new, unsupported mechanism. The model used here is an extension of our previous two-layer model of developmental refinements between retina/thalamus and V1, in which spontaneous activity refined feedforward receptive fields from thalamus to V1. In that study, we specifically quantified how receptive field refinement led to sparsification of cortical activity in V1 over development, including reduced event amplitude, reduced event size/participation, and reduced pairwise correlations (Wosniack et al., 2021). Thus, the example shown in Figure 3J is consistent with a mechanism that has already been systematically characterized in the simpler two-layer setting.

      In the present manuscript, the central modeling results concern the emergence of topography, alignment, and the balance of unimodal and bimodal RL neurons. We therefore have softened the corresponding statement and explicitly refer to Figure 3J as an illustrative example.

      (9) Figure 5C is a bit difficult to interpret. The corresponding text states, "However, when activity across V1 and S1 is moderately correlated, having some unimodal RL neurons can achieve a higher total maximum fraction of variance for both V1 and S1 compared to the purely bimodal case (Figure 5C)", from which I infer that these dots represent networks resulting from "moderate correlations." However, the exact range of correlations considered should be mentioned in the text or figure caption. Additionally, I find it unusual that some networks with close to 0% bimodal cells perform quite well in reconstructing both S1 and V1. Many data points overlap, but I notice quite a few pale dots in the upper right of the plot. I believe this should be addressed in the main text.

      We thank the reviewer for this helpful comment. We have added the correlation values used for the simulations in Figure 5C to the figure caption and clarified the interpretation in the Results. The high reconstruction performance for some networks with relatively few bimodal cells arises because, when V1 and S1 activity are not perfectly correlated, unimodal RL neurons can provide unambiguous information about activity in one sensory area. In contrast, a purely bimodal population can make it more difficult to distinguish whether one or both primary sensory cortices were active. Thus, for moderately correlated inputs, a mixture of unimodal and bimodal RL neurons can reconstruct both sensory areas better than a population composed entirely of bimodal neurons. We have revised the main text to make this point explicit.

      (10) The network schematics in Figures 3A and 5A could be improved to better illustrate the network setup using a similar approach as the one used by this research group in Wosniack et al. (2021). Adding arrowheads to the lines from V1/S1 to RL would clarify that these are purely feedforward inputs. It would also be helpful to depict that V1 and S1 are driven by correlated events that are spatially structured.

      We thank the reviewer for this helpful suggestion. We have revised the schematics in Figures 3A and 5A to make the feedforward nature of the model clearer by adding arrowheads to the projections from V1 and S1 to RL. We have also clarified the depiction and description of the input activity. Specifically, Figure 3C shows the spontaneous events driving V1 and S1 in the model, including shared events that are both temporally correlated and spatially structured across corresponding topographic locations in the two primary sensory areas. These shared events activate matched contiguous groups of neurons in V1 and S1, while independent events activate randomly chosen contiguous groups within each area. We have clarified this point in the Results and Methods.

      General comments regarding the text (including typos):

      (1) In Statistical analysis, "In wide-field calcium imaging (we re-analyzed data from [46] (Figure 1))..." should be referencing Figure 2.

      Typo fixed.

      (2) Right before Table 1, the authors mention that they ran the simulations for 500,000 milliseconds, which is 500 seconds. This doesn't seem long enough for the weights to approach their steady-state values given the inter-event interval. Since the example simulations in Figure 3 are 1,000 seconds long, I'm guessing this is a typo.

      Typo fixed. Indeed the simulations in Figure 3 were 1,000 ms (1 s) long.

      (3) The specific time step used for the simulations should be specified. Currently, the text only mentions "sufficiently small time steps".

      We have now specified the simulation time step in the Methods.

      (4) In the Rate-based network model section, you write "These biased weights decay with a Gaussian profile with increasing distance (Figure 3)), with amplitude a and spread s", but Table 1 denotes these parameters differently.

      We have corrected the notation so that the parameter names are consistent between the Methods and Table 1.

      (5) Currently, all differential equations are written as 1/tau*df/dt. Based on the units of your time constants (seconds), I believe these equations should be tau*df/dt.

      We have corrected the differential-equation notation.

      (6) Equations 5-6 and 8 should be differential equations.

      We have corrected these equations so that they are written as differential equations. These mistakes happened because we changed formats between from Word to Latex.

      (7) The expectation in Equation 8 is not clearly defined and I would think here that the W_ij's should be within expectations. In the next paragraph, the authors specify that they are interested in a specific case of W_ij's, but this condition has not been introduced yet.

      We thank the reviewer for pointing out this ambiguity. We have revised the text around Equation 8 to define the expectation more clearly and to introduce the specific steady-state connectivity configuration before it is used. Because the expectation is taken over the input activity statistics at steady state, the weights are fixed quantities in this calculation. Including W_ij inside the expectation would therefore not change the result, but we have revised the notation and explanatory text to make this clearer.

      (8) The expectation in Equation 8 is not clearly defined, and I believe that the W_ij’s should be included within the expectations (in the following paragraph, the authors mention that they are interested in a specific case of W_ij’s, but this condition has not yet been introduced).

      This comment is the same as the one above. Please see the point above for the reply.

      (9) At the start of "Optimal weight matrix for correlated input populations", you write that the vector X is M x 1. If that is the case X'X would be a 1x1 matrix. I'm not sure if you meant to write X as 1 x M or to examine XX'.

      We thank the reviewer for pointing out this dimensional inconsistency. We have corrected the notation in the Methods. The concatenated input vector X=[v; s] has size M x 1, so the relevant input covariance matrix is X X^T not X^T X. This covariance matrix has size M x M, as required for the eigenvector analysis. We revised the corresponding equations and explanatory text accordingly.

      (10) Equation 11 has an s_i on the right-hand side that should be a \mu_s.

      Typo fixed.

      Reviewer #2 (Recommendations for the authors):

      Some sentences may require more scientific rigor. For instance: "We found that activity between the visual and the somatosensory cortex is often, but not always, temporally synchronized.

      We have revised the Results to state the quantitative observations more explicitly. Specifically for this example, we now report that the average activity in V1 and S1 across PN9PN12 animals showed a range of Pearson correlation coefficients with a mean of approximately 0.5. We also describe the examples in Figure 2B-D as near-synchronous but spatially distinct activation of subregions in V1, RL, and S1, and we use the lagged cross-correlation analysis in Figure 2G to support the conclusion that V1-S1 correlations peak near zero lag rather than reflecting stereotyped propagation with a fixed delay.

      Reviewer #3 (Recommendations for the authors):

      Minor suggestions on how to improve some specific aspects of the manuscript.

      Introduction:

      (1) What do the authors mean when they write "Higher-order areas (HOAs) situated between primary sensory areas"? This sentence might need some editing.

      We have revised the sentence to clarify that we are referring to higher-order cortical areas that receive and combine inputs from multiple primary sensory areas. We now also state explicitly that some of these areas, including RL, are anatomically positioned between the primary sensory cortices whose inputs they integrate.

      (2) In later portions of the manuscript, it becomes clear what the authors mean when they write “whereby sensory neurons converge onto higher-order cortex while preserving space”, but I think that it would be beneficial if this statement would be better explained also in the introduction.

      This has been clarified in the introduction. Specifically, we now clarify that topographic convergence means that neurons representing corresponding regions of sensory space in different primary sensory areas can project to overlapping or nearby locations in higher-order cortex. In the case of RL, this means that visual and tactile representations with corresponding spatial organization can converge onto RL neurons, including bimodal neurons.

      (3) Could the authors provide some more information about RL and the rationale as to why it was chosen as the HOA that they investigated in the study?

      We have expanded the Introduction to make the rationale for focusing on RL more explicit. We now introduce RL as a higher-order cortical area located between V1 and S1 that receives topographically organized input from both primary sensory cortices. We also explain that RL contains overlapping visual and tactile representations, including bimodal neurons, and that corresponding locations in visual and whisker space can converge in RL. In addition, we now note that RL has been implicated in visually guided tactile behavior and cross-modal generalization. These anatomical and functional properties make RL a particularly suitable model system for studying how aligned multisensory representations emerge.

      Results:

      (1) "RL was found to slightly lag behind V1 and S1". On what evidence is this statement based upon? As far as I can understand, there are no significant differences between V1 and RL besides amplitudes being higher in RL, which I don't think can be univocally interpreted as a sign that RL lags behind V1 in the developmental profile.

      The evidence for a delayed RL maturation relative to V1 and S1 is limited and comes from the pattern of coefficient estimates in the linear mixed models, now shown in Supplementary Tables S1-S3, rather than from a robust difference across all measured activity features. We have therefore revised the Results to state more conservatively that RL and V1 develop more similarly during the second postnatal week, while S1 shows more mature activity features earlier in development. The full linear mixed-model comparisons using V1, S1, and RL as reference areas are provided in Supplementary Tables S1-S3.

      (2) Figure 1H is very hard to read.

      (a) The slopes and the intercepts have values that differ by orders of magnitude, so the slopes get squeezed and become invisible. Further, the different parts of the plots (e.g. the one of amplitude and duration) are almost overlapping, which is a bit confusing. Slopes and intercepts should also have different units of measure (see Equation 3), so I wonder how they can lie on the same axis. Can the authors try to plot the data in a manner that is easier to visually inspect?

      (b) Including the "reference" (V1) intercept in H is also a bit misleading, as one might intuitively interpret it as a difference between V1 and other brain areas. Perhaps the overall differences between brain areas (regardless of age) might be best represented in a plot without age on the x-axis (only brain area). Alternatively, one might point them out directly on the plots in DG.

      (c) In D-G, what do the individual dots represent? The legend states N=10 animals, but I only see ~6 dots per plot.

      We thank the reviewer for these helpful points. We have revised the caption of Figure 1H and added Supplementary Tables S1-S3, which provide the full linear mixed-model estimates for each choice of reference area. These tables report the intercepts, slopes, interaction terms, confidence intervals, and significance levels in a format that avoids placing quantities with different units and scales on the same visual axis.

      For the caption of Figure 1H: The V1 value corresponds to the model intercept at PN8, whereas the age coefficient corresponds to the slope for V1. The S1, RL, Age: S1, and Age: RL terms represent differences relative to this reference model. To avoid the impression that the V1 intercept represents a difference between areas, we now explicitly state that the coefficients in Figure 1H are interpreted relative to V1 at PN8, and that the complete comparisons using S1 and RL as reference areas are provided in Supplementary Tables S2 and S3.

      Finally, we clarified that the individual points in Figure 1D–G represent animal-level averages for each cortical area at the indicated age. The value N = 9 refers to the total number of animals included across the dataset, not to the number of animals at each postnatal age. Because recordings were distributed across ages and some points overlap visually, fewer points are visible in individual panels than the total N.

      (3) Figure 2B-C: at which lag does this correlation peak? Is it at 0ms? Or does one brain area precede/follow the other one?

      We thank the reviewer for this comment. We have revised Figure 2 to include a lagged V1–S1 cross-correlation analysis. The V1–S1 correlation peaks close to zero lag and decreases for both positive and negative lags, indicating that the dominant temporal relationship is near-synchronous rather than consistent with fixed-delay propagation from one primary sensory cortex to the other. The curves show a mild asymmetry, with somewhat stronger correlations when S1 precedes V1, but because the dominant peak is near zero lag, we interpret the data primarily as evidence for near-synchronous, spatially structured coactivity across areas rather than stereotyped travelling-wave propagation. We have added this interpretation to the Results and clarified the temporal-lag convention in the Figure 2 legend.

      (4) Figure 2D-E: in the methods section the authors report that "The actual color of each pixel represents the highest coefficient of correlation value across the three channels." I think that this important information should be included in the main text or the legend of the figure.

      We have changed Fig. 2 now to clarify the quantification of the functional correlation maps and also added the information requested by the reviewer to the figure legend.

      (5) Figure 3D: I think that it would be beneficial if the authors would highlight directly in the figure that those connectivity matrices are between V1/S1 and RL.

      This information has been added to the figure.

      (6) Figure 3I: does the vertical line correspond to the "critical amount of temporal correlation" (eq. 2)? If so, could the authors provide this information in the figure or the figure legend?

      This line was unintentional and has been removed.

      (7) It would be nice if the data that was generated for this study (and the data that has already been published and was used to generate Figure 2) would be made publicly available on an open-access repository.

      We agree that open data sharing is important. We have made the code used for the model and figure generation available in the repository listed in the Data and Code Availability section. At present, we are not able to deposit the complete raw imaging datasets in an open repository because the wide-field and two-photon imaging files are very large, amounting to multiple terabytes, and we do not currently have a sustainable hosting solution for these raw data. We will share data upon request, and we will deposit the raw imaging datasets in an appropriate open repository if a feasible long-term hosting solution becomes available.

      References

      M. Chini, T. Pfeffer, and I. Hanganu-Opatz. An increase of inhibition drives the developmental decorrelation of neural activity. eLife, 11:e78811, 2022.

      P. Golshani, J. T. Gonçalves, S. Khoshkhoo, R. Mostany, S. Smirnakis, and C. PorteraCailliau. Internally mediated developmental desynchronization of neocortical network activity. Journal of Neuroscience, 29(35):10890–10899, 2009.

      A. Gribizis, X. Ge, T. L. Daigle, J. B. Ackman, H. Zeng, D. Lee, and M. C. Crair. Visual cortex gains independence from peripheral drive before eye opening. Neuron, 104(4):711–723.e3, 2019.

      S. Lakhera, E. Herbert, and J. Gjorgjieva. Modeling the emergence of circuit organization and function during development. Cold Spring Harbor Perspectives in Biology, 17(2):a041511, 2025.

      A. H. Leighton, J. E. Cheyne, G. J. Houwen, P. P. Maldonado, F. De Winter, C. N. Levelt, and C. Lohmann. Somatostatin interneurons restrict cell recruitment to retinally driven spontaneous activity in the developing cortex. Cell Reports, 36(1):109316, 2021.

      H. Matsumoto, T. Murakami, and K. Ohki. Topographic correspondence between retinotopic and whisker somatosensory map in mouse higher visual area and its development. Frontiers in Neural Circuits, 19:1552130, 2025.

      T. Murakami, T. Matsui, M. Uemura, and K. Ohki. Modular strategy for development of the hierarchical visual network in mice. Nature, 608:578–585, 2022.

      N. L. Rochefort, O. Garaschuk, R.-I. Milos, M. Narushima, N. Marandi, B. Pichler, Y. Kovalchuk, and A. Konnerth. Sparsification of neuronal activity in the visual cortex at eyeopening. Proceedings of the National Academy of Sciences of the United States of America, 106(35):15049–15054, 2009.

    1. eLife Assessment

      This study in the Drosophila antennal lobe, which contains multiple non-equivalent sensory channels, provides valuable new insight into how early-life sensory experience can produce lasting, cell-type-specific changes in neural circuit function. The work demonstrates that glial-mediated pruning during a defined developmental window leads to persistent suppression of odor responses in one olfactory neuron type, while sparing another. The evidence is convincing and supported by multiple complementary approaches, although some mechanistic interpretations remain speculative and would benefit from additional functional testing.

    2. Reviewer #1 (Public review):

      Summary:

      This study builds on earlier work showing that early-life odor exposure can trigger glial-mediated pruning of specific olfactory neuron terminals in Drosophila. Moving from indirect to direct functional imaging, the authors show that pruning during a narrow developmental window leads to long-lasting suppression of odor responses in one neuron type (Or42a) but not another (Or43b). The combination of calcium and voltage imaging with connectomic analysis is a strength, though the voltage imaging results are less straightforward to interpret and may not reflect synaptic output changes alone.

      Strengths:

      Biologically, one of the main strengths of this work is the direct comparison between two odor-responsive OSN types that differ in their long-term adaptation to early-life odor exposure. While Or42a OSNs undergo pruning and remain persistently suppressed into late adulthood, Or43b OSNs, which also respond to the same odor, show little lasting change. This contrast not only underscores the cell-type specificity of critical-period plasticity but also points to a potential role of inhibitory network architecture in determining susceptibility. The persistence of the Or42a suppression well beyond the developmental window provides compelling evidence that early glia-mediated pruning can imprint a stable, life-long functional state on selected sensory channels. By situating these functional outcomes within the context of detailed connectomic data, the study offers a framework for linking structural connectivity to long-term sensory coding stability or vulnerability.

      Comments on revised version:

      I thank the authors for their careful revision and thoughtful responses to the reviewers' comments. The revised manuscript addresses my previous concerns in a satisfactory manner, and the interpretation of the findings has been appropriately clarified and balanced. I have no further major comments.

    3. Reviewer #2 (Public review):

      Recent work from the authors identified the synaptic changes and glial reaction that occurs during exposure of a Drosophila odorant receptor neuron population to continued exposure of a stimulating odorant. This work markedly advanced our understanding of cellular response to critical periods. This current Advance manuscript carries that work forward and examines the non-autonomous responses to constant odorant exposure. The authors discover that the changes to ORN populations are not accompanied by changes to either PN dendrite or PN axon volume, nor are they concurrent with changes in postsynaptic PN structures. These changes are, however, notable accompanied by changes in Ca2+ and voltage responses in ORNs. Importantly, this set of responses is specific for the Or42a ORNs (that are highly sensitive to the odorant in question, ethyl butyrate) and not the Or43b ORNs (which respond to ethyl butyrate, but not as drastically). Finally, the authors include connectomics analyses showing that Or43b and Or42a ORNs differ in their synaptic input/output relationships.

      This is an excellent use of the Advance mechanism for the journal as these are important follow-up findings for the parent story. The non-autonomous effects (or lack thereof) on PNs is an important part of the story as is the functional response of Or42a ORNs and the differing response of similarly (but not identically) sensitive Or43b ORNs. The experiments are well conceived, controlled, and conducted. Where the story falters a bit, though, is with the connectomics analysis. The authors show distinct differences between Or43b and Or42b ORN input output relationships and suggest that those differences may underlie the differences observed in their response to ethyl butyrate exposure during the critical period. This is certainly a possibility, but as it stands now, it is too disconnected to offer significant proof. There would have to be additional experiments to address this. Right now, the inclusion of the connectomics work feels like a distraction at best, and a complete non sequitur at worst. To be clear, the connectomics work is well done and I have no issues with its validity, but is not helpful to the central thesis of the work. I would suggest the authors either remove it entirely or strongly rethink how it fits into the paper.

      Comments on revised version:

      I appreciate the consideration of my comments and the authors' responses. The additional data on PN synapse number is intriguing (and welcome) as is the text discussing potential postsynaptic compensatory mechanisms. I respect the authors' decision in retaining the connectivity analysis, but despite the textual changes, I still feel that it is peripherally related to the main thesis of the work and would best be omitted from the paper and included in a separate, more relevant study. Ultimately, though, that is their choice.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study builds on earlier work showing that early-life odor exposure can trigger glial-mediated pruning of specific olfactory neuron terminals in Drosophila. Moving from indirect to direct functional imaging, the authors show that pruning during a narrow developmental window leads to long-lasting suppression of odor responses in one neuron type (Or42a) but not another (Or43b). The combination of calcium and voltage imaging with connectomic analysis is a strength, though the voltage imaging results are less straightforward to interpret and may not reflect synaptic output changes alone.

      Strengths:

      Biologically, one of the main strengths of this work is the direct comparison between two odor-responsive OSN types that differ in their long-term adaptation to early-life odor exposure. While Or42a OSNs undergo pruning and remain persistently suppressed into late adulthood, Or43b OSNs, which also respond to the same odor, show little lasting change. This contrast not only underscores the cell-type specificity of critical-period plasticity but also points to a potential role of inhibitory network architecture in determining susceptibility. The persistence of the Or42a suppression well beyond the developmental window provides compelling evidence that early glia-mediated pruning can imprint a stable, life-long functional state on selected sensory channels. By situating these functional outcomes within the context of detailed connectomic data, the study offers a framework for linking structural connectivity to long-term sensory coding stability or vulnerability.

      Weaknesses:

      The narrative begins with the absence of changes in PN dendrites and axons. While this establishes specificity, it is a relatively weak starting point compared to the novel OSN functional results.

      We agree that switching the order of Figures 1 and 2 recontextualizes the negative PN morphology findings to make their significance more clear, especially with the addition of PN odour-evoked activity data (see Figures 2A, B of the revised manuscript).

      Calcium imaging with GCaMP, though widely used, is an indirect measure of synaptic function, and reduced signals could reflect changes in non-synaptic calcium influx as well as release probability. The interpretation of the voltage imaging results is also unclear: if suppression were solely due to impaired synaptic release, one might expect action potential-evoked voltage signals to remain unchanged. The reported changes raise the possibility of deficits in action potential initiation or propagation, which would shift the mechanistic explanation.

      Although it is true that non-synaptic Ca<sup>2+</sup<> influx could contribute to odour-evoked signals in OSN axon terminals, it seems likely to be a relatively small contribution when compared to Ca<sup>2+</sup> influx via voltage-gated Ca<sup>2+</sup> channels at the active zone. Given the observation that synaptic markers are eliminated during this form of critical period plasticity and remain decreased even after OSNs regrow their terminals days later (consistent with our observed continued decrease in odour-evoked responses), the most parsimonious explanation is that we are seeing a reduction in synaptic Ca<sup>2+</sup> influx. We cannot dismiss the possibility that there is a decreased voltage signal arising from fewer action potentials being elicited by the odour stimulation. However, the reduction in voltage signal must arise at least in part from the observed reduction in Ca<sup>2+</sup> influx. We have therefore provided additional text to this effect in the results section.

      The difference between Or42a and Or43b OSNs is attributed to varying inhibitory input densities from connectome data, but this remains speculative without functional tests such as manipulating GABA receptor expression in OSNs. In Or43b, there is essentially no strong phenotype, making it premature to ascribe the absence of suppression solely to inhibitory connectivity.

      We have tempered our conclusions to posit additional mechanisms that could explain the more mild pruning that occurs for Or43b OSNs. While the pruning phenotype for Or43b OSNs is not as strong as Or42a, it is not absent. To further explore the contribution of inhibition as a candidate mechanism underlying differences in susceptibility of Or42a and Or43b to this form of critical period plasticity we compared the relative impact of knocking down expression of GABA-A receptor (called “rdl”) in Or42a and Or43b OSNs. Consistent with the degree of pruning being regulated inhibition, knocking down expression of rdl enhanced pruning for both Or42a and Or43b OSNs. However, because the magnitude of the enhancement was similar between both OSN types, we agree with the reviewer that inhibitory connectivity cannot be the sole mechanism that explains the difference and have therefore tempered our language appropriately.

      Finally, the study does not connect circuit-level changes to behavioral outcomes; assays of odor-guided attraction or discrimination could place the findings in an organismal context.

      We agree that behavioral assays will be a critical component for understanding the functional consequences of this form of critical period plasticity. However, the goal of this study was to extend our prior work to determine the longevity and selectivity of the critical period pruning. Behavioral assays testing the consequences of this form of early life plasticity will be a component of future studies.

      Some introduction material overlaps with the authors' 2024 paper, and the novelty of the present study could be signposted more clearly.

      We have included text to highlight the novelty of the present study.

      Reviewer #2 (Public review):

      Recent work from the authors identified the synaptic changes and glial reaction that occur during exposure of a Drosophila odorant receptor neuron population to continued exposure of a stimulating odorant. This work markedly advanced our understanding of cellular response to critical periods. This current Advance manuscript carries that work forward and examines the non-autonomous responses to constant odorant exposure. The authors discover that the changes to ORN populations are not accompanied by changes to either PN dendrite or PN axon volume, nor are they concurrent with changes in postsynaptic PN structures. These changes are, however, notable, accompanied by changes in Ca2+ and voltage responses in ORNs. Importantly, this set of responses is specific to the Or42a ORNs (that are highly sensitive to the odorant in question, ethyl butyrate) and not the Or43b ORNs (which respond to ethyl butyrate, but not as drastically). Finally, the authors include connectomics analyses showing that Or43b and Or42a ORNs differ in their synaptic input/output relationships.

      This is an excellent use of the Advance mechanism for the journal, as these are important follow-up findings for the parent story. The non-autonomous effects (or lack thereof) on PNs is an important part of the story, as is the functional response of Or42a ORNs and the differing response of similarly (but not identically) sensitive Or43b ORNs. The experiments are well-conceived, controlled, and conducted. Where the story falters a bit, though, is with the connectomics analysis. The authors show distinct differences between Or43b and Or42b ORN input-output relationships, and suggest that those differences may underlie the differences observed in their response to ethyl butyrate exposure during the critical period. This is certainly a possibility, but as it stands now, it is too disconnected to offer significant proof. There would have to be additional experiments to address this. Right now, the inclusion of the connectomics work feels like a distraction at best, and a complete non sequitur at worst. To be clear, the connectomics work is well done and I have no issues with its validity, but it is not helpful to the central thesis of the work. I would suggest the authors either remove it entirely or strongly rethink how it fits into the paper.

      We have tempered our stated interpretations of the connectivity analysis and include new experiments examining the impact of GABA signaling on pruning. We have therefore opted to retain the connectivity analysis as we feel that it has been better integrated into the overall narrative of the paper.

      Major Concerns:

      (1) The examination of PN axon terminals in the MB and LH is interesting, but it is only one possibility. Oftentimes, the volume of neurons remains constant with perturbation, while the synapse number is affected. Figure 1C and E would be greatly helped by examining synapse number (via Brp or Brp-Short) in the PN axons.

      We agree that the counting synapse number would provide greater resolution information about synapse function relative to axon volume and have added this analysis to what is now Figure 2.

      (2) The use of dlg1[4K] is a strong use of a new tool, but the result is surprising. The presynaptic ORN synapse number onto the PNs is notably changed, but that is not reflected in a postsynaptic PSD-95 change. That suggests a compensatory mechanism that the authors might explore. A good proportion of PN puncta should be postsynaptic to those ORNs, so why aren't they adjusted?

      We agree that this result suggests that a compensatory mechanism may be present. We have therefore added new text to point out this observation and potential explanation.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The interpretation of the voltage imaging results would benefit from clarification. If these signals are reduced because of upstream action potential changes rather than synaptic release, this should be explicitly discussed and illustrated with representative raw traces for both OSN types. The proposed link between inhibitory connectivity and selective vulnerability could be tested more directly, for example, by manipulating GABA receptor function in OSNs.

      We have now tested the link between inhibitory connectivity and susceptibility to glial pruning by testing the effects of GABA receptor knockdown in either Or42a or Or43b OSNs (fully described above).

      Adding an intermediate post-exposure time point for Or42a responses could help resolve whether suppression is immediate or develops over time.

      The suppression of Or42a odour-evoked responses is present immediately after the 2 day exposure period and responses remain suppressed until 25 days post-eclosion, indicating that the suppression is immediate and sustained. We therefore respectfully disagree that another physiological time point will help resolve whether the suppression is immediate or develops over time.

      In terms of presentation, the introduction could be tightened to reduce overlap with the 2024 paper, figures should have clear axis labels and consistent terminology for neuron types and glomeruli, and a schematic summarising key inhibitory connections for Or42a vs. Or43b would aid clarity

      We have now streamlined the introduction, improved clarity on axis labels and checked for consistency of terminology.

      Minor Concerns:

      (1) The dlg1[4K] is made with a V5 epitope but the authors have it labeled mCD8::GFP in Figure 1F. This is likely a typo and should be corrected.

      This typo has now been corrected.

      (2) Can the responses be separated in Figures 2A, C, and E? It is difficult to see the differences in oil and EB exposure. This would make it much more straightforward to tell the difference if both traces were clearly visible.

      Overlaying the averaged response traces for in Figure 2C, E and G (now Figures 1C, E and G) enables the reader to make direct visual comparisons between the responses of OSNs from flies in each condition to both mineral oil and ethylbutyrate. Separating the individual traces would make it much more difficult to make these comparisons.

    1. eLife Assessment

      This is an important study that provides evidence that GATA6-dependent programming of peritoneal macrophages helps to regulate lipid metabolism and influences eosinophil survival. The evidence linking GATA6 deficiency to altered lipid profiles and eosinophil accumulation is solid, although the proposed mechanistic pathway connecting sphingolipid remodelling, LTE4 production and eosinophil survival currently remains incomplete. Strengthening these links and/or appropriately calming the conclusions would help to increase the impact of the study.

    2. Reviewer #1 (Public review):

      Summary:

      Recent findings have established that macrophage function is tailored to individual tissues through upregulation of tissue-specific transcription factors in response to local microenvironmental signals. However, how these transcriptional pathways affect macrophage lipid metabolism and the importance of this for homeostasis of neighbouring immune cells remains relatively uncharted. One exemplary pathway is the specific expression of GATA-6 by macrophages within the serous cavities that is triggered by local retinoic acid production. Here, Czubala et al have used mice with macrophage-specific deletion of GATA6 (GATA-6KO-mye) to study the importance of tissue-specific macrophage programming in regulating the macrophage and tissue lipidome and the functional importance of this for the regulation of eosinophil numbers in the tissue.

      Strengths and Weaknesses:

      The authors show accumulation of lipid-rich vesicles in the absence of GATA6, which lipidomic analysis suggests are largely comprised of sphingolipids and glycophospholipids. Using published transcriptional data identifies candidate genes in GATA6-deficient cells that may underlie these changes. Manipulating two of these candidate genes, Gba2 and Smpd1, in a macrophage cell line leads to similar changes in sphingolipid composition to those in GATA6-deficient macrophages in vivo, supporting the hypothesis that tissue specialisation of peritoneal macrophages induces transcriptional changes via GATA6 that directly control sphingolipid metabolism. GATA6 deficiency is then shown to affect the oxylipin content of peritoneal macrophages and peritoneal fluid, including higher levels of LTE4 in fluid. Elevated expression of the Ltc4s in GATA6-deficient cells is predicted as the likely mechanism leading to elevated LTE4.

      To determine the functional effects of altered lipid metabolism, and specifically LTE4, the authors focus on the elevated accumulation of peritoneal eosinophils previously reported to occur in GATA-6KO-mye mice. They show that eosinophils undergo less apoptosis in these mice and the absence of a measurable increase in known eosinophil chemokines leads them to conclude that eosinophil numbers arise through increased longevity. However, this point remains to be formally demonstrated, and directly measuring the longevity of eosinophils in the cavity would greatly strengthen their conclusions. The authors then examine known regulators of eosinophil survival, IL-5 and GM-CSF. They convincingly demonstrate a role for IL-5 in the regulation of peritoneal eosinophil numbers but conclude that survival factors other than IL-5 and GM-CSF likely control the differential numbers in control and GATA-6KO-mye mice, given IL-5 was observed to be a general survival signal in both genotypes and that no difference in the levels of these growth factors was observed in lavage fluid between genotypes. The authors then blocked production of prostaglandins using the inhibitor indomethacin. This treatment also led to a general reduction in survival and number of eosinophils in both control and GATA-6KO-mye mice, leading to the conclusion that altered prostaglandin production is not the underlying mechanism regulating elevated eosinophil numbers in the absence of GATA6.

      One weakness in these conclusions is that if the GATA-6-KO-mye phenotype does lead to increased production of a homeostatic growth factor for eosinophils, then inhibition/blockade of such a factor would be expected to lead to loss of eosinophils in both WT and GATA-6KO-mye mice. Furthermore, cytokines, chemokines, and lipid mediators can be rapidly bound and removed or metabolised in vivo by their receptors, meaning detecting an increase in production in body fluids can be difficult.

      Finally, they block production of LTE4 using an inhibitor of the upstream enzyme 5-LO. This treatment reduces eosinophil survival and number in GATA-6KO-mye mice, from which the key conclusion is drawn that elevated LT4E is responsible for the increased survival and accumulation of eosinophils in GATA-6KO-mye mice. The major weakness here is that the equivalent experiment in control mice to determine if inhibition of 5-LO leads to a general reduction in survival/number of eosinophils or if this effect is restricted to the GATA-6KO-mye appears not to have been performed.

      Impact and context:

      Overall, this study demonstrates key alterations in lipid metabolism and lipid mediator release resulting from loss of GATA6 expression in peritoneal macrophages, and links this to the elevated survival/accumulation of eosinophils that occurs concurrently in GATA6-KO-mye mice. The role of endogenous LTE4 in regulation of eosinophil survival and/or migration into tissues is exciting and opens up a new avenue of research for understanding the importance of this pathway in regulation of eosinophils across tissues and during disease. Furthermore, unlike in the mouse, GATA6-expressing macrophages represent only a minor proportion of macrophages in the human peritoneal cavity, while the dominant GATA6-negative population is more equivalent to the GATA6-KO-mye cells studied here (PMID: 38102487). Hence, the data presented in the current manuscript could have important implications for how eosinophil numbers and lipid metabolism may be regulated by these cells in people.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript examines how GATA6-dependent programming of resident peritoneal macrophages regulates their lipidome and, in turn, eosinophil homeostasis, combining lipid imaging, mass spectrometry, transcriptional analysis and in vivo pharmacology. BODIPY/CARS microscopy with targeted lipidomics convincingly shows substantial lipid changes following myeloid GATA6 deficiency, particularly in sphingolipids, with Smpd1 and Gba2 manipulations providing mechanistic support.

      Strengths:

      The authors also connect these changes to eosinophil biology, confirming increased peritoneal eosinophils in Gata6-deficient mice with reduced apoptosis (via two methods) rather than increased production. Testing of alternative explanations (chemokines, IL-5, prostaglandins, 12/15-LOX products) strengthens the argument by narrowing candidate mechanisms. The identification of increased LTE4 is notable as it correlates with eosinophil abundance, and zileuton reduces LTE4, eosinophil numbers, and increases apoptosis. This supports a role for 5-LOX/cysteinyl-leukotriene signalling.

      Weaknesses:

      The principal weakness is specificity: zileuton affects the broader leukotriene pathway, not LTE4 alone, so correlation with LTE4 doesn't establish causality. This matters more given the LTE4 receptor remains unidentified (to the best of my knowledge). Similarly, the proposed transcellular biosynthesis mechanism (ImmGen data suggesting complementary enzyme expression across cell types converting LTC4 to LTE4) is inferential; direct evidence of cellular source and transfer is lacking.

      Design limitations include reliance on pooled animals in some lipidomic measurements, small replicate numbers, and a stronger eosinophil phenotype in females that shifts subsequent analysis toward females. Indeed, this sex dependence deserves more discussion given it limits generalisability.

      Overall, this is a technically strong, conceptually interesting study. The core conclusions, that GATA6-dependent regulation of the macrophage lipidome and a role for cystl Lts in eosinophil survival, are well supported. The more specific claim that LTE4 is the causal factor via a defined transcellular pathway is plausible but not yet firmly established. Experimental strengthening or moderated claims would improve the study.

    4. Reviewer #3 (Public review):

      Summary:

      The authors sought to define how GATA6-dependent programming of resident peritoneal macrophages regulates lipid metabolism and, in turn, eosinophil survival. By integrating a myeloid-restricted GATA6-deficiency model with cellular phenotyping, lipidomic analyses, and measurements of lipid mediators, the study attempts to connect macrophage transcriptional identity to sphingolipid and cysteinyl leukotriene pathways that may shape eosinophil persistence. The work also appears intended to provide a mechanistic bridge between prior observations from this group and others regarding GATA6-positive macrophages, lipid metabolism, and eosinophil homeostasis.

      Strengths:

      (1) The study addresses an important and understudied question: how tissue-resident macrophage identity controls the local lipid environment and thereby influences eosinophil survival.

      (2) The use of a genetically defined GATA6-deficiency model provides a biologically relevant framework for testing the contribution of macrophage programming.

      (3) The lipidomic data broaden the analysis beyond a single mediator and identify coordinated changes in sphingolipids and glycerophospholipids that may generate useful hypotheses for the field.

      (4) The finding that GATA6 deficiency promotes eosinophil survival is clear, potentially important, and consistent with prior work cited by the authors.

      (5) The study is performed by a knowledgeable team and brings together macrophage biology, eosinophil biology, and lipid metabolism in a way that is likely to interest several research communities.

      Weaknesses:

      (1) The central mechanistic chain-GATA6 deficiency leading to altered sphingolipid abundance, altered LTE4 production, and consequently increased eosinophil survival-is not fully demonstrated. The data support associations among these features, but the causal order remains uncertain.

      (2) The cited literature linking sphingolipid and cysteinyl leukotriene biosynthesis does not substitute for direct testing in this model. Perturbation or rescue experiments targeting sphingolipid synthesis and cysteinyl leukotriene production would be needed to establish necessity and directionality.

      (3) The broader lipidomic changes complicate the emphasis on sphingolipids. Because glycerophospholipids are also increased, the phenotype may reflect more extensive membrane-lipid remodeling, altered phospholipase activity, changes in the Lands cycle, or shifts in free fatty-acid availability.

      (4) The manuscript would benefit from a clearer distinction between observations made directly in GATA6-deficient peritoneal macrophages and mechanistic inferences extrapolated from prior studies.

      (5) The physiological and pathological relevance is not yet sufficiently established. It remains unclear whether enhanced eosinophil survival translates into altered eosinophil accumulation, activation, or tissue injury during inflammatory disease in the peritoneal cavity or lung.

    5. Author response:

      We would like to thank all the reviewers and the editors for their considerate evaluation of our study.

      We are pleased that overall the reviewers were positive about the bulk of our study establishing a role of tissue macrophage programming/specialisation in regulating the macrophage lipidome, in the peritoneum, including the exemplar sphingolipid class. The reviewers raise understandable issues about the specificity of the available inhibitory compounds, such as zileuton meaning that conclusive statements about the role of LTE4 are not possible.

      In a revised manuscript, we will address all points but predominantly focus on the second aspect of the study, ensuring that reviewers comments are addressed appropriately, detailing and weaknesses, or ambiguities, with our study. This will include, but will not be limited to:

      - Further commentary on the regulation of eosinophil numbers within the tissue;

      - Addressing the specificity of zileuton and the implications of this for interpretation of our results with respect to eosinophil biology;

      - More careful framing of the transcellular biosynthesis potential;

      - A detailed discussion of sex dependency with regard to eosinophil numbers in general and any potential effect on the reported Gata6-dependent phenomenon;

      We are grateful for the constructive comments.

    1. eLife Assessment

      Avoidance of UV and blue light by the nematode C. elegans is mediated by the unusual transmembrane protein LITE-1, a non-canonical photoreceptor. In this valuable work, the authors report the surprising finding that LITE-1 function is also required for avoidance of very high concentrations of the food-associated odorant diacetyl. Although no molecular evidence is provided, the paper reports convincing studies indicating that LITE-1 can serve as a chemoreceptor for very high concentrations of diacetyl, adding an unexpected layer of complexity to the function of this unusual protein.

    2. Reviewer #1 (Public review):

      Summary:

      This paper describes an interesting phenotype of C. elegans lite-1 mutants. Previous work showed that lite-1 mutants lose a violet / blue light avoidance response. The authors show here that lite-1 mutants also show a defect in negative diacetyl chemotaxis. While wild-type worms avoid diacetyl at high concentrations, lite-1 mutants are instead *attracted* to it. The authors go on to perform Ca2+ imaging in sensory neurons and find that ADL and ASK neurons show altered Ca2+ responses to diacetyl in lite-1 mutants, suggesting LITE-1 is required for these responses. As unc-13 mutants with defective synaptic transmission show similar diacetyl Ca2+ responses as wild-type, this suggests these neurons respond cell autonomously to diacetyl. Indeed, expression of LITE-1 in ADL from a specific promoter shows phenotypic rescue. The authors then use a strain that expresses LITE-1 in the body wall muscles and show this expression is sufficient to engender them with sensitivity to diacetyl, as measured through altered swimming, hypercontractility, and egg laying. The authors interpret this result as LITE-1 may act as a diacetyl receptor. The authors test whether a structurally similar molecule, 2,3 pentanedione shows similar effects, and they find it does. Alpha-fold modeling and molecular docking analysis show where diacetyl might bind to the LITE-1 protein. They then test whether lite-1 mutants show chemotaxis defects to other molecules as seen with diacetyl.

      Strengths:

      Overall, the study follows up on an interesting and useful result. The experiments as presented are generally well-conceived and performed. The authors use a variety of behavior and imaging approaches to test how LITE-1 mediates diacetyl avoidance. The author revisions addressed the concerns I raised previously.

      Weaknesses:

      In response to the first submission, Reviewer 3 raised the possibility that light facilitates the production of diacetyl which then activates LITE-1. The authors helpfully revised the manuscript to incorporate this mechanisms. However, is it possible that diacetyl and 2,3-pentanedione are instead (or also) acting as photosensitizers, generating an(other) activator of LITE-1? Diacetyl has been previously shown to have chemical reactivity which is enhanced by light (citations below). I realize that the experiments have ruled out a role for acute light exposure in causing phenotypes in some of the experiments, but it is formally possible that prior light exposure may have caused diacetyl to generate peroxides or other photo-products that have the observed biological effect which is then lost in the lite-1 mutant. That is, what if the relevant molecule is already present in the diacetyl bottle / stock solution? At that point, further light exposure may not matter. This possibility was not really addressed in the manuscript.

      -Huang CY, Li J, Liu W, Li CJ. Diacetyl as a "traceless" visible light photosensitizer in metal-free cross-dehydrogenative coupling reactions. Chem Sci. 2019 Apr 8;10(19):5018-5024. doi: 10.1039/c8sc05631e. PMID: 31183051; PMCID: PMC6530541.<br /> -Pengcheng Lian, Ruyi Li, Xiao Wan, Zixin Xiang, Hang Liu, Zhiyu Cao, Xiaobing Wan Acetylation of alcohols and amines under visible light irradiation: diacetyl as an acylation reagent and photosensitizer. Organic Chemistry Frontiers 2022, 9 (2), 311-319.<br /> -Rowell, Keiran N & Kable, Scott & Jordan, Meredith J. T. (2022). An assessment of the tropospherically accessible photo-initiated ground state chemistry of organic carbonyls. Atmospheric Chemistry and Physics. 22. 929-949. 10.5194/acp-22-929-2022.

    3. Reviewer #2 (Public review):

      Summary:

      Koh and colleagues investigate the broader sensory role of LITE-1, a gustatory receptor previously linked to UV light detection in C. elegans. Their study explores whether LITE-1 also mediates avoidance of specific chemical stimuli-namely, high concentrations of diacetyl and 2,3-pentanedione. They show that LITE-1 is required in the ADL and ASK neurons for calcium responses to diacetyl, and that its expression in body-wall muscles is sufficient to trigger hypercontraction upon odorant exposure. Molecular docking suggests both odorants may directly bind to LITE-1 with micromolar affinity. These findings suggest LITE-1 may act as a multimodal receptor for both light and chemical stimuli.

      Strengths:<br /> • Methodological Precision: The study is technically strong, with well-executed calcium imaging and quantitative behavioral assays that clearly show neural and muscular responses to chemical stimuli.<br /> • Novelty and Scope: The work presents a compelling case for LITE-1 functioning as a multimodal sensor, which is an intriguing expansion of its known role.<br /> • Potential Impact: If validated, the findings could significantly advance the understanding of sensory integration in C. elegans, and the tools developed may be broadly useful to the research community.<br /> • Relevance to the Field: The study adds to evidence that C. elegans uses non-canonical sensory pathways and may inspire further exploration of multimodal receptor functions in other systems.

      Weaknesses:<br /> • Lack of Rescue Experiments: The absence of rescue experiments makes it difficult to definitively link the observed phenotypes to loss of lite-1.<br /> • Single Loss-of-Function Approach: The reliance on a single genetic mutant limits interpretability. Additional strategies such as RNAi (e.g., neuron-specific knockdown) would provide stronger evidence.<br /> • Unclear Neuronal Contribution: While calcium responses in ADL and ASK are reduced, it's unclear which neuron(s) are necessary for behavioral avoidance. Cell-specific rescue or knockdown experiments are needed.<br /> • Unvalidated Docking Data: The molecular docking predictions lack experimental validation. Site-directed mutagenesis would be needed to support claims of direct interaction.<br /> • Limited Odorant Specificity Testing: Docking analysis does not include non-binding odorants, making it difficult to assess binding specificity.<br /> • Incomplete Quantification: Some calcium imaging results (e.g., in AWA neurons of unc-13 mutants) lack statistical comparisons, which limits their interpretive value.

      Comments on revisions:

      I thank the authors for their thorough revision. The manuscript is substantially improved, and most of the concerns raised in my original review have been addressed.

      The strongest improvement is the addition of new genetic evidence supporting a role for LITE-1 in high-concentration diacetyl avoidance. The use of multiple independent lite-1 alleles strengthens the conclusion that the phenotype is specifically due to loss of lite-1 function, and the ADL-specific rescue experiment is an important addition. While the rescue is not complete, it is convincing and supports the idea that LITE-1 activity in ADL contributes to the avoidance response.

      The neuronal analysis is also stronger. The new statistical analysis across the sensory neurons addresses my previous concerns regarding quantification, and the revised interpretation of the calcium imaging data is more balanced. The revised title of this section is also more consistent with the data and appropriately focuses on ADL and ASK rather than ASH.

      I also appreciate the additional controls addressing possible effects of ambient light. The light versus dark experiments make it unlikely that the observed behavioral phenotypes are secondary to unintended activation of LITE-1 by environmental illumination.

      The expanded odorant analysis and inclusion of docking predictions for additional compounds are useful additions. Together with the ectopic expression experiments in body-wall muscles, these data strengthen the argument that LITE-1 can respond to diacetyl and related compounds.

      My main remaining concern is the same one raised in the initial review: the docking results remain largely computational predictions and have not been tested experimentally through binding-site mutagenesis or other functional validation. As a result, the manuscript still does not demonstrate direct ligand binding to LITE-1. However, I think the authors have strengthened the indirect evidence considerably, and the conclusions are now generally written in an appropriately cautious manner. I would encourage the authors to continue framing the docking results as supportive of a direct interaction rather than definitive proof of one.

      Overall, I believe the manuscript has been significantly strengthened by the revision. The remaining limitation is largely mechanistic and does not, in my opinion, undermine the central conclusions of the study.

    4. Reviewer #3 (Public review):

      In this work, Brown and colleagues report that the photosensor protein LITE-1 of the nematode C. elegans may also a chemosensor that can be activated by high concentrations of the compound diacetly. LITE-1 was described as a putative ion channel of the gustatory receptor family, which is mainly constituted by insect odorant receptors. These form tetrameric ion channels that can be activated by odorant. Specificity is achieved by forming heteromeric channels from three copies of the odorant receptor co-receptor (ORCO) and another subunit that resembles ORCO in the pore-forming C-terminus, but brings in a binding site for the respective odorant. LITE-1 has a very similar structure, according to Alphafold3 predictions, and also carries a binding pocket. In LITE-1, this was proposed to be occupied by a light-absorbing molecule that activates the channel when a photon is absorbed. Alternatively, compounds generated by absorption of high-energy photons may be formed in vivo and bound by the LITE-1 binding pocket. Koh et al. now demonstrate that another, non-light activated compound, diacetyl, at high concentrations, can activate cells expressing LITE-1. Such (chemosensory) cells are also responsible for the avoidance of high concentrations of diacetyl. For this aspect, the protein seems to act in neurons, which are not necessarily the same cells in which LITE-1 is evoking the photophobic response. LITE-1 activation in excitable cells, i.e muscles, causes strong body contraction and paralysis, and the authors show that this is also the case when diacetyl is present. This action is surprisingly rapid, i.e. within 10 seconds after adding diacetyl, raising the question of how the compound can enter the body so quickly. The authors further present molecular docking studies showing that diacetyl could occupy the binding pocket of LITE-1. Last, they show that another compound chemically resembling diacetyl, i.e. 2,3-pentanedione, can also induce avoidance in a LITE-1 dependent manner, though not as potently.

      The data are intriguing and the demonstration of LITE-1 being a diacetyl chemosensor is interesting. Following the first submission and review, the authors addressed most of the questions that this reviewer had and improved the paper significantly. It will add to the further understanding of the still-mysterious multimodal sensory ion channel LITE-1.

      The authors identified mutants lacking diacetyl responses. In their chemotaxis assay (Fig. 1A, B), they show that lite-1 mutants do not avoid high concentrations of diacetyl. However, the animals actually show attraction, as the chemotaxis index was positive. If the lite-1 animals were insensitive, they should be indifferent and the chemotaxis index should be close to zero. The authors now showed that other neurons contribute to the avoidance response that are not themselves bona fide chemosensory neurons, as the avoidance behavior remained in a tax-4 mutant, that is lacking most sensory neuron responses. The authors further showed that diacetyl responses of ADL can be rescued by expressing LITE-1 specifically in this neuron in a lite-1 mutant background, thus demonstrating that LITE-1 acts cell-autonomously in ADL to affect avoidance behavior.

      The effect of diacetyl on muscle cells (Fig. 3C) is pretty rapid. As shown in the initial submission, already during 1 minute after application the animals are almost maximally contracted. The authors now provide a time course with data points every 10 seconds. This shows that contraction is maximal already after 10 seconds of exposure (provided the zero time point is the one where diacetyl is added). This is remarkable, as the compound would have to either pass the worm cuticle, or enter through the gut and diffuse through the body to reach the muscle cells. It would be of interest to compare this to time courses of other pharmacological agents that need to enter the worm's body for their action. As a comparison, often sodium azide is used to paralyze worms for imaging purposes. This molecule is even smaller than diacetyl, but azide action typically requires more time for maximal effects. Could there be active transport mechanisms involved? Maybe diacetyl can pass through some transporters for related molecules into / through intestinal cells quickly, to then reach the body fluid and muscle cells.

      One alternative explanation could be that other mechanisms may be at play. E.g. diacetyl may be immediately sensed by ciliated chemosensory neurons that might release a signaling molecule that leads to activation of LITE-1 in muscles, or that sensitizes it somehow, responding to light used for filming animals. The authors addressed these concerns by repeating their assay in a lite-1 mutant background. The authors had tested unc-13 mutants to rule out indirect effects on the neurons recorded. Likewise, eliminating neuropeptide signaling via unc-31 mutants would have been informative, as neuropeptide signaling plays a role in LITE-1-mediated light avoidance behavior (PMID 40238937, 39489735).

      Molecular docking studies are now described in more detail. The authors also provide a structural model of the diacetyl-docked LITE-1. In this model, only one of the four putative binding sites carries diacetyl. Maybe the authors could test how structure is affected if all four sites contain diacetyl? Mutations of LITE-1 that lack aminoacids shown to be contacted by diacetyl have been described (C300, R222). It would have been insightful to test if such mutants are still activated by diacetyl. This would also verify that there is not an unknown breakdown product of diacetyl that could affect LITE-1 function through oxidative pathways, as diacetyl has been implicated as a photosensitizer in the past.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper describes an interesting phenotype of C. elegans lite-1 mutants. Previous work showed that lite-1 mutants lose a violet/blue light avoidance response. The authors show here that lite-1 mutants also show a defect in negative diacetyl chemotaxis. While wild-type worms avoid diacetyl at high concentrations, lite-1 mutants are instead *attracted* to it. The authors go on to perform Ca2+ imaging in sensory neurons and find that ADL and ASK neurons show altered Ca2+ responses to diacetyl in lite-1 mutants, suggesting LITE-1 is required for these responses. As unc-13 mutants with defective synaptic transmission show similar diacetyl Ca2+ responses as wild-type, this suggests these neurons respond cell autonomously to diacetyl. However, whether lite-1 also acts cell-autonomously is not discussed. Indeed, because unc-13 and lite-1 mutants show different ADL and ASK Ca2+ responses, it seems the diacetyl response regulated by LITE-1 is likely acting outside of those cells. An interesting result that is not commented on is the switching of the valence of the ASK Ca2+ response in lite-1 mutants. ASK neurons still respond to diacetyl, but instead of a strong increase in Ca2+, diacetyl appears to drive it strongly lower. This may be consistent with the switch in valence in the diacetyl chemotaxis assay. It also argues against the idea that LITE-1 is a low-affinity diacetyl receptor that drives avoidance or the Ca2+ responses in ASK, since it is still present in lite-1 mutants. The authors then use a strain that expresses LITE-1 in the body wall muscles and show this expression is sufficient to engender them with sensitivity to diacetyl, as measured through altered swimming and hypercontractility. The authors interpret this result as LITE-1 may act as a diacetyl receptor. The authors test whether a structurally similar molecule, 2,3-pentanedione, shows similar effects, and they find it does. Alpha-fold modeling and molecular docking analysis show where diacetyl might bind to the LITE-1 protein. They then test whether lite-1 mutants show chemotaxis defects to other molecules, as seen with diacetyl. Generally, they find that the observed diacetyl responses are unique, although lite-1 mutants do lose their avoidance response to 2,3-pentanedione. However, unlike the acquisition of diacetyl attraction in lite-1 mutants, 2,3 pentanedione avoidance is *lost*; it is not switched to attraction. Overall, I felt the description of the results and their implications could have been more in-depth. Further, the evidence that LITE-1 is a chemoreceptor itself, rather than acting in some way to shape chemoreceptor responses (via light or otherwise), remains unclear, as conceded by the authors.

      Strengths:

      Overall, the study follows up on an interesting and useful result. The experiments as presented are generally well-conceived and performed. The authors use a variety of behavioral and imaging approaches to test how LITE-1 mediates diacetyl avoidance.

      Weaknesses:

      The study is missing experiments needed to resolve whether LITE-1 is doing what they propose. The evidence that LITE-1 is a diacetyl receptor is lacking support since lite-1 mutants have their avoidance and calcium responses flipped, which would not be expected if it were acting solely as an avoidance receptor. Presumably, the authors are concluding that the attractive response that is left in the lite-1 mutant is mediated by ODR-10, but that experiment is not shown.

      We interpret the shift from avoidance to attraction in lite-1 mutants as consistent with the loss of an aversive sensory component in the presence of an underlying attractive response to diacetyl. We initially hypothesised that this residual attraction was mediated predominantly by ODR-10. To test this, we now generated and analysed lite-1; odr-10 double mutants. The double mutants retained an attractive response to diacetyl, indicating that ODR-10 alone does not account for the attraction observed in the absence of LITE-1 and that additional receptors or sensory pathways are likely to contribute. This finding is consistent with previous studies where loss of ODR-10 did not lead to a complete loss of diacetyl responsiveness.

      Similarly, the authors concede that "the use of lite-1 point mutants that affect specific LITE-1 function, such as light sensing, channel gating, or binding pocket, could further elucidate LITE-1 mechanisms." This reviewer agrees, and such experiments designed to localize diacetyl binding site(s) would be necessary to conclude definitively that LITE-1 is a diacetyl receptor. The body wall muscle assay used or some other heterologous experimental system could work for such a structure-function analysis. A concern is whether the extensive number of LITE-1 point mutants described in the literature affect cell surface expression vs. receptor function, which might complicate the interpretation of a result showing loss of diacetyl responses.

      We agree that structure-function analysis using LITE-1 point mutants could help identify regions or residues that contribute to the diacetyl response and is an important future direction for research, which we have included in the discussion.

      Reviewer #2 (Public review):

      Summary:

      Koh and colleagues investigate the broader sensory role of LITE-1, a gustatory receptor previously linked to UV light detection in C. elegans. Their study explores whether LITE-1 also mediates avoidance of specific chemical stimuli-namely, high concentrations of diacetyl and 2,3-pentanedione. They show that LITE-1 is required in the ADL and ASK neurons for calcium responses to diacetyl, and that its expression in body-wall muscles is sufficient to trigger hypercontraction upon odorant exposure. Molecular docking suggests both odorants may directly bind to LITE-1 with micromolar affinity. These findings suggest LITE-1 may act as a multimodal receptor for both light and chemical stimuli.

      Strengths:

      (1) Methodological Precision: The study is technically strong, with well-executed calcium imaging and quantitative behavioral assays that clearly show neural and muscular responses to chemical stimuli.

      (2) Novelty and Scope: The work presents a compelling case for LITE-1 functioning as a multimodal sensor, which is an intriguing expansion of its known role.

      (3) Potential Impact: If validated, the findings could significantly advance the understanding of sensory integration in C. elegans, and the tools developed may be broadly useful to the research community.

      (4) Relevance to the Field: The study adds to evidence that C. elegans uses non-canonical sensory pathways and may inspire further exploration of multimodal receptor functions in other systems.

      Weaknesses:

      (1) Lack of Rescue Experiments: The absence of rescue experiments makes it difficult to definitively link the observed phenotypes to loss of lite-1.

      We have now performed the rescue experiment expressing lite-1 in ADL, and showed that LITE-1 in ADL is sufficient for avoidance, although it is not a complete rescue to wild-type levels.

      (2) Single Loss-of-Function Approach: The reliance on a single genetic mutant limits interpretability. Additional strategies such as RNAi (e.g., neuron-specific knockdown) would provide stronger evidence.

      We observed the loss of avoidance in three independent lite-1 alleles. Combined with the new cell-specific rescue experiment, we think this provides sufficient support for the conclusion that the phenotype is due to loss of lite-1 function.

      (3) Unclear Neuronal Contribution: While calcium responses in ADL and ASK are reduced, it's unclear which neuron(s) are necessary for behavioral avoidance. Cell-specific rescue or knockdown experiments are needed.

      We have expressed lite-1 genomic DNA under the ADL-specific promoter srh-220, which restored the avoidance phenotype, although it is not a complete rescue of wild-type behaviour. Together with calcium imaging data, this suggests that proper avoidance likely requires input from both ADL and ASK neurons.

      (4) Unvalidated Docking Data: The molecular docking predictions lack experimental validation. Site-directed mutagenesis would be needed to support claims of direct interaction.

      We agree that the docking data does not in itself establish direct binding (we think the muscle expression and paralysis provides stronger evidence). Based on previously reported docking experiments, we wanted to check if diacetyl could occupy the same binding pocket. We have now also included docking data of the other odorants from the chemotaxis assays in the manuscript.

      (5) Limited Odorant Specificity Testing: Docking analysis does not include non-binding odorants, making it difficult to assess binding specificity.

      We agree and have now included docking data of the other odorants from the chemotaxis assays. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      (6) Incomplete Quantification: Some calcium imaging results (e.g., in AWA neurons of unc-13 mutants) lack statistical comparisons, which limits their interpretive value.

      We have generated the scatter plots of calcium imaging responses across the different sensory neurons, and the statistical significance was assessed using two-sided t-tests with FDR correction, which is now included in the manuscript.

      Reviewer #3 (Public review):

      In this work, Brown and colleagues report that the photosensor protein LITE-1 of the nematode C. elegans may also be a chemosensor that can be activated by high concentrations of the compound diacetyl. LITE-1 was described as a putative ion channel of the gustatory receptor family, which is mainly constituted by insect odorant receptors. These form tetrameric ion channels that can be activated by odorants. Specificity is achieved by forming heteromeric channels from three copies of the odorant receptor co-receptor (ORCO) and another subunit that resembles ORCO in the pore-forming C-terminus, but brings in a binding site for the respective odorant. LITE-1 has a very similar structure, according to Alphafold3 predictions, and also carries a binding pocket. In LITE-1, this was proposed to be occupied by a light-absorbing molecule that activates the channel when a photon is absorbed. Alternatively, compounds generated by absorption of high-energy photons may be formed in vivo and bound by the LITE-1 binding pocket. Koh et al. now demonstrate that another, non-light-activated compound, diacetyl, at high concentrations, can activate cells expressing LITE-1. Such (chemosensory) cells are also responsible for the avoidance of high concentrations of diacetyl. LITE-1 activation in excitable cells, i.e, muscles, causes strong body contraction and paralysis, and the authors show that this is also the case when diacetyl is presented. The authors further present molecular docking studies showing that diacetyl could occupy the binding pocket of LITE-1. Last, they show that another compound chemically resembling diacetyl, i.e., 2,3-pentanedione, can also induce avoidance in a LITE-1 dependent manner, though not as potently.

      The data are intriguing, and the demonstration of LITE-1 being a diacetyl chemosensor is interesting. Yet, there are a few questions arising that the authors should address.

      The authors identified mutants lacking diacetyl responses. In their chemotaxis assay (Figures 1A, B), they show that lite-1 mutants do not avoid high concentrations of diacetyl. However, the animals actually showed attraction, as the chemotaxis index was positive. If the lite-1 animals were insensitive, they should be indifferent, and the chemotaxis index should be close to zero. This means, other neurons contribute to the diacetyl response, and the result of these neurons being activated means/remains attraction? If so, the authors need to rule out any effects of these neurons on the effects they attribute to LITE-1 in the other assays.

      We have tested tax-4 mutants in the chemotaxis assay and found that, contrary to the predicted chemotaxis index of zero, these animals retained strong avoidance of high concentrations of diacetyl. This indicates that tax-4 mutants are not chemosensory null for this stimulus and that TAX-4 independent sensory pathways contribute to high diacetyl avoidance. We agree that these experiments cannot completely rule out indirect neuronal effects. We have therefore revised the text to acknowledge this limitation. Nevertheless, the rapid paralysis and contraction observed when LITE-1 is expressed specifically in body-wall muscle in a lite-1 mutant background support the idea that LITE-1 is sufficient to confer a diacetyl-evoked response in these cells.

      The effect of diacetyl on muscle cells (Figure 3C) is pretty rapid, i.e., already during 1 minute after application, the animals are almost maximally contracted. How fast is it really? Can the authors provide a time course with more time points during the first minute? This is a relevant question, as the compound would have to either pass the worm cuticle or enter through the gut and diffuse through the body to reach the muscle cells. Can one expect this to occur within (less than) a minute? In this context, the authors need to rule out that other mechanisms may be at play. E.g., diacetyl may be immediately sensed by ciliated chemosensory neurons that might release a signaling molecule that leads to activation of LITE-1 in muscles, or that sensitizes it somehow, responding to light used for filming animals. The authors should repeat this assay in a lite-1 mutant background.

      We repeated the paralysis assays under red-filtered illumination to minimise potential effects of light, with animals maintained in darkness from hatching to adulthood. We also included lite-1 mutants to assess whether neuronal LITE-1 contributed to the paralysis response. In addition, the assay was repeated with more frequent time points, revealing that paralysis and body contraction occurred within 10 s and neuronal LITE-1 does not contribute to the effect.

      Furthermore, the authors tested unc-13 mutants to rule out indirect effects on the neurons recorded. Likewise, they should eliminate neuropeptide signaling via unc-31 mutants (a recent paper cited by the authors showed involvement of neuropeptide signaling in LITE-1-mediated light avoidance behavior).

      We agreed and have acknowledged and discuss in the manuscript that contributions from gap junction-mediated communication, neuropeptide signalling and other chemosensory pathways cannot be excluded.

      Last, to demonstrate that effects are not indirect in response to chemosensory neurons, the authors should repeat the contraction or swimming assay in a tax-4 mutant, which largely lacks chemosensation. This also applies to the chemotaxis assay. Animals should exhibit a chemotaxis index to diacetyl of zero, then.

      We have tested tax-4 mutants, and like wild-type animals, retained strong avoidance of high concentrations of diacetyl, indicating that TAX-4-independent sensory pathways contribute to this response. This indicates that tax-4 mutants are not chemosensory null for this stimulus and that TAX-4-independent sensory pathways contribute to high diacetyl avoidance. Therefore, repeating the contraction or swimming assay in a tax-4 background would not completely exclude indirect input from other chemosensory neurons. In addition, rapid paralysis and contraction were observed when LITE-1 is expressed specifically in body-wall muscle in a lite-1 mutant background, and together with the calcium imaging and rescue data, they support a role for ADL and ASK in mediating high diacetyl avoidance. The tax-4 chemotaxis data is now included in the manuscript.

      Does diacetyl activate other neurons expressing LITE-1? A number of cells express LITE-1 at high levels, which the authors have not tested (they restricted their analyses to chemosensory neurons). This is important to address because it leaves the possibility that LITE-1 requires a specific partner only present in these chemosensory neurons to detect diacetyl. This partner would have to be present also in muscles, where diacetyl could activate ectopically expressed LITE-1. According to CeNGEN scRNAseq data, cells expressing LITE-1 can be identified. The ADL and ASH neurons actually come up only at the lowest threshold, so some of the other cells showing much higher levels of LITE-1 mRNAs, i.e., AVG, ALM, PLM, ASG, PHA, PHB, AVM, RIF, or some pharyngeal neurons, should be tested. ASG was among the cells the authors recorded from, but this neuron did not show a response.

      We have acknowledged and discuss in the manuscript that other non-sensory neurons may contribute to the avoidance behavioural, and which should be the future direction for investigation.

      The authors need to show that diacetyl responses of ADL and/or ASK can be rescued by expressing LITE-1 specifically in these neurons in a lite-1 mutant background.

      We have expressed lite-1 genomic DNA under the ADL-specific promoter srh-220, which restored the avoidance phenotype, although it is not a complete rescue of wild-type behaviour. Together with calcium imaging data, this suggests that proper avoidance likely requires input from both ADL and ASK neurons.

      Molecular docking studies are not described in detail. How was this done?

      Molecular docking was performed in two stages. First, diacetyl was docked to the tetrameric LITE-1 model using DynamicBind without a predefined binding pocket. The generated complexes were ranked using the DynamicBind confidence score, and the highest ranked poses were used to identify the candidate binding site. The top DynamicBind pose was then used to define the box region for redocking with Gnina. Gnina poses were ranked using the CNN score. A more detailed molecular docking procedure has now been updated in the Methods section.

      Diacetyl is a very small molecule. How well can docking algorithms assess this at all?

      We agree that the small size of diacetyl limits the precision of docking scores because it forms relatively few protein contacts. However, its small size and limited conformational flexibility also simplify pose sampling. To increase robustness, we used two conceptually different docking approaches. First, DynamicBind was used without a predefined binding pocket to identify candidate binding regions while allowing ligand-associated protein conformational adjustments. Second, the resulting pocket was subjected to focused redocking and CNN-based pose ranking with Gnina. The results are interpreted as a structural hypothesis for the probable binding site and relative affinity ranking, rather than as definitive proof of binding or an accurate quantitative affinity measurement.

      Did the authors preselect the binding pocket, or did the algorithm sample the entire molecular surface of the LITE-1 model and end up with the binding pocket?

      The binding pocket was not predefined. Diacetyl was first docked to the tetrameric LITE-1 model using DynamicBind without specifying pocket residues or grid coordinates. DynamicBind therefore performed global, pocket-agnostic docking. The highest-ranked poses identified a candidate pocket, which was then used for focused redocking with Gnina.

      The latter would be very convincing. The authors should provide control docking experiments with other molecules that caused avoidance in their hands (i.e. benzaldehyde, 2,4,5,trimethlythiazole, isoamyl alcohol, nonanone, octanone), but did not activate LITE-1. Also, they should try docking molecules related to diacetyl, and if there are some that do not dock under the same conditions, such molecules should be used in a behavioral experiment. Ideally, they should also not activate LITE-1. Examples could be, e.g., diacetyl monoxime or 2,4-pentanedione.

      We have now included docking data of the other odorants from the chemotaxis assays. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      Last, the authors should provide a PDB file with the docked diacetyl to allow readers to assess the binding for themselves. Since a large number of mutations of LITE-1 have been reported, it may be that amino acids shown to be essential for LITE-1 function are also required for diacetyl binding. If so, this could be backed up with an experiment.

      We agree that structure-function analysis using LITE-1 point mutants could help identify regions or residues that contribute to the diacetyl response and which we have highlighted in the discussion as an important future direction for research. Additionally, we have now provided the PDB file containing a representative DynamicBind derived docking pose of diacetyl within the LITE-1 binding pocket.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) lite-1 mutant animals, as described, fail to avoid' high concentrations of diacetyl, but isn't it more accurate to say that the valence of the response is changed from avoidance to attraction? Do the authors believe that attraction is mediated by ODR-10? Can you build an odr-10; lite-1 double mutant and determine if they lose this attraction to high concentrations of diacetyl and/or 2,3 pentanedione?

      Our initial interpretation was that, in the absence of LITE-1-mediated avoidance, attraction to high concentrations of diacetyl is driven by the low-concentration receptor ODR-10. Interestingly, however, the lite-1; odr-10 double mutants remained strongly attracted to high concentrations of diacetyl, suggesting that this phenotype is independent of ODR-10 and may instead be mediated by other, less specific odorant receptors. This is consistence with the odr-10 mutants not completely losing attraction to low concentration of diacetyl, suggesting the involvement of other potential/putative receptors (Sengupta et al., 1996; Taniguchi et al., 2015). The lite-1; odr-10 double mutant data is now added in the results section (lines: 82 to 90; Figure S1B).

      (2) For ADL, yes, it seems like lite-1 mutants have a reduced diacetyl response, but the ASK response seems... different. While it goes up (slowly) in wild-type, cellular Ca2+ levels in ASK (and maybe ADL) are *reduced* by diacetyl in lite-1 mutants. Can the authors comment on this, and the behavioral responses change in valence?

      ADL and ASK are involved in both attractive and aversive responses so it is possible that the attractive component of the diacetyl response suppresses their activity. In the absence of LITE-1 activation, the observed decrease in calcium responses may reflect the unopposed inhibitory input. The slower decay of the calcium signal in ASK neurons of unc-13 mutants further supports the presence of additional inhibitory signals influencing their activity. This explanation is now included in the results section (lines: 101 to 117).

      (3) ADL and ASK calcium traces in unc-13 mutants look generally similar and lack the effects seen in lite-1 mutants. Does that mean the lite-1 effect is in cells other than ADL or ASK? Can the authors spend more time discussing these differences?

      We acknowledge that LITE-1 is expressed in multiple cell types beyond chemosensory neurons, and that non-chemosensory neurons may also contribute to the observed phenotype. In the previous version, we had highlighted the interneuron AVG as a potential contributor, given its role in light-induced escape. In the revised discussion, we have now expanded this section to include additional possible contributors such as the LITE-1-expressing phasmid neuron PHA and the pharyngeal interneurons I2, both of which have been implicated in hydrogen peroxide sensing (lines: 177 to 181).

      (4) Chemotaxis responses to diacetyl and 2,3-pentanedione in lite-1 mutants are rather different. Diacetyl switches from repulsive (CI < 0) to *attractive* (CI > 0), which is not what would be expected for mutations that eliminate a receptor. In contrast, the 2,3-butanedione responses are more what would be predicted: diacetyl goes from inhibitor to no effect (CI ~0). Again, if the authors feel that this is because of the ODR-10 function, can they discuss whether 2,3-butanedione is predicted to bind ODR-10 like diacetyl?

      We think a switch to attraction is expected for the removal of a receptor for an aversive signal, in an attractive background signal. We did indeed think that this attraction was mediated by odr-10, but the double mutant results now show other receptors must be involved. This is also not wholly unexpected since previous work has identified other receptors of high-concentration diacetyl and the original odr-10 paper didn’t report a complete absence of diacetyl response, suggesting the presence of other receptors that mediate attraction to diacetyl (Sengupta et al., 1996; Taniguchi et al., 2014). The new data is now added in the results section (lines: 82 to 90; Figure S1B; Supplementary video 1 and 2).

      (5) Were the behavior experiments performed in the dark? I realize the calcium imaging experiments and some of the video behavior recordings are not possible in complete darkness, but maybe the authors made efforts to exclude visible light effects (e.g., infrared illumination, etc.) in some assays that might help determine whether light plays *no* role in the effects observed. Alternatively, the authors could try repeating their chemotaxis experiments in the dark or at least communicate in the methods that this was not viewed as a concern (and why). As the authors propose and discuss LITE-1 modulating diacetyl responses via light sensation as a possibility, it is incumbent upon them to communicate the steps they took to overcome this concern for themselves.

      We acknowledge this and have performed a chemotaxis experiment to compare assays performed under dark and ambient light conditions, and no significant differences were observed (results section: lines 75 to 80; Figure S1A, material and methods section: 241 to 246). Therefore, subsequent chemotaxis assays were carried out under ambient light while avoiding exposure to strong illumination.

      Paralysis assays were repeated under red-filtered illumination to minimise light effects, with animals maintained in darkness from hatching to adulthood. Additionally, the assay was expanded to include lite-1 mutants, ruling out contributions from neuronal LITE-1 to paralysis. The new data is now incorporated into the result section (diacetyl; lines: 135 to 137, figures 4 and S3; 2,3-pentanedione; lines: 152 to 154, figures 4 and S6), materials and methods section have been updated to reflect these changes (lines: 250 to 253 and 257 to 260).

      Reviewer #2 (Recommendations for the authors):

      Minor Issues:

      (1) Pmyo-3::LITE-1 worms shrink in the absence of odorants (Figures 3C, 4D); possible effects of ambient light should be discussed.

      We acknowledge the possibility that worms expressing LITE-1 in body-wall muscle experience minor contractions under ambient light, though it is not sufficient to cause paralysis.

      To minimise potential light-induced effects, the paralysis assays were repeated with red-filtered illumination to reduce light stimulation of LITE-1. Animals were maintained in darkness from hatching to adulthood, and in the updated assay, worm length remained relatively constant.

      The new data (Figures 3C, 4E, S4C and S6C) and materials and methods section has been updated accordingly (lines: 250 to 253 and 257 to 260).

      (2) The title is misleading, as ASH does not show altered activity in lite-1 mutants and should be removed from the claim.

      ASH has been removed from the title (line: 97).

      (3) Specific Kd values should be provided for the reported micromolar binding affinities.

      The values from DynamicBind and Gnina are provided in Figure S5.

      Recommendations:

      (1) LITE-1, a member of the gustatory receptor family, was previously shown to mediate UV light responses in C. elegans. In this study, Koh and colleagues demonstrate that LITE-1 is also required for the nematode's avoidance of high concentrations of diacetyl - an odorant that is attractive at low levels but aversive at higher concentrations. Using calcium imaging, the authors show that LITE-1 is necessary in the sensory neurons ADL and ASK for calcium transients in response to high concentrations of diacetyl. Additionally, they find that expressing LITE-1 in body-wall muscles causes hypercontraction upon diacetyl exposure. Similar LITE-1-dependent responses were observed for 2,3-pentanedione, another structurally related odorant. Molecular docking analyses suggest that both diacetyl and 2,3-pentanedione directly bind to LITE-1 with micromolar affinity.

      These findings are intriguing and have the potential to significantly advance our understanding of LITE-1 as a multimodal sensory receptor. However, several major issues need to be addressed to support the authors' conclusions:

      (1) Rescue experiments are missing. The authors should rescue at least one lite-1 mutant to confirm that the observed avoidance defects are specifically due to loss of lite-1.

      Because the avoidance defect was observed in three independent lite-1 alleles, we think background mutations are unlikely to be causal. We have now also performed a rescue experiment with lite-1 expressed in ADL neurons. In this strain, attraction is restored. The new data have been incorporated into the results section (lines: 119 to 121; Fig. 2D).

      (2) Alternative loss-of-function approach. To strengthen the findings, the authors should use a different method to disrupt lite-1 function-such as RNAi by feeding or cell-specific RNAi driven by the lite-1 promoter (see PMID: 17459615).

      We believe the multiple alleles (Fig. 1B) and new cell-specific rescue experiment (Fig. 2D) provide sufficient support for the conclusion that the phenotype is due to loss of lite-1 function.

      (3) Clarify the role of ADL and ASK neurons. While calcium imaging data show reduced activity in these neurons in lite-1 mutants, it remains unclear whether lite-1 is required in ADL, ASK, or both for avoidance behavior. Cell-specific rescue or RNAi experiments, along with additional calcium imaging, are needed to determine the contribution of each neuron.

      We agree that more in-depth work will be required to dissect the neuronal pathways involved in LITE-1-mediated diacetyl avoidance. We have avoided making specific comments on how exactly the observed imaging results relate to the behavioural phenotype. The new rescue experiment expressing lite-1 in ADL does at least show that LITE-1 in ADL is sufficient for avoidance, although it’s not a complete rescue to wild-type levels (lines: 119 to 121; Fig. 2D).

      (4) Validation of molecular docking results. While molecular docking suggests direct binding of odorants to LITE-1, experimental validation is needed. Mutations that reduce predicted binding affinity (engineered in transgenes or via CRISPR) should be tested for functional impact on avoidance behavior.

      We agree the docking does not in itself establish direct binding (we think the muscle expression and paralysis provides much stronger evidence). Based on previously reported docking experiments, we were simply curious whether diacetyl would be predicted to occupy the same binding pocket. We have now updated the discussion in the use of lite-1 mutants to test for impact on diacetyl avoidance (lines: 188 to 191).

      (5) Include analysis of non-binding odorants. Docking results should also be presented for odorants that did not elicit LITE-1-dependent avoidance, to help establish specificity.

      We have now included docking data of odorants from the chemotaxis assays, with the corresponding docking values shown in Supplementary Figure 5. 2-butanone, which is avoided by lite-1 mutants, was predicted to have a slightly higher binding affinity for LITE-1 than 2,3-pentanedione. This highlights the need to interpret the in silico docking data together with real experimental data, rather than using the computational predictions alone to infer functional receptor activation.

      (6) Figure 2C concerns. In neurons such as AWA, calcium transients in unc-13 mutants appear reduced compared to wild-type. A statistical comparison for all the neurons should be included to assess significance.

      Scatter plots of calcium imaging responses across the different sensory neurons were generated, and statistical significance was assessed using two-sided t-tests with FDR correction. Only ADL and ASK neurons showed significant differences between lite-1 mutants and wild-type animals. No significant differences were observed in any neuronal pairs between unc-13 mutants and wild-type, including AWA neurons, although the difference is close to significant (p = 0.07). The scatter plots were now included as Supplementary Figure 3.

      Minor points:

      (1) Figures 3C and 4D: Pmyo-3::LITE-1 worms appear to shrink even without diacetyl or 2,3-pentanedione. Could this be due to ambient light? The authors should discuss this possibility.

      We acknowledge the possibility that worms expressing LITE-1 in body-wall muscle experience minor contractions under ambient light, though it is not sufficient to cause paralysis. We repeated the paralysis assays with red-filtered illumination to reduce light stimulation of LITE-1. Animals were maintained in darkness from hatching to adulthood, to minimise potential light-induced effects, and in the updated assay, worm length remained relatively constant.

      The new data (Figures 3C, 4E, S4C and S6C) and materials and methods section has been updated accordingly (lines: 250 to 253 and 257 to 260).

      (2) Title revision needed: The title "Chemosensory neurons ADL, ASK, and ASH are involved in avoidance of diacetyl" is misleading, as calcium transients in ASH appear unaffected in lite-1 mutants. The title should reflect the actual data.

      ASH have been removed from the title (line: 97).

      (3) Binding affinity clarification: The authors report micromolar binding affinity for LITE-1 but should provide specific dissociation constants for clarity and completeness.

      The values from DynamicBind and Gnina are now provided in Supplementary Figure 5.

      Reviewer #3 (Recommendations for the authors):

      How did the authors measure body length if the animals were swimming in the diacetyl solution? Standard 6-well plates have an area of roughly 10 cm², meaning that if one adds 1 ml, the liquid level should be 1 mm. The animals would be able to move in 3D, so it is likely that animals swim up and down and do not move in one flat plane, i.e., head and tail would be out of focus, and only a projection image would be recorded that would lead to an underestimation of actual worm length.

      We acknowledge that this issue may led to an underestimation of worm length in the previous assay. However, every effort was made to exclude worms that moved out of focus. The paralysis assay has since been modified to include spreading a thin layer of solution across the worms, which keeps them mostly in focus, particularly those that are paralysed. Worms that were partially out of focus were excluded from the analysis.

      We have incorporated the new data into the results section, reflected in the updated Figures 3C, 4E, S4C, and S6C. Corresponding revisions have also been made in the materials and methods (lines: 250 to 253 and 257 to 260).

      Could diacetyl be a compound that results from UV absorption in cells? This may be worth discussing. What could be the precursor molecule?

      We are not aware of such a precursor, but we cannot rule it out. Even if UV absorption in cells leads to diacetyl production, it is likely that the resulting diacetyl levels are insufficient to activate LITE-1, as our data suggest that LITE-1 functions as a receptor for high concentrations of diacetyl. We have updated the discussion accordingly (lines: 169 to 173).

      In lines 57-61, the references to Edwards 2008 and Ward 2008 do not seem to fit the statements made in this sentence.

      The inclusion of Edwards et al., 2008 was an error, and it has now been removed. Ward et al., 2008 demonstrated that ASJ phototransduction requires cGMP and CNG channels, stating that “Our studies indicate that C. elegans photoreceptor cells also employ CNG channels and the second messenger cGMP for phototransduction.”

    1. eLife Assessment

      This study provides a valuable contribution to our understanding of the neural basis of perceptual decision-making by jointly modeling behavioral outcomes and EEG signals in a contrast comparison task. The methods and analyses are solid, systematically comparing standard models assuming continuous evidence accumulation with models that track evidence without temporal integration (extrema detection). The authors show that behavior and neural signals are equally consistent with both alternatives, highlighting limitations in current modeling approaches and questioning the generality of evidence accumulation mechanisms.

    2. Reviewer #1 (Public review):

      Summary:

      This paper examines whether humans use protracted temporal integration in a noise-free, deferred-response contrast discrimination task, using a covert evidence-duration manipulation combined with EEG (SSVEP, CPP, Mu/Beta). The key finding is that evidence for protracted sampling is behaviorally and neurally supported, but even joint CPP + behaviour fitting cannot fully discriminate a standard integration (DDM) model from a novel "extremum-flagging" non-integration model. The paper is transparent about this outcome.

      Strengths:

      This is a well-conducted and well-written study that makes a genuine contribution to the perceptual decision-making literature by introducing a clean experimental design for probing temporal integration without participants adapting their strategy and demonstrating for the first time that a non-integration model (extremum-flagging) can replicate CPP waveform dynamics that have long been considered hallmarks of evidence accumulation. The transparent treatment of equivocal modelling outcomes is commendable.

      Weaknesses:

      My main concerns relate to statistical power, the under-specification of the and the extremum-flagging mechanism. Addressing these would greatly strengthen the paper.

      (1) The sample of 16 participants (15, after the exclusion of one participant) is described as "close to similar EEG studies" with no formal power analysis. Given that the paper's core claim rests on subtle quantitative differences between two model classes - differences that are, by the authors' own admission, not sufficient to declare a winner - even a modest increase in sample size might yield a more decisive outcome. At minimum, the authors should report a sensitivity analysis or post-hoc power calculation to indicate what effect sizes the current N could reliably detect, particularly for the rmANOVA comparisons and the neural constraint fitting.

      (2) The Extremum-flagging model is the paper's most novel contribution, yet its physiological basis is underspecified. The model posits that each decision-terminating bound-crossing triggers a stereotyped, half-sine-shaped centroparietal signal, but no neural circuit or computational mechanism is proposed for how the brain could detect the first bound-crossing event in a non-accumulating evidence stream or generate a temporally precise, fixed-amplitude signal in response. Possible connections to P3b theories of context updating and response facilitation are acknowledged, but these are vague functional descriptions rather than mechanistic accounts. I think the discussion should engage more directly with potential neural substrates that could generate this flagging signal, and whether these are consistent with the known generators of the CPP/P3b. Without this, the extremum-flagging model risks being viewed as a mathematical convenience rather than a biologically plausible alternative.

      (3) The Integration model at the preferred neural weighting estimates a high-to-low contrast drift rate ratio of 8.7, whereas the empirical Mu/Beta lateralization slopes suggest a ratio of approximately 3.5. The authors attribute this discrepancy to the nonlinear contrast response function of early visual cortex and the salience of the high-contrast evidence onset, but these explanations are speculative. These outcomes are arguably the most quantitatively damaging result for the integration model, so they deserves more than a brief discussion. I would recommend that the authors (a) estimate what range of contrast response nonlinearities would be required to close this gap, (b) test whether an alternative drift rate parameterization (e.g., scaling drift rates directly by SSVEP amplitude rather than contrast) reduces the discrepancy, or (c) be more explicit about treating this as a point against the Integration account.

      (4) The sensitivity analysis over neural constraint weightings (w = 0.1 to 1000) is thoughtful, but the paper ultimately acknowledges that the preferred weighting is w=10, chosen because it achieves "a good fit to CPP dynamics without substantively sacrificing behavioral fit" - a qualitative criterion. No principled statistical framework is used to select the optimal weighting or to compare models at a given weighting. A Bayesian model comparison could provide a more formal framework for combining behavioral and neural fit components, and would allow a clearer statement about the relative posterior probability of each model.

      Comments on revisions:

      In reply to my comments, the authors have added a post-hoc power analysis that provides adequate justification for the sample size, a more nuanced discussion of neural mechanisms that could support an extremum flagging model, and several supplementary analyses that show the generality of findings across parameter levels. These are welcome additions that strengthen confidence in key findings while acknowledging nuances involved in quantitative analyses and model fitting.

    3. Reviewer #2 (Public review):

      The manuscript by Hajimohammadi, Mohr, O'Connell and Kelly is intended to demonstrate that participants integrate evidence over time to make a decision, even in a noise-free, static decision context. This is validated by the observation that 1) participant accuracy improves with increased exposure to the stimulus; and 2) there is a correlation between participant accuracy and a neural index of evidence accumulation, as measured by centro-parietal positivity (CPP).

      Strengths:

      (1) Joint modelling of accuracy and CPP dynamics is a significant achievement, as behaviour alone often cannot distinguish between competing theories of decision-making. In the case of protracted sampling in particular, the absence of reaction times (RT) due to the delayed nature of the response makes this method highly appealing.

      (2) The experimental manipulations and the method used to extract the different neural indices are well chosen, enabling the mapping of putative cognitive processes such as evidence accumulation and motor preparation onto the recorded EEG with clarity.

      (3) The in-depth discussion of the results clearly articulates those reported by the authors and in previous works.

      Weaknesses:

      (1) Regarding the first point I raised in the first version of the manuscript, I'm satisfied with the author's response.

      (2) For the second point, however, I'm still unconvinced, noting that my understanding of the fitting routine is relatively limited. To me, the comparison of the behavioral and the joint-modelling is relatively weak in terms of evidence. Despite the revision, it is still unclear whether the behavioral models are well conditioned given the very weak constraint exerted by (aggregated, see below) binary decision data only. To my previous comment, the authors reply, "We can instead address identifiability somewhat indirectly through parameter estimate consistency across the 10 fits we conducted with different instantiations of noise". I'm worried that the different instantiations of noise (limited to 10), generated for the parameter search, have any impact at all. A better quantification of the uncertainty in parameter estimates, as well as the AIC, is needed, for example through a cross-validation scheme, or bootstrapping/jack-knifing at the level of the participants. Without this, I remain unconvinced that the behavioral models are suited for the comparison with the joint-neural models:

      (3) Relatedly, regarding the response of the authors to my minor comment 4, after the revision of the manuscript, it is clearer that the behavioral models were fitted only on 6 data points (accuracies averaged over participants for each condition) for models with up to 3 parameters. I understand that the authors average behavior to make a comparison with the joint neural model, whose neural signals are noisy at the participant level. However, if the neural model, or the chosen linking function, needs this aggregation to outperform the behavioral model, that questions a bit whether the joint model is really useful in practice.

      While I see these two last points as serious for the comparison between behavioural and joint-neural models, this does not change the main addition of the paper: that is, the joint model and the evidence for protracted sampling.

    4. Reviewer #3 (Public review):

      Summary:

      The authors aim to compare proposal models of perceptual decision making using a joint modeling approach, where they fit models to both behavioral outcomes as well as CPP. Most notably, they compare a standard evidence accumulation model with models that track the evidence without integrating it over time (extrema detection). The authors report that the joint CPP-behavioral data do not discriminate between two of their proposals.

      Strengths:

      This is an interesting finding that reinforces the idea that what we believe to see based on aggregation over trials may not be what happens on every single trial. The models are creative and the simulations are convincing, relating the models to multiple neural markers of decision formation. These include the CPP but also mu/beta power spectra.

      Initial weaknesses:

      Contrary to the original draft, the current version now clarifies the goals of the study as well as the role of internal vs external noise.

      The authors clarify their stance on the optimality of extrema detection in the Public Review, where they make some good points. Specifically, they argue that models that have studied optimal decision making in the past have considered a quite narrow definition of optimality, for example, ignoring energetic costs.

      The authors clarified the fixed value of the scaling parameter (which apparently was already mentioned in the initial draft, which I had missed). My comment about the discrepancy between the bounded integration model and the EZ diffusion model can thus be ignored. However, I do think that the similarity of the bounded integration model and the EZ diffusion model could be acknowledged in the manuscript somewhere.

      I still think the choice of modeling the grand average accuracy and EEG signal may distort the results. As Figure 1 reveals, there is in fact substantial variation in the change in accuracy across conditions. I can imagine that one model may be preferred over another based on the aggregate data, while the other model is preferred for some participants. On the group level, this may lead to different conclusions, at least quantitatively (e.g., in terms of G2), but potentially even qualitatively. This is especially true for the neurally-informed models because of the covariation between behavioral and neural data.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper examines whether humans use protracted temporal integration in a noise-free, deferred-response contrast discrimination task, using a covert evidence-duration manipulation combined with EEG (SSVEP, CPP, Mu/Beta). The key finding is that evidence for protracted sampling is behaviorally and neurally supported, but even joint CPP + behaviour fitting cannot fully discriminate a standard integration (DDM) model from a novel "extremum-flagging" non-integration model. The paper is transparent about this outcome.

      Strengths:

      This is a well-conducted and well-written study that makes a genuine contribution to the perceptual decision-making literature by introducing a clean experimental design for probing temporal integration without participants adapting their strategy and demonstrating for the first time that a non-integration model (extremum-flagging) can replicate CPP waveform dynamics that have long been considered hallmarks of evidence accumulation. The transparent treatment of equivocal modelling outcomes is commendable.

      Weaknesses:

      My main concerns relate to statistical power, the under-specification of the and the extremum-flagging mechanism. Addressing these would greatly strengthen the paper.

      (1) The sample of 16 participants (15, after the exclusion of one participant) is described as "close to similar EEG studies" with no formal power analysis. Given that the paper's core claim rests on subtle quantitative differences between two model classes - differences that are, by the authors' own admission, not sufficient to declare a winner - even a modest increase in sample size might yield a more decisive outcome. At a minimum, the authors should report a sensitivity analysis or post-hoc power calculation to indicate what effect sizes the current N could reliably detect, particularly for the rmANOVA comparisons and the neural constraint fitting.

      We appreciate the reviewer’s concern regarding sample size and statistical sensitivity. To address statistical robustness throughout the paper, we have now reported effect sizes for our statistical tests (e.g. η2 for rmANOVA; including Tables S1 and S2), and we provide error-shading around the ERP waveforms to indicate the reliability of the key patterns our models are aimed at capturing (i.e. the dramatically higher and earlier CPP peak for high-contrast, and very little systematic differences across the four low-contrast durations - see revised Figure 3). We also conducted an indicative post-hoc power analysis using the G*Power software based on the behavioural data. Using the observed partial η2 = 0.44 for the test of duration effect on accuracy among only the low-contrast conditions, and the final sample of 15 participants, this amounts to a statistical power of 0.998.

      On the model comparison, while we agree that larger sample sizes are generally beneficial for population-level inferences, we respectfully maintain that our current sample size is sufficient to support the core claim that qualitative dynamics of neural signatures of decision formation, usually assumed to reflect temporal integration, can be successfully reproduced using non-integration models in the delayed-response task conditions we examine here. Any marginal changes in quantitative fit resulting from having a higher N contribute to grand averages are unlikely to substantively alter this conclusion of the model comparison. The statistical reliability of the data to which our models are fitted is also bolstered by the number of trials (about 256 per condition per participant). It is common in behavioural modelling studies for data to be collected from a much smaller sample (e.g., fewer than 10 subjects) but with a high trial yield - a relevant precedent for us being Stine et al., (2020), who provided a compelling demonstration of similar model fits for Extrema and Integration models using only 6 subjects. In sum, the key qualitative data patterns of accuracy improvements with duration and broader, lower and duration-invariant low-contrast CPPs are statistically robust and provide a strong basis to reveal the fundamental principle that the Extremum-flagging and Integration models are both able to produce these key qualitative dynamics.

      (2) The Extremum-flagging model is the paper's most novel contribution, yet its physiological basis is underspecified. The model posits that each decision-terminating bound-crossing triggers a stereotyped, half-sine-shaped centroparietal signal, but no neural circuit or computational mechanism is proposed for how the brain could detect the first bound-crossing event in a non-accumulating evidence stream or generate a temporally precise, fixed-amplitude signal in response. Possible connections to P3b theories of context updating and response facilitation are acknowledged, but these are vague functional descriptions rather than mechanistic accounts. I think the discussion should engage more directly with potential neural substrates that could generate this flagging signal, and whether these are consistent with the known generators of the CPP/P3b. Without this, the extremum-flagging model risks being viewed as a mathematical convenience rather than a biologically plausible alternative.

      We thank the reviewer for this constructive comment. While the focus of this paper was indeed on simple mathematical descriptions in the spirit of classical cognitive modelling, we agree that expanding on the potential neural substrates of the Extremum-flagging model strengthens its utility as an alternative framework. We have revised the Discussion to engage with potential biological mechanisms, particularly those previously proposed to underlie the P300/P3b, such as Nieuwenhuis’ (2005) proposal that it reflects a phasic arousal response mediated by the LC/NE system that serves to activate task-relevant areas following completion of a decision. We agree that aside from this, accounts of ERP component functions over the years have often been vague and non-mechanistic, but the idea that they reflect discrete neural activations marking an internal cognitive event in a stereotyped way persists, and remains a basic assumption of several new and influential ERP signal analysis toolboxes (e.g. Ehinger 2019; Weindel 2024). If the flagging signal’s fixed amplitude seems physiologically implausible, all-or-nothing neural activation events are not generally unheard of in neurophysiology, and, again, we are taking an approach favouring parsimony in the spirit of cognitive modelling, and we found that we did not need to assume any variation in the amplitude of the flagging signal in order to capture the key decision signal dynamics alongside behavioural accuracies in this particular case.

      We also discuss the study of Latimer et al. (2015), who demonstrated that discrete, step-function state transitions that on single trials may mark extrema detection events, can produce ramp-like signals when trial-averaged. While the biological plausibility of such step-function dynamics remains a subject of debate, it serves as another example of how continuous evidence integration is not the only way to reproduce the ramping neural signals traditionally observed in grand-average neural signals.

      (3) The Integration model at the preferred neural weighting estimates a high-to-low contrast drift rate ratio of 8.7, whereas the empirical Mu/Beta lateralization slopes suggest a ratio of approximately 3.5. The authors attribute this discrepancy to the nonlinear contrast response function of early visual cortex and the salience of the high-contrast evidence onset, but these explanations are speculative. These outcomes are arguably the most quantitatively damaging result for the integration model, so they deserve more than a brief discussion. I would recommend that the authors (a) estimate what range of contrast response nonlinearities would be required to close this gap, (b) test whether an alternative drift rate parameterization (e.g., scaling drift rates directly by SSVEP amplitude rather than contrast) reduces the discrepancy, or (c) be more explicit about treating this as a point against the Integration account.

      We agree that the quantitative discrepancy we demonstrated between the empirically observed buildup rate ratio in motor preparation signals (3.5) and the greater drift-rate ratio (8.7) required by the Integration model to fit the CPP waveforms is an important one that should be emphasised and discussed with greater depth and clarity. As we said, nonlinear contrast response functions and a boosting effect of the salient high-contrast step-change are two plausible ways that a drift rate might scale disproportionately more steeply with contrast, but in principle, assuming straightforward transmission of evidence accumulation to the motor level, Mu/Beta lateralization slopes should then reflect this steeper drift rate scaling, or at least approach it even when allowing for some temporal blurring. We have thus put more emphasis on the discrepancy by confirming that if we constrain the drift rates to be directly proportional to contrast, the Integration model is indeed significantly hampered in its ability to produce the much steeper CPP buildup for higher-contrast trials, much more so than the Extremum-flagging model (Figure 4 - Supplement 7). We have also applied a temporal blurring equivalent to the short-time Fourier Transform to the simulated motor preparation waveforms (convolving with a boxcar of the same duration as the Fourier window) in Figure 4N-P so that the real and simulated traces are on an equal footing in this respect. We have also revised the Discussion to elaborate on how this quantitative discrepancy represents a point against the Integration account, and possible ways it might be reconciled with an Integration account. One reason, for example, why the relative steepness of the Centroparietal ERP in the high-contrast condition so far exceeds that of Mu/Beta might be that additional processes are evoked by the very salient step-change, which may make a positive-polarity contribution to the centroparietal ERP waveform and hence cause overestimation of how early and steeply the underlying, high-contrast CPP decision signal rises. We looked into this by carefully examining time courses and topographies through the initial period of buildup, with no additional smoothing low-pass filter applied, now presented in Figure 3 - Supplementary Figure 1. While the smoothed waveforms that we show in the main paper and to which we fit models could be seen to have a brief inflection during the main buildup for the high-contrast condition, removing the smoothing shows that this arises not from random noise but from a distinct bimodal morphology, with a distinct early peak and lull during the buildup, which temporally coincides with a very strong bilateral occipital N2 (associated with a low-level evidence-onset detection or ‘target selection’ process - Loughnane et al 2016), in a way that suggests that the positive tail-end of the dipolar neural generators of the N2 may contribute to the initial part of the positive centro-parietal buildup. It is difficult to estimate the extent to which the neurally-constrained model estimate of high-contrast drift rate is inflated by this initial overlapping potential, because we can’t precisely know the ground truth of the N2 tail’s contribution, but this analysis provides a potential explanation that can be explored in future (e.g. through softening the strong-evidence onset with a ramp or use of auditory evidence). We thank the reviewer for raising this as we feel that this extra discussion positively adds to the theme of the paper to highlight methodological challenges with neurally-constrained modelling. In the process, we have updated the methods section to present in full detail the centroparietal electrode selection and waveform smoothing that was applied to provide the models with a relatively uninterrupted buildup signal to capture, which is important for readers to appraise the potential impact of this overlapping potential.

      (4) The sensitivity analysis over neural constraint weightings (w = 0.1 to 1000) is thoughtful, but the paper ultimately acknowledges that the preferred weighting is w=10, chosen because it achieves "a good fit to CPP dynamics without substantively sacrificing behavioral fit" - a qualitative criterion. No principled statistical framework is used to select the optimal weighting or to compare models at a given weighting. A Bayesian model comparison could provide a more formal framework for combining behavioral and neural fit components, and would allow a clearer statement about the relative posterior probability of each model.

      We agree with the reviewer that theoretically, the Bayesian framework provides a principled way to combine behavioural and neural evidence by weighting each source according to its statistical reliability. However, a Bayesian formulation typically quantifies reliability through across-trial variance, which applies quite differently for accuracy and EEG data. While the precision of EEG measurements can be estimated empirically (e.g., from noise characteristics), we currently lack a formal measure of uncertainty for the linking function itself, that is, the theoretical mapping between neural signatures and latent decision processes. This represents an unresolved methodological issue rather than a straightforward parameter estimation problem. Second, although hierarchical Bayesian approaches are well established for standard diffusion models, the mechanisms examined here for extremum flagging do not currently have tractable closed-form formulations suitable for Bayesian integration. Developing a dedicated hierarchical Bayesian framework for these non-standard mechanisms would require substantial methodological work and is beyond the scope of this research.

      Thus, rather than imposing a single assumed reliability relationship between neural and behavioural data, we chose to perform a systematic sweep across weighting values. We view this approach as a transparent sensitivity analysis that accommodates different scientific priors regarding the relative contribution of neural versus behavioural constraints. By presenting the full range of w (including in supplemental tables and figures), readers can directly evaluate how model behaviour changes when emphasis is shifted between behavioural data and neural data, transparently revealing how the behavioural and neural signal fits can trade against one another.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Hajimohammadi, Mohr, O'Connell and Kelly is intended to demonstrate that participants integrate evidence over time to make a decision, even in a noise-free, static decision context. This is validated by the observation that (1) participant accuracy improves with increased exposure to the stimulus; and (2) there is a correlation between participant accuracy and a neural index of evidence accumulation, as measured by centro-parietal positivity (CPP).

      Strengths:

      (1) Joint modelling of accuracy and CPP dynamics is a significant achievement, as behaviour alone often cannot distinguish between competing theories of decision-making. In the case of protracted sampling in particular, the absence of reaction times (RT) due to the delayed nature of the response makes this method highly appealing.

      (2) The experimental manipulations and the method used to extract the different neural indices are well chosen, enabling the mapping of putative cognitive processes such as evidence accumulation and motor preparation onto the recorded EEG with clarity.

      (3) The in-depth discussion of the results clearly articulates those reported by the authors and in previous works.

      Weaknesses:

      (1) One main issue to support the interpretation of the authors toward the need for protracted sampling is the timing of the evidence. By design, participants believe that the signal is present for 1.6 seconds (reinforced by the fact that easy trials were displayed for 1.6 seconds). However, the difference in stimuli is turned off either 1.4, 1.2, 0.8 or 0 seconds before the cue to respond. While this makes sense in the context of the authors' question, it also raises the possibility that participants will focus on the last samples before answering. Even if participants apply equal weighting, this still favours them delaying evidence accumulation until they are sufficiently certain that the evidence should be present (e.g. participants might start accumulating after the stimulus has disappeared in the 0.2 condition). I do not see an easy way to test these alternative explanations outside of running a study in which the evidence is always offset before the go cue.

      This is a reasonable question about the design - if participants were under the impression that they had a whole 1.6 sec of stimulation, couldn’t they afford to wait until later into the stimulus to start sampling? However, the task was designed to be so difficult that participants would be deterred from ignoring any initial evidence, and the fixed and explicitly instructed lead-in period as well as the interleaved easy trials, would have continually reinforced their ability to time their sampling onset quite precisely. Indeed, key aspects of the data confirm they did not appreciably delay sampling. First, accuracy in even the shortest (0.2 s) condition was reliably above chance (t(15) = 2.60, p = 0.0201) and improved steadily across evidence durations (Figure 1B). This places an upper bound on the accumulation onset: participants cannot have delayed accumulation until after the evidence disappeared and still achieve above-chance performance; if they only used the ‘last samples,’ at the end of the stimulus, they would have performed at chance level for all durations except 1.6 sec. Second, we fit a model that allowed for such a delayed sampling onset, captured in the parameter ‘sampT,’ which, across the range of neural weightings (Tables S3, S5-8), consistently landed within a few tens of msec of evidence onset (often slightly before rather than delayed), and improved the overall model fit very little relative to the addition of starting point variabilities or collapsing bound. The Methods section now addresses these aspects of task design.

      (2) Regarding the behavioural models, are these identifiable based on accuracy data alone? This should be addressed using a parameter recovery study, in which a set of parameters is used to generate data, and the same fitting routine used for the real data is used to estimate the parameters. This would enable us to determine what can be inferred from the model comparison presented. This is not a serious problem for the manuscript, as it specifically aims to go beyond behaviour. It is, however, worth noting that such a parameter recovery addition could be used to demonstrate the need for a joint modelling framework to answer the question of protracted sampling on delayed response times (RT).

      As the reviewer notes, we did have the specific aim of going beyond behaviour, and the need to do so is demonstrated in the inability to adjudicate between the alternative models based on behaviour alone. We took this as sufficient justification without a formal parameter recovery test to assess the degree to which behaviour-only models could accurately estimate parameter values. Still, we agree that it is valuable to address parameter identifiability in some way. Since a full parameter recovery covering the full possible parameter space for each of the many models would be too great in volume to add to this paper, we can instead address identifiability somewhat indirectly through parameter estimate consistency across the 10 fits we conducted with different instantiations of noise; we now provide the standard deviations alongside the mean parameter values for the D1, D2 and B parameters of each of the behaviour-only models in Table 1 - Table Supplement 1, which indicates that the parameter estimates were reliable across 10 different instantiations. 

      Minor comments:

      (1) I would advise authors to fix the D1 parameter and use it as a scaling parameter across all models. Currently, as I understand it, the models are scale-free, meaning the same fit is achieved by multiplying all parameters by two, for example. This makes the fit more complex (bounds on parameter values are required) and means that the models are less comparable in terms of their estimates. Perhaps I'm missing something, but I would have thought that fixing D1 (the common parameter across all models) would solve these issues.

      The models are not scale-free because they are constrained relative to a fixed sampling noise parameter value of s = 0.1; All tables in the main text have now been updated to make this more immediately clear. Aside from this being standard in diffusion modelling (Ratcliff & Smith, 2004), this enabled us to replicate the observation by Stine et al., (2020) that since the non-integration models depend on the magnitude of individual evidence samples rather than an integration of many, the drift rate values must be set much higher to achieve the same choice accuracy as the integration models (Table 1).

      (2) Why is the snapshot model so bad despite being a good model in Stine et al 2020? Can the authors speculate in the discussion?

      We thank the reviewer for querying this. We had originally thought that the poor performance of the snapshot model made sense because the continued presentation of zero contrast difference for short-evidence trials renders it a bad strategy. Because our main purpose was to briefly substantiate the principle that accuracies alone are an insufficient basis for model comparison and move on to the main goal of jointly modelling accuracies and CPP dynamics, we did not take the same level of care to ensure we attained the very best fit of the behaviour-only models, as we did for the neurally-constrained models. In the neurally-constrained modeling, we took care to check for every parameter whether the range of allowed values (Table S4) was narrow enough to avoid the optimisation algorithm getting lost in untenable parts of parameter space, yet wide enough to include the optimum point, and wherever we saw parameter values landing at or near the edge of the allowed range we expanded that range and re-ran the model fit. Applying these same checks to the behaviour-only fitting, we found that the SnapShot model needed a wider range on drift rate and when we applied this, the fit was much more competitive, in line with Stine et al., (2020), though it remained the worst-fitting model among all two-drift-rate behaviour-only models (see updated Table 1). We similarly conducted these checks across all behaviour-only models and re-ran them. The extrema detection model with last-sample default when no bound is hit also improved its fit, though again it did not fit better than the version with guess default. Thus, the point we were making with this section, that behaviour alone can be captured competitively by a range of integration and non-integration models, is bolstered by the updated model fits. Since the last-sample default was competitive in the behaviour-only fits, we also ran a version of the Extremum-flagging model jointly fit to accuracies and CPP dynamics with a last-sample rather than random guess default when a bound was not reached, and show in new Figure 4 - Figure Supplement 8 that the conclusions are the same. Again, thank you for prompting us to look back at those fits.

      (3) The meaning of the flag width is unclear. Figure 4 provides the reader with an intuitive understanding of the model that the authors have in mind. However, the tables in the appendices report values between 0.2 and 0.9. I understand that these values represent the width of the half-sine in seconds. This suggests that the actual estimated values for these flag events are much broader than those displayed in Figure 4. While this is probably fine for most models, it can be problematic for the extremum-flagging model, as it means that the rise to the peak takes between 0.1 and 0.45 seconds. While strictly speaking, this is still a 'flag' model, such a slow rise to the peak, given the usual expectation of evidence accumulation, would place this model closer to a smooth integration model than to a boundary-crossing flagging mechanism.

      We thank the reviewer for raising this about the flag width parameter. In so doing, they enabled us to catch that our schematic depiction of the model in Figure 4 was misleading, and have now revised it to make clear that the flag signal is a post-decision one triggered by the bound crossing, and we have updated explanations accordingly (in ‘Neurally-constrained models’ and Discussion). The reviewer is correct that the reported values in the supplemental materials (approximately 0.2–0.9 s) correspond to the width of the half-sine kernel used to model the post-decision flag event. However, the flag signal is stereotyped, evidence-independent, and is triggered once the decision threshold has already been crossed, so it does not share the key characteristics of evidence integration, regardless of how wide the model estimates it. In the extremum-flagging model, the boundary crossing remains a discrete event. The width parameter instead captures the temporal extent of the neural process that follows this commitment event. Such a post-decision neural process unfolding over several hundred milliseconds is in line with some classic theories of the centroparietal P300/P3b component, and we now expand our discussion point on this to address proposed neural substrates (e.g. Nieuwenhuis et al’s (2005) implication of a phasic noradrenaline system response).

      (4) In the modelling section, it is not clear overall (i.e. for G<sup>2</sup> and R<sup>2</sup>) how the participant dimension is taken into account. Are these individually fitted models, and if so, how are the secondary statistics generated from the individual estimates? Or were these fitted over all participants?

      All models were fitted to the grand-average neural and behavioural data across participants, rather than to individual participant data. We chose this approach as the CPP signal at the individual level is highly noisy, which can introduce substantial instability and noise into the model fitting procedure. We have revised the Modelling section to explicitly state that the reported G<sup>2</sup> and R<sup>2</sup> values are derived from models fitted to the grand-average data, and in the revised discussion acknowledged this as a limitation of the current modelling framework.

      (5) On page 7, in the last sentence of the first paragraph of the section titled 'Decision-Related Neural Signals', the authors state that 'this stable contrast-difference encoding suggests that a constant (i.e. non-adapting) drift rate is a reasonable simplifying model assumption'. However, I am not sure how this is true given that SSVEP quantifies encoding, yet the drift rate can vary due to non-sensory aspects (e.g. attention).

      The reviewer makes a good point - even if sensory encoding is stable, non-sensory factors like attention could cause dynamic changes in the effective drift rate independently of the sensory representation itself. However, our point in that section, which we have revised to put more clearly, was to test for one particular well-known time-varying effect that could impact drift rate, namely sensory adaptation, a well-established phenomenon behaviorally and at the level of sensory neuronal responses, where prolonged stimulation produces reductions over time in sensory neural activity. If strong adaptation were present in the sensory evidence representation indexed by the SSVEP, we would expect corresponding temporal changes in the signal. The absence of such changes lends support to the simplifying assumption (as in most accumulation models) that the drift rate is approximately stationary over time, even if we cannot be sure there isn’t a time-varying effect downstream.

      (6) The mu/beta lateralisation does indeed favor the integration model more, but in terms of boundary estimation and starting-point analyses, both models are pretty far apart. Providing an interpretation of this observation, e.g. regarding alternative linking functions for mu/beta, would add to the manuscript.

      In response to this comment, we revised the manuscript in the Discussion to say that in the current analyses, we implicitly assume an approximately linear mapping between Mu/Beta amplitude and decision units. However, the true relationship may instead reflect another monotonic transformation (e.g., involving power rather than amplitude, logarithmic scaling such as dB units, or a nonlinear saturating function). This uncertainty could affect the apparent correspondence between the neural signal and the model-derived estimates of boundary position or urgency dynamics. While our analyses support a close relationship between Mu/Beta lateralisation and the evolving decision process, the precise quantitative mapping remains uncertain. One possibility is that urgency itself evolves nonlinearly (e.g., decelerating over time), even if the measured neural trajectory appears approximately linear under the current transformation assumptions.

      Reviewer #3 (Public review):

      Summary:

      The authors aim to compare proposal models of perceptual decision making using a joint modeling approach, where they fit models to both behavioral outcomes as well as CPP. Most notably, they compare a standard evidence accumulation model with models that track the evidence without integrating it over time (extrema detection). The authors report that the joint CPP-behavioral data do not discriminate between two of their proposals.

      Strengths:

      This is an interesting finding that reinforces the idea that what we believe to see based on aggregation over trials may not be what happens on every single trial. The models are creative, and the simulations are convincing, relating the models to multiple neural markers of decision formation. These include the CPP but also mu/beta power spectra.

      Weaknesses:

      The paper makes some strong points, and the work seems generally well-executed. The weaknesses that I identified are twofold:

      (1) Embedding in the literature/exposition of the main argument.

      The focus in the introduction is on the noise-free nature of the stimulus and the prolonged presentation time. However, after reading the paper, I felt these were mostly experimental design choices that enable comparison of the different models using the CPP. Perhaps my misreading of the goals of the paper stems from two other observations:

      (a) The fact that the stimulus is noise-free does not entail that perception is noise-free. Thus, the argument that using a noise-free stimulus precludes the necessity of temporal integration seems not completely valid. Of course, one could argue that noise is limited in this case, but that makes a noise-free stimulus more of a design choice.

      (b) The focus on prolonged stimulus presentation, but at the same time the contrast with expanded judgement, did not make sense to me. Perhaps, as a non-native speaker, I am misreading the subtle difference between "protracted sampling" and "longer sampling", but again, the longer duration seems mostly a design choice.

      We thank the reviewer for this impression, which has helped us revise the introduction to more clearly motivate the paradigm as an interesting case for close examination. The primary driver of our choice of stimulus and task parameters was not to enable model comparison using the CPP; it was to examine a decision scenario that exists in everyday life but that has not been examined in terms of underlying decision mechanisms because it offers only sparse behavioural data - the scenario in which plainly visible objects (without noise or stochasticity, as in daylight conditions) need to be examined for a subtle feature difference to guide a later action. The reviewer echoes our point in the Intro, that despite the absence of physical noise in the stimulus, perceptual processing itself is not noise-free. Therefore, temporal integration is certainly not precluded, but its benefit is minimised and less obvious to the decision maker. Given the examples we raise where integration was found to not be employed to its optimal extent (e.g. bound setting foregoing accuracy improvements with duration), and the various theoretical accounts citing energy costs associated with integration and the fleeting nature of many natural environments where prolonged deliberation about a static stimulus is not the norm (e.g. Uchida et al 2006), it is quite hard to guess a priori whether humans will engage in protracted sampling and integration in this case, in practice, even if it is optimal under basic assumptions. As we make clear in our revised Intro, this theoretical interest in the uncertain case of long, noise-free stimuli where perfect, unbounded integration may be optimal but seems doubtful given extant empirical findings, is coupled with a methodological interest in the extent to which neural signatures of decision formation can ‘come to the rescue’ and provide grounds for reliable adjudication between competing mathematical models, when behavioural data fall short.

      More could be said about the optimality of the extrema detection methods. In particular, decades of work (centuries?) have shown that evidence integration is an optimal decision-making procedure: For example, the Sequential Probability Ratio Test is Bayes-optimal wrt mean RT (Wald, 1946); evidence accumulation together with collapsing threshold serves to maximize rewards in repeated choices (e.g., Bogacz et al., PsychRev, 2006; Boehm et al. APP, 2020). Given all this work, why would the brain have evolved to adopt a different mechanism? I realize that the paper is not about optimal decision making, but some discussion of this point seems warranted.

      We had a similar impression initially when reading Stine et al., (2020) where extrema-detection was pitted against integration - is extrema detection so suboptimal that it is too implausible to even consider? Ditterich (2006) argued that signal-to-noise ratio would have to be implausibly high for extrema-detection to produce the behaviour observed on typical decision tasks. However, the fact is, we do not know the effective signal-to-noise ratio, nor can we precisely quantify the costs associated with prolonged evidence accumulation, such as attentional or energetic costs (Drugowitsch et. al., 2012). Even if the extrema detection strategy appears implausibly suboptimal, it is an important principle to demonstrate how not only behavioural but also neural decision signal dynamics can be so nicely consistent with integration yet technically can be quantitatively captured with non-integration mechanisms.

      (2) Modeling choices.

      The authors introduce a parameter, sampT, that represents uncertainty in the sampling onset time. It was not clear to me whether this parameter represented an offset of all trials, or a distribution (probably the latter). I wonder how exactly this parameter was integrated into the models, and in particular, if and how it interacts with the starting-point parameters. My intuition is that on a single-trial, IF early sampling occurs, you can model that with either a negative sampT and z at 0, or with sampT at 0 but a shift in z. This would suggest trade-offs between these parameters, making them hard to estimate independently. Since the paper does not depend on the identification of parameter estimates, this may not be a huge problem, but nevertheless it is good to explore the consequences.

      We thank the reviewer for raising an important question regarding the relationship between sampT and starting-point variability (sz). Mechanistically, early accumulation onset can indeed generate effects that resemble starting-point variability: if accumulation begins during a period containing only zero-mean noise, then by the time informative evidence appears, the decision variable will already have diffused away randomly from zero. In this sense, negative sampT can induce variability in the state of the accumulator at evidence onset. However, the two mechanisms are not mathematically equivalent. The sz parameter assumes a uniform distribution over starting points, whereas the variability induced by early accumulation onset would instead reflect the distribution resulting from integrating zero-mean Gaussian noise over variable durations. Aside from this distinction between distribution shapes, the reviewer is correct that these parameters could partially trade off with one another when sampT takes negative values. In our model, however, sampT was allowed to take either positive or negative values. Positive values delay the onset of evidence integration relative to the evidence, thereby ignoring the first samples, very different from the effect of starting point variability. Nevertheless, to the extent that they can partially trade off each other to some degree, the consequent problem this might cause to accurately estimating both parameters is part of the reason we do not fit a model that includes both simultaneously.

      The way the Bounded Integration model (BIntg) is formulated seems very close to the EZ-diffusion model (Wagenmakers et al., PBR, 2007). This model states that the proportion of correct responses Pc = 1/(1+exp(-B*D/s^2), with B and D the bound and drift rate parameters, respectively. However, filling in the numbers for the high contrast condition from Table 2, and assuming that s=2 (because the model description states that dt=2, with s undefined), I get a Pc of 80% for the 1.6H condition. This seems substantially less than what Figure 2 suggests.

      As we had stated in the Methods section, the model used “a standard deviation of 0.1 arbitrary units” for the evidence. We now make it more explicitly clear that this corresponds to setting within-trial noise s = 0.1 as the scaling parameter (the first paragraph of ‘Model Fits to behaviour only’ and the first paragraph of ‘Integration models’ in Methods). Replacing (s = 0.1) in the suggested calculation yields a predicted accuracy close to 1 for the high-contrast condition, consistent with both the behavioural data and the model predictions shown in Figure 2. We have now clarified this parameter explicitly in the revised Methods and updated Table 1 and Table 2 to avoid confusion.

      On some occasions, it is unclear to me what modeling choices are being made:

      (a) It seems as if the models are fit on accuracy data alone (before introducing the neural data). This seems suboptimal given that the authors do report differences in RT.

      Because of the delayed-report feature of the task, RTs do not directly reflect decision termination time, which is the basis of the use of RT in cognitive modelling normally. Here, whether the decision process has concluded during the stimulus or not, indeterminate response-cue detection and motor execution processes intervene between the stimulus and RT, which would necessitate complicating the models with additional mechanisms, of which there are several possibilities as reflected in the response to the reviewer’s final comment about the RT effects below. This could potentially obscure the core mechanisms of decision formation during the stimulus itself, which was the focus of the study.

      (b) Are the models fit on all data combined, or on the data of individual participants? Fitting individual participant data is preferred, as combined or aggregated data may be distorted by individual differences.

      Because of the noise and variability of EEG data at the single-participant level, we model data averaged across participants, which we have ensured is clear in the revised paper. We provide individual accuracy trends in Figure 1, to verify that the accuracy improvements with increasing evidence duration seen on average are representative of the vast majority of individual subjects. We also added a comment on the limitation this incurs regarding individual difference analysis in the revised discussion.

      (c) The authors seem to suggest that the diffusion coefficient s is estimated (in the section "Integration models"). Most likely, however, this is set to a fixed value. Obviously, it matters for the model comparison using AIC whether this parameter was freely estimated or not.

      As noted in our response to an earlier Comment, the diffusion coefficient was fixed at s=0.1, and to make this explicit, we have entered it for all models in the revised Table 1 and Table 2.

      Not really a weakness, but I wondered about the effect of stimulus duration on RT. In particular, what hypothesis (or post hoc explanation) do the authors have for these RT effects? I could think of at least three hypotheses that are consistent with the behavioral data:

      (a) H1: The shorter the evidence duration, the more likely participants are to require a double-check before response execution, reflecting their uncertainty about their decision.

      (b) H2: There is a collapsing threshold that initiates at stimulus offset, leading to quicker responses on trials where there is more evidence.

      (c) H3: motor preparation is correlated with the evidence signal, which leads to faster responses on trials with more evidence.

      We thank the reviewer for these hypotheses. We agree that the RT effects admit multiple possible interpretations, and while we are cautious not to overinterpret them mechanistically in the paper given our focus on the decision process during the stimulus preceding these response-cue-triggered responses, we do take them to signify that the decision process has not always fully completed and been transformed to a finalised action plan by the time of response cue (start of Results section). To consider these interesting possibilities further:

      We agree that the longer RTs for shorter duration, more uncertain trials could reflect a “double-checking” process (H1), but it could alternatively reflect the fact that if a bound has already been reached during the stimulus, this commitment can be translated to fully-selected action plan that only needs triggering, whereas if a bound has not been reached by stimulus offset, which would occur more often for shorter evidence-duration trials, more of the motor action-selection process would yet need to be completed to initiate the action, causing the slight delay in RT. In other words, on longer-duration trials, the accumulated evidence is more likely to have already reached the bound before the response cue, allowing participants to both commit to a choice and prepare the associated motor response in advance. Therefore, RTs would be shorter as the remaining processes after the cue primarily involve cue detection and motor execution.

      This interpretation is broadly compatible with the reviewer’s H3 account, in the sense that motor preparation may track the evolving decision variable/evidence state. It is also possible that collapsing bounds are set on a post-stimulus, cue-evoked process (H2), which is not mutually exclusive with the above possibilities. It would be hard to determine whether such a process is primarily a response cue-detection decision process that is modulated by uncertainty state at stimulus offset, or a cue-triggered “double-check” process that perhaps operates on the iconic memory of the evidence, and our paradigm does not allow these alternatives to be cleanly dissociated.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you can see, the reviewers are positive about the work and highlight several strengths, while at the same time offering recommendations for improvement. Once these points are satisfactorily addressed, this may also lead to a revision of the eLife assessment below.

      Reviewer #1 (Recommendations for the authors):

      As outlined in my public review, my most important recommendations relate to sample size and possible neural bases of the extremum-flagging model.

      Minor Comments:

      (1) p. 4 rmANOVA statistic is listed as 12.35.31.

      This has been corrected, with thanks for spotting it.

      (2) The stable d-SSVEP amplitude during the evidence period is used to justify a constant (non-adapting) drift rate assumption. This is a reasonable inference, but the SSVEP reflects early sensory encoding rather than the decision variable per se. Neural adaptation or gain changes at later processing stages could still produce a non-constant effective drift rate even with a stable sensory representation. This inference should be qualified.

      This is true. We have clarified that these checks for one important potential source of a time-varying drift rate, namely adaptation at the level of early sensory representation, but admit that other effects may happen downstream.

      (3) The Methods describe a leaky accumulation extension; the Results note leak ≈ 0.0002 at w=10, which is effectively zero. This is a positive result (evidence against leaky integration) that should be stated more explicitly in the Results or Discussion rather than appearing only in supplementary tables.

      We thank the reviewer for highlighting this. We have pointed to this result now in the first paragraph of the Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) The panels in Figure 4 H, I, J are not discussed in the Neurally-constrained models Section, while I believe they are probably more informative than Figure 4E alone.

      The manuscript has been revised to explain Figure 4H,I, J fully under ‘Neurally-constrained models’.

      (2) In the method section, we don't know how many trials were rejected based on the chosen threshold.

      The Method section is updated with the rejected trials after preprocessing. It now reads, “After preprocessing, 12017 trials remained across all conditions, with an average rejection rate of 15% (± 12.9%) across participants.”

      (3) Figure 1: For b and c, data are mean {plus minus} s.e.m. after between-participant variance was factored out. -> reference or detailed method.

      We have now explained in the caption that this is done by subtracting the overall mean of each individual from their data and adding back the grand mean, retaining the between-condition differences - that is, we remove the component of variance that repeated-measures tests ignore.

      (4) Figure 2: Shouldn't the snapshot model only feature one sample, as in Stine et al. 2020? This figure and others would also benefit from a better resolution.

      The reviewer is correct regarding the schematic of the snapshot model presented in Stine et al., (2020). We used small, light orange dots to represent the evidence samples presumably being encoded and a larger orange dot to indicate the single randomly chosen sample used as evidence for the decision. In our revised figure, we have increased the visual distinction between the evidence dots and the selected dot, and pointed this out in the caption. We have also improved the figure resolutions.

      (5) Typo:

      - Semi-saturation Table S4.

      - pi missing in the text of the G^2 equation.

      Thank you for catching these.

      Reviewer #3 (Recommendations for the authors):

      Small, random points:

      (1) How was it ensured that participants indeed did not detect the change in contrast throughout the 1.6s interval? In previous work (Winkel et al., PBR, 2014) we did something similar in a random-dot motion task, but observed that participants always observed the change, if we did not slowly change the coherence of the stimulus (unfortunately, Winkel et al., 2014 is not explicit about the exact parameters of the change, but Figure 2 suggests that the change in coherence lasted 50ms, independent of stimulus strength).

      It is true that abrupt changes in stimulus strength are more salient and detectable than a ramped change. The experimenters tested the stimulus during the task design subjectively, to satisfy themselves that they could not tell when the contrast stepped back to baseline, but this was not verified systematically with psychometrics, nor can we be sure that a more sensitive observer couldn’t sometimes detect the change. However, based on the task design, stimulus properties, and both behavioural and neural data, we are confident that participants are very unlikely to have been sufficiently confident in detecting the step-back in contrast to ceasing their contrast-comparison decision process at that point:

      (1) Participants were naive to the underlying manipulation. They were informed that trials would naturally vary in difficulty. It was normal for them to perceive some trials as harder than others without suspecting a mid-trial structural change.

      (2) The contrast difference in the hard condition was very subtle, and though it is possible that the step-down in contrast could be detected with above-chance accuracy if instructed to do so, given there was no instruction on whether and when the step-down would happen, it is very unlikely they could be detected with sufficient confidence to be certain there is no remaining evidence in the stimulus. Furthermore, the rapid, flickering nature of the stimulus would have helped to mask the transition point, in comparison with a sudden change in a continuously-playing random-dot motion stimulus (as in Winkel et al., 2014).

      (3) Our behavioural post-cue RT data imply the participants did not cease decision formation at evidence offset. If they had, then they would have been afforded the most time to prepare their chosen action in advance of the response cue in the case of the earlier evidence offset (i.e. shorter durations), yet these were the conditions with the longest, not the shortest post-cue RTs.

      (4) The low-contrast CPP traces remained elevated for the full 1600 ms interval, unperturbed by the evidence offsets. If participants were explicitly detecting a sudden change in contrast, we would expect to see a transient evoked response locked to that change, marking that detection. Instead, the sustained elevation of the CPP is characteristic of a continuation of the decision process, uninterrupted, through the subtle offsets.

      (2) Could you include the regression coefficients of the statistical modeling of the behavioral data?

      The regression coefficient is now added to the second paragraph of the Results section.

      (3) I felt Figure 5B was a bit confusing: Are the dashed lines here the ipsilateral sides or the non-linear bounds? This was confusing because "data" only has a solid line in the legend.

      We agree that the subtle nonlinearity, which appears visually close to the linear model, may have caused this confusion. To improve clarity, we have added arrows to Figure 5B to explicitly indicate which legend refers to which panel, and distinguish the dashed nonlinear bounds from the other traces.

      (4) "amplitude variations [...] to be used as an independent evaluation of model fit": Could you refer to where these model predictions are presented? I think this is in Figure 4 - Sup 5?

      We thank the reviewer for raising this. The model-predicted waveforms showing amplitude variations across durations for all neural weightings are presented in Figure 4 - Supplementary Figures 2 and 5. We have clarified this in the revised manuscript.

      References

      Ditterich, J. (2006). Evidence for time‐variant decision making. European Journal of Neuroscience, 24(12), 3628–3641. https://doi.org/10.1111/j.1460-9568.2006.05221.x

      Drugowitsch, J., Moreno-Bote, R., Churchland, A. K., Shadlen, M. N., & Pouget, A. (2012). The cost of accumulating evidence in perceptual decision making. Journal of Neuroscience, 32(11), 3612–3628.

      Ehinger, B. V., & Dimigen, O. (2019). Unfold: An integrated toolbox for overlap correction, non-linear modeling, and regression-based EEG analysis. PeerJ, 7, e7838.

      Latimer, K. W., Yates, J. L., Meister, M. L. R., Huk, A. C., & Pillow, J. W. (2015). Single-trial spike trains in parietal cortex reveal discrete steps during decision-making. Science, 349(6244), 184–187. https://doi.org/10.1126/science.aaa4056

      Loughnane, G. M., Newman, D. P., Bellgrove, M. A., Lalor, E. C., Kelly, S. P., & O’Connell, R. G. (2016). Target selection signals influence perceptual decisions by modulating the onset and rate of evidence accumulation. Current Biology, 26(4), 496–502.

      Nieuwenhuis, S., Aston-Jones, G., & Cohen, J. D. (2005). Decision making, the P3, and the locus coeruleus–norepinephrine system. Psychological Bulletin, 131(4), 510.

      Stine, G. M., Zylberberg, A., Ditterich, J., & Shadlen, M. N. (2020). Differentiating between integration and non-integration strategies in perceptual decision making. Elife, 9, e55365.

      Uchida, N., Kepecs, A., & Mainen, Z. F. (2006). Seeing at a glance, smelling in a whiff: Rapid forms of perceptual decision making. Nature Reviews Neuroscience, 7(6), 485–491.

      Weindel, G., van Maanen, L., & Borst, J. P. (2024). Trial-by-trial detection of cognitive events in neural time-series. Imaging Neuroscience, 2, imag–2.

      Winkel, J., Keuken, M. C., Van Maanen, L., Wagenmakers, E.-J., & Forstmann, B. U. (2014). Early evidence affects later decisions: Why evidence accumulation is required to explain response time data. Psychonomic Bulletin & Review, 21(3), 777–784. https://doi.org/10.3758/s13423-013-0551-8

    1. eLife Assessment

      This fundamental study, supported by convincing evidence, identifies the specific neuronal components that enable HIF-1 signaling in specific serotonergic neurons to regulate effectors in peripheral tissues. Specifically, stabilizing HIF-1 in ADF and NSM serotonergic neurons is sufficient to extend lifespan and healthspan in C. elegans, and these physiological benefits require several oxygen-sensing neurons, as well as neurotransmitters (GABA and tyramine), and a neuropeptide (NLP-17). The strengths of the study include rigorous experimental design, high quality data, the use of multiple genetic tools, and functional dissection using the well-validated C. elegans model.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this study by Kitto et al., the authors set out to identify specific signaling components regulating the hypoxic response from the neurons to the periphery and which components are required for lifespan extension. Their previous work had shown that expression of a stabilized HIF-1 mutant in the nervous system extends lifespan through the serotonin receptor SER-7 and leads to the induction of fmo-2 in the intestine. In the current study, they mapped the precise neural circuits required for this response, as well as the signaling mediators. Their work reveals that neurotransmitters GABA and tyramine, and the neuropeptide NLP-17, act downstream of neuronal HIF-1 to convey a "hypoxic signal" to peripheral tissues. Through cell-type-specific expression studies, targeted knockouts, and comprehensive lifespan analysis, the authors provide robust evidence to support their conclusions. The insights gained from the study are both moving the field forward as they advance our understanding of neuro-peripheral hypoxic signaling, but they also lay the groundwork for potential therapeutic strategies aimed at the modulation of such signaling pathways.

      Strengths:

      (1) This study provides new evidence further delineating signaling components required for hypoxic signaling-mediated longevity, from the nervous system to the periphery. Using a rigorous approach where they express stabilized HIF-1 mutant selectively in ADF, NSM, and HSN serotonergic neurons, followed by cell-type-specific tph-1 knockouts to pinpoint ADF-dependent serotonin signaling as essential for both lifespan extension and intestinal fmo-2 induction.

      This was followed by generating 11 transgenic lines that drive SER-7 expression under distinct neuron-specific promoters, to systematically tease out in which of 27 candidate neurons SER-7 functions to mediate hypoxia-induced longevity. This ultimately highlighted the RIS interneuron as the required signaling hub.

      (2) As the intestine lacks direct neuronal innervation, the authors employ neuron-specific RNAi (TU3311 strain) and dense core vesicle analyses to identify that the neuropeptide NLP-17 is required to transmit the hypoxic signal from RIS to induce fmo-2 in the intestine.

      (3) Overall, the paper is very well written. The experiments were carried out carefully and thoroughly, and the conclusions drawn are also well supported by the results they are showing.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aimed to identify the specific neurons, neurotransmitters, and neuropeptides that mediate the longevity effects of the hypoxic response in C. elegans. By genetically dissecting the pathway downstream of HIF-1, they define a neural circuit involving ADF serotonergic neurons, the SER-7 receptor in the RIS interneuron, tyraminergic signaling from RIM, and neuropeptide NLP-17, ultimately linking neuronal hypoxic sensing to pro-longevity signaling in the intestine.

      Strengths:

      The study employs a diverse genetic toolkit, including neuron-specific transgenes, tissue-specific knockouts and rescues, RNAi knockdowns, allowing the authors to pinpoint causality, sufficiency, and necessity with high resolution. The comprehensive mapping of cell-nonautonomous signaling adds depth to our understanding of how HIF and serotonin signaling interface with aging pathways. The conclusions are supported by consistent survival assays and fmo-2 gene expression analyses.

    4. Reviewer #3 (Public review):

      Summary:

      This study found that ADF serotonergic neurons have a significant role in extending lifespan mediated by HIF-1, as well as serotonin receptor SER-7 in the GABAergic RIS interneurons. The author focuses on the sufficiency and necessity of components from the central nervous system and how they contribute to aging upon hypoxia.

      Previous work from the lab has identified that the stabilization of HIF-1 in neurons is sufficient to extend lifespan through the serotonin receptor, SER-7, which subsequently activates fmo-2 in the intestine and leads to lifespan extension. Building on this, the author sought to determine which serotonergic neurons are involved and found that serotonin signaling in ADF neurons is required for lifespan extension mediated by HIF-1.

      The author next tested which subset of neurons requires Ser-7 expression to rescue hypoxic response. They found that ser-7 expression in multiple neurons is sufficient to induce fmo-2, with the top candidate being the RIS neuron. Ablation of the RIS neuron did not extend lifespan, suggesting that ser-7 expression in the RIS neuron is required for lifespan extension, positioning it as a key component in the longevity signaling pathway.

      The author also investigated neurotransmitters and found that GABA and tyramine are important components in this circuit. They showed that the tyramine receptor called tyra-3 is required for vhl-1-mediated longevity. Given that tyra-3 is expressed in oxygen- and carbon dioxide-sensing neurons, the author demonstrated that these sensing neurons work downstream of serotonin signaling. Lastly, the author screened neuropeptide/receptor binding pairs and identified NLP-17 as playing a role in hypoxia-mediated longevity.

      Originality and Significance:

      This research is significant in that it uncovers components that are sufficient and necessary for lifespan extension via the hypoxic response. It provides comprehensive data supporting longevity induced by HIF-1-mediated hypoxic response, in conjunction with fmo-2, a longevity gene, as demonstrated in previous work from the lab. Moreover, it provides a number of new transgenic worm tools for C. elegans and aging communities.

      Conclusions:

      This study provides insights into how hypoxic response regulates aging in a cell non-autonomous manner, outlining a potential circuit involving neurons, neurotransmitters, and neuropeptides.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We thank the reviewers and editors for their time and valuable input into improving our manuscript. The reviewers recognized the value of this work while identifying places where further explanation and/or additional experiments could strengthen the manuscript. We greatly appreciate this feedback and have addressed reviewer comments through additional experiments and necessary textual edits.

      In response to the reviewer comments, our principal data-driven changes to revise the manuscript include: 1) testing how stabilizing HIF-1 in the ADF and NSM neurons affects healthspan; 2) measuring for interactions between HIF-1 stabilization in serotonergic neurons and the mt-UPR; and 3) measuring nlp-17 expression downstream of HIF-1 stabilization in serotonergic neurons. We also made textual changes including: 1) clarifying that our study focuses specifically on the vhl-1-mediated genetic hypoxic response; 2) summarizing which parts of the working model are experimentally validated vs. speculative; and 3) expanding our discussion of future directions needed to understand the epistasis of the many signals acting in this circuit. Our responses to each reviewer comment are below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study by Kitto et al., the authors set out to identify specific signaling components regulating the hypoxic response from the neurons to the periphery and which components are required for lifespan extension. Their previous work had shown that expression of a stabilized HIF-1 mutant in the nervous system extends lifespan through the serotonin receptor SER-7 and leads to the induction of fmo-2 in the intestine. In the current study, they mapped the precise neural circuits required for this response, as well as the signaling mediators. Their work reveals that neurotransmitters GABA and tyramine, and the neuropeptide NLP-17, act downstream of neuronal HIF-1 to convey a "hypoxic signal" to peripheral tissues. Through cell-type-specific expression studies, targeted knockouts, and comprehensive lifespan analysis, the authors provide robust evidence to support their conclusions. The insights gained from the study are both moving the field forward as they advance our understanding of neuro-peripheral hypoxic signaling, but they also lay the groundwork for potential therapeutic strategies aimed at the modulation of such signaling pathways.

      We appreciate the reviewer’s positive assessment of the topic and general interest in this work.

      Strengths:

      (1) This study provides new evidence further delineating signaling components required for hypoxic signaling-mediated longevity, from the nervous system to the periphery. Using a rigorous approach where they express stabilized HIF-1 mutant selectively in ADF, NSM, and HSN serotonergic neurons, followed by cell-type-specific tph-1 knockouts to pinpoint ADF-dependent serotonin signaling as essential for both lifespan extension and intestinal fmo-2 induction.

      This was followed by generating 11 transgenic lines that drive SER-7 expression under distinct neuron-specific promoters, to systematically tease out in which of 27 candidate neurons SER-7 functions to mediate hypoxia-induced longevity. This ultimately highlighted the RIS interneuron as the required signaling hub.

      (2) As the intestine lacks direct neuronal innervation, the authors employ neuron-specific RNAi (TU3311 strain) and dense core vesicle analyses to identify that the neuropeptide NLP-17 is required to transmit the hypoxic signal from RIS to induce fmo-2 in the intestine.

      (3) Overall, the paper is very well written. The experiments were carried out carefully and thoroughly, and the conclusions drawn are also well supported by the results they are showing.

      Weaknesses:

      Overall, I don't see many weaknesses. One point relates to their read-outs, which rely heavily on lifespan measurements and fmo-2 induction without evaluating other physiological processes that serotonin or NLP-17 might affect. For translational relevance, it would be valuable to assess or mention potential adverse effects, such as changes in reproduction, pharyngeal pumping, or proteostasis capacity (proteostasis capacity specifically in the tissue showing fmo-2 upregulation).

      We thank the reviewer for the positive review and fully agree and acknowledge that the primary readouts used in this work, fmo-2 induction and lifespan, may not reflect other important elements of health and physiology. To address this weakness, we have performed three measurements of healthspan (pumping, thrashing, and maximum velocity), in the ADF and NSM HIF-1 stabilized strains at young adulthood and at middle age. We focused on examining these HIF-1 stabilized strains rather than our nlp-17 knockout animals because stabilizing HIF-1 in the NSM or ADF neurons is sufficient to extend lifespan. nlp-17 is necessary, but it remains unclear whether nlp-17 signaling is sufficient to extend lifespan. Our new data show that both ADF- and NSM-specific HIF-1 stabilization had no effect on pumping rate in young (day 1 of adulthood) worms. In aged animals (day 12 of adulthood), however, NSM-, but not ADF-, specific HIF-1 stabilization rescued the pumping rate decline in hif-1 knockout compared to WT worms (new Fig. S1C). Similarly in the new thrashing data, NSM-, but not ADF-, specific HIF-1 stabilization rescued the thrashing rate decline in the hif-1 knockout young and aged worms (new Fig. S1D). Lastly, our new data show that ADF and NSM HIF-1 stabilization had no effect on average or maximum movement speed at days 1 and 5 of adulthood (new Fig. S1E-F). Together, these results indicate that genetic activation of the hypoxic response in NSM neurons but not in the ADF neurons could improve healthspan.

      While we did not examine reproductive capacity in these strains, we have expanded our discussion section to mention the importance of fully characterizing other elements of physiology, health, and behavior in future work.

      “Another limitation of this work is that it uses lifespan as the main readout for organismal health. While genetic manipulations that extend lifespan often improve stress resistance and healthspan [8,80], longevity manipulations can also have adverse effects on reproduction [81,82] and behavior [83,84]. In this study, we find that HIF-1 stabilization in the ADF neurons does not prevent the deleterious effects of hif-1 knockout on mobility, but that HIF-1 stabilization in the NSM neurons may attenuate age-related decline in pumping and thrashing (Fig. S1). However, future work should examine whether other modifications to this pathway, such as manipulations to RIM, RIS, or NLP-17 signaling, influence healthspan in addition to lifespan. It will be important for future studies to determine whether various components of this pathway affect both longevity and the response to different types of stressors like oxidative stress, proteotoxic stress, and infection, as HIF-1 activity also interacts with multiple stress responses [41,42,39,40,43].”

      While lifespan assays and fmo-2 expression do provide strong evidence, incorporating additional markers of stress resistance could strengthen the link between hypoxic signaling and organismal health as well.

      We also measured hsp-6 expression via qPCR in the serotonergic neuron-specific HIF-1 stabilized strains to determine whether these conditions that lead to upregulated fmo-2 may also affect the mt-UPR. Interestingly, we find that stabilizing HIF-1 in either the ADF or NSM serotonergic neurons decreases hsp-6 expression relative to WT worms (new Fig. S1G). This could suggest either that the mt-UPR response is impaired in these worms, or that HIF-1 stabilization decreases proteotoxic stress leading to a lower basal level of hsp-6. Although this method of measurement did not allow us to interrogate whether these changes occur in the specific tissues where fmo-2 is upregulated, we have expanded our discussion to emphasize that further investigation of cell- and tissue-specificity within the hypoxic response should be a focus of future work.

      Finally, we agree that it is important to examine whether activating the hypoxic response promotes stress resistance in addition to longevity. We did not focus on this element of the hypoxic response in this work because HIF activity is known to promote adaptive stress-responses to some stressors like infection [7,8] and oxidative stress [9,10], while simultaneously impairing the proteosasis stress response [11]. We have added this important information to our introduction section, and have expanded our discussion section to emphasize that a key future direction will be to test which specific components of this pathway also facilitate stress resistance.

      Introduction Section Modification:

      “However, the physiological changes induced by the hypoxic response are broad and involve adaptations such as increased vascularization, metabolic rewiring, and changes in cell survival pathways. In mammals, some of these same adaptations can be detrimental, as mutations in components of the hypoxic response have been linked to conditions like cancer and cardiovascular disease [12-15]. Additionally, HIF activity is essential to promote some forms of stress resistance to infection [9,10] and oxidative stressors [7,8] but can have detrimental effects on proteostasis [11].”

      Discussion Section Modification:

      “Another limitation of this work is that it uses lifespan as the main readout for organismal health. While genetic manipulations that extend lifespan often improve stress resistance and healthspan [1,2], longevity manipulations can also have adverse effects on reproduction [3,4] and behavior [5,6]. In this study, we find that HIF-1 stabilization in the ADF or NSM neurons has little effect on mobility in young and aged animals. However, future work should examine whether other modifications to this pathway, such as manipulations to RIM, RIS, or NLP-17 signaling, influence healthspan in addition to lifespan. It will be important for future studies to determine whether various components of this pathway affect both longevity and the response to different types of stressors like oxidative stress, proteotoxic stress, and infection, as HIF-1 activity also interacts with multiple stress responses [7,8,9-11].”

      Reviewer #2 (Public review):

      Summary:

      The authors aimed to identify the specific neurons, neurotransmitters, and neuropeptides that mediate the longevity effects of the hypoxic response in C. elegans. By genetically dissecting the pathway downstream of HIF-1, they define a neural circuit involving ADF serotonergic neurons, the SER-7 receptor in the RIS interneuron, tyraminergic signaling from RIM, and neuropeptide NLP-17, ultimately linking neuronal hypoxic sensing to pro-longevity signaling in the intestine.

      Strengths:

      The study employs a diverse genetic toolkit, including neuron-specific transgenes, tissue-specific knockouts and rescues, RNAi knockdowns, allowing the authors to pinpoint causality, sufficiency, and necessity with high resolution. The comprehensive mapping of cell-nonautonomous signaling adds depth to our understanding of how HIF and serotonin signaling interface with aging pathways. The conclusions are supported by consistent survival assays and fmo-2 gene expression analyses.

      Weaknesses:

      A key limitation is the lack of clear evidence showing epistasis of so many identified molecular/neuronal components downstream of HIF-1 and serotonin. Thus, the mechanisms of how a diverse set of molecules/neurons coordinate and mediate neuronal HIF-1 effects on intestinal fmo-2 and longevity remain murky.

      We thank the reviewer for these important points. We agree that the epistatic relationships between ADF serotonin, RIM tyramine, RIS GABA, and neuropeptide NLP-17 signaling remain unclear within this pathway. Determining the epistatic relationships of each signal within this complex pathway will require: 1) generating genetic manipulations to each identified signaling component that may mimic vhl-1 knockout to promote longevity; 2) crossing these new strains into multiple genetic knockouts we identified as required for vhl-1 mediated longevity; and 3) measuring the lifespans of each double and triple mutant. We are very interested in testing the epistatic relationships between all of these molecules and neurons, and believe this extensive follow-up exploration will generate significant future results.

      To better address these limitations of the current work, we have added a summary table to Fig. 7 as well as a paragraph to our discussion section. This table and paragraph better explain which components of our working model have been tested for necessity, sufficiency, and epistasis, and which components of this model remain unclear (Fig. 7). See updated Fig. 7 with added table.

      Updated discussion section detailing this limitation:

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1 depletion, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant. This approach would also narrow down which signals are downstream of the genetic activation of the hypoxic response, and which are sufficient to extend lifespan upstream of the hypoxic response in a normoxic environment. One notable target for further exploration is the SER-7 expressing RIS neuron, which plays a role in sleep [16] and stress resistance [17], and can extend lifespan when optogenetically activated under normoxic conditions [18].”

      Some rescue strategies may inadvertently cause non-physiological expression.

      This is a great point. We bring attention to this limitation in the discussion section. Our cell-specific knockouts (ADF tph-1 KO, Fig. 1E) or ablations (RIS ablation, Fig. 2C) data showed that the neurons identified using rescue strains (i.e., tph-1 in the ADF and ser-7 in RIS) are likely not false positives. However, we did not generate a RIM-specific knockout or ablation strain to confirm our tdc-1 results and have added this limitation to the discussion of these results.

      Discussion of rescue strategy limitations and the need for a RIM-specific knockout:

      “The circuit-mapping approaches employed in this work are also impacted by limitations in cell-specific genetic modifications and in the use of RNAi knockdown. For example, cell-specific rescue constructs can sometimes lead to unintended rescues in other cell types due to cell-nonautonomous signaling. Because all serotonin-producing neurons also express the serotonin reuptake transporter mod-5, serotonin produced by one cell in our rescue strains could be taken up by other serotonin-producing neurons, leading to unintended signaling effects. This may also be true of the uv1 and RIM tyraminergic rescue strains, although little is known about tyramine reuptake in C. elegans. Two cells identified via tissue-specific rescue experiments (ADF and RIS, Fig. 1G and Fig. 2D) were also found to be necessary via cell-specific knockout (ADF tph-1 KO, Fig. 1E; RIS ablation, Fig. 2C), decreasing the likelihood of a false positive from the rescue strain technique. However, the role of the RIM neuron was identified via a tdc-1 rescue strain and was not validated using a RIM-specific knockout (Fig. 7B). Therefore, the contribution of RIM signaling to this circuit is less well-validated, and a RIM-specific tdc-1 knockout strain should be examined in future work.”

      Additionally, environmental hypoxia was not tested in parallel, so the claim on "hypoxia response" throughout the manuscript is not justified by genetic manipulation alone, and the translational relevance of the genetic manipulations remains somewhat uncertain.

      We thank the reviewer for identifying the need for additional specificity in our language. We agree that it is critical to make it clear that this paper only examines genetic mimics of hypoxia, rather than environmental hypoxia, and have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (e.g., “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We discuss this in Discussion and are excited to interrogate these differences in future work.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      Reviewer #3 (Public review):

      Summary:

      This study found that ADF serotonergic neurons have a significant role in extending lifespan mediated by HIF-1, as well as serotonin receptor SER-7 in the GABAergic RIS interneurons. The author focuses on the sufficiency and necessity of components from the central nervous system and how they contribute to aging upon hypoxia.

      Previous work from the lab has identified that the stabilization of HIF-1 in neurons is sufficient to extend lifespan through the serotonin receptor, SER-7, which subsequently activates fmo-2 in the intestine and leads to lifespan extension. Building on this, the author sought to determine which serotonergic neurons are involved and found that serotonin signaling in ADF neurons is required for lifespan extension mediated by HIF-1.

      The author next tested which subset of neurons requires Ser-7 expression to rescue hypoxic response. They found that ser-7 expression in multiple neurons is sufficient to induce fmo-2, with the top candidate being the RIS neuron. Ablation of the RIS neuron did not extend lifespan, suggesting that ser-7 expression in the RIS neuron is required for lifespan extension, positioning it as a key component in the longevity signaling pathway.

      The author also investigated neurotransmitters and found that GABA and tyramine are important components in this circuit. They showed that the tyramine receptor called tyra-3 is required for vhl-1-mediated longevity. Given that tyra-3 is expressed in oxygen- and carbon dioxide-sensing neurons, the author demonstrated that these sensing neurons work downstream of serotonin signaling. Lastly, the author screened neuropeptide/receptor binding pairs and identified NLP-17 as playing a role in hypoxia-mediated longevity.

      Originality and Significance:

      This research is significant in that it uncovers components that are sufficient and necessary for lifespan extension via the hypoxic response. It provides comprehensive data supporting longevity induced by HIF-1-mediated hypoxic response, in conjunction with fmo-2, a longevity gene, as demonstrated in previous work from the lab. Moreover, it provides a number of new transgenic worm tools for C. elegans and aging communities.

      We thank the reviewer for the positive assessment of the manuscript. We appreciate all suggestions for further improving the work and have made changes based on these suggestions (see details below).

      Data and Methodology:

      (1) The experiments were thoroughly conducted, especially the generations of strains using different neuron-type promoters and crossing into mutant strains to demonstrate sufficiency and necessity.

      (2) Some figure legends from the text do not match what the data show. (Figure 6E, F, G).

      We have made changes to the legends accordingly to make sure figures and legends are consistent.

      (3) The lifespan graph legends are confusing and could use some revamping for better clarification.

      We have updated the lifespan graph legends to provide the statistics in the same section as the labels, indicating which line is which condition (see updated Figure 1E). We hope this change better clarifies the lifespan graph legends.

      Conclusions:

      This study provides insights into how hypoxic response regulates aging in a cell non-autonomous manner, outlining a potential circuit involving neurons, neurotransmitters, and neuropeptides.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As suggested by the Reviewers 2 & 3, including environmental hypoxia will broaden the impact of the study. If not, the authors should consider clarifying their response as "genetic activation of the hypoxic response" (see Reviewer 2 below).

      We thank the editor and reviewers 2 and 3 for this clarification. We agree that it is critical to make it clear that this paper only examines genetic mimics of hypoxia, rather than environmental hypoxia, and have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (e.g., “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We have also expanded our discussion of this important future direction.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      Suggested minor changes by Reviewer 3 should also be made to improve the readability of the paper.

      We have made changes suggested by Reviewer 3 to improve the readability of the paper.

      Reviewer #2 (Recommendations for the authors):

      (1) Suggestions for additional experiments or analyses:

      (a) To clarify the hierarchical relationships among the components of the identified circuit, epistasis experiments between serotonin, tyramine, NLP-17, and oxygen-sensing neurons (e.g., double mutants or sequential rescues) would strengthen the proposed model and help determine whether these signals act in parallel or downstream of each other.

      We thank the reviewer for identifying this important caveat. As described in the public review response, we completely agree that understanding the epistasis of serotonin, tyramine, NLP-17, and oxygen-sensing neuron signaling within this pathway is important to fully test our working model. We hope to address these questions about epistasis and interactions between different signals in upcoming projects.

      To clarify that the order of many of these signals remains speculative in our working model, we have added a table to Fig. 7, that summarizes what is known and unknown about the epistasis of these pathway components. We have also expanded our discussion of this limitation. See updated Fig. 7 with added table.

      Updated discussion section detailing this limitation:

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1 depletion, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron we found to be required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant.”

      (b) Testing whether environmental hypoxia (e.g., 0.5-1% O₂ exposure) elicits similar neuronal requirements and fmo-2 induction as the genetic HIF-1 stabilization would validate that the described pathway is relevant to the actual hypoxic response and improve translational relevance.

      We thank the reviewer for identifying the need for clarification. As also discussed in the public review section, we agree that it is important to emphasize that this paper exclusively examines genetic mimetics of hypoxia, rather than actual exposure to a hypoxic environment. To clarify this point, we have replaced every occurrence of “the hypoxic response” with either 1) “genetic activation of the hypoxic response”, or 2) describing the specific manipulation used in that experiment (ie “vhl-1 mediated longevity”). We also agree that examining the similarities and differences between genetic and environmental activation is a key next step towards evaluating the translational potential of this pathway. We discuss this in discussion and are excited to interrogate these differences in future work.

      “While this study identifies many neural signals required for vhl-1 knockdown or knockout to extend lifespan, one key limitation of this work is the potential differences between genetic and environmental methods of inducing the hypoxic response. While vhl-1 knockdown or knockout leads to HIF-1 stabilization by blocking its proteasomal degradation, it also results in hydroxylated HIF-1. This contrasts with environmental hypoxia, in which HIF-1 remains stable because it cannot be hydroxylated. While HIF-1 is stabilized and localized to the nucleus in both cases, there are differences in transcriptional outcomes between stable hydroxylated and unhydroxylated states [19,20]. Additionally, we did not explore the effects of alternative genetic activators of the hypoxic response such as HIF-1 hydroxylase PHD/EGL mutants, which may provide further insight into how different manipulations of the hypoxic response impact longevity. Future work should interrogate the similarities and differences between the circuits driving longevity in response to environmental hypoxia, vhl-1 knockdown, HIF-1 stabilization, and PHD/EGL knockdown.”

      (c) Functional readouts beyond lifespan and fmo-2 induction (e.g., neuronal activity monitoring or optogenetic modulation of key neurons) could help clarify how information flows through the circuit.

      We thank the reviewer for this valuable suggestion. We are hoping to be able to implement these experimental techniques in future work. We agree that in combination with genetic epistasis analyses, direct measurements of neuronal signaling will greatly improve our understanding of the directionality and interactions between signals in this circuit. Our discussion section recommends employing these techniques in future studies.

      “Finally, while the use of RNAi knockdown and genetic knockouts establishes the necessity of many signals within the vhl-1-mediated longevity circuit, the exact directionality of these signals remains unclear. It is possible that increased, decreased, or pulsatile changes in signaling through these bioamines and neuropeptides are required for genetic activation of the hypoxic response to extend lifespan. Work on C. elegans reversal behavior has also revealed an antagonistic relationship between RIM and RIS activity facilitated by both chemical (neuropeptide and tyramine) and electrical (gap junction) signaling [16,21]. This known interaction should also be interrogated in the context of how these cells may communicate following genetic induction of the hypoxic response. Future work in this area could use tools to measure or modify neuronal activity, such as calcium imaging or optogenetics, to begin answering these questions.”

      (2) Recommendations for improving the writing and presentation:

      (a) The manuscript is well written overall, but clarity would be improved by explicitly stating in the abstract and introduction that the study is based on genetic activation of the hypoxic response, rather than environmental hypoxia.

      We appreciate this valuable suggestion and have updated the abstract and introduction to clarify that this work examines genetic activation of the hypoxic response.

      Updated sentences from the abstract:

      “Here, we interrogate the cell-nonautonomous signaling pathway downstream of genetic activation of the hypoxic response.” 

      “Together, these insights develop a circuit for how genetic induction of the hypoxic response cell-nonautonomously modulates ageing and suggests valuable targets for modulating ageing in mammals.”

      Updated sentences from the introduction:

      “In this study, we uncover key neural components of the longevity circuit initiated by genetic induction of the hypoxic response. Within this circuit, we identify individual cells, signals, and receptors necessary and/or sufficient to extend lifespan downstream of genetic activation of the hypoxic response. More specifically, we find serotonin signaling in the ADF serotonergic neurons is both necessary and sufficient to extend lifespan through genetic activation of the hypoxic response. This pathway signals through the serotonin receptor SER-7 in the RIS interneuron. We further demonstrate additional neurotransmitters (GABA and tyramine), and a neuropeptide (NLP-17) are critical for mediating these longevity effects. Finally, we identify that oxygen sensing neurons (URX, AQR, PQR and BAG) act downstream of neuronal HIF-1 in this circuit. Our insights into this longevity pathway provide a mechanistic understanding of how genetic activation of the hypoxic response delays aging and improves health.”

      (b) In the discussion, clearly delineating which parts of the proposed pathway are firmly established versus inferred would aid interpretation.

      We thank the reviewer for this idea, and have added the following text to the discussion:

      “Evidence for the necessity, sufficiency, and epistatic relationships between each signal are summarized in new Fig. 7B. In brief, all signaling molecules presented in this work are necessary for vhl-1 to extend lifespan. Rescuing ADF serotonin production, RIS ser-7 expression, and RIM tyramine production is sufficient for vhl-1 to extend lifespan. The sufficiency of oxygen sensing neurons BAG and UPA/PQR/AQR, and the neuronal signals of GABA, NLP-17, and TYRA-3 to restore vhl-1 mediated longevity remains unclear. All signals act downstream of vhl-1. ADF HIF-1 stabilization and SER-7 signaling act upstream of fmo-2 induction, and the oxygen sensing neurons (BAG, UPA/PQR/AQR) act upstream of or in parallel to ADF HIF-1 stabilization.”

      (c) Adding a summary table or schematic that visually distinguishes necessity vs. sufficiency for each component (e.g., ADF, RIS, RIM, NLP-17) would make the overall model more accessible.

      We thank the reviewer for this great suggestion and have added a table to new Fig. 7B that summarizes what is known about necessity for vhl-1, sufficiency for vhl-1, and sufficiency to extend lifespan independently of vhl-1 for each signal (table included in response to Reviewer 2, comment 1a).

      (3) Minor corrections and clarifications:

      (a) Define or replace "hypoxic response" with "genetically induced hypoxic response" where appropriate to avoid conflating genetic manipulations with actual environmental hypoxia.

      We appreciate this valuable suggestion and have replaced “the hypoxic response” and with “genetic activation of the hypoxic response” or “genetically induced hypoxic response” throughout the manuscript to clarify this point.

      (b) All genes and alleles should be italicized per worm nomenclature.

      We thank the reviewer for this comment and have reviewed the manuscript to italicize all gene names and alleles. In some locations, the protein is referred to instead of the gene using the conventional uppercase non-italicized format.

      Reviewer #3 (Recommendations for the authors):

      (1) Using vhl-1 RNAi as the sole approach to demonstrate hypoxic response appears somewhat limited, as vhl-1 is also involved in HIF-1 independent processes that can influence lifespan in C. elegans. Including additional downstream effectors of HIF-1, such as egl-9, or having HIF-1 nondegradable strain as validation could strengthen the findings.

      We thank the reviewer for this valuable comment and agree that an important next step is to test whether these signals are also required for other genetic (HIF-1 stabilized, egl-9) and environmental activators of the hypoxic response to extend lifespan. We have worked to clarify that this paper focuses primarily on vhl-1 mediated longevity throughout the text and have also included this limitation in our discussion section (for details, please see response to Reviewer 2 public review).

      (2) Investigating how the healthspan is affected by serotonergic neuron-specific hypoxic responses would be interesting and could enhance understanding of the physiological mechanisms underlying lifespan extension. 

      We appreciate this suggestion, and have performed three measurements of healthspan (pumping, thrashing, and maximum velocity), in the ADF and NSM HIF-1 stabilized strains at young adulthood and at middle age. We found that both ADF- and NSM-specific HIF-1 stabilization had no effect on pumping rate in young (day 1 of adulthood) worms. In aged animals (day 12 of adulthood), however, NSM-, but not ADF-, specific HIF-1 stabilization rescued the pumping rate decline in hif-1 knockout compared to WT worms (new Fig. S1C). Similarly, NSM-, but not ADF-, specific HIF-1 stabilization rescued the thrashing rate decline in the hif-1 knockout young and aged worms (new Fig. S1D). ADF and NSM HIF-1 stabilization also had no effect on average or maximum movement speed at days 1 and 5 of adulthood (new Fig. S1E-F). Together, these results indicate that genetic activation of the hypoxic response in NSM neurons but not in the ADF neurons could improve healthspan.

      (3) While the experiments were thoroughly performed, the connections between components such as NLP-17, GABA, and tyramine in regulating aging appear critical for establishing a "cell non-autonomous circuit." Additionally, how the potentially antagonistic roles of RIM and RIS neurons influence this axis could be interesting to further explore.

      We thank the reviewer for identifying this important caveat. As described in the public review response to Reviewer 2, we completely agree that understanding the epistasis of serotonin, tyramine, NLP-17, and oxygen-sensing neuron signaling within this pathway is important to fully test our working model. Our current data showed that intestinal fmo-2 is required for neuronal HIF-1 stabilization to extend lifespan, indicating information must be communicated between the nervous system and the intestine through serotonin, tyramine, NLP-17 and responsible neurons using a “cell non-autonomous circuit” [22]. However, we will continue to address questions about epistasis and interactions between different signals in upcoming projects to fully establish the circuit.

      With respect to RIM and RIS, we value this suggestion and agree that there could be interesting signaling occurring between RIM and RIS in this circuit, as is observed in initiation of reversal behaviors [16,21]. We have updated the discussion section to mention this interesting antagonistic relationship between RIM and RIS signaling in the context of reversal behaviors:

      “Finally, while the use of RNAi knockdown and genetic knockouts establishes the necessity of many signals within the vhl-1-mediated longevity circuit, the exact directionality of these signals remains unclear. It is possible that increased, decreased, or pulsatile changes in signaling through these bioamines and neuropeptides are required for genetic activation of the hypoxic response to extend lifespan. Work on C. elegans reversal behavior has also revealed an antagonistic relationship between RIM and RIS activity facilitated by both chemical (neuropeptide and tyramine) and electrical (gap junction) signaling [16,21]. This known interaction should also be interrogated in the context of how these cells may communicate following genetic induction of the hypoxic response. Future work in this area could use tools to measure or modify neuronal activity, such as calcium imaging or optogenetics, to begin answering these questions.”

      (4) Does serotonergic neuron-specific rescue impact the mitochondrial unfolded protein response (mtUPR), given that serotonin signaling has been shown to modulate mtUPR?

      We appreciate this question and suggestion. To determine whether activating the hypoxic response in serotonergic neurons modifies the mt-UPR, we measured hsp-6 expression via qPCR in the ADF and NSM-specific HIF-1 stabilized strains. Interestingly, we find that stabilizing HIF-1 in either the ADF or NSM serotonergic neurons decreases hsp-6 expression relative to WT worms (new Fig. S1G. This could suggest either that the mt-UPR response is impaired in these worms, or that HIF-1 stabilization decreases proteotoxic stress leading to a lower basal level of hsp-6. Although this method of measurement did not allow us to interrogate whether these changes occur in the specific tissues where fmo-2 is upregulated, we have expanded our discussion of these results to emphasize that further investigation of this response should be a focus of future work.

      (5) To remain consistent with the flow of Figure 1, the authors should include fmo-2 expression in NSM:HIF-1S in addition to the ADF:HIF-1S (Figure 1I).

      We appreciate this suggestion and have added NSM::HIF-1S data to Fig. 1I. We found that stabilizing HIF-1 in either the ADF or the NSM has a similar effect on fmo-2 induction.

      (6) It would greatly strengthen the RIS observation if the authors demonstrated that ablation of another neural subtype from their screen does not abolish lifespan extension by vhl-1 RNAi. This would be a good supplemental figure, but it is not necessary for the overall story.

      We thank the reviewer for this suggestion and agree that ablating a ser-7 expressing neuron that was not a hit from our screen would be an excellent additional control. We did find that ablating a non-ser-7-expressing neuron (the RIC, Fig. 3E) did not affect vhl-1 mediated longevity, suggesting that impairing the signaling of any interneuron is not sufficient to disrupt the phenotype. In addition, ablating a ser-7 expressing neuron other than RIS would provide much stronger support for this finding. We have suggested this approach to further validate this working model in the future directions section of our discussion.

      “Finally, additional genetic controls could better support the role of RIS-specific ser-7 expression in genetic activation of the hypoxic response. For example, a ser-7 expressing neuron that was not a hit in our screen could also be ablated and tested for necessity in vhl-1 mediated longevity. This experiment would test whether the ability of RIS ablation to attenuate vhl-1-mediated longevity is not a false positive driven by any disruption to ser-7 expression.”

      (7) For consistency with the rest of the manuscript, it would strengthen the hypothesis if modulating the expression of nlp-17 or its receptors impacted the intestinal activation of fmo-2 transcription.

      This is a great point. We attempted this experiment, but were unable to achieve consistent results (see Author response image 1). This result could be due to indirect effects of the overexpression of nlp-17 signaling modulating fmo-2 induction in a complicated circuit, variability in expression of its receptor, or other complexities within the circuit.

      Author response image 1.

      (8) It would strengthen the manuscript to determine whether serotonergic, GABA, and/or tyramine signaling activate the expression or secretion of this nlp-17 neuropeptide.

      We thank the reviewer for this great idea of experiment. To address this suggestion, we performed qPCR to measure nlp-17 mRNA in WT, hif-1 KO, ADF HIF-1 stabilized, and NSM HIF-1 stabilized strains. Compared to WT and hif-1 KO controls, we observed no change in nlp-17 expression when HIF-1 was stabilized in the ADF or NSM serotonergic neurons (new Fig. S5F). This could suggest either that nlp-17 signaling acts in parallel to serotonergic signaling following genetic activation of the hypoxic response. Alternatively, neuronal HIF-1 stabilization may modify nlp-17 splicing or translation without resulting in detectable differences in mRNA levels. Together, these data indicate NLP-17 signaling is required for longevity following genetic activation of the hypoxic response, although whether this peptide is synthesized or released in response to hypoxic response remains unclear.

      We agree that it is also important to connect nlp-17 expression and/or secretion to other components of this pathway. However, we believe the most effective experiment to confirm a connection between GABA and tyramine signaling and nlp-17 in the context of hypoxia would be to measure nlp-17 expression in strains that manipulate GABA and/or tyramine signaling in a manner that mimics vhl-1 knockdown and extends lifespan. Because we have not yet validated hypoxic-response mimetics for these specific signals, we hope to first generate these strains and then measure their effect on nlp-17 expression in future work. The importance of identifying manipulations to GABA and tyramine signaling that promote longevity has been added to our discussion section. 

      “While many individual neurosignaling components are essential for genetic activation of the hypoxic response to extend lifespan, their epistasis is unclear (Fig. 7B). Most components of the pathway identified in this work act downstream of vhl-1, and upstream of fmo-2 induction (summarized in Fig. 7A-B). However, the order of each signal between these two endpoints is only predicted based on C. elegans neural wiring and the overlap between various identified signals and cells. For example, we hypothesize in our working model that GABA may be produced by the RIS neuron in this circuit because RIS is the primary GABAergic neuron required for vhl-1-mediated longevity. Alternatively, it is possible that GABA is produced by a different cell that either acts in series or in parallel with RIS signaling. In order to determine the order of each signaling component, future studies should generate genetic manipulations to each signaling component that may mimic vhl-1 knockout to promote longevity, cross these new strains into knockouts of other signals required for vhl-1 mediated longevity; and measure the lifespans of each double and triple mutant. This approach would also narrow down which signals are downstream of the genetic activation of the hypoxic response, and which are sufficient to extend lifespan upstream of the hypoxic response in a normoxic environment. One notable target for further exploration is the SER-7 expressing RIS neuron, which plays a role in sleep [16] and stress resistance [17], and can extend lifespan when optogenetically activated under normoxic conditions [18].”

      References:

      (1) Calabrese, E. J., Dhawan, G., Kapoor, R., Iavicoli, I. & Calabrese, V. What is hormesis and its relevance to healthy aging and longevity? Biogerontology 16, 693-707 (2015). https://doi.org/10.1007/s10522-015-9601-0

      (2) Zhou, I. K., Pincus, Z. & Slack, J. F. Longevity and stress in Caenorhabditis elegans. Aging 3, 733-753 (2011). https://doi.org/10.18632/aging.100367

      (3) Yuan, R., Hascup, E., Hascup, K. & Bartke, A. Relationships among Development, Growth, Body Size, Reproduction, Aging, and Longevity - Trade-Offs and Pace-Of-Life. Biochemistry (Mosc) 88, 1692-1703 (2023). https://doi.org/10.1134/S0006297923110020

      (4) Mautz, B. S., Lind, M. I. & Maklakov, A. A. Dietary Restriction Improves Fitness of Aging Parents But Reduces Fitness of Their Offspring in Nematodes. J Gerontol A Biol Sci Med Sci 75, 843-848 (2020). https://doi.org/10.1093/gerona/glz276

      (5) Duric, V., Clayton, S., Leong, L. M. & Yuan, L.-L. Comorbidity Factors and Brain Mechanisms Linking Chronic Stress and Systemic Illness. Neural Plasticity 2016, 1-16 (2016). https://doi.org/https://doi.org/10.1155/2016/5460732

      (6) Mariotti, A. The Effects of Chronic Stress On Health: New Insights Into the Molecular Mechanisms of Brain–Body Communication. Future Science OA 1 (2015). https://doi.org/https://doi.org/10.4155/fso.15.21

      (7) Bellier, A., Chen, C.-S., Kao, C.-Y., Cinar, H. N. & Aroian, R. V. Hypoxia and the Hypoxic Response Pathway Protect against Pore-Forming Toxins in C. elegans. PLoS Pathog 5, e1000689 (2009).

      (8) Palazon, A., Goldrath, W. A., Nizet, V. & Johnson, S. R. HIF Transcription Factors, Inflammation, and Immunity. Immunity 41, 518-528 (2014). https://doi.org/https://doi.org/10.1016/j.immuni.2014.09.008

      (9) Vora, M. et al. The hypoxia response pathway promotes PEP carboxykinase and gluconeogenesis in C. elegans. Nature Communications 13 (2022). https://doi.org/https://doi.org/10.1038/s41467-022-33849-x

      (10) Nakazawa, S. M., Keith, B. & Simon, C. M. Oxygen availability and metabolic adaptations. Nature Reviews Cancer 16, 663-673 (2016). https://doi.org/https://doi.org/10.1038/nrc.2016.84

      (11) Fawcett, M. E., Hoyt, M. J., Johnson, K. J. & Miller, L. D. Hypoxia disrupts proteostasis in Caenorhabditis elegans. Aging Cell 14, 92-101 (2015). https://doi.org/https://doi.org/10.1111/acel.12301

      (12) Ohh, M., Taber, C. C., Ferens, F. G. & Tarade, D. Hypoxia-inducible factor underlies von Hippel-Lindau disease stigmata. Elife 11 (2022). https://doi.org/10.7554/eLife.80774

      (13) Wind, J. J. & Lonser, R. R. Management of von Hippel-Lindau disease-associated CNS lesions. Expert Rev Neurother 11, 1433-1441 (2011). https://doi.org/10.1586/ern.11.124

      (14) Lonser, R. R. et al. von Hippel-Lindau disease. Lancet 361, 2059-2067 (2003). https://doi.org/10.1016/S0140-6736(03)13643-4

      (15) Kaelin, W. G. Molecular basis of the VHL hereditary cancer syndrome. Nat Rev Cancer 2, 673-682 (2002).

      (16) Costa, S. W. et al. A GABAergic and peptidergic sleep neuron as a locomotion stop neuron with compartmentalized Ca2+ dynamics. Nature Communications 10 (2019). https://doi.org/10.1038/s41467-019-12098-5

      (17) Wu, Y., Masurat, F., Preis, J. & Bringmann, H. Sleep Counteracts Aging Phenotypes to Survive Starvation-Induced Developmental Arrest in C. elegans. Curr Biol 28, 3610-3624.e3618 (2018). https://doi.org/10.1016/j.cub.2018.10.009

      (18) Busack, I. & Bringmann, H. A sleep-active neuron can promote survival while sleep behavior is disturbed. PLOS Genetics 19, e1010665 (2023). https://doi.org/10.1371/journal.pgen.1010665

      (19) Powell-Coffman, J. A. & Coffman, C. R. Apoptosis: Lack of oxygen aids cell survival. Nature 465, 554-555 (2010). https://doi.org/10.1038/465554a

      (20) Kruempel, J. C. P. et al. Hypoxic response regulators RHY-1 and EGL-9/PHD promote longevity through a VHL-1-independent transcriptional response. Geroscience 42, 1621-1633 (2020). https://doi.org/10.1007/s11357-020-00194-0

      (21) Bach, M., Bergs, A., Mulcahy, B., Zhen, M. & Gottschalk, A. (2023).

      (22) Leiser, S. F. et al. Cell nonautonomous activation of flavin-containing monooxygenase promotes longevity and health span. Science (2015). https://doi.org/10.1126/science.aac9257

    1. eLife Assessment

      This important study combines anatomical tracing, tissue clearing, and functional manipulations to demonstrate lateralized brainstem control of hepatic glucose metabolism and identify a site of sympathetic nerve crossover supplying the liver. The evidence supporting the anatomical organization of hepatic sympathetic innervation is compelling, and the functional studies provide solid support for a role of asymmetric sympathetic outflow in regulating glucose homeostasis. While some uncertainty remains regarding the contribution of sensory innervation and the extent to which these findings generalize beyond mice, the work provides an invaluable advance in understanding neural regulation of liver metabolism.

    2. Reviewer #2 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      The manuscript by Wang and colleagues aims to determine whether hepatic glucose metabolism is differentially regulated by the left and right sides of the LPGi and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. The authors then activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, and in changes in protein expression in the liver lobes. These data suggested lobe-specific modulation of HGP. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) Image clarity was improved in some cases, but not in others. For example, Figure 3I, showing c-Fos expression, is not convincing due to the image quality and lack of orientation.

      (2) The methods section states that 8-week-old male mice were used in the experiments without specifying the experiments (e.g., brain injection with AAVs or PRV organ inoculation). The authors should include these details.

      (3) The authors should use the exact location of pre- and postganglionic neurons, as they often refer to neurons in the sympathetic chain. Their findings should be compared with the existing literature on the location of preganglionic cells.

      (4) Figure legends should be revised and matched with the text.

    3. Reviewer #4 (Public review):

      Summary of General Strengths & Weaknesses:

      The studies here are highly informative for anatomical tracing and sympathetic nerve function in the liver in relation to glucose levels, but because they are conducted in a single species, it is challenging to translate them to humans or determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies provides mechanistically informative. Denervation studies lack proper controls, and sensory innervation in the liver is overlooked.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also labels a portion of sensory fibers that need to be ruled out in whole-mount imaging data.

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. The Discussion currently mentions this only briefly, but comparing methods and observations would be helpful.

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases in tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though that is clearly assumed. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      Comments on the revised version.

      Across all reviewer comments, the revised resubmission has adequately addressed all concerns.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether hepatic glucose metabolism is differentially regulated by the left and right sides of the LPGi and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi, which were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, and in changes in protein expression in the liver lobes. These data suggested lobe-specific modulation of HGP. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) Image clarity was improved in some cases, but not in others. For example, Figure 3I, showing c-Fos expression, is not convincing due to the image quality and lack of orientation.

      We sincerely apologize for the insufficient image clarity and anatomical orientation in the original Figure 3I. To resolve this issue, we have performed the following revisions in the revised Figure 3I:

      (1) Replaced the original panels with the high-resolution confocal images showing clear c-FOS immunofluorescence in the LPGi.

      (2) Included explicit anatomical orientation indicators (Bregma −6.75 mm) to clearly demarcate the boundaries of the LPGi.

      (3) Added ROI outlines surrounding the LPGi region.

      (2) The methods section states that 8-weeks-old male mice were used in the experiments without specifying the experiments (e.g., brain injection with AAVs or PRV organ inoculation). The authors should include these details.

      We thank the reviewer pointing out this oversight. We have updated the Methods section under "Animals" and specific procedure subsections to clearly state the exact age of animals.

      (1) For retrograde trans-synaptic PRV tracing, 8-week-old mice received intrahepatic viral injections and were sacrificed 5 days post-injection.

      (2) For chemogenetic and optogenetic manipulations, stereotaxic AAV injections were performed at 8 weeks of age. Mice were allowed 4 weeks for viral expression and recovery before undergoing metabolic tests or light stimulation at 12 weeks of age.

      (3) For chemical denervation (6-OHDA), 8-week-old mice were injected into targeted lobes and examined 7 days post-denervation.

      (4) For postnatal innervation mapping, neonatal mice at postnatal week 0 (P0), week 1 (P7), and week 2 (P14) were harvested for tissue clearing.

      (3) The authors should use the exact location of pre- and postganglionic neurons as they often refer to neurons in the sympathetic chain. Their findings should be compared with the existing literature on the location of preganglionic cells.

      We appreciate the reviewer for this feedback. We agree that our original description lacked precise anatomical localization regarding the pre- and postganglionic neurons, and it was inaccurate to state that descending fibers pass through the sympathetic chain (SyC).

      Based on our whole-mount tissue clearing data, we observed that the preganglionic neurons of the brain-liver sympathetic circuit are primarily located in the T6–T12 segments of the thoracic spinal cord. Accordingly, we have revised the text in Results 4 to specify these exact locations.

      Manuscript Revision (Results 4):

      "Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the thoracic spinal cord (T6–T12) send descending fibers via the splanchnic nerves to innervate postganglionic neurons in the CG-SMG (Figure 4A)."

      Furthermore, following your valuable suggestion to compare our findings with existing literature, we reviewed a recent study published in Nature Communications (Harima, Yukiko et al. Parallel labeled-line organization of sympathetic outflow for selective organ regulation in mice. Nat Commun. 2024;15(1):10478). In that study, researchers injected retrogradely transducible AAVs directly into the CG-SMG and traced the preganglionic neurons predominantly to the T8–T13 segments. Their results are largely consistent with our findings. Interestingly, the broader anatomical range observed in our trans-synaptic liver-to-brain mapping (T6–T12) compared to their CG-SMG-specific tracing (T8–T13) reveals a slight discrepancy. This observation suggests an intriguing anatomical hypothesis: a subset of sympathetic preganglionic nerves may bypass the CG-SMG relay entirely and project directly to the liver.

      (4) Figure legends should be revised and matched with the text.

      We apologize for the oversight. We have conducted a comprehensive audit of all figure and legends to ensure precise matching between the main text and the figures.

      Specifically, we have corrected a typographical error in the Figure 1 Legend where panel (C) was mistakenly labeled as a second panel (B), and we fixed a spelling error ("LPG" corrected to "LPGi"). Additionally, we corrected a miscitation in Results (Section 3) regarding Figure 3. In the original text, Figure 3C was incorrectly grouped with blood glucose data, whereas it actually displays the Western blot validation of sympathetic denervation.

      We have revised the corresponding sections in the manuscript as follows:

      Manuscript Revision (Figure 1 Legend):

      “(C) Quantification of PRV-labeled neurons in left and right LPGi across different hepatic lobes: left lateral, median, right posterior, right anterior, caudate, and porta hepatis (n = 3).

      (D) Sankey diagram showing projection patterns from left and right LPGi to individual hepatic lobes. (E and F) Representative slices of EGFP+ and mRFP+ neurons in left (top) and right (bottom) LPGi following PRV-EGFP (right anterior lobe) and PRV-mRFP (median lobe) injections. Proportions of EGFP+, mRFP+, and co-labeled neurons in left and right LPGi (F, n = 3). Scale bars, 100 μm.”

      Manuscript Revision (Results 3):

      “Despite the absence of directly sympathetic input to denervated lobes, systemic blood glucose levels were unchanged compared with controls (Figures 3A-3B, Figure S5A), indicating functional compensation through the remaining intact liver.”

      Reviewer #4 (Public review):

      Summary of General Strengths & Weaknesses:

      The studies here are highly informative for anatomical tracing and sympathetic nerve function in the liver in relation to glucose levels, but because they are conducted in a single species, it is challenging to translate them to humans or determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies provides mechanistically informative. Denervation studies lack proper controls, and sensory innervation in the liver is overlooked.

      We sincerely thank the reviewer for their time and evaluation. We respectfully note that these comments mirror those raised during the previous round of review. We would like to kindly direct the reviewer to the extensive revisions we implemented in our previous resubmission, which directly and comprehensively addressed these exact concerns. These revisions remain intact in the current version of the manuscript. Below, we briefly summarize how each point was previously addressed for your convenience.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      As addressed in our previous revision, we fully agree with this suggestion. We updated the title of the manuscript to explicitly include the species: "Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice." We also clarified the species used throughout the main text to ensure accuracy.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also labels a portion of sensory fibers that need to be ruled out in whole-mount imaging data.

      As detailed in our previous response, we acknowledge this important limitation. In our prior revision, we addressed this concern through both additional data analysis and text revisions:

      (1) We provided SyGlass 3D reconstruction data demonstrating that the TH-positive nerve fibers originate from the celiac-superior mesenteric ganglia (CG-SMG), a well-established sympathetic ganglion (Figure S5F).

      (2) In parallel, we collected dorsal root ganglia (DRG) from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While the T7-12 DRG segments are historically known to contain the sensory neurons that innervate the liver (Anat Rec A Discov Mol Cell Evol Biol. 2004; Auton Neurosci. 2024), we detected only a remarkably sparse number of PRV-positive neurons in these segments. This effectively functionally distinguishes this efferent pathway from primary sensory afferents (Supplementary figure B).

      (3) We explicitly added this methodological limitation to the Discussion section (paragraph 6) of the current manuscript, noting that more selective approaches, such as genetic targeting of sympathetic lineages, will be important for future validation."

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      As outlined in our previous response, we incorporated this crucial context into our revised manuscript. Specifically, we expanded the Discussion section (paragraph 3) to contrast our precise cell-type-specific chemogenetic and optogenetic approaches with historical studies that relied on coarse electrical stimulation. This addition highlights how our current methodology reveals the contralateral and lobe-specific architecture of brain-liver sympathetic control that was previously obscured.

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases in tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though that is clearly assumed. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We appreciate this insightful physiological perspective, which we addressed comprehensively in our previous revision. As we previously agreed, the central nervous system integrates a broad range of afferent signals, and compensatory sensory or parasympathetic mechanisms likely contribute to the observed LPGi activation following hepatic sympathetic denervation.

      To address this, we significantly expanded our Discussion section (paragraph 4) in the prior revision. We explicitly proposed a model wherein hepatic glucose production is regulated by an integrated afferent-central-efferent loop, acknowledging that our current study primarily resolves the efferent component. We clearly noted the lack of direct assessment of sensory or parasympathetic innervation as a limitation and highlighted this dynamic crosstalk as a critical avenue for future investigation.

      Comments on the revised version.

      Across all reviewer comments, the revised resubmission has adequately addressed all concerns.

      Recommendations for the authors:

      Reviewer #4 (Recommendations for the authors):

      No further recommendations aside from tempering the CGRP language, as marking all sensory fibers.

      We appreciate the reviewer for pointing out this important anatomical distinction. We entirely agree that CGRP specifically labels peptidergic sensory afferents and does not represent the entirety of the sensory nervous system.

      We have carefully reviewed the entire manuscript and tempered our language accordingly. Wherever CGRP is mentioned, we have clarified that it serves as a marker for peptidergic sensory fibers, rather than functioning as a pan-sensory marker.

      Manuscript Revision (Results 1):

      "Unlike the NTS, a well-established hepatic sensory center served here as a positive control, the LPGi contained few CGRP-positive cell bodies (Figure S1G), indicating a lack of peptidergic sensory projections."

    1. eLife Assessment

      This important study links blood-derived dietary content to sustained increases in sleep in the mosquito Aedes aegypti. Using multiple independent approaches, the authors provide convincing evidence for blood-induced changes in sleep. These findings have broad implications for understanding how specialized diets regulate sleep across species and for mosquito vector biology.

    2. Reviewer #1 (Public review):

      Summary:

      The presented investigation aims to expand the sleep definition and its relationship with blood meal and/or circadian clock in the mosquito, Aedes aegypti. The authors exhausted the established sleep analytical paradigm and three behaviour toolkits: LAM10, EthoVision, and DART. They also investigated the potential underlying molecular mechanism by using dsRNA injection (LkR) and KO mosquito (Cyc-/-).

      Strengths:

      The authors presented a very solid dataset showing posture changes and increase in the arousal threshold of mosquito after 10 minutes of immobility. This is major clarification and extension to our understanding in insect sleep beyond Drosophila. Inclusion of analytical parameters such as bout length, waking activity and pDoze/Wake provide critical reminder for other investigators of the steps needed for defining sleep in a new species. The investigation, with its technical span in behaviour assays, therefore, establish a good standard for mosquito sleep analysis to the same quality seen in the landmark studies (Shaw et al 2000 and Hendricks et al 2000) for Drosophila sleep. The pioneering data showing clear effect of blood meal and LkR reduction on locomotion and sleep provides an entry point for further investigations. The author has addressed previous concern on coincidence of sleep increase and locomotion reduction by using their two high-res. video tracking velocity or pDoze/Wake, showing that the "sleepy" mosquitos remain capable to reach high speed locomotion albeit less frequently. The authors also discuss the possibility of ATP and alternative explanation regarding sugar content in diet.

    3. Reviewer #2 (Public review):

      Summary:

      Zhang et al. investigate how blood feeding and dietary protein influence sleep in the mosquito Aedes aegypti. The authors first establish a behavioural definition of sleep using postural analysis and arousal threshold measurements, then demonstrate that both blood meals and a bovine serum albumin (BSA)-based protein diet increase sleep for several days. They further show that RNAi-mediated knockdown of the leucokinin receptor (Lkr) enhances sleep, implicating neuropeptide signalling in the regulation of postprandial sleep.

      Strengths:

      The central question is well-motivated, and the experimental approach is systematic. The use of multiple independent methods to characterise sleep - postural analysis, infrared activity monitoring, videography, and arousal threshold - provides converging evidence. The 10-minute immobility criterion is grounded in the arousal threshold data, bouts exceeding 10 minutes corresponding to the first bin at which a significant effect emerges. The demonstration that the sleep increase is already detectable before oviposition establishes that the phenotype begins with feeding rather than with the completion of the reproductive cycle. The BSA feeding experiment is a particularly effective demonstration that dietary protein, rather than other blood components, is a key regulator of the sleep increase. The conservation of leucokinin signalling in sleep regulation between Drosophila and Ae. aegypti is a noteworthy finding that adds comparative depth. The "opportunistic versus determined" host-seeking distinction is appropriately framed as a hypothesis for future testing rather than as a conclusion drawn from the present data, and the limits of the design with respect to reproductive physiology are stated explicitly.

      Weaknesses:

      (1) Confound of reproduction and sleep. Blood and BSA both support egg development, so neither condition isolates nutrient sensing from reproductive physiology. The relative contributions of diet, egg development and post-reproductive recovery remain undetermined.

      (2) Sleep versus reduced locomotion. The pDoze and pWake measures are defined here as proportions of time above or below a velocity threshold, rather than as the per-minute transition probabilities of the established definition (Wiggin et al. 2020, PNAS). So defined, they are equivalent to percent sleep and percent wake and cannot distinguish a sleep-like state from the mechanical consequences of engorgement.

      (3) Data availability. Raw data are stated to be available on request rather than deposited in a public repository, which makes independent reanalysis less straightforward than it need be.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We thank both reviewers for their thoughtful and constructive evaluations of our manuscript. We are grateful that both reviewers found the study to provide a strong behavioral framework for defining sleep in Aedes aegypti and appreciated the breadth of the behavioral and genetic approaches used. We also appreciate the reviewers’ careful identification of several issues requiring clarification, particularly regarding the interpretation of post-blood-meal sleep, the support for the 10-min sleep threshold, possible nutritional confounds in the BSA experiments, the framing of the host-seeking model, and the description of statistical analyses. In the revised manuscript, we have addressed these concerns by clarifying our rationale, tempering several conclusions, revising the statistical reporting and methods, explicitly stating sample sizes, and expanding the Discussion to better acknowledge limitations and alternative interpretations. Where appropriate, we have also revised the text to distinguish more clearly between increased sleep and reduced locomotion, and to frame mechanistic conclusions more cautiously.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Conventionally, a coincidence of sleep increase and locomotion reduction would weaken the certainty of a sleep increase assessment. The authors implied this concurrence observed after blood meal is derived from internal "drowsy" neural state instead of physical "cripple", but they did not use their two high-resolution video tracking velocity or pDoze/Wake to clarify this.

      Thank you for addressing this point. We understand the need to validate locomotion when used as a readout of sleep. We note that analysis of waking activity is normalized to time spent awake, and therefore should be separate from the time spent inactive that is classified as sleep. Based on the reviewers’ suggestions we have reanalyzed some data and revised the relevant sections in include this analysis.: In brief we performed pDoze/pWake analyses on the two high-resolution tracking video from EthoVision XT system. A velocity threshold of 0.4 mm/s was used, with velocities above 0.4 mm/s defined as wake/activity and velocities below 0.4 mm/s defined as doze/sleep state. pWake and pDoze were defined as proportional time metrics of wake/active (velocity > 0.4 mm/s) and doze/sleep (velocity > 0.4 mm/s) within each LD cycle. The conclusion that sleep is increased following blood feeding is supported by these data. We also note (as described in response to Reviewer 2, that this paper represents a step towards describing sleep in mosquitoes. We hope that future application of approaches used in Drosophila, such as brain imaging and indirect calorimetry will further refine our understanding. Along these lines, we have also included a section in the Discussion about how additional measures, including systems like FlyVista might be applied in the future.

      (2) The major molecular component underlying blood meal effect on sleep/locomotion is less certain, because the BSA solution used for feeding contains ATP, which itself is able to enter haemolymph and potentially exerts sleep/locomotion effect. Additionally, the basal or control sleep recording is done after sucrose feeding. It is, however, unclear from the method if this is 10% too? And if the observed sleep level increase after a blood meal is a result of sugar level reduction in the blood (~0.1%).

      We thank the reviewer for raising this important issue. We think it is unlikely that the small amount of ATP used for feeding is driving the sleep phenotype. We have now included this point as a caveat within the discussion, and explained its inclusion.

      (2) Sucrose concentration in controls

      We apologize that this was not clearly stated. Yes, the control mosquitoes were maintained on 10% sucrose, and we have now clarified this explicitly in the Methods and figure legends where relevant.

      (3) Could the effect reflect reduced sugar intake rather than blood/protein?

      We think this is unlikely, however it cannot be ruled out based on the experiments we have run. We have added discussion of this point. However, we note in fruit flies, this has been studied extensively, and loss of sugar under certain contexts reduces sleep. The points above highlight the need for systematic analysis of the dietary components that contribute to sleep in mosquitoes. While we regret being unable to include them in this manuscript, we note that many of these experiments are challenging (with many controls) and have been ongoing for over a decade (with contributions from many labs) in Drosophila.

      Reviewer #2 (Public review):

      (1) The authors settle on a 10-minute immobility threshold, but their own data do not convincingly support this choice… A 15-minute threshold would be better supported by the data as presented.

      We appreciate this evaluation of the sleep threshold. We chose 10 minutes because the first significance in arousal threshold is at the time-point of 10-15 minutes. Therefore, we believe that sleep bouts longer than 10 minutes should be qualified as sleep. We are particularly interested in why arousal threshold continues to increas at 15 minutes. This is either incomplete sleep between minutes 10 and 15 or the presence of multiple sleep states. We have established a new system in the lab using Zantiks that we believe will allow for simultaneous recording of posture and arousal threshold. We now explicitly comment on this in the discussion, and the need for further analysis of the timeframe for which sleep is defined. Nevertheless, we believe we have honed in on a period of 10-15 minutes that serves as a good proxy for sleep regulation. We hope that this initial description of sleep in mosquitoes provides an initial step towards defining sleep, and that future studies that include techniques applied in Drosophila including brain imaging, indirect calorimetry and additional videography will define more nuanced changes in sleep. We have written in limitations and future opportunities to better define sleep throughout the manuscript.

      (2) The primary experimental paradigm measures sleep beginning at Day 4 post-blood feeding, immediately after oviposition... what is being measured as ‘sleep’ could reflect post-reproductive quiescence or recovery rather than diet-induced sleep per se. The BSA experiment partially addresses this, but since BSA also triggers vitellogenesis and egg production, the confound persists.

      We agree this is an important concern. Our intent in measuring sleep after oviposition was to isolate prolonged post-feeding effects from the well-established transient suppression of host-seeking that occurs during the first ~72 h after blood feeding. However, as the reviewer notes, this design does not by itself distinguish post-feeding sleep from other physiological processes associated with reproduction, including vitellogenesis, oviposition, or post-reproductive recovery. To address this issue, we included the experiment measuring sleep immediately after blood feeding, before oviposition. We agree, however, that this rationale should have been stated more clearly and that the limitation remains relevant, particularly because BSA can also support egg development. In the revised manuscript, we have therefore: In the current version we have clarified more explicitly that the immediate post-blood-meal recording was included to show that the sleep increase begins before oviposition; We have also tempered our interpretation of the Day 4–5 phenotype to avoid implying that it is purely diet-driven and fully independent of reproductive state; and expanded the Discussion to acknowledge that blood feeding, protein feeding, and reproductive physiology are closely linked in female mosquitoes and that our current experiments do not fully disentangle these processes. These changes frame the data more cautiously: blood/protein feeding is sufficient to induce a sleep-promoting state that begins immediately after feeding and persists into the post-oviposition period, but the relative contributions of nutrient sensing, egg development, and reproductive recovery remain to be determined.

      (3) The opportunistic vs. determined host-seeking hypothesis… requires actual measurement of host-seeking alongside sleep to be substantiated, or at least the caveats need to be discussed more explicitly.

      We agree with the reviewer. Our intention was to present this as a conceptual model motivated by the temporal dissociation between published host-seeking recovery and the prolonged sleep phenotype observed here, not as a demonstrated behavioral framework directly tested in this study. In the revised manuscript, we have substantially softened this section by clarifying that we did not directly measure host-seeking behavior in the current study; adding explicit caveats that the proposed framework remains speculative until sleep and hostseeking are measured simultaneously in the same animals across the same post-feeding time course. We appreciate this comment and agree that the distinction should be presented as a model for future testing rather than as a central conclusion established by the current data.

      (4) The methods describe ‘one-way ANOVA, followed by Mann-Whitney tests with Welch’s correction,’ which is an internally inconsistent combination…

      We thank the reviewer for catching this lack of clarity. We apologize for this inconsistency. We have fixed this error. In the revised manuscript, we have carefully rewritten the statistical analysis section to specify: which datasets were analyzed using parametric tests (e.g., ANOVA, with appropriate post hoc comparisons where assumptions were met), which datasets were analyzed using non-parametric tests (e.g., Mann-Whitney), and where Welch’s correction was applied, specifically for unequal-variance t-tests, not Mann-Whitney tests. We have also revised Methods, Figure legends and reporting throughout to ensure that the statistical test named in the text matches the reported test statistics. The changes include statistical methods rewritten for consistency and accuracy, and updated figure legends that include exact sample sizes. In addition, one summary spreadsheet of statistical analysis throughout this study is provided and will be submitted as a supplementary file.

    1. eLife Assessment

      This important study demonstrates how individual taste preferences change over time, how these changes are reflected in cortical activity, and how sensory experience contributes to reshaping both. The evidence is convincing and broadly supports the main conclusions. The findings should be of interest to neuroscientists studying sensory processing and cortical plasticity.

    2. Reviewer #1 (Public review):

      Summary:

      Maigler et al. set out to test the hypothesis that individual differences in taste preferences are (in part) due to individual differences in central taste processing. They first tested rats' preferences for a variety of taste stimuli on multiple days. They then recorded responses of neurons in taste cortex to the same tastes on two consecutive days.

      Strengths:

      The authors collected high-resolution behavioral data from the same animals across multiple days, allowing for a detailed characterization of individual variation in taste preferences. They then performed recordings from the same set of animals in response to the same stimuli, allowing them to draw parallels between behavioral and neural responses.

      Weaknesses:

      (1) The authors collect extensive behavioral data and show that preference vary between animals and days, but little insight is provided into what underlies these changes and to what extent they reflect "preference". Two animals drank equal amounts of sucrose and quinine on day one of preference testing, suggesting that behavior does not reflect preference but (lack of) habituation to/proficiency with the testing environment.

      (2) Recordings were performed only after multiple days of preference testing, and preferences were not tested in between/following recording sessions. This design precludes a direct comparison between neural and behavioral responses.

      (3) Similarly, correlations between neural responses and behavioral measures are not analyzed/reported on an animal-by-animal basis.

    3. Reviewer #2 (Public review):

      Summary:

      The study from Maigler et al investigates how between- and within-animal differences in taste preference relate to differences in neural responsiveness. The experiments rely on an elegant combination of behavioral assays to measure preference (e.g., repeated brief access testing, BAT) and electrophysiological recordings to monitor the activity of ensembles of neurons in the gustatory cortex (GC) of rats.

      BAT with distinct batteries of tastants revealed pronounced variability in preference (measured as licking bout size) across individuals. This variability across individuals persisted after repeated testing. Repeated BAT also revealed that each individual rat's preference for different tastants changed across time.

      Electrophysiological responses of GC neurons to batteries of tastants showed that firing in the "late epoch" of taste processing (i.e., 500ms post taste delivery) correlated more strongly with the individualized rat's BAT preference rather than with a canonical preference ranking. Importantly, this correlation was stronger for the last BAT session compared to the first. Finally, the authors show that the correlation disappeared in a second, consecutive recording session, indicating that exposure to tastants reconfigure preferences.

      Strengths:

      (1) The experimental design allows for an unprecedented look at the relationship between individual variability in taste preferences and neural processing.

      (2) The study demonstrates that taste preference variability is not mere experimental noise but reflects the dynamic nature of taste. A key strength is the clear evidence that behavioral variability is reflected in neural activity patterns, establishing a strong correlation between brain and behavior.

      (3) The evidence that simple exposure to familiar tastes can reconfigure preferences and taste representations is interesting.

      Weaknesses:

      The authors appropriately addressed the weaknesses in the revision process.

    4. Reviewer #3 (Public review):

      Summary:

      Maigler & Lin et al present a convincing set of behavioral and electrophysiological experiments and analyses exploring how individual differences in taste preference map onto neural responses in the gustatory cortex (GC). They go on to examine how both preferences and neural responses shift following intervening taste experience. Their experiments are strengthened by examining tastes of distinct identities and palatability (sweet, sour, salty, bitter) and correspond each animal's individual preference to the palatability-related late phase of the neural response.

      Strengths:

      (1) They demonstrate a relationship between the behavioral expression of taste preference and palatability-related GC neural responses. The direct correlation of expression of taste preference with GC neural responses indicates that taste preference behavior may be less noisy than previously thought, reflecting actual neural activity.

      (2) They address the stability of individual taste preference by comparing within and between session expression. This finding indicates that individual preference on any given trial or test session can differ from canonical palatability.

      (3) The animal's preferences for various tastes are reflected in the neurophysiological recordings, despite that behavior and physiology sessions are separated by several weeks, with an intervening surgery, across multiple delivery methods (active licking vs passive delivery), and across homeostatic states (thirsty vs sated).

      (4) They provide evidence that representational drift in palatability coding may arise from sensory experience rather than from the passive passage of time.<br /> The findings are novel and impactful, and the results are relatively complete.

      Weaknesses:

      (1) The authors state that "no differences in effects were observed between taste batteries" (Methods), but it is not clear which analyses were performed to determine this, especially considering that many of the analyses are within-animal. Without more clarity, it is difficult to evaluate whether the interaction of different tastes within the sets of stimuli bias the main conclusions.

      (2) It is not clear how the analyses in Figure 3 emerge from the behavioral patterns across Figures 2A2 and 2B2. Along these lines, citric acid responses for R1 and R2 and saccharine responses for R7 and R8 are not shown in Figures 2A2 and 2B2, thus, we cannot determine if they change from the first to final BAT sessions. Salt, even at low concentrations, may not be palatable when rats are in a water-restricted state. It would be interesting to see Figure 3's analysis exclusively for sweet or for bitter tastes.

      (3) A potential reason why a single taste experience session causes taste palatability responses in GC neurons to revert to a reflection of canonical palatability ranks, as is seemingly shown in Figure 6, is not discussed.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear whether there are any systematic changes in preferences over the course of testing that could explain the observed changes in correlation with neural responses, such as changes due to learning (e.g., flavor nutrient conditioning, relief of neophobia), changes in deprivation state, or habituation to/proficiency with the BAT setup.

      For the revision, we have added analysis, including a new figure (Figure 3) between what are now Figures 2 & 4, testing the hypothesis that preference changes across testing days are non-random in direction (e.g., that they reflect attenuation of neophobia). This new analysis failed to reveal evidence supporting the hypotheses that: 1) preference for palatable tastes increases with experience (a result that would make sense given research on neophobia; 2) the preference for aversive tastes decrease with experience; or 3) absolute consumption of any particular taste changes in a reliable direction from session to session (lines 142-157 and new Figure 3).

      A secondary point is whether any changes in preference are attributed to internal individual versus external contextual factors. Both types of variation (i.e., across individuals and across time within an individual) are mentioned in the introduction, but it is not clear what the authors believe about the nature or neural representation of these sources of variation.

      While we assume that differences between rats are due to internal factors (given the controlled home-cage environment), we can’t be sure that some subtle, subthreshold (for us as observers) factor impacts taste preferences. Similarly, while changes across time within an individual is categorically within the individual, we cannot be sure whether some subtle facet of their experiences determines how preferences change (as opposed to it being purely internal). We have added prose to the Discussion session on this topic—including citation of Hilary Schiff’s recent work showing nurture-related preference changes as part of this new prose (lines 387-398).

      With respect to neural data analysis, no individual animal/day data are shown, making it difficult to assess the extent to which differences in correlation match individual differences in preferences and/or changes in preference with time within individuals.

      The revision now explicitly includes Figure panels (with analysis) showing the relationships between individual neural responses and consumption in the first and last BAT tests for a representative rat (lines 172-198; Figures 4A and 4D). As requested in the non-public comments, we have also added waveforms recorded for the representative neuron in an inset to Figure 4B.

      The correlation analysis is also lacking control for the fact that there is a certain degree of "chance" associated with behavioral and neural measures having matching ranks.

      Certainly chance cannot explain our results, which consist centrally of within-rat differences in match (that is, regardless of chance match levels, what we observed was specifically an enhancement of that match for the most recent behavioral assessment compared to an earlier assessment in the same rat)—a finding that is all the more surprising given that: 1) 2 weeks separate that behavior test and the electrophysiology session; and that 2) that gap between the ephys test and the (less well-matched) first behavioral test is only 1-3 days longer. Nonetheless, in appreciation of Reviewer 1’s concern, we have added an independent, convergent analysis to the revision, testing whether the observed pattern vanishes when we shuffle the preference ranks between tastes with neighboring ranks in the behavioral data (a more conservative test than complete shuffles among tastes). The results of this analysis, which are in the new Figure 5, provide further proof that our result is not based on chance—that they specifically reflect a match between neuronal activity and behavior (lines 242-251).

      Finally, …it is unclear to what extent changes in correlation may be attributed to overall changes in responsiveness of the neural population.

      We include several new analyses in the revision that test the hypothesis that the reduction in match between behavioral rankings and neural responses in the second electrophysiology sessions reflects spontaneous or taste-driven changes in neural excitability. These additional analyses reveal no clear between-session differences in baseline and/or taste-evoked responses, or in the percentages of neurons that are taste responsive and/or palatability-related (lines 292-309; Figure 7).

      Reviewer #2 (Public review):

      The manuscript could use additional corollary analyses to provide a more complete picture of the phenomenon. For instance, how many neurons (per animal and in total) have significant correlations with the final BAT patterns? And with the first BAT? Can a time course of such counts be provided? Can some decoding analyses be performed at a single session level to reconstruct a rat's behavioral preference pattern from its neural activity?

      These are all really good ideas. As noted in our response to Reviewer 1, we have implemented all but the last of the suggested analyses, which did not produce evidence suggesting that our results can be explained by changes in neuronal properties between the two recording sessions (lines 292-309; Figure 7). We have also made attempts to apply the decoding analysis; unfortunately, we don’t have large enough samples to obtain stable results such a subtle decoding task (reflecting the last BAT session’s preference pattern is significantly better than the first session’s pattern).

      The manuscript could benefit from additional polishing, both in the text as well as in the figures.

      An extensive holistic edit has been done, starting with suggestions made by Reviewer 2 in the non-public comments.

      Reviewer #3 (Public review):

      Without a behavioral measure collected after recording day 1 intraoral exposure, it is not possible to determine whether taste preference was altered by that experience…The authors' conclusion would be strengthened by adding an intervening brief access test between recording days 1 and 2.

      We very much appreciate Reviewer 3’s suggestion. Alas, the primary authors involved in data collection on this project have moved on, and we won’t be able to collect the additional dataset that would be required. Instead, we have softened the conclusion that we reached in the last section, and suggested the proposed experiment as a future direction (lines 366-374).

      The current experimental design exposes animals to 3 distinct sets of substances … [that] differ in identity … and concentration. Because palatability is known to be comparative depending on the other substances available and concentration-dependent, this introduces challenges to interpretation, [and] without more clarity, it is difficult to evaluate whether the interaction of different tastes within the sets of stimuli biases the main conclusions.”

      This is an interesting point. Analyzing each set of batteries separately and performing between-battery comparisons would require a larger number of experimental subjects then we have in our current sample size. That said, while we acknowledge that taste preference ranking is relative, we believe the ranking system used here deviates little, if any, from the 'true' ranking (and is therefore significantly relevant to gustatory activity). This is supported by our newly obtained result in response to Reviewer 1 & Reviewer 2 (see above), where an ancillary shuffle analysis (Figure 5C) showed that swapping adjacent preference orders eliminated the experimental effects across all batteries.

      Responses to sweet tastes are not reported in the electrophysiology data. This is seemingly the case because rats given set 1 received no sweet stimulus while rats given set 2 received to 2 distinct sweet tastes. Finally, rats given set 3 did not receive quinine, yet quinine is reported in electrophysiology data.

      We are unsure of the source of this confusion—in every case, the rat received the same tastes in the electrophysiology sessions that were delivered in the BAT preference tests—but in appreciation of Reviewer 2’s concern, we have modified the text and table to ensure: 1) that panels reflecting data from single example rats (panels that therefore necessarily include only a subset of possible tastes) are clearly marked as such; and 2) that the nature of which taste batteries were delivered is more explicit (lines 104-112; 172-178).

      The choice of reporting average lick cluster size is problematic because the authors use thirsty rats with 10-second-long trials. Thirsty rats are likely to lick in relatively long clusters, especially for neutral and palatable tastes. If the rat is mid-cluster when the trial ends, the final cluster would be cut off prematurely, resulting in shorter overall average lick cluster size, disproportionately affecting neutral and palatable tastes over aversive tastes.

      We have ourselves been deeply concerned with this issue, and in fact have recently published a paper that includes within it a direct test demonstrating that calculations of lick bout lengths from 10-sec BAT trials result in taste palatability estimates that are identical to (and less noisy than) those generated from more classically-used 15-min ad lib licking. We now cite this paper (Stone, Lin, et al., 2026) in the Methods section, along with text clarifying how we calculated lick clusters. We also conducted an additional analysis that estimates taste preference after removing these “prematurely ended bouts” without changing the observed pattern of results (lines 494-510).

      Of course, even if this last analysis had changed things, the result of clusters being cut short by the end of a trial would be an underestimation of the preference for the palatable tastes (which drive far more licking than aversive tastes and are therefore more likely to be mid-bout at the end of a trial). Such an underestimation would in turn be expected to reduce the observed neural-behavioral correlation. This fact highlights the robustness of our findings.

      Canonical palatability rankings may not apply to the concentrations selected in every stimulus set. This is particularly true for set 1, which included two concentrations of citric acid and quinine for the behavior. It is also not clear which concentrations are reported in Figures 3A2 and 3B2. Meanwhile, the concentrations of quinine and citric acid used for electrophysiology are quite low.

      In the revised Methods section, we explicitly motivate our reasoning (including citations) behind canonical rankings for each taste battery used (lines 513-522). Every taste used was of agreed-upon preference levels, and in the rare case that two concentrations of the same taste were used, both were known to have distinct palatabilities (e.g., 0.1M NaCl is preferred to 0.05M NaCl). This careful selection of tastes ensured that it was trivial to avoid misordering of canonical palatability rankings.

      And even if mistakes in canonical rankings had been made, the impact of these inaccuracies in these rankings would have been minimal. Our findings are primarily driven by high levels of inter-individual (between different rats) and intra-individual (day-to-day fluctuations within the same rat) preference differences. Given this variability, the fact that the brain-behavior correlation were consistently worse using these rankings almost certainly means that the neural activity matches preference behavior—our thesis. This conclusion is further supported by our shuffle analysis, which demonstrated that randomizing the order did not yield superior correlations between taste ranking and GC activity.

    1. eLife Assessment

      This useful study introduces MULTI i2, a robust and high-throughput method to measure Plasmodium falciparum viability in the presence of drugs. This new assay has the premise to offer significant time and cost savings over the traditional Parasite Reduction Rate (PRR) assay and should enable faster screening of drug combinations, which is urgently needed in the field. The assay is well validated, with convincing data showing it can reproduce known drug interactions and identify new interaction patterns, but the study falls short of demonstrating broad applicability.

    2. Reviewer #1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRRv2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify time-dependent onset of parasiticidal action compared with the PRRv2.

      This is a very well-executed and well-described body of work with a comprehensive set of analyses. The revised manuscript provided more context in comparing this method to other methods in the field, with a clearer explanation that the current MULT-i2 assay focuses on assessing viability, whereas other methods are more focused on assessing growth inhibition. The new assay is also faster than the prior PRR methodology. The current design will be useful to stay multi-drug combinations. The authors state that this parasite line will be available from BEI with an accompanying MTA. Other earlier comments and concerns were very well addressed by the authors.

    3. Reviewer #2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRRv2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compare their MULT-i2 assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i2 assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i2 assay.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.<br /> Measurement of parasite viability in the MULT-i2 assay was achieved by extrapolating chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (eg ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRRv2 assay?

      The addition of an inducible element is an improvement of their earlier lacZ/β-galSENSOR (PMID: 41575867), however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRRv2, they fail to compare it to their own non-inducible lacZ/β-galSENSOR system. Their non-inducible system already showed superiority to the cPRR assays and it would be good to show how they compare and what the advantages of the new system are over the old. Eg how is the signal to noise improved? How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h but it would be very useful to the community to look at earlier timepoints as well and compare it to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h and 24h).

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      Comments on revised version:

      The authors adequately addressed our comments and the resulting manuscript describes a specialized resource for antimalarial drug development.

    4. Reviewer #3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULTI-i2, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i2 assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULTI-i2 assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULTI-i2 provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULTI-i2 methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1) The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS) e.t.c.? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      (2) Related to that above, how would MULTI-i2 perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      (3) Given the stated cost and labor efficiency of MULTI-i2, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i2 method more attractive. In particular, it would be nice to see if one could use MULTI-i2 for studies of triple combinations as enthusiastically suggested.

      (4) Throughout the manuscript, the authors claim that MULTI-i2 is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      Comments on revisions:

      I have no more comments.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [<sup>3</sup>H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we have also explicitly stated this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we mention and describe alternative approaches for assessing parasite viability. This also includes the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We have revised the text and now explicitly mention the use of dual staining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al., 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non-commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer #2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We have clarified this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We have added this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post-antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We have revised the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      The authors observed differences between their i-lacZ assay and conventional PRR assays attributable to the accumulation of lacZ enzyme at higher levels of surviving parasites, followed by degradation. Have the authors tested how long lacZ remains stable in standard or overgrown parasite cultures?

      We thank the reviewer for this important question. We assessed the stability of β-galactosidase activity in parasite lysates stored under different conditions and observed that the enzymatic activity remained stable for up to 21 days when lysates were stored at either -20°C or 37°C (Hellingman et al., 2024).

      We have not systematically characterized β-galactosidase stability in standard or overgrown parasite cultures. However, in experiments involving overgrown cultures, we observed that the β-galactosidase-derived signal decreased rapidly in overgrown culture settings, suggesting that enzyme stability in overgrown cultures is lower than in standard cultures and parasite lysates.

      At what parasitemia were the counts reported in Figure 1D conducted at?

      We thank the reviewer for this clarification request. The measurements shown in Figure 1D were performed at approximately 3% parasitemia, assuming an erythrocyte infection rate of 10-fold within 48 hours as parasite cultures were initiated at 0.3% parasitemia and incubated for 48 hours under rapamycin before the measurements were performed.

      Figure 2: Is the increasing background in DMSO-treated parasites attributable to leakage of the di-cre system contributing to a background level of LacZ induction? To what extent would this impact results in the PRR assay format?

      We thank the reviewer for this important observation. We agree that low-level leakage of the loxP-DiCre system may contribute to the increased lacZ signal observed in DMSO-treated parasites under overgrowth conditions. However, this effect was only observed when parasites were allowed to proliferate extensively in the absence of effective drug pressure.

      In the context of the MULT-i<sup>2</sup> assay, these conditions correspond to compound concentrations that do not affect parasite survival or replication. Such concentrations are outside the range of interest for evaluating antimalarial activity, as they represent inactive treatment conditions. Therefore, although reporter leakage may contribute to background signal under extreme overgrowth conditions, we expect this effect to have a negligible impact on the interpretation of MULT-i<sup>2</sup> assay results.

      Please define the abbreviations used (e.g., NONMEM and DV).

      We thank the reviewer for pointing this out. We have revised the manuscript to define all abbreviations at their first occurrence in the text and have added the relevant terms to the abbreviation list.

      Line 297: cPRR assay - give citation.

      We thank the reviewer for pointing this out. We have added the appropriate citation for the cPRR assay at the indicated location in the revised manuscript.

      The 2 in MULT-i<sup>2</sup> is not always superscripted.

      We thank the reviewer for pointing this out. We have corrected the formatting throughout the manuscript to ensure that the “2” in MULT-i<sup>2</sup> is consistently presented as a superscript where appropriate.

      Reviewer #3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the MULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1) The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We have added more explanations to the Discussion including the strengths and limitations.

      Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. eLife Assessment

      This study reports a novel function for syntaxin 11, a specialized SNARE protein critical for the immune system whose mutations cause familial hemophagocytic lymphohistiocytosis type 4. The data convincingly show that depletion of STX11 impairs store-operated calcium entry in Jurkat T cells and that this defect is recapitulated in primary cells from a patient suffering from the disease; the authors further show that the syntaxin interacts with the pore subunit of the ORAI1 channel and propose that it primes the channel by promoting the assembly of multimers before activation by its endogenous ligand, the ER Ca2+ sensing protein STIM1. This is a conceptually important claim that challenges the prevailing view that all structural transitions in ORAI1 are STIM-driven. The high-quality data strongly support the interpretations, but the discussion would be strengthened if the authors suggest alternative mechanisms and mention earlier studies reporting ORAI1 regulation by vesicular trafficking.

    2. Reviewer #1 (Public review):

      Summary:

      Patients with STX11 mutations develop familial hemophagocytic lymphohistiocytosis Type 4, a fatal immune disorder marked by defective T and NK cell cytotoxicity and cytokine storm. The conventional explanation attributes this to impaired cytotoxic granule release, but this has never fully accounted for the broader disease picture. This study proposes an alternative mechanism. The authors show that STX11 is required for store-operated calcium entry through ORAI1 channels, which are essential for both cytotoxic killing and NFAT-driven gene expression in T cells. In STX11-deficient cells, ORAI1 currents drop, NFAT nuclear translocation fails, IL-2 expression is suppressed, and degranulation is impaired. These defects are largely rescued by ionomycin or a constitutively active ORAI1 mutant, placing the primary lesion at calcium signaling rather than the fusion machinery. Mechanistically, STX11 binds the C-terminal tail of ORAI1 via its Habc domain and maintains ORAI1 in a state competent for productive assembly prior to STIM1-dependent gating, a step the authors call "priming."

      Strengths:

      The paper identifies a novel and disease-relevant role for STX11 in calcium channel regulation and raises the possibility of using channel agonists as a therapeutic strategy in the disease. The biochemical and functional data are of high quality and generally consistent with the interpretation. The proposal that a non-conventional syntaxin directly interacts with ion channels to prime its activation is novel and interesting. Additional experiments now exclude the possibility that STX11 acts as a SNARE to sustain calcium fluxes by promoting the delivery of additional functional channels.

      Weaknesses:

      Previous studies reporting regulation of ORAI1 by vesicular trafficking are ignored and alternative mechanisms are not considered.

    3. Reviewer #2 (Public review):

      Summary:

      Vig's lab delineates a critical role for STX11 in CRAC channel function, particularly in the context of the fatal immune disorder familial hemophagocytic lymphohistiocytosis type 4 (FHL4). They demonstrate that Syntaxin 11 directly binds and regulates Orai1, and that STX11 depletion abolishes CRAC currents and downstream signaling. Loss of STX11 reduces IL2 gene expression and impairs degranulation, both of which are rescued by the constitutively active Orai1 mutant H134S, whereas a gain‑of‑function mutant targeting the C‑terminus fails to restore these defects. The authors conclude that STX11 primes Orai1 for optimal local assembly that is independent of STIM1 yet required for CRAC channel gating.

      Strengths:

      This study is firmly grounded in disease biology and demonstrates that STX11 downregulation leads to profound functional defects. Using a comprehensive suite of methods and analyses, the authors interrogate the co-regulation of STX11 and Orai1 and present a near-complete view of STX11's modulatory role in CRAC channel function and downstream signaling pathways. The figures are clear, and the statistical analyses are rigorous and convincing.

      Weaknesses:

      The authors conclude that Syntaxin 11 directly binds Orai1. This conclusion is well supported by a multifaceted approach-including co-immunoprecipitation (co-IP), molecular dynamics simulations, co-localization/FRET assays, and targeted mutational analysis-all of which are thoroughly executed. While the interaction appears reasonably strong in co-IP experiments, the STX11-Orai1 interaction is comparatively weaker in pull-down assays, which the authors attribute to instability of the purified His-STX11 protein. A remaining gap is direct evidence of interaction in live cells; this is understandably challenging given that fluorescent tagging of STX11 is not feasible. Fully resolving this question lies beyond the scope of the present study and will require more advanced approaches to capture STX11 binding dynamics.

      Comments on revised version.

      The authors have addressed all my comments highly satisfactorily!

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study reports a novel function for syntaxin 11, a specialized SNARE protein critical for the immune system whose mutations cause familial hemophagocytic lymphohistiocytosis type 4. The data convincingly show that depletion of STX11 impairs store-operated calcium entry in Jurkat T cells and that this defect is recapitulated in primary cells from a patient suffering from the disease; the authors further show that the syntaxin interacts with the pore subunit of the ORAI1 channel and propose that it primes the channel by promoting the assembly of multimers before activation by its endogenous ligand, the ER Ca2+ sensing protein STIM1. This is a conceptually important claim that challenges the prevailing view that all structural transitions in ORAI1 are STIM-driven. The data are high-quality and broadly consistent with the interpretation, but alternative mechanisms for the defects are not considered; additional work should rule out vesicular trafficking, discuss other mechanisms, and address methodological issues.

      We thank the editor and reviewers for assessing our work. We have now included additional experiments in a new main Figure 2, which directly rule out any general or Orai1 plasma membrane trafficking defects in Syntaxin11-depleted cells. There are additional experiments and/or analysis in many other figures, throughout the paper. We have included new and missing methods, quantifications and calibrations, and provided response to each of the reviewer’s comments below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Patients with STX11 mutations develop familial hemophagocytic lymphohistiocytosis Type 4, a fatal immune disorder marked by defective T and NK cell cytotoxicity and cytokine storm. The conventional explanation attributes this to impaired cytotoxic granule release, but this has never fully accounted for the broader disease picture. This study proposes an alternative mechanism. The authors show that STX11 is required for store-operated calcium entry through ORAI1 channels, which are essential for both cytotoxic killing and NFAT-driven gene expression in T cells. In STX11-deficient cells, ORAI1 currents drop, NFAT nuclear translocation fails, IL-2 expression is suppressed, and degranulation is impaired. These defects are largely rescued by ionomycin or a constitutively active ORAI1 mutant, placing the primary lesion at calcium signaling rather than the fusion machinery. Mechanistically, STX11 binds the C-terminal tail of ORAI1 via its Habc domain and maintains ORAI1 in a state competent for productive assembly prior to STIM1-dependent gating, a step the authors call "priming."

      Strengths:

      The paper identifies a novel and disease-relevant role for STX11 in calcium channel regulation and raises the possibility of using channel agonists as a therapeutic strategy in the disease. The biochemical and functional data are of high quality and generally consistent with the interpretation. The proposal that a non-conventional syntaxin directly interacts with ion channels to prime its activation is novel and interesting.

      Weaknesses:

      For readers to appreciate the value of patient experiments derived from a single individual, the authors should quote prior studies showing that STX11 protein levels are abolished in all known human STX11 mutations. The priming model, while functionally well-supported, rests on indirect structural evidence, and the precise conformational transition involved remains to be defined. These are acknowledged limitations, but alternate mechanisms have not been explored and formally excluded. More direct evidence should be provided to exclude the possibility that STX11 could act as a conventional SNARE and sustain calcium fluxes by promoting the delivery of additional ORAI1 channels from vesicles.

      In the revised version, we have included references for all those prior STX11 human mutations that have been biochemically characterized till date. The reviewer has correctly pointed out that STX11 protein levels were almost abolished in almost all previously reported mutations. See line 168-173. Therefore, the prior STX11 patient mutations are essentially comparable to the frameshift mutation characterized in this study, in terms of STX11 protein depletion and, therefore, the mechanisms underlying the phenotypic defects reported here as well as earlier. We, therefore, believe that our data from even a single FHLH4 patient, with severely depleted STX11 levels, and additional knockdown studies across three different cell lines, are representative of majority of STX11 mutant FHLH4 patients that have been previously characterized.

      Regarding the Reviewers’ concern that absence of STX11 as a conventional SNARE could affect Orai1 channel delivery from intracellular vesicles. We would like to point out the following:

      (1) In Miao et al. 2013 (1), Figure 3C-D, we showed that expression of a dominant-negative mutant of NSF, a non-redundant protein in vesicle trafficking, impaired vesicle trafficking but did not affect SOCE. This experiment had essentially ruled out a role for vesicle trafficking in SOCE. In the same paper, we had also shown that Orai1 levels in the PM do not increase post-store depletion (Figure 3-figure supplement 2).

      (2) SNAP23/25 form a four helical bundle with R- and Q-SNAREs in orchestrating vesicle fusion. In this paper, we have ruled out a direct role for SNAP23, SNAP25 and SNAP29 in SOCE (Figure 2-figure supplement 3).

      (3) In v1 of this manuscript, we had shown that U2OS cells stably expressing Orai1-BBS-YFP have identical levels of Orai1 in the PM with and without STX11 depletion (Supplementary Figure 3B). This showed that the biosynthesis or delivery of Orai1 to the PM is not affected by STX11 depletion. The levels were also assessed in store-depleted U2OS cells but not included because in Miao et al. 2013 we had already established that levels of PM Orai1 remain essentially equal in resting versus store-depleted cells.

      In the revised version, we have included the data from store-depleted cells in U2OS and also done quantification of PM Orai1 in HEK293 and Jurkat T cells. In addition, we have added three independent membrane trafficking/ vesicle secretion assays performed in STX11-depleted cells (new Figure 2 and associated supplements). In all cases, we find no evidence of Orai1 in intracellular vesicles or a general defect in membrane trafficking/ secretion in STX11-depleted cells. Orai1 is constitutively and stably expressed in the PM in resting as well as store-depleted cells in three different cell lines.

      (4) Most importantly, in Figure 7I-J of this manuscript, we showed that calcium influx from a constitutively active mutant Orai1 (Orai H134S) is identical between STX11-depleted and scramble control cells. If wildtype Orai1 was indeed stuck in vesicles in STX11-depleted cells, then how would mutant H134S Orai1 be able to rescue the defect in SOCE? We have included the quantification of PM levels of Orai1 mutants w.r.t WT Orai1 in new Figure 8-figure supplement 3B and 3D.

      In summary, we have now done several new experiments to directly measure Orai1 levels in the PM and general vesicle trafficking assays in HEK293 and Jurkat T cells and have found no defects in these upon STX11 depletion.

      Regarding STX11 induced precise conformational transition, we are trying to setup collaborations with scientists who might be able to visualize this in situ. Please note that while purification of isolated pore subunits of ion channels followed by crystallization or expression in synthetic membranes for cryo-EM is currently considered a gold standard in the analysis of ion channel pore subunits, we have shown that ion channels are dynamic macromolecular complexes, in vivo (2), where synaptic proteins dynamically bind to induce conformational changes and affect their stoichiometry (2). Please also see (3) and (4). More advanced approaches, therefore, need to be developed to enable visualization of the dynamics of ion channel macromolecular complexes in their native environment in situ. In the absence of such approaches, the structural insights obtained from detergent-purified isolated subunits will remain incomplete.

      Reviewer #2 (Public review):

      Summary:

      Vig's lab delineates a critical role for STX11 in CRAC channel function, particularly in the context of the fatal immune disorder familial hemophagocytic lymphohistiocytosis type 4 (FHL4). They demonstrate that Syntaxin 11 directly binds and regulates Orai1, and that STX11 depletion abolishes CRAC currents and downstream signaling. Loss of STX11 reduces IL2 gene expression and impairs degranulation, both of which are rescued by the constitutively active Orai1 mutant H134S, whereas a gain‑of‑function mutant targeting the C‑terminus fails to restore these defects. The authors conclude that STX11 primes Orai1 for optimal local assembly that is independent of STIM1 yet required for CRAC channel gating.

      Strengths:

      This study is firmly grounded in disease biology and demonstrates that STX11 downregulation leads to profound functional defects. Using a comprehensive suite of methods and analyses, the authors interrogate the co-regulation of STX11 and Orai1 and present a near-complete view of STX11's modulatory role in CRAC channel function and downstream signaling pathways. The figures are clear, and the statistical analyses are rigorous and convincing.

      Weaknesses:

      The authors conclude that Syntaxin 11 directly binds Orai1. This conclusion is well supported by a multifaceted approach, including co-immunoprecipitation (co-IP), molecular dynamics simulations, co-localization/FRET assays, and targeted mutational analysis-all of which are thoroughly executed. While the interaction appears reasonably strong in co-IP experiments, the STX11-Orai1 interaction is comparatively weaker in pull-down assays, which the authors attribute to instability of the purified His-STX11 protein. A remaining gap is direct evidence of interaction in live cells; this is understandably challenging given that fluorescent tagging of STX11 is not feasible. Fully resolving this question lies beyond the scope of the present study and will require more advanced approaches to capture STX11 binding dynamics.

      We thank the reviewer for acknowledging that analysis of the dynamic binding of STX11 will require standardization of advanced techniques which are beyond the scope of the present study. We plan to continue developing methods that will allow us to visualize the binding and unbinding of STX11 to Orai1 in vivo.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Mechanistic issues:

      (1) More direct evidence should be provided to exclude the possibility that STX11 could act as a conventional SNARE and sustain calcium fluxes by promoting the delivery of additional functional channels, stored in secretory vesicles or recycling endosomes, to the membrane. A significant fraction of the ORAI1 channel is in vesicles, and the mobilization of this intracellular pool regulates the rates of calcium fluxes in HEK-293 cells (PMID 26116575) and effector T cells (PMID: 35217583). Mobilization of this pool could account for part or all of the functional effects reported here. The only evidence that STX11 depletion does not impact the plasma membrane availability of the channel relies on one single flow cytometry profile (Supplementary Figure 3). This experiment is performed in U2OS cells stably expressing a fusion protein containing an extracellular bungarotoxin site. This cellular system is only used here; the other data are obtained either in HEK-293 cells or in Jurkat T cells. The level of endogenous STX11 in U2OS cells is unknown, and the efficiency of protein depletion has not been assessed. The efficiency of depletion should be shown, and the total number of channels assessed by comparing the expression levels of permeabilized and non-permeabilized cells. A flow cytometry profile of cells treated with thapsigargin should be included to match the experimental conditions of functional recordings. It would be valuable to repeat this experiment in Jurkat T cells by expressing ectopically the channel tagged on the extracellular side. To further exclude the involvement of vesicular trafficking, the authors should also evaluate the contribution of VAMP8, as this R-SNARE has been proposed to interact with STX11 to regulate the exocytosis of specialized granules in cytotoxic T cells (PMID: 26124288).

      We have not observed significant levels of intracellular Orai1, as described in the PMID 26116575 paper. Multiple technical reasons could explain the artefactual appearance of intracellular Orai1. HEK293 is an embryonic kidney cell line, with most cells showing a distinct spindle shape and filopodia as shown on the ATCC website (https://www.atcc.org/products/crl-1573). The HEK cells shown throughout the Hodeify et al. 2015 paper (PMID 26116575) lack this typical morphology that majority of the HEK cells should show. Transfecting cells with high amounts of DNA and liposomes can severely affect the health and morphology of the cells leading to artifacts where PM proteins appear to be stuck intracellularly. The plasma membrane of the cell in movie 1 of PMID 26116575, for instance, also shows membrane blebs or membrane ruffles. The authors should have used a PM marker such as WGA (wheat germ agglutinin) to distinguish PM Orai1 from any intracellular Orai1 to establish whether what appear as intracellular Orai1 vesicles are not blebs of PM Orai1. Similarly, no endo/exocytotic vesicle marker is used to distinguish them from Orai1’s presence in apoptotic vesicles of unhealthy cells. In view of the overall abnormal morphology and any PM or intracellular vesicle marker, the claim that Orai1 resides in intracellular vesicles is unfounded.

      Similarly, the PMID: 35217583, quoted by reviewer #1, is about in vitro differentiated primary mouse T cells. The study lacks any detail of the generation of the HA-tagged mouse Orai1 plasmid, used in the study, or even a back reference to show how this plasmid was validated for normal expression in primary T cells earlier. Ectopic expression of CMV promoter-driven plasmids in mouse primary T cells is extremely challenging and often results in poor cell health and incomplete and selective expression in only 5-10% of cells. It is unclear if the same HA-tagged human Orai1 plasmid that was used in PMID 26116575 is being used in this study to express in primary mouse cells. In the absence of all this information, it is unclear whether there was an issue with the generation of a new HA-tagged mouse Orai1 construct, which made the protein get stuck intracellularly. Or potentially the expression of a human protein in mouse primary T cells is the problem. The functional verification of the construct by showing rescue of SOCE in Orai1-deficient primary mouse cells is an absolutely essential control but is missing. In the absence of these controls, one cannot disregard decades of robust data on the localization of Orai1 in the PM from multiple labs and papers that have established its PM localization conclusively (5) (6).

      Most importantly, in both the PMID 26116575 and PMID: 35217583, HA tagged-Orai1 is detected using a bivalent antibody, followed by secondary antibody. It is well known that cross-linking of cell surface receptors or proteins using bivalent antibodies has a major caveat that involves antibody-mediated clustering and capping, especially in lymphocytes, which typically induces rapid internalization of the entire antigen-antibody complex. The paper by Sekine-Aizawa et al., in 2004, showed that tagging of receptors/ channels with bungarotoxin binding site (BBS) followed by labelling with bungarotoxin (BTX) bypasses this confounding factor and, therefore, allows accurate estimation and localization of PM versus intracellular proteins. The artefactual bivalent antibody-induced endocytosis continues even when cells are incubated on ice as endocytosis is only slowed but not stopped on ice.

      Due to this challenge, we have generated BBS tagged Orai1-YFP. We use flow cytometry to show PM Orai1 because it is an unbiased and quantitative way of showing surface Orai1 expression (estimated by measuring the intensity of surface-bound BTX), simultaneously, in thousands of cells with no potential for visual bias in the selection and imaging of cells. BTX labelling is done on ice and cells are washed and fixed right after labelling with BTX-A647 to stop endocytosis. In the revised version, we have used both complementary approaches of flow cytometry and microscopy, showing representative images of cells, alongside quantifications. All new experiments were performed in HEK293 and Jurkat T cells to estimate surface versus total expression in a new main Figure 2. The post-store-depletion data have been added to the existing U2OS experiment in new Figure 2-figure supplement 2A-E.

      Our BBS-tagged Orai1 construct also has a YFP tag at the C-terminus. Since the flow cytometry experiment involves gating on YFP-positive cells, the total number of channels (biosynthesis) can be compared by looking at the YFP intensities in the scramble versus STX11-depleted groups. These data have now been included in the previous and new experiments. In none of the cases could we detect any difference in YFP or BTX-A647 intensities, pre- or post-store-depletion in scr or STX11-depleted cells which would indicate defects in biosynthesis or Orai1’s presence in vesicles in any group.

      Regarding a role for VAMP8, in Miao et al. eLIFE 2013 (1), Figure 3C-D, we showed that expression of a dominant negative mutant of NSF, a non-redundant protein in the vesicle trafficking pathway, impaired Transferrin receptor recycling within 20 hours but did not affect SOCE at all. This experiment had conclusively ruled out any role for vesicle trafficking in SOCE and therefore assessment of the role of each of the individual proteins involved in membrane trafficking becomes redundant. We have additionally ruled out a role for SNAP23/ SNAP25/ SNAP29 (old supplementary Figure 12, new Figure 2-figure supplement 3A-C). SNAP23/25 form a four helical bundle with most R-SNARE and Q-SNAREs to orchestrate vesicle fusion. R-SNAREs, typically, need SNAP23/25 to interact with Q-SNAREs. Since a role for these non-redundant proteins has also been ruled out by us, it is unlikely that VAMP8 plays a role in modulating the effects of STX11 in SOCE.

      (2) Both ORAI1 and STX11 are S-Acylated on cysteine residues, and this post-translational modification promotes their recruitment to the immune synapse (PMID: 24910990, 34913437). One possibility that should be discussed is that STX11 could enhance the recruitment of the ORAI1 channel into lipid domains rich in cholesterol, thereby favoring its activation. This type of priming would still require direct interaction between the two proteins but involve a different mechanism than the one discussed by the authors. This could be experimentally tested by expressing a STX11 mutant lacking the cysteine residues required for its S-Acylation. It would also be interesting to test whether depletion of STX11 impairs the recruitment of ORAI1 to the immune synapse forming between Jurkat T cells and antigen-presenting cells.

      In our experiments reported in this paper, we have used soluble anti-CD3 as well as plate-coated anti-CD3 in combination with soluble anti-CD28 to stimulate Jurkat and primary T cells. These antibodies are routinely used to stimulate T cells and this type of stimulus doesn’t depend on the formation of a classical immune synapse with an antigen-presenting cell (APC). Despite the absence of synapse, the T cells get fully activated and functional, as seen by NFAT translocation and secretion of cytokines, such as IL-2 in new Figure 4. Therefore, whether there is a defect in the recruitment of Orai1, or STX11, to T cell synapse formed with an APC is not within the scope of this study. Furthermore, accurate analysis of protein localization within the immune synapse requires a dedicated study employing sub-diffraction resolution microscopy approaches.

      Regarding PMID: 24910990, and the mechanism of recruitment of STX11 to the membranes. We believe this remains unknown. The frameshift mutant used in our study lacked all terminal cysteines, which have been earlier proposed to be crucial for membrane targeting, as well as a terminal part of the SNARE domain and yet it localized to the PM just as well as wild-type STX11 (See new Figure 5E). We have added this result in the text line 302-305 and removed the line stating that post-translational modifications of the terminal cysteines target STX11 to the PM as previously claimed in PMID: 24910990 and mentioned in v1 of this paper. In view of these new data, it currently remains unknown whether potential attachment to PIP2 in PM via various basic residues, spread throughout the sequence (7, 8), or binding to another protein targets STX11 to the PM (PMID: 26771955). There is no obvious poly-basic stretch in STX11 sequence, therefore, a systematic and focused deletion and mutagenesis study will be needed to individually assess the above possibilities which is outside the scope of the present study.

      Methodological issues

      (3) Since STX11 colocalize with Orai1 already in basal conditions, independently of STIM1, it could influence basal calcium levels. This cannot be appreciated from the data presented, because all the SOCE protocols start in calcium-free conditions, preventing baseline comparison between WT and STX11-deficient cells. A potential difference in basal calcium levels should be explored, and the impact of STX11 depletion on basal calcium fluxes should be documented by calcium shifts (2 mM → 0 mM → 2 mM) or manganese quenching approaches.

      The cells are typically loaded with Fura2 in 2mM calcium containing Ringer’s buffer. We switch the cells to 0mM right at the start of the SOCE protocol and start imaging within 5-10 seconds. Therefore, in our experience, the baselines of scramble versus STX11-depleted cells should show a difference even in the SOCE protocol if the basal calcium levels are affected because Fura2 is already present and bound to basal calcium present in the cytosol at the start of the protocol. Still, we have done the experiment suggested by the reviewer as shown in Author response image 1. The assay started with cells in 2 mM extracellular Ca<sup>2+</sup> followed by addition of 10 mM EGTA, which according to the following equation quenches the 2 mM extracellular Ca<sup>2+</sup> (https://somapp.ucdmc.ucdavis.edu/pharmacology/bers/maxchelator/CaEGTA-TS.htm).

      where, [Ca2+]<sub>Free</sub> is the free/unbound Ca2+, [Ca2+]<sub>Total</sub> is the total Ca2+, [EGTA]<sub>Total</sub> is the total EGTA concentration and Kd is the dissociation constant between Ca2+ and EGTA at 37°C and pH 7.4.

      To confirm complete sequestration of the extracellular Ca<sup>2+</sup>, we also repeated the assay with 20 mM EGTA but did not notice any difference between the 10 mM and 20 mM EGTA conditions. We do not see any differences in basal calcium, under any condition, between Scr and STX11 shRNA treated cells HEK or Jurkat T cells.

      Author response image 1.

      Representative Fura-2 traces of Scr (black) and STX11 (red) shRNA-treated HEK293 (A) and Jurkat (B) cells, where the cells were incubated with 2 mM Ca<sup>2+</sup> followed by addition of 10 mM EGTA to quench the 2 mM Ca<sup>2+</sup>. Since we do not have access to a perfusion system, we could not test Fura2 response after re-addition of 2mM calcium to the existing EGTA and Ca<sup>2+</sup> mixture. However, the transition from 0mM to 2mM is already shown in Figure 8.

      (4) The quantification of the calcium imaging data is problematic and requires clarification. In most figures, the data are shown normalized to the control condition. According to the method section (lines 721-724), 100% is the maximum value of the scramble shRNA group (amongst the three experiments). But what was measured here? The slope during calcium readmission? The peak amplitude after calcium readmission? Expressed as a ratio or as calcium values? The recordings are presented in micromolar calcium concentration. This implies calibration, but a calibration procedure is not mentioned. Please clarify. For the recordings of constitutive calcium entry in Figure 7, this normalization is not performed, and the data are expressed as ratio values. Here, it looks like the parameter quantified and compared is the absolute ratio value after calcium readmission. This is inappropriate. The trace in Figure 7G shows that the basal levels differ by more than two ratio units between control and STX11-depleted cells. Normalizing the data in Figure 7G to the basal ratio value would show no difference in the peak response amplitude between the two conditions. These data should be re-analyzed, and both the slope and the amplitude of the response should be presented, with statistics performed on independent recordings, not on individual cells pooled from different experiments. Cells from the same recording are not experimentally independent samples but rather replicates of the same experiment.

      We have updated the relevant method section with more details in lines 823-858. The peak amplitude after calcium readmission was measured and compared to the baseline as described in the updated methods. The Fura 2 calibration method has also been added to the revised version, we apologize for this omission earlier. Separate calibration was done for Fura 2 experiments in all figures except Figure 7 from version 1 (new Figure 8). The experiments in new figure 8 were done using a different objective (20X, water) and, therefore, although a separate calibration was done for these experiments it was not applied to the data. We apologise for this omission on our part and have now applied the respective calibration to these experiments.

      The trace in Figure 7G of version 1 should look different even in 0mM calcium in our opinion. The reason is the same as that explained in point 3 above; CAD-mediated constitutive activation of Orai1 is one of the strongest. The cells, even when they are being loaded with Fura2 in Ringer’s buffer with 2mM calcium, are constitutively recruiting calcium ions. This should result in a shift in Fura2 excitation due to higher levels of basal calcium. When the cells are switched to 0mM calcium and imaged within 4-5 seconds, the intracellular Fura 2 is still bound to all this extra calcium in the cytosol and therefore the baselines should show a significant difference. If the cells were imaged for several minutes in 0mM calcium, we might have seen the difference in basal ratios slowly reducing. However, we switched to 2mM calcium within 120sec. At this point any free Fura2 would be expected to bind incoming calcium again. For the same reason, the cytosol of Scr cells which would still have relatively higher levels of intracellular calcium concentration compared to STX11-depleted cells will show a smaller further increase in 2mM due to calcium-dependent inhibition of CRAC currents within the time frames we have measured. We have now added the Fura-2-calibrated response in the new Figure 8G-H. We have shown below the baseline-subtracted (normalized) response for Figure 8G. As you can see, there is still a significant difference in 2mM calcium between Scr and STX11-depleted cells but in this representation the important difference at 0mM is masked, we have therefore chosen to retain the original figure with Fura-calibrated values at 0 as well as 2mM calcium in Figure 8G-H.

      The cells shown in the old Figure 7G of version 1 were not from the same recording but from three different experiments. Because the cells were imaged with 20X objective in these experiments to allow selection of Orai1-CFP or mutant Orai1-CFP and CAD-YFP double-positive cells, the number of cells analyzed per experiment was less compared to other experiments. We have now shown Fura 2 calibrated values in the new Figure 8G-L. We have also re-done the statistical analysis on three independent experiments from each. As shown in Author response image 2, the difference is still statistically significant whether we show merged cells from all three experiments or single representative experiment out of three repeats. We believe merged cells have more information to offer and therefore have retained the same figures with the original analysis in the main Figure 8G-L.

      Author response image 2.

      Box plots representing quantification of individual repeats of constitutive calcium influx in Scr and STX11 shRNA-treated HEK293 cells expressing YFP-CAD and Orai1-CFP without (A) and with baseline subtraction (B). (C-D) Quantification of individual repeats of Scr and STX11 shRNA-treated HEK293 cells expressing Orai1-H134S (C) and Orai1-ANSGA (D) mutants.

      (5) Quantification of pull-down experiments. The binding data in Figures 4F, 5F, and 5J are presented largely qualitatively. Densitometric quantification with statistical comparisons across wild-type and mutant conditions would make these results more convincing, particularly given that the authors themselves acknowledge the interaction appears relatively weak in vitro.

      This has been done and included alongside the respective panels in the new Figure 6G and 6L (for old Figure 5F and 5J of version 1, where differences appeared relatively small in some experiments). The differences across lanes in both figures and their repeats were statistically significant.

      Figure 4F showed a clear and visually significant difference in binding across lanes and repeats and therefore no quantification is needed for these experiments in our opinion.

      Limitations of the study and mechanistic inferences.

      (7) Interpretation of the ORAI:ORAI FRET and crosslinking data. STX11 depletion increases basal ORAI:ORAI FRET (Figure 7A-C) and shifts crosslinked species toward higher molecular weights (Figure 7D-E). The authors interpret this as ORAI1 being trapped in an unprimed state, but higher FRET and higher-order species would conventionally suggest increased rather than decreased assembly. The paper needs a clearer mechanistic explanation of what "unprimed" looks like structurally. Is this aberrant crowding, non-productive oligomerization, or something else? The distinction between a change in intermolecular distance within existing oligomers versus an increase in oligomer density matters here and should be addressed.

      Higher ORAI: ORAI FRET and a shift in the size of crosslinked Orai1 oligomers, when analyzed together, suggests formation of ‘non-functional’ higher-order oligomers. Higher order does not necessarily translate to better function in the case of ion channels, it can also lead to non-selectivity or formation of ‘non-productive’ oligomers, as mentioned by the reviewer. It was shown by us earlier in Li et al. (2016) (2) that bigger oligomer size revealed by higher number of photobleaching steps of Orai1 did not translate to better function but led to non-selectivity.

      In new Figure 8, crosslinking with BS3, which has a spacer arm and working distance of ~11 Å, very likely reflects a change in the number of subunits within individual oligomers and not crosslinking of independent existing oligomers. This is because we show that neither total Orai1 expression nor Orai1 expression in the PM change in any group in new Figure 2. FRET works best within 1 to 10 nm distance, and therefore, in theory, can lead to energy transfer between neighbouring Orai1 oligomers in high Orai1-expressing cells. However, because there was no change in Orai1 abundance in the PM (new Figure 2) or distribution within PM (new Figure 7E,F,J,L) of any group, FRET changes also likely reflect intra-oligomer changes rather than inter-oligomer interactions. FRET changes can also arise from a change in the respective orientation of fluorophore pairs but when analyzed together with crosslinking studies, changes in pore assembly likely coincide with conformational shifts in Orai1 protomers. Furthermore, FRET has been used earlier to show shifts in conformation of other ion channels (9). Therefore, we believe that, when used together, these two approaches strongly suggest an intermediate conformational state along with a change in number of Orai1 subunits per channel since there was no evidence of overcrowding in the PM or obvious segregation of Orai1 in specific regions of PM in new Figures 2 and 7E,F,J,L.

      We could not assess whether the oligomers of Orai1 formed in the absence of STX11 possess an intact pore. The presence or absence of pore in STX11-depleted cells will require extraction of Orai1 oligomers from native membranes and performing systematic structural analysis using cryo-EM or related approaches which is outside the scope of this study.

      (8) The ANSGA versus H134S discrepancy. H134S ORAI1 rescues calcium influx in STX11-depleted cells (Figure 7I-J), but the ANSGA mutant does not (Figure 7K-L). The authors conclude from this that STX11 induces molecular shifts within ORAI1 transmembrane helices, and not in its C-terminal tail. This is an important mechanistic inference that needs more discussion. What does this imply about the conformational state of primed ORAI1? And why is straightening of the tails not sufficient for full opening without the correct TM helix arrangement? This distinction has implications for how the STX11-ORAI1 interaction should be modelled and should be engaged with more thoroughly.

      There is no discrepancy here, please also see our response to reviewer 2’s comment #4. The experiment implies that the conformational state of primed Orai1 involves shifts in the TM region of Orai1 and is different from unprimed state. The structural similarities between H134S and ANSGA Orai1 mutants have not been formally established. Unlike H134S, no structure exists for the ANSGA mutant. In the absence of this, it is impossible to comment on whether the two constitutively active mutants are structurally comparable or whether there are multiple ways to stabilize open states of CRAC channel pore, especially when using TM mutants of Orai1.

      The goal of this experiment was to determine what kinds of structural shits STX11 potentially induces in native Orai1. Using previously characterized constitutively active mutants and fusion proteins from the CRAC field, we have ruled out a potential role for STX11 in simply changing the orientation of Orai1 C-terminal tails. A discussion on the topic of why tail straightening of Orai1 is insufficient to open Orai1 is outside the scope of this paper. As pointed by reviewer 2, it is possible that C-term tails already exist pointing towards the cytosol in native, resting Orai1, although this has not been shown in any study using structure of full-length WT Orai1 and is purely speculative at this point. We prefer to not engage in speculative structural insights.

      Other points

      (9) Figures 1G and 1H. The patient-derived mutant STX11 band runs at approximately 37 kDa rather than the predicted 39.5 kDa. The authors suggest instability or reduced antibody reactivity, but premature translation termination is also a possibility that should be acknowledged.

      We have added this point in line 168.

      (10) Figure 2B. The traces and current voltage relationships should be rescaled to show the rectification and inactivation profile of the current in cells depleted of STX11.

      This has been done and modified in new Figure 3 (Figure 2 of version 1).

      While preparing source data files for all figures, we noticed an error in the value of the SE in the STX11-depleted group of old Figure 2C, which has now been corrected. The SE value in the older version was erroneously pasted from an adjacent data column.

      Similarly, in old Figure 3 (version 1), new Figure 4C, we noticed that some data points in the STX11 group were pasted twice in the same excel column. These cells were removed and additional cells were analyzed from the same experiment and added to this group. The overall result remains the same but the distribution of data points looks a bit different.

      (11) Figure 4B: This experiment should be repeated in cells treated with thapsigargin to deplete intracellular calcium stores, and the extent of colocalization quantified by measuring the Pearson's correlation coefficient.

      This has been done. Pearson’s correlation coefficient is included in new Figure 5D.

      (12) Figure 4F. Why is there no detectable band in the input lane of the left blot?

      Western blots show relative intensities of bands of proteins across lanes. A faint band in the input lane of old Figure 4F suggests that the IP/ co-IP/ pull down was robust. If we increase the exposure, the input band would become stronger but the pull-down band would become over-saturated and the difference in the intensities would not be linear. The faint non-specific bands in other lanes represent a fraction of soluble STX11 that tends to crash out of solution over time and gets spun down with the beads. See lines 535-541 explaining this.

      (13) Figure 5. Immunofluorescence data showing the membrane staining of the mutated syntaxin and channel should be included, as well as calcium recordings of cells expressing YFP-CAD with WT and mutated ORAI1.

      In version 2 Figure 6A, we have now also shown co-localization of mutant STX11 with Orai1-YFP in resting and store-depleted cells, in addition to WGA. Pearson’s correlation (not shown) did not show any significant difference in the localization of mutant synatxin 11 w.r.t Orai1. Calcium recordings of CAD-induced constitutive calcium influx from wild-type versus mutant Orai1 are now shown in new Figure 6O-P.

      (14) Figure 6B. A clear colocalization of CFP-O1 and STIM1-YFP is visible on the images, yet the authors conclude from morphometric analysis that the channel is not recruited into ER-PM clusters. Please show the difference in colocalization quantified by measuring the Pearson's correlation coefficient. Pictures should also be provided with the C-terminally tagged construct.

      The quantification of CFP-Orai1 localization inside Stim1-YFP puncta was already shown in old Figure 6E and F. The residence of Orai1 inside STIM1 puncta versus total Orai1 in the PM of STX11-depleted groups was clearly reduced. We have now also shown Pearson’s correlation coefficient for Stim Orai co-localization inside puncta in new Figure 7F.

      TIRF microscopy images of C-terminally tagged Orai1 were already included in Supplementary Figure 10. No defect in co-clustering of C-terminally tagged Orai1-YFP and N-terminally tagged CFP-Stim1 was seen and yet SOCE was inhibited. Therefore, we never concluded from Figure 6 that Orai1 and Stim1 fail to co-localize. We said, they fail to form ‘functional’ clusters. We have now moved the representative TIRF images from supplementary figure 10 to the new main Figure 7G. The Pearson’s correlation coefficient for Stim Orai1 co-localization inside puncta is shown in new Figure 7L.

      (15) Figure 6E and 6F show the same data.

      Figure 6E showed fraction of Orai1 inside Stim1 puncta divided by total Orai1, and 6F showed fraction of Orai1 outside puncta divided by total Orai1. The plots are different but we agree that the data are coming from same cells. We have removed old panel 6F and replaced it with Pearson’s correlation coefficient of Stim1:Orai1 colocalization in puncta in new Figure 7F.

      (16) Figure 7G-L. The difference in constitutive calcium fluxes should be confirmed by Manganese quench recordings. The surface expression of the Orai1 mutants should be shown.

      We have now shown the quantification of surface expression of Orai1 mutants for each respective mutant in the new Figure 8-figure supplement 3B and 3D. The Orai1 mutants we have used in this paper are well established in the literature, they showed clear surface localization and the differences in calcium influx between Scr and STX11 treated cells upon overexpression of Orai1 mutants in HEK are robust. Therefore, we do not see any compelling reason for repeating all of the experiments from Figure 7G to 7L to also show manganese quench recordings, as suggested by the reviewer. We have applied Fura 2 calibration done for these experiments to calculate intracellular calcium. These have been shown in the revised and new Figure 8G-L, where F340/380 ratios of representative calcium assays have also been replaced with the calibrated intracellular calcium concentration.

      (17) Supplementary Figure 12. The recordings show a very large variability between experiments. The different SNAREs that are depleted here could compensate for each other, accounting for this variability. It would be interesting to show the effect of the combined silencing of all the SNARES tested here. The efficiency of the protein depletion should also be documented.

      Genome-wide high- or medium-throughput screens are inherently noisy. None of the genome-wide high- or medium-throughput screens show evidence of protein depletion for each gene in any of the published screens to our knowledge. We chose to only characterize the candidates that reproducibly showed > 70% inhibition of SOCE, others were ignored as noise.

      Silencing of all SNAREs together will definitely lead to loss of morphology and early lethality as all membrane trafficking will be stopped. We never analyze cells that do not show normal morphology and have compromised viability for ablation of SOCE.

      (18) Lines 236-238. The authors note that STX11 harbors a stretch of C-terminal cysteines proposed to be essential for its membrane localization, but do not elaborate on the underlying mechanism. It would strengthen the discussion to explicitly acknowledge that this membrane anchoring is mediated by S-acylation of these cysteines PMID: 24910990 and to connect this to the known enrichment of Orai1 in lipid rafts and the immune synapse PMID 34913437. Both observations are relevant to understanding how STX11 and Orai1 are brought into proximity at the plasma membrane, and their omission leaves an explanatory gap in the proposed interaction model.

      Please see our response to point #2 above. We do not think C-terminal cysteines target STX11 to the PM. We have corrected this claim based on an earlier study, PMID: 24910990, in the revised version of this paper. Analysis of immune synapse and lipid rafts are outside the scope of this paper. The mechanism of PM targeting of STX11 is currently unestablished and will require a systematic and focused mutational analysis which is outside the scope and main focus of this paper.

      (19) Line 351. The statement that syntaxin depletion does not alter the structure or proximity of junctional ER to the plasma membrane is not supported by data. Neither electron microscopy nor TIRF imaging has been performed, which would be required to back up this claim.

      Because Stim1 itself can be used as a marker of ER-PM junctions, this statement was supported by data shown in Figure 6C, D, G, H of version 1 of this paper where the intensity and area of Stim1 clusters was assessed using TIRF microscopy and found to be indistinguishable between STX11 and scramble control cells. The imaging done in Figure 6G, H was TIRF imaging and this was already specified in the legend. We have now also done TIRF imaging of GFP-Mapper-expressing scr and STX11-depleted cells. Mapper is a genetically encoded fluorescent protein that was previously shown to mark ER-PM junctions (10). We found no significant difference in the area or intensity of GFP-Mapper puncta (new Figure 7O-Q), just like Stim1 puncta didn’t show any defect in STX11-depleted cells. Please see modified text from 416-423.

      (20) The molecular dynamics methods need more detail: force field, simulation length, water box dimensions, and convergence criteria should all be specified to allow replication. The supplementary RMSD plots (Supplementary Figure 5B) should also show individual replicate trajectories rather than averages only.

      We had already mentioned the force field (OPLS4) and simulation length (500ns) in the methods section. Also, the RMSD plots in Supplementary Figure 5B already showed individual replicates in version 1.

      We have now updated the methods with following additions:

      The OPLS4 force field was used for all 500 ns simulations in an orthorhombic water box with a buffer distance of 10 Å beyond the solute in each direction. Simulation stability was assessed based on the protein backbone RMSD over simulation time.

      Trajectory clustering was performed using the trajectory clustering tool in Schrödinger, which applies affinity propagation to the pairwise backbone RMSD-based similarity matrix. Within each affinity propagation run, convergence was defined as no change in the set of exemplar frames for 15 consecutive iterations, with a maximum of 400 iterations per run. If convergence was not reached, the damping factor was increased from 0.5 in increments of 0.01 until convergence.

      Reviewer #2 (Recommendations for the authors):

      Overall, this is a timely and impactful study supported by a broad set of methods and cell types. Before publication, the manuscript should address the following points.

      Major:

      (1) The authors note that STX11 contains cysteine residues that enable membrane association. What is the specific mechanism of membrane attachment? Could it occur via S-acylation (palmitoylation)? Both Orai1 and STIM1 are known to undergo S-acylation, which raises the possibility that this modification might also facilitate STX11 membrane anchoring and/or co-residence with Orai1. Is STX11 constitutively membrane-associated, or does it show preferential localization to specific membrane subdomains, particularly in proximity to Orai1?

      STX11 is constitutively membrane-associated and does not show any preferential localization to specific membrane subdomains in confocal images. Figure 4, panel A and B from version1 clearly showed this. In an earlier paper by Hellewell et al. 2014, PMID: 24910990, S-acylation of terminal cysteines of STX-11 was proposed to be crucial for membrane attachment of STX11 and its recruitment to the immune synapse. However, please see our response to reviewer 1’s comment #2 and a new Figure 5E for the localization of the frameshift FHLH4 mutant characterized in this paper. The frameshift mutant that we have characterized lacked all terminal cysteines as well as a short terminal part of the SNARE domain. Cloning and ectopic expression of this mutant still showed constitutive localization to PM and did not show preferential distribution to any specific regions. Therefore, we do not think that terminal cysteines of STX11 contribute to its membrane attachment, we have accordingly modified lines 291-293, 302-305, 540 in the revised version. Also see our response to your point#6 below.

      (2) Is there a possibility to monitor a dynamic change in STX11 co-localization from before to after store-depletion?

      We did not observe any change in the overall distribution of STX11 in cells expressing STX11 alone or co-expressing Orai1 with STX11, pre- or post-store-depletion (please see new figure 5B-C). In cells co-expressing ORAI1, STX11 and STIM1 (see new Figure 5M-N), we could not capture the dynamic segregation of STX11 into regions of PM devoid of STIM:ORAI puncta and therefore have only pre- or post-store-depletion images. Dynamic change in STX11 distribution would require live, multi-colour, high-resolution imaging of diffraction-limited ER-PM junctions and adjacent regions which is technically extremely challenging, and especially due to our inability to tag STX11 with a fluorescent tag without disrupting its localization. Also see our response to your point#6 below.

      (3) The authors use CAD to prove that the interaction with the R289A_E272A_E275A_E278A mutant is normal as for the wild-type. Does this also hold for STIM1 wild-type full-length?

      This is also true for full-length STIM1. The data have now been added to the new Figure 7-figure supplement1.

      (4) The authors state that STIM1 binds both the N- and C-termini of Orai1. While STIM1 binding to the Orai1 C-terminus is well established, the nature of its interaction with the N-terminus remains debated. Fragment-based assays suggest direct binding to the N-terminus; however, direct interaction with full-length Orai1 has not been conclusively demonstrated. This point should be phrased more cautiously to reflect the current uncertainty.

      We have re-phrased the sentence as follows in line 543: “The individual relevance of Orai1 N- versus C-terminus in the trapping versus gating of Orai1 remains unclear”

      (4) In the discussion, the authors report: "Though crucial for trapping and gating, Orai1 tails were missing from early structures of Drosophila Orai [28]". A previous NMR structure suggested that the C-terminal tails of two adjacent Orai1 subunits bend and pair with each other in an antiparallel fashion, and sit closely apposed to PM [37]. However, in recent structures of constitutively active H134 mutant Orai, the C-terminal tails were found to orient away from the membrane [28]. In STX11-depleted cells, switching the CFP-tag from the Orai1 N- to the C-terminus could rescue its clustering but not gating by Stim1. Furthermore, STX11 depletion inhibited the constitutively active ANSGA mutant of Orai1 [29], where the tails of Orai1 are proposed to be constitutively unlatched. These data essentially reinforce our conclusions that STX11 induced molecular shifts encompass Orai1 transmembranes.' However, the information provided here is not fully correct. The early Drosophila Orai structure lacks the full N-terminus but retains most of the C-terminus. It was the X‑ray, not cryo‑EM, structure that suggested an antiparallel arrangement of the Orai1 C-termini. Although the "open" X‑ray structure shows unlatching and straightening of TM4-C-termini, it remains uncertain whether these features reflect physiological gating or crystallization artifacts. It is also unclear whether the Orai1 ANSGA gain‑of‑function mutant adopts a similar unlatching; however, prior work indicates that ANSGA impairs proper coupling to the C‑terminal binding interface (in contrast to Orai1 H134S, which maintains effective coupling). This raises the key question: why do H134S and ANSGA respond differently to STX11 depletion? One possibility is that these mutants stabilize distinct conformations of the TM4-C-termini ("latched" vs "unlatched" states) that differentially dictate the requirement for STX11 in channel assembly or gating. We recommend refining the discussion

      We have changed the word ‘missing’ to ‘truncated’ in line 545 and 546.

      We agree that there is no evidence in literature that establishes similarity between H134S and ANSGA mutation-induced conformations of Orai1. It is, however, implied in most previous studies of mutant Orai1s that there is only one possible open state/conformation. We have added the suggested point and modified the discussion in line 551-556.

      (5) The authors highlight: 'A major problem with this interpretation is that even though amplification of CRAC currents was shown, none of the previous patch clamp studies established whether the higher currents resulted from a greater number of active channels or unchecked conductance per channel by performing single channel recordings.' It should be noted that CRAC channels have extremely low single‑channel conductance, making direct single‑channel recordings challenging. As a result, estimates of open probability and channel number typically rely on fluctuation (noise) analysis rather than direct measurements of single‑channel events (see https://doi.org/10.1085/jgp.200609588). We suggest acknowledging this limitation in the discussion to contextualize the interpretation of gating and channel density.

      We acknowledge how challenging it is to record the single-channel conductance from CRAC channels. We have added this fact to the discussion and the reference that the reviewer has suggested in line 570-572.

      (6) STX11 appears to shift Orai1 localization into puncta. Activated STIM1 is known to engage plasma membrane PIP2 to facilitate Orai1 coupling. How, if at all, is STX11 linked to PIP2 or PIP2-rich microdomains? Is there evidence for direct PIP2 binding by STX11, or for indirect recruitment via PIP2-binding partners? Any available data on STX11's lipid interactions or its enrichment within PIP2-enriched regions would help clarify this mechanism.

      We have not claimed that STX11 shifts Orai1 into puncta. We already showed in old supplementary figure 10 and Figure 6G-J of version1 (v1) of this paper that the C-terminally tagged Orai1-CFP can very well form puncta and co-localize with STIM1 in STX11-depleted cells. To avoid this confusion, we have moved the old Supplementary Figure 10 from v1 to the main figure in revised version, see new Figure 7 panel G. Despite the presence of ORAI1-CFP in puncta with YFP-Stim1, the SOCE was inhibited in STX11-depleted cells. Please also see new Pearson’s correlation coefficient for Stim1 and Orai1 colocalization in Figure 7 panel L. Therefore we concluded that, Orai1 forms ‘nonfunctional’ clusters with Stim1 in STX11 depleted cells, please see modified lines 408-412, clearly explaining this.

      Although syntaxin 1A has been shown to interact with cholesterol (11) as well as PIP2 (7, 8) using either a stretch of polybasic residues or basic residues spread throughout several domains. To our knowledge, these have not been proposed to recruit or segregate STX11 in membranes. STX11 doesn’t contain an obvious stretch of poly-basic residues in its sequence, either, to quickly mutate and address this question. Please also see our response to your point #1 and #2 above. Answering this question will require a systematic and dedicated mutagenesis study.

      Minor:

      (1) Please indicate in Figure 1 in the respective graphs in which cell type the Ca2+ imaging studies have been performed.

      Done.

      (2) Figure 4C, D: Why are the input bands so weak?

      Please see our response to reviewer #1’s similar comment 12 above.

      (3) Figure 5A: Please clarify what WGA is.

      WGA is wheat germ agglutinin which is used to mark PM in imaging experiments. It binds to N-acetyl-D-glucosamine and sialic acid residues found in mammalian cell membranes and glycoproteins. We have added the explanation to the new Figure 6A legend.

      The authors state that "the constitutively active ANSGA (261-265) mutant of Orai1 (Supplementary Figure 11G), which harbors 4 consecutive mutations in the Orai1 C-terminus ...". Please clearly state this is the nexus region connecting the C-terminus with TM4. The 5 aa stretch is not the C-terminus; it is just close to the C-terminus.

      We have modified this, as suggested, in line 470-471.

      References:

      (1) Miao Y, Miner C, Zhang L, Hanson PI, Dani A, Vig M. An essential and NSF independent role for alpha-SNAP in store-operated calcium entry. Elife. 2013;2:e00802.

      (2) Li P, Miao Y, Dani A, Vig M. alpha-SNAP regulates dynamic, on-site assembly and calcium selectivity of Orai1 channels. Mol Biol Cell. 2016;27(16):2542-53.

      (3) Chorev DS, Baker LA, Wu D, Beilsten-Edmands V, Rouse SL, Zeev-Ben-Mordehai T, et al. Protein assemblies ejected directly from native membranes yield complexes for mass spectrometry. Science. 2018;362(6416):829-34.

      (4) Dorwart MR, Wray R, Brautigam CA, Jiang Y, Blount P. S. aureus MscL is a pentamer in vivo but of variable stoichiometries in vitro: implications for detergent-solubilized membrane proteins. PLoS Biol. 2010;8(12):e1000555.

      (5) Vig M, Peinelt C, Beck A, Koomoa DL, Rabah D, Koblan-Huberson M, et al. CRACM1 is a plasma membrane protein essential for store-operated Ca2+ entry. Science. 2006;312(5777):1220-3.

      (6) Feske S, Gwack Y, Prakriya M, Srikanth S, Puppel SH, Tanasa B, et al. A mutation in Orai1 causes immune deficiency by abrogating CRAC channel function. Nature. 2006;441(7090):179-85.

      (7) Murray DH, Tamm LK. Clustering of syntaxin-1A in model membranes is modulated by phosphatidylinositol 4,5-bisphosphate and cholesterol. Biochemistry. 2009;48(21):4617-25.

      (8) van den Bogaart G, Meyenberg K, Risselada HJ, Amin H, Willig KI, Hubrich BE, et al. Membrane protein sequestering by ionic protein-lipid interactions. Nature. 2011;479(7374):552-5.

      (9) Miranda P, Contreras JE, Plested AJ, Sigworth FJ, Holmgren M, Giraldez T. State-dependent FRET reports calcium- and voltage-dependent gating-ring motions in BK channels. Proc Natl Acad Sci U S A. 2013;110(13):5217-22.

      (10) Chang CL, Chen YJ, Liou J. ER-plasma membrane junctions: Why and how do we study them? Biochim Biophys Acta Mol Cell Res. 2017;1864(9):1494-506.

      (11) Lang T, Bruns D, Wenzel D, Riedel D, Holroyd P, Thiele C, et al. SNAREs are concentrated in cholesterol-dependent clusters that define docking and fusion sites for exocytosis. EMBO J. 2001;20(9):2202-13.

    1. eLife Assessment

      This valuable study provides insights into the role of steroid signaling during tumorigenesis in the adult male drosophila accessory gland (functional equivalent of the prostate gland in mammals), hinting at a possible counterintuitive anti-tumoral role of sex hormones during prostate cancer in certain patients. While the Drosophila model provides an elegant way to study the hypothesis derived from the Cancer Atlas analysis, the analyses of public prostate cancer expression data are incomplete and critical knowledge on patients' treatment modalities and normalization across different datasets is missing. This work would be of interest to prostate cancer researchers as it suggests that the absence of androgen receptor signaling in humans could constitute a mechanism promoting tumor escape.

    2. Reviewer #1 (Public review):

      Summary:

      In this article, Vialat and his colleagues examine the early stages - which remain largely unknown - of the tumor escape process, particularly the basal extrusion of tumor cells following endocrine therapy for prostate cancer.

      They first used the "Prostate Cancer Atlas" database, which provides access to a vast amount of transcriptomic data, to perform high-throughput analyses. Interestingly, analyzing a series of Androgen Receptor (AR) target genes in castration-resistant prostate cancers, they concluded that the loss of the canonical AR signaling pathway may contribute to tumor resistance.

      Using a well-established model in Drosophila, they then replicated in vivo an endocrine therapy targeting the accessory gland by genetically inhibiting the expression of ecdysone, the only sex steroid present in Drosophila. These experiments induced basal extrusion similar to the mechanism observed in tumor escape in humans.<br /> These results suggest that the deprivation of sex steroids may play an important role in tumor progression.

      However, although the data from the "Prostate Cancer Atlas" constitutes a powerful tool that serves as the basis for this new concept, clinical validation using carefully selected human tumor samples would help strengthen the authors' conclusions.

      Strengths:

      (1) The Prostate Cancer Atlas is a comprehensive collection of clinical data derived from RNA sequencing and serves as a powerful tool for conducting high-throughput analyses in this paper.

      (2) The Drosophila model used in this article is well established and has already been the subject of publications by the team. In addition to being an in vivo model, Drosophila offers a threefold advantage for this study: i) the presence of an accessory gland, similar to the prostate, which allows for the simulation of tumor formation and, in particular, extrusion mechanisms; ii) its regulation by a single sex steroid, ecdysone; iii) the genetic ability to modulate or inactivate ecdysone expression, which allows for a parallel to be drawn with hormonal deprivation in humans.

      (3) This study presents interesting and original findings. The data are, for the most part, of high quality.

      Weaknesses:

      (1) The Prostate Cancer Atlas, which is an essential tool in this study, was described only briefly - if at all - in both the introduction and the "Materials and Methods" section. The selection criteria used to distinguish CRPC or NEPC from adenocarcinoma in the Atlas or as determined by the authors, as well as the analytical methods, were not specified. It is therefore difficult to be convinced by the results, particularly those presented in Figures 1 and 2.

      (2) Although the hypothesis put forward by the authors - that the deprivation of sex hormones contributes to tumor progression - is strongly supported by the Drosophila model and by the in silico analysis of transcriptomic data from the Atlas, this concept still needs to be clinically validated by analyzing a series of prostate cancer samples, either through transcriptomic analysis or by tracking gene expression in histological sections.

      (3) With regard to the cells responsible for tumor escape, stem cells have been described as "candidates for the initiating resistant tumor growth" (lanes 50-55), but it is also essential to address the recent concept of "persistent cells". Indeed, these cells have been primarily associated with their tolerance to treatment (chemotherapy) and are referred to as "drug-tolerant cells". However, persistent cells could also correspond to cells that evade hormone therapy in the case of prostate cancer. This possibility should be discussed in the article.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, Vialat and collaborators study the role of steroid hormone signalling on the development of prostate cancer (patients) and of accessory gland tumours in Drosophila, a tissue functionally equivalent to the prostate. Mining publicly available prostate cancer expression data and using gene expression signatures, they uncover that androgen signalling is actually down-regulated in castration resistant prostate cancers (CRPC) compared to "primary" cancers, leading the authors to wonder whether down-regulation of canonical androgen signalling could represent an important event increasing tumour aggressiveness. They then take advantage of their recently published tumour model in the accessory gland of Drosophila adult males, in which cells are primed for tumorigenesis by the constitutive activation of the EGFR receptor, to test directly this hypothesis. They show that the genetic invalidation of ecdysone reception and signalling increases the aggressiveness of the "pre-cancerous" lesions, and that ecdysone-insensitive tumours present higher proliferation and initiate basal extrusion.

      Strengths:

      The authors bring original observations on the role of ecdysone signalling to prevent male accessory gland tumour development in Drosophila

      Weaknesses:

      (1) The link between the human data mining and Drosophila model is not straightforward.<br /> (2) Important information, in particular clinical information, is missing in the presentation of the cancer patients' data, making it complicated to grasp the solidity of the claims.<br /> (3) Data-mining insights should be validated by orthogonal approaches.<br /> (4) Ecdysone signalling activity should be monitored.

      While the two parts of the study both investigate the role of steroid signalling on tumour growth, the link remains slightly artificial. I think starting with Drosophila and then opening with some patient data would be better suited to the level of proof reached here, implying that the anti-tumoral role of steroids observed experimentally in the fly might be conserved based on data mining in patients, rather than trying to prove in the fly the hints gained from public data mining. Indeed, there are many important differences between the mammalian prostate and the fly accessory gland, as well as between sex hormone androgen signalling and developmental timing ecdysone signalling.

      The prostate cancer data mining and re-evaluation brings some interesting observations that appear to challenge the androgen driver, contrary to the vast amount of literature. Indeed, the authors observe an apparent decrease in androgen signalling in the more advanced states of the disease, in particular CRPC. In order to better evaluate its clinical relevance, more background on the tumours analysed should be provided.

      What treatments were received by the patients? Hormonotherapy? LH/RH analogues? +/- anti-androgens? Are these treatments still given when CRPC emerge and tissues were banked? Metastatic disease? Are these only primary tumours in situ? Are there metastases included in the analyses?

      Frequently, castration resistance is associated with alternatively spliced variants of the AR (AR-V7) that become constitutive and could bind to new AR-sensitive enhancers, even in the absence of androgen. Is the splice variant status of patients known, or could it be inferred from the expression data? Would there be different responses according to AR-V7 status?

      Regarding the signature used. Why not monitor PSMA, one of the major prostate cancer markers, which is regulated by AR?

      Finally, to consolidate the surprising observation that AR signalling is repressed in CRPCs, the authors should back these in silico predictions with orthogonal approaches such as histochemistry on patients' TMA or tissues from mouse models, monitoring AR activity.

      Regarding the fly experiments, the observation that ecdysone signalling depletion cooperates with EGFR-lambda activation to generate big overgrowths that delaminate basally without passing through the muscular sheet is interesting. However, several important controls need to be provided in order to support the claims:<br /> a) The authors should use an ecdysone reporter (ERE-LacZ, ERE-GFP...) to monitor and show that Ecdysone signalling is indeed lower in the tumours after genetic manipulations, or that it is higher in EGFR-lambda small clones.<br /> b) EcR is normally a repressor, which is turned into an activator in the presence of 20-hydroxyecdysone. The removal of EcR could lead to de-repression of genes and thus slightly activate the pathway. Monitoring ecdysone signalling activity is thus critical.<br /> c) The authors should also monitor the expression of Phantom, Shadow, Shade, and EcR in the different accessory glands (wild-type, EGFR-lambda, EGFR-lambda & EcR-RNAi). It is extremely surprising that systemic ecdysone has so little role since Phantom, Shadow, or Shade RNAi appear as potent as EcR-RNAi. This quantification has actually been performed for Sad in Figure S5, which is not even mentioned in the text. It should be done for Phtm.

      A UAS-yellow-RNAi (or similarly irrelevant RNAi) rather than UAS-GFP should be used as a control for the EcR, Phtm, Sad, Shd, and Tub RNAi. Indeed, loading the RNAi machinery could have some unexpected effects not controlled by the UAS-GFP.

      The authors should not use the term "sex steroid" when referring to ecdysone. It is a steroid hormone important for developmental timing and rate of growth, but is not a sex hormone, as sex is cell autonomously genetically determined in the fly.

    4. Author response:

      eLife Assessment

      This valuable study provides insights into the role of steroid signaling during tumorigenesis in the adult male drosophila accessory gland (functional equivalent of the prostate gland in mammals), hinting at a possible counterintuitive anti-tumoral role of sex hormones during prostate cancer in certain patients. While the Drosophila model provides an elegant way to study the hypothesis derived from the Cancer Atlas analysis, the analyses of public prostate cancer expression data are incomplete and critical knowledge on patients' treatment modalities and normalization across different datasets is missing. This work would be of interest to prostate cancer researchers as it suggests that the absence of androgen receptor signaling in humans could constitute a mechanism promoting tumor escape.

      We thank the reviewers for the time they have spent on the manuscript, the production of a public review and their useful recommendations. As a general goal for the corrected version, we will try to provide more insight on the data (especially the human data), and more controls, to strengthen our conclusions. We are aware of the lack of a definitive proof of the role of the apparent decrease in AR signaling on tumour progression, but hope that this manuscript will encourage medical scientists to test/challenge its counterintuitive results in large cohorts of tissues and mouse/human models.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this article, Vialat and his colleagues examine the early stages - which remain largely unknown - of the tumor escape process, particularly the basal extrusion of tumor cells following endocrine therapy for prostate cancer.

      They first used the "Prostate Cancer Atlas" database, which provides access to a vast amount of transcriptomic data, to perform high-throughput analyses. Interestingly, analyzing a series of Androgen Receptor (AR) target genes in castration-resistant prostate cancers, they concluded that the loss of the canonical AR signaling pathway may contribute to tumor resistance.

      Using a well-established model in Drosophila, they then replicated in vivo an endocrine therapy targeting the accessory gland by genetically inhibiting the expression of ecdysone, the only sex steroid present in Drosophila. These experiments induced basal extrusion similar to the mechanism observed in tumor escape in humans.

      These results suggest that the deprivation of sex steroids may play an important role in tumor progression.

      However, although the data from the "Prostate Cancer Atlas" constitutes a powerful tool that serves as the basis for this new concept, clinical validation using carefully selected human tumor samples would help strengthen the authors' conclusions.

      Strengths:

      (1) The Prostate Cancer Atlas is a comprehensive collection of clinical data derived from RNA sequencing and serves as a powerful tool for conducting high-throughput analyses in this paper.

      (2) The Drosophila model used in this article is well established and has already been the subject of publications by the team. In addition to being an in vivo model, Drosophila offers a threefold advantage for this study: i) the presence of an accessory gland, similar to the prostate, which allows for the simulation of tumor formation and, in particular, extrusion mechanisms; ii) its regulation by a single sex steroid, ecdysone; iii) the genetic ability to modulate or inactivate ecdysone expression, which allows for a parallel to be drawn with hormonal deprivation in humans.

      (3) This study presents interesting and original findings. The data are, for the most part, of high quality.

      Weaknesses:

      (1) The Prostate Cancer Atlas, which is an essential tool in this study, was described only briefly - if at all - in both the introduction and the "Materials and Methods" section. The selection criteria used to distinguish CRPC or NEPC from adenocarcinoma in the Atlas or as determined by the authors, as well as the analytical methods, were not specified. It is therefore difficult to be convinced by the results, particularly those presented in Figures 1 and 2.

      As the tool has been published in different articles, we chose to limit its description. However, we agree that explanation are necessary, that will be added in the new version. First, we initially used here just basic categories, in order to avoid any possible bias; so mCRPC includes rare DNPC and NECP patients. We will also put the data with true ARPC, with essentially the same results.

      For human data, we also expect to use transcriptomic data from an independent cohort to check whether the same loss of AR signaling occurs during progression. Furthermore, we consider to add the data showing that decrease in canonical AR signaling also (logically with the previous results) correlates with castration status or ADT exposure (these info are available on ProstateCancerAtlas). Interestingly, and this can be put in supplementary data, prostatecanceratlas detects changes in EMT genes or proliferation genes that are coherent with what is known about cancer progression, indicating that the apparent decrease in AR signaling should correspond to a real phenomenon.

      (2) Although the hypothesis put forward by the authors - that the deprivation of sex hormones contributes to tumor progression - is strongly supported by the Drosophila model and by the in silico analysis of transcriptomic data from the Atlas, this concept still needs to be clinically validated by analyzing a series of prostate cancer samples, either through transcriptomic analysis or by tracking gene expression in histological sections.

      There are many indirect evidences that loss of AR signaling induces tumor progression in mouse (as stated in the intro or the discussion of the manuscript). However, as suggested in the introduction of the letter, we believe that medical scientists are the most qualified to prove that sex steroid deprivation indeed induces tumor progression in human. We will add in any case data to at least reinforce this puzzling finding of a decrease in AR canonical signaling during progression.

      (3) With regard to the cells responsible for tumor escape, stem cells have been described as "candidates for the initiating resistant tumor growth" (lanes 50-55), but it is also essential to address the recent concept of "persistent cells". Indeed, these cells have been primarily associated with their tolerance to treatment (chemotherapy) and are referred to as "drug-tolerant cells". However, persistent cells could also correspond to cells that evade hormone therapy in the case of prostate cancer. This possibility should be discussed in the article.

      This is an interesting suggestion, which can be discussed: on the one hand, intrabasal cells may not have accumulated mutations to survive the loss of EcR signaling, as would do persistent cells. On the other hand, they strongly proliferate, and show no sign of senescence, behaving more like resistant cells. So, it does not look to us that we induced the appearance of persistent cells in the Drosophila accessory gland, except if these cells are quickly reactivating to give rise to intrabasal cells.

      Reviewer #2 (Public review):

      Summary:

      In this study, Vialat and collaborators study the role of steroid hormone signalling on the development of prostate cancer (patients) and of accessory gland tumours in Drosophila, a tissue functionally equivalent to the prostate. Mining publicly available prostate cancer expression data and using gene expression signatures, they uncover that androgen signalling is actually down-regulated in castration resistant prostate cancers (CRPC) compared to "primary" cancers, leading the authors to wonder whether down-regulation of canonical androgen signalling could represent an important event increasing tumour aggressiveness. They then take advantage of their recently published tumour model in the accessory gland of Drosophila adult males, in which cells are primed for tumorigenesis by the constitutive activation of the EGFR receptor, to test directly this hypothesis. They show that the genetic invalidation of ecdysone reception and signalling increases the aggressiveness of the "pre-cancerous" lesions, and that ecdysone-insensitive tumours present higher proliferation and initiate basal extrusion.

      Strengths:

      The authors bring original observations on the role of ecdysone signalling to prevent male accessory gland tumour development in Drosophila

      Weaknesses:

      (1) The link between the human data mining and Drosophila model is not straightforward.

      (2) Important information, in particular clinical information, is missing in the presentation of the cancer patients' data, making it complicated to grasp the solidity of the claims.

      (3) Data-mining insights should be validated by orthogonal approaches.

      (4) Ecdysone signalling activity should be monitored.

      While the two parts of the study both investigate the role of steroid signalling on tumour growth, the link remains slightly artificial. I think starting with Drosophila and then opening with some patient data would be better suited to the level of proof reached here, implying that the anti-tumoral role of steroids observed experimentally in the fly might be conserved based on data mining in patients, rather than trying to prove in the fly the hints gained from public data mining. Indeed, there are many important differences between the mammalian prostate and the fly accessory gland, as well as between sex hormone androgen signalling and developmental timing ecdysone signalling.

      This is an interesting suggestion. Actually, we first wrote the manuscript by starting with Drosophila data and then going to patients data, and previous reviewers said that this was not possible to directly go from Drosophila to human. So, we suppose that the real way to solve this will be by the validation or refutation of the data by other teams in different models.

      The prostate cancer data mining and re-evaluation brings some interesting observations that appear to challenge the androgen driver, contrary to the vast amount of literature. Indeed, the authors observe an apparent decrease in androgen signalling in the more advanced states of the disease, in particular CRPC. In order to better evaluate its clinical relevance, more background on the tumours analysed should be provided.

      This point is also important to Reviewer 1, and will be implemented.

      What treatments were received by the patients? Hormonotherapy? LH/RH analogues? +/- anti-androgens? Are these treatments still given when CRPC emerge and tissues were banked? Metastatic disease? Are these only primary tumours in situ? Are there metastases included in the analyses?

      Most CRPC come from metastatic sites. Most of the CRPC were treated by ADT. We chose to have an approach that included all the samples; but we will provide insight, whenever available, on these absolutely relevant questions.

      Frequently, castration resistance is associated with alternatively spliced variants of the AR (AR-V7) that become constitutive and could bind to new AR-sensitive enhancers, even in the absence of androgen. Is the splice variant status of patients known, or could it be inferred from the expression data? Would there be different responses according to AR-V7 status?

      As a first approach, from the cohorts that were used, it seems that in the PCA patients, there are around or less than 15% of patients harboring the AR-V7 driver. It could be of interest to test their behavior regarding the same set of genes, and it will be done if we can identify the patients.

      Regarding the signature used. Why not monitor PSMA, one of the major prostate cancer markers, which is regulated by AR?

      It will be done. PSMA behaves as the others, even though the drop between primary samples and ARPC samples is very limited and just statistically significant.

      Finally, to consolidate the surprising observation that AR signalling is repressed in CRPCs, the authors should back these in silico predictions with orthogonal approaches such as histochemistry on patients' TMA or tissues from mouse models, monitoring AR activity.

      As said previously, we believe that this specific work will be better done by medical scientists.

      Regarding the fly experiments, the observation that ecdysone signalling depletion cooperates with EGFR-lambda activation to generate big overgrowths that delaminate basally without passing through the muscular sheet is interesting. However, several important controls need to be provided in order to support the claims:

      Considering the comments regarding the fly experiments, we agree that, if experiences are taken individually, controls are lacking. However, we have to explain our strategy and why the results taken in their entirety have a significance. In our model of epithelial tumorigenesis, we have started to explore the EcR pathway after years of work on other pathways. At the first experiment (with the EcR RNAi line), we were struck by the intrabasal phenotype that did not occurred in our previous experiments, and especially for the 14 RNAi lines that we published in two independent articles on Ras/MAPK, Pi3K/Akt pathways and cholesterol metabolism. As justly said in the review, many unexpected effects can happen, so we decided to explore the role of five other genes of the same pathway to be sure of the reproducibility of the phenotype when we block the EcR pathway. The odds of having the same specific phenotype for 6 lines of the EcR pathway when there is always another phenotype for 14 lines targeting other pathways can be calculated: p = 0.00000494. So, the best control we offer, and it is largely significant, is the repetition of the experiments intended to downregulate the EcR pathway, that produce the same phenotypes independently of the target.

      Furthermore, all lines we used were previously tested, validated and most of the time published in other scientific works. This is essential to us, as one complexity of doing rare clones in an otherwise normal tissue is that decreasing an mRNA in less than 5% of the cells of course difficultly leads to a detectable drop in overall expression in the whole gland. This is also the reason why we always tried pairs of fly lines to block the receptor activity itself (RNAi EcR, RNAi Shd), the receptor's downstream targets (RNAi HR3, RNAi HR4), and the production of ecdysone (RNAi Sad, RNAi Phtm). As we validated RNAi Sad, we can try anyway to validate at least another RNAi of another category. Furthermore, we did use a RNAi White control: it behaves in the same way as the GFP control. We will put the results comparing the two lines in supplementary data.

      (a) The authors should use an ecdysone reporter (ERE-LacZ, ERE-GFP...) to monitor and show that Ecdysone signalling is indeed lower in the tumours after genetic manipulations, or that it is higher in EGFR-lambda small clones.

      This would be of interest to validate that the 6 lines are behaving in the same way (at least, they give similar phenotypes). However, in the adult accessory gland, ERE activity is largely lower than during development (DOI: 10.1016/j.jinsphys.2011.03.027), and to be able to decrease it, authors had to express notoriously strong dominant-negative EcR-DN. We can try the experiment but are really not persuaded that we will be able to see a drop of activity with only a decrease of expression of the gene. If we can think of another solution that could be more efficient, we will try it as the idea is of course interesting.

      (b) EcR is normally a repressor, which is turned into an activator in the presence of 20-hydroxyecdysone. The removal of EcR could lead to de-repression of genes and thus slightly activate the pathway. Monitoring ecdysone signalling activity is thus critical.

      Actually, there are different EcR isoforms. EcR-B1 is generally considered as the main activator of the pathway, as EcR-A is a repressor of the pathway. The EcR RNAi line which was used does not target a specific isoform.

      (c) The authors should also monitor the expression of Phantom, Shadow, Shade, and EcR in the different accessory glands (wild-type, EGFR-lambda, EGFR-lambda & EcR-RNAi). It is extremely surprising that systemic ecdysone has so little role since Phantom, Shadow, or Shade RNAi appear as potent as EcR-RNAi. This quantification has actually been performed for Sad in Figure S5, which is not even mentioned in the text. It should be done for Phtm.

      The levels of ecdysone are tenths of times lower in adult compared to the peaks during embryogenesis or metamorphosis. And one source of production is the epithelial cells of the accessory glands themselves. Considering that EcR is expressed in all the cells of the accessory gland (epithelial cells and muscle cells), it seems plausible that there is only a very little amount of ecdysone that can in fact be available for the other epithelial cells.

      A UAS-yellow-RNAi (or similarly irrelevant RNAi) rather than UAS-GFP should be used as a control for the EcR, Phtm, Sad, Shd, and Tub RNAi. Indeed, loading the RNAi machinery could have some unexpected effects not controlled by the UAS-GFP.

      The NLSGFP line we used here is the one we already published twice (and we compared it to RNAi lines), and this is the reason why we used this already validated control. However, we tested a RNAi White line, and it behaves in the same way.

      The authors should not use the term "sex steroid" when referring to ecdysone. It is a steroid hormone important for developmental timing and rate of growth, but is not a sex hormone, as sex is cell autonomously genetically determined in the fly.

      In human, sex hormones control sexual differentiation (up to adult characteristics) and sexual reproduction. In Drosophila, ecdysone controls sexual reproduction in both sexes and sexual differentiation at least in the female (DOI: 10.1007/s004270050186). From these results, we do not think that saying it is a sex steroid (not a sex hormone) is ill suited. We intend to precise what we put in this term in the introduction to avoid overinterpretation from our part.

    1. eLife Assessment

      In this solid work, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role for two conserved acidic residues rather than a single one. This valuable study used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. We thank the authors for revising the manuscript according to the reviewers' comments. We have no further comments.]

      In this manuscript the applicants study two residues in the GHKL ATPase active site of Aq MutL and GyrB, and argue that the catalytic base function is shared between two conserved acidic residues that are 3 residues apart.

      In the manuscript, the authors generated mutant versions in MutL and GyrB (both ala and the appropriate Asn/Gln version) and performed ATPase analysis. They also generated high resolution crystal structures of the GyrB NTD with AMPPnP for WT and mutants of the two acidic residues. The data show that mutation in either of these residues does not fully kill activity (with the exception of the Alanine mutation of the first of the two, that interferes with ATP (or AMPPnP) binding). When the acidic residues are mutated to Asn/Gln, the catalytic water can still be positioned, and hence these mutants are more active than the Ala mutants. In both cases the double mutation is catalytic dead.

      The authors then perform phylogenetic analysis and ancestral gene reconstruction and based on this they argue that HSP90 forms a different class of GHKL ATPases, and lost rather than gained this separate status.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role of two conserved acidic residues rather than one. The authors have used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism.

      Through a detailed re-analysis of their previously published structure of the aqMutL NTD (ATPase domain) in complex with AMPPCP, they identified Glu29 and Glu32 as interacting with nucleophilic water for the catalysis. The authors carefully dissected the respective roles of these two acidic residues with a series of site-directed mutations. Mutations at Glu29 impaired ATPase activity without affecting protein secondary structure or ATP binding in the case of the E29Q mutant. Moreover, mutations at Glu32 did not affect secondary structure (except for E32G) but reduce ATPase activity. Activity was abolished when both residues (E29Q/E32Q) are mutated.

      The authors extended their study to another GHKL ATPase, aqGyrB. Their findings further supported the cooperative function of the corresponding acidic residues in aqGyrB (Glu48 and Asp51) during ATP hydrolysis. Mutation of these residues partially impaired ATP hydrolysis without affecting protein secondary structure. ATPase activity was completely lost in the double mutant E48Q/D51M. While the E48Q mutant retained the ability to bind ATP, the E48A mutant did not. High-resolution structures of the WT and E48A, E48Q, D51A and D51N mutants of the aqGyrB NTD demonstrated that nucleophilic water positioning depended on these residues. E48 played a dominant role in water positioning and is critical for stabilising ATP lid formation and associated conformational changes, whereas D51 contributed cooperatively to catalysis.

      The authors investigated the functional impact of mutating the corresponding residues in the human MutL homologs PMS2 and MLH1. Clinical variants consistently exhibited reduced or abolished ATPase activity, providing a potential molecular basis for Lynch syndrome, through impaired DNA mismatch repair.

      Lastly, through evolutionary analysis, the authors inferred that the second acidic residue was likely present in the common ancestor of MutL, GyrB, and MORC proteins, but was lost in the case of Hsp90.

      Strengths:

      (1) This study contains a detailed structural and biochemical analysis of a biologically important set of GHKL ATPases. The authors identify a second acidic residue that is conserved and contributes to catalysis in a large subset of GHKL ATPases. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed, which involves cooperative and partially overlapping roles for the catalytic residue pair. This revised mechanistic model is invaluable for the interpretation of clinical variants of GHKL ATPases such as PMS2 and MLH1.

      (2) The work described was performed to an excellent and rigorous technical standard. The structural and biochemical data are sound. The evidence supporting the claims is compelling.

      Weaknesses:

      (1) The identification in this study of a second acidic residue contributing to catalysis but not absolutely essential for catalysis is a useful finding. However, given that many structures of GHLK ATPases have been determined with different nucleotide analogs bound and that the essential role of the first acidic residue is well established, the importance and scope of the advances described here remain focused within the field of study of GHKL ATPases.

      (2) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of any effect of disease-linked mutations in GHKL ATPases would have strengthened this study.

      (3) The effect of other aqMutL NTD E32 mutants, particularly, the E32K mutant on ATP binding remains unclear, although experimental assessment of nucleotide binding would be challenging due to the high protein concentrations required for the equilibrium dialysis assay.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript the applicants study two residues in the GHKL ATPase active site of Aq MutL and GyrB, and argue that the catalytic base function is shared between two conserved acidic residues that are 3 residues apart.

      In the manuscript, they generated mutant versions in MutL and GyrB (both ala and the appropriate Asn/Gln version) and performed ATPase analysis. They also generated high resolution crystal structures of the GyrB NTD with AMPPnP for WT and mutants of the two acidic residues. The data show that mutation in either of these residues does not fully kill activity (with the exception of the Alanine mutation of the first of the two, that interferes with ATP (or AMPPnP) binding). When the acidic residues are mutated to Asn/Gln, the catalytic water can still be positioned, and hence these mutants are more active than the Ala mutants. In both cases the double mutation is catalytic dead.

      The authors then perform phylogenetic analysis and ancestral gene reconstruction and based on this they argue that HSP90 forms a different class of GHKL ATPases, and lost rather than gained this separate status.

      Strengths:

      The biochemical analysis seems solid.

      Weaknesses:

      A major question that remains, is why the mutations have so much more detrimental effect in MutL (100-fold lower kcat/KM) than they do in GyrB (3-fold lower). Can the authors explain this? Doesn't this argue against the proposed catalytic conservation?

      The authors need to discuss this issue explicitly to make it clear that conservation of the mechanism is not complete and that other interpretations are possible.

      The structure figures all have omit maps for just the AMPPnP and the water, whereas the density for the the acidic residues and their mutants are not shown.

      This has been addressed.

      There are some issues with figure S2B and S5.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Fukui et al. re-examined the ATP hydrolysis mechanism in GHKL ATPases, revealing a cooperative role of two conserved acidic residues rather than one. The authors have used a range of biochemical and structural techniques on various mutants from different members of the GHKL ATPase family to test and validate their proposed mechanism.

      Through a detailed re-analysis of their previously published structure of the aqMutL NTD (ATPase domain) in complex with AMPPCP, they identified Glu29 and Glu32 as interacting with nucleophilic water for the catalysis. The authors carefully dissected the respective roles of these two acidic residues with a series of site-directed mutations. Mutations at Glu29 impaired ATPase activity without affecting protein secondary structure or ATP binding in the case of the E29Q mutant. Moreover, mutations at Glu32 did not affect secondary structure (except for E32G) but reduce ATPase activity. Activity was abolished when both residues (E29Q/E32Q) are mutated.

      The authors extended their study to another GHKL ATPase, aqGyrB. Their findings further supported the cooperative function of the corresponding acidic residues in aqGyrB (Glu48 and Asp51) during ATP hydrolysis. Mutation of these residues partially impaired ATP hydrolysis without affecting protein secondary structure. ATPase activity was completely lost in the double mutant E48Q/D51M. While the E48Q mutant retained the ability to bind ATP, the E48A mutant did not. High-resolution structures of the WT and E48A, E48Q, D51A and D51N mutants of the aqGyrB NTD demonstrated that nucleophilic water positioning depended on these residues. E48 played a dominant role in water positioning and is critical for stabilising ATP lid formation and associated conformational changes, whereas D51 contributed cooperatively to catalysis.

      The authors investigated the functional impact of mutating the corresponding residues in the human MutL homologs PMS2 and MLH1. Clinical variants consistently exhibited reduced or abolished ATPase activity, providing a potential molecular basis for Lynch syndrome, through impaired DNA mismatch repair.

      Lastly, through evolutionary analysis, the authors inferred that the second acidic residue was likely present in the common ancestor of MutL, GyrB, and MORC proteins, but was lost in the case of Hsp90.

      Strengths:

      (1) This study contains a detailed structural and biochemical analysis of a biologically important set of GHKL ATPases. The authors identify a second acidic residue that is conserved and contributes to catalysis in a large subset of GHKL ATPases. An updated and extended mechanistic model of ATP hydrolysis by this class of enzymes is proposed, which involves cooperative and partially overlapping roles for the catalytic residue pair. This revised mechanistic model is invaluable for the interpretation of clinical variants of GHKL ATPases such as PMS2 and MLH1.

      (2) The work described was performed to an excellent and rigorous technical standard. The structural and biochemical data are sound. The evidence supporting the claims is compelling.

      Weaknesses:

      (1) The identification in this study of a second acidic residue contributing to catalysis but not absolutely essential for catalysis is a useful finding. However, given that many structures of GHLK ATPases have been determined with different nucleotide analogs bound and that the essential role of the first acidic residue is well established, the importance and scope of the advances described here remain focused within the field of study of GHKL ATPases.

      (2) The authors assessed the consequences of variants in the human MutL homologs PMS2 and MLH1, but various other human GHKL ATPases contain clinically relevant variants, some of which have stronger disease associations than the mutations examined in this study. A broader analysis of any effect of disease-linked mutations in GHKL ATPases would have strengthened this study.

      (3) The effect of other aqMutL NTD E32 mutants, particularly, the E32K mutant on ATP binding remains unclear, although experimental assessment of nucleotide binding would be challenging due to the high protein concentrations required for the equilibrium dialysis assay.

      We are grateful to the Editors and reviewers for their careful assessment of our revised manuscript and for identifying the remaining points that required clarification. We have addressed each of these comments in the present revision. We believe that these revisions have resolved the remaining concerns and have further improved the clarity and accuracy of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please discuss the large difference in the effect of the mutants on activity explicitly

      According to the reviewer’s suggestion, we have added the following discussion to the revised manuscript:

      “Although mutation of the two acidic residues impaired the ATPase activity in both aqMutL and aqGyrB, the magnitude of the effects differed substantially, with much greater reduction in the catalytic efficiency in aqMutL than in aqGyrB (Table 1). The molecular basis for this quantitative difference is currently unclear. One possible explanation is that subtle differences in the active-site architecture and surrounding residues alter the relative contribution of each acidic residue to catalysis, allowing aqGyrB to tolerate perturbation of either residue more effectively than aqMutL.” (p.6 line 257-262 in the revised manuscript)

      (2) Figure S2B is a completely different view from the other panels, please provide the correct one.

      We have revised Supplementary Fig. S2B so that the E48A structure is now shown from a viewpoint as similar as possible to those used in the other panels. We note, however, that the E48A structure cannot appear completely identical to the other panels because the E48A mutant does not bind the ATP analog and therefore does not undergo the nucleotide-binding-associated conformational changes observed in the other structures.

      (3) S5 : It is not clear to me what is meant by " The scale bar indicates the number of amino acid substitutions per site." : there is a 'Tree Scale 1" but no other numbers in my version.

      We thank the reviewer for pointing out that the scale bar in Supplementary Fig. S5 was insufficiently explained. The value “1” in the tree scale corresponds to a branch length of one amino acid substitution per site. To avoid ambiguity, we have revised the scale-bar label in Supplementary Fig. S5 to explicitly indicate “1 substitution/site” and have clarified its meaning in the figure legend:

      “Branch lengths are proportional to the evolutionary distances inferred by IQ-TREE. The scale bar represents an evolutionary distance of one amino acid substitution per site.” (p. 22 line 767-769 in the revised manuscript)

      Reviewer #2 (Recommendations for the authors):

      (1) P. 9, in the "Data Accessibility Statement", all three PDB codes (23UX, 23UY, and 23UZ) should be listed.

      We thank the reviewer for this comment. We carefully rechecked the Data Accessibility Statement and confirmed that all three PDB accession codes (23UX, 23UY, and 23UZ) are included in the statement.

      (2) Supplementary Figures S2 and S3. The authors have written "Asn33" and "Asn52", instead of "Glu32" and "Asp51" in both the figure and figure legend of Supplementary Figure S3. They have also written "TND" instead of "NTD" in the figure legend.

      “Asn33” and “Asn52” in Supplementary Fig. S3 are not typographical errors. Asn33 in aqMutL and Asn52 in aqGyrB are the residues that directly coordinate the Mg<sup>2+</sup> ion and are distinct from the acidic residues discussed in this paper. To avoid confusion, we have added the following sentence to the legend of Supplementary Fig. S3:

      “These Mg<sup>2+</sup>-coordinating asparagine residues are adjacent to, but distinct from, the second acidic residues Glu32 in aqMutL and Asp51 in aqGyrB examined in this study.” (p. 21 line 749-751 in the revised manuscript)

      We have also corrected the typographical error “TND” to “NTD” in the figure legend. (p. 21 line 749 in the revised manuscript)

    1. eLife Assessment

      This work provides important findings that reveal a mechanistic model of how VDAC2 influences BAX activation in mitochondria during BAX-mediated apoptosis. The experimentally supported model is convincing, with several experiments strengthening its validity. Major strengths include the complementary biochemical, biophysical, structural, and computational approaches, as well as the study's broad relevance.

    2. Reviewer #1 (Public review):

      Summary:

      The authors elegantly demonstrate a biochemically reconstituted approach to showcase the VDAC2-BAX interaction using lipid nanodiscs. The reconstitution method is specific to VDAC2 (and not VDAC1) and can capture several structural conformations. The authors show that the VDAC2-BAX heterodimer is sufficient for the direct capture and stabilization of BAX on the outer mitochondrial membrane by VDAC2. Their structural model demonstrates that a GXXXA motif within α-helix 9 of BAX drives its interaction with the β-barrel interface of VDAC2 in the membrane. AlphaFold 3 models suggest BAX adopts several distinct conformations, notably including both a strongly pore-occluding state and a loosely pore-occluding state. Functionally, electrophysiology experiments suggest that the addition of BAX modulates the voltage-gating function of VDAC2 by reducing conductance. Finally, conformation-specific antibodies and cross-linking mass spectrometry capture these structural rearrangements of BAX, reinforcing the proposed structural model.

      Strengths:

      Overall, the manuscript provides a solid structural and molecular rationale for BAX recruitment to VDAC2 and its subsequent oligomerization.

      Weaknesses:

      The authors have not sufficiently discussed key protein modifications during apoptosis in detail, especially regarding residues implicated in phosphorylation and their impact on VDAC2 association.

      Overall, the manuscript is well written and presents an elegant biochemical and biophysical approach to identifying key functional states of the VDAC2-BAX complex. However, certain key functions of the complex are not extensively discussed or accounted for in the final model. For instance, components of this complex are phosphorylated in response to apoptotic or anti-apoptotic cues. Specifically, phosphorylation of S184 (located within the critical α-helix 9), T167, and S163 has key functions in promoting or preventing outer mitochondrial membrane translocation. How do the authors reconcile their structural models with the functional states of the complex generated in response to these signaling cues? This is particularly relevant given that the expression systems used here presumably yield proteins lacking these post-translational modifications (PTMs). The authors should consider running AlphaFold 3 predictions that incorporate key PTMs and discuss their potential functional impact. In its current state, the manuscript implies that unmodified BAX is sufficient for membrane translocation. Clarifying how PTMs influence pore occlusion and 6A7 epitope accessibility would significantly enrich this body of work.

    3. Reviewer #2 (Public review):

      In this study, the authors aimed to elucidate the precise molecular details underlying the BAX-VDAC2 interaction and subsequent BAX activation. To achieve this, they successfully combined AlphaFold3 structural modeling, cross-linking mass spectrometry, site-directed mutational screening, biochemical assays, and functional electrophysiology experiments.

      The authors demonstrate that the direct interaction between BAX and VDAC2 is fully autonomous, occurring independently of any additional mitochondrial or cellular proteins. Their findings suggest that BAX exists on the outer mitochondrial membrane in two distinct populations: loosely membrane-associated, and tightly stabilized via its specific interaction with VDAC2. Crucially, the data overturn historical assumptions by demonstrating that BAX does not insert into the internal VDAC2 channel pore. Instead, the BAX α9 helix docks onto the lipid-facing outer surface of the VDAC2 β-barrel.

      Interestingly, while the anchor is external, the soluble domain of BAX physically blocks the pore opening, leading to the observed occlusions of the VDAC channel. Following this docking event, the N-terminal 6A7 epitope of BAX becomes exposed, signaling a conformationally active state. However, the authors show that this structural activation does not trigger an immediate release from VDAC2 or prompt immediate oligomerization. Rather, BAX is maintained in a pre-oligomeric, primed intermediate state while bound to VDAC2. What ultimately regulates the release of this primed intermediate from VDAC2 to allow full oligomerization and pore formation remains an open question.

      Altogether, this study provides pivotal mechanistic insights, clearly defining VDAC2 as a key checkpoint regulator of mitochondrial apoptosis.

    4. Reviewer #3 (Public review):

      Summary:

      The authors are trying to provide the molecular basis for the emerging role of VDAC2 in mitochondrial apoptosis. They use multiple approaches from biochemistry, cell biology, and structural biology.

      Strengths:

      Isolating the VDAC2-BAX complex and providing the molecular basis of this interaction in mitochondrial apoptosis is pretty innovative and significant.

      The authors have tried to validate their results using multiple approaches, which corroborates the quality of the study.

      Weaknesses:

      The scientific data and its presentation could be improved.

    1. eLife Assessment

      This paper provides a useful assessment of the role of macrophages in the development of functional paw withdrawal, which is a feature in the mouse model of vincristine-induced peripheral neuropathy (VIND). The authors suggest a role for E-selectin as a critical molecule promoting VIPN, acting via immune cell recruitment and inflammasome-mediated signalling; however, their data on NLRP3 inflammasome activation in vivo remains somewhat incomplete. This study will be of interest for immunologists.

    2. Reviewer #1 (Public review):

      This paper looks at the effect of vincristine-induced peripheral neuropathy (VIND), a common effect of cancer therapy. The authors performed in vivo experiments in mice by injecting them with vincristine sulphate i.p. (+/- various inhibitors or antibodies) or E-selectin intraplantar (i.pl.), and in vitro experiments using dorsal root ganglia (DRG) neurons and bone marrow derived macrophages (BMDMs).

      Inhibition of E-selectin with antibodies or genetic depletion reduced the accumulation of F4/80+ macrophages in the DRG and sciatic nerves (located beside the spine) after vincristine administration, and attenuated the mechanical hypersensitivity (paw withdrawal).

      The authors went on to perform spatial transcriptomics on isolated DRG neurons and found some pathways changed. E-selectin injected directly intraplantar (i.pl.) mimicked the effect of vincristine administration on the mechanical hypersensitivity. Whereas chlodronate depletion of myeloid cells reduced these changes in the E-selectin model. Using LPS priming before vincristine in BMDMs in vitro, the authors demonstrate an increase in many cytokines, including IL-1beta (typically associated with the formation of an inflammasome) and elevated p-NFkB. Finally, treatment of mice with anakinra (which neutralizes IL-1beta) also attenuated the E-selection-induced reduction in mechanical hypersensitivity when injected i.pl.

      The authors address an important aspect that after cancer therapy, there can be peripheral nerve damage that has lasting consequences for patients, although the precise mechanism is unknown. The authors delineate that E-selectin has an important role in the mouse model, where depletion or inhibition attenuated the negative effect of vincristine (i.p.) on mechanical hypersensitivity (paw withdrawal). Administration of E-selectin into the foot (i.pl.) also mimicked the changes observed in the vincristine-treated mice. There seems to be a role for macrophages, as they were associated with DRGs in vivo, and depletion attenuated motor deficits in the E-selectin injection model.

      However, I am not convinced by the data supporting some of the conclusions drawn by the authors, particularly on the role of the NLRP3 inflammasome in their in vivo model.

      Main points:

      (1) The initial experimental paradigm looks at the DRG neurons, which are located by the spine, and from there the foot pad is examined in subsequent experiments. It would be relevant to show whether the foot pad is altered in the vincristine-treated mice and whether the infiltration of myeloid cells that was demonstrated at DRGs is also observed in the foot in the vincristine model. Otherwise, the mechanism being investigated in the vincristine model, which might have similar functional results (paw withdrawal), but the mechanism behind both could be completely different.

      (2) The rationale of performing spatial sequencing on DRG neurons isolated from vincristine mice is unclear. It is likely that more information could have been obtained from looking at sections from these animals, and there would be a better link to the experiments on BMDMs which follow afterwards. Indeed, the spatial data does not seem to play a key role in the study. There is not a clear link between it (which was carried out on DRGs) and the later focus on macrophages and indeed the NLRP3 inflammasome.

      (3) The authors suggest that the E-selectin is enhancing NFkB-induced priming of the NLRP3 inflammasome. LPS+vincristine increased IL-1b release from BMDMs in vitro, which was elevated in the presence of E-selectin. E-selectin also increased ASC speck formation by approx. 20% in vitro. The ASC speck formation in vitro was blocked by MCC950, a specific NLRP3 inhibitor, but the authors went on to use anakinra in vivo using the E-selectin i.pl. model. It is really unclear why the switch to anakinra occurred for the in vivo work, as blocking IL-1b is central to many inflammatory pathways, not just NLRP3. Use of MCC950 would have been more appropriate to demonstrate that negative effects on mechanical function are mediated by the NLRP3 inflammasome. As there were no readouts of NLRP3 inflammasome activity measured in any of the mice in vivo (e.g. local ASC specks, IL-1b release, western blot of typical inflammasome components such as IL-1b, caspase-1, ASC or gasdermin D) either at the DRG site or the foot, we cannot say that the cell culture data mimics or models the in vivo conditions at this time.

      (4) Additionally, the reliance on the E-selectin administration models for the second half of the paper is curious. It would have been relevant to test whether the immune-modulating inhibitors could also attenuate the vincristine-induced effects on mechanism hypersensitivity, to better link the E-selectin model with the vincristine one.

      (5) Details are missing from the figure legends and the methods. The concentrations of compounds used in cell culture and exposure times are not clear.

    3. Reviewer #2 (Public review):

      Summary:

      Using antibody treatments, genetic models and in vitro studies, the authors convincingly show that E-Selectin is a driver of VIPN.

      Strengths:

      In vivo studies are robust. Antibody studies as well as the inflammasome-related work are very well done.

      Weaknesses:

      (1) The spatial transcriptomics data need to be improved in terms of visualization.

      The authors should plot genes that define the cell types in the sequencing dataset, and also show the cellular map on the cut, not only UMAPs. Can the authors not use the n=3 samples per condition to perform some statistical analyses? While the CellChat analyses are informative, they are difficult to read. Fold change over control when comparing conditions and pathways may be a better way of visualization.

      (2) In Figure 1B, the % DAB is not easy to understand. It would be better to do the staining also via IF, and then maybe to count nuclei. The corresponding figures shown in the supplemental figure are not convincing. If F4/80 is not working well, Iba1 could be an alternative.

      (3) The authors should define the background of all the mice used. Have they been back-crossed to B6j mice? One cannot compare C57BL6J mice with full knockouts if they are not littermate controls. Especially immunological responses are completely dependent on the background of mice. See e.g. PMID: 40568896.

      (4) Could the authors comment on the role of ICAM2? Why was this not tested as well?

    1. eLife Assessment

      This important study examines the dynamics of late endosomes and lysosomes (LEL) in cultured astrocytes, leading to a model of how synaptic activity might regulate astrocytic specializations surrounding synapses via the lysosomal calcium channel TRPML1, LEL motility, and proteins involved in regulating actin-membrane association. The experimental work is convincing, including appropriate controls and combining different tools (multiple markers, pharmacology, genetic knock-downs) to test each conclusion, although issues with the specificity of some of the drugs or confounding factors arising from genetic knock-downs limit the strength of some conclusions. Dysfunction in glia-neuron crosstalk as well as lysosomal dysfunction are increasingly recognized as highly relevant for neurodegenerative disease development, so this manuscript will be of interest to the wider neuroscience community, as well as scientists interested in the cell biology of astrocytes and lysosomes.

    2. Reviewer #1 (Public review):

      Summary:

      Late endosomes and lysosomes (LEL) are dynamic organelles with critical roles in cell physiology via transport of cargos to various destinations, degrading cargos, and as calcium stores. The latter is a less studied function of LELs, and virtually nothing is known about LEL function and transport in astrocytic processes. This manuscript investigates the dynamics of LELs in astrocyte processes co-cultured with neurons and finds that the lysosomal calcium channel Trpml1 regulates their positioning near astrocytic specializations (PAPs) downstream of synaptic activity.

      Strengths:

      Rigorous and well-controlled study of an understudied area of cellular neuroscience, namely regulation of organelle transport in astrocytes to shape synaptic environment and functionality.

      Weaknesses:

      Some of the same mechanistic links have been probed in neurons and other cell types, but astrocyte cell biology is still less extensively studied, making this an important contribution. Currently, only cultured astrocytes are being investigated.

    3. Reviewer #2 (Public review):

      Summary:

      Many ion channels/transporters in endosomes and lysosomes remain poorly understood. Even though the functional importance of endo-lysosomes in astrocytes has been recently recognized in many neurological diseases including lysosomal storage disorders and neurodegenerative diseases, the basic biology has not been much studied. In this manuscript, Spivey et al. showed how one of the key ion channels i.e., TRPML1, regulates endosomes and lysosomes (LELs) in astrocytes -which are also not much investigated in the field, compared to neurons- and its activity may further affect the structure of astrocytes and synaptic activity of neighboring neurons. The authors elegantly use multiple controls from agonists, antagonists, and TRPML1 knockdown to corroborate the results. The data are very strong and well supported. I believe that revising a few parts of the manuscript will greatly augment the significance of this work.

      Strength:

      The rigor of the study and data using various controls. The quality of the data is also very impressive.

      Weakness:

      A limitation of the study is that the mechanistic experiments rely primarily on overexpression or genetic knockdown of TRPML1 approaches, both of which alter TRPML1 abundance on endolysosomal membranes and potentially introduce artifacts related to protein level, affecting luminal ion homeostasis. While the overall conclusions are convincing, validation in a genetic MCOLN1 knockout mouse model would provide more definite evidence for the proposed mechanism.

    4. Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors assess how TRPML1 influences lysosome trafficking and function in astrocytes using a neuron-astrocyte coculture system. Specifically, it was found that activation of TRPML1 reduces LEL motility in astrocyte branches while TRPML1 increases it, and a model was proposed in which TRPML1 coordinates lysosomal positioning along astrocyte branches, enabling LELs to locally regulate actin-membrane linkers that influence PAP (peripheral astrocyte processes) structure and plasticity. The authors primarily use pharmacology to buttress their findings.

      Strengths:

      TRPML1 is currently under investigation as a potential therapeutic target for lysosomal storage disorders and neurodegenerative diseases such as AD and PD. The manuscript is hence timely and important as the crosstalk between glia cells and neurons is likewise increasingly recognized as disease-relevant.

      Weaknesses:

      Unfortunately, experiments performed to investigate the role of TRPML1 are based purely on pharmacological tools. No bona fide KO data are available. At least for some experiments, this should be done. MLIV iPSC lines are available.

      If no KO controls are being provided, at least the authors shall use ML1-SA1 (EVP-169) as an agonist instead of ML-SA1, because ML-SA1 activates all three TRPML channels, which would be a major problem for this study. Likewise, the ML-SI3 antagonist also has effects on other TRPMLs. EDME is a blocker with higher specificity for TRPML1.

      When using pharmacological tools, the original papers relating to these tools may be cited. For example:

      - For ML-SA1 https://www.nature.com/articles/ncomms1735 and for MK6-83 https://www.nature.com/articles/ncomms5681

      - ML-SI3 as a blocker for TRPML1 seems problematic: https://pubmed.ncbi.nlm.nih.gov/33187805/ ; an alternative may be: https://www.nature.com/articles/s41598-021-87817-4

      No reference is made to other important regulators of lysosomal cation homeostasis such as TPC1 or TPC2, two other Ca2+/Na+ permeable cation channels, or other TRPML channels. The authors should provide some data supporting or excluding a role of these channels.

    1. eLife Assessment

      This study provides a useful advance in generating mouse oligodendrocytes by direct lineage conversion from cortical astrocytes. The authors demonstrate that Sox10 converts astrocytes to MBP+ oligodendrocytes, whereas Olig2 expression converts astrocytes to PDFRalpha+ oligodendrocyte progenitor cells. The data supporting the conclusions are convincing with comprehensive transcriptomic analysis.

    2. Reviewer #1 (Public review):

      Bajohr and colleagues propose a transcription factor-driven approach to generating bonafide oligodendrocyte lineage cells (OLCs) from primary mouse astrocytes. Ectopic expression of Olig2, Sox10, or Nkx6.2 in isolated astrocytes produced a range of OLC-like cell states, with Sox10 emerging from lineage tracing and single cell RNA sequencing experiments as the most successful transcription factor in driving direct lineage reprogramming. The authors strengthened their claims with an unbiased, deep learning perturbation model to predict genetic drivers of the astrocyte cluster to OLC cluster transition observed in their scRNA seq dataset. Here, Sox10 surfaced in the top ten correlated genes, and the top transcription factor, mediating this fate shift. Altogether, this paper presents an interesting approach to generate OLCs, a cell type historically difficult to procure, from primary mouse astrocytes to study this lineage in development and disease and perhaps repopulate it in dysmyelinating conditions. While this certainly addresses a technical gap in the field, authors defined iOLCs as ones with lineage-specific gene expression and morphological characteristics, lacking any functional analysis to assess the reprogrammed cells' capacity to myelinate. This comment and other critiques are discussed below.

      While Sox10 and Mbp expression in iOLCs, as confirmed by IHC, is a promising result suggesting that ectopic Sox10 instructs transduced cells to develop into cells of myelinating potential, functional confirmation is essential. As mentioned in the discussion, the absence of a substrate for myelination may have also contributed to the low DLR efficiency. Co-culturing Sox10 iOLCs with primary neurons and examining the cells' potential to engage and enwrap axons would greatly strengthen the authors' claim that this could be an effective therapeutic approach to myelin regeneration in vivo, or even a technical approach to studying myelin dynamics in vitro.

      In Figure 1B, it appears that Mbp expression in tdTomato+ cells decreases in Sox10 transduced iOLs during the observed time period. Can the authors elaborate on this result, given that MBP expression is crucial for myelination and should, if anything, increase with time?

      The authors acknowledge that there is a conversion of tdTomato- zsGreen+ cells with an astrocyte-like morphology to OLC cells expressing Mbp following Sox10 induction (Supplementary figure 5C,D). While they note the diversity of the astrocyte lineage in the discussion, further analysis should be applied to this subset of cells to confirm the subset of astrocyte or progenitor-like cell type that gives rise to their cell endpoint of interest (Sox10-driven Mbp+ iOLs).

      Finally, ectopic expression of Olig2 and Sox10 in primary astrocytes resulted in very different OLC subtypes, as evidenced by OLC marker expression seen in IHC and the subclustering of these cell types in scRNA seq. Although this diversity in OLC type and generation efficiency follows with previous reports showing that these two transcription factors vary in effect, might the authors further discuss this discrepancy given that the two transcription factors regulate one another (as mentioned in the introduction) and should theoretically give rise to more similar cells? Perhaps due to the lower specificity of Olig2 in marking a pure OLC population relative to Sox10?

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Bajohr and colleagues propose a transcription factor-driven approach to generating bonafide oligodendrocyte lineage cells (OLCs) from primary mouse astrocytes. Ectopic expression of Olig2, Sox10, or Nkx6.2 in isolated astrocytes produced a range of OLC-like cell states, with Sox10 emerging from lineage tracing and single cell RNA sequencing experiments as the most successful transcription factor in driving direct lineage reprogramming. The authors strengthened their claims with an unbiased, deep learning perturbation model to predict genetic drivers of the astrocyte cluster to OLC cluster transition observed in their scRNA seq dataset. Here, Sox10 surfaced in the top ten correlated genes, and the top transcription factor, mediating this fate shift. Altogether, this paper presents an interesting approach to generate OLCs, a cell type historically difficult to procure, from primary mouse astrocytes to study this lineage in development and disease and perhaps repopulate it in dysmyelinating conditions. While this certainly addresses a technical gap in the field, authors defined iOLCs as ones with lineage-specific gene expression and morphological characteristics, lacking any functional analysis to assess the reprogrammed cells' capacity to myelinate. This comment and other critiques are discussed below.

      While Sox10 and Mbp expression in iOLCs, as confirmed by IHC, is a promising result suggesting that ectopic Sox10 instructs transduced cells to develop into cells of myelinating potential, functional confirmation is essential. As mentioned in the discussion, the absence of a substrate for myelination may have also contributed to the low DLR efficiency. Co-culturing Sox10 iOLCs with primary neurons and examining the cells' potential to engage and enwrap axons would greatly strengthen the authors' claim that this could be an effective therapeutic approach to myelin regeneration in vivo, or even a technical approach to studying myelin dynamics in vitro.

      In Figure 1B, it appears that Mbp expression in tdTomato+ cells decreases in Sox10 transduced iOLs during the observed time period. Can the authors elaborate on this result, given that MBP expression is crucial for myelination and should, if anything, increase with time?

      The authors acknowledge that there is a conversion of tdTomato- zsGreen+ cells with an astrocyte-like morphology to OLC cells expressing Mbp following Sox10 induction (Supplementary figure 5C,D). While they note the diversity of the astrocyte lineage in the discussion, further analysis should be applied to this subset of cells to confirm the subset of astrocyte or progenitor-like cell type that gives rise to their cell endpoint of interest (Sox10-driven Mbp+ iOLs).

      Finally, ectopic expression of Olig2 and Sox10 in primary astrocytes resulted in very different OLC subtypes, as evidenced by OLC marker expression seen in IHC and the subclustering of these cell types in scRNA seq. Although this diversity in OLC type and generation efficiency follows with previous reports showing that these two transcription factors vary in effect, might the authors further discuss this discrepancy given that the two transcription factors regulate one another (as mentioned in the introduction) and should theoretically give rise to more similar cells? Perhaps due to the lower specificity of Olig2 in marking a pure OLC population relative to Sox10?

      We thank the editor and reviewers for their additional comments, which have significantly improved our manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors use Aldh1l1+ astrocyte with a GFAP promoter linked to TF expression, claiming that the converting cells are cortical astrocytes. However, during early mouse development, radial glia that can give rise to multiple cell types, including oligodendrocytes, also express Aldh1l1 and low levels of GFAP. Therefore, it is not proven whether the resulting iOLs came from mature astrocyte or a radial glia population. Especially since in Figure 3B,G it is shown that a majority of the D0 population expresses high amounts of Vimentin and Nestin, both markers associated with radial glia and immature astrocytes. It would be beneficial for the authors to confirm that either there are no contaminating radial glia or that the radial glia don't express the TF, especially since the iOL population is so small.

      Thank you for this comment. We agree that previous studies have demonstrated Aldh1l1 expression in radial glia cells [1]. However, our dissection protocol to obtain the postnatal astrocytes is cortex-specific and does not take portions of the VZ/SVZ, preventing radial glia contamination.

      Nevertheless, to confirm the astrocytic identity of our Aldh1l1+ starting cells we used AUCell enrichment scoring [2]. First, Aldh1l1+ cells were subclustered from our starting culture single cell dataset (Author response image 1A). We then defined two gene signature modules: an “astrocyte” module, comprised of canonical astrocyte markers (Aqp4, Gja1, Slc1a2, Glul, Aldoc, S100b, Nfia, Thbs1, Cst3, Clu), and a “radial glia” module, containing common radial glia and progenitor markers (Pax6, Fabp7, Sox2, Hes1, Prom1, Top2a, Mki67, Ube2c, Cdk6, Mcm2). When we scored all Aldh1l1+ cells (n=869) for enrichment of each signature, no cells were classified as radial glia (Author response image 1B). Instead, the astrocyte signature predominated, with 68.3% classified as astrocytes (Author response image 1B,C). The remaining cells (36.7%), were classified as transitional, reflecting substantial expression of genes from both modules (Author response image 1B,C). Therefore, although the Aldh1l1 astrocytes do express genes common to radial glia, there are no cells that express only progenitor markers. This is consistent with literature showing that many radial glia genes are commonly found in astrocytes [3], [4], [5].

      Taken together, our stringent dissection protocol and bioinformatic profiling of our starting Aldh1l1 cells suggests that the resulting Aldh1l1+iOLs are originating from astrocytes, rather than radial glia.

      In Figure 2D, Sox10 and Nkx6.2 have an n=4 while the Cre control has an n=3. Why is this the case? Was the 4th point excluded? Similar attention should be given to other panels in Figure 2 for consistency.

      We thank the reviewer for highlighting this. No outliers or datapoints were excluded from the analysis. Rather, the fourth culture for our control treated cells was not viable for analysis.

      Author response image 1.

      Astrocyte gene expression in Aldh1l1+ cells. (A) UMAP clustering of Aldh1l1+ cells in our starting cultures (0DPT, non-transduced). (B) UMAP clustering from (A) overlayed with cell classification based on AUCell gene signature expression scoring. (C) Feature plots showing AUCell enrichment scores (darker purple indicates higher enrichment) for the astrocyte signature (left), radial glia signature (middle), and the differential enrichment score (right, Astro_AUCell - RG_AUCell) (darker purple scores indicate higher astrocyte signature and negative scores (gray) indicate higher radial glia signature).

      Representative images in Figure 2 do not convincingly support the argument by the authors. It appears that some of the cells highlighted by the arrows are just background (e.g. PDGFRa and td Tomato in Figure 2E, or zsGreen in Figure 2F). Additionally, the authors should show a different representative image depicting astrocyte morphology in Figure 2G 7DPT.

      Thank you to the reviewer for this comment. We have replaced the images in Figure 2E,F to better represent our findings (updated manuscript Figure 2E,F). We have also adjusted the representative image in Figure 2G 7DPT to better visualize the astrocyte morphology (updated manuscript Figure 2G) as well as included as supplementary additional examples of pre-conversion astrocyte morphology to supplement our morphology analysis (Author response image 2).

      Author response image 2.

      Lineage tracing confirms true conversion of astrocytes to oligodendrocyte lineage cells. Representative images of astrocyte morphology observed prior to cell conversion (arrow indicates converting cells, scale bar =50um).

      References

      (1) L. C. Foo and J. D. Dougherty, “Aldh1L1 is expressed by postnatal neural stem cells in vivo,” Glia, vol. 61, no. 9, pp. 1533–1541, Sep. 2013, doi: 10.1002/glia.22539.

      (2) S. Aibar et al., “SCENIC: Single-cell regulatory network inference and clustering,” Nat Methods, vol. 14, no. 11, pp. 1083–1086, Nov. 2017, doi: 10.1038/nmeth.4463.

      (3) M. Götz and Y.-A. Barde, “Radial Glial Cells: Defined and MajorIntermediates between EmbryonicStem Cells and CNS Neurons,” Neuron, vol. 46, no. 3, pp. 369– 372, May 2005, doi: 10.1016/j.neuron.2005.04.012.

      (4) P. Malatesta, I. Appolloni, and F. Calzolari, “Radial glia and neural stem cells,” Cell and Tissue Research, vol. 331, no. 1, pp. 165–178, 2008, doi: 10.1007/s00441-0070481-8.

      (5) S. Clavreul, L. Dumas, and K. Loulier, “Astrocyte development in the cerebral cortex: Complexity of their origin, genesis, and maturation,” Front Neurosci, vol. 16, p. 916055, Sep. 2022, doi: 10.3389/fnins.2022.916055.

    1. eLife Assessment

      This study presents a useful database resource containing protein conformations generated through molecular dynamics simulations, with extensive quality evaluation and benchmarking. While the database is well-constructed and professionally organized, the evidence supporting its claimed representation of protein conformational landscapes is incomplete, as the short simulation times and starting structure bias prevent true Boltzmann sampling of the conformational space.

    2. Reviewer #1 (Public review):

      Summary:

      The authors describe a new database that rigorously explores protein conformations.

      Strengths:

      It is extremely well done, using state-of-the-art tools by a group at the top of the field of structural modeling. The evaluation of qualities and the benchmarking of the structures are outstanding, and it is expected that the new database will have a significant impact on the field.

      Weaknesses:

      The authors are using MD simulation to generate some of the structure, and therefore should have access to standard MD energies. I am surprised that no evaluation is provided based on these energies that can be extended to free energies.

    3. Reviewer #2 (Public review):

      Summary:

      The authors developed a dataset of protein conformations by running molecular dynamics simulations starting from both native and decoy conformations for a large number of proteins. These conformations were put together as a dataset for querying and downloading, along with their energies under different force fields. The authors suggest that such conformations represent the proteins' conformational landscape, so that they will be useful for evaluating methods generating multiple conformations of proteins.

      Strengths:

      The dataset is online and working. It has good documentation for others to use.

      Weaknesses:

      The biggest weakness is that the collected conformations very likely do not represent the true conformational landscape. To represent the conformational landscape, the structures need to be sampled based on the Boltzmann distribution. However, in this study, conformations are generated by running very short (125ps to 375ps) MD simulations starting from near-native conformations and decoys. Such short simulations will produce small fluctuations around the starting conformations, so the distribution of conformations is largely dominated by the distribution of the initial conformations, which by one means are Boltzmann distributed. A conformation might be physically plausible, but it might have very small weight in the Boltzmann distribution. On the other hand, conformations with large weights might not be in the dataset.

    4. Reviewer #3 (Public review):

      Summary:

      This manuscript describes a web-based tool that allows researchers to compare large numbers of representative ("plausible") conformations of proteins. It also includes energetic analysis from multiple widely used structure-prediction methods.

      Strengths:

      This tool will likely be useful for students who want to learn more about the ensemble properties of proteins. The resource is well organized and it represents a large amount of computing resources.

      Weaknesses:

      It is not entirely clear how the database may be utilized by other groups to advance research. It could be helpful if the authors add a short section that provides example use cases that illustrate how this database can support new strategies for studying protein dynamics.

    5. Author response:

      eLife Assessment

      This study presents a useful database resource containing protein conformations generated through molecular dynamics simulations, with extensive quality evaluation and benchmarking. While the database is well-constructed and professionally organized, the evidence supporting its claimed representation of protein conformational landscapes is incomplete, as the short simulation times and starting structure bias prevent true Boltzmann sampling of the conformational space.

      We thank the editors for recognizing the usefulness of ProteinConformers and the value of its quality evaluation and benchmarking. We will revise the manuscript to clarify that ProteinConformers provides large-scale, energetically profiled descriptions of protein conformational landscapes, with broad coverage of locally stereochemically valid and energetic compatible structures from non-native to near-native regions, rather than a complete equilibrium sampling of all conformational states. These revisions will better define the scope of the resource while preserving its intended use for benchmarking and data-driven studies of protein conformational variability.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a new database that rigorously explores protein conformations.

      Strengths:

      It is extremely well done, using state-of-the-art tools by a group at the top of the field of structural modeling. The evaluation of qualities and the benchmarking of the structures are outstanding, and it is expected that the new database will have a significant impact on the field.

      We thank Reviewer #1 for the positive evaluation of our work and for recognizing the potential impact of the ProteinConformers resource.

      Weaknesses:

      The authors are using MD simulation to generate some of the structure, and therefore should have access to standard MD energies. I am surprised that no evaluation is provided based on these energies that can be extended to free energies.

      We thank the reviewer for this helpful suggestion. We reprocessed the original MD energy files and extracted three MD-derived energy terms, including total energy, system potential energy, and protein-only potential energy. These energy terms have been added to the ProteinConformers resource and the web portal. We will update the manuscript to describe these additional energetic annotations.

      Reviewer #2 (Public review):

      Summary:

      The authors developed a dataset of protein conformations by running molecular dynamics simulations starting from both native and decoy conformations for a large number of proteins. These conformations were put together as a dataset for querying and downloading, along with their energies under different force fields. The authors suggest that such conformations represent the proteins' conformational landscape, so that they will be useful for evaluating methods generating multiple conformations of proteins.

      Strengths:

      The dataset is online and working. It has good documentation for others to use.

      We appreciate Reviewer #2’s positive assessment of the online resource and documentation.

      Weaknesses:

      The biggest weakness is that the collected conformations very likely do not represent the true conformational landscape. To represent the conformational landscape, the structures need to be sampled based on the Boltzmann distribution. However, in this study, conformations are generated by running very short (125ps to 375ps) MD simulations starting from near-native conformations and decoys. Such short simulations will produce small fluctuations around the starting conformations, so the distribution of conformations is largely dominated by the distribution of the initial conformations, which by one means are Boltzmann distributed. A conformation might be physically plausible, but it might have very small weight in the Boltzmann distribution. On the other hand, conformations with large weights might not be in the dataset.

      We thank the reviewer for this important and constructive comment. We agree that the conformations in ProteinConformers should not be interpreted as an equilibrium ensemble sampled according to the Boltzmann distribution. Because the MD simulations used here are short, the resulting snapshots around each seed mainly reflect local relaxation and limited thermal fluctuation from that seed, rather than exhaustive equilibrium sampling. Therefore, the relative population of conformations in our dataset should not be interpreted as a Boltzmann weight, and some thermodynamically important states may be underrepresented or absent.

      Our goal in this work is different from conventional long-timescale MD studies that aim to estimate equilibrium populations from one or a few initial structures. ProteinConformers was designed as a large-scale, multi-seed, MD-refined conformer resource. The broad structural coverage comes primarily from initiating simulations from many diverse seed decoys for each protein, while the short all-atom MD protocol is used to relax structures under a molecular mechanics force field, remove structures that fail to converge, reduce steric clashes and unrealistic local geometries, and generate energetically annotated conformers. We will revise the manuscript to make this distinction clearer and to avoid implying that ProteinConformers provides a rigorous Boltzmann-sampled representation of the underlying thermodynamic landscape.

      We also agree that longer simulations and enhanced sampling methods, such as replica-exchange MD, metadynamics, or umbrella sampling, would be necessary to estimate equilibrium populations and improve sampling of rare but thermodynamically relevant states. We will add this point as a limitation and future direction in the revised manuscript. Thus, ProteinConformers should be viewed as a broad, energetically annotated, MDrefined conformer library for benchmarking, data-driven modeling, and descriptions of protein conformational landscapes, rather than as a complete equilibrium ensemble itself.

      Reviewer #3 (Public review):

      Summary:

      This manuscript describes a web-based tool that allows researchers to compare large numbers of representative ("plausible") conformations of proteins. It also includes energetic analysis from multiple widely used structure-prediction methods.

      Strengths:

      This tool will likely be useful for students who want to learn more about the ensemble properties of proteins. The resource is well organized and it represents a large amount of computing resources.

      We thank Reviewer #3 for the positive assessment of the ProteinConformers resource and for recognizing its potential value for community education.

      Weaknesses:

      It is not entirely clear how the database may be utilized by other groups to advance research. It could be helpful if the authors add a short section that provides example use cases that illustrate how this database can support new strategies for studying protein dynamics.

      We thank the reviewer for this constructive suggestion. We agree that the manuscript should more explicitly explain how other groups can use ProteinConformers to advance research. In the revised Discussion, we will add a short section describing concrete use cases of the database. In particular, we will emphasize that the benchmark analysis already presented in this manuscript provides a worked example of how ProteinConformers can be used by other groups. ProteinConformers-lite, together with the released evaluation metrics and codes, can serve as a standardized reference set for testing new multi-conformation or protein ensemble generation methods. Other groups can generate conformational ensembles for the same targets, compare their coverage of low-energy regions using the diversity metrics reported in Table S1, and evaluate the agreement of residue-pair geometric statistics using the plausibility metrics reported in Table S2.

      We will also describe additional use cases enabled by the full ProteinConformers resource. The dataset can be used as a training, validation, or pretraining resource for conformation generators, energy-aware ranking models, and model quality assessment methods, because each conformer is paired with structural similarity annotations and multiple energetic scores. The broad coverage from non-native to near-native conformations also enables systematic analysis of how local stereochemical validity, global structural similarity, and energetic evaluations co-vary across diverse conformational perturbations, and may provide useful structural proxies or starting points for modeling flexible or disordered-like protein states, where experimentally resolved structural data are often limited. In addition, the interactive portal allows users to filter protein-specific conformers by structural similarity, energetic annotations, and secondary-structure features, making it possible to construct customized subsets for downstream biomolecular modeling, hypothesis generation, and educational exploration.

    1. eLife Assessment

      This valuable study investigates communication from the anterior cingulate cortex to the hippocampus during learning. The optogenetic evidence supporting the functional connection including the interneurons likely to be involved, is convincing. However, the evidence linking this pathway to learning and memory and certain aspects of the statistical interpretation remain limited. Overall, the work will be of interest to neuroscientists studying cortical-hippocampal interactions and learning and memory.

    2. Reviewer #2 (Public review):

      This work is composed of two largely independent parts. The first part (Figures 1-4) attempts to study correlations between the anterior cingulate cortex (ACC) and hippocampal area CA1 in the context of learning and memory; a number of issues including missing controls make this part inconclusive and hard to interpret. The second part (Figures 5 and 6) presents evidence for a pathway in which inputs from the ACC indirectly inhibit pyramidal cells in the superficial sublayer of CA1. The optogenetic evidence demonstrating the functional connection, including the interneurons likely to be involved, is convincing, making the second part of the manuscript a valuable contribution to neuroscience. However, I do not see evidence for this connection in the correlational analyses in the first part of the study, making the involvement of this pathway in learning and memory uncertain.

      Strengths:

      The biggest strength of the work is the optogenetic manipulation experiments in the second part of the study (Figures 5 and 6), which convincingly demonstrate that stimulation of ACC pyramidal neurons activates an interneuron population with symmetric spike waveforms, and inhibits parvalbumin interneurons and pyramidal cells in CA1sup, while CA1deep cells remained largely unaffected by the stimulation.

      Weaknesses:

      The main weakness is the disconnected nature of the two parts of the study. The second part convincingly shows that ACC provides a net inhibitory drive to the hippocampus (at least to CA1sup pyramidal and PV cells, while CA1deep cells were mostly unaffected). However, the first part investigates positive cross-correlations between pre-ripple ACC activity and subsequent CA1 ripple activity. This can be observed in Figure 1-supplement 1, where CA1 cells' activity peaks around 70ms after ACC spikes. Moreover, the GLM analyses were also based on positive ACC cell-CA1 cell pair correlations as the authors reported no bias towards negative weights for the GLM analyses (see the rebuttal letter). Thus, the correlational and GLM analyses in the first part primarily characterize a positive ACC-CA1 relationship, rather than the inhibitory influence demonstrated in the second part; the two parts of the manuscript therefore investigate different phenomena (possibly confounding inputs and network effects in part 1 versus the direct ACC-CA1 connection in part 2).

      The key problem is that the main results of the two parts - namely, a dampening of the positive cross-correlations following learning in part 1 and the inhibitory ACC-CA1 connection revealed in part 2 - would be contradictory if they were interpreted as describing the same phenomenon. If the inhibitory ACC-CA1 connection was key to the downregulation of CA1 activity after learning as the authors suggest in the discussion, then we would expect ACC activity driving this change to be particularly predictive of CA1 activity in this post-learning period. Indeed, because prediction gain measures how well ACC spiking can predict subsequent CA1 spiking, any additional predictive information from the direct ACC->CA1 pathway should increase prediction gain. Instead, prediction gain decreased following learning. Thus, the positive (dampened after learning) ACC-CA1 correlations observed in the first part cannot be explained by the inhibitory ACC-CA1 pathway demonstrated in the second part. The most likely explanation is therefore that the cross-correlations studied in part 1 are dominated by other factors (such as shared inputs from other areas or coordination of cortical rhythms) and reported changes in prediction gain therefore primarily reflect changes in these factors, while the contribution of the direct ACC-CA1 pathway is drowned out and undetectable using this approach. As they stand, the two halves of the paper cannot be reconciled into the same framework.

      The second weakness is the lack of control for learning. The main result of part 1 of the study is that there is dampening of the (positive) CA1 response to ACC pre-ripple activity after learning. However, nothing indicates this is due to learning as there is no control data with no learning. Moreover, the pre- and post- task periods were not matched for duration and sleep depth, so it is entirely possible that the observed dampening could be due to reduced recruitment of some cells in ripples. An appropriate control would therefore be important for attributing this to learning

      The final weakness is statistical and goes beyond the lack of hierarchical statistics (which is also an issue with this work). The failure of a test to reach significance cannot be interpreted as evidence for the opposite. Yet the authors interpret it as such: for example, the lack of significant correlation between prediction gain values in pre- and post-task sleep in Figure 3C (p=0.14) is incorrectly interpreted as proof that ACC-CA1sup communication has reorganized as a result of learning. The claims of reorganization (mentioned multiple times in the abstract) hinge solely on this failed statistical test. Yet a failure to reach significance does not successfully demonstrate reorganization as it could result from a number of other reasons, including lack of statistical power or noisy estimates. To demonstrate reorganization, one would need to show that the observed change is greater than expected under an appropriate control (e.g. control task with no learning; or sleep data split in two halves), but this is missing from this manuscript.

      Note that in the entire manuscript, the only differences between CA1sup and CA1deep are reported as two independent tests, one of which is significant and the other does not reach statistical significance. However, this is not evidence for different effects in CA1sup and CA1deep and statements like "we uncovered a pathway-specific difference" to describe these findings are unwarranted and not supported by the data; only direct statistical comparison between the two effects could support such claims. The exception to this weakness is the optogenetic experiments in Figure 5 where CA1sup and CA1deep responses to optogenetic ACC stimulation were directly compared and found to be different.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This work by Hall et al provides a novel and important new finding about communication between the anterior cingulate cortex (ACC) and the CA1 region of the dorsal hippocampus: there is a clear ability of ACC to predict CA1 activity, and that is modulated by learning/experience. Furthermore, they have some evidence that the modulation differs by whether the CA1 neurons were in the deep versus superficial sub-layer of CA1. The evidence is suggestive of new and exciting findings, but some gaps and weaknesses remain to be addressed before I believe all of the authors' claims can be supported. The figures also need to be slightly better organized, and the discussion is missing a major dimension in my opinion. Overall, this is a strong submission, but with some gaps to fill.

      Strengths:

      (1) This is a well-written manuscript - the introduction was especially clear, well-cited, and motivating.

      (2) The sub-layer specific communication between ACC and CA1 represents the discovery of a novel and functionally impactful piece of neurobiology.

      (3) Optogenetics was an important verification of ACC-CA1 communication, as was the analysis of neurons by waveform type.

      Weaknesses:

      (1) Figure 2: Why are the data separated into two groups from the outset? If all data are combined, is there a general drop in prediction gain from pre to post?

      Thank you for bringing this to our attention. In Figure 1F, all data is combined for GLM decoding. We found a significant pre-to-post decrease in prediction gain specifically using the –200 to 0 ms window to predict CA1 spiking during ripples. Figure 3 builds upon these findings to examine how prediction gain changes relate to task engagement.

      (2) 2b and 2c are important since they are complementary means to show the same thing, and it is important that they cross-validate each other, especially since the non-significant task active neuron difference in 2b appears to be nearly as strong as the significant difference to its left. A more holistic analysis can be done to compare these dimensions.

      We appreciate this feedback. In light of this comment, as well as similar concerns raised by other reviewers regarding the binary classification of neurons as task-active or task-inactive (modulation index > 0 or < 0), we adopted a more comprehensive analytical approach. Rather than relying on an arbitrary threshold, we first examined the continuous relationship between modulation index and prediction gain change across all neurons (Figure 2C). We then divided neurons into modulation index quartiles to assess whether prediction gain changes varied across different degrees of task modulation (Figure 2D). Follow-up analyses focused on the extreme quartile comparisons that contributed to the observed effects (Figure 2E). Overall, this new approach better captured the continuous nature of task-related modulation while avoiding the inclusion of a large population of neurons with modulation indices near zero that may not meaningfully differ in task engagement. However, we were unable to replicate modulation index changes for neurons split by prediction gain score percentile and removed those figures from the manuscript. We have adjusted the text accordingly to account for this new effect.

      (3) Sup vs deep neuron definition: Did the authors have any means to validate this anatomical separation using histology or otherwise? I don't believe they described anything like that, and instead use physiology to infer anatomical location. I understand anatomy-based methods may be practically impossible with tetrodes, but this limitation should at least be mentioned, and it should be explained that without something like silicon probes or histological validation, anatomy had to be inferred from physiology.

      We think this is an important limitation to address and thank you for bringing this to our attention. Given our technical restraints, we only validated radial position through physiological properties. This remains an outstanding limitation of this study. We have since added text into the main discussion bringing attention to this limitation. See below:

      “Lastly, our CA1 sublayer classification does come with its own caveats. Tetrode identification of CA1 sublayers is not a trivial matter. We implemented guidelines informed by past research (see methods) to help ensure we isolated CA1sup and CA1deep groups, only including neurons where we had the greatest confidence (Berndt et al., 2023; Mizuseki et al., 2011). In doing so, we excluded neurons which classification was uncertain. It is possible this excludes some meaningful populations of CA1 sublayers. Additionally, despite our approach, some cross-inclusion of sublayers may remain. Therefore, interpretations of CA1 sublayers difference should be considered with these limitations in mind. That said, the number of neurons and animals tested does provide overall confidence regarding our results. Future studies investigating this ACC to CA1 sublayer specific line of communication would benefit from the use of silicon probes or neuropixels that enable precise radial localization.”

      (4) Superficial vs deep differences in firing rate ratio based on PG: there are many fewer CAdeep neurons, but in 4c, the trends appear to be the same pre-training, top PG lower than others. It seems the lack of difference in CA1deep in 4c may be due to the much lower power/n. This should be discussed or addressed.

      We appreciate this feedback and since have added more recordings to address these lower Ns (CA1sup: Previous N = 71, Revisions N = 89; CA1deep: Previous N = 21, Revisions N = 61). Notably, we find that previous firing rate ratio (now called modulation index) is no longer significant with the inclusion of more neurons and is reported accordingly (Figure 2—figure supplement 1).

      (5) In Figure 5, the term "firing rate ratio" is used, and it sounds the same as in previous figures, but this is a different ratio (based on modulation by opto stim, not task).

      To improve clarity and avoid confusion with task-related modulation metric used across the paper, we renamed “Firing Rate Ratio" throughout the manuscript to "Modulation Index". We also relabeled the Figure 5D Y-axis as "Z-scored Firing Response" to more accurately reflect the plotted data and avoid confusion.

      (6) I would like to learn more about these v-type neurons. I understand we do not yet know about their molecular or morphologic correlate, but more analysis can be done with the current data.

      We thank you for this feedback. We have included further analysis into V-type properties. Namely, we performed autocorrelegrams, theta phase modulation, and burst index analyses. See Figure 6 and Figure 6—figure supplement 1.

      Additionally, we performed cross-correlogram analyses to examine whether V-type interneurons exhibited consistent temporal relationships with PV interneurons or other CA1 neurons. However, V-type and PV interneurons were sparse throughout our recordings, with most sessions containing two or fewer identified interneurons, which limited our ability to perform meaningful cross-correlogram analyses. Nevertheless, we examined the available recordings but found no consistent evidence of correlated firing between interneuron classes or between V-type interneurons and CA1 pyramidal neurons.

      (7) I would like more discussion of ACC-CA1 connectivity.

      We have since added greater discussion of ACC-CA1 connectivity into the discussion section. See below:

      “Finally, an important caveat to mention is that it remains an ongoing debate whether ACC directly projects to CA1 (Andrianova et al., 2023; Rajasethupathy et al., 2015; Shi et al., 2022). One lab reported clear monosynaptic ACC-to-CA1 connection (Rajasethupathy et al., 2015), while another lab replicated those same experiments and were unable to come to the same conclusions (Andrianova et al., 2023). Further studies report no direct connection (Shi et al., 2022). The contention in connectivity may arise from differences in targeting strategies, injection coordinates, or viruses used. Our findings reported an excitatory response in CA1 V-type interneurons in response to ACC stimulations, proposing another possibility for ACCàCA1 connectivity. Interestingly V-type interneurons responded with extremely low-latency as fast as 4.2 ms after stimulations, compatible with monosynaptic timing (Cho et al., 2013; Petreanu et al., 2007; Wang et al., 2009). If the ACC→V-Type connection was monosynaptic pathway, it could help explain discrepancies in the field, as the relative sparsity of V-type interneurons may reduce the likelihood of detecting ACC→CA1 connectivity. However, future anatomical studies are needed to conclusively determine connectivity.

      Alternatively, ACC→CA1 communication may be mediated by multiple intermediate structures (Behzadi et al., 1990; Oh et al., 2014; Shi et al., 2022; Souza et al., 2022). The ACC sends monosynaptic projections to the nucleus reuniens (RE) and median raphe (MnR), both of which project directly to CA1 (Oh et al., 2014; Shi et al., 2022). The RE has a known role in contextual discrimination learning and memory specificity (Ramanathan & Maren, 2019; Ramanathan et al., 2018; Ratigan et al., 2023; Silva et al., 2021; Xu & Südhof, 2013). Interestingly, RE→CA1 activity tuned to immobility (freezing) emerges only after shocks are presented, suggesting a learning-induced modification between regions, similar to that seen in our ACC–CA1 data (Ratigan et al., 2023). As for MnR, it receives dense inputs from the ACC (Behzadi et al., 1990; Souza et al., 2022), and its projections to the CA1 are predominantly glutamatergic (Jackson et al., 2009; Senft et al., 2021; Szonyi et al., 2016). Notably, these glutamatergic MnR inputs directly target CA1 interneurons, including CCK basket cells (Miettinen & Freund, 1992; Morales & Bloom, 1997; Senft et al., 2021), while avoiding PV interneurons and pyramidal neurons (Acsady et al., 1993; Freund et al., 1990; Halasy et al., 1992; Miettinen & Freund, 1992; Papp et al., 1999; Turi et al., 2019). This connectivity suggests that the ACC may indirectly modulate CA1 activity through the MnR, potentially suppressing PV interneuron and pyramidal neuron activity via local inhibitory circuits, thereby contributing to the regulation of hippocampal oscillations and memory consolidation (Huang et al., 2022; Wang et al., 2015). Ultimately, future experiments combining pathway-specific manipulations with simultaneous recordings will be necessary to distinguish direct from polysynaptic mechanisms.”

      (8) Some elements may be missing from the discussion, relating baseline functioning versus post-learning function.

      We thank the reviewer for this feedback and their recommendation for possible alternate explanations. We have added these discussions into the main text. See below:

      “Alternatively, ACC→CA1 communication may contribute to the homeostatic downscaling of memory-unrelated synapses during sleep. Evidence finds that slow-wave sleep is strongly linked to downscaling of non-learning related neuron activity (Gulati et al., 2017; Liu et al., 2010; Tononi & Cirelli, 2003; Tononi & Cirelli, 2006; Watson et al., 2016). Slow-wave sleep ripples in particular depotentiate memory-unrelated synapses (Gulati et al., 2017; Norimoto et al., 2018). In our study, we find that learning-related reduction in communication between ACC and CA1 were selective for task-inactive neurons. Therefore, ACC→CA1sup communication may not simply weaken following learning but rather becomes selectively disengaged from task-inactive neurons, enabling homeostatic downscaling while preserving behaviorally relevant synapses (Liu et al., 2010; Norimoto et al., 2018; Tononi & Cirelli, 2003; Tononi & Cirelli, 2006; Watson et al., 2016). Still, behavioral recruitment alone cannot account for the observed remodeling of ACC→CA1 communication, as CA1sup and CA1deep neurons did not display significant differences in task-related activity (Figure 2—figure supplemental 2). Instead, these learning-related changes of task-inactive neurons appear sublayer-specific.”

      Reviewer #2 (Public review):

      Summary:

      This study uncovers an inhibitory pathway from the anterior cingulate cortex (ACC) to pyramidal cells in the superficial sublayer of hippocampal area CA1 (CA1sup). As ACC neuron spiking tends to precede hippocampal ripples, this presents the intriguing possibility that ACC inputs are selectively inhibiting particular CA1sup neurons, which could play a role in the reactivation of task-related ensembles known to take place during hippocampal ripples. Indeed, through a generalized linear model (GLM) analysis, the authors demonstrate that the ACC activity within the 200ms immediately preceding the ripple is predictive of the ripple content.

      Strengths:

      The biggest strength of the work is the optogenetic manipulation experiments, which convincingly demonstrate that stimulation of ACC pyramidal neurons activates an interneuron population with symmetric spike waveforms, and inhibits parvalbumin interneurons and pyramidal cells in CA1sup but not CA1deep sublayer.

      An additional strength in the GLM analysis which consistently shows that ACC activity preceding the ripple is predictive of hippocampal activity during the ripple considerably more than in shuffled data for all cells and periods tested.

      Weaknesses:

      The major weakness of this work is that the link with learning and memory is not very well supported.

      The only evidence of rebalancing and reorganization appears to be a single statistical test (the test in Figure 1f, p=0.013) demonstrating a decrease of the GLM prediction gain from pre-task sleep to post-task sleep; the same test is repeated for subsets of the data in the rest of the figures. As the idea of rebalancing and reorganization is central to the paper as currently written, exploring it through another measure, independent of the GLM prediction gain, should be expected. The notion that this pathway is suppressed in sleep following learning can be supported by demonstrating a decrease in any of the following measures: ACC spike-triggered average CA1sup responses, cross-covariances (Wierzynski et al 2009) between ACC and CA1sup cells in post-task sleep, or ripple-triggered cross-correlations (Sirota et al. 2009).

      We thank the reviewer for this helpful feedback. We have added an additional analysis the reviewer pointed out to address this concern. Specifically, we performed an ACC spike‑triggered analysis. The ACC spike‑triggered average further supported the learning‑related decrease in ACC-to-CA1 activity (see Figure 1—figure supplement 1). We did not include a separate cross-covariance analysis because the ACC spike-triggered average captures essentially the same temporal relationship between ACC and CA1 activity.

      The differences between task-active and task-inactive neurons are not convincing. The separation between task-active and task-inactive neurons is to divide a distribution that is far from bimodal into what appears to be two arbitrary groups. Similarly, the authors divide cells relative to their prediction gain ("Top PG" and "Bottom PG" in Figure 2c), which fails to select for the population of significantly predicted cells (relative to the shuffle). Within CA1sup cells, after learning, there is a significant decrease in the prediction gain for "task-inactive" cells but not "task-active" cells, but it is important to keep in mind that the "task-active" group contains only 24 neurons, and there was no difference between the two groups of cells ("task-active" vs "task-inactive") when directly compared.

      We agree with this concern. To address this, we removed conclusion based on those arbitrary criteria instead opting for a more continuous approach. Specifically, we adopted a more comprehensive analytical approach. Rather than relying on an arbitrary threshold, we first examined the continuous relationship between modulation index and prediction gain change across all neurons (Figure 2C). We then divided neurons into modulation index quartiles to assess whether prediction gain changes varied across different degrees of task modulation (Figure 2D). Follow-up analyses focused on the extreme quartile comparisons that contributed to the observed effects (Figure 2E). Overall, this new approach better captured the continuous nature of task-related modulation while avoiding the inclusion of a large population of neurons with modulation indices near zero that may not meaningfully differ in task engagement.

      Finally, it is not clear whether the identity of the pathway-responsive CA1sup neurons is fixed or whether it may change with learning. A deeper analysis into the cell pair cross-correlations or the weights of the GLM analysis may reveal whether there is a reorganization of CA1sup responses (some cells that were inhibited are no longer inhibited, and vice versa) or a dampening (the same CA1sup cells are inhibited in both cases, but the inhibition is less-pronounced in post-task sleep). The possibility of a rigid circuit dampened immediately following fear conditioning, is not discussed by the authors.

      We appreciate this feedback. To address this concern without weight analysis, we examined the stability of prediction gain scores between pre- and post-training sleep. Preservation of neuronal prediction-gain rankings would suggest that learning weakens existing predictive communication while maintaining the relative contribution of individual neurons, consistent with a dampening response. In contrast, poor preservation of prediction-gain rankings would be more indicative of a reorganization of predictive relationships across the population. This led to interesting findings regarding sublayer differences: ACC→CA1sup communication is more dynamic and evolving following learning, whereas ACC→CA1deep communication remains comparatively stable (See Figure 2A&B; Figure 3 C–F).

      Reviewer #3 (Public review):

      Summary:

      In this study, Hall and colleagues investigate how the coupling of activity from ACC to CA1is altered by fear learning, showing that during sleep immediately before learning, there is evidence for increased coupling of ACC activity with neurons that will subsequently be inhibited during the learning process. They go on to show that this effect seems to be mediated most by a subpopulation of neurons in the superficial layer of CA1. This fits with previous reports suggesting that these superficial neurons are key for the flexible updating of memory. The authors then go on to show that artificial activation of ACC using optogenetics results in varied effects in CA1, including a subtle decrease in activity of superficial neurons that lasts longer than the stimulus itself. Finally, the authors present some preliminary data suggesting that different interneurons may be recruited by this optogenetic stimulation in different ways and at different times.

      Overall, this is an interesting paper, but much of the analysis is very preliminary, and much of the crucial data about the learning effects and alterations to cell firing are not presented clearly and fully. This is further confounded by a rather opaque description of the results and analysis in the text. Overall, there is something very interesting here, but there needs to be a substantial series of extra analyses to clearly say what this is. In many cases, more robust analysis may render the results underpowered, which could dramatically change the conclusions of the paper.

      Strengths:

      The authors performed difficult, dual-location recordings across a multi-day learning paradigm, which seems like it could be a really nice dataset. They delve into the circuit basis of an interesting finding regarding ACC to CA1 connectivity and how this changes before and after fear conditioning. They provide data to suggest this connectivity may be through specific and distinct subcircuits in CA1.

      Weaknesses:

      (1) There is essentially no information in the text or figures about what the actual learning was, how it was done, how individual animals performed, and how any of these metrics related to learning. Looking at the methods, the authors did a number of things never mentioned anywhere in the text or figures, including novel arena exposure, contextual reexposure in extinction after learning, etc. It seems that this is a very rich dataset that has not been presented at all. I would recommend at the very least:

      We appreciate the reviewers’ feedback and have worked to address these concerns. See below our response.

      (a) Plot all of the behavioural training data, and how each mouse relates to one another - did the mice learn? At this stage, we don't know!

      We have now plotted all contextual fear conditioning behavioral data for each mouse (Figure 2—figure supplement 1). All mice exhibited high level of freezing during the contextual fear test, suggesting successful learning of the context–shock association.

      (b) Explain in the text in detail exactly what was done and why, and what this tells us about the neuronal activity.

      We have now added text to describe in detail the behavioral results and how that may relate to our GLM analyses. We have also more clearly detailed our experimental objectives (what was done and why) utilizing contextual fear conditioning,

      “In this approach, we were able to examine ACC–CA1 communication prior, during, and after learning, enabling us to examine how this communication evolves across fear learning. Specifically, we emphasized investigation into communication changes between pre- and post-training sleep to understand whether functional connectivity undergoes learning-related reorganization.”

      “Lastly, we examined whether PG scores correlated with the freezing response in mice during recall. Across all mice, freezing was significantly higher during recall than pre-shock baseline during training (Figure 2—figure supplement 2A). Overall, we found no correlation between PG and freezing (Figure 2—figure supplement 2B–D). However, there was a trend for a positive correlation (p = .07) between ΔPG and freezing percentage. An important consideration is that the uniformly high levels of freezing in mice limited behavioral variability, potentially reducing our ability to detect relationships between ACC–CA1 communication decoding and behavioral differences.”

      (c) If there is variance in learning and or conditioning, does this relate to features in the analysis, such as the GLM result.

      We examined this question by first investigating whether prediction gain scores in pre-training, post-training, or overall ΔPG correlated with freezing percentage. We found no significant correlation between any of the variables (Figure 2—figure supplement 1). We speculate this may be a result of a relatively robust freezing response limiting the ability for our fine-grained GLM decoding analyses to detect those differences. We have added this consideration to the main text.

      (2) Along similar lines, a key metric for most of the paper is that neurons most coupled with ACC are more likely to be inhibited during training. However, there is nothing anywhere in the paper showing these data. How do neurons in general respond to contextual shocks? The methods describe this as the average firing rate during training, normalised to pre-sleep activity. This metric seems a bit coarse and may obscure really important task-relevant dynamics. Are the neurons active at specific times, are they tuned to relevant parts of the task, and do any of these features of the cell activity also relate to the coupling with ACC? Similarly, how did the authors mitigate the influence of electrical artefacts caused by the foot shock in their recordings? Again, there is a huge amount of data here that is not being described, and likely holds very valuable information about what is actually happening. The paper would really benefit from the inclusion of these data in an accessible form, such as heatmaps of spiking, how these patterns change over time, and around e.g., foot shock, etc. Also key is how these features are altered by the variability of learning across subjects.

      We thank reviewer for this feedback. As pointed out, electrical artifacts caused by the footshocks prevents our ability to examine, with temporal sensitivity, neurons’ responses to footshocks. Therefore, we are left to examine activity changes across longer timescales. We acknowledge that our current modulation index analysis is a bit coarse. One reason is that our preliminary analyses using more temporally sensitive approaches did not reveal robust CA1 activity associated with specific behaviors, such as freezing or transitions between mobility and immobility. Thus, we chose a more holistic approach looking at the full CFC session to include all components that CA1 may be encoding during the training session. For example, although the pre-shock baseline period does not contain any footshock stimuli, it serves a key part in the process as mice begin to encode their environment around them. Nevertheless, we have added an additional analysis to examine how modulation changes across the pre-shock versus post-shock window (Figure 5—figure supplement 1). Although this provides greater insight into how ACC and CA1 activity changes across different dimensions in the task, further investigation utilizing casual manipulations is necessary to elucidate which phase ACC→CA1 activity is most involved.

      (3) A number of the effects are presented by comparing a statistically significant effect to a non-statistically significant effect (e.g. in Figure 2b, Figure 2d, Figure 4 b,c, and others). This isn't really valid - the key test that the two groups are different is either with a direct test of the difference or an interaction term in an e.g., ANOVA test. In some places, I am not sure the same conclusions will be drawn from the data with these tests.

      We want to thank the reviewer for this critical feedback. We have since added the appropriate statistical measure including linear mixed-effects models and ANOVA tests and for our analysis to avoid our previous statistical errors.

      (4) To what extent is defining superficial and deep CA1 neurons solely by ripple waveform an accepted method? Of the two papers referenced for this approach, one is a 2-photon calcium imaging paper that does not do electrical recordings (as far as I am aware), and the second uses this as a descriptor after defining the positions of units on an array. It would be good to clarify how accepted this is, and also how robust this is. At the very least, some kind of metric or walkthrough in the supplement as to how this was done, and how well each cell was classified and with what confidence, or some metric of how distinct and separate the two populations were (or was it just a smudge).

      We appreciate this feedback. While the Berndt 2023 paper implemented 2-photon calcium imaging, they also used tetrode classifications for radial axes in that paper which help informed our approach. Ultimately, our tetrode classification remains an outstanding limitation which we have since added to the main text (See below).

      “Lastly, our CA1 sublayer classification does come with its own caveats. Tetrode identification of CA1 sublayers is not a trivial matter. We implemented guidelines informed by past research (see methods) to help ensure we isolated CA1sup and CA1deep groups, only including neurons where we had the greatest confidence (Berndt et al., 2023; Mizuseki et al., 2011). In doing so, we excluded neurons which classification was uncertain. It is possible this excludes some meaningful populations of CA1 sublayers. Additionally, despite our approach, some cross-inclusion of sublayers may remain. Therefore, interpretations of CA1 sublayers difference should be considered with these limitations in mind. That said, the number of neurons and animals tested does provide overall confidence regarding our results. Future studies investigating this ACC to CA1 sublayer specific line of communication would benefit from the use of silicon probes or neuropixels that enable precise radial localization.”

      (5) In the optogenetic experiment in Figure 5, the effect on the CA1 sup neurons seems to be driven by changes in a small subpopulation of this group, with no change in the others. Related to point 2, is there anything else in the data that can pull out what these cells are? More detailed analysis of the firing of these neurons might pull out something really interesting.

      We thank the reviewer for this feedback. Firstly, we want to clarify that optogenetic experiments were performed in a separate cohort of mice that did not undergo contextual fear conditioning. We have adjusted the text accordingly to make this distinction clearer. We have also added a per-animal separation of CA1 heatmap responses to ACC stimulations to demonstrate suppression is preserved across animals (Figure 5—figure supplement 3; Figure 6—figure supplement 2). Finally, our optogenetic experiments were primarily focused on understanding anatomical connectivity. Consequently, common waking behaviors (e.g., exploration and feeding) were not standardized across animals, and our analyses were therefore restricted to comparisons between slow-wave sleep and wakefulness more broadly.

      (6) Related to this - a number of comparisons simply pool neurons across mice and analyse them as if independent. This is done a lot in the past, but it would be better if an approach that included the interdependence of neurons recorded from the same mouse at the same time were used (such as a hierarchical model). While this is complex, a simpler approach would just be to plot the summary data also per mouse. For example, in Figure 5, how do the neurons inhibited by ACC activation spread across the different mice? Is the level of inhibition related to how well the mice learned the CS-US association?

      For dual-site analysis we have now added animal-level comparison for some key analyses (See Figure—figure supplement 1&2). As for optogenetic experiments, we have added per-animal heatmaps for ACC stimulation response (Figure 5—figure supplement 3; Figure 6—figure supplement 2).

      (7) Figure 6 is interesting, but very preliminary. None of the effects are quantified, and one of the cell types is not identified. I think some proper analysis needs to be done, again across mice, to be able to draw conclusions from these data.

      We thank the reviewer for this feedback. Reviewer 1 had a similar concern, and we have since added additional analyses to the revisions. Specifically, we performed autocorrelegrams, theta phase modulation analyses and a burst index analysis for V-Type interneurons. Importantly, however, these approaches still collapse neurons across mice. Given the sparse nature of V-type and PV interneurons, files often contain 2 or fewer interneurons making within-animal comparison difficult. That said, we have added supplemental figures displaying per-animal changes in response to ACC stimulations (Figure 6—figure supplement 1).

      (8) Finally, in general, I felt that the way the paper was written was very hard to follow, often relying on very processed levels of analysis that were hard to relate back to the raw traces and their biological meaning. In general taking more words to really simply and fully explain each analysis, and taking the words and figures to walk through how each analysis was done and what it tells us about the neuronal data/biology would be really beneficial, especially to someone who is not an extracellular electrophysiologist or immersed in the immediate field.

      We thank the reviewer for this feedback. Throughout the manuscript, we have revised the text to improve clarity in explaining our approaches and their results.

      In summary, while this manuscript explores an intriguing hypothesis about pre-learning circuit dynamics, it is currently held back by insufficient clarity in behavioural analysis, data presentation, and statistical quantification. Addressing these core issues would greatly improve interpretability and confidence in the findings.

      Additional comment:

      For the optogenetic experiments, we reprocessed and resorted the neuronal dataset to ensure accurate cell classification. Following this re-analysis, the principal findings remained unchanged. However, we found that sublayer-specific differences in response to ACC stimulation were restricted to the first second following ACC stimulation. Consequently, we removed the previous Figure 5E, which examined firing rate changes across successive 1-sec time bins, as the additional time windows did not provide further evidence of sublayer-specific effects.

      Acsady, L., Halasy, K., & Freund, T. F. (1993). Calretinin is present in non-pyramidal cells of the rat hippocampus--III. Their inputs from the median raphe and medial septal nuclei. Neuroscience, 52(4), 829-841. https://doi.org/10.1016/0306-4522(93)90532-k

      Andrianova, L., Yanakieva, S., Margetts-Smith, G., Kohli, S., Brady, E. S., Aggleton, J. P., & Craig, M. T. (2023). No evidence from complementary data sources of a direct glutamatergic projection from the mouse anterior cingulate area to the hippocampal formation. eLife, 12, e77364. https://doi.org/10.7554/eLife.77364

      Behzadi, G., Kalén, P., Parvopassu, F., & Wiklund, L. (1990). Afferents to the median raphe nucleus of the rat: Retrograde cholera toxin and wheat germ conjugated horseradish peroxidase tracing, and selective<span class="small">d</span>-[<sup>3</sup>H]aspartate labelling of possible excitatory amino acid inputs. Neuroscience, 37(1), 77-100. https://doi.org/10.1016/0306-4522(90)90194-9

      Berndt, M., Trusel, M., Roberts, T. F., Pfeiffer, B. E., & Volk, L. J. (2023). Bidirectional synaptic changes in deep and superficial hippocampal neurons following in vivo activity. Neuron, 111(19), 2984-2994.e2984. https://doi.org/10.1016/j.neuron.2023.08.014

      Cho, J. H., Deisseroth, K., & Bolshakov, V. Y. (2013). Synaptic encoding of fear extinction in mPFC-amygdala circuits. Neuron, 80(6), 1491-1507. https://doi.org/10.1016/j.neuron.2013.09.025

      Freund, T. F., Gulyas, A. I., Acsady, L., Gorcs, T., & Toth, K. (1990). Serotonergic control of the hippocampus via local inhibitory interneurons. Proc Natl Acad Sci U S A, 87(21), 8501-8505. https://doi.org/10.1073/pnas.87.21.8501

      Gulati, T., Guo, L., Ramanathan, D. S., Bodepudi, A., & Ganguly, K. (2017). Neural reactivations during sleep determine network credit assignment. Nat Neurosci, 20(9), 1277-1284. https://doi.org/10.1038/nn.4601

      Halasy, K., Miettinen, R., Szabat, E., & Freund, T. F. (1992). GABAergic Interneurons are the Major Postsynaptic Targets of Median Raphe Afferents in the Rat Dentate Gyrus. Eur J Neurosci, 4(2), 144-153. https://doi.org/10.1111/j.1460-9568.1992.tb00861.x

      Huang, W., Ikemoto, S., & Wang, D. V. (2022). Median Raphe Nonserotonergic Neurons Modulate Hippocampal Theta Oscillations. J Neurosci, 42(10), 1987-1998. https://doi.org/10.1523/JNEUROSCI.1536-21.2022

      Jackson, J., Bland, B. H., & Antle, M. C. (2009). Nonserotonergic projection neurons in the midbrain raphe nuclei contain the vesicular glutamate transporter VGLUT3. Synapse, 63(1), 31-41. https://doi.org/10.1002/syn.20581

      Liu, Z.-W., Faraguna, U., Cirelli, C., Tononi, G., & Gao, X.-B. (2010). Direct Evidence for Wake-Related Increases and Sleep-Related Decreases in Synaptic Strength in Rodent Cortex. The Journal of Neuroscience, 30(25), 8671. https://doi.org/10.1523/JNEUROSCI.1409-10.2010

      Miettinen, R., & Freund, T. F. (1992). Convergence and segregation of septal and median raphe inputs onto different subsets of hippocampal inhibitory interneurons. Brain Res, 594(2), 263-272. https://doi.org/10.1016/0006-8993(92)91133-y

      Mizuseki, K., Diba, K., Pastalkova, E., & Buzsáki, G. (2011). Hippocampal CA1 pyramidal cells form functionally distinct sublayers. Nature Neuroscience, 14(9), 1174-1181. https://doi.org/10.1038/nn.2894

      Morales, M., & Bloom, F. E. (1997). The 5-HT3 receptor is present in different subpopulations of GABAergic neurons in the rat telencephalon. J Neurosci, 17(9), 3157-3167. https://doi.org/10.1523/JNEUROSCI.17-09-03157.1997

      Norimoto, H., Makino, K., Gao, M., Shikano, Y., Okamoto, K., Ishikawa, T., Sasaki, T., Hioki, H., Fujisawa, S., & Ikegaya, Y. (2018). Hippocampal ripples down-regulate synapses. Science, 359(6383), 1524-1527. https://doi.org/10.1126/science.aao0702

      Oh, S. W., Harris, J. A., Ng, L., Winslow, B., Cain, N., Mihalas, S., Wang, Q., Lau, C., Kuan, L., Henry, A. M., Mortrud, M. T., Ouellette, B., Nguyen, T. N., Sorensen, S. A., Slaughterbeck, C. R., Wakeman, W., Li, Y., Feng, D., Ho, A., . . . Zeng, H. (2014). A mesoscale connectome of the mouse brain. Nature, 508(7495), 207-214. https://doi.org/10.1038/nature13186

      Papp, E. C., Hajos, N., Acsady, L., & Freund, T. F. (1999). Medial septal and median raphe innervation of vasoactive intestinal polypeptide-containing interneurons in the hippocampus. Neuroscience, 90(2), 369-382. https://doi.org/10.1016/s0306-4522(98)00455-2

      Petreanu, L., Huber, D., Sobczyk, A., & Svoboda, K. (2007). Channelrhodopsin-2–assisted circuit mapping of long-range callosal projections. Nature Neuroscience, 10(5), 663-668. https://doi.org/10.1038/nn1891

      Rajasethupathy, P., Sankaran, S., Marshel, J. H., Kim, C. K., Ferenczi, E., Lee, S. Y., Berndt, A., Ramakrishnan, C., Jaffe, A., Lo, M., Liston, C., & Deisseroth, K. (2015). Projections from neocortex mediate top-down control of memory retrieval. Nature, 526(7575), 653-659. https://doi.org/10.1038/nature15389

      Ramanathan, K. R., & Maren, S. (2019). Nucleus reuniens mediates the extinction of contextual fear conditioning. Behavioural brain research, 374, 112114. https://doi.org/https://doi.org/10.1016/j.bbr.2019.112114

      Ramanathan, K. R., Ressler, R. L., Jin, J., & Maren, S. (2018). Nucleus Reuniens Is Required for Encoding and Retrieving Precise, Hippocampal-Dependent Contextual Fear Memories in Rats. The Journal of Neuroscience, 38(46), 9925. https://doi.org/10.1523/JNEUROSCI.1429-18.2018

      Ratigan, H. C., Krishnan, S., Smith, S., & Sheffield, M. E. J. (2023). A thalamic-hippocampal CA1 signal for contextual fear memory suppression, extinction, and discrimination. Nat Commun, 14(1), 6758. https://doi.org/10.1038/s41467-023-42429-6

      Senft, R. A., Freret, M. E., Sturrock, N., & Dymecki, S. M. (2021). Neurochemically and Hodologically Distinct Ascending VGLUT3 versus Serotonin Subsystems Comprise the r2-Pet1 Median Raphe. J Neurosci, 41(12), 2581-2600. https://doi.org/10.1523/JNEUROSCI.1667-20.2021

      Shi, W., Xue, M., Wu, F., Fan, K., Chen, Q. Y., Xu, F., Li, X. H., Bi, G. Q., Lu, J. S., & Zhuo, M. (2022). Whole-brain mapping of efferent projections of the anterior cingulate cortex in adult male mice. Mol Pain, 18, 17448069221094529. https://doi.org/10.1177/17448069221094529

      Silva, B. A., Astori, S., Burns, A. M., Heiser, H., van den Heuvel, L., Santoni, G., Martinez-Reza, M. F., Sandi, C., & Gräff, J. (2021). A thalamo-amygdalar circuit underlying the extinction of remote fear memories. Nature Neuroscience, 24(7), 964-974. https://doi.org/10.1038/s41593-021-00856-y

      Souza, R., Bueno, D., Lima, L. B., Muchon, M. J., Gonçalves, L., Donato, J., Jr., Shammah-Lagnado, S. J., & Metzger, M. (2022). Top-down projections of the prefrontal cortex to the ventral tegmental area, laterodorsal tegmental nucleus, and median raphe nucleus. Brain Struct Funct, 227(7), 2465-2487. https://doi.org/10.1007/s00429-022-02538-2

      Szonyi, A., Mayer, M. I., Cserep, C., Takacs, V. T., Watanabe, M., Freund, T. F., & Nyiri, G. (2016). The ascending median raphe projections are mainly glutamatergic in the mouse forebrain. Brain Struct Funct, 221(2), 735-751. https://doi.org/10.1007/s00429-014-0935-1

      Tononi, G., & Cirelli, C. (2003). Sleep and synaptic homeostasis: a hypothesis. Brain Research Bulletin, 62(2), 143-150. https://doi.org/https://doi.org/10.1016/j.brainresbull.2003.09.004

      Tononi, G., & Cirelli, C. (2006). Sleep function and synaptic homeostasis. Sleep Med Rev, 10(1), 49-62. https://doi.org/10.1016/j.smrv.2005.05.002

      Turi, G. F., Li, W. K., Chavlis, S., Pandi, I., O'Hare, J., Priestley, J. B., Grosmark, A. D., Liao, Z., Ladow, M., Zhang, J. F., Zemelman, B. V., Poirazi, P., & Losonczy, A. (2019). Vasoactive Intestinal Polypeptide-Expressing Interneurons in the Hippocampus Support Goal-Oriented Spatial Learning. Neuron, 101(6), 1150-1165 e1158. https://doi.org/10.1016/j.neuron.2019.01.009

      Wang, D. V., Yau, H.-J., Broker, C. J., Tsou, J.-H., Bonci, A., & Ikemoto, S. (2015). Mesopontine median raphe regulates hippocampal ripple oscillation and memory consolidation. Nature Neuroscience, 18(5), 728-735. https://doi.org/10.1038/nn.3998

      Wang, J., Hasan, M. T., & Seung, H. S. (2009). Laser-evoked synaptic transmission in cultured hippocampal neurons expressing channelrhodopsin-2 delivered by adeno-associated virus. J Neurosci Methods, 183(2), 165-175. https://doi.org/10.1016/j.jneumeth.2009.06.024

      Watson, Brendon O., Levenstein, D., Greene, J. P., Gelinas, Jennifer N., & Buzsáki, G. (2016). Network Homeostasis and State Dynamics of Neocortical Sleep. Neuron, 90(4), 839-852. https://doi.org/https://doi.org/10.1016/j.neuron.2016.03.036

      Xu, W., & Südhof, T. C. (2013). A Neural Circuit for Memory Specificity and Generalization. Science, 339(6125), 1290-1295. https://doi.org/doi:10.1126/science.1229534

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggested fixes (these correlate with the numbered points in the Weaknesses section of the Public Review):

      (1) That seems like the larger finding to start with, before dissecting by Firing Activity Index. I'd first show the overall finding, then dissect it.

      We thank the reviewer for this recommendation. In our manuscript, figure 1F examines the overall pre-to-post prediction gain changes prior to any neuron separation. Further separation of neurons’ characteristics and classification occurs in subsequent figures.

      (2) The overall picture painted by Figure 2 suggests a correlation analysis should be carried out to search for a general property: prediction score gain vs firing activity index. Are those two variables considered significant by Pearson correlation? Did the authors try that and it didn't work, so they did these analyses? If there is no significant correlation, what does a more detailed look at those two variables on an x-y plot teach us? I suggest considering showing such a plot to readers, at least in a Supplement.

      We thank the reviewer for this feedback and have incorporated this approach in our revised manuscript. Specifically, we performed a correlation analysis between prediction gain change and modulation index (formerly called firing activity index). We uncovered a significant positive correlation between variables, suggesting that task engagement modifies ACC→CA1 communication (Figure 2C).

      (3) (No additional comments).

      (4) This weakness may be able to be addressed by doing a correlation of depth (LFP amplitude of sharp wave) versus the firing rate ratio. For example, the threshold used for deep may have been such that it reduced the number of detected deep neurons, but if a general relationship between depth and degree of FRR is found, it can remove issues from this difference in statistical power.

      We thank the reviewer for this comment. We were unable to perform a reliable link between correlation depth and modulation index. LFP amplitude can vary substantially between tetrodes due to differences in electrode impedance, placement, and recording conditions, making direct comparisons across animals difficult. While within-animal analyses could largely circumvent these issues, many recording sessions did not include tetrodes spanning the full superficial-to-deep CA1 axis, preventing a reliable assessment of this relationship.

      (5) I would give it a different name - "opto-modulation index" or "opto firing rate ratio" perhaps. This would make it clear that you are not measuring task-based modulation of firing.

      We have since modified our wording to improve clarity. Specifically, we renamed “Firing Rate Ratio" throughout the manuscript to "Modulation Index". We also relabeled the Figure 5D Y-axis as "Z-scored Firing Response" to more accurately reflect the plotted data and avoid confusion

      (6) Specifically: can the post-opto lag of v-type versus wide-waveform and PV-type neurons be analyzed? Are the V-type neurons increasing firing before the others decrease? What about cross correlograms between v-type and pyramidal neurons, either at baseline or post-stim?

      We appreciate this feedback. We have added a figure showing differences in response lags to the optostimulation. We demonstrate that V-Type interneurons clearly fire prior to PV and pyramidal cells (Figure 6—figure supplement 1E&F). Additionally, we performed cross-correlogram analyses to examine whether V-type interneurons exhibited consistent temporal relationships with PV interneurons or other CA1 neurons. However, V-type and PV interneurons were sparse throughout our recordings, with most sessions containing two or fewer identified interneurons, which limited our ability to perform meaningful cross-correlation analyses. Nevertheless, we examined the available recordings but found no consistent evidence of correlated firing between interneuron classes or between V-type and pyramidal neurons.

      (7) Can the authors discuss the candidate pathways for connectivity from ACC to CA1?

      We thank the reviewer for this feedback. We have since added discussions on ACC-to-CA1 connectivity and discussed possible relay brain regions between ACC and CA1.

      (8) There is mounting evidence about the role of sleep oscillatory events playing homeostatic roles, not only memory-based. The authors bring this up, but do not offer it as an explanation for their findings, but I believe they probably should. For example, Norimoto et al 2018 cited by the authors. Also, Gulati/Gunguly et al 2017 Nature Neuroscience suggests downscaling as a default NonREM activity. Gulati and also Roux/Buzsaki NatNeuro 2017 show that certain privileged or tagged neurons can be protected from this. This therefore reflects that default activity in nonREM may have a homeostatic role, but then learning may alter that default. I believe this should be discussed as a possible reason for the dissociation between ACC and CA1, the authors observe after CFC.

      In more detail, the authors state that CFC worsened ACC ability to predict ripple spike rate vectors. The authors suggest this may reflect "worsened" communication from ACC to HPC. It could also reflect a SHIFT (not worsening) in the information state of the hippocampus, where ripples reflect novel information and/or information coming from other brain regions. Essentially, the novel information may out-compete usual information flows. For example, ACC may be a default "feeder" into ripples (for example as part of default mode network) when there was no recent highly salient information, but under non-default conditions such as after CFC, ripple content may be fed from other sources (be they internal or external to the HPC). I believe this should be discussed.

      For example, were CA1 sup task inactive neurons basically DMN-active neurons? Figure 2 shows neurons with the highest pre-training ACC prediction were the ones that dropped the most in training - again suggesting these neurons may be tuned to internal or default dynamics rather than CFC (or other novel experiences).

      This shift from a default communication mode to a more experience-based one should probably be discussed as an alternative explanation, rather than simply "worsening" of communication.

      We thank the reviewer for this feedback and their recommendation for possible alternate explanations. We have incorporated many of the listed citations and ideas they discussed into our discussion section proposing homeostatic downscaling and a shift in the default mode network as possible explanations for our results.

      Minor Weaknesses:

      (1) Introduction Line 52: "during replays" should probably be "during replay events".

      Changed.

      (2) Introduction Line 76: "how communications" should be "how communication"

      Changed.

      (3) 200-0, 400-200, 600-400 time bins are a bit unclear in Figure 1f. Are they really negative times, rather than positive? Perhaps negative signs could be put into the legend of 1f, or the time windows can be shown on the left side of 1e. Or potentially 1f could use the same -0.6, -0.4, etc as 1e so readers understand they are linked (if I understand correctly).

      They are negative in the sense the occur before the ripple event. Figure 1e now displays the negative signs.

      (4) I don't believe the methods describe how many tetrodes are put into the ACC. It would seem this should be put in the ACC portion of the "Stereotaxic surgery" section.

      We now clearly explain the number of tetrodes (8) in the "Stereotaxic surgery" section.

      (5) Results line 125: "learning induced" should be "learning-induced".

      Changed.

      (6) In terms of display, deep and sup are swapped in the various figures in terms of which is shown first/second (at least for readers assuming left is first). I suggest putting deep first or sup first in all figures. To me, sup first seems more natural, but homogeneity seems best regardless. This will help readers easily track results.

      We adjusted the figures so that superficial is typically displayed first with some exceptions. For example, in Figure 3A, CA1deep is shown first to preserve the anatomical (dorsal-ventral) relationship.

      (7) Results line 178: "optogenetics stimulation" should be "optogenetic stimulation".

      Changed.

      (8) Results line 180: "upon stimulations" should be "upon stimulation".

      Changed.

      (9) Results line 180: "to different capacities" could be "to different degrees" or "in different manners".

      Changed.

      Reviewer #2 (Recommendations for the authors):

      The sleep scoring procedure is not described clearly. The text references delta waves and ripple oscillations, but the accompanying citation (Wang et al. 2015) does not use such a procedure. Since the post-task rest sessions are called "sleep sessions", there is some confusion about whether the data was restricted to slow wave sleep or not. If data from each sleep session were taken without restricting to actual sleep, that would be problematic because animals may be less likely to sleep immediately following fear conditioning, which could introduce some sleep/wake bias into the comparisons. In particular, the relationship between cortex and ripple activity has been reported to dramatically change between awake and sleep states (Tang & Jadhav, 2019). I am not including this point in the public review in case the data was in fact restricted for sleep, and it simply needs to be clarified in the text.

      Thank you for this feedback. The recordings were in fact restricted to sleep. We have added text to the manuscript to make this clearer. Moreover, we added further discussion on how sleep was calculated.

      There is a puzzling paragraph in the discussion, arguing that "Here, we add to this understanding with CA1sup neurons having a diminished role in fear memory formation". Sparse task-related activity in CA1sup does not imply that CA1sup is not involved in memory. Indeed, while the median of CA1sup neurons' firing rate ratio was below zero, there is a substantial proportion of neurons that are recruited, and these could be extremely important for memory. Most studies on reactivation and replay would only concentrate on cells sufficiently active in the task, and observe whether these patterns of activity are enhanced in post-task sleep.

      We thank the reviewer for this feedback and have removed the text claiming sublayer difference in fear conditioning.

      If the ACC is indeed inhibiting the CA1sup pyramidal cells through V-type interneurons, then one would expect the average GLM weights predicting the activity of those best-predicted CA1sup cells to be negative. If that is true, that could nicely tie the prediction effect to the optogenetic results, demonstrating that ACC's relationship to pyramidal cells is inhibitory in natural conditions as well.

      We thank the reviewer for this suggestion. Unfortunately, our primary GLM analysis did not properly save weight coefficients to perform such analyses. To address this as closely as possible, we performed a preliminary analysis using a modified version of our GLM to examine whether the coefficients predicting CA1sup pyramidal neuron activity exhibited a bias toward negative weights. While this modified analysis did not generate coefficients directly comparable to the prediction gain values reported in the manuscript, it allowed us to assess whether an overall difference in coefficient sign was evident between CA1 sublayers. We found no significant bias toward negative coefficients and no clear differences between CA1sup and CA1deep neurons. Although this result does not provide additional support for an inhibitory relationship under natural conditions, it does not necessarily contradict our optogenetic findings. GLM coefficients quantify statistical dependencies between neural activities and reflect not only direct interactions but also indirect network effects, shared inputs, and the model structure. Consequently, the sign of a GLM coefficient should not be interpreted as a direct measure of whether the underlying synaptic relationship is excitatory or inhibitory.

      Figure 1f is strangely missing comparisons for positive delays. If such windows were to be included and if the reactivation gain is lower for them, that could really drive home the point that communication takes place in the ACC->CA1 direction more than in the CA1->ACC direction.

      Our goal of this study was to examine how incoming information from the cortex may differentially drive CA1 sublayer activity. While we think examining the reverse direction offers a compelling future direction, it was beyond the scope of our present manuscript.

      There is some confusion about the N-s. There's a total of 190 CA1 neurons (Figure 2a legend). 21 of them are deep, and 77 are sup (Figure 3b legend), so presumably 92 would be neither. The legend of Figure 4b agrees with this: 24 task-active CA1sup and 53 task-inactive CA1sup cells, while in CA1deep, there were 14 task-active and 7 task-inactive cells, but in Figure 3c, there is a comparison of N=24 CA1deep cells and n=94 CA1sup cells (so 72 neither).

      For one dual-site animal, the CFC recording file was corrupted, while the pre- and post-training recordings remained intact. As a result, this animal was included only in the GLM analyses, leading to slight differences in sample size across analyses. We have clarified these sample sizes in the revised manuscript and highlighted this discrepancy in the Methods section.

      There appears to be a typo on line 213, the reference should be to Figure 6f.

      Changed.

      Reviewer #3 (Recommendations for the authors):

      To what extent do the authors think that this is learning dependent, as opposed to stress dependent? Not that I want an experiment here, but it is important to note that both of these regions are very much involved in stress responses. Do you think you would get the same result with a purely appetitive learning paradigm? Or is this specific to stress? It might be nice to add this to the discussion.

      We thank the reviewer for this raising this point of discussion. Future experiments utilizing appetitive behavioral tasks could help address these important questions. While we have avoided speculating too much, we agree it is valuable to acknowledge this possibility. Accordingly, we have added a brief discussion point to at least call attention to this possibility for the reader.

      “Another caveat to mention is that both the ACC and CA1 are involved in the stress response (Kim et al., 2015; Lamotte et al., 2021). Future experiments utilizing appetitive learning paradigms, rather than the aversive contextual fear conditioning used here, will help disentangle learning-related remodeling of ACC→CA1 communication from changes driven by stress.”

      There are a number of typos that confuse the message - for example, in the abstract, the authors say that ACC suppresses superficial CA1 interneurons. This seems most likely an error - I think the authors mean superficial CA1 neurons? Or PV interneurons? Similar errors exist throughout, as well as odd combinations of bold and italics across and within words, etc. Overall, especially in consideration of my final main point above regarding clarity of the text, it would be good to have a proper proofread to make sure the text is as clear as possible

      We thank the reviewer for identifying these issues. The specific error in the abstract has been corrected, and we have since carefully proofread the revised manuscript.

    1. eLife Assessment

      This valuable study uses EEG and computational modeling to investigate hemispheric oscillatory asymmetries in unilateral spatial neglect. The work benefits from rare patient data and a careful multimethod approach. The study manuscript provides convincing evidence contributing to our understanding of the neural dynamics underlying spatial neglect.

    2. Reviewer #2 (Public review):

      This study investigates how altered neural oscillations may contribute to unilateral spatial neglect (USN) following right-hemisphere stroke. By combining steady-state visual evoked potentials (SSVEPs), phase-amplitude coupling (PAC), transfer entropy (TE), and computational modeling, the authors aim to show that USN arises from disrupted hemispheric synchronization dynamics rather than simply from lesion extent. The integration of empirical EEG data with a mechanistic model is a major strength and offers a valuable new perspective on how frequency-specific neural dynamics relate to clinical symptoms.

      The work has several notable strengths. The combination of experimental and modeling approaches is innovative and powerful, and the findings provide a coherent mechanistic framework linking abnormal neural entrainment to attentional deficits. The study also provides concrete compelling evidence supporting the potential for frequency-specific neuromodulatory interventions, which could have translational relevance.

      In the revised manuscript, the authors have carefully and comprehensively addressed the concerns raised during the first round of review. In particular, the additional characterization of lesion distribution and volume provides important anatomical context for the electrophysiological findings, while the rationale for the choice of electrodes and clinical correlation analyses is now much clearer. The methodological description has also been improved substantially, including clarification of the SSVEP measure, analysis procedures, and potential confounds related to transfer entropy and volume conduction. In addition, the discussion now provides a more nuanced account of the relationship between stimulus-locked responses and intrinsic oscillatory activity, as well as the potential contribution of alpha lateralization to attentional dysfunction.

      Overall, I consider the revised manuscript to provide compelling evidence for an important contribution to our understanding of the neural dynamics underlying spatial neglect. The authors have addressed my previous concerns satisfactorily, and the manuscript now provides a clearer and more balanced account of both the strengths and limitations of the findings. It should serve as a valuable reference for future work on oscillatory mechanisms in stroke and attention.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Okazaki et al. showed flickering stimuli to patients with unilateral spatial neglect (USN) and measured EEG responses. They compared this with another patient group (post-stroke, but no USN) and healthy controls. The author's rationale was to entrain intrinsic brain rhythms using the flicker of different frequencies (3-30 Hz). Effects found unique to the 9-Hz stimulation condition differentiate USN patients from the other groups, leading them to conclude that USN can be characterized by increased hemispheric alpha asymmetry, driven by a relatively increased response in the intact hemisphere.

      Strengths:

      This study is principled empirical work that benefits from access to special patient groups of considerable size (about 60 stroke patients in total, and 20 USN). The authors use state-of-the-art established methods to (1) deliver and (2) quantify the responses to the flicker stimulation in the EEG recordings. In addition, they use phase-coupling measures to investigate cross-frequency coupling (here: alphagamma) and a measure of directed connectivity between brain areas, transfer entropy. The results are supported by means of simulations using a coupled oscillators model.

      Weaknesses:

      In my eyes, the major conceptual weakness of the study is that the authors make the a priori assumption that the flicker stimulation entrains intrinsic brain rhythms, especially alpha (9 Hz). To date, there is no direct (and only equivocal indirect) evidence that alpha rhythms can be entrained with periodic visual stimulation. In the present study, the assumption of alpha entrainment permeates some analytical decisions - where it would be possible to separate stimulus-driven from intrinsic rhythms more strongly than is currently the case, potentially yielding deeper insights into the oscillopathy of USN - and, ultimately, the interpretation of the results. Another potential issue to consider here is the analysis of gamma rhythms in EEG data, absent a control of miniature eye movements, a known problem (YuvalGreenberg et al., 2008, https://doi.org/10.1016/j.neuron.2008.03.027) that may be exacerbated here, given that USN patients could show different auxiliary gaze behaviour.

      We thank Reviewer #1 for the careful and constructive evaluation of our study, and for recognizing the strengths of the patient cohort and our combined empirical and computational approach. We also appreciate the reviewer’s concern that our original wording could be read as assuming that flicker stimulation necessarily entrains intrinsic alpha rhythms. In the revised manuscript, we have clarified that our interpretation is based on frequency-specific stimulus-locked responses and model-based inference, rather than on an a priori assumption of entrainment. We have also revised the relevant parts of the Introduction and Discussion to distinguish more clearly between stimulus-locked responses and intrinsic oscillatory dynamics. In addition, we have expanded our discussion of the potential influence of miniature eye movements on gamma-band activity and PAC. These points are addressed in detail in our responses to the Recommendations for the Authors below.

      Reviewer #2 (Public review):

      This study investigates how altered neural oscillations may contribute to unilateral spatial neglect (USN) following right-hemisphere stroke. By combining steady-state visual evoked potentials (SSVEPs), phase-amplitude coupling (PAC), transfer entropy (TE), and computational modeling, the authors aim to show that USN arises from disrupted hemispheric synchronization dynamics rather than simply from lesion extent. The integration of empirical EEG data with a mechanistic model is a major strength and offers a valuable new perspective on how frequency-specific neural dynamics relate to clinical symptoms.

      The work has several notable strengths. The combination of experimental and modeling approaches is innovative and powerful, and the findings provide a coherent mechanistic framework linking abnormal neural entrainment to attentional deficits. The study also provides concrete evidence to support the potential for frequency specific neuromodulatory interventions, which could have translational relevance.

      At the same time, there are areas where the evidence could be clarified or contextualized further. The manuscript would benefit from more detailed characterization of lesions, since differences in lesion topography (white vs. gray matter, occipital vs. parietal areas) could greatly improve our understanding of the physiopathology causing unilateral spatial neglect and the altered neural oscillations reported. Methodological choices, such as focusing analyses on occipital electrodes rather than parietal sites, and the potential influence of volume conduction in transfer entropy analyses, also need clearer justification/elaboration. In addition, while the authors report several neural metrics, it is not always clear why SSVEP power was chosen as the primary correlate of clinical severity over other measures. More broadly, the manuscript would be strengthened by clearer definitions of dependent variables and reporting of software and toolboxes used.

      Overall, the study makes a significant contribution by demonstrating that USN can be conceptualized as a disorder of disrupted oscillatory dynamics. With some clarifications and expansions, the paper will provide readers with a clearer understanding of both the strengths and the limitations of the evidence, and it will stand as a valuable reference for future work on oscillatory mechanisms in stroke and attention.

      We thank Reviewer #2 for the positive assessment of our integrated empirical and computational approach, and for highlighting the potential contribution of our findings to understanding oscillatory mechanisms in USN.

      In response to the reviewer’s comments, we have added new supplementary figures showing lesion overlap maps and lesion-volume analyses (Supplementary Figure 1), additional analyses related to electrode selection and SSVEP topography (Supplementary Figure 2), and clinical-correlation analyses of hemispheric imbalance measures (Supplementary Figure 3). We have also revised the manuscript to better contextualize lesion topography and lesion extent, clarify the rationale for the occipital-electrode and clinical-correlation analyses, and provide additional methodological details. These issues are addressed in detail in our point-by-point responses to the Recommendations for the Authors below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I found your manuscript well-written and hence easy to follow. Allow me to provide some recommendations to tackle the "weaknesses":

      (1) I have mentioned that the entrainment assumption is not warranted based on the current evidence, but I also think that the analysis and interpretation is critically constrained by this. The way the analysis is carried out conflates stimulus-driven and intrinsic brain rhythms, especially in the alpha band. In other words, spectral representations of the EEG data will likely be dominated by natural alpha rhythms, whereas the stimulus-driven signals could be accentuated by a different analysis, time-locked to the stimulation. This is explained in greater detail in Keitel et al. (2019, https://doi.org/10.1523/JNEUROSCI.1633-18.2019). Looking into alpha and stimulus-driven responses separately, without the assumption of entrainment, may actually allow a more complete picture of the impact of USN because the stimulus driven SSVEPs are taken to indicate different cortical processes than alpha (see e.g., Duecker et al., 2021, https://doi.org/10.1523/JNEUROSCI.3134-20.2021, though for gamma). Importantly, this all does not exclude the possibility that alpha rhythms were indeed entrained here, but given the current situation, this should be an outcome of the study rather than an a-priori assumption. I suggest re-framing the manuscript this way.

      We agree that the current wording could be read as presuming alpha entrainment. We have revised the framing and terminology to reflect that entrainment is inferred from the 9-Hz–specific stimulation effect and the absence of a corresponding hemispheric bias at rest, with further support from the resonance mechanism demonstrated by our computational model. We also point readers to the Discussion “Alpha frequency-specific hemispheric bias and entrainment in USN patients”, where we explain why the present 9-Hz–specific findings are interpreted in terms of alpha-range resonance/phase alignment and how this is expressed in EEG. We revised the Introduction SSER description (p.4) to:

      “SSERs are elicited by rhythmic sensory stimulation and provide a measure of stimulus-locked neural responses. When the stimulation frequency is close to the system’s intrinsic resonance frequency, these responses may include an entrainment component, reflecting the alignment of endogenous oscillations to external input (Pikovsky, Rosenblum, and Kurths 2003; Okazaki et al. 2021)”

      We have also revised the Introduction hypothesis statement (p.5) to avoid implying an a priori entrainment assumption, replacing it with:

      “We hypothesized that frequency-specific stimulus-locked synchronization dynamics in response to rhythmic stimulation would be selectively disrupted in one hemisphere, resulting in an interhemispheric imbalance in USN.”

      Finally, to ensure consistent framing in the Discussion, “Alpha frequency-specific hemispheric bias and entrainment in USN patients”, we have revised the opening sentence (p.20) as follows:

      “We observed a hemispheric bias in stimulus-locked responses to flickering stimuli in USN patients in the alpha range, which corresponds to the natural frequency of the visual system (Rosanova et al. 2009; Okazaki et al. 2021).”

      (2) Transfer Entropy is used as a measure of directed connectivity and applied to narrow-band filtered EEG signals. Methodological issues have been pointed out with regard to that (Daube et al., 2022, https://doi.org/10.48550/arXiv.2201.02461). Has this been considered?

      We thank the reviewer for raising this important methodological concern. As Daube et al. (2022) noted, Transfer Entropy (TE) can be overestimated when applied to narrow-band signals with strong autocorrelation. However, in our analysis, TE was computed not from the narrow-band 9-Hz waveform itself but from its instantaneous amplitude (amplitude envelope), which fluctuates nonperiodically on a slower timescale, thereby reducing the risk of spurious causality driven by sinusoidal autocorrelation. We clarified this explicitly in the Methods, “Transfer entropy (TE)” (p.8) by adding:

      “After applying an 8.5–9.5 Hz FIR bandpass filter to the EEG responses to 9-Hz flickering stimuli, we extracted the instantaneous amplitude (amplitude envelope) from the analytic signal using the Hilbert transform. Because this amplitude envelope fluctuates nonperiodically at a slower timescale than the carrier 9-Hz oscillation, it provides a broadband measure of signal dynamics while minimizing the strong autocorrelation inherent in narrow-band periodic signals that can spuriously inflate TE estimates (Daube C et al., 2022).”

      We have also clarified the relationship between TE and volume conduction in the same section:

      “Importantly, because TE evaluates time-lagged prediction (from Y(t) to X(t+τ)), it is not designed to capture zero-lag common-source correlations (i.e., volume conduction) and therefore characterizes directed, nonzero-lag dependencies rather than an instantaneous coupling.”

      Finally, we have added an explicit interpretation emphasizing the direction-specific nature of the effect in the Discussion, “Biased information transfer in USN patients” (p.23):

      “This directional asymmetry argues against spurious overestimation of TE due to autocorrelation or volume conduction (Daube, Gross, and Ince 2022), because such pseudo-causal effects would be expected to manifest more symmetrically in both directions. In addition, by computing TE from the amplitude envelope of the 9-Hz activity, we reduced the influence of strong autocorrelation inherent in narrow-band oscillatory signals and thereby minimized the conditions that Daube et al. identified as leading to TE overestimation. Taken together, these points suggest that the observed TE asymmetry is unlikely to be explained solely by methodological artifacts and may reflect a genuine directional imbalance in interregional communication following right-hemisphere damage.”

      (3) If a closer control of miniature eye movements is not possible, I suggest removing any gamma analysis, or at least prominently mentioning the caveat that gamma activity may be contaminated by eye movement artifacts.

      We agree that the contribution of miniature eye movements cannot be completely excluded. However, we consider it unlikely that the hemispheric asymmetry in alpha– gamma PAC reported in this study mainly arises from eye-movement artifacts, for the following reasons. First, SP (saccadic spike potential)-related PAC would require saccade timing to be tightly phase-locked to the 9-Hz cycle. However, SPs are time-locked to saccade onset and typically cluster around 200–300 ms after stimulus onset (Yuval-Greenberg et al., 2008; Keren et al., 2010). Thus, it is unlikely that they would be consistently phase-locked to the 9-Hz cycle (≈111 ms), and even modest temporal jitter would markedly blur PAC. Second, because SPs are brief spike-like transients with broadband high-frequency components, periodic SP contamination would be expected to yield a broadband gamma profile, rather than the relatively narrow band (35–45 Hz) observed here. In addition, we directly compared gamma-band power (35–45 Hz) at O1 and O2 during 9-Hz stimulation using the same gamma range as in the PAC analysis and found no significant hemispheric differences in any group (Author response image 1), arguing against a systematic unilateral increase in gamma power driven by asymmetric saccade behavior. Taken together, these considerations make it difficult to attribute the observed alpha–gamma PAC asymmetry primarily to SPs arising from eye movements. Nonetheless, residual eye-movement artifacts cannot be completely ruled out, and we now explicitly state this limitation in the Discussion, “Limitations” (p. 25) by adding:

      “Fourth, the hemispheric bias in alpha–gamma PAC should be interpreted in light of potential contamination from miniature saccades (Yuval-Greenberg, Tomer, Keren, Nelken, & Deouell, 2008). However, several observations make it unlikely that such artifacts are the primary source of the effect. For saccadic spike potentials to account for the PAC under 9-Hz stimulation, they would need to occur in a highly periodic and phase-locked manner relative to the 9-Hz cycle. This scenario is unlikely given the stimulus-locked dynamics of miniature saccades. (i.e., post-stimulus inhibition followed by a rebound around 200–300 ms) (Yuval-Greenberg & Deouell, 2009). Moreover, SP-related contamination would be expected to yield a broadband gamma profile (~20–90 Hz) (Keren, Yuval-Greenberg, & Deouell, 2010; Yuval-Greenberg & Deouell, 2009), rather than the relatively narrow band (35–45 Hz) observed here. In addition, we found that gamma-band power in the 35–45 Hz range did not show any hemispheric difference between O1 and O2 during 9-Hz stimulation (data not shown).”

      Author response image 1.

      Hemispheric differences in gamma-band (35–45 Hz) power during 9-Hz flicker stimulation. Left (O1) − Right (O2) gamma power did not differ from zero in any group (non-USN: p = 0.45; USN: p = 0.34; healthy: p = 0.35). Error bars represent 2 standard errors of the mean.

      (4) Please provide more methodological detail on the resting state recordings. When, how, and under which circumstances were these recorded?

      Thank you for pointing this out. We clarified when the resting-state interval was taken in the Methods, “Steady-state visual evoked potential (SSVEP)” (p.7) by adding:

      “For the resting-state interval, power was estimated using the same procedure from the ‘off’ interval immediately preceding the 3-Hz flicker block (see Fig. 1)”

      (5) Provide power spectra of the EEG data for illustration - ideally for resting state and stimulation conditions. These should allow the reader to visually evaluate the effects of different stimulation frequencies, as well as differences between participant groups.

      Figure 2 already presents spectra normalized to the resting-state baseline. To make this explicit for readers, we clarified this point in the Figure 2 caption (p.11) by adding:

      “Spectra are normalized to the baseline from the resting-state interval.”

      Reviewer #2 (Recommendations for the authors):

      (1) The authors indicate L.608 "the precise extent and topography of brain lesions could not be fully homogenized across patients", but the manuscript would highly benefit from any additional detail that could be obtained from characterization of the lesion sites. At minimum, it would be important to indicate for USN and non-USN groups whether lesions predominantly affected white matter or gray matter, and whether occipital versus parietal cortices were involved (e.g., using MRI atlas templates). If this coarse information can be obtained, the authors could test whether any of these anatomical details can distinguish are different between USN and nonUSN patients. This information could help the reader evaluate whether the reported neural asymmetries might be driven by lesion topography rather than purely by oscillatory dynamics.

      We understand this comment as raising the important concern that lesion location and lesion volume may differ between the USN and non-USN groups and could contribute to the observed neural asymmetry. We agree that the presence of USN and the alteration of oscillatory dynamics should be interpreted in relation to which regions and networks are damaged, and to what extent. In response to the reviewer’s suggestion, we generated lesion overlap maps based on the available structural images and additionally quantified lesion volume for each patient (replaced Supplementary Figure 1). This analysis confirmed that lesion volume was significantly larger in the USN group than in the non-USN group. Thus, lesion volume is an important anatomical factor to consider when interpreting group differences in USN and neural responses.

      At the same time, the present results suggest that lesion volume and coarse lesion topography alone are not sufficient to explain the 9 Hz-specific imbalance in interhemispheric synchrony. Interhemispheric synchrony depends on distributed network functions involving multiple cortical and subcortical regions and their connecting pathways, and similar functional imbalances may arise from different patterns of anatomical damage. Moreover, although the newly added lesion maps confirmed more extensive lesions in the USN group, lesions in both groups predominantly involved the right MCA territory, and lesion extent and location varied across patients. Importantly, the hemispheric asymmetry in neural responses emerged selectively in the 9 Hz condition, whereas SSVEP responses at other frequencies were largely balanced between hemispheres. If lesion volume or coarse lesion topography alone were sufficient to explain the effect, one might expect a more uniform reduction across frequencies or a simpler pattern corresponding to lesion extent.

      Accordingly, we do not treat lesion topography and oscillatory dynamics as competing explanations. Instead, we regard them as hierarchically related: anatomical damage alters network components, and this in turn gives rise to a frequency-specific imbalance in synchronization capacity. To clarify this point, we replaced the previous Supplementary Figure 1 with a new figure showing lesion overlap maps and lesion volume information, and revised the Discussion section “Distinct neural responses in non-USN and USN patients” (p. 24) as follows.

      “To further characterize the anatomical background of these group differences, we generated lesion overlap maps and quantified lesion volume in the USN and non-USN groups (Supplementary Figure 1). Both groups predominantly showed lesions involving the right MCA territory, but lesion volume was significantly larger in the USN group than in the non-USN group (USN: 72,419 ± 73,778 mm<sup>3</sup>; non-USN: 11,739 ± 23,907 mm<sup>3</sup>; Mann–Whitney U = 361.0, p = 1.41 × 10<sup>-5</sup>). These anatomical differences indicate that lesion extent is an important factor associated with USN. However, they do not by themselves fully explain the frequency-specific neural effect observed here. Importantly, this frequency specificity coincides with the intrinsic alpha frequency of the visual system. This correspondence suggests that the present finding may not simply arise from lesion location or lesion volume alone, but may instead reflect a more complex mechanism involving selective functional disruption of oscillatory networks. From this perspective, lesion topography and oscillatory dynamics should be regarded not as competing explanations, but as different levels at which the same pathological condition can be understood. The key question is which network components are affected and how their dysfunction gives rise to the frequency-selective hemispheric imbalance in synchronization capacity at 9 Hz. Thus, even if USN and non-USN patients differ in lesion extent and aspects of lesion topography, this does not undermine the present interpretation, but rather highlights the need to examine how anatomical damage relates to frequency-specific network dysfunction.”

      We have also referred to our computational account and clarified its implication for the functional mechanism in the same subsection (p. 24):

      “Our computational model illustrates a plausible mechanism for such a process. When two coupled oscillators sharing the same intrinsic alpha frequency are connected via asymmetric interhemispheric couplings, the model selectively produces an imbalance in synchrony at the resonant frequency, whereas responses at non-resonant stimulation frequencies remain relatively balanced between hemispheres. This model result suggests that post-lesion asymmetry in interhemispheric coupling may bias alpha-band information processing (e.g. phase-dependent sampling/synchrony) between hemispheres, and may consequently manifest as systematic biases in perceptual and attentional allocation.”

      We have also revised the Limitations (p. 25) to acknowledge that the present lesion analyses characterize the distribution and extent of lesions across groups, but do not directly identify which anatomical network disruptions give rise to the 9 Hz-specific imbalance in interhemispheric synchrony.

      “Second, although we added lesion overlap maps and quantified lesion volume, these analyses primarily characterize where and how extensively lesions were distributed across the two patient groups. They do not directly identify which anatomical network disruptions give rise to the 9 Hz-specific imbalance in interhemispheric synchrony. Future studies with larger cohorts will be necessary to combine detailed lesion-symptom mapping and assessments of white-matter disconnection with neural synchrony analyses to determine which anatomical network disruptions lead to frequency-specific alterations in oscillatory dynamics after stroke.”

      (2) The rationale for focusing on O1 and O2 electrodes should be clarified. Given that neglect is classically associated with parietal dysfunction, one would expect analyses of parietal electrodes to be informative. Could the authors justify their choice, and possibly report whether similar effects were (or were not) observed at parietal sites?

      Our primary SSVEP analyses focused on O1 and O2 because flicker stimulation is designed to drive the visual system, and the fundamental SSVEP component is typically maximal over occipital electrodes in healthy participants (Norcia et al., 2015). We also note that hemispheric differences at central and frontal electrodes are already shown in Fig. 4 and described in the Results, demonstrating no significant left–right differences at these sites across any stimulation frequency conditions, including 9 Hz. In addition, we added a supplementary figure showing the scalp topography of the mean fundamental-frequency SSVEP power averaged across all stimulation conditions, which displays the expected occipital maximum and thus a typical SSVEP spatial profile (Supplementary Fig. 2, shown below). We added the rationale and the reference in the Methods, “Steady-state visual evoked potential (SSVEP)” (p. 7) section by adding:

      “SSVEP power at the fundamental (stimulated) frequency was then extracted for subsequent analyses. Because the fundamental SSVEP component is typically maximal over occipital electrodes in healthy participants (Norcia et al., 2015), we evaluated occipital electrodes (O1/O2) as primary sites. The scalp topography of the mean fundamental-frequency SSVEP power averaged across stimulation conditions is shown in Supplementary Fig. 2.”

      (3) For the mutual information and transfer entropy analyses, the potential influence of volume conduction should be acknowledged. Numerous studies mitigate this issue by applying source reconstruction or connectivity metrics that are insensitive to zerolag correlations. Even if the present study did not use such approaches, the authors should discuss the extent to which volume conduction might confound their results, and ideally provide some justification for why their findings remain valid.

      Regarding directed connectivity, our primary analysis uses Transfer Entropy (TE), which evaluates time-lagged prediction and is therefore, by definition, not directly sensitive to zero-lag common-source correlations. Consistently, our key finding is direction-specific (feedforward only), which is not readily explained by symmetric, zero-lag relationships typical of volume conduction. We stated this explicitly in the Methods, “Transfer entropy (TE)” (p.8) by adding:

      “Importantly, because TE evaluates time-lagged prediction (from Y(t) to X(t+τ)), it is not designed to capture zero-lag common-source correlations (i.e., volume conduction) and therefore characterizes directed, nonzero-lag dependencies rather than instantaneous coupling.”

      We have also clarified the direction-specific logic in the Discussion, “Biased information transfer in USN patients” (p.22):

      “This hemispheric difference appeared only in the Feedforward (visual-to-frontal) direction, and no such difference was observed in the Feedback (frontal-to-visual) direction. This directional asymmetry argues against spurious overestimation of TE due to autocorrelation or volume conduction (Daube C, Gross J and Ince RAA, 2022), because such pseudo-causal effects would be expected to manifest more symmetrically in both directions. In addition, by computing TE from the amplitude envelope of the 9-Hz activity, we reduced the influence of strong autocorrelation inherent in narrow-band oscillatory signals and thereby minimized the conditions that Daube et al. identified as leading to TE overestimation. Taken together, these points suggest that the observed TE asymmetry is unlikely to be explained solely by methodological artifacts and may reflect a genuine directional imbalance in interregional communication following right-hemisphere damage.”

      (4) In line 579, the authors write 'patients with USN may have larger lesions than those without USN.' It would be useful to clarify whether this claim is supported by the present data (i.e., lesion size comparisons between the two patient groups) or whether it reflects prior literature. If based on the current sample, please provide the corresponding statistical evidence.

      This issue has been addressed as part of our response to comment (1), with corresponding revisions made to the Discussion (p. 24), the Limitations (p. 25), and Supplementary Figure 1.

      (5) The dependent variable for SSVEP analyses is not clearly described, which makes it difficult to understand the results (e.g., L.214, L285-297). It seems that for all stimulation conditions, one value was extracted (and then compared between O1 and O2 electrodes), but is that the amplitude at the stimulated frequency (e.g., for a visual stimulation at f Hz, comparing the amplitude of the power spectrum at f Hz between O1 and O2)?

      Thank you for pointing this out. We clarified that the dependent variable for the SSVEP analyses was the power at the fundamental (stimulated) frequency f for each condition. We added this explicitly to the Methods, “Statistical analysis” (p.9):

      “The SSVEP power at the fundamental (stimulated) frequency f was analyzed for each condition as the dependent variable...”

      (6) Most analyses are described relatively precisely mathematically, but no toolbox or software is mentioned. Were they implemented using custom scripts? If toolboxes/software was used, the authors should indicate which ones and their versions, and add references. For instance, I might be wrong, but I guess the linear mixed model analyses were not performed using custom code. MATLAB is mentioned for the simulations, but no version is indicated. This is particularly important for reproducibility since the authors did not provide direct access to their scripts.

      The EEG preprocessing, PAC, and cluster-based permutation tests were conducted using the FieldTrip toolbox integrated with custom MATLAB scripts. Linear mixed-effects models were performed in SPSS. We added this information to the Methods, “Preprocessing” (p.7):

      “All EEG analyses were implemented using custom scripts in MATLAB (MathWorks, Natick, MA, USA) and the FieldTrip toolbox (Oostenveld R et al., 2011).”

      We have also specified the tools used for the linear mixed-effects models and the cluster-based permutation test in the Methods, “Statistical analysis” (p.9):

      “…using a linear mixed model in SPSS… a cluster-based permutation test (Maris E and Oostenveld R, 2007) using FieldTrip.”

      (7) The rationale for correlating BIT scores specifically with SSVEP power (rather than with the hemispheric imbalance measure, or with MI/TE results) is not clearly articulated. The results section highlights several neural metrics (SSVEP imbalance, PAC, TE), so it would strengthen the manuscript if the authors explained why SSVEP power was prioritized for correlation analyses. Is there a theoretical or empirical reason for this choice?

      While our main group-level results demonstrate a frequency-selective hemispheric imbalance, Fig. 5 suggests that the lesioned and intact hemispheres are not necessarily simple mirror images of each other and can show distinct response profiles. Therefore, to clarify which hemisphere’s response changes are associated with symptom severity, we assessed correlations with BIT using hemisphere-specific SSVEP power. At the same time, it is also informative to directly test the extent to which symptom severity covaries with hemispheric imbalance metrics (i.e., left–right differences in SSVEP, PAC, and TE). Accordingly, in the revised manuscript we additionally analyzed correlations between BIT and hemispheric imbalance measures of SSVEP, PAC, and TE, and report these results in the Supplementary Information (Supplementary Fig. 3). We added this clarification to the Results, “Correlation between SSVEP power and USN severity” (p.17):

      “In addition, correlations between BIT and the hemispheric imbalance of SSVEP power are provided in the Supplementary Information (Supplementary Fig. 3A). We also report, as additional exploratory analyses, correlations between BIT and hemispheric imbalance measures derived from PAC and TE (Supplementary Fig. 3B and 3C). None of these correlations reached statistical significance after correction, although the feedforward TE imbalance to the left frontal region showed a marginal uncorrected association with BIT score (r = -0.51, uncorrected p = 0.05).”

      We have also added the following statement to the Discussion, “Biased information transfer in USN patients” (p.23).

      “This interpretation is also consistent with the trend-level correlation shown in Supplementary Figure 3C. Although the correlation did not survive correction for multiple comparisons, patients with a stronger feedforward TE bias toward the left frontal region tended to show higher BIT scores. This observation raises the possibility that asymmetric information transfer from the right visual cortex to the intact left frontal cortex may contribute to compensatory network reorganization associated with milder neglect symptoms.”

      (8) The discussion could be enriched by considering whether the observed abnormal alpha-band entrainment relates to the well-known alpha-lateralization phenomenon in spatial attention. In healthy individuals, covert attentional orienting is typically accompanied by lateralized modulations of alpha power (increased ipsilateral, decreased contralateral). Could the asymmetric alpha entrainment reported here be interpreted as a pathological exaggeration of this mechanism, thereby linking the electrophysiological findings more directly to attentional orientation deficits in neglect?

      We thank the reviewer for raising this valuable point. To integrate our results with the alpha-lateralization literature, we added a new paragraph to the Discussion, “Local hemispheric bias of alpha-band entrainment and attentional dysfunction in USN patients” (p.22) subsection

      “Clinically, USN presents as neglect of the left visual field and is generally thought to reflect a relative dominance of rightward orienting. Because alpha power is often interpreted as reflecting functional inhibition (Worden et al. 2000; Kelly et al. 2006; Foxe and Snyder 2011), classic alpha-power lateralization associated with rightward orienting would typically predict increased alpha power over the right hemisphere and decreased alpha power over the left. However, in our data, right-hemisphere alpha power during stimulation was comparable to that of controls, whereas the intact left hemisphere exhibited stronger stimulus-locked synchronization (SSVEP) and enhanced alpha–gamma coupling. Thus, the present findings do not appear to reflect a simple amplification of tonic spatial alpha-power lateralization. Rather, they suggest a dynamic, temporally selective form of alpha-mediated inhibition that organizes the timing of local excitability. From this perspective, alpha-power lateralization and stimulus-locked entrainment may reflect related but distinct aspects of alpha-based inhibitory control, with the former regulating spatial gating and the latter regulating the timing of sensory processing (Jensen and Mazaheri 2010).

      Moreover, such a timing-based bias may arise even in the absence of explicit flicker. Visual input is continuously sampled under the influence of intrinsic alpha activity, and perceptual sensitivity fluctuates with the phase of ongoing oscillations (Romei et al. 2008; Iemi et al. 2017; VanRullen 2016). As a consequence, processing is relatively facilitated when incoming events coincide with high-excitability phases and relatively suppressed otherwise (Busch, Dubois, and VanRullen 2009; Mathewson et al. 2010). Thus, a hemispheric bias in alpha-band synchronization capacity could bias the timing with which sensory input is sampled across the two hemispheres, even without explicit rhythmic stimulation. This view is also consistent with the possibility that the bias is less apparent during eyes-closed rest with minimal visual input, yet becomes behaviorally expressed in everyday settings where continuous input is sampled in a phase-dependent manner (Landau and Fries 2012; VanRullen 2016). Overall, a hemispheric bias in synchronization capacity within the alpha range likely disrupts the temporally coordinated sampling of visual input required for balanced spatial attention, contributing to the attentional deficits observed in USN.”

      (9) All reported t-tests should include degrees of freedom (df). This is essential for transparency and allows readers to assess the robustness of the statistical results.

      We added the degrees of freedom for the t-tests reported in the Results, “Hemispheric imbalance of TE” (p.15):

    1. eLife Assessment

      The study presents a valuable conceptual framework by classifying pattern-forming gene subnetworks into three established categories. By meticulously enumerating and simulating various network topologies, the authors provide a systematic catalog of these patterning motifs. However, the supporting evidence remains incomplete, as the mathematical generalizations rely on simplified assumptions that may not hold in more complex or realistic scenarios.

    2. Reviewer #1 (Public review):

      Summary:

      The authors tackle a long-standing question in developmental theory: given a gene-regulatory network that includes extracellular signaling, which topologies are even capable of transforming an initial spatial profile into a genuinely new pattern? Building on the classical reaction-diffusion framework in one dimension, but imposing biologically motivated constraints, they prove that every one-signal sub-network must be either Hierarchical (H), self-activating (L+), or self-inhibiting (L-). They further demonstrate that only three composite classes of full networks - pure H, a coupled L+ L- "Turing" pair, and an L- module fed by an intracellular positive loop ("noise-amplifying")-can create non-trivial spatial transformations. Analytical criteria and illustrative simulations are provided, together providing a closed taxonomy, which is supposed to be relevant for real systems.

      Strengths:

      - Useful classification framework. Reducing a vast number of possible gene circuits to three canonical pattern-forming motifs is a valuable organizing insight for both theorists and experimentalists.

      - Practical interpretability. Given a reaction network diagram, one can now decide (assuming the model applies to real systems) whether spatial patterning is even possible, saving experimental effort on in silico screens that could never succeed.

      Weaknesses:

      - Theoretical limitations in the application of Linear Stability Analysis (LSA): I remain uncertain about the framework's reliance on LSA as a necessary condition for non-trivial pattern transformation, especially for large initial perturbations ("spikes"). The revised manuscript itself states that spike amplitudes must be sufficiently small for the linearization to hold. In the rebuttal, the authors argue that large spikes can nevertheless be treated because their influence is initially small outside the spike. However, linear stability of a homogeneous steady state only describes the response to infinitesimal perturbations around that state; it does not generally exclude finite-amplitude perturbations from entering a different nonlinear basin of attraction and producing a heterogeneous stationary state, e.g., as in subcritical Turing patterns. Thus, I do not think the rebuttal establishes the stronger claim that a linearly stable network cannot produce a non-trivial pattern regardless of nonlinear terms.

      - Presentation: The manuscript remains difficult to follow. The argument is distributed across many named requirements and topology classes, long prose descriptions of network structures, and repeated cross-references to the Supplementary Information. Given that the main contribution is a conceptual classification, I think the logical hierarchy should be considerably easier to reconstruct.

      Discussion:

      The study offers a solid conceptual organization of pattern-forming networks. However, the theoretical bridge between infinitesimal linear stability and macroscopic, non-linear pattern emergence still presents some uncertainties. The way the current framework formally treats large initial perturbations leaves some questions open regarding its broad analytical applicability to real biological tissues.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors tackle a long-standing question in developmental theory: given a gene-regulatory network that includes extracellular signaling, which topologies are even capable of transforming an initial spatial profile into a genuinely new pattern? Building on the classical reaction-diffusion framework in one dimension, but imposing biologically motivated constraints, they prove that every one-signal sub-network must be either Hierarchical (H), self-activating (L+), or selfinhibiting (L-). They further demonstrate that only three composite classes of full networks - pure H, a coupled L+ L- "Turing" pair, and an L- module fed by an intracellular positive loop ("noise-amplifying")-can create non-trivial spatial transformations. Analytical criteria and illustrative simulations are provided, together providing a closed taxonomy, which is supposed to be relevant for real systems.

      Strengths:

      (1) Useful classification framework. Reducing a vast number of possible gene circuits to three canonical patternforming motifs is a valuable organizing insight for both theorists and ---experimentalists.

      (2) Practical interpretability. Given a reaction network diagram, one can now decide (assuming the model applies to real systems) whether spatial patterning is even possible, saving experimental effort on in silico screens that could never succeed.

      Weaknesses:

      (1) After the resubmission, I still have concerns regarding the formal definition of "non-trivial transformations" (P1/P2) and its application to noisy or multi-dimensional systems. The criteria rely on counting "new" critical points (maxima/minima). In their response, the authors argue that the diffusion operator instantly smooths discontinuous white noise, allowing critical points to be properly defined. However, this very smoothing process passively generates a landscape of new, smooth local extrema from the initial noise. Consequently, trivial diffusive regularization could inadvertently fulfil the criteria for a "non-trivial" transformation, leaving the definition conceptually problematic.

      That is indeed the case: diffusion alone can generate new concentration maxima and minima; but these would be transient unless there are some self-activatory loops (as we detail over the article). If these concentration maxima and minima are transient in time, they do not count for pattern transformation. In the two version of the article we explicitly stated (in P1 in the introduction when defining pattern transformation) that we only consider resulting patterns that are stable in time. This implies that the concentration maxima and minima that may transiently arise from diffusion alone do not count for the definition of pattern transformation. In the current new version (the third) and after the reviewer’s suggestion, we insist (in red text) that the new concentration maxima and minima need to be stable in time (i.e. non-transient).

      Furthermore, when extending the framework to 2D/3D, the manuscript assumes that starting from a central "spike" will robustly preserve radial symmetry, yielding concentric rings or shells. This overlooks the fundamental nature of macroscopic mean-field models like reaction-diffusion equations. The realization of the final multidimensional pattern depends strictly on the stability of the solution against ubiquitous perturbations (including angular modes) rather than solely on the deterministic symmetry of the initial condition. It remains unclear how the current framework accounts for spontaneous symmetry breaking in cases where these angular modes become unstable, challenging the assumption that radial symmetry will strictly dictate the outcome. We note that the authors' use of noise as an initial condition does not resolve this fundamental issue. Reaction-diffusion equations inherently describe mean-field dynamics, meaning that microscopic fluctuations are continuously present in any real system, regardless of whether explicit stochastic terms are written into the equations. Ultimately, if a symmetric mean-field solution is structurally unstable to these inherent fluctuations, it simply cannot be realized in nature.

      If we understood right the reviewer is saying that there are always fluctuations everywhere and that, thus, the spike initial pattern should also have noise everywhere. In the discussion (and partially in the introduction) we have now added a discussion (in red) on how the resulting patterns possible from such spike-with-noise initial pattern can actually be understood, to a large extent, from those of the homogeneous-with-noise and spike-without noise. In that case, as the reviewer suggests, the resulting patterns do not necessarily have radial symmetry. We have kept the results on spike initial patterns without noise because they are very helpful to understand the pattern transformations possible from spike-with-noise initial patterns. Moreover, as we discuss now, the spatial fluctuations that occur everywhere can be really small compared to the concentration in the spike and, thus, the spike initial pattern without noise can be a good approximation in some cases, at least worth considering.

      (2) Theoretical limitations in the application of Linear Stability Analysis (LSA): I remain uncertain about the framework's reliance on LSA to categorize macroscopic transformations, especially those arising from large initial perturbations (spikes). In their rebuttal letter, the authors justify this by assuming the perturbation remains small over a short time interval. However, because the study aims to describe stationary, asymptotic states, applying a linear approximation that relies on transient t->0 conditions to predict long-term global stability is not fully resolved.

      We are not trying to predict long-term stability and we never intended to. We never claimed that to be the case. We are interested in stable in time spatial patterns but the long-term stability of resulting patterns is not something we intend to see from the LSA. The LSA just provides a necessary (but not sufficient) condition for non-trivial pattern transformations: any network unable to sustain the growth of small perturbations (i.e., any linearly stable network topology) cannot lead to a non-trivial pattern transformation, regardless of the nonlinear terms in the reaction term f(g). From the previous suggestions by the reviewer it could be the case that he/she thinks that since there is always noise everywhere and all the time we can never apply the LSA or that it cannot be applied along time. In fact, we do not try to applied over time (that would make no sense for our purposes). The LSA is only applicable and only informative when applied to the initial pattern (that is at time 0) to see gene network that cannot transform initial patterns into other patterns. As we discuss in the previous version and we further stress in the current one (in red in the LSA), large spikes do not invalidate this approach. Even if the spike would be large, it would only affect cells outside the spike through the diffusion of gene products from the spike. Thus, if a small enough time is considered, large spikes can be considered as small concentration perturbations outside the spike and, thus, in the the worse case scenario our whole approach may not be applicable inside the spike but outside of it (that is reasonable since spikes are by definition narrow).

      Here it is important to stress that many things that used the LSA in the original version of the article do not used it in the current version of the article. In fact, in this article we address two main questions: (i) which gene network topologies can produce non-trivial pattern transformations; and (ii) what can we say about the stationary patterns they produce. We acknowledge that a linear stability analysis alone is not enough to fully characterize the long-term behavior of the system, and this is why we only use LSA as an aid to answer question (i).

      This was not clear enough in the first version of the article but it was explicitly stated in the last version (in the Gene network classification and Linear stability analysis section). This allows us to discard many topologies, but further analysis is still required to assess whether linearly unstable network topologies can actually produce non-trivial pattern transformations.

      It is in this further study that question (ii) comes to play. Here we do not rely on LSA, but instead impose a series of requirements (R1-R5) on the reaction term f(g) that constrain the nonlinear dynamics in a biologically motivated way that prevents pathological behaviors (particularly, boundedness of solutions (R4) and monotonicity of the reaction (R5)). These requirements enable the qualitative analysis of the different unstable network topologies in order to say some things about the possible stationary patterns. This is complemented with numerical simulations and quantitative analysis of some prototypical examples in the supplementary information.

      (3) In the previous round of the review, I suggested that a biomolecular sink, such as A+B -> AB reaction, could break the approach. In their response letter, the authors defend their approach by arguing that such reactions can be accommodated by their abstract constraints (R1-R5) as long as the signs of the Jacobian elements remain invariant. However, the problem I see here is not the sign of the interactions, but the severe loss of spatial homogeneity.

      When a macroscopic initial perturbation (a "spike" of morphogen) is introduced into a domain with a strong bimolecular sink, it will inevitably cause massive local depletion of the consumed substrate near the source. Consequently, the background state of the system will rapidly evolve into a profile with macroscopic spatial gradients long before any spontaneous pattern-forming instability takes over. Mathematically, this dictates that the system no longer possesses a homogeneous steady state, and the Jacobian matrix becomes explicitly space-dependent, which should break the classical LSA approach.

      This criticism seems to be intimately related to the previous one. If we understand correctly, the argument is that in a very non-linear system, such as in the sink described, the spike will rapidly lead to a local change in the concentration of other gene products and that then the LSA is not applicable. This is true but what we care about is whether the initial pattern (that is the system at time zero) is actually stable or not. We care about it because as we explain, pattern transformation is only possible if the perturbation in the initial pattern (spike or noise) is unstable. This is simply a necessary condition (an initial pattern may be unstable to perturbation and still not produce nontrivially transformations). So whether the system will be suitable for a LSA some time after the initial pattern, as the reviewer suggests, is not something we need to know for our classification. Related to the other comments the reviewer may be concerned with whether other perturbations occurring everywhere (and all the time), that is noise, may actually affect the possible resulting patterns (that we described in the new section of the discussion).

      We want to thank the reviewer for this clarification, as we believe we had not fully understood their concern in the previous round of review. In our framework, the bimolecular sink the reviewer suggests corresponds to a three-gene-product network where A and B mutually inhibit each other and both activate AB.

      First it is important to consider that unless something else is specified an A+B→AB system with homogeneous initial pattern (with or without noise) is just not stable: the A and B gene products will decay to zero concentration and AB to a maximal concentration (that would be homogeneous over space if there is no noise). Besides the final stable state is totally stable since with no A or B left the concentration of AB cannot change. We explicitly state in the article (in the LSA section) that we apply the LSA to systems that, when unperturbed, are stable. For other systems we just wait for the system to stabilize and then ask whether pattern transformation is possible from that state if there is some perturbation, where we can now apply a LSA). So strictly speaking the system the reviewer is suggesting is outside the scope of the article and indeed unsuitable for LSA. Nevertheless, the system the reviewer proposes cannot lead to non-trivial pattern formation, as we detail below.

      The reviewer does not specify whether A, B or AB diffuse, we then consider all possibilities. We also assume that A is the gene product in the spike.

      Case in which no molecule diffuses. In this case a spike of A will simply lead to a valley of B (since A reacts with B to deplete B), and a spike of AB in the exact same location of the spike (i.e., no non-trivial pattern transformation occurs). If there is noise, each small fluctuation in the concentration of A or B would lead to a similar fluctuation in AB. Notice this case does not lead to non-trivial pattern transformations (the peaks in the initial pattern and resulting pattern are in the same places). This latter situation in fact we explain in the “Pattern formations from homogeneouswith-noise initial patterns in H networks section.”

      Case in which AB diffuses. Since nothing is promoting the production of A and B, their concentration will inevitably decay to zero and since AB diffuses its concentration on the long-term will inevitably become homogeneous (irrespectively of which initial pattern there may be).

      Case in which only A diffuses. In this case the spike of A will initially lead to a valley of B (since A consumes B) and a peak of AB (since this consumption leads to AB). However, since for each molecule of AB a molecule of both A and B are required and the concentration of B is homogeneous (since B does not diffuse), having more of A around the spike would not lead to more AB in the peak than elsewhere. The concentration of AB would thus become homogeneous over time. The same applies if B is the only molecule that diffuses.

      Case in which A and B diffuse and AB does not. In this case the spike of A will lead to a peak of AB. This would deplete B around the spike but since B can diffuse, new molecules of B would arrive and lead to a further growth in the peak of AB. As a result a stable peak of AB will form (even if A and B will ultimately decay to zero), just around the initial spike of A and both A and B will decay to zero (so no new peaks or valleys form and thus, no non-trivial pattern transformation).

    1. eLife Assessment

      This important study uses a combination of experimental and modeling approaches to investigate the role of actomyosin in epithelial invagination during Ciona siphon tube morphogenesis. Several types of convincing quantitative analyses and modeling approaches are presented that support a model in which bidirectional relocation of actomyosin drives invagination. Since epithelial invagination contributes to the morphogenesis of many developing organs, this work has the potential to appeal to both cell biologists and developmental biologists.

    2. Reviewer #2 (Public review):

      Summary:

      The authors propose that bidirectional redistribution of actomyosin drives tissue invagination in Ciona siphon tube formation. They suggest a two-stage model where actomyosin first accumulates apically to drive a slow initial invagination, followed by redistribution to lateral domains to accelerate the invagination process through cell shortening. They have shown that actomyosin activity is important for invagination - modulation of myosin activity through expression of myosin mutants altered the timing and speed of invagination; furthermore, optogenetic inhibition of myosin during the transition of the slow and fast stages disrupted invagination. The authors further developed a vertex model to validate the relationship between contractile force distribution and epithelial invagination.

      Strengths:

      (1) The authors employed various techniques to address the research question, including optogenetics, use of MRLC mutants, and vertex modelling.

      (2) The authors provide quantitative analyses for a substantial portion of their imaging data, including cell and tissue geometry parameters as well as actin and myosin distributions. The sample sizes used in these analyses appear appropriate.

      (3) The authors combined experimental measurements with computer modeling to test the proposed mechanical models, which represents a strength of the study. It provides a framework to explore the mechanical principles underlying the observed morphogenesis.

      Comments on revised version.

      The authors have adequately addressed my previous concerns regarding the optogenetic experiments, and the addition of the new modeling analysis further strengthens the study.

    3. Author response:

      The following is the authors’ response to the previous reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This is an extensively revised version of a previously submitted manuscript that, as detailed in their 20-page response to the first reviews, satisfactorily addresses the reviewers' comments. In particular, the revised manuscript makes it much clearer how this work fits into and advances the field. The added experiments strengthen the rigor of the manuscript as well. Overall, this paper is ready to go.

      We thank the reviewer for the positive evaluation and for recognizing the improvements made in our revised manuscript.

      Reviewer #2 (Public review):

      The revised manuscript has been substantially improved. The authors have addressed many of my previous concerns through the addition of new data, analyses, and discussion. The characterization of epithelial folding in the ascidian Ciona provides valuable insight into a comparatively less explored morphogenetic system, and the imaging and quantitative analyses are overall compelling. That said, a few important points remain to be addressed.

      We thank the reviewer for the positive assessment and for acknowledging the improvements made in our revised manuscript. We are also grateful for the reviewer’s continued constructive suggestions.

      One remaining issue concerns the mechanistic novelty of the actomyosin redistribution described in this study. The authors emphasize that the key novelty lies in the stepwise translocation of actomyosin from the lateral membrane to the apical domain during the initial stage (apical constriction), followed by redistribution from the apical domain back to the lateral domain during the accelerated stage (invagination). I agree that the dynamic redistribution itself is potentially interesting and may represent an underexplored aspect of epithelial morphogenesis. However, as I discussed in my previous review comments, from a mechanics perspective, the role of apical actomyosin in driving apical constriction and of lateral actomyosin in contributing to tissue folding/invagination have already been demonstrated in multiple systems, although to varying extents depending on the model. Therefore, while the current study convincingly documents a distinct spatiotemporal sequence of actomyosin localization in Ciona atrial siphon tube formation, it could be clarified further to what extent this work advances new mechanical principles underlying epithelial folding, as opposed to revealing a variation in the deployment of previously described force-generating modules.

      Importantly, I think the manuscript has the potential to provide deeper conceptual insight if the authors more explicitly consider the significance of the "redistribution" process itself. Redistribution does not only involve the appearance of actomyosin at a new membrane domain; it also necessarily involves its disappearance from the previous domain. The latter aspect has, in my view, been much less explored in the literature. For example: Is the removal of lateral actomyosin during the early phase important for efficient apical constriction? Conversely, is the reduction of apical actomyosin during the later accelerated phase important for proper invagination mechanics? These questions are particularly interesting because they address whether redistribution between domains serves an active mechanical regulatory role, rather than focusing on the role of force-generating actomyosin at a given location.

      I acknowledge that addressing these questions experimentally could be technically challenging. One potentially powerful way to address this would be through the revised computational model. For example, the authors could test whether tissue folding is altered when actomyosin is allowed to accumulate at a new domain without being concomitantly depleted from the original domain. Such analyses could help distinguish whether redistribution itself has functional mechanical importance, rather than merely reflecting sequential recruitment to different cellular regions. In my opinion, incorporating this aspect would substantially strengthen the conceptual and mechanistic novelty of the study.

      We thank the reviewer for raising this important point. We agree that the mechanical significance of actomyosin redistribution, beyond the individual roles of apical and lateral contractility, deserves further clarification. In our study, quantitative analysis of F-actin dynamics revealed a bidirectional reorganization of the actomyosin network during siphon morphogenesis: the increase of F-actin intensity in one domain was accompanied by its reduction in the other domains during both the initial and accelerated invagination stages. These observations suggest that actomyosin redistribution may represent an active mechanical regulatory process rather than merely sequential recruitment to different cellular domains. Following the reviewer’s suggestion, we used the computational model developed in this study to examine its functional significance.

      Specifically, we modified the temporal dynamics of actomyosin activity by maintaining apical contractility during the accelerated stage or by prematurely enhancing lateral contractility during the initial stage in simulations. We found that sustained apical tension during the accelerated stage primarily induced stronger central cell elongation (Figure 6—figure supplement 1A-C), whereas elevating lateral actomyosin activity during the initial stage (14–16 hpf) suppressed central cell elongation and reduced the inward movement of surrounding cells toward the central region (Fig. 6—figure supplement 1D-F), resulting in earlier bending deformation with a flatter invaginating morphology (Fig. 6—figure supplement 1F). Although both perturbations eventually converged to similar final shapes, likely because they reached similar final actomyosin distributions, their distinct morphogenetic trajectories demonstrate the mechanical importance of the temporal sequence of actomyosin redistribution. These results further clarify the significance of the previously identified apico-basal tension imbalance and lateral contraction by revealing how their sequential activation coordinates tissue deformation. The initial dominance of apical contractility, together with limited lateral contraction, promotes cell elongation and convergence of the active region, whereas the subsequent shift toward lateral contractility facilitates cell shortening and deep tissue invagination.

      Thus, our results highlight that the bidirectional redistribution of actomyosin is not merely a consequence of morphogenesis, but contributes to the dynamic regulation of epithelial folding. We have incorporated these new simulation results and analyses into the revised manuscript as Figure 6—figure supplement 1.

      My other concern relates to the new optogenetic data presented in Figure 4-figure supplement 2. In the "Dark" samples, active myosin does not appear to be clearly enriched along the membrane, but instead seems relatively diffuse within the cytoplasm. This appears distinct from the images shown in Figure 2, where active myosin exhibits clear membrane enrichment. Could the authors provide top-view images for the samples shown in Figure 4-figure supplement 2? This would help clarify whether active myosin is indeed enriched along the apical membrane at 16 hpf and along the lateral membrane at 17 hpf in the "Dark" condition.

      We thank the reviewer for this careful observation. We agree that the p-MLC signal in the "Dark" samples of the original Figure 4—figure supplement 2 appeared less membrane-enriched than in Figure 2. This was due to a technical adjustment: because the optogenetic system occupied the 568 nm channel, we have to switch the p-MLC signal from Alexa Fluor 568 anti-rabbit IgG to Alexa Fluor 647 anti-rabbit IgG, which yielded relatively weaker membrane signal under our imaging conditions. To address this, we have replaced the cross-sectional images with better-representative examples and added top-view images (Figure 4—figure supplement 2). These new panels clearly showed that active myosin was enriched at apical junctions (16 hpf) and lateral membranes (17 hpf) in the Dark condition, consistent with that in Figure 2. Since the Dark and Light groups were processed identically, the relative comparison remains valid.

      In addition, the tissue morphology in the "17 hpf Light 1 hr" panel of Figure 4-figure supplement 2 appears noticeably different from that shown in Figure 4. Specifically, the apical side of the tissue in Figure 4 appears substantially more relaxed than in Figure 4-figure supplement 2. Based on the authors' interpretation of the optogenetic experiments, apical active myosin is not strongly affected by the treatment described in Figure 4. If so, one would expect apical constriction to remain largely intact. However, the more relaxed apical domain shown in Figure 4 seems to suggest that apical constriction may in fact be perturbed by the optogenetic manipulation. This apparent discrepancy complicates the interpretation of the experiment and seems somewhat inconsistent with the authors' main conclusion from this figure.

      We thank the reviewer for this very careful and insightful observation. We agree that the tissue morphology in the "17 hpf Light 1 hr" panel of Figure 4—figure supplement 2 appears less relaxed than that in Figure 4B.

      We acknowledge that optogenetic inhibition of myosin activity might indeed have a partial effect on activity of apical myosin, which can lead to apical relaxation after the initial constriction, as shown in the representative embryo in Figure 4B. However, because Ciona embryos are not always perfectly synchronized at 16 hpf (the time point when light illumination was initiated), individual embryos exhibit slight variations in the extent of apical constriction that has already been achieved during the initial stage (13.5–16.0 hpf). For embryos that had completed a relatively stronger apical constriction by 16 hpf, the apical domain can maintain its constricted morphology even after light exposure (Figure 4—figure supplement 2). Embryos with relatively weaker apical constriction at the time of light onset are more prone to exhibit apical relaxation upon optogenetic manipulation, as illustrated in Figure 4B. This developmental heterogeneity is the primary reason for the morphological variability observed between individual embryos in the optogenetic groups. Importantly, despite this morphological variability, the quantitative comparison of apical p-MLC intensity between the Light and Dark groups in Figure 4—figure supplement 2B showed no statistically significant difference (t-test, ns), which is consistent with the fact that during normal development, apical myosin activity naturally declines after 16 hpf (Figure 2B). In contrast, lateral p-MLC intensity was significantly reduced in the Light group compared to the Dark control (Figure 4—figure supplement 2B). This reduction in lateral contractility is the key factor responsible for the blockade of invagination progression.

      We have revised the statements accordingly. Hopefully, these clarifications have adequately addressed the reviewer's concern.

      Reviewer #3 (Public review):

      Concerns raised in the initial submission were addressed in the revised manuscript.

      We thank the reviewer for the encouraging feedback and for acknowledging our revisions.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further revisions suggested.

      We thank the reviewer again for the positive assessment.

      Reviewer #2 (Recommendations for the authors):

      Here are several additional comments and suggestions in addition to the concerns described in the public review.

      Line 86 - 88: the authors state "However, this transition is from apical to basolateral (Sherrard et al., 2010), rather than a bidirectional redistribution between apical and lateral domains." This statement feels somewhat out of context because the concept of "bidirectional redistribution" has not yet been introduced. It may fit more naturally in the Discussion section, after the relevant observations and interpretations have been fully presented.

      We thank the reviewer for this suggestion. We have revised the sentence in the Introduction to avoid introducing the concept of "bidirectional redistribution" prematurely.

      Line 92 - 95: the authors raised the questions of "Whether a bidirectional redistribution of actomyosin between apical and lateral domains operates as a core mechanism for sequential invagination, and whether lateral contractility is essential for the accelerated phase, remain unclear." These questions also feel somewhat out of context, for the same reason mentioned above.

      We thank the reviewer for this suggestion. We have revised the text to avoid prematurely introducing the concept of "bidirectional redistribution" in the Introduction.

      Line 162 - 163: The authors state that "This redistribution pattern was consistent with that of F-actin in the corresponding phases." This conclusion should be revised, as the reported increase in apical F-actin and reduction in lateral F-actin during the initial stage do not appear to reach statistical significance, which is different from that of active myosin.

      We thank the reviewer for this careful observation. We agree that the F-actin changes during the initial stage did not reach statistical significance, unlike the active myosin data. We have revised the sentence to state that the myosin redistribution showed a similar trend to F-actin, while acknowledging the lack of statistical significance for F-actin.

      Line 349 - 350: "and the invagination speed (represented by slopes of curves in Figure 5A) gradually slows down at later stages." It seems that Figure 5A should be Figure 5B.

      Thank you for pointing out this typo. We have corrected it in the revised manuscript (now Figure 5C).

      Reviewer #3 (Recommendations for the authors):

      We appreciate the efforts made by the authors to address the questions and comments raised in the initial submission. The only remaining concerns are regarding grammar and spelling and a couple of minor errors.

      (1) Lines 346-347, "These trends are consistent with the experimental mutant data (Figure 5A, B)."

      This was confusing - was this meant to be Figure 3A, B?

      We sincerely apologize for this confusion. We have corrected it in the revised manuscript (now Figure 3B).

      (2) Lines 349-350, "... the invagination speed (represented by the slopes of curves in Figure 5A) gradually slows down at later stages..."

      Similar to above, was this meant to be Figure 5C?

      Thank you for pointing out this typo. We have corrected it in the revised manuscript (now Figure 5C).

      (3) There is a typo in Line 336 ("dimmish") and some minor grammatical concerns in the introduction.

      We have corrected the typo "dimmish" to "diminish". We have also carefully proofread the entire manuscript and fixed any remaining grammatical issues.

    1. eLife Assessment

      The main contribution of this potentially useful manuscript is a standardized comparative MRI resource covering multiple avian clades, complemented by histological cross-validation and diffusion-based connectivity analyses. Such a resource could significantly enhance accessibility to internal avian brain anatomy for evolutionary neurobiology. Its central claim that avian brain compartments exhibit modular diversification aligns with the existing literature. However, uncertainties regarding data availability, concerns with the tractography pipeline, and insufficient recognition of prior work lead to the conclusion that the methods, data, and analyses are incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      The study presents a novel analysis of MRI resources for 16 avian species, spanning major (though not all) clades and ecological niches. This is a significant step towards large-scale datasets on internal parcellation and long-range connectivity, central to evolutionary studies for understanding the evolution of the bird brain.

      Strengths:

      The integration of high-resolution T2-weighted and diffusion-weighted MRI with histological validation (Nissl and Luxol Fast Blue staining) provides a strong, cross-validated framework for studying avian brain anatomy. Data on long-range connectivity are particularly useful for understanding how relationships between brain components evolved. The approach is also scalable, allowing for more detailed evolutionary analyses compared to what is currently possible.

      Weaknesses:

      The sampling supports evidence of modular evolution in the bird brain, but it is limited for broad evolutionary claims, as the effects of sizes and phylogenies can be hard to disentangle without enough species per clade.

      Tractography-based claims should be treated cautiously without sensitivity analyses. This is particularly important when comparing brains with different sizes and tissue properties.

      Existing literature is not acknowledged sufficiently. This makes some claims of novelty misleading, and prevents readers from understanding the current state of knowledge in this research area.

    3. Reviewer #2 (Public review):

      This manuscript presents a comparative MRI dataset from 16 avian species and uses MRI and tractography to examine variation in brain organization across birds. The authors argue that their analyses support mosaic brain evolution and provide a framework for comparative neuroanatomy. Although the dataset represents a useful resource, particularly given the inclusion of understudied species like penguins, toucans, and hornbills, I have substantial concerns regarding the novelty of the study, the anatomical interpretation of the results, and the validity of the tractography analyses. In its current form, I do not believe the manuscript provides sufficient new biological insight to support many of its conclusions.

      Major Concerns

      (1) The authors repeatedly state that comparative neuroanatomical studies in birds have largely been unable to examine internal brain organization or "internal parcellation". This claim is inaccurate and reflects limited engagement with a substantial body of literature. For decades, comparative studies have examined variation in the size of major avian brain subdivisions as well as specific sensory, motor, and associative nuclei. For example, the extensive work of Andrew Iwaniuk and colleagues has documented variation in numerous brain regions across birds and related these differences to ecology, behavior, and sensory specialization (e.g., Gutierrez-Ibanez et al., 2009; Iwaniuk et al., 2006, 2008, 2010; Corfield et al., 2015). Other authors have also made important contributions in this area (e.g., Boire and Baron, 1994; Burish et al., 2004; Moore and DeVoogd, 2011, 2017). Importantly, previous work has already examined variation in major subdivisions of the avian brain using relatively standardized datasets (e.g., Iwaniuk et al., 2004; Iwaniuk and Hurd, 2005), including datasets that contain more species and greater taxonomic diversity than the current study. In other words, these studies have already provided detailed analyses of internal brain organization across broad taxonomic samples.

      The manuscript should therefore be reframed as providing a new MRI-based resource rather than introducing the first comparative framework for studying internal avian brain organization. The current framing significantly overstates the novelty of the work.

      (2) A second significant concern is the lack of anatomical specificity in the tractography analyses. The authors repeatedly refer to regions such as "anterior cortex," "dorsal cortex," and "temporal cortex." These terms are not standard anatomical designations in avian neuroanatomy and provide little information about the actual structures being analyzed.

      For example, the "temporal cortex" could potentially include portions of the nidopallium (including the caudolateral nidopallium, NCL), mesopallium, and arcopallium. Similarly, the "anterior cortex" could correspond to the somatosensory or visual Wulst, the anterior nidopallium, or several other structures. The designation "dorsal cortex" is similarly difficult to interpret. Because these seed regions may encompass multiple functionally distinct systems, it is impossible to evaluate the biological significance of the reported connectivity patterns.

      I strongly encourage the authors to define their seed regions using accepted avian neuroanatomical terminology and to provide detailed anatomical maps. More informative analyses would focus on well-defined structures with known connectivity, such as the Wulst, arcopallium, entopallium, or NCL. As currently presented, the tractography results are too coarse to support meaningful biological conclusions.

      (3) I am not an MRI specialist, but I have concerns regarding the interpretation of the tractography results. Bird brains are small, and diffusion MRI tractography is already known to be challenging even in substantially larger brains. The manuscript provides limited information regarding image resolution, diffusion sampling, and the expected accuracy of tract reconstruction in these specimens. More importantly, there is little validation of the tractography results. Diffusion tractography is prone to both false positives and false negatives, and reconstructed pathways cannot be assumed to represent true anatomical connections.

      The authors should provide evidence that their tractography pipeline can accurately recover known pathways. For example, they could compare reconstructed tracts with well-established anatomical pathways such as the anterior commissure or major visual pathways, which would substantially strengthen confidence in the results. Without such validation, it is difficult to determine whether the observed species differences reflect biological variation or methodological artifacts.

      (4) I also have some methodological concerns regarding the comparisons of anterior commissure (AC) size and cerebellar foliation. First, the authors measure the AC in a coronal section. I would recommend measuring the AC area in a midsagittal section instead. Furthermore, the authors use the cross-sectional area of the same coronal section as the scaling variable. This seems problematic because the area of any given section will depend on the angle of sectioning and other technical factors. If the objective is to compare the relative size of the AC, then total brain volume or telencephalon volume would be more appropriate scaling variables.

      With respect to cerebellar foliation, the authors developed their own metric. I would encourage them to use methods already established in the literature, such as the foliation index described by Iwaniuk et al. (2006). Their approach may yield similar results, but using the foliation index would facilitate direct comparisons with existing datasets and would allow incorporation of additional published data (e.g., Cunha et al., 2021, which includes foliation index measurements for 54 bird species). The authors should also be aware that the foliation index scales with body size. Consequently, the high degree of foliation observed in penguins may not necessarily indicate cerebellar expansion or increased demands for sensorimotor integration associated with their specialized locomotion. I therefore believe that the conclusions regarding variation in AC size and cerebellar foliation should be re-evaluated after more appropriate analyses are performed.

    4. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study presents a novel analysis of MRI resources for 16 avian species, spanning major (though not all) clades and ecological niches. This is a significant step towards large-scale datasets on internal parcellation and long-range connectivity, central to evolutionary studies for understanding the evolution of the bird brain.

      Strengths:

      The integration of high-resolution T2-weighted and diffusion-weighted MRI with histological validation (Nissl and Luxol Fast Blue staining) provides a strong, cross-validated framework for studying avian brain anatomy. Data on long-range connectivity are particularly useful for understanding how relationships between brain components evolved. The approach is also scalable, allowing for more detailed evolutionary analyses compared to what is currently possible.

      Weaknesses:

      The sampling supports evidence of modular evolution in the bird brain, but it is limited for broad evolutionary claims, as the effects of sizes and phylogenies can be hard to disentangle without enough species per clade.

      We appreciate the reviewer highlighting this limitation and agree that it warrants explicit consideration. Although the 16 species included in the current study encompass diverse avian lineages and ecological characteristics, our sampling was not designed to rigorously disentangle the effects of phylogeny and brain size within individual clades.

      To improve the taxonomic coverage of our dataset, we plan to expand the MRI dataset in the revised manuscript by including additional specimens that are currently available to us. Specifically, we will acquire and analyze MRI data from the slaty-backed gull (Larus schistisagus), black-tailed gull (Larus crassirostris), budgerigar (Melopsittacus undulatus), and crow (Corvus sp.). In addition, we plan to incorporate publicly available MRI data from the ostrich (Struthio camelus). These additions will broaden the phylogenetic and anatomical coverage of the comparative dataset.

      Nevertheless, we fully acknowledge the reviewer’s point that, even with these additional species, the number of species sampled within individual clades will remain insufficient to rigorously separate phylogenetic effects from size-related effects or to support broad generalizations across all avian lineages. We will therefore explicitly state this limitation in the revised manuscript and temper evolutionary claims that extend beyond what can be supported by the present taxonomic sampling.

      Accordingly, we will more clearly position the primary contribution of this study as the establishment of an MRI-based comparative resource and analytical framework applicable across diverse avian species, rather than as a comprehensive phylogenetic test of avian brain evolution. Within this scope, we will present the observed variation among brain compartments as evidence consistent with mosaic/modular diversification, while carefully limiting the broader evolutionary interpretation of these patterns.

      Tractography-based claims should be treated cautiously without sensitivity analyses. This is particularly important when comparing brains with different sizes and tissue properties.

      The reviewer raises an important point regarding the interpretation of tractography-based comparisons. We agree that particular caution is required when comparing datasets across species that differ in brain size, tissue properties, and imaging characteristics.

      In the revised manuscript, we will perform additional sensitivity analyses to evaluate the robustness of the major tractography-derived patterns. Specifically, we will examine the effects of varying the FA threshold used for tract reconstruction, rather than relying solely on the current threshold of 0.1, and assess whether the major reconstructed trajectory patterns and cross-species differences are robust to this parameter.

      We will also re-analyze the available diffusion data using Generalized Q-Sampling Imaging (GQI) as an alternative reconstruction approach and compare the resulting trajectory patterns with those obtained using the current DTI-based analysis. We recognize that the ability to resolve complex fiber configurations depends on the underlying diffusion acquisition parameters, including b-value and spatial and angular resolution. We will therefore interpret these comparisons within the limitations of the currently available datasets, without implying that either reconstruction approach provides a definitive representation of the underlying fiber architecture.

      In addition, we will provide a more detailed description of the relevant diffusion MRI acquisition parameters and spatial and angular resolution of the datasets so that potential technical differences among species can be more clearly evaluated. We will also discuss how such differences may affect cross-species tractography comparisons.

      Finally, we will revise the interpretation of the tractography results throughout the manuscript to more clearly distinguish diffusion MRI-derived reconstructed trajectories from anatomically demonstrated neuronal connections. We will avoid treating reconstructed streamlines as direct evidence of anatomical connectivity and will explicitly discuss the potential for false-positive and false-negative tract reconstruction, as well as other limitations inherent to diffusion MRI tractography.

      Existing literature is not acknowledged sufficiently. This makes some claims of novelty misleading, and prevents readers from understanding the current state of knowledge in this research area.

      We fully agree with this assessment. Although substantial comparative neuroanatomical work has established important principles of avian brain evolution, the current manuscript does not sufficiently cite or incorporate this literature into the discussion. As a result, some statements regarding the novelty of the present study are overstated.

      In the revised manuscript, we will substantially expand our discussion and citation of the relevant literature, including the extensive comparative studies of individual avian brain regions and brain subdivisions highlighted by the reviewers. We will revise the Introduction and Discussion accordingly to more accurately describe the current state of knowledge in comparative avian neuroanatomy and to clearly distinguish the contributions of previous studies from those of the present work.

      We will also carefully revise statements regarding the novelty of our study. Rather than implying that our study provides the first comparative framework for examining internal avian brain organization, we will more appropriately emphasize its contribution as an MRI-based comparative resource that enables standardized visualization and analysis of internal brain anatomy and connectivity across diverse avian species.

      Reviewer #2 (Public review):

      This manuscript presents a comparative MRI dataset from 16 avian species and uses MRI and tractography to examine variation in brain organization across birds. The authors argue that their analyses support mosaic brain evolution and provide a framework for comparative neuroanatomy. Although the dataset represents a useful resource, particularly given the inclusion of understudied species like penguins, toucans, and hornbills, I have substantial concerns regarding the novelty of the study, the anatomical interpretation of the results, and the validity of the tractography analyses. In its current form, I do not believe the manuscript provides sufficient new biological insight to support many of its conclusions.

      Major Concerns

      (1) The authors repeatedly state that comparative neuroanatomical studies in birds have largely been unable to examine internal brain organization or "internal parcellation". This claim is inaccurate and reflects limited engagement with a substantial body of literature. For decades, comparative studies have examined variation in the size of major avian brain subdivisions as well as specific sensory, motor, and associative nuclei. For example, the extensive work of Andrew Iwaniuk and colleagues has documented variation in numerous brain regions across birds and related these differences to ecology, behavior, and sensory specialization (e.g., Gutierrez-Ibanez et al., 2009; Iwaniuk et al., 2006, 2008, 2010; Corfield et al., 2015). Other authors have also made important contributions in this area (e.g., Boire and Baron, 1994; Burish et al., 2004; Moore and DeVoogd, 2011, 2017). Importantly, previous work has already examined variation in major subdivisions of the avian brain using relatively standardized datasets (e.g., Iwaniuk et al., 2004; Iwaniuk and Hurd, 2005), including datasets that contain more species and greater taxonomic diversity than the current study. In other words, these studies have already provided detailed analyses of internal brain organization across broad taxonomic samples.

      The manuscript should therefore be reframed as providing a new MRI-based resource rather than introducing the first comparative framework for studying internal avian brain organization. The current framing significantly overstates the novelty of the work.

      We appreciate the reviewer drawing attention to this issue. Although a substantial body of comparative neuroanatomical work has already examined internal brain organization in birds, the current manuscript does not sufficiently cite or incorporate this literature into the discussion. As a consequence, some statements regarding the novelty of our study are overstated.

      In the revised manuscript, we will expand our discussion and citation of previous work, including the studies highlighted by the reviewer on interspecific variation in major avian brain subdivisions, sensory and motor nuclei, and other anatomically defined brain regions. We will revise the Introduction and Discussion to more accurately represent the existing body of comparative avian neuroanatomy and to clarify how the present study complements and extends these established approaches.

      Most importantly, we will reframe the manuscript so that its primary contribution is presented as the establishment of an MRI-based comparative resource for visualizing and quantitatively analyzing internal brain anatomy across diverse avian species. We will remove or revise statements implying that the present study provides the first comparative framework for examining internal avian brain organization. Instead, we will emphasize the complementary advantages of the MRI-based approach, particularly its ability to provide non-destructive three-dimensional visualization of internal brain structures in intact specimens and to enable comparison of multiple brain compartments within a common analytical framework across diverse avian species.

      We believe that this revised framing will more accurately position the contribution of our study within the existing literature and clarify the specific value of the dataset.

      (2) A second significant concern is the lack of anatomical specificity in the tractography analyses. The authors repeatedly refer to regions such as "anterior cortex," "dorsal cortex," and "temporal cortex." These terms are not standard anatomical designations in avian neuroanatomy and provide little information about the actual structures being analyzed.

      For example, the "temporal cortex" could potentially include portions of the nidopallium (including the caudolateral nidopallium, NCL), mesopallium, and arcopallium. Similarly, the "anterior cortex" could correspond to the somatosensory or visual Wulst, the anterior nidopallium, or several other structures. The designation "dorsal cortex" is similarly difficult to interpret. Because these seed regions may encompass multiple functionally distinct systems, it is impossible to evaluate the biological significance of the reported connectivity patterns.

      I strongly encourage the authors to define their seed regions using accepted avian neuroanatomical terminology and to provide detailed anatomical maps. More informative analyses would focus on well-defined structures with known connectivity, such as the Wulst, arcopallium, entopallium, or NCL. As currently presented, the tractography results are too coarse to support meaningful biological conclusions.

      This is an important concern, and we agree that the anatomical nomenclature and definition of the regions used in the tractography analyses require substantial improvement.

      In the revised manuscript, we will carefully re-evaluate the anatomical description of each region with reference to established avian neuroanatomical terminology, anatomical atlases, and available histological information. Where the analyzed regions can be reliably assigned to established anatomical structures, we will replace broad mammalian-style positional terminology such as “anterior cortex,” “dorsal cortex,” and “temporal cortex” with more appropriate avian neuroanatomical terminology. For example, we will re-examine whether the region currently referred to as the “optic lobe” can be more precisely defined as the optic tectum, while the cerebellum can be retained as an anatomically well-defined region.

      At the same time, we recognize an important limitation of the present tractography analysis. The spatial resolution and analytical framework of the present comparative datasets do not necessarily permit reliable assignment of all analyzed regions or reconstructed trajectory patterns to fine pallial subdivisions such as the entopallium, arcopallium, or NCL across species. We therefore do not intend to assign such specific anatomical identities where they cannot be supported with sufficient confidence.

      Instead, for regions that cannot be unambiguously assigned to a single established anatomical subdivision, we will define their location and extent using reproducible anatomical landmarks and clearly indicate the level of anatomical resolution supported by the data. We will also provide revised anatomical maps showing the locations of the analyzed regions and their relationship to major avian brain subdivisions. This will allow readers to evaluate more clearly which anatomical structures may contribute to the reconstructed trajectory patterns.

      Accordingly, we will revise the biological interpretation of the tractography results to match this anatomical resolution. Rather than attributing the reconstructed patterns to specific fine-scale pallial structures or functional systems when these cannot be reliably distinguished, we will interpret them more conservatively as broad patterns of fibre organisation among anatomically defined brain regions. These revisions will improve the anatomical transparency of the analysis while avoiding anatomical or functional interpretations that exceed the resolution of the present datasets.

      (3) I am not an MRI specialist, but I have concerns regarding the interpretation of the tractography results. Bird brains are small, and diffusion MRI tractography is already known to be challenging even in substantially larger brains. The manuscript provides limited information regarding image resolution, diffusion sampling, and the expected accuracy of tract reconstruction in these specimens. More importantly, there is little validation of the tractography results. Diffusion tractography is prone to both false positives and false negatives, and reconstructed pathways cannot be assumed to represent true anatomical connections.

      The authors should provide evidence that their tractography pipeline can accurately recover known pathways. For example, they could compare reconstructed tracts with well-established anatomical pathways such as the anterior commissure or major visual pathways, which would substantially strengthen confidence in the results. Without such validation, it is difficult to determine whether the observed species differences reflect biological variation or methodological artifacts.

      We agree with the reviewer that the tractography results require cautious interpretation and that further evaluation of the robustness and anatomical plausibility of the tractography pipeline would strengthen the study.

      In the revised manuscript, we will provide a more detailed description of the diffusion MRI acquisition parameters, spatial and angular resolution, diffusion reconstruction, and tractography procedures so that the methodological limitations of the analyses can be more clearly assessed.

      Recent studies using DTI under imaging conditions comparable to those employed in the present study have demonstrated successful reconstruction of fiber trajectories in the mouse brain, which is smaller than the avian brains examined here (Janz et al., eLife, 2017). We therefore consider that the size of the avian brain itself does not preclude DTI-based fiber reconstruction and that such analyses are technically feasible at the brain sizes examined in the present study.

      Importantly, we will perform additional sensitivity analyses to evaluate the robustness of the reconstructed trajectory patterns. Specifically, we will examine the effects of varying the FA threshold used for tract reconstruction. We will also re-analyse the available diffusion data using Generalized Q-Sampling Imaging (GQI) as an alternative reconstruction approach and compare the resulting trajectory patterns with those obtained using the current DTI-based analysis. We recognize that the ability to resolve complex fiber configurations depends on the diffusion acquisition parameters, including b-value and spatial and angular resolution. We will therefore interpret the comparison between reconstruction approaches within the limitations of the currently available datasets, without implying that either approach provides a definitive reconstruction of the underlying fiber architecture.

      We will also evaluate the robustness of the k-means clustering used to summarize tractography patterns. Because the current choice of k = 10 was not based on an independently established biological criterion, we will examine alternative values of k and assess whether the major trajectory patterns and cross-species differences are robust to the choice of cluster number.

      In addition, we will examine anatomically well-characterized pathways, including major commissural and visual pathways, and assess whether the reconstructed trajectories are consistent with known avian neuroanatomy. We will use these comparisons as an assessment of anatomical plausibility rather than as definitive validation of tractography accuracy.

      Finally, we will revise the manuscript to more clearly distinguish diffusion MRI-derived reconstructed trajectories from anatomically demonstrated neuronal connections. We will avoid interpreting reconstructed streamlines as direct evidence of anatomical connectivity and will explicitly discuss the possibility of false-positive and false-negative tract reconstruction, as well as other limitations inherent to diffusion MRI tractography.

      (4) I also have some methodological concerns regarding the comparisons of anterior commissure (AC) size and cerebellar foliation. First, the authors measure the AC in a coronal section. I would recommend measuring the AC area in a midsagittal section instead. Furthermore, the authors use the cross-sectional area of the same coronal section as the scaling variable. This seems problematic because the area of any given section will depend on the angle of sectioning and other technical factors. If the objective is to compare the relative size of the AC, then total brain volume or telencephalon volume would be more appropriate scaling variables.

      With respect to cerebellar foliation, the authors developed their own metric. I would encourage them to use methods already established in the literature, such as the foliation index described by Iwaniuk et al. (2006). Their approach may yield similar results, but using the foliation index would facilitate direct comparisons with existing datasets and would allow incorporation of additional published data (e.g., Cunha et al., 2021, which includes foliation index measurements for 54 bird species). The authors should also be aware that the foliation index scales with body size. Consequently, the high degree of foliation observed in penguins may not necessarily indicate cerebellar expansion or increased demands for sensorimotor integration associated with their specialized locomotion. I therefore believe that the conclusions regarding variation in AC size and cerebellar foliation should be re-evaluated after more appropriate analyses are performed.

      These methodological suggestions are very helpful. We agree that both the anterior commissure analysis and the assessment of cerebellar foliation should be re-evaluated to enable more anatomically and quantitatively appropriate comparisons across species.

      Anterior commissure:

      We agree that the current analysis, in which AC area was measured from the coronal section showing the largest cross-sectional profile and normalized to whole-brain area in the same section, may be influenced by differences in brain geometry and section orientation among species. In the revised manuscript, we will re-evaluate the method used to quantify AC size, including measurements from sagittal or midsagittal views where anatomically appropriate. We will also examine normalization against volumetric measures, such as total brain or telencephalic volume, rather than relying solely on a single-section whole-brain area. The corresponding results and interpretations will be revised accordingly.

      Cerebellar foliation:

      We also agree that our current branch-counting approach should be considered in relation to established quantitative measures of avian cerebellar foliation. In the revised manuscript, we will evaluate whether the cerebellar foliation index described by Iwaniuk et al. (2006), or a comparable standardized measure applicable to our midsagittal MRI datasets, can be used to re-analyze cerebellar foliation. This will also allow us to place our observations more directly in the context of the larger comparative datasets reported previously, including that of Cunha et al. (2021).

      Importantly, we recognize that cerebellar foliation is strongly influenced by allometric and phylogenetic factors. We will therefore re-evaluate the interpretation of interspecific differences in cerebellar foliation in light of brain/body size relationships and the existing comparative literature. In particular, we will revise the interpretation of the pronounced cerebellar foliation observed in penguins and other species and avoid attributing these differences directly to locomotor or sensorimotor specialization unless supported by the revised analyses.

    1. eLife Assessment

      This work has much novelty, supportive data, and the potential for changing thinking about melanoma biology with the discovery of MRGPRX4 as a relatively melanoma-specific G protein-coupled receptor. MRGPRX4 is normally restricted to peripheral sensory neurons, but it is aberrantly expressed in human melanomas; melanomas with neural-crest-like/invasive features were most likely affected. When ectopically expressed in murine melanocytes, MRGPRX4 drives melanoma development with metastasis; the mechanism appears to involve ligand-independent receptor signaling. Overall, the text presents this as a novel model of oncogenesis driven by re-expression of a sensory GPCR - this is an important study with compelling data.