10,000 Matching Annotations
  1. Aug 2026
    1. Reviewer #2 (Public review):

      Summary:

      Using the 5xFAD model in combination with GPR34 mice, the authors explore the function of microglia in the context of neurodegeneration. Using a broad spectrum of methodology, they show that DAM signatures are increased in KO 5xFAD mice. Using several KO clones of GPR34 KO iMGLs and another set of broad methodologies, the authors show that GPR34 is important for microglia homeostasis,<br /> phagocytosis, specifically of myelin. GPR34 KO iMGLs also show a distinct transcriptional response to myelin. Together, they propose that GPR34 limits microglial activation in neurodegeneration.

      Strengths:

      All methods are state-of-the-art, and the combination of mouse and human microglia responses is a particular strength.

      Weaknesses:

      No weaknesses were identified by this reviewer.

    2. Author response:

      We greatly thank the Reviewing Editor, Senior Editor and the reviewers for their constructive and thoughtful feedback, as well as for recognizing the significance and strengths of our study. We are very encouraged by overall positive assessment and appreciate these insights to strengthen the manuscript. Below, we outline our plans to address the key points raised in reviewer#1’s public review. We also note that reviewer#2 did not identify any weaknesses in the study and are thankful for this positive evaluation.

      Point 1: Relating in vivo and in vitro findings and discussing the relevance of myelin clearance in AD. As suggested, we will elaborate our discussion to cover the relationship between our in vivo and in vitro findings and present a more unified picture of GPR34 function. We will also highlight the relevance of myelin clearance in Alzheimer’s disease independent of amyloid plaque burden.

      Point 2: in vivo assessments and phenotypes. While we recognize the value of expanding our in vitro and cellular findings to in vivo and cognitive measures, we believe this additional assessment is beyond the scope of this current study. At least, we will expand our discussion to relate our findings of GPR34-medilated myelin pathology in the contexts of neurodegeneration and cognitive deficiency in AD.

      Point 3: Myelin engulfment versus lysosomal degradation. We agree that the clear distinction between impaired myelin engulfment and altered lysosomal degradation is an important point. Although additional experiments to address these scenarios are beyond the scope of the study, we will perform targeted pathway analysis of existing transcriptomic datasets focusing on the phagosome and lysosomal degradation pathways along with the in vivo findings related to CD68, which could offer an interpretation.

      Point 4: Comparison of transcriptional signatures. As suggested, we will compare the transcriptional signatures induced by myelin exposure in WT iMGLs with our 5xFAD RNA-seq datasets, as well as publicly available datasets from human Alzheimer’s disease microglia, to unravel the potential converged and distinct features.

      Point 5: Reconciling our findings with previous GPR34 studies. We will expand our discussion to compare and clarify our current findings with the previous studies and cover potential reasons for those different observations. We will also cover the current limitations of existing GPR34 pharmacological tools compounds, including their limited blood-brain permeability, which precludes the effective use of those tools for proposed in vivo studies.

    1. eLife Assessment

      This is a fundamental study of individual variation and the contribution of learning to behavioural individuality. The experimental design of massively parallel behavioural phenotypes is outstanding and the conclusions are supported by a compelling and rigorous analysis across a large number of experiments in thousands of individuals across genotypes and conditions. The dataset further represents an advance in studying visual associative learning thanks to the ability to make longitudinal measurements of many behavioural decisions within the same animals. These results are a major contribution to the understanding of the sources of behavioural individuality.

    2. Reviewer #1 (Public review):

      "Learning is a fundamental source of individuality," by Manna and colleagues, interrogates different sources of variation in individual behavior. The authors place individual flies in a Y-shaped arena, which is a common design in the field, and illuminate the arms of the Y with blue versus green light. They track the color preference of individual animals and also perform operant conditioning, meaning that they teach the fly to avoid a particular color/arm by generating a foot shock when the fly enters that arm. There are a number of things that are impressive about this setup: The authors are able to collect data on thousands of individual flies of many different strain backgrounds, and they demonstrate a strong change in color preference after conditioning. This is nice, because in past papers visual learning ability has been modest and difficult to study. To put a number on it, in this paper animals on average don't show a color preference at the start of the assay, spending around 30% of their time in the one arm illuminated green, and the remaining time in the two arms illuminated blue. After conditioning, the average animal spends only 23% of its time in the green arm.

      The authors run 64 animals through the assay for each of 88 wild type strains (maybe? see Major Point 1 below) and see considerable strain-specific (genetic) variation in the change in time spent in the shocked color after conditioning. Some strains show no learning, while others spend <10% of their time in the shocked color after conditioning. They also, I believe, see that some strains have more variability across individuals, which would suggest that some strains have stronger canalization at the development or circuit function level than others-i.e. some genotypes produce more consistent copies of the individual, others less consistent copies. (Or, some genotypes produce robust circuits, and others produce noisy circuits.)

      Finally, the authors argue statistically that learning itself increases variability in individual performance. This makes a lot of sense to me intuitively. Learning changes the physical/chemical properties of circuits in the brain, and because it evolves over time and interacts with environmental variables, it seems like it should send different animals down different channels. Or, at a conceptual level, if I learn to play the piano and my sister doesn't (because of some genetic difference between us or something stochastic), this learning experience will cause all sorts of other differences in our behavior as time passes. I also think the authors do have enough data to be able to make this finding. However, the presentation of the argument in this portion of the paper is hard for me to understand, and I am not an expert in statistics, so the strength of the result is difficult for me to evaluate.

      Major points:

      (1) It's difficult to track through the paper the number of animals tested for different assays. At the beginning, it says N=5632, which works out to 64 flies for each of the 88 DGRP strains. 64 happens to be the number of parallel Y arenas they have. Later in the methods, there's description of more variation within the set of 64 for each strain-two different parent sets per strain, different sexes, conditioned and un-conditioned. And, while the results text focuses on the color learning, the methods discuss additional assays (place learning, multi-day learning).

      Given the numbers, does each run of the 64 mazes include all the tested flies of one strain, or are flies of many strains included in each batch? Do different flies do different assays (color, place, multi-day) or do they all do all the assays? Perhaps there is a table including this information already in the supplement, but I recommend making it much clearer in the main results text and methods. While the dataset is large, if it is split over many conditions and/or if batch and genotype confound each other, this will affect the robustness of the results and how strong the conclusions can be.

      (2) The data presentation in Figure 1 is elegant and easy to follow, but getting into Figure 2 and subsequently, I get lost in the statistics and have trouble understanding what is being measured. My understanding of the big picture is that while genetics and individual randomness contribute a lot to behavior, the evidence for learning as an amplifier of individuality is that variance in behavior among animals of the same strain increases over time in the conditioned group (i.e. the group that is doing the most learning, or a specific kind of learning), but not in the control group. This idea is illustrated in the flattening distributions in the cartoons in Figure 1A. The authors should include graphs of the real data that use the same format as in that cartoon. Instead, the graphs present "residuals," and I don't know what those are. I suspect it's "variation left over after accounting for effects of strain and individual stochasticity." I see the residuals being tracked per strain over time in Figure 2H, but I don't see the change over time in other graphs. I'm looking for something simple like, "variation within the strain at the beginning of learning and at later time points in learning." (But I'm not sure exactly what instantaneous measurement would be the focus in longitudinal analyses of learning behavior.)

      (3) Figure 3 is a cool stab at tracking down the precise mechanism by which stochastic environment interacts with learning to send individuals along different behavioral routes. But again, like in Figure 2, I don't have the sophisticated understanding of statistics to understand exactly what the graphs are telling me, or how they relate to the underlying measurements. I'm relying on the results text alone to reach a conceptual understanding and just taking the graphs on trust.

      So, overall, the authors have a very nice body of work here and with the potential to add a new facet to our understanding of the origins of diversity in animal behavior. In addition to the interpretations they focus on here, this dataset also represents an advance in studying visual associative learning in general, and quite an amazing ability to make longitudinal measurements of many behavioral decisions within the same animals. Improving the data presentation to make it easier to follow for a larger swathe of researchers, especially in figures 2 and 3, will increase its potential impact.

      Comment on revised version:

      The authors have addressed my main points, including adding description of their statistical analyses and providing more detail about the different assays run and which animals were included in the same assay batches.

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to test the extent to which differences in learning capacity and experience contribute to behavioural variation in a genetically identical population under identical environmental conditions.

      Strengths:

      The authors developed and used a scaled-up version of a simple two-choice behavioural paradigm allowing them to test thousands of individuals across multiple genotypes. They then deployed clever and powerful statistical analysis methods and provided compelling evidence for a role of variability in learning in the expression of behavioural variation.

      Weaknesses:

      There are no major weaknesses, although some level of longitudinal analysis to strengthen the evidence for a strict definition of individuality would be a welcome extension of a future study. In addition, it would have been very interesting, although understandably beyond the current scope, to delineate a potential source of learning variability in the brain.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      "Learning is a fundamental source of individuality," by Manna and colleagues, interrogates different sources of variation in individual behavior. The authors place individual flies in a Y-shaped arena, which is a common design in the field, and illuminate the arms of the Y with blue versus green light. They track the color preference of individual animals and also perform operant conditioning, meaning that they teach the fly to avoid a particular color/arm by generating a foot shock when the fly enters that arm. There are a number of things that are impressive about this setup: The authors are able to collect data on thousands of individual flies of many different strain backgrounds, and they demonstrate a strong change in color preference after conditioning. This is nice, because in past papers, visual learning ability has been modest and difficult to study. To put a number on it, in this paper, animals on average don't show a color preference at the start of the assay, spending around 30% of their time in the one arm illuminated green, and the remaining time in the two arms illuminated blue. After conditioning, the average animal spends only 23% of its time in the green arm.

      The authors run 64 animals through the assay for each of 88 wild-type strains (maybe? see Major Point 1 below) and see considerable strain-specific (genetic) variation in the change in time spent in the shocked color after conditioning. Some strains show no learning, while others spend <10% of their time in the shocked color after conditioning. They also, I believe, see that some strains have more variability across individuals, which would suggest that some strains have stronger canalization at the development or circuit function level than others, i.e., some genotypes produce more consistent copies of the individual, others less consistent copies. (Or, some genotypes produce robust circuits, and others produce noisy circuits.)

      Finally, the authors argue statistically that learning itself increases variability in individual performance. This makes a lot of sense to me intuitively. Learning changes the physical/chemical properties of circuits in the brain, and because it evolves over time and interacts with environmental variables, it seems like it should send different animals down different channels. Or, at a conceptual level, if I learn to play the piano and my sister doesn't (because of some genetic difference between us or something stochastic), this learning experience will cause all sorts of other differences in our behavior as time passes. I also think the authors do have enough data to be able to make this finding. However, the presentation of the argument in this portion of the paper is hard for me to understand, and I am not an expert in statistics, so the strength of the result is difficult for me to evaluate.

      Major points

      (1) It's difficult to track through the paper the number of animals tested for different assays. At the beginning, it says N=5632, which works out to 64 flies for each of the 88 DGRP strains. 64 happens to be the number of parallel Y arenas they have. Later in the methods, there's a description of more variation within the set of 64 for each strain, two different parent sets per strain, different sexes, conditioned and unconditioned. And, while the results text focuses on the color learning, the methods discuss additional assays (place learning, multi-day learning).

      Given the numbers, does each run of the 64 mazes include all the tested flies of one strain, or are flies of many strains included in each batch? Do different flies do different assays (color, place, multi-day), or do they all do all the assays? Perhaps there is a table including this information already in the supplement, but I recommend making it much clearer in the main results text and methods. While the dataset is large, if it is split over many conditions and/or if batch and genotype confound each other, this will affect the robustness of the results and how strong the conclusions can be.

      (2) The data presentation in Figure 1 is elegant and easy to follow, but getting into Figure 2 and subsequently, I get lost in the statistics and have trouble understanding what is being measured. My understanding of the big picture is that while genetics and individual randomness contribute a lot to behavior, the evidence for learning as an amplifier of individuality is that variance in behavior among animals of the same strain increases over time in the conditioned group (i.e., the group that is doing the most learning, or a specific kind of learning), but not in the control group. This idea is illustrated in the flattening distributions in the cartoons in Figure 1A. The authors should include graphs of the real data that use the same format as in that cartoon. Instead, the graphs present "residuals," and I don't know what those are. I suspect it's "variation left over after accounting for effects of strain and individual stochasticity." I see the residuals being tracked per strain over time in Figure 2H, but I don't see the change over time in other graphs. I'm looking for something simple like, "variation within the strain at the beginning of learning and at later time points in learning." (But I'm not sure exactly what instantaneous measurement would be the focus in longitudinal analyses of learning behavior.)

      (3) Figure 3 is a cool stab at tracking down the precise mechanism by which a stochastic environment interacts with learning to send individuals along different behavioral routes. But again, like in Figure 2, I don't have the sophisticated understanding of statistics to understand exactly what the graphs are telling me, or how they relate to the underlying measurements. I'm relying on the results text alone to reach a conceptual understanding, and just taking the graphs on trust.

      So, overall, the authors have a very nice body of work here, and with the potential to add a new facet to our understanding of the origins of diversity in animal behavior. In addition to the interpretations they focus on here, this dataset also represents an advance in studying visual associative learning in general, and quite an amazing ability to make longitudinal measurements of many behavioral decisions within the same animals. Improving the data presentation to make it easier to follow for a larger swathe of researchers, especially in figures 2 and 3, will increase its potential impact.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to test the extent to which differences in learning capacity and experience contribute to behavioural variation in a genetically identical population under identical environmental conditions.

      Strengths:

      The authors developed and used a scaled-up version of a simple two-choice behavioural paradigm, allowing them to test thousands of individuals across multiple genotypes. They then deployed clever and powerful statistical analysis methods and provided compelling evidence for a role of variability in learning in the expression of behavioural variation.

      Weaknesses:

      There are no major weaknesses, although some level of longitudinal analysis to strengthen the evidence for a strict definition of individuality would be a welcome extension of a future study. In addition, it would have been very interesting, although understandably beyond the current scope, to delineate a potential source of learning variability in the brain.

      Following our provisional response to the reviewers, we have implemented these additions to the manuscript:

      (1) We have added 7 additional tables (Table 1-6 and table 8) to the supplementary that detail how many individual flies were used in which of the seven separate experiments, how the individuals were distributed across genotypes, replicates and sexes, and how many were filtered out before the final analysis. At the bottom of each table, we added a short description of the type of experiment and a brief explanation of the filtering. The four smaller experiments where we tested the two mutant lines and one wild-type DGRP line were used primarily to test and validate the experimental platform and the behavioural paradigm used for the main experiment. In these four experiments we tested green place learning, blue place learning, green colour learning, blue colour learning in four separate batches of flies. In each of these experiments we used 192 individuals (64 individuals x 3 genotypes x 4 experiments = 768 individuals in total). In the multiday experiment we used 64 flies per genotype per each of the four groups of sequences of learning paradigms, in total 512 individual flies. Here, each of these 512 individuals were retested in different learning paradigms over 4 days (Table 5). The main experiment was the green place learning (Table 6) where we tested all 88 DGRP lines and again the two mutant lines was used to obtain the majority of the main results and conclusions (64 individuals x 90 genotypes = 5760 individuals). Lastly, additional 896 individuals were measured in the experiment using blue place learning paradigm to test the consistency of learning behaviour as opposed to colour bias within genotype (Table 8). In summary, in all experiments, we have always measured behaviour in 64 flies per genotype (full loading of the behavioural platform), and they were distributed almost entirely evenly across replicates, sexes, and conditions (control vs conditioned). No individual was reused across experiments. In most cases, after filtering the data, 60 or fewer individuals were used in final analyses. For the very few deviations from this experimental design (which occurred due to unforeseen events such as dropped/sick vials, flies flying away or accidentally squished during setup, skewed number of males and females etc.) we added a short explanation in the text below the tables. In total, across all reported experiments in this study, we measured behaviour in 7936 individuals.

      (2) We have added a schematic visual representation of classical measurement of individuality (variance of the distribution of behaviour within genotype where genetically identical individuals are raised in the same environment), entropy-based measurement of individuality (residual individuality) and the change in residual individuality, as we use them in this study (Figure 2D). We also provide a list of different DH<sub>resid</sub> measures and what distributions are being compared across the DH<sub>resid</sub> in the same figure. We hope this will serve as a more intuitive explanation of individuality and help readers interpret and follow more easily the results that we report after this figure.

      (3) In the same vein, we added another schematic visual representation to Figure 3 (Figure 3F) where we depict how distributions of individual behaviour may change with every decision and how this change translates to (or can be read out from) the change in residual individuality. We have also renamed the X axis of Figure 3E to “DH<sub>resid</sub> Start”, so that it is clearer what is measured here and matches the explanation in Figure 2D.

      (4) We have added two additional supplementary figures where the reader can inspect in more detail how the distributions of individual behaviour change longitudinally across time for each genotype in control and conditioned (Figure 3 – Figure supplement 2 and  Figure 3 – Figure supplement 3). From these figures one can glance how variance as well as the shapes of the distributions change as the flies learn in the conditioned setting, and how they remain largely the same in the control where flies behave spontaneously. We have added a sentence in the main text to introduce these figures: “We found that the distributions of individual behaviour were broader and their shapes changed substantially over the course of the experiment for the conditioned flies, and not for the control flies Figure 3 - figure supplement 2, Figure 3 - figure supplement 3).”

      (5) As noted in the first provisional response to reviewers, we changed the sentence “In every individual, behaviour is shaped by deterministic, genetic factors and by environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.” to “In every individual, behaviour is shaped by fixed genetic factors and by variable environmental events throughout lifetime, which may be stochastic and can occur at the molecular, cellular, organismal and even population scales.”

      (6) Some sentences were edited in the results so that we can correctly refer to the newly added tables and figures. Context, meaning or interpretation of the results in these sentences was not altered.

      (7) While adding the new table references to the text, we noticed a typo that propagated in the previous version where the number of flies used in the main experiment was stated to be N= 5238, when in fact it should have been N=5239. This is now fixed.

      We once again thank the reviewers for their comments and suggestions – we believe their suggestions helped us improve the presentation and interpretability of our study and we hope the reviewers and readers will agree with this as well.

    1. eLife Assessment

      This short report is an important study that visual acuity declines nonlinearly with cone dropout, while eye motion partially compensates by improving sampling from remaining cones. The method for experimentally simulating cone dropout is compelling, leveraging state-of-the-art imaging and testing in human subjects.

    2. Reviewer #1 (Public review):

      The authors demonstrate an innovative approach to investigate the effect of cone dropout on visual acuity using their newly developed Oz platform. By systematically reducing the coverage of real-world input to the cone photoreceptor mosaic ("cone dropout condition"), the authors are able to assess how having less cones leads to reduced vision, in comparison to existing approaches ("pixel dropout condition").

      The observation of visual acuity maintenance with cone dropout has been a longstanding mystery since the 2013/2018 papers by Ratnam and Foote. The authors should be commended for their approach to address this important question. However, there are some simplifications and assumptions being applied to make this jump (i.e. that a 50% reduction in cone stimulation in a healthy eye is comparable to a 50% reduction in cone density in a patient). It seems unlikely that in a patient eye, with cone dropout, that there will be gaps in the mosaic. Not considering any other non-photoreceptor related reasons for visual acuity loss which can occur in patients, the cone aperture acceptance angle may be different due to changes in cone size or packing; the sensitivity of individual cones may also be reduced due to deficits in the visual cycle recovery which could be affected in disease. Some of these limitations could be addressed and acknowledged more explicitly.

      The capture of a rich dataset including both cone imaging and eye motion is valuable. Since the C stimulus test relies on foveal fixation, and there is a high degree of subject-to-subject variation in peak cone density, the authors may wish to report on peak cone density measurements of the subjects being included in this study. In addition, evaluating whether the eye motion is affected by simulated cone dropout condition can help to rule out whether these observed effects can be attributed to eye motion.

      Overall, this is an impressive study incorporating state-of-the-art technology to probe the fundamental limits of human vision.

      Comments on revised version.

      The authors have nicely addressed my concerns. The additional clarifications and revised text have strengthened the paper. Thank you also for pointing out the inaccuracy of referring to the system as the olo system; this has been corrected.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors demonstrate an innovative approach to investigate the effect of cone dropout on visual acuity using their newly developed olo system. By systematically reducing the coverage of real-world input to the cone photoreceptor mosaic ("cone dropout condition"), the authors are able to assess how having fewer cones leads to reduced vision, in comparison to existing approaches ("pixel dropout condition").

      The capture of a rich dataset, including cone imaging and eye motion, is valuable. Benchmarking with the prior literature, suggesting that good visual acuity can be maintained despite a 50% loss in cone density, is impressive. However, it is known that cone density varies dramatically from the peak cone density location in the foveal center to even a location a few degrees outside of the fovea. In addition, there is a high degree of subject-to-subject variation in peak cone density. Given that the C stimulus is hollow in the middle, the stimulus does not actually hit the location of the peak cone density but must land slightly outside of it. Therefore, considering the actual cone density of where the stimulus lands will be important to discuss and/or analyze.

      The reviewer is correct that the cone density will vary dramatically with distance from the foveal center. However, importantly, in our experiment the Landolt C stimulus is fixed in the world and the eye is free to move across it. Therefore, the subject can direct their gaze to any part of the letter, rather than it being fixed to the hollow center of the letter. In the worst case, if the subject kept their gaze fixed at the center of the letter and were viewing the largest letter corresponding to the worst acuity measured in our experiments (20/100), the cone density on average would be 13% lower at the edge of the letter than at the center. However, it is unlikely that a subject would have fixated in this manner, and the vast majority of letters shown during the experiments were much smaller than 20/100.

      For completeness, we have calculated the peak cone densities for each of our subjects using the cone density centroid method described by Reiniger et al (2021). We have added these numbers and a description of the method to the Subjects section in Methods and Materials on lines 339-344.

      The observation of visual acuity maintenance with cone dropout has been a longstanding mystery since the 2013/2018 papers by Ratnam and Foote. The authors should be commended for their approach to addressing this important question. However, there are some simplifications and assumptions being applied to make this jump (i.e., that a 50% reduction in cone stimulation in a healthy eye is comparable to a 50% reduction in cone density in a patient). It seems unlikely that, in a patient's eye, with cone dropout, there will be gaps in the mosaic. Not considering any other non-photoreceptor-related reasons for visual acuity loss, which can occur in patients, the cone aperture acceptance angle may be different due to changes in cone size or packing; the sensitivity of individual cones may also be reduced due to deficits in the visual cycle recovery, which could be affected in disease. Some of these limitations could be addressed and acknowledged more explicitly.

      Cone loss does manifest differently in different retinal degenerative diseases, and in this work we implement dropout on a cone-by-cone level. To address the reviewer’s points, we have added a description of the range of spatial manifestations of cone loss across a range of diseases to the Discussion section on lines 320-328, and emphasize that we focus on one particular manifestation in this paper.

      Overall, this is an impressive study incorporating state-of-the-art technology to probe the fundamental limits of human vision.

      We thank the reviewer for their helpful comments and constructive feedback.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The patient recruitment limitation seems to be a bit artificial here. This is indeed a limitation, but perhaps not the primary limitation or motivating factor. Consider removing/rephrasing this motivation.

      We agree with the reviewer’s comment, and have removed it from the abstract and removed its framing as a limitation in the Introduction section in two instances on lines 31 and 36-37.

      Was the peak cone density quantified, and the location of the peak cone density determined? Reporting the range of eccentricities over which the C stimulus lands relative to the peak cone density location, as well as the actual cone density that is being used to sample the C stimulus on the retina, seems to be important for contextualizing this study. It is a bit too simple to only consider the percentage of cones that are reduced.

      We have added the peak cone density for each subject to the Subjects section in Methods and Materials. This experiment did not require fixation; rather, the Landolt C stimulus was fixed in space and the subject could move their eye freely across it, meaning that different parts of the fovea may have sampled the letter on different trials. At the highest dropout percentage, where acuity was the worst, the letter size was 20/100 at threshold, or 25 arcmin. If the subject were to fixate with their peak cone density at the center of the letter, we have computed that the average decrease in cone density at the edge of the letter (12.5 arcmin away) would be 13%. The majority of trials in the experiment showed letters that were much smaller than this, and would have been subject to even less variation in cone density.

      Do the authors have any idea about the approximate size of the cones in healthy subjects compared to diseased eyes? Importantly, if the cones in patients are larger due to the dropout of their neighbors, then the retinal coverage area would be larger due to their larger size, and the amount of light that can be coupled into larger cones may also be larger. Can this be modeled or discussed?

      In our implementation, we did not emulate a change in cone size, and rather modeled the loss as discrete holes in an otherwise intact retina. We have added text to the Discussion section on lines 320-328 to make the distinction between this form of cone loss and other forms where cones appear to fill in for their neighbors resulting in a contiguous mosaic of lower density overall.

      Acknowledging some of the shortcomings of this approach for simulating the patient condition could be improved. It may be worthwhile to tone down the premise of this paper if these cannot be adequately explained.

      In order to tone down the premise of the paper, we have made the following changes to the text.

      We now emphasize on lines 65-69 in the Introduction section that we focus specifically on the impact of cone loss on acuity without modeling downstream factors.

      In addition to the description of other diseases that we added in response to a previous comment, we have also added the following text on lines 313-318 of the Discussion section:

      “... factors beyond the photoreceptors also play a role in shaping vision under retinal degeneration. In this work, we did not model any downstream factors such as shorter outer segments (Foote 2018), retinal rewiring (Jones 2016, Lee 2021), or ganglion cell hyperactivity (Kramer 2023). Instead, we sought to characterize vision in the presence of cone loss at the lowest possible level, considering only the decrease in sampling power at the retinal input.”

      In the methods, it is not completely clear the rationale for determining the appropriate size of the C stimulus. How is visual acuity determined if the C stimulus size is not changed?

      A more careful explanation of how the C stimulus size is set is warranted.

      In the experiments measuring visual acuity, the C stimulus size was selected by a QUEST staircase on each trial. For each dropout condition, we ran 4 interleaved QUEST staircases with 20 trials each. This is described in the “Acuity Threshold Experiment” section in the main text (lines 99-100) and in Materials and Methods (line 407). To clarify further, we have updated the following sentence on line 423:

      “For each condition, we ran 4 interleaved QUEST staircase procedures (Watson and Pelli, 1983) with 20 trials per staircase, which varied the size of the Landolt C on each trial.”

      What is the clinical visual acuity of the subjects being tested? It seems important to report this if the authors want to use their C stimulus as a proxy for clinical visual acuity.

      The subjects being tested have excellent acuity. In Figure 1, we can see that their adaptive-optics-corrected acuity for the baseline 0% dropout condition ranges from approximately 20/10 to 20/12.5 across the 4 subjects. We have added the following statement to the “Subjects” section in Materials and Methods (line 339):

      “All subjects self-reported to have normal vision.”

      Given that the title of the paper emphasizes the role of eye motion, it seems that a more careful analysis of the magnitude and type(s) of eye motion could be added. There are eye motion data provided in the supplemental figure, but it is not completely clear how this eye motion data is actually being used to derive meaningful information about visual acuity.

      We performed analyses to determine whether there seemed to be a significant difference in eye motion patterns between the cone and pixel dropout conditions, which was described in the section “Analysis of Eye Motion Data” and in Supplementary Figure S1. In that figure, we show that for all 4 subjects there is no significant difference in the iso-density contour area containing 68% of their eye motion data. We suggest in the paper that due to the pseudorandom presentation of trials and the limited duration of those trials, subjects were unlikely to adapt or adjust their eye movement, and that instead their natural eye motion served as a data collector that improved acuity.

      What is the accuracy of the eye motion and cone dropout stimulation delivery in the fovea? Given the small size of the cones, it seems that this is one of the most challenging locations of the eye to test with this new olo technology.

      Eye tracking and targeted light delivery are crucial in the AOSLO system and the reviewer is correct to point out that this is most difficult to achieve at the foveal center. To address this concern, we have done some simple modeling and have added the following text to the Cone-by-Cone Stimulation section in the Methods and Materials.

      “This latency, combined with other factors such as diffraction and residual aberrations, limit the ability to restrict the light to only the targeted cone. Considering a 543-nm focus through a 7.2 mm pupil, a random tracking error with a full-width-at-half maximum (FWHM) of 0.5 arcminutes (Harmening et al. (2014)), a 0.0125 diopter residual defocus error (maximum error given the step sizes of 0.025 diopters in the AOSLO defocus controller), an average cone spacing of 0.5 arcminutes (Wang et al. (2019)), and a Gaussian cone acceptance aperture with a FWHM that is 0.5 times the inner segment diameter (Macleod et al. (1992)), we estimate that each targeted cone receives 5.41 times more light than its nearest neighbor. This means that the ’dead’ cones cannot be fully excluded from the visual processing. Furthermore, the light leakage reported in Fong et al. (2025) further adds to the signal of non-targeted cones.

      Nevertheless, it is important to point out that the information about the stimulus (Landolt C in our case) is sampled at the targeted cone’s location and so, although nearby stimulated cones might detect light, they do not contribute to any increases in the sampling process. This is analogous to adding defocus blur to letters in the pixel dropout condition as neither situation will improve the spatial information.”

      References

      Reiniger, J.L., Domdei, N., Holz, F.G., Harmening, W.M.: Human gaze is systematically offset from the center of cone topography. Current Biology 31(18), 4188–4193 (2021)

    1. eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is convincing, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management.

    2. Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where 2 herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables, and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section).

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications on the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists. Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015). It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats). Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

    3. Reviewer #2 (Public review):

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via use of shared resources.

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good of demonstrating the potential conservation impacts of their research.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence that plateau pikas, at moderate densities, can facilitate yak nutrition by suppressing a poisonous plant, offering a helpful perspective on reciprocal interactions between small mammal ecosystem engineers and large herbivores. The evidence is solid, supported by a manipulative field experiment and appropriate measurements of intermediary ecological processes, although some claims about density dependence, competition, and stress-gradient mechanisms are not fully supported by the experimental design. The work will be of interest to ecologists, conservation biologists, and rangeland managers, particularly those studying grassland herbivore interactions and livestock management on the Qinghai-Tibetan Plateau.

      Thank you very much for these positive assessments of our work. Below, we provide point-by-point responses to the comments from the two peer reviewers, and we hope these revisions are satisfied by you and the reviewers.

      Reviewer #1 (Public review):

      Summary:

      This is important and significant work because it helps describe the complexity of interactions between system components where two herbivores interact with vegetation. Whereas other studies have shown that the larger ungulate (yaks, Bos grunniens, in this case) can facilitate the abundance and population growth of the smaller (the semi-fossorial lagomorph, Ochotona curzoniae, plateau pika hereafter), this study flips the tables and shows that, at least under some conditions, moderate densities of the plateau facilitate the nutritional condition of yaks.

      The study was not designed to investigate the reasons that pikas clip Stellera chamaejasme. That said, based on other studies and general knowledge of the ecology of these pikas, it is likely that they clip (although do not eat) this plant because its relatively large size hinders predator detection. This species of pika does better where vegetation height is low than where it is higher.

      Strengths:

      Notably, the strong inference the authors can claim for their results is supported by the careful experimental design. A weaker paper would have simply noted correlations between pika burrow density and yak feeding efficiency without experimental removal. This paper, to its credit, not only used experimental removals but also documented the various intermediary results that support the ultimate conclusions. The statistical approaches used appear to be appropriate. (Readers are encouraged to read the full Materials and Methods, which are available in the Supplementary Materials section.)

      We appreciate these positive comments on our work.

      Weaknesses:

      Although the study was well designed and executed, and its conclusions appear strongly supported, readers interested in the management implications of the Qinghai-Tibetan Plateau should be mindful of its limitations. First, the study site, at approximately 3,200 m elevation, was relatively low by Qinghai-Tibetan Plateau standards. Stellera chamaejasme becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. Thus, it would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace S. chamaejasme as the problematic plant for pastoralists.

      Thank you for this suggestion. We have acknowledged this limitation in the Discussion by adding the paragraph below (see the Third point):

      “Despite of these, several questions deserve further investigation. First, our study examined pika–yak interactions only during the summer period, when food resources are most abundant. Whether such facilitative effects weaken or even shift toward competition under more stressful conditions—for example, when forage becomes limited during autumn or winter—remains to be tested. Second, if the documented facilitation of yak nutrition by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition? Third, our study site is located at approximately 3,200 m elevation, relatively low by Qinghai-Tibetan Plateau standards. Stellera becomes less common at elevations > 4,000 m, where a majority of livestock grazing occurs. It would be instructive to learn, through follow-up studies, whether similar facilitation occurs where unpalatable (and mildly poisonous) species in such genera as Astragalus, Oxytropis, and Thermopsis replace Stellera as the problematic plants for pastoralists (Lu et al., 2012; Li and Zhao, 2025). Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 272-286 in the Discussion section.

      Second, the authors make no mention of wild ungulates, so it is unclear what, if any, role they may have played in this system. At least one study in Qinghai Province, albeit at a slightly higher elevation, showed that not only pikas, but also Tibetan gazelles (Procapra picticaudata), which were commonly observed on grazed pastures, grazed more frequently on some dicots avoided by domestic sheep than did the livestock themselves (Harris et al. 2015).

      Citation:

      Harris RB, Wang, WY, Badinqiuying , Smith AT, Bedunah DJ (2015) Herbivory and Competition of Tibetan Steppe Vegetation in Winter Pasture: Effects of Livestock Exclosure and Plateau Pika Reduction. PLoS ONE 10(7): e0132897. doi:10.1371/journal.pone.0132897

      Thank you for this suggestion. We have added more details about the study site, particularly regarding wild ungulates, in the Methods section. Specifically, we have included the sentence of “Wild ungulates, such as Tibetan gazelles (Procapra picticaudata) (Harris et al., 2015), and other small mammals such as rabbits and zokors, occur rarely in the area.”

      See these revisions in Line 333-335 in the Methods section.

      It would also be instructive to learn if similar facilitation as observed here applied to the other principal livestock species in the area, domestic sheep (which are often herded together with smaller numbers of domestic goats).

      Thank you for the suggestion. We have acknowledged this limitation in the Discussion, by adding a paragraph as: “Finally, it is unclear whether similar facilitation as observed here applied to the other principal livestock species in the area, such as domestic sheep and goats.”

      See these revisions in Line 284-286 in the Discussion section.

      Finally, as suggested by this study, the interactions between all components of the system are complex and interactive. If pika facilitation of yak nutrition at the densities documented results in herders increasing yak density, might the increased herbivory from the domestic animals provide the conditions for the pika population to increase beyond the densities observed here, and thus toward the levels where facilitation yields to competition?

      Thank you for your suggestion. We have acknowledged this limitation in the Discussion, by adding the paragraph as “Second, if the documented facilitation of yaks by pikas prompts herders to increase yak densities, could the resulting rise in livestock herbivory push pika populations beyond the levels observed here, potentially toward the threshold where facilitation gives way to competition (Yang et al., 2026)?”

      See these revisions in Line 276-279 in the Discussion section.

      Reviewer #1 (Recommendations for the authors):

      Although no doubt a bit sensitive, it would have been better to reveal a bit more about how pikas were removed.

      We have provided more details about how pikas were removed in the no-pika treatment, by adding “For the no-pika treatment, pikas were trapped once every two weeks using 30 live traps (25 cm high × 25 cm wide × 40 cm long) within each plot and relocated elsewhere in the study site.” in the Methods section. We didn’t recorded how many pikas were removed from the corresponding plots, so no data were available for this point.

      See these revisions in Line 411-413 in the Methods section.

      The authors also missed a few relevant papers worth citing, including Badingqiuying, R. B. Harris, and A. T. Smith. 2018. Summer habitat use of plateau pikas (Ochotona curzoniae) in response to winter livestock grazing in the alpine steppe Qinghai-Tibetan Plateau. Arctic, Antarctic, and Alpine Research 50 (1): e1447190

      We have cited this key paper in Line 75 in the Introduction section.

      Reviewer #2 (Public review):

      Summary:

      This study uses a combination of field sampling and manipulative experiments to test for facilitative impacts of pikas on yaks via suppression of a poisonous forb. The authors found that, when Stellera forbs were present, yak weight increases over the growing season were greater in the presence of pikas compared to in their absence. This occurred because, although pikas do not consume Stellera, they clip it and use it in nest/burrow construction, thereby decreasing its relative abundance in the plant community. Thus, overall, the study contributes to our understanding of how herbivores of different size classes indirectly affect each other via the use of shared resources.

      Strengths:

      It is well known that large herbivores on grasslands impact smaller animals, but the reciprocal interaction is rarely tested. Thus, this study asks a valuable question, and the experiment is well-designed to test it. The authors also do a good job of demonstrating the potential conservation impacts of their research.

      We appreciate these positive comments on our work.

      Weaknesses:

      What the authors tested is really cool, but their claims go far beyond what they can say based on their experimental design. For example, the authors claim to show that pika impacts on yaks display density-dependent transitions from competition to facilitation. However, their experiment only looked at the presence (at moderate densities) and absence of pikas, and they only tested for facilitation, not competition.

      The paper would also benefit from changes to the framing in the introduction and discussion. For example, the authors pitch the work as a test of the stress-gradient hypothesis. However, there is no abiotic stress gradient in the study, which is an essential component of the SGH. They also pitch the work in terms of density dependence, but there is no significant variation in population densities beyond the presence-absence binary. The paper would be stronger if they focused their framing around the literature on facilitative interactions across mammals of different size classes, especially indirect facilitation via use of shared resources, which is what this paper is really about.

      We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the Stress Gradient Hypothesis (SGH). Thus, we deleted the description on SGH. However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate densities of pikas has the best beneficial effects on yak growth. We added a separate paragraph in discussion to have a clear discussion.

      We have made the major revisions below to address these concerns.

      (1) We have revised the title into “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands ”.

      (2) We have deleted all the statements about facilitation and competition predicted by the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We added a paragraph in discussion (Line 248-259) to have a clear discussion on the humped relation between yak weight gains and pika burrow densities as “Because of the natural variations in pika density in the pika-present treatment, we were able to obtain a hump-shaped relationship between yak weight gains and pika burrow densities in these plots. Compared with the absence of pikas, the facilitative effect reached its maximum at approximately 200 burrows/ha but became competitive at densities exceeding 400 burrows/ha (Figure 3C). This result reveals that pika density modulates the net outcome for yak weight gain, with a facilitation peak at ~200 burrows/ha and a competition onset above 400 burrows/ha. Our findings offer empirical evidence for the non-monotonicity theory, under which the competition-facilitation balance varies with population density: facilitation dominates at low densities, competition at high densities, and these density-dependent shifts may underpin community stability and productivity (Zhang, 2003; Zhang et al., 2015). The theory further holds that the facilitation threshold, not the competition-facilitation transition, is the critical factor governing the stability of interacting species or communities (Zhang et al., 2015).”

      Most importantly, there are inconsistencies in what is visualized in the figures compared to the model results. For example, the results section in several places notes a lack of significant interaction terms in the model but shows interactions in the p-values on the figures.

      In the Results section, there are only two places where we discussed non-significant interactions: Line 175–177 “Pikas and Stellera had no interactive effects on abundance of sedges, forbs, and neutral detergent fiber (NDF) of total forage for yaks (Figure 3F, I and Appendix 1—figure 1, table 5,8).” and Line 190–192 “Pikas and Stellera had no interactive effects on yaks’ foraging efficiency on forbs (Appendix 1—figure 2, table 10).”.

      We have cross-checked both the Results section and the Figures sections mentioned above, and confirmed that they are consistent now.

      The authors also plot smoothed lines rather than their model results and then draw interpretations from those lines that cannot be tested in the models that they used.

      Thank you for the suggestion. Now we have added the Appendix 1—table 3 and Appendix 1—table 7 for the model results of generalized additive models (GAMs) for Figure 2C and Figure 3C that plotted with smoothed lines in Appendix 1.

      There are also missing details that are important for model interpretation, including the distributions used and the sample sizes.

      We have provided the Appendix 1—table 13 to summarize all statistical models used in the study, including the distributions used and the sample sizes in the Appendix 1.

      We have also added a sentence of “A summary of all statistical models used in the study is available in Appendix 1 table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Another major concern with experimental design is in the forage nutrient analyses. The authors picked plants along a grazing trail, then measured nutrient content without standardizing based on plant species, so any differences across treatments could be because of what they happened to grab rather than overall forage quality.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g). We have revised this section as below.

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Figure 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) Introduction

      Line 53 - I wouldn't describe small mammals like rodents as keystone species. They can have strong impacts on ecosystems, but not disproportionate relative to population size, which is a key part of that definition. It may be true when you talk specifically about pikas later on, but not small mammals as a general category.

      We have replaced this term with “consumers” here, see Line 63.

      Lines 58-61 - Good hook

      Thank you for this positive comment.

      Lines 69-74 - I don't think the stress gradient hypothesis is the right pitch for this work. The SGH posits that facilitation increases with abiotic stress, but no abiotic stressors were measured in this study. Population density interacts with abiotic factors in the SGH, but population density in and of itself is not an abiotic stressor. The papers you cite here all look at the interaction between population density and abiotic stressors (e.g., water availability). So these lines set me up to expect a stress gradient in your experiment that didn't exist, then left me confused later on. It would be better to highlight the strength of your work (lines 70-76 pose interesting questions and predictions) rather than trying to make it fit within the SGH.

      Thank you for pointing out this problem. We have deleted the statements about SGH here.

      See these revisions in Line 88-91 in the Introduction section.

      (2) Materials and Methods

      Lines 498-509 - Did you verify beforehand that no plants within the enclosures had been grazed on? How did you know that the consumption or clipping was specifically from those pikas?

      We have clarified here by adding “Before cage installation, we carefully checked the plants within each plot and removed those that had been previously grazed or damaged by herbivores.” in the Methods section. In this case, we made it sure that the consumption or clipping was specifically from those pikas with the cages.

      See the revisions in Line 349-350 in the Methods section.

      Lines 509-512 - I would be careful calling this preference. It's really just a record of what they consumed along paths they were walking, which could be about accessibility and convenience as much as preference.

      We have replaced “diet preferences” with “diet composition” here, see Line 359 in the Methods section.

      Line 514 - Clarify the specific question or hypothesis you're testing with this field survey, beyond just generally testing associations.

      We have modified the sentences here as “In July 2021, we investigated the potential facilitation of pikas on yaks mediated by the poisonous Stellera forbs under unmanipulated field conditions in the study site.”.

      See Line 365-366 in the Methods section.

      Lines 532-534 - The intro for this paper sets it up to be about density-dependent movement from competition to facilitation, but the experiment here is set up to compare pika presence/absence. The framing of the paper needs to be adjusted to better align with this experimental design.

      The same issue as mentioned above. We agree that our work had explored only the facilitative effects of pikas on yaks, rather than the balance between competition and facilitation as predicted by the Stress Gradient Hypothesis (SGH). However, the finding of a humped relation between yak weight gains and pika burrow densities (Figure 3C) is very important which provides evidence that moderate density of pika has the best benefical effect on yak. We added an separate paragraph in discussion to have a clear discussion about this point in the Discussion section.

      We have made the major revisions below to address this concern.

      (1) We have revised the title as “Small mammalian herbivores at moderate densities facilitate livestock growth by improving vegetation composition in grasslands”.

      (2) We have deleted the statements about facilitation and competition and the SGH in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph about SGH was removed here), and the References sections.

      (3) We kept the discussion on the humped relation between yak weight gains and pika burrow densities (Figure 3C). We added an separate paragraph in discussion to have a clear discussion about this point. For details, see Line 248-259 in the Discussion section.

      Lines 566-569 - Why did you need to simulate pika clipping when you already had pika presence/absence treatments?

      We conducted Stellera removal treatment by simulating poisonous plant clipping behaviors of pikas because we want to confirm that the shifts in abundance of this dominant poisonous plant species is the key mechanism in driving pika-yak facilitation in our system. If we simply looked the differences in yak weight gain in the pika presence/absence treatments, it should be difficult to secure the underlying mechanism. In addition to the reduction in abundance of the poisonous Stellera, pikas may cause a variety of shifts vegetation properties including plant productivity and diversity, and soil disturbances that can exert direct and indirect effects on yak foraging activities, and thus their weight gains.

      Lines 566-569 - Is there another citation you can give showing that Stellera forbs taller than 20cm are both preferred by pikas and exert greater impacts on plant/animal communities? Those are big assumptions that need to be better supported or explained. If there was a logistical reason that you didn't remove all of the smaller forbs, that needs to be laid out as well.

      We have added one new citation here to support this method here. We have revised this section as “To simulate the clipping behavior of pikas, we clipped only those Stellera forbs exceeding 20 cm in height. This threshold was chosen based on previous observations that pikas preferentially target large forbs of this size (Liu et al., 2009).”

      Liu W, Zhang Y, Wang X, Zhao JZ, Xu QM, Zhou L. 2009. The relationship of the harvesting behavior of plateau pikas with the plant community (In Chinese). Acta Theriologica Sinica 29:40-49.

      Also, we have deleted the sentence of “and can exert significant impacts on the plant community and on yak grazing behaviors (Z.Z., field observations)” mentioned above, because these patterns were observed only by the authors in the field and lack supporting data.

      See these revisions in Line 419-422 in the Methods section.

      Line 587 - sample size per treatment? Was it consistently one species that you measured, and if so, which one? If you collected different species or a mix of species across treatments, then you can't really compare the nutrient values because you haven't accounted for interspecific variation.

      We have revised this section to provide more details on how forage samples were collected and their quality were analyzed. Specifically, five forage samples were collected per grazing plot, focusing on the two dominant plant species—one sedge and one grass—that were most frequently grazed by yaks. To ensure comparability across plots and treatments, we mixed the two species at equal dry mass (5 g).

      We have revised this section as below:

      “To assess forage quality, five forage samples were collected from each grazing plot to quantify their nutritive values. To obtain samples that reflect the forage actually consumed by yaks, we tracked the animals along their grazing paths and collected the plant tissues of the two most frequently consumed species: the dominant sedge Kobresia humilis and the dominant grass Elymus nutans (Fig. 2B; Pan et al., 2019). The collected tissues of each species were dried in a forced-air oven at 60 °C for 48 h, then ground through a 1-mm mesh. Subsequently, 5 g of each dried and ground species were combined in a 1:1 dry mass ratio, and the resulting mixture was stored in plastic bags for subsequent analyses.”

      See these revisions in Line 439-447 in the Methods section.

      Lines 603-613 - Please provide the dependent variables in each model, as well as any interaction terms, in addition to the random effects. Please also state explicitly what you were trying to test with each of these models.

      There are two models that use the tweedie family (forb and sedge bite rate). Indeed, we need to include the tweedie power parameter to help understand the mixture of the three families. We have included the p-value (power) in the Appendix 1—table 13 in the Appendix 1. We chose to use tweedie because a normal gaussian family model fitted the results poorly.

      We have added the Appendix 1—table 13 in the Appendix 1, which provides the summary of all statistical models used in the study, including response variables, model type, distribution family (with Tweedie power parameter where applicable), interaction terms, random effects structure, and sample sizes.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Lines 613-614 - Tweedie is a category of distributions that includes quite a few different options, including Gaussian, Poisson, and Gamma distributions, some of which are normal and some of which are not. So, justifying the use of Tweedie distributions in your model structure doesn't really make sense, and it doesn't really tell me which distribution each model pulled from. Please clarify specifically which distributions you used for each model and why.

      We have now added the reason why we used Tweedie distributions by adding the sentence of “There were two models that used the tweedie family (forb and sedge bite rate). We chose to use tweedie because a normal gaussian family model fitted the results poorly” in Line 468-470 in the Statistical analyses section.

      We have also included the p-value (power) for forb and sedge bite rate in the Appendix 1—table 13 in the Appendix 1.

      Lines 615-618 - Provide citations for R packages described in the text.

      We have provided all the related citations for all R packages described in the text, as listed below.

      glmmTMB: Brooks, M. E., Kristensen, K., van Benthem, K. J., Magnusson, A., Berg, C. W., Nielsen, A., Skaug, H. J., Maechler, M., & Bolker, B. M. (2017). glmmTMB Balances Speed and Flexibility Among Packages for Zero-inflated Generalized Linear Mixed Modeling. The R Journal, 9(2), 378–400. https://doi.org/10.32614/RJ-2017-066

      mgcv: Wood, S.N. (2017). Generalized Additive Models: An Introduction with R (2nd edition). CRC Press. AND Wood, S.N. (2011). Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 73(1), 3–36. https://doi.org/10.1111/j.1467-9868.2010.00749.x

      DHARMa: Hartig, F. (2022). DHARMa: Residual Diagnostics for Hierarchical (Multi-Level / Mixed) Regression Models. R package version 0.4.6. https://CRAN.R-project.org/package=DHARMa

      tidyverse: Wickham, H., Averick, M., Bryan, J., Chang, W., McGowan, L.D., François, R., Grolemund, G., Hayes, A., Henry, L., Hester, J., Kuhn, M., Pedersen, T.L., Miller, E., Bache, S.M., Müller, K., Ooms, J., Robinson, D., Seidel, D.P., Spinu, V., Takahashi, K., Vaughan, D., Wilke, C., Woo, K., & Yutani, H. (2019). Welcome to the tidyverse. Journal of Open Source Software, 4(43), 1686. https://doi.org/10.21105/joss.01686

      See these revisions in Line 472-475 in the Statistical Analyses section.

      (3) Results

      Line 127 - Somewhere in the intro or methods, describe what clipping is and why the pikas do it.

      We have provided more details about the clipping behaviors of pikas by adding “Notably, pikas often clip (although do not eat) the wolf poison S. chamaejasmehas because its relatively large size hinders predator detection (Fan et al., 1998).” in Line 330-332 in the Methods section.

      Line 128 - Change from "preferred" to "consumed greater proportions of"

      See the correction in Line 146 in the Results section.

      Line 136 - You measured yak weight once a month, so give the result in monthly weight gain rather than daily.

      Here we preferred to keep the unit of daily weight gains (converted from monthly ones), as this is the standard presentation for livestock growth performance, also see Fig. 1 in Odadi et al., 2011 Science’s paper.

      W. O. Odadi, M. K. Karachi, S. A. Abdulrazak, T. P. Young, Science 333, 1753–1755 (2011).

      Lines 138-140 - The linear models, as you described them in the methods (lines 603-620), don't test for hump-shaped relationships. Please update the methods to explain how you tested the density relationship and how you got this interpretation.

      The description for Figure 3C here showed estimates from a GAM (family: gaussian), and we did not use a linear model in this figure.

      We have now added a new model summary Appendix 1—table 7 for this Figure 3C in the Appendix 1.

      Line 145 - I don't think you can claim that the total available forage was more nutritious for yaks, as you picked plants that yaks happened to be chewing along a grazing path. You would need to take samples from a consistent set of plant species at random locations to make this claim.

      Sorry for this confusion. We modified “the total available forage” as “the major available forage” here (see Line 172) because we collected the same forage plant species and analyzed their nutrients.

      The same as mentioned above, to clarify the sampling methods, we have also revised this point in the Methods section to provide more detail on how forage samples were collected and their quality were analyzed. See these revisions in Line 439-447 in the Methods section.

      (4) Discussion

      Line 166 - Need to address inconsistency in how you talk about pika density. Here you talk about the impacts of pikas at moderate densities, which I think is a fair claim. Elsewhere, you talk about density-dependence, which I don't think you really measured, given that all your sampling was either in the absence of pikas or within a narrow window of densities that can all be categorized as moderate.

      Done! As mentioned above, we have removed the term of “density-dependence” in the whole manuscript, but keep the term of “moderate density” in the Abstract (Line 56-59), Introduction (Line 88-91), Discussion (Line 231-233, the whole paragraph was removed here), and the References section.

      Line 176-178 - This is a really cool finding.

      Thank you for this positive comment.

      Lines 179-181 - You didn't test competition between plant species, and you didn't measure light, soil moisture, or soil nutrients. So you can suggest competition as a potential mechanism, but you can't say definitively that's what is happening.

      We now have lowered our tone here as “We speculate that these improvements in food availability and nutrition for yaks may be due to the release of grasses and sedges from competition with the forbs for limiting above- and below-ground resources”.

      See the revision in Line 208-211 in the Discussion section.

      Lines 184-186 - You did a good job of it here, suggesting a likely potential mechanism at play without claiming it is for sure happening when it hasn't been measured.

      Thank you for this positive comment!

      Lines 188-199 - Strong paragraph. The impacts of large herbivores on smaller animals are well-studied, but reciprocal impacts are often overlooked.

      Thank you for this positive comment!

      Lines 201-205 - You didn't test the stress gradient hypothesis because there was no abiotic gradient. You also did not take any measurements during outbreaks, so you cannot claim to have compared low-moderate to outbreak pika densities. I think the paper would be much stronger if you removed the stress gradient hypothesis and instead focused more on the literature around facilitation between mammals of different body sizes, as you do in lines 205-209.

      We agreed that our work didn’t specifically design to test the stress gradient hypothesis (SGH) between pikas and yaks, so we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Line 209-215 - Again, you didn't test a competition-facilitation balance because you never tested or demonstrated competition. One of the main strengths of this paper is demonstrating facilitation, so build on that strength rather than referencing things you didn't measure.

      The same as mentioned above, we have deleted this paragraph here, see Line 231-233 in the Discussion section.

      Lines 218-222 - Not an accurate description of the relationship between herbivore diet and body size. Larger herbivores typically tolerate lower-quality plants in order to consume sufficient calories, but plenty of them do this via mixed feeding. Grazing in large herbivores and livestock is usually due to specifics of the digestive tract (e.g., hindgut fermentation) rather than specifically about body size.

      We have deleted this description of the relationship between herbivore diet and body size here. Instead, we have modified this statement as “The coexistence of a diverse of herbivore species with different diet selections and size classes can lead to an “compensatory effect” on grass and forb biomass that helps to maintain a balance and diverse plant community” in the Discussion section.

      See these revisions in Line 234-237 in the Discussion section.

      Lines 245-251 - Paragraph addresses an important point. Lines 247-249, though, overstate what you measured. There's no measurement of livestock production or biodiversity in the study.

      We have replaced the term of “livestock production and biodiversity” with “livestock growth performance” here. See Line 288-298 in the Discussion section.

      (5) Figures

      Figure 2C-D - This applies to all figures, but you need to plot the best-fit line generated from your model instead of using geom_smooth, which is what these lines look like. You can do this using functions like predict or ggpredict. Using these smoothed lines implied non-linear relationships that you didn't actually test for.

      We have redrawn Figures 2C and 2D to use model estimates directly and have included the related model summaries as Appendix 1—table 3 and table 4 in the Appendix 1.

      Figure 3B - This figure doesn't show yak weight gain in the presence of pikas. Instead, it shows weight loss when pikas are absent. It's a subtle difference, but very important for interpretation. Yak can maintain weight just fine without pikas as long as Stellera are absent, too. Your results consistently show no Pika x Stellera interactions, but that doesn't match your significance values here. Need to double-check and explain that.

      We have revised the descriptions for Figure 3 and 4 in the Results section, by emphasizing that the absence of pikas REDUCED weight gains of yaks, INCREASED toxic plant abundance, and REDUCED the quantity and quality of palatable grasses and sedges. We have revised these sections as below:

      Abstract (see Line 51-53)

      “Compared to the pika-present treatment, pika removal dramatically increased cover of the poisonous Stellera forbs by two-fold, reducing the abundance and protein content of palatable grasses and sedges, yak foraging efficiency, and yak weight gain by up to 42%.”

      Results (see Line 153-167, Line 169-175, Line 184-192)

      Also, we did find significant Pika x Stellera interactions for yak weight gains, we have provided these details in the Appendix 1—table 5 and table 6 in the Appendix 1.

      (6) Recommendations

      Lines 245-257 - Would recommend combining the last two paragraphs into one.

      We have combined the last two paragraphs, see Line 288-298 in the Discussion section.

      Line 530 - Can you replace large with a more precise measure of area?

      We are unable to provide a more precise measure of area here, so we have deleted the description of “in a large area”, but we have also added the note of “in the study site” by the end of the sentence to better describe the location of the plots.

      See the revision in Line 413 in the Methods section.

      Line 541 - Does this mean +/- 7.8 standard deviations? If so, how big is that range in kg?

      Here should be “115±7.8 kg”, we have done this correction in Line 392 in the Methods section.

      Figure 3 - Would be helpful to use "Stellera/pika present" and "Stellera/pika absent" rather than saying "No Stellera/pika" since you also use the No. abbreviation for numbers a lot in this figure.

      We have revised all the related Figures for this issue in Figure 3, 4, and Appendix 1—figure 1,2.

      Figure 3F and I - Show letters for significance on these two plots as well, even if it is just a row of a's.

      We have redrawn Figure 3F,I to address this issue.

      Figure S1 - Applies to all boxplots. Be consistent about showing significance, even if it is a row of a's indicating no difference between treatments.

      We have re-drawn Appendix 1—figure 1,2 to address this issue.

      Table S1 - Something happened with the line numbers, so they are in the table instead of on the left side.

      We have fixed this problem for Appendix 1—table 1 in the Appendix 1.

      Table S1 - Applies to all tables. Include the type of model that you ran (including distribution if not Gaussian) in the table legend.

      We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

      Table S9 - This legend has a good description, including the type of model you used and what you were testing. Apply this more detailed legend to the rest of the tables.

      Again! We have provided an summary of all statistical models used in the study in Appendix 1—table 13 in the Appendix 1.

      Also, we have added a sentence of “A summary of all statistical models used in the study is available in Appendix 1—table 13.” in Line 475-476 in the Statistical analyses section to indicate this information.

    1. eLife Assessment

      This important study provides new insights into the regulation of cell organization and division in Trypanosoma brucei through the phosphorylation-dependent control of a kinesin motor protein by a polo-like kinase. The authors present convincing evidence, combining rigorous biochemical, cell biological, and imaging analyses, demonstrating that phosphorylation modulates kinesin localization and function, thereby influencing cellular organization and cytokinesis. The findings advance our understanding of the molecular mechanisms governing trypanosome cell division and will be of broad interest to researchers studying trypanosomes, cytoskeletal regulation, and eukaryotic cell division.

    2. Reviewer #1 (Public review):

      I thank the authors for the revised manuscript and for the detailed responses.

      I think the main points raised in the review have now been addressed. In particular, the new experiment with TbPLK inhibition and mass spectrometry is an important addition, as it provides direct evidence that phosphorylation of KIN-G at Thr301 and Ser569 depends on TbPLK activity in cells.

      I also appreciate that the authors have toned down the interpretation of the Golgi phenotype. The revised text now makes clear that the fluorescence data show altered Golgi/ERES organization or duplication, but do not prove a structural defect in Golgi biogenesis.

      The added discussion of the T301A result is also helpful. The finding that only a small fraction of KIN-G is phosphorylated at Thr301 in asynchronous cells makes the lack of a strong T301A phenotype more understandable.

      Overall, I am happy with the revision of the beautiful manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid, and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei support the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?<br /> a. The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.<br /> b. Published work links PLK to cell division, FAZ elongation, etc... The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc....<br /> c. Some experiments or at least commentary on points a and b above would strengthen the paper.<br /> - The authors have now addressed this question by assessing what % of KING is phosphorylated at T301 and adding commentary on this point in the revised paper.<br /> - I would suggest that the model (new figure 8) include a dephosphorylation step, as that is proposed by the authors in the text. Also include in the legend some commentary on the role of phosphorylation, which is the center point of this paper, but not currently mentioned.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?<br /> a. The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.<br /> - The authors have addressed this question by demonstrating that T301 phosphorylation is reduced upon treatment with a PLK inhibitor, thus supporting that PLK phosphorylated T301 in vivo. It is noted that one might consider an alternate kinase is also able to phosphorylate T301 in absence of PLK activity, as that could explain the relatively low (~27%) reduction in phosphorylation by PLK inhibitor treatment.

    4. Reviewer #3 (Public review):

      Summary:

      Here the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. They further confirm that inhibition of TbPLK in cells reduces KIN-G phosphorylation at T301 (and S569). Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during early S-phase of the cell cycle. Centin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescene to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      The authors have addressed prior weaknesses in the manuscript through additional experimentation and rewording of the conclusions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides new insight into the regulation of cell organization and division in Trypanosoma brucei through the control of a kinesin motor protein by a polo-like kinase. The authors present solid evidence from rigorous biochemical and imaging analyses showing that phosphorylation modulates kinesin function and cellular organization. However, direct in vivo evidence that PLK phosphorylates kinesin-G is lacking.

      We performed experiments to investigate the effect of PLK inhibition on the phosphorylation of KIN-G in vivo in trypanosome cells by immunoprecipitation and mass spectrometry. The new results showed that treatment of trypanosome cells with GW843682X, a potent TbPLK inhibitor validated previously in procyclic trypanosomes, reduced the phosphorylation levels on Thr301 and Ser569 of KIN-G by ~27% and 100%, respectively. The partial reduction in Thr301 phosphorylation after GW843682X treatment could be attributed to slower dephosphorylation of phosphorylated Thr301 after GW843682X was added to the cell culture. Nonetheless, these new results demonstrated that KIN-G is an in vivo substrate of PLK.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript identifies the orphan kinesin KIN-G as a substrate of Polo-like kinase (TbPLK) in Trypanosoma brucei and demonstrates that phosphorylation of Thr301 inhibits KIN-G microtubule binding and disrupts its cellular function. Using a combination of in vitro kinase assays, phosphosite mapping, microtubule binding and gliding assays, and in vivo complementation with phosphomimetic and phosphodeficient mutants, the authors link TbPLK-mediated regulation of KIN-G to defects in centrin arm integrity, FAZ elongation, Golgi organization, flagellum positioning, and division plane placement. The study provides a mechanistic advance in understanding how TbPLK regulates centrin arm biogenesis and integrates KIN-G into the growing regulatory network controlling hook complex and FAZ assembly. Overall, the work is technically strong, internally consistent, and builds logically on previous studies from this group and others.

      Strengths:

      A major strength of the manuscript is the clear mechanistic link between phosphoryltion of Thr301 and loss of microtubule binding activity. The use of phosphomimetic (T301D) and phosphodeficient (T301A) mutants in an RNAi-rescue framework provides a clean and convincing demonstration of functional relevance in vivo. The integration of biochemical assays with detailed cell biological phenotyping (centrin arm length, FAZ elongation, basal body segregation, and cytokinesis markers) is particularly effective and makes the central conclusion robust. The observed phenotypic cascade from centrin arm defects to FAZ and division plane abnormalities is also well aligned with existing models of trypanosome morphogenesis.

      Weaknesses:

      My (more or less main) concern relates to the interpretation of the Golgi phenotype. The conclusion that phosphorylation of KIN-G "impairs Golgi biogenesis" is currently based on fluorescence microscopy using TbGRASP and Sec13 markers and on quantification of the number and distribution of Golgi/ERES puncta in binucleated cells. While these data convincingly demonstrate altered Golgi/ERES number and spatial organization, they do not distinguish between true defects in Golgi biogenesis or duplication and alternative possibilities such as fragmentation, vesiculation, or mislocalization of Golgi membranes. Given the central role of Golgi-centrin arm organization in the proposed model, ultrastructural analysis (for example, by EM or electron tomography) would greatly strengthen this aspect of the study by providing direct evidence for structural alterations of the Golgi and its association with the centrin arm and ERES. Such data would elevate this part of the manuscript from a descriptive fluorescence phenotype to a true structural cell biological insight. I appreciate that this experiment goes beyond the current dataset, but it would substantially enhance the mechanistic depth of the Golgi-related conclusions and strengthen the causal chain linking centrin arm defects to Golgi abnormalities. However, I have to confess, the inclusion of such data would make this reviewer particularly enthusiastic about the work. If this is not feasible, I would recommend tempering the wording of "Golgi biogenesis" to a more conservative description, such as altered Golgi organization or duplication, and explicitly acknowledging the limitations of fluorescence-based analysis for this conclusion.

      Thanks for these very constructive comments, which are very well taken. We totally agree with this reviewer on these points. Since it is not feasible for us to perform EM or electron tomography, we have revised the manuscript to describe the effect of KIN-G phosphorylation on the Golgi as “altered Golgi duplication” rather than “Golgi biogenesis”. We also explicitly acknowledge the limitations of fluorescence-based analysis of the Golgi for this conclusion and suggest that further characterization with EM or electron tomography would allow one to reveal the potential structural alterations of the Golgi and its association with the centrin arm.

      An additional conceptual point concerns the dual role of TbPLK in centrin arm regulation. TbPLK is known to promote centrin arm biogenesis through phosphorylation of TbCentrin2, yet in this study, TbPLK phosphorylation of KIN-G negatively regulates centrin arm assembly. This dual positive and negative regulatory role is intriguing but could be discussed more explicitly. The manuscript would benefit from a clearer conceptual framework addressing how phosphorylation of KIN-G might serve as a temporal or spatial switch to restrain KIN-G activity at specific stages of centrin arm assembly.

      This is a great point. However, we are not sure whether the previous work on TbPLK phosphorylation of TbCentrin2 could lead to the conclusion that TbPLK promotes centrin arm biogenesis through phosphorylation of TbCentrin2. In the published work (de Graffenried et al., MBoC, 2013), trypanosome cells expressing the phospho-deficient mutant TbCentrin2-S54A only showed minor growth defects, exhibiting growth defects after 5 days (de Graffenried et al., MBoC, 2013). Cells expressing the phosphomimic mutant TbCentrin2S54D, however, showed very strong growth defects (de Graffenried et al., MBoC, 2013). The effects of TbPLK phosphorylation on TbCentrin2 appear to be quite similar to that of TbPLK phosphorylation on KIN-G, although the KIN-G-T301A mutant does not have growth defects (up to 5 days in our experiments). It appears that the primary role of TbPLK in regulating TbCentrin2 and KIN-G is negative regulation. Nonetheless, we have added more discussion on these regulatory roles of TbPLK in the revised manuscript.

      Finally, a schematic model summarizing the proposed regulatory pathway from TbPLK phosphorylation of KIN-G to centrin arm assembly, FAZ elongation, division plane placement, and Golgi organization would aid the reader.

      Thanks for this suggestion. We made a schematic model to summarize the roles of KIN-G and its regulation by TbPLK. This is included in Figure 8.

      Reviewer #2 (Public review):

      Summary:

      The authors identify KIN-G as an in vitro substrate for phosphorylation by TbPLK and show that several of the in vitro P-ated sites, including T310, overlap with P-ation sites seen in live cells. The authors further show that PLK-mediated P-ation inhibits KIN-G binding to microtubules in vitro, as does a KIN-G-T301D mutant, and that expression of a KIN-G-T301D Phospho-mimic in T. brucei phenocopies KIN-G RNAi knockdowns, producing defects in cell division, morphogenesis of the centrin arm, FAZ and other cellular structures, as well as a misplaced cytokinesis furrow.

      Understanding cytoskeletal rearrangements that drive cell division in T. brucei is an important and unresolved problem, so the work addresses important questions that are of great interest. PLK and KIN-G have previously been shown to be important for cell division and morphogenesis of cytoskeletal structures that drive cell division in T. brucei. The current work advances our understanding by suggesting a potential mechanism by which PLK and KIN-G might participate, namely through PLK-dependent P-ation to control KIN-G MT binding activity.

      Strengths:

      The authors use a rigorous combination of biochemistry, phosphoproteomics, cell biology, and mutant analysis to support their conclusion that PLK-mediated P-ation of KIN-G negatively regulates KIN-G microtubule binding, and this may explain the observation that a KIN-G T301 phosphomimic mutant blocks cell division and perturbs biogenesis of cytoskeletal structures that drive cell division and morphogenesis. Combining rigorous and informative in vitro studies with mutant analysis in live cells is a great strength. The work is solid and important, though a few pieces are needed to fully connect the in vitro findings with the in vivo observations, as detailed below.

      Weaknesses:

      Overall, I find this work to be solid and to provide an important addition to our understanding of mechanisms controlling cell division in T. brucei. The biochemistry, in particular, is rigorous and convincingly demonstrates PLK can P-ate KIN-G, altering its MT-binding ability. Analysis of phospho-mutants of KIN-G in live T. brucei supports the conclusion that P-ation of KIN-G at T301 negatively affects KIN-G function in vivo. I think, however, that the results fall short of supporting the title, because, although the data convincingly show that PLK can phosphorylate KIN-G at T301 in vitro, and that T301 is P-ated in vivo, they do not formally demonstrate (nor even test) whether PLK is the kinase responsible for this phosphorylation in vivo (experiments to address this seem quite feasible). I also do not see where the authors try to reconcile the absence of phenotype for KIN-G-T301A with the implied importance of KIN-G phosphorylation by PLK in cell division, which calls into question the need for P-ation of KIN-G-T301 in cell division. Suggestions for addressing these concerns are provided below.

      My two main questions are:

      (1) What is the biological relevance of KIN-G P-ation at T301?

      (a) The authors report no defect for the KIN-G-T301A mutant, so what then is the need for T301 P-ation, if the cell gets along fine without it? One step toward addressing this would be to ask what fraction of KIN-G shows P-ation at T301. Although published studies indicate P-ation at T301, it isn't known what percentage of KIN-G in the cell is P-ated. One might anticipate, for example, that T301-P is a small minority of the population in asynchronous cultures and that T301 P-ation increases at specific cell cycle stages.

      This is a great point that is very well taken. We also had been puzzled by the observation of no growth defects of T301A mutant. This comment enlightened us. From the new experiments we performed to compare the phosphorylated peptides of KIN-G in cells treated and non-treated with the PLK inhibitor GW843682X, we calculated the percentage of phosphorylated Thr301 in non-GW843682X-treated cells. We found that the percentage of peptides containing the phosphorylated Thr301 is ~14% of the total Thr301-containing peptides (Fig. 1H). This result indicates that T301-P is indeed a small minority of the population in the asynchronous trypanosome cells. It is possible that T301 phosphorylation may occur at a specific cell cycle stage such as early S-phase, during which PLK and KIN-G co-localize at the centrin arm.

      (b) Published work links PLK to cell division, FAZ elongation, etc.. The current work suggests that one role of PLK is to P-ate KIN-G at T301. In contrast, however, the current work also indicates that P-ation of KIN-G at T301 is unnecessary for normal cell division, FAZ elongation, etc..

      Yes, previous work discovered essential roles of TbPLK in basal body segregation, centrin arm biogenesis, FAZ elongation, and cytokinesis. These functions of TbPLK correlate with TbPLK’s localization to multiple subcellular structures, the basal body, the centrin arm, and the new FAZ tip, and are attributed to the regulation of its substrates at these structures. At the basal body, TbPLK phosphorylates SPBB1, which is required for basal body segregation. Defects in basal body segregation can lead to defective flagellum positioning and FAZ elongation. At the new FAZ tip, TbPLK regulates the cytokinesis regulator CIF1, which is required for cytokinesis. At the centrin arm, TbPLK phosphorylates TbCentrin2 at S54 and KIN-G at T301 (and S569, which was newly identified as an in vivo TbPLK site and has not yet been characterized). However, cells expressing TbCentrin2-S54A have very weak growth defects, and cells expressing KIN-G-T301A have no detectable growth defects. In contrast, cells expressing TbCentrin2-S54D and cells expressing KIN-G-T301D have strong growth defects. Therefore, the essential role of TbPLK in centrin arm biogenesis apparently is not attributed to the phosphorylation of TbCentrin2 and KIN-G. It is possible that phosphorylation of other centrin arm-localized protein(s) by TbPLK may be essential for centrin arm biogenesis, but this possibility remains to be explored.

      (c) Some experiments or at least commentary on points a and b above would strengthen the paper.

      We performed experiments and presented the data in Fig. 1H. We also included commentary in the revised manuscript on the points about the potential role of T301 phosphorylation. Thanks for these great comments that significantly improved the manuscript.

      (2) Is PLK the kinase that P-ates Kin-G T301 in vivo?

      (a) The authors show PLK P-ates T301 (and other residues) in vitro, and that T-301 is P-ated in vivo. To bring the analysis full circle, it would be informative to examine KIN-G P-ation in a PLK mutant or upon inhibition of PLK with published inhibitors. This seems to be a very doable experiment with the tools available.

      We treated trypanosome cells with a potent PLK inhibitor GW843682X, which was previously demonstrated to inhibit TbPLK activity in vitro and mimic TbPLK knockdown in vivo in trypanosomes, and immunoprecipitated KIN-G for mass spectrometry. We compared the KIN-G peptides identified by mass spectrometry from trypanosome cells treated with or without GW843682X, and found that two phosphosites (T301 and S569) were reduced by ~27% and 100%, respectively, after GW843682X treatment (Fig. 1G). These results provided evidence to support that TbPLK phosphorylates KIN-G in vivo.

      Reviewer #3 (Public review):

      Summary:

      Here, the authors investigate the role of the Trypanosoma brucei polo-like kinase TbPLK in the function of flagellum-associated cellular structures in trypanosomes. They set out to test the hypothesis that a key substrate of TbPLK is the kinesin protein KIN-G, and that TbPLK phosphorylation of KIN-G regulates its functions in cells.

      Strengths:

      Using in vitro biochemistry with purified proteins, the authors convincingly demonstrate that TbPLK phosphorylates KIN-G at 29 sites. Moreover, they convincingly show that phosphorylation at one site, T301, impairs the binding of purified KIN-G to purified microtubules. Using immunofluorescence-based imaging approaches, they also show that TbPLK colocalizes with KIN-G at centrin arms during the early S-phase of the cell cycle. Centrin arms are structures that are located near the basal body and flagellum and are important for new flagellum biogenesis, Golgi positioning, and cell division. To evaluate the function of KIN-G phosphorylation in cells, they depleted KIN-G by RNAi, simultaneously expressed phospho-mimetic (T301D) and phospho-ablative mutant proteins, and used immunofluorescence to examine the impact on flagellum-associated cellular structures. They show that expression of the phospho-mimetic mutant KIN-G-T301D causes the following defects: reduced cell proliferation, disruption of centrin arm and Golgi biogenesis, impairment of FAZ elongation and flagellum positioning, and misplacement of the cell division plane. The data convincingly support the conclusion that KIN-G phosphorylation on T301 plays an important role in regulating the cellular functions of this kinesin motor protein.

      Weaknesses:

      Some of the broader conclusions are not directly supported by the data. For example, the title states "Polo-like kinase phosphorylation of the orphan kinesin KIN-G negatively regulates centrin arm biogenesis in Trypanosoma brucei," but the data do not directly address the specific role of TbPLK in phosphorylating KIN-G in cells. Moreover, some of the more specific conclusions in the paper, for example, that "phosphorylation of KIN-G" causes various cellular defects, are a bit of an overstatement. The supporting data rely on the expression of a phospho-mimetic mutant of KIN-G. Presumably, phosphorylation in cells is a normal part of KIN-G regulation, and it is not just phosphorylation, but rather hyperphosphorylation that is being mimicked by the mutant. Some rewording of the specific conclusions is warranted, and the broader conclusion would be better supported with additional experimental evidence.

      This is a great point that is very well taken. We performed new experiments to address the in vivo phosphorylation of KIN-G by TbPLK. We treated cells with a potent PLK inhibitor, GW843682X, and then immunoprecipitated KIN-G for mass spectrometry to identify changes in phosphorylation. We found that the phosphorylation levels of T301 and S569 were reduced by ~27% and 100%, respectively, confirming that these two sites are in vivo TbPLK phosphosites.

      We also calculated the ratio of phospho-T301 versus non-phospho-T301 in non-treated cells and found that phospho-T301 accounts for ~14% of the total KIN-G protein. This new result suggests that it is the hyperphosphorylation that causes growth defects. We have revised the manuscript accordingly to reflect this point.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Several statements use rather strong causal language (for example, "thereby impairing Golgi biogenesis, FAZ elongation, and division plane placement"). While the phenotypic correlations are convincing, direct causality is largely inferred from prior literature. Slightly tempering this wording could improve precision. It would also be helpful to state explicitly in figure legends the number of cells analyzed per condition and the number of independent experiments for each quantified phenotype.

      Thanks for these constructive comments, which we agree and appreciate greatly. We have revised the manuscript accordingly to improve precision. The total number of cells analyzed, and the number of independent experiments were included in the figure legends.

      Reviewer #2 (Recommendations for the authors):

      Minor comments for improving the text are:

      (1) The paper overall is clearly written. However, the Discussion starts with a solid sentence, then becomes a bit diffuse in discussing a wide range of PLK activities that were not addressed in the current work. That detracts attention a bit from the central contributions of this paper.

      Thanks for this comment. We have deleted the discussion about TbPLK activities that were published previously.

      (2) At least two places in the text state apparent contradictions.

      (a) p.5 and Figure 2C. The authors say microtubule gliding speed was "...insignificantly reduced..." by the TbPLK-K70R mutant, yet they then state that motility was "interfered with". If the effect is "insignificant", why do they claim there is an effect?

      (b) p6 and Figure 3C. The authors report KIN-G-T301A impact on microtubule gliding activity is insignificant, but then say this mutation reduces the motility of KIN-G. These statements are contradictory.

      We meant to say that there was a slight but insignificant effect. We agree that such statements are somewhat contradictory and, hence, have been deleted. Thanks.

      (3) p. 8, and Figure 7. "ventral side" and "leading edge" are not defined but are used to describe the KIN-G RNAi phenotype.

      We have deleted the wording “ventral side”, as it is not necessary. Thanks.

      (4) Figure 7B. Please explain the labeling - the new flagellum daughter is indicated as having the old posterior, while the old flagellum daughter cell is indicated as having the new cell posterior. This is counterintuitive to a reader not intimately familiar with the T. brucei cell division process.

      Thanks very much. We included two sentences in the revised manuscript to explain this point.

      The sentences read as follows: “The nascent posterior is formed near the mid-portion of the NFD cell through microtubule bundling and cytoskeleton remodeling during late stages of the cell cycle (Wheeler et al., 2013). Consequently, the NFD cell inherits the old, existing cell posterior, whereas the OFD cell inherits the newly formed or nascent cell posterior.”

      (5) Figure 4, 5, and 7: "% Cells" is reported. Please indicate the total number of cells that were examined.

      The total number of cells were included in the figure legends.

      Reviewer #3 (Recommendations for the authors):

      (1) The manuscript should be carefully edited for minor grammatical errors.

      Thanks. We have carefully proofread the manuscript and corrected the grammatical errors.

      (2) A general conclusion is that TbPLK phosphorylation of KIN-G in cells is critical for regulating its motor activity. However, this relies on the expression of phospho-mimetic mutants, which bypass TbPLK. Thus, there really is no direct evidence provided to support the specific role of TbPLK other than the in vitro phosphorylation data. Some additional experiments to assess the specific role of TbPLK in phosphorylating KIN-G in cells would lend support for the general conclusion. Is it possible to deplete or inhibit TbPLK and show that this impacts the phosphorylation of KIN-G in cells?

      This is a great point that is very well taken. We performed a new experiment by inhibiting TbPLK with a potent PLK inhibitor, GW843682X, and then immunoprecipitating KIN-G for mass spectrometry to identify the phosphorylation levels before and after GW843682X treatment. We were able to confirm that T301 and S569 phosphorylation was reduced after treatment. This confirms that TbPLK phosphorylates T301 and S569 of KIN-G in vivo in trypanosome cells.

      Specific:

      (1) Figure 2: For microtubule gliding assays, representative videos should be included as supplementary data. Also, when the data do not show a significant difference between KIN-G and KIN-G + TbPLD-K70R, then the authors should not state there is a slight difference, as this is not supported by the data.

      We have deleted the statement. We have included the representative videos for all the experiments presented in Figures 2 and 3. Thanks.

      (2) Figure 3: As stated above, for microtubule gliding assays, representative videos should be included. Moreover, for KIN-G-T301A, the authors say that the gliding activity was "moderately, but insignificantly, reduced." Again, if the difference is not significant, it cannot be concluded that there is a difference compared with the wild-type protein.

      We have deleted the statement. Thanks.

      (3) Some of the headings in the Results section are not accurate. Regarding the in vivo results, the heading "Phosphorylation of Thr301 in KIN-G by TbPLK causes defective cell proliferation" seems to be an overstatement. It may be more accurate to state that "Expression of a phospho-mimetic mutant of KIN-G causes defective cell proliferation." The same comment applies to the other headings that follow this one. Presumably, there is a population of phosphorylated KIN-G in cells, and phosphorylation/dephosphorylation is a normal part of its regulation. In the text, the authors might more accurately conclude that hyperphosphorylation causes the defects they are seeing.

      This is a great point. We agree and as we responded above, we have revised the manuscript, from the title to the main text, to reflect the point that hyperphosphorylation of T301 by TbPLK causes the defects. Thanks very much for these great comments!

    1. eLife Assessment

      This important study demonstrates that nutrient resorption efficiency in the widespread wetland grass Phragmites australis is largely canalized by phylogeographic lineage, ecotype and geographic origin, rather than responding plastically to short-term salt stress. The findings have implications beyond plant ecophysiology because they suggest that predictions of wetland nutrient cycling under increasing salinization should account for the genetic composition of plant populations. The evidence supporting the central conclusions is compelling, based on a well-designed common-garden experiment involving 110 genotypes, paired control and salinity treatments, and convergent metabolomic, ionomic and whole-plant evidence confirming that the treatment imposed substantial physiological stress. The interpretation is nevertheless limited to one growing season and a single, moderate salinity treatment, leaving open whether chronic, more severe or multigenerational exposure might induce plastic or transgenerational responses; further clarification of the relationships among phylogeographic lineage, ecotype and latitude, and a more cautious interpretation of the variance explained by latitude, would strengthen the manuscript.

    2. Reviewer #1 (Public review):

      Summary:

      This study demonstrates that nutrient resorption efficiency (NuRE) in Phragmites australis is genetically canalized rather than plastic to salt stress. Using 110 genotypes in a common garden, the authors show that intraspecific variation in NuRE is explained by phylogeographic lineage, ecotype, and latitude, not by effective salinity. Element-specific regulatory strategies further reveal how N, P, and K resorption are differentially controlled. At the population level, this is an important study that fundamentally advances our understanding of plant functional trait evolution and its implications for ecosystem nutrient dynamics under global change.

      Strengths:

      This study is the first to demonstrate genetic determination of a key nutrient conservation trait under effective salt stress in a widespread macrophyte, directly testing the 'plastic acclimation versus inherent conservatism' paradigm in a non-nutrient stress context. The experimental design is rigorous: each genotype was paired across control and salt treatments, and multilevel stress effectiveness (metabolomics, biomass, Na accumulation) was confirmed before evaluating NuRE. The large sample size of a macrophyte and dual classification (phylogeography + ecotype) allow robust disentangling of genetic versus plastic sources of variation.

      The analysis comprehensively tests three resorption control hypotheses using appropriate SMA regression, revealing element-specific and condition-dependent patterns. The latitudinal gradient and variation partitioning provide strong evidence that genetic origin and geographic context outweigh short-term plasticity, with important implications for predicting ecosystem nutrient cycling under global change. This study provides a clear empirical demonstration that a key nutrient conservation trait can remain homeostatic under non-nutrient stress, and that intraspecific variation is primarily a product of population differentiation rather than short-term plasticity.

      Weaknesses:

      First, the salinity treatment spanned only one growing season. The conclusion of genetic canalization therefore specifically refers to the absence of plasticity to an acute salt shock. Whether long-term, multigenerational chronic salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question for further research. Likewise, the physiological mechanisms underlying the observed lack of plastic increase in NuRE (for example, phloem loading or senescence gene expression) are not directly resolved, leaving some inference about trade-offs versus true unresponsiveness. These points do not weaken the study's main conclusion. Instead, they suggest productive future directions, such as longer-term field manipulations and targeted molecular investigations.

      Second, the test of nutrient limitation control relies on resorbed N:P and N:K ratios as proxies, an established but indirect approach. Direct nutrient addition experiments would provide stronger causal evidence. Also, the metabolomic analysis is used primarily to validate stress effectiveness; deeper integration of specific metabolites with NuRE variation across genotypes could have offered mechanistic insights but was not pursued. Additionally, the potential collinearity between ecotype and phylogeographic lineage among Chinese populations is not quantitatively addressed. None of these considerations undermines the main finding, which is supported by a robust experimental design and widely accepted analytical approaches.

    3. Reviewer #2 (Public review):

      Summary:

      The study finds that nutrient resorption efficiency in Phragmites australis shows no plastic response to salinity stress but is canalized by phylogeographic lineage, ecotype, and latitude. In a common garden with 110 genotypes, salinity induced stress, yet no plastic change occurred for N, P, or K resorption. The authors conclude that intraspecific variation is historical and geographic; thus, predictions of wetland nutrient cycling need to account for phylogeographic composition.

      Strengths:

      The core finding that NuRE shows no plastic response to salinity, but is instead evolutionarily canalized by lineage and latitude, challenges a key assumption of broad trait plasticity. This conclusion is firmly supported by a robust common garden design with 110 genotypes, rigorous multi-level stress validation, and element-specific resorption analyses. The work provides compelling evidence that intraspecific variation in this critical nutrient cycling trait is shaped by phylogeographic history rather than short-term acclimation. The implications for predicting wetland responses to salinization are significant, as ecosystem-level nutrient dynamics may be constrained by the genetic composition of plant populations.

      Weaknesses:

      The experiment covers only one growing season, with salinity applied in June and measurements in December. While the stress is clearly effective, longer-term or multi-year stress might reveal acclimation or epigenetic effects that are not captured. Given the author team's expertise in parental and transgenerational effects in clonal plants, this limitation is particularly relevant and warrants more thorough discussion in the manuscript.

      The salinity treatment uses a single moderate level of 10 ppt, which does not allow assessment of whether more extreme stress might trigger a plastic response. A dose-response design across a gradient would have provided stronger inference about the threshold at which NuRE canalization might be overcome. Additionally, the ecotype analysis in Figure 4 applies only to Chinese populations, as classification was not available for non-Chinese populations, which should be stated more explicitly in the Results.

      The variation partitioning shows latitude as a significant predictor, but the R² values are relatively low, indicating that much variance remains unexplained. The manuscript should avoid overinterpreting latitude's explanatory power and more openly acknowledge the role of unmeasured factors. The interpretation of slopes greater than 1 for the resorbed N:P versus green N:P relationship, labeled as "inverted limitation", also needs further explanation regarding its functional significance.

    4. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback. We are grateful that the experimental design and the evidence for genetically canalized nutrient resorption efficiency (NuRE) were recognized as compelling and important, with clear implications for predicting wetland nutrient cycling under salinization. We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them, with particular attention to the weaknesses noted.

      Reviewer #1 raised two important concerns regarding the scope and generality of our conclusions. First, the salinity treatment spanned only one growing season, so the genetic canalization we document specifically refers to the absence of a plastic response to an acute salt shock; whether long-term, chronic or multigenerational salinity could act as a selective agent or induce transgenerational plasticity remains an open and interesting question. Second, the test of nutrient limitation control relies on resorbed N:P and N: K ratios as proxies rather than direct nutrient manipulation, and the metabolomic analysis is used primarily to validate stress effectiveness. We accept these criticisms and will address them as follows.

      Regarding the temporal scope of the treatment, we will explicitly state in the Discussion that our conclusion of canalization pertains to short-term acclimation to an acute salt shock, and we will discuss the scenarios under which chronic, more severe or multigenerational exposure could trigger plastic, acclimatory or transgenerational responses. Given our team’s prior work on parental and transgenerational effects in clonal plants, we will frame these as testable hypotheses for future research. We will also acknowledge that the physiological mechanisms underlying the lack of a plastic increase in NuRE (for example, phloem loading or senescence-associated gene expression) are not directly resolved in the present study and will propose targeted molecular investigations as a natural next step.

      Regarding the nutrient limitation tests, we will clearly acknowledge in the Discussion that the resorbed N:P and N:K ratios provide an established but indirect proxy for nutrient limitation, and we will discuss how direct nutrient addition experiments could provide stronger causal evidence, while deepening the integration of the metabolomic profiles with genotype-level NuRE variation where feasible. We will also quantitatively assess the potential collinearity between ecotype and phylogeographic lineage among the Chinese populations.

      Reviewer #2 raised three substantive issues regarding the interpretation and presentation of our results. First, although latitude emerged as a significant predictor in the variation partitioning, the relatively low R² values indicate that much variance remains unexplained, and our interpretation should be more cautious; in addition, the functional significance of the “inverted” nutrient limitation (slopes greater than 1 for resorbed versus green N:P) needs further explanation. Second, the single moderate salinity level (10 ppt) does not allow assessment of whether more extreme stress might trigger a plastic response. Third, the ecotype analysis applies only to the Chinese populations, as ecotype classification was not available for non-Chinese populations, and this should be stated explicitly in the Results. We fully agree with these points and will revise accordingly.

      To address the concern about latitude, we will temper the language in the Results and Discussion, explicitly noting the limited proportion of variance explained by latitude and acknowledging the role of unmeasured factors, while retaining the study’s central message that genetic and phylogeographic origin outweigh short-term plasticity in shaping NuRE. We will also expand the Discussion to explain the functional significance of the “inverted” nutrient limitation, that is, why P and K are resorbed more completely relative to N, whether as a strategy to maintain optimal N:P:K ratios or a reflection of the higher costs and lower availability of N.

      To address the salinity gradient concern, we will acknowledge in the Discussion that the single moderate salinity level may not have been severe enough to trigger a plastic response and will justify future dose-response experiments to identify the threshold at which canalization of NuRE might be overcome.

      To address the ecotype limitation, we will explicitly state in the Results that the ecotype analysis in Figure 4 is based only on the Chinese populations, for which ecotype classification was available.

    1. eLife Assessment

      This important study introduces an open-source software package that improves the computational efficiency of whole-brain modelling and facilitates individualised model fitting in large neuroimaging cohorts. The evidence is convincing for the main computational and practical claims, supported by evaluations across multiple optimisation strategies, speed benchmarks, and demonstrations using human neuroimaging data. Its modular design and documentation should make it a resource that is of value to researchers interested in scalable and individualised brain network modelling.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript introduces cuBNM, a GPU‑accelerated Python package for whole‑brain modeling. The authors demonstrate that running simulations on GPUs provides substantial benefits in computational speed, cost-efficiency, and scalability compared to traditionally used CPUs, making large‑scale and individualized brain network modeling computationally feasible. The usage of cuBNM has been demonstrated by running optimization of group-level and individualized low- and high-dimensional models. By investigating the test-retest reliability and heritability of simulated and empirical measures in the Human Connectome Project dataset, the authors showed that simulated features were fairly reliable and significantly heritable.

      Strengths:

      This study is timely and presents an important contribution to the field of whole-brain computational modeling. A major strength is that the authors go beyond introducing a GPU-accelerated framework by demonstrating its utility through comprehensive benchmarking and biologically relevant applications, including individualized model fitting, comparisons of homogeneous and heterogeneous models, and analyses of test-retest reliability and heritability.

      The computational performance is evaluated comprehensively, assessing speed, computational cost, energy consumption, and scalability across different simulation settings. The Human Connectome Project dataset is used to demonstrate that the software enables individualized whole-brain modeling in large datasets.

      Finally, the software is modular, open-source, and well-documented, and can facilitate the broader adoption of GPU-accelerated whole-brain modeling within the neuroscience community.

      Weaknesses:

      The test-retest reliability and heritability are estimated using high-quality Human Connectome Project data. The manuscript would benefit from discussion and/or demonstrations regarding how the software performs under more challenging conditions, such as clinical datasets, shorter data acquisitions, higher-motion datasets, or multi-site datasets.

      Apart from demonstrating the benefits of GPUs over CPUs, the manuscript would benefit from a more direct comparison between cuBNM and other whole-brain modeling software, such as The Virtual Brain.

      The manuscript demonstrates that heterogeneous models improve the fit to empirical functional connectivity. However, the biological interpretation of this improvement could be expanded. The heterogeneous models are also more complex than homogeneous models, and some improvement in model fit may be explained by the increased model flexibility.

      In whole-brain brain network modeling, different parameter combinations can result in similar empirical functional connectivity measures. The manuscript would benefit from a discussion of how this influences the interpretation of individualized parameter estimates.

    3. Reviewer #2 (Public review):

      Summary:

      The authors aim to address a major problem in brain network modeling: the high computational cost of simulating and fitting brain activity models, particularly for large samples, individualized models, and broad parameter searches. They introduce cuBNM, an open software package that uses graphics processing units to accelerate model simulation, fitting, and calculation of simulated brain activity features.

      The manuscript is primarily a methods and software contribution, rather than a paper providing novel neurobiological insights. The authors demonstrate the tool using human imaging data, showing examples of group-level and individualized model fitting, comparisons between homogeneous and heterogeneous model parameterizations, and analyses of repeated-measurement stability and genetic influences of simulated features. They also provide speed and scaling tests to support the claim that the software can make large-scale and individualized brain network modeling more practical for the field.

      Strengths:

      A major strength of this work is that it addresses a clear computational bottleneck in brain network modeling. The authors provide an open software package that combines a user-friendly Python interface with an accelerated back-end, making large numbers of simulations and model fits more practical for other researchers.

      The demonstrations are broad and relevant to real use cases. The authors show group-level and individualized model fitting, different optimization strategies, and comparisons between homogeneous and heterogeneous models, rather than limiting the paper to a narrow technical benchmark. The benchmarking and openness of the work further increase its value. The comparisons across hardware and network sizes give readers a practical sense of the tool's speed and scalability, while the availability of code, documentation, tutorials, and containers should make the method easier for the community to test and adopt.

      Weaknesses:

      (1) The benchmarking provides solid evidence for substantial speed improvements within the authors' implementation, but the generality of the performance claims is more limited. The largest reported speed-ups are measured relative to a single central processing unit thread, and the study does not fully benchmark cuBNM against other optimized brain modeling frameworks. This makes the results useful as evidence of strong acceleration in the tested setting, but less definitive as a general comparison across available implementations.

      (2) The comparison between homogeneous and heterogeneous models is informative, but it is not fully controlled for model complexity. The best-fitting node-based heterogeneous model has more free parameters than the homogeneous and map-based alternatives, so its improved fit may partly reflect greater flexibility rather than a more biologically valid parameterization. As a result, the model comparison supports the conclusion that this parameterization fits better under the current setup, but not necessarily that it is generally superior or more biologically realistic.

      (3) The reliability and heritability analyses are valuable demonstrations of what scalable individualized modeling can enable, but they do not establish the simulated features as validated biological mechanisms. Because these simulated features are derived from individualized structural and functional imaging data, their stability across repeated measurements and genetic influences may partly reflect information already present in the empirical inputs or fitting targets. These results therefore support a more cautious conclusion: the simulated features retain stable and genetically structured variation, but their biological interpretation remains model dependent.

      (4) The empirical demonstrations are narrower than some of the broader claims made in the manuscript. Most analyses rely on one human imaging dataset, one cortical parcelation, one main brain model, and a specific fitting objective, while broader claims refer to diverse populations, dense networks, high-dimensional models, and biological applications. The current results show that cuBNM is a useful and scalable tool in the tested setting, but the extent to which the findings generalize across datasets, model classes, network resolutions, or clinical contexts remains to be established.

    1. eLife Assessment

      This is a valuable study addressing a debated question in brain repair research by testing whether NeuroD1 can convert brain immune cells into nerve cells using a virus-free genetic approach and live imaging. The solid evidence presented here supports that, under the conditions tested, NeuroD1-expressing cells do not become nerve cells and instead retain their original identity; however, some aspects of expression level, injury timing, and cell-type composition would benefit from further clarification. This work will be of interest to neuro-immunologists.

    2. Reviewer #1 (Public review):

      Summary:

      This study revisits an important and controversial question in brain repair: whether NeuroD1 can convert brain immune cells into nerve cells in vivo. Using a virus-free genetic system, in vivo imaging, injury experiments, and single-cell profiling, the authors provide convincing evidence that NeuroD1-expressing cells do not become nerve cells under the tested conditions. Instead, these cells largely retain their original immune-cell identity, and some appear to undergo cellular stress or loss.

      Strengths:

      The main strength of the work is that it tests this question with a cleaner genetic strategy, avoiding some of the concerns associated with viral delivery and unintended cell labeling. Although the overall conclusion is consistent with the authors' previous work, the current study adds useful independent evidence, particularly through the virus-free fate-mapping system and live imaging in the brain.

      Weaknesses:

      There are some limitations. In the injury experiment, the labeled cells may include both resident brain immune cells and blood-derived immune cells recruited after injury, so the authors should be cautious when referring to all labeled cells as microglia. The level of NeuroD1 expression achieved by the genetic system is also not fully defined, which matters because the effects of such a cell-fate regulator may depend on expression level. Finally, the tested time window may not fully address very delayed or incomplete neuronal differentiation.

      Overall, this is a useful and careful study that supports the conclusion that NeuroD1 does not drive brain immune cells to become nerve cells in the tested settings. It should be valuable for researchers studying brain repair, cell fate conversion, and genetic fate mapping, and it provides a clear caution against overinterpreting reprogramming results based only on viral labeling.

    3. Reviewer #2 (Public review):

      Summary:

      In vivo glia-to-neuron conversion emerges as a potential regeneration-based therapeutic strategy for neural injuries and diseases. However, controversies exist in this exciting field, largely arising from the non-stringent methods employed for analyzing in vivo neuronal conversions. The study by Li et al. directly addressed this controversy regarding Neurod1-mediated microglia-to-neuron conversion. They took advantage of two transgenic mouse lines to specifically express Neurod1 in the microglia of adult mouse brains. Results from immunohistochemistry, in vivo live-cell imaging, and scRNA-seq convincingly demonstrate that microglia cannot be converted in vivo to neurons by ectopic Neurod1 expression under both normal and injury conditions. Instead, it induces microglia death, consistent with their earlier findings. These solid results, though negative, are critical additions to the field and further support that stringent lineage tracing methods are essential for studying in vivo cell reprogramming. Overall, the studies are rigorously designed and executed. Only minor issues need to be dealt with.

    1. eLife Assessment

      This study uses a technically challenging long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The study provides solid evidence that ORNs have a circadian pattern of spontaneous firing that is Orco-dependent and yet that Orco transcript abundance does express a circadian rhythmicity and that cAMP can modulate all Orco-dependent activity. Together with computational work, the work is valuable in proposing a provocative hypothesis that an Orco-centred post-translational feedback loop generates the circadian rhythm.

    2. Joint Public Review:

      The revised manuscript is much clearer, and the additional analyses address several of the original concerns. RAIN analyses (Rhythmicity Analysis Incorporating Nonparametric methods) now detects circadian rhythmicity in 7/11 recordings under light-dark conditions and 8/12 recordings in constant darkness, compared with 2/12 following treatment with the Orco antagonist. This supports circadian modulation of spontaneous firing and a role for Orco in its normal expression. The expanded qPCR analysis of Orco also supports the conclusion that Orco transcript abundance is not circadian, and the cAMP experiment shows that cAMP can modulate Orco-dependent activity.

      The remaining issue concerns the mechanistic interpretation. The lack of rhythmic Orco transcript abundance does not distinguish an autonomous post-translational feedback-loop (PTFL) clock from a model in which the canonical transcriptional-translational (TTFL) clock acts upstream through cAMP, calcium, kinases, phosphatases, channel trafficking, or related pathways to regulate Orco.

      Similarly, the new Figure 10 provides a useful representation of the authors' hypothesis, but the proposed delayed feedback and coupling mechanisms are not experimentally demonstrated.

      We do not think any further experiments are necessary for the present study. Instead, we recommend that the manuscript should clearly distinguish between what the data show and what remains proposed. The data support circadian modulation of ORN firing a role for Orco in its normal expression, non-circadian Orco transcript abundance, and cAMP-sensitive modulation of Orco-dependent activity. The proposal that an Orco-centred membrane feedback loop generates the rhythm is intriguing and may remain a hypothesis generated through this study that needs formal testing in the future. This should be explicitly stated. While this has been done in the discussion section, elsewhere, including in the abstract and elsewhere, the original claim remains.

    3. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study uses technically compelling long-term in vivo recordings and computational modeling to investigate whether hawkmoth olfactory receptor neurons show circadian modulation of spontaneous firing. The authors further propose the provocative model that post-translational mechanisms, rather than the transcriptional-translational processes, may contribute to circadian regulation of neuronal excitability.

      We are pleased that our study is recognized as valuable and that our technically very challenging long-term in vivo recordings and computational modeling are appreciated. We agree that we are proposing a provocative model that opposes the current hypothesis in chronobiology, which suggests that all observed biological circadian rhythms are outputs of a transcriptional-translational feedback loop (TTFL) clock. Instead, we suggest that a cell comprises, in addition to the TTFL clock, other posttranslational feedback loop (PTFL) clocks without the need of daily transcription and daily degradation of its core elements. While the circadian TTFL clock is entrained to the daily light-dark cycle, the circadian PTFL clocks are suggested to be entrained to other daily cues such as to the availability of pheromone, or the daily changing levels of hormones and second messenger levels. Our novel hypothesis proposes that the TTFL and PTFL clocks are coupled and linked, constituting an adaptive, plastic network that can tune and phase-lock to different external and internal Zeitgeber signals. However, we certainly do not claim that the TTFL circadian clock is not at all involved in the circadian control of the ORN’s circadian membrane potential rhythms. We clarified our manuscript accordingly.

      However, the evidence for circadian firing in these neurons […] remains incomplete.

      As requested by the reviewers during the previous round of review, we had provided the results of RAIN analysis (Thaben and Westermark, 2014) of individual animals in our first revision (Fig. 4A), which clearly shows that two-thirds of the population express circadian rhythms in key attributes that we used to characterize the spontaneous spiking activity. In the initial review it was assumed by the reviewers that phase alignment of dispersed rhythms would bias interpretations of rhythmicity. After having established with RAIN that individuals show circadian rhythms, albeit dispersed across the population due to the lack of a zeitgeber in DD conditions, phase-alignment of recordings from DD animals is a valid next step to prepare the data for statistical analysis across the population. Phase-alignment of desynchronized rhythms is a generally accepted and necessary method employed in chronobiology (e.g., for insect ORNs: Gosh et al., 2024). It is proven as prerequisite to find and statistically analyze rhythmicity in complex, desynchronized data.

      As we explained in the previous rebuttal and in our first revision, rhythms in electrical activity of insect ORNs cannot be easily synchronized by the light-dark cycle alone, but appear to require daily cycles of pheromone, as shown in other moth species (Gosh et al., 2024). Please be aware that the animals that we used here have never been exposed to pheromone, as we state in the Methods. Therefore, this lack of pheromone exposure can explain why about one third of our experimental population is not expressing any daily or circadian rhythmicity in spiking attributes. This is an important result of our manuscript, reported for the first time for Manduca sexta, providing evidence for our hypothesis that it is not the LD-entrained circadian TTFL clock that governs electrical activity rhythms in ORNs. As we explained in the first revision, and clarified further here in the second revision, we cannot phase-synchronize our animals with cycles of pheromone application in our experimental paradigm because we are researching circadian rhythms in spontaneous spiking activity and not pheromone responses. We failed to obtain phase-alignment with a single pheromone pulse the night before the experiments started. These data were added as supplementary Figure to Fig. 3 in the first revision. Here, we further revised our manuscript to clarify this important finding.

      Thus, as requested by the reviewers in the initial review, we could successfully confirm our previous results of circadian firing in ORNs and the disruption of these circadian rhythms with Orco antagonist OLC15 with RAIN. In the current review, the reviewers raise no further specific critical points or comments that would doubt our careful rhythm analysis of our long-term recordings. Thus, we conclude that we provided clear evidence for our central, exciting new finding. For the first time we demonstrated an unexpected new task for Orco: Orco controls the circadian firing pattern in the spontaneous activity, and thus, of the ORN’s membrane potential, via its property as leak/pacemaker channel.

      However, the evidence […] for post-translational modification of Orco as the underlying mechanism remains incomplete.

      We agree with the reviewers that there are many more experiments and combined efforts of biochemists, structural biologists, and electrophysiologists required to provide complete evidence for post-translational modification of Orco and to reveal the underlying mechanism of its circadian control. It is beyond the scope of the current manuscript to provide all details of post-translational control of Orco.

      The reviewers asked previously for additional evidence that Orco transcription is not controlled via the TTFL clock. As requested, we provided extended qPCR–based evidence that Orco, in contrast to timeless, is not controlled by the TTFL clock on the transcriptional level (Fig. 6 in Revision 1). Furthermore, we added a new result in Revision 1 to demonstrate cAMP-dependent post-translational modulation of Orco open-time probability (Fig 9 in Revision 1) at a ZT at which antennal cAMP levels are low (Schendzielorz et al., 2015). We already showed in Flecke et al., 2010, that the addition of cAMP at different ZTs increases the spontaneous spiking activity only at specific ZTs. Here, we show that the effect of cAMP depends on Orco. Since Orco’s circadian role is not mediated via TTFL control, it can be concluded that post-translational mechanisms provide daily/circadian temporal control. In this additional Figure we provide statistically significant proof that, in agreement with our model-prediction, Orco’s circadian control of the ORN spontaneous activity could be mediated via the second messenger cAMP. ZT-dependent input for Orco would be provided via daily changes in cAMP levels (Schendzielorz et al., 2015).

      In contrast, the study does provide strong evidence that the application of cyclic nucleotides can modulate Orco-dependent activity at a single time point, and reports that the temporal pattern of Orco transcript abundance is not circadian.

      We appreciate that the reviewer confirms that our revised manuscript with additional experiments now provides strong evidence that cAMP modulates Orco-dependent spontaneous activity of M. sexta ORNs. Since we already published that cAMP levels expresses daily rhythms in M. sexta antennae (Schendzielorz et al., 2015), and in vivo cAMP infusion increases spontaneous activity and sensitizes pheromone detection (Flecke et al., 2010), and our computational model here proves that circadian modulation of open time probability of Orco is sufficient to explain our experimental data, it is sufficient for our conclusions to test just the one specific zeitgeber time when endogenous cAMP levels are low and pharmacological cAMP increase has the strongest impact. To further reveal complete ZT-dependence of cAMP modulation of Orco´s control of spontaneous activity is beyond the scope of the current manuscript and not part of the current research question. The structure of Orco is extraordinarily conserved during evolution, thus, the cited experimental results from other laboratories and other species showing that Orco is a hub for posttranslational modification are very likely generalizable to different insect species. We clarified the manuscript accordingly. In Drosophila, Orco has at least 5 phosphorylation sites for protein kinase C (PKC), is cAMP-dependently sensitized, and has a Ca<sup>2+</sup>/calmodulin binding site that orchestrates the localization of the OR-Orco heteromer to the cilia. However, in fruit fly and other insects, so far, it can only be speculated how circadian control is provided for Orco, since there are no other publications that examine the circadian regulation of Orco in detail. We clarified our manuscript accordingly.

      To summarize, the logical conclusion based on our newly provided data is that the current hierarchical hypothesis in chronobiology based solely on a circadian TTFL clock that controls Orco transcription does not explain our findings in hawkmoth ORNs. Therefore, we suggest a new systemic hypothesis based upon coupled TTFL and PTFL circadian clocks that can also reconcile otherwise inconsistent data published for insect and mammalian circadian clocks (please see reviewed data in: Stengl and Schneider, 2024). We clarified our manuscript in the second revision and added a new Figure 10 to illustrate our novel hypothesis.

      However, the findings are incomplete to exclude a role for transcriptional-translational mechanisms and their associated multi-layered controls in circadian regulation.

      We certainly do not imply excluding a role for the TTFL clock in (indirectly) affecting circadian control of the membrane potential of ORNs. The new qPCR experiments added in the first revision clearly show that the circadian control mediated via Orco is not an output of the TTFL clock via transcriptional control of Orco. Instead, we predict links between a PTFL membrane clock comprising Orco as hub to integrate posttranslational control and the TTFL nuclear clock. We clarified our manuscript accordingly, adding a new Figure 10 to further illustrate and visualize our hypothesis. The predictions of this systemic hypothesis will be challenged in further experiments that, however, are beyond the scope of the current manuscript.

      Joint Public Review:

      This manuscript puts forward the provocative idea that a posttranslational feedback loop regulates daily and ultradian rhythms in neuronal excitability. The authors used in vivo long-term tip recordings of the long trichoid sensilla of male hawkmoths to analyze spontaneous spiking activity indicative of the ORNs' endogenous membrane potential oscillations. This firing pattern was disrupted by pharmacological blockade of the Orco receptor. They then use these recordings together with computational modeling to predict that Orco receptor neuron (ORN) activity is required for circadian, not ultradian, firing patterns. Orco did not show a circadian expression pattern in a qPCR experiment, and its conductance was proposed to be regulated by cyclic nucleotide levels. This evidence led the authors to conclude that a post-translational feedback loop (PTFL) clockwork, associated with the ORN plasma membrane, allows for temporal control of pheromone detection via the generation of multi-scale endogenous membrane potential oscillations. The findings will interest researchers in neurophysiology, circadian rhythms, and sensory biology. However, the manuscript has limited experimental evidence to support its central hypothesis and is undermined by several assumptions that underlie their data analysis and model builds, as well as insufficient biological data including critical controls to validate and/or fully justify the model the authors are proposing.

      We want to remind our reviewers that we used “ORN” as abbreviation for olfactory receptor neuron (= sensory receptor neurons, a.k.a. olfactory sensory neuron (OSN)) and not for Orco receptor neuron, although we focus on the function of Orco. Accordingly, our central finding is that Orco as ion channel is required for the daily/circadian modulation of spontaneous action potential activity generated by the olfactory receptor neurons in the absence of pheromone stimulation.

      We do not understand the specific basis for the conclusions of the reviewers. Therefore, we ask to please specify what experimental evidence is missing to support our central hypothesis that Orco is not directly TTFL- but PTFL clock-controlled, and to name specifically what the “several assumptions” are that undermine our careful data analysis and model builds. Which specific argument in our previous rebuttal was wrong, was not conclusive? Furthermore, please specify your claim that “critical controls are missing”. Which controls are missing for which experiments? We did add a new figure panel in the first revision to demonstrate that neither the addition of DMSO (the OLC15 solvent) nor the repeated attachment of the recording electrode, which was necessary to obtain paired datasets, altered the spontaneous spiking activity (Fig 1B in Revision 1). Furthermore, we expanded the time series of qPCR data and added tim as positive control to Orco (Fig 7), strengthening our argument that Orco expression is not under TTFL control. Dose-response curves of various Orco agonists and antagonists have been published before (see our references in Revision 1) and are therefore not repeated here.

      Our newly added data confirm what our modeling predicted: cAMP increases spontaneous ORN activity dependent on Orco. Previous publications provide evidence for daily rhythms in cAMP concentrations in hawkmoth antennae (Schendzielorz et al., 2012).

      As is true for any other hypothesis, a hypothesis can only be falsified but not validated and needs to be tested by many experiments from many laboratories over a long time until it will be replaced by the next hypothesis that better explains accumulating contradicting evidence. We are very much looking forward to experimental challenges of our provocative new hypothesis by colleagues in the field of olfaction and of chronobiology. We are convinced that our manuscript will greatly stimulate the field, possibly provoking a paradigm switch in chronobiology and in olfactory research.

      Strengths:

      The authors raise several intriguing model-based hypotheses regarding the mechanisms that underlie the generation of olfactory rhythms. The electrophysiological approach and the long-term recording paradigm are elegant and technically impressive. In the revised version, the authors have added additional qPCR data supporting the lack of rhythmic Orco transcript expression and included a new figure suggesting that cAMP can modulate Orco conductance.

      We thank the reviewers for their acknowledgement of our careful work and hope that our further revisions and clarifications help to argue our case.

      Major weaknesses:

      (1) The cAMP experiment was only conducted at one time-point, which is insufficient to support the central claim that "AMP and cGMP may have ZT-dependent effects on Orco conductivity".

      We agree with the reviewers and revised our discussion accordingly to clarify that in this manuscript it is not our central claim that cAMP and cGMP may have ZT-dependent effects on Orco conductivity. Instead, our data show for the first time that Orco controls circadian rhythms of spontaneous activity of ORNs and that the circadian rhythmicity of spontaneous activity is lost when Orco is blocked. Therefore, we provide novel experimental evidence that Orco is a prerequisite to the circadian rhythmicity of spontaneous activity and thus, to circadian rhythms in the membrane potential of ORNs. Furthermore, as requested by the reviewers we provided clear evidence in the first revision that Orco is not controlled at the transcriptional level by the TTFL clock, in contrast to the TTFL clock protein TIMELESS. Thus, it follows logically that Orco is under post-transcriptional control. Since cAMP levels show circadian oscillations and Orco is gated by cAMP (Fig 9 in Revision 1), we used our computational model to show that a cAMP-dependent increase in Orco conductance alone, via daily oscillating concentrations of cAMP, is sufficient to explain our findings. Therefore, we propose here that daily/circadian oscillations of cAMP modify spontaneous spiking activity via Orco on a posttranscriptional level. But it certainly does not provide all evidence for respective mechanisms of how this cAMP modulation of Orco is obtained, since this is beyond the scope of the current manuscript.

      Since we realized that it is difficult for our readers to visualize a circadian PTFL membrane clock we added a new hypothesis-Figure (Figure 10) and considerably focused and clarified especially the discussion of our manuscript. We pointed out that a membrane-associated signalosome that comprises delayed negative feedback mechanisms, and, thus, constitutes an oscillator, a “membrane clock” that generates oscillations. Based upon our data we propose a membrane-associated signalosome constituting a PTFL circadian clock with Orco as central element. This PTFL membrane clock generates superimposed ultradian and circadian rhythms in its outputs: rhythms in the membrane potential, Ca<sup>2+</sup>, and cAMP levels. The PTFL clock comprises positive feedforward elements that upregulate its outputs, resulting in more depolarization, higher Ca<sup>2+</sup>- and higher cAMP levels. Via the clock’s delayed negative feedback mechanisms these outputs are downregulated, again, resulting in hyperpolarization, decreasing Ca<sup>2+</sup>- and cAMP levels. This signalosome comprises the pacemaker channel Orco as a central hub that is controlled via changes in voltage, Ca<sup>2+</sup>, and cAMP levels. Nevertheless, we predict coupling between the multiscale PTFL membrane clock and the TTFL circadian clock in the nucleus to obtain stable circadian rhythms. As likely mechanism of coupling we predict that Ca<sup>2+</sup>- and cAMP-dependent kinases interlink both types of clocks, thereby obtaining robust and at the same time flexible interlinked cellular rhythms.

      We hope to now successfully clarify and to visualize our central hypothesis of our manuscript that Orco is not directly controlled by a TTFL circadian clock but is a central element of a membrane-associated posttranslational feedback loop clock (PTFL) clock that is linked to but not forced by the TTFL clock which is predicted to control intracellular Ca<sup>2+</sup> homeostasis in a circadian rhythm.

      (2) The revised manuscript continues to rely heavily on prior publications or defers key mechanistic questions (or important manipulations) to future studies. In its current form, the evidence presented remains insufficient to support the central claim that a PTFL constitutes the primary underlying circadian clock mechanism. The proposed model is intriguing, but the data provided do not yet directly demonstrate the novel mechanism.

      We do not understand why the reviewers considers it to be problematic that we “continue to rely heavily on prior publications”. Certainly, we built upon previous publications of our lab as well as on manifold experimental data published by other laboratories in the field of insect olfaction. Our ample citations demonstrate that we have an overview both of the current state of literature and relevant previous literature, dating back to the very first experiments that pioneered pheromone transduction in insects. Based upon our extensive knowledge and experimental data collected in insect olfaction and based on very careful, critical, rigorous analysis of our data and data published by others, we were able to come up with a novel interpretation of the current literature about insect olfaction that differs considerably from the current main views. We consider this to be our strength and judge it as good scientific practice and not a flaw of our work. However, since we do not focus on OR-Orco heteromers and their function in pheromone/odor transduction in the cilia in the current manuscript, we considerably shortened this part of the discussion, avoiding pointing out that highly sensitive moth pheromone transduction greatly differs from less sensitive general odor transduction in Drosophila. Furthermore, since here we focus on cAMP, but not on cGMP-dependent modulation of Orco, we also deleted/considerably shortened this part of our discussion.

      We certainly agree with the reviewers that, while the data provided in the current manuscript are a logical basis for developing our novel hypothesis, they are not a direct and sufficient demonstration of proof and we are not able yet to directly demonstrate and explain the novel mechanism predicted. We would like to point out that if we provided this final proof, it would not be any more a novel hypothesis, but only a novel finding.

      We agree with the reviewers that our provocative hypothesis requires rigorous testing by many further experiments, hopefully not only by our laboratory, but hopefully stimulating new experimental challenges by other laboratories employing different species. But certainly, these experiments with proof-of-principle will take many years and are beyond the scope of our current research paper.

      As per eLife’s assessment system we would like to ask the reviewers to provide detailed feedback as to which experiments/results within the scope of this manuscript would complete this work, or how they think this study should be framed in the light of the results that we obtained. Nevertheless, we hope that with our current careful review the reviewers will be more convinced by our arguments and experiments as valid basis for our provocative new hypothesis.

    1. eLife Assessment

      This fundamental work significantly advances our understanding of how contact-dependent antagonism enables keystone bacteria to establish and maintain their niche over time. The evidence obtained is convincing, supporting most of the conclusions drawn. This work will be of significant interest to the microbiome research community.

    2. Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate the physiological role of the Type VI secretion system (T6SS) in a naturally evolved gut microbiome derived from wild mice (the WildR microbiome). Focusing on Bacteroides acidifaciens, the authors use newly developed genetic tools and strain replacement strategies to test how T6SS-mediated antagonism influences colonization, persistence, and fitness within a complex gut community. They further show that the T6SS resides on an integrative and conjugative element (ICE), is distributed among select community members, and can be horizontally transferred, with context-dependent effects on colonization and persistence. The authors conclude that the T6SS stabilizes strain presence in the gut microbiome while imposing ecological and physiological constraints that shape its value across contexts.

      This study is likely to have significant impact on the microbiome field by moving experimental tests of T6SS function out of simplified systems and into a naturally co-evolved gut community. The WildR system, together with the strain replacement strategy, ICE-seq approach, and genetic toolkit, represents a powerful and reusable platform for future mechanistic studies of microbial antagonism and mobile genetic elements in vivo.

      The datasets-including isolate genomes, metagenomes, and ICE distribution maps-will be valuable community resources, particularly for researchers interested in strain-resolved dynamics, horizontal gene transfer, and ecological context dependence. Even where mechanistic resolution is incomplete, the work provides a strong experimental foundation upon which such questions can be directly addressed.

      Overall, this study occupies a space between system building and mechanistic dissection. The authors demonstrate that the T6SS influences persistence and community structure in vivo, but the physiological basis of these effects remains unresolved. Interpreting the results as evidence of fitness costs or selective advantage therefore requires caution, as multiple ecological and host-mediated processes could produce similar abundance trajectories.

      Placing the findings within the broader literature on microbial antagonism, particularly work emphasizing measurable costs, benefits, and tradeoffs, would help readers better contextualize what is directly demonstrated here versus what remains an open question. Viewed in this light, the principal contribution of the study is to show that such questions can now be addressed experimentally in a realistic gut ecosystem.

      Strengths:

      A major strength of this study is that it directly interrogates the physiological role of the T6SS in a naturally evolved gut microbiome, rather than relying on simplified pairwise or in vitro systems. By working within the WildR community, the authors advance beyond descriptive surveys of T6SS prevalence and address function in an ecologically relevant context.

      The authors provide clear genetic evidence that Bacteroides acidifaciens uses a T6SS to antagonize co-resident Bacteroidales, and that loss of T6SS function specifically compromises long-term persistence without affecting initial colonization. This temporal separation is well designed and supports the conclusion that the T6SS contributes to maintenance rather than establishment within the community.

      Another strength is the identification of the T6SS on an integrative and conjugative element (ICE) and the demonstration that this element is distributed among, and exchanged between, community members. The use of ICE-seq to track distribution and transfer provides strong support for horizontal mobility and adds mechanistic depth to the study.

      Finally, the transfer of the T6SS-ICE into Phocaeicola vulgatus and the observation of context-dependent colonization benefits followed by decline is a compelling result that moves the study beyond simple "T6SS is beneficial" narratives and highlights ecological contingency.

      Weaknesses:

      Despite these strengths, there is a mismatch between the precision of the claims and the precision of the measurements, particularly regarding fitness costs, physiological burden, and mechanistic role of the T6SS.

      First, while the authors conclude that the T6SS "stabilizes strain presence" and that its value is constrained by fitness costs, these costs are not directly measured. Persistence, abundance trajectories, and eventual loss are informative outcomes, but they do not uniquely identify fitness tradeoffs. Decline could arise from multiple non-exclusive mechanisms, including community restructuring, host-mediated effects, incompatibilities of the ICE in new hosts, or ecological retaliation, none of which are disentangled here.

      Second, the manuscript frames the T6SS as having a defined physiological role, yet the data do not resolve which physiological processes are under selection. The experiments demonstrate that T6SS activity affects persistence, but they do not distinguish whether this occurs via direct killing, resource release, niche modification, or higher-order community effects. As a result, "physiological role" remains underspecified and risks being conflated with ecological outcome.

      Third, although the authors emphasize context dependence, the study offers limited quantitative insight into what aspects of context matter. Differences between native and recipient hosts, or between early and late colonization phases, are described but not mechanistically interrogated, making it difficult to generalize beyond the specific cases examined.

      Fourth is the lack of engagement with recent experimental literature demonstrating functional roles of the T6SS beyond simple interference competition. While the authors focus on persistence and competitive outcomes, they do not adequately situate their findings within recent work demonstrating that T6SS-mediated antagonism can serve additional physiological functions, including resource acquisition and DNA uptake, thereby linking killing to measurable benefits and tradeoffs. The absence of this literature makes it difficult to place the authors' conclusions about physiological role and fitness cost within the current conceptual framework of the field. Without this context, the physiological interpretation of the results remains incomplete, and alternative functional explanations for the observed dynamics are underexplored.

      A further limitation concerns the taxonomic scope of the functional analysis. The authors state the role of the T6SS in the murine environment is functionally investigated using genetically tractable Bacteroides species, citing lack of genetic tools for Mucispirillum schaedleri. While this is a reasonable practical choice, it means that a substantial fraction of T6SS-encoding species in the WildR community are not experimentally interrogated. Consequently, conclusions about the role of the T6SS in the murine gut necessarily reflect the subset of taxa that are genetically accessible and may not fully capture community-level or niche-specific functions of T6SS activity. Given that M. schaedleri is represented as a metagenome-assembled genome, its isolation and genetic manipulation would be technically challenging. Nonetheless, explicitly acknowledging this limitation and slightly tempering claims of generality would strengthen the manuscript.

      Finally, several interpretations would benefit from more cautious language. In particular, claims invoking fitness costs, selective advantage, or physiological burden should be explicitly framed as inferences from persistence dynamics, rather than as direct measurements, unless supported by additional quantitative fitness or growth assays.

      Comments on revised version.

      The authors have addressed my main concerns by more clearly distinguishing ecological outcomes from directly measured physiological mechanisms. They have moderated claims about fitness costs and benefits, replaced "physiological" with "ecological" where appropriate, expanded the discussion of potential downstream benefits of T6SS-mediated killing, and acknowledged the limited taxonomic scope of the functional analyses. The persistence trajectories support context-dependent relative fitness effects, although they do not identify the specific physiological basis of those effects. The revised manuscript now generally maintains this distinction. These revisions substantially improve the precision and balance of the manuscript.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to determine how a contact-dependent bacterial antagonistic system contributes to the ability of specific bacterial strains to persist within a complex, native gut community derived from wild animals. Rather than focusing on simplified or artificial models, the authors aimed to examine this system in a biologically realistic setting that captures the ecological complexity of the gut environment. To achieve this, they combined controlled laboratory experiments with animal colonization studies and sequencing-based tracking approaches that allow individual strains and mobile genetic elements to be followed over time.

      Strengths:

      A major strength of the work is the integration of multiple complementary approaches to address the same biological question. The use of defined but complex communities, together with in vivo experiments, provides a strong ecological context for interpreting the results. The data consistently show that the antagonistic system is not required for initial establishment but plays a critical role in long-term strain persistence, an insight that moves beyond traditional invasion-based views of microbial competition. The observation that transferable genetic elements can confer only temporary advantages, and may impose longer-term costs depending on community context, adds important nuance to current understanding of microbial fitness.

      Weaknesses:

      Overall, the study is not a lack of evidence, but a deliberate trade-off between ecological realism and mechanistic resolution, which leaves some causal pathways open to interpretation.

      Comments on revised version.

      The authors have addressed all previous concerns thoroughly and satisfactorily.

    4. Reviewer #3 (Public review):

      Summary:

      In this work, the authors investigate the contribution of the type VI secretion system of Bacteroidales to gut microbiome assembly and the targeting of closely related species. They demonstrate that B. acidifaciens relies on T6SS-mediated antagonism to prevent displacement by co-resident Bacteroidales and other members of the microbiome, allowing it to persist in the gut. They also developed new tools for analyzing the distribution of mobile genetic elements. This study advances our understanding of how molecular systems contribute to shaping complex microbial communities.

      Strengths:

      The use of a gnotobiotic model colonized with a wild-mouse microbiome is a significant strength of this study. This approach allows tracking of microbiome changes over time and evaluating the targeting by Bacteroidales carrying T6SS in a more natural setting. The development of ICE-seq for mapping the distribution of the T6SS in the microbiome is remarkable, enabling the study of how this bacterial weapon is transferred between microbiome members without requiring long-read metagenomics methods.

      Weaknesses:

      Some conclusions are based on a limited number of mice per condition. This could be due to the complexity of using a gnotobiotic mouse model, but this should be considered when interpreting the data.

      Overall, the authors successfully achieved their objectives, and their experimental design and results support their findings. As mentioned in the discussion, it would be important to investigate the role of the T6SS in resilience to microbiome disturbances, such as antibiotics, diet, or pathogen invasion. This work represents a step forward in understanding how contact-dependent competition influences the gut microbiome in relevant ecological contexts.

    5. Author response:

      The following is the authors’ response to the original reviews.

      We appreciate that the reviewers provided an overall positive assessment of our manuscript and offered constructive suggestions for improvement. All three reviewers noted that a key strength of our study is the implementation of a gut microbiome model for the characterization of interbacterial antagonism pathways such as the type VI secretion system (T6SS) that approaches natural complexity. They note our work represents a significant advance in microbiome research, and generates resources that will be of use to many researchers in the field. Two of the reviewers point out that the complexity of our model limits the nature of measurements we can make, and suggest we temper the strength of the some of the conclusions we draw. As noted in more detail below, in our revised manuscript, we have used more precise wording to characterize our findings, and we are more explicit about the connection between the measurements we made and what we can conclude about the physiological role of the T6SS in the gut microbiome.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate the physiological role of the Type VI secretion system (T6SS) in a naturally evolved gut microbiome derived from wild mice (the WildR microbiome). Focusing on Bacteroides acidifaciens, the authors use newly developed genetic tools and strain-replacement strategies to test how T6SS-mediated antagonism influences colonization, persistence, and fitness within a complex gut community. They further show that the T6SS resides on an integrative and conjugative element (ICE), is distributed among select community members, and can be horizontally transferred, with context-dependent effects on colonization and persistence. The authors conclude that the T6SS stabilizes strain presence in the gut microbiome while imposing ecological and physiological constraints that shape its value across contexts.

      This study is likely to have a significant impact on the microbiome field by moving experimental tests of T6SS function out of simplified systems and into a naturally coevolved gut community. The WildR system, together with the strain replacement strategy, ICE-seq approach, and genetic toolkit, represents a powerful and reusable platform for future mechanistic studies of microbial antagonism and mobile genetic elements in vivo.

      The datasets, including isolate genomes, metagenomes, and ICE distribution maps, will be a valuable community resource, particularly for researchers interested in strainresolved dynamics, horizontal gene transfer, and ecological context dependence. Even where mechanistic resolution is incomplete, the work provides a strong experimental foundation upon which such questions can be directly addressed.

      Overall, this study occupies a space between system building and mechanistic dissection. The authors demonstrate that the T6SS influences persistence and community structure in vivo, but the physiological basis of these effects remains unresolved. Interpreting the results as evidence of fitness costs or selective advantage, therefore, requires caution, as multiple ecological and host-mediated processes could produce similar abundance trajectories.

      Placing the findings within the broader literature on microbial antagonism, particularly work emphasizing measurable costs, benefits, and tradeoffs, would help readers better contextualize what is directly demonstrated here versus what remains an open question. Viewed in this light, the principal contribution of the study is to show that such questions can now be addressed experimentally in a realistic gut ecosystem.

      We thank the reviewer for this thoughtful summary of our study. We were glad to read they conclude our work will have a significant impact on the microbiome field and that the resources we have developed will be of value to the community.

      Strengths:

      A major strength of this study is that it directly interrogates the physiological role of the T6SS in a naturally evolved gut microbiome, rather than relying on simplified pairwise or in vitro systems. By working within the WildR community, the authors advance beyond descriptive surveys of T6SS prevalence and address function in an ecologically relevant context.

      The authors provide clear genetic evidence that Bacteroides acidifaciens uses a T6SS to antagonize co-resident Bacteroidales, and that loss of T6SS function specifically compromises long-term persistence without affecting initial colonization. This temporal separation is well designed and supports the conclusion that the T6SS contributes to maintenance rather than establishment within the community.

      Another strength is the identification of the T6SS on an integrative and conjugative element (ICE) and the demonstration that this element is distributed among, and exchanged between, community members. The use of ICE-seq to track distribution and transfer provides strong support for horizontal mobility and adds mechanistic depth to the study.

      Finally, the transfer of the T6SS-ICE into Phocaeicola vulgatus and the observation of context-dependent colonization benefits followed by decline is a compelling result that moves the study beyond simple "T6SS is beneficial" narratives and highlights ecological contingency.

      We appreciate this detailed and nuanced characterization of the strengths of our study.

      Weaknesses:

      Despite these strengths, there is a mismatch between the precision of the claims and the precision of the measurements, particularly regarding fitness costs, physiological burden, and the mechanistic role of the T6SS.

      We acknowledge that in some places, our manuscript could benefit from greater precision in the language we use when linking the outcomes we observe in our study to their potential underlying causes. Specific revisions we made to address this concern are described below.

      First, while the authors conclude that the T6SS "stabilizes strain presence" and that its value is constrained by fitness costs, these costs are not directly measured. Persistence, abundance trajectories, and eventual loss are informative outcomes, but they do not uniquely identify fitness tradeoffs. Decline could arise from multiple nonexclusive mechanisms, including community restructuring, host-mediated effects, incompatibilities of the ICE in new hosts, or ecological retaliation, none of which are disentangled here.

      We agree that multiple mechanisms could explain why populations of certain species carrying a T6SS decline over time, and why for others, the T6SS contributes to long-term persistence. Our use of the term “fitness cost” to describe the phenomenon of decline observed for P. vulgatus carrying the T6SS was not meant to imply any particular underlying mechanism, but was rather our attempt to characterize the phenotypic outcome we observed in simplified terms. We note that ecological context is an important determinant of the fitness cost or benefit of any given trait, and our study sheds light on the importance of the presence of the WildR community and the mouse intestinal environment to the fitness contribution of the T6SS to B. acidifaciens and P. vulgatus. Nonetheless, to avoid implying an overly simplistic interpretation of our results, we have modified our language in the manuscript in several places when describing the role of the T6SS in species persistence in mice colonized with the WildR community.

      Second, the manuscript frames the T6SS as having a defined physiological role, yet the data do not resolve which physiological processes are under selection. The experiments demonstrate that T6SS activity affects persistence, but they do not distinguish whether this occurs via direct killing, resource release, niche modification, or higher-order community effects. As a result, "physiological role" remains underspecified and risks being conflated with ecological outcome.

      We acknowledge that our study does not fully resolve the physiological processes under selection that mediate role of the T6SS in maintaining B. acidifaciens populations in WildR-colonized mice. Indeed, several of the outcomes of T6SS activity the reviewer lists, such as target cell killing and nutrient release, are inextricably linked and thus inherently difficult to disentangle. We note that we did attempt to measure higher-order community effects of T6SS activity with metagenomic sequencing, but acknowledge that this approach may not have been sufficiently sensitive to detect small community shifts mediated by a relatively low-abundance species. To address the concern that our current framing implies more of a mechanistic understanding that our study achieves, we have substituted “ecological” for “physiological” where appropriate throughout the manuscript.

      Third, although the authors emphasize context dependence, the study offers limited quantitative insight into what aspects of context matter. Differences between native and recipient hosts, or between early and late colonization phases, are described but not mechanistically interrogated, making it difficult to generalize beyond the specific cases examined.

      We are not entirely clear what the reviewer means by “differences between native and recipient hosts”, but we agree that additional quantitative studies will be needed to address the generalizability of our findings. Future studies are also needed to address the mechanistic basis for the difference in the benefit conferred by the T6SS that we observed between P. vulgatus and B. acidifaciens.

      Fourth is the lack of engagement with recent experimental literature demonstrating functional roles of the T6SS beyond simple interference competition. While the authors focus on persistence and competitive outcomes, they do not adequately situate their findings within recent work demonstrating that T6SS-mediated antagonism can serve additional physiological functions, including resource acquisition and DNA uptake, thereby linking killing to measurable benefits and tradeoffs. The absence of this literature makes it difficult to place the authors' conclusions about physiological role and fitness cost within the current conceptual framework of the field. Without this context, the physiological interpretation of the results remains incomplete, and alternative functional explanations for the observed dynamics are underexplored.

      We thank the reviewer for specifically highlighting the potential pertinence of this literature to our study. Indeed, we did not cite studies indicating a link between T6SS activity and the uptake of DNA and other resources released by targeted cells. As we note above, the release of intracellular contents from target cells is an inevitable consequence of the delivery of lytic effectors. Thus, distinguishing between fitness benefits conferred from the elimination of competitor species and those arising from scavenging the nutrients released during this process is not straightforward. Measuring the benefits deriving from the uptake of certain released molecules, such as DNA, was not immediately feasible in the system employed in this study and instead we focused on the direct lytic consequences of the effectors delivered via the T6SS. We revised the Discussion to include reference to these possible downstream benefits of T6SS activity (Lines 476-479).

      A further limitation concerns the taxonomic scope of the functional analysis. The authors state that the role of the T6SS in the murine environment is functionally investigated using genetically tractable Bacteroides species, citing the lack of genetic tools for Mucispirillum schaedleri. While this is a reasonable, practical choice, it means that a substantial fraction of T6SS-encoding species in the WildR community are not experimentally interrogated. Consequently, conclusions about the role of the T6SS in the murine gut necessarily reflect the subset of taxa that are genetically accessible and may not fully capture community-level or niche-specific functions of T6SS activity. Given that M. schaedleri is represented as a metagenome-assembled genome, its isolation and genetic manipulation would be technically challenging. Nonetheless, explicitly acknowledging this limitation and slightly tempering claims of generality would strengthen the manuscript.

      The reviewer points out that studying the T6SS activity in M. schadleri would potentially expand the generality of our claims. We agree that having an isolate of this species along with genetic tools for its manipulation would allow us to probe the importance of the T6SS in the gut microbiome more broadly. At the suggestion of the reviewer, we have added explicit mention of the potential benefit of studying the T6SS in this organism to the Discussion (lines 538-539), an endeavor that lies outside of the scope of the current study.

      Finally, several interpretations would benefit from more cautious language. In particular, claims invoking fitness costs, selective advantage, or physiological burden should be explicitly framed as inferences from persistence dynamics, rather than as direct measurements, unless supported by additional quantitative fitness or growth assays.

      We agree with the reviewer that invoking fitness costs, selective advantages or physiological burdens should be done cautiously, and have made revisions to our manuscript where we acknowledge that more precise language was needed (line 43, 416, line 417). However, we would also argue invoking fitness costs and benefits when describe strain persistence dynamics in mice has substantial precedent in the literature (Feng et al. 2020, Brown et al. 2021, Park et al. 2022, Segura Munoz et al. 2022), to list a handful of representative examples published by different groups). It is unclear to us what additional in vivo growth measurements could be taken to substantiate our claim that the T6SS provides a fitness benefit to B. acidifaciens during prolonged gut colonization, or that carrying the ICE imposes a fitness cost on P. vulgatus during longterm colonization. Our in vitro experiments evaluating the competitiveness conferred by T6SS activity provide a measure of insight into its fitness benefits, but as our in vivo strain persistence data and the work of many others show, in vitro measurements do not necessarily capture in vivo parameters.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to determine how a contact-dependent bacterial antagonistic system contributes to the ability of specific bacterial strains to persist within a complex, native gut community derived from wild animals. Rather than focusing on simplified or artificial models, the authors aimed to examine this system in a biologically realistic setting that captures the ecological complexity of the gut environment. To achieve this, they combined controlled laboratory experiments with animal colonization studies and sequencing-based tracking approaches that allow individual strains and mobile genetic elements to be followed over time.

      Strengths:

      A major strength of the work is the integration of multiple complementary approaches to address the same biological question. The use of defined but complex communities, together with in vivo experiments, provides a strong ecological context for interpreting the results. The data consistently show that the antagonistic system is not required for initial establishment but plays a critical role in long-term strain persistence. This insight that moves beyond traditional invasion-based views of microbial competition. The observation that transferable genetic elements can confer only temporary advantages, and may impose longer-term costs depending on community context, adds important nuance to current understanding of microbial fitness.

      We thank the reviewer for the positive feedback and are glad they agree our study provides new insight into the role of interbacterial antagonism in natural communities.

      Weaknesses:

      Overall, there is not a lack of evidence, but a deliberate trade-off between ecological realism and mechanistic resolution, which leaves some causal pathways open to interpretation.

      The reviewer makes a good point that the complexity of the experimental system we employ precludes some lines of experimentation that would yield more mechanistic information. As the reviewer notes, we were aware of the tradeoff between mechanistic resolution and ecological realism when selecting our experimental system. Our deliberate choice to favor biological complexity over mechanistic clarity in this study stemmed from our perception that a major gap in understanding of the T6SS and other antagonism pathways lies in defining their ecological function in complex microbial communities.

      Reviewer #3 (Public review):

      Summary:

      Shen et al. investigate the contribution of the type VI secretion system of Bacteroidales in the gut microbiome assembly and targeting of closely related species. They demonstrate that B. acidifaciens relies on T6SS-mediated antagonism to prevent displacement by co-resident Bacteroidales and other members of the microbiome, allowing B. acidifaciens to persist in the gut.

      Strengths:

      Using a gnotobiotic model colonized with a wild-mouse microbiome is a significant strength of this study. This approach allows tracking of microbiome changes over time and directly examining targeting by Bacteroidales carrying T6SS in a more natural setting. The development of ICE-seq for mapping the distribution of the T6SS in the microbiome is remarkable, enabling the study of how this bacterial weapon is transferred between microbiome members without requiring long-read metagenomics methods.

      We thank the reviewer for their enthusiasm toward our study.

      Weaknesses:

      Some conclusions are based on only four mice per condition. The author should consider increasing the sample size.

      We agree that in some experiments it would be beneficial to increase the sample size from four mice. However, the experiments we performed for this study are time and resource-intensive. Additionally, the experiments on which we base our primary conclusions were all independently replicated with similar results. Given these factors, we determined that the extra confidence that might be afforded by increasing our sample size did not merit the delay in publication and investment in resources that would be required.

      Overall, the authors successfully achieved their objectives, and their experimental design and results support their findings. As mentioned in the discussion, it would be important to investigate the role of the T6SS in resilience to disturbances in the microbiome, such as antibiotics, diet, or pathogen invasion. This work represents a step forward in understanding how contact-dependent competition influences the gut microbiome in relevant ecological contexts.

      We agree that investigating the role of the T6SS during perturbations of the microbiome is a key next step for this work and thank the reviewer for highlighting this important future direction.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Beth A. Shen et al. present a comprehensive and carefully executed study investigating the ecological role of the type VI secretion system (T6SS) in maintaining bacterial strains within a native, complex gut microbiome derived from wild mice. By integrating genetic manipulation, metagenomic and sequencing-based tracking approaches, in vitro competition assays, and gnotobiotic colonization experiments, the authors provide compelling evidence that the T6SS functions primarily as a persistence factor rather than a determinant of initial colonization.

      The study is conceptually strong and addresses an important gap in our understanding of how interbacterial antagonistic systems operate in complex, native microbial communities. The manuscript is generally well organized, the data are clearly presented, and the main conclusions are supported by robust experimental evidence. In particular, the demonstration that T6SS-encoding ICEs confer context- and hostdependent fitness effects, including transient benefits and potential long-term costs, adds important nuance to prevailing models of microbial competition.

      That said, several aspects of the study would benefit from clarification and deeper mechanistic discussion. Addressing the points below would further strengthen the rigor and interpretability of the work. Overall, this is a strong and interesting manuscript, requiring some revisions.

      We greatly appreciate the positive summary of our work by the reviewer, which highlights the multi-faceted approach we took to address gaps in our understanding of interbacterial antagonism in the microbiome. It is our hope that the reviewer agrees that our revisions of the manuscript, based on their feedback, clarify our methods and strengthen the interpretability of our work.

      Major comments

      (1) The competition assays in Figure 2B suggest that T6SS-dependent fitness effects are most pronounced among members of the order Bacteroidales. However, these experiments primarily measure population-level competitive outcomes rather than direct T6SS-mediated targeting events. In addition, the limited number of non-Bacteroidales strains included in the assay makes it difficult to conclude that T6SS activity is strictly restricted to closely related taxa.

      The authors should either temper their conclusions regarding target specificity or clarify that these data reflect competitive outcomes rather than direct evidence of targeting. Expanding the discussion to acknowledge these limitations would improve interpretative accuracy.

      With regards to the measurement we employed for assessing T6SS-mediated targeting, we acknowledge that this is, to a degree, an indirect way of determining T6SS targeting. However, there is extensive precedent in the literature for the use of similar assays in assessing targeting by many contact-dependent antagonism systems including the T6SS in Bacteroidales (Russell et al. 2014, Chatzidaki-Livanis et al. 2016, Wexler et al. 2016) and many Proteobacteria (e.g. (Hood et al. 2010)), the T4SS in Xanthomonas citri (Souza et al. 2015), the CDI system in Escherichia coli (Aoki et al. 2005), and the Esx system in Streptococcus intermedius (Whitney et al. 2017). In these studies, targeting was demonstrated by specific depletion of the competitor strain in the presence of a strain encoding an active antagonism system. We acknowledge that the competitive index we report Figure 2B reflects the relative population levels of both species in the assay, and thus does not directly show target species depletion. We opted to use this metric to display the data in the manuscript as a way of efficiently encapsulating and comparing many strain combinations in a single figure, and because the competitive index differences we observed in these derive from differences in target species growth yields (see Author response image 1, indicating growth yields from a representative strain pairing).

      Author response image 1.

      The T6SS of B. acidifaciens targets a WildR-derived P. vulgatus strain. CFUs indicate populations of the indicated strains after co-culture of wild-type or T6SS-inactivated B. acidifaciens with P. vulgatus. Data represent means and standard errors (n=3, *P<0.01, t-test with log ><0.01, t- test with log transformed data)

      We additionally acknowledge that more extensive testing is needed to fully understand the target range of the Bacteroidales T6SS. In our study, we assessed targeting of every WildR species that was readily culturable, which to the best of our knowledge, represents the broadest panel of targets for the Bacteroidales T6SS to be tested to date. We limited our testing to these strains, as the goal of these experiments was to gain insight into which co-residents of the WildR could be targeted by B. acidifaciens. We agree that testing of a broader cross-section of potential targets has merits, but this would require targeted cultivation strategies to obtain these organisms, and lies outside the scope of the current study. We have revised the manuscript to clarify that the target range testing encompassed the diversity of isolates available (p. 10, lines 231-236).

      (2) The bae1 gene encoded in Bacteroides caecimuris F12 contains a frameshift mutation. It would be valuable for the authors to comment on whether such frameshift mutations are a common genomic feature among gut-associated Bacteroides species in murine models. In addition, comparative analysis of human gut metagenomic datasets could reveal whether homologous effector proteins are present in commensal Bacteroides populations, and whether these homologs exhibit similar disruptive mutations.

      More broadly, the manuscript would benefit from a discussion of whether expression of a fully functional bae1 effector might impose a fitness cost on Bacteroidales members, for example, through metabolic burden or altered resource allocation. This is particularly relevant in light of recent studies demonstrating that T6SS effectors can drive physiological trade-offs by modulating metabolic dynamics (PMID: 40592326). Integrating this perspective would strengthen the evolutionary interpretation of effector mutagenesis.

      We agree with the reviewer that the functional and evolutionary significance of the point mutation in bae1 merits further investigation. Following the reviewer's suggestion, we looked in our own datasets and available public datasets from mouse and human microbiomes for evidence of bae1 inactivation. Unfortunately, the gene is present at a low enough frequency that these analyses were inconclusive. In our own metagenomic data from WildR mice, we did not obtain sufficient sequencing depth to assess the frequency at which bae1 is inactivated across genomes. We found a single complete copy of bae1 identical to that of B. acidifaciens in one published mouse microbiome-derived MAG, and detected fragments of the gene in a number of publicly available isolate and MAG genomes, but these were too low of quality to assess whether or not the gene was intact.

      As to whether or not bae1 expression imposes a fitness cost in the producing organism, we think this is unlikely to be significant, given that the impacts of Bae1 will be neutralized by the accompanying immunity protein. We speculate that the point mutation in the B. caecimuris gene is more likely to have arisen through genetic drift than as a result of selection.

      (3) Quantification and tracking of ICE transfer in vivo. In Figure 4D, the authors assess the abundance of resident P. vulgatus populations in germ-free mice co-gavaged with wild-type strains and derivatives carrying either the intact ICE or ICE ΔtssC. Because both ICE variants are capable of horizontal transfer, it is essential to clearly describe how the authors distinguish between (i) the original wild-type strain, (ii) engineered donor strains, and (iii) recipient strains that have newly acquired the ICE or ICE ΔtssC.

      Clarification of the specific molecular or sequencing-based strategies used to discriminate these populations is necessary to ensure accurate interpretation of the colonization dynamics.

      In this experiment, the P. vulgatus strains we introduced which carried the ICE (either the wild-type version or ICE DtssC) also contained an erythromycin resistance cassette (ermG) inserted distal to the ICE insertion site. Populations of the ICE-containing strain were quantified by either qPCR targeting the ermG gene (Fig. 4D, Supplemental Fig. 4D) or by plating on erythromycin-containing media (Fig. 4F). Endogenous P. vulgatus populations were quantified by qPCR targeting the ermG insertion site, which is disrupted in the marked strain. These methodological details have been added to the figure legend for clarity. We acknowledge that transfer of the ICE between introduced and endogenous populations is possible, and would not be detected by these metrics. To assess whether this occurs, we performed ICE-seq analyses on samples collected from mice colonized by the WildR and P. vulgatus ICE at early (7 days) and late (56 days) time points. These analyses revealed that overall, ICE distribution in this experiment was similar to that observed in mice colonized with the WildR alone (Figure 4A and Author response image 2). They additionally provided corroborating evidence that the population of ICE-containing P. vulgatus declined over the course of the experiment. Importantly, the only ICE insertion site we detected in P. vulgatus in these samples was that found in the introduced P. vulgatus strain. Previous studies show that GA1-containing ICE can insert at numerous locations in Bacteroides sp. genomes, a finding supported by our mapping of the ICE insertion sites from in vitro transfer experiments (Supplemental Fig. 4C) (Garcia-Bayona et al. 2021). Thus, our ICE-seq detection of a sole P. vulgatus ICE insertion site indicates that transfer of the element between P. vulgatus populations is likely not occurring in our experiments.

      Author response image 2.

      ICE-seq analysis indicates that introduction of P. vulgatus ICE into WildR-colonized mice has little impact on ICE distribution among endogenous strains. Graphs show frequency of mapped ICE junctions deriving from the indicated species as determined by 5¢ or 3¢ ICE-Seq analysis of DNA extracted from fecal samples collected either 7 or 56 days post-gavage of the WildR and P. vulgatus ICE into germ-free mice.

      (4) The analysis of fitness trade-offs associated with ICE acquisition in P. vulgatus convincingly demonstrates that the benefits of ICE transfer are transient and contextdependent. However, the mechanistic basis of these trade-offs remains underexplored. While the study primarily attributes both benefits and costs to T6SS-mediated antagonism, the ICE likely encodes additional genes that could influence metabolism, regulation, or stress responses.

      We agree with the reviewer that there are many mechanistic questions remaining regarding the benefits and costs associated with ICE acquisition, and acknowledge that we have not investigated the fitness contributions of ICE-encoded genes other than the T6SS. Indeed, as we noted in our discussion of the results from introducing P. vulgatus carrying the ICE into WildR-carrying mice, our data suggest that ICE genes outside the T6SS may be beneficial (lines 463-465). At the reviewer’s suggestion, we have reiterated the importance of considering the fitness contribution of genes beyond the T6SS in determining ICE distribution in the WildR community (line 527).

      Minor comments:

      (1) In lines 319 and 333, the manuscript refers to "Supplemental Figure 3F" and "Supplemental Figure 3G," respectively. However, the provided Supplemental Figure 3 appears to end at panel E. Please clarify or correct these references.

      We have modified the text to reference the correct figure panels.

      (2) Line 1043: The notation for "OD600" should be corrected for consistency and accuracy.

      The notation for OD600 has been updated to be consistent throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      Minor comments:

      (1) Line 144. I would be careful of using "strong correlation, in this sentence. Although it shows a higher correlation than lab mice. Also, the labels in Figure 1A for mouse WildRF7 are confusing and not well explained in the figure legend.

      We modified line 147 (new line in edited manuscript) to say “positive correlation” rather than “strong correlation” to better represent the result. We also revised the legend for Figure 1A to better explain the samples of WildR F7 that were analyzed.

      (2) Line 155. It's unclear which strains were isolated from the WildR community, and the reason for isolating only 15 strains. Also, Supplemental Figure 1 shows 17 isolates, not 15.

      We apologize for the confusion here. We isolated 17 strains, which is the number of distinct strains we were able to readily culture from this community. We obtained genome sequences for 15 of these, and were able to assemble a genome for one more of the strains from metagenomic data.

      (3) Line 236. Is it known what makes B. uniformis resistant to B. acidifaciens carrying a T6SSS? Does it have an orphan immunity protein?

      We do not know why B. uniformis is not targeted by B. acidifaciens under the conditions of our experiments. It does not encode homologs of the immunity genes bai1 or bai2, and does not appear to be intrinsically resistant to targeting by this T6SS given that it is effectively targeted by P. vulgatus carrying the ICE (Fig.4b).

      (4) Line 295. There is a consistent decline in C. acid abundance after 27 days in Figures 3B and 3C. How do you explain this? Is the endogenous B. acid expanding to outcompete C. acid exo since the total C. acid exo remains constant when gavaging 100x B. acid exo?

      We believe that the eventual decline in the introduced population of B. acidifaciens is likely due to a fitness cost imposed by the erm resistance marker we employed. We noted this phenomenon when describing the results depicted in Fig. 3F-H, but neglected to include this explanation earlier. This oversight has been corrected (lines 315-317).

      References

      Aoki, S. K., R. Pamma, A. D. Hernday, J. E. Bickham, B. A. Braaten and D. A. Low (2005). "Contact-dependent inhibition of growth in Escherichia coli." Science 309(5738): 1245–1248.

      Brown, E. M., H. Arellano-Santoyo, E. R. Temple, Z. A. Costliow, M. Pichaud, A. B. Hall, K. Liu, M. A. Durney, X. Gu, D. R. Plichta, C. A. Clish, J. A. Porter, H. Vlamakis and R. J. Xavier (2021). "Gut microbiome ADP-ribosyltransferases are widespread phage-encoded fitness factors." Cell Host Microbe 29(9): 1351-1365 e1311.

      Chatzidaki-Livanis, M., N. Geva-Zatorsky and L. E. Comstock (2016). "Bacteroides fragilis type VI secretion systems use novel effector and immunity proteins to antagonize human gut Bacteroidales species." Proc Natl Acad Sci U S A 113(13): 3627– 3632.

      Feng, L., A. S. Raman, M. C. Hibberd, J. Cheng, N. W. Griffin, Y. Peng, S. A. Leyn, D. A. Rodionov, A. L. Osterman and J. I. Gordon (2020). "Identifying determinants of bacterial fitness in a model of human gut microbial succession." Proc Natl Acad Sci U S A 117(5): 2622-2633.

      Garcia-Bayona, L., M. J. Coyne and L. E. Comstock (2021). "Mobile Type VI secretion system loci of the gut Bacteroidales display extensive intra-ecosystem transfer, multispecies spread and geographical clustering." PLoS Genet 17(4): e1009541.

      Hood, R. D., P. Singh, F. Hsu, T. Guvener, M. A. Carl, R. R. Trinidad, J. M. Silverman, B. B. Ohlson, K. G. Hicks, R. L. Plemel, M. Li, S. Schwarz, W. Y. Wang, A. J. Merz, D. R. Goodlett and J. D. Mougous (2010). "A type VI secretion system of Pseudomonas aeruginosa targets a toxin to bacteria." Cell Host Microbe 7(1): 25–37.

      Park, S. Y., C. Rao, K. Z. Coyte, G. A. Kuziel, Y. Zhang, W. Huang, E. A. Franzosa, J. K. Weng, C. Huttenhower and S. Rakoff-Nahoum (2022). "Strain-level fitness in the gut microbiome is an emergent property of glycans and a single metabolite." Cell 185(3): 513-529 e521.

      Russell, A. B., A. G. Wexler, B. N. Harding, J. C. Whitney, A. J. Bohn, Y. A. Goo, B. Q. Tran, N. A. Barry, H. Zheng, S. B. Peterson, S. Chou, T. Gonen, D. R. Goodlett, A. L. Goodman and J. D. Mougous (2014). "A type VI secretion-related pathway in Bacteroidetes mediates interbacterial antagonism." Cell Host Microbe 16(2): 227–236.

      Segura Munoz, R. R., S. Mantz, I. Martinez, F. Li, R. J. Schmaltz, N. A. Pudlo, K. Urs, E. C. Martens, J. Walter and A. E. Ramer-Tait (2022). "Experimental evaluation of ecological principles to understand and modulate the outcome of bacterial strain competition in gut microbiomes." ISME J 16(6): 1594-1604.

      Souza, D. P., G. U. Oka, C. E. Alvarez-Martinez, A. W. Bisson-Filho, G. Dunger, L. Hobeika, N. S. Cavalcante, M. C. Alegria, L. R. Barbosa, R. K. Salinas, C. R. Guzzo and C. S. Farah (2015). "Bacterial killing via a type IV secretion system." Nat Commun 6: 6453.

      Wexler, A. G., Y. Bao, J. C. Whitney, L. M. Bobay, J. B. Xavier, W. B. Schofield, N. A. Barry, A. B. Russell, B. Q. Tran, Y. A. Goo, D. R. Goodlett, H. Ochman, J. D. Mougous and A. L. Goodman (2016). "Human symbionts inject and neutralize antibacterial toxins to persist in the gut." Proc Natl Acad Sci 113(13): 3639–3644.

      Whitney, J. C., S. B. Peterson, J. Kim, M. Pazos, A. J. Verster, M. C. Radey, H. D. Kulasekara, M. Q. Ching, N. P. Bullen, D. Bryant, Y. A. Goo, M. G. Surette, E. Borenstein, W. Vollmer and J. D. Mougous (2017). "A broadly distributed toxin family mediates contact-dependent antagonism between gram-positive bacteria." Elife 6(Jul 11): e26938.

    1. eLife Assessment

      This study presents a valuable insight into the mechanism of action of With-No-lysine (K) kinases (WNK) in the insulin-dependent signaling in the context of learning and memory. The study uses solid methodological approaches that align with the current standards in the field. The revised manuscript, which includes diverse models ranging from mouse to cell systems assessed by a broad range of complementary techniques, convincingly argues that WNK signaling contributes to neuronal function in insulin-sensitive brain regions such as the hippocampus through its effects on GLUT4 trafficking.

    2. Reviewer #1 (Public review):

      [Editor's Note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. When experimentally feasible, the authors have adequately addressed the concerns of the reviewers in the revised manuscript to support the conclusions of the study.

      Summary:

      The study by Akita B. Jaykumar et al. explored an interesting and relevant hypothesis whether serine/threonine With-No-lysine (K) kinases (WNK)-1, -2, -3, and -4 engage in insulin-dependent glucose transporter-4 (GLUT4) signaling in the murine central nervous system. The authors especially focused on the hippocampus as this brain region exhibits high expression of insulin and GLUT4. Additionally, disrupted glucose metabolism in the hippocampus has been associated with anxiety disorders, while impaired WNK signaling has been linked to hypertension, learning disabilities, psychiatric disorders or Alzheimer's disease. The study took advantage of selective pan-WNK inhibitor WNK 643 as the main tool to manipulate WNK 1-4 activity both in vivo by daily, per-oral drug administration to wild-type mice, and in vitro by treating either adult murine brain synaptosomes, hippocampal slices, primary cortical cultures, and human cell lines (HEK293, SH-SY5Y). Using a battery of standard behavior paradigms such as open field test, elevated plus maze test, and fear conditioning, the authors convincingly demonstrate that the inhibition of WNK1-4 results in behavior changes, especially in enhanced learning and memory of WNK643-treated mice. To shed light on the underlying molecular mechanism, the authors implemented multiple biochemical approaches including immunoprecipitation, glucose-uptake assay, surface biotylination assay, immunoblotting, and immunofluorescence. The data suggest that simultaneous insulin stimulation and WNK1-4 inhibition results in increased glucose uptake and the activity of insulin's downstream effectors, phosphorylated Akt and phosphorylated AS160. Moreover, the authors demonstrate that insulin treatment enhances the physical interaction of the WNK effector OSR1/SPAK with Akt substrate AS160. As a result, combined treatment with insulin and the WNK643 inhibitor synergistically increases the targeting of GLUT4 to the plasma membrane. Collectively, these data strongly support the initial hypothesis that neuronal insulin- and WNK-dependent pathways do interact and engage in cognitive functions.

      In response to our initial comments, the authors mildly revised the manuscript, which did not improve the weaknesses to a sufficient level. Our follow-up comments are labeled under "Revisions 1".

      Strengths:

      The insulin-dependent signaling in the central nervous system is relatively understudied. This explorative study delves into several interesting and clinically relevant possibilities, examining how insulin-dependent signaling and its crosstalk with WNK kinases might affect brain circuits involved in memory formation and/or anxiety. Therefore, these findings might inspire follow-up studies performed in disease models for disorders that exhibit impaired glucose metabolism, deficient memory, or anxiety, such as Diabetes mellitus, Alzheimer's disease, or most of psychiatric disorders.

      The graphical presentation of the figures is of high quality, which helps the reader to obtain a good overview and to easily understand the experimental design, results, and conclusions.

      The behavioral studies are well conducted and provide valuable insights into the role of WNK kinases in glucose metabolism and their effect on learning and memory. Additionally, the authors evaluate the levels of basal and induced anxiety in Figures 1 and 2, enhancing our understanding of how WNK signaling might engage in cognitive function and anxiety-like behavior, particularly in the context of altered glucose metabolism.

      The data presented in Figures 3 and 4 are notably valuable and robust. The authors effectively utilize a variety of in vivo and in vitro models, combining different treatments in a clear manner. The experimental design is well-controlled, efficiently communicated, and well-executed, providing the reader with clear objectives and conclusions. Overall, these data represent particularly solid and reproducible evidence on the enhanced glucose uptake, GLUT4 targeting, and downstream effectors' activation upon insulin and WNK/OSR1 signaling crosstalk.

      Weaknesses:

      (1) The study used a WNK643 inhibitor as the only tool to manipulate WNK1-4 activity. This inhibitor seems selective; however, it has been reported that it exhibits different efficiency in inhibiting the individual WNK kinases among each other (e.g. PMID: 31017050, PMID: 36712947). Additionally, the authors do not analyze nor report the expression profiles or activity levels of WNK1, WNK2, WNK3, and WNK4 within the relevant brain regions (i.e. hippocampus, cortex, amygdala). Combined, these weaknesses raise concerns about the direct involvement of WNK kinases within the selected brain regions and behavior circuits. It would be beneficial if the authors provided gene profiling for WNK1, 2, 3, and -4 (e.g. using Allen brain atlas). To confirm the observations, the authors should either add results from using other WNK inhibitors or, preferentially, analyze knock-down or knock-out animals/tissue targeting the single kinases.

      Revisions 1: The authors added Fig. S1A during the revisions to show expression of Wnt1-4. While the expression data from humans is interesting, the experimental part of the study is performed in mice. It would be more informative for the authors to add expression profiles from mice or overview the expression pattern with suitable references in the introduction to address this point. The authors did not add data from knock down or knockout tissue targeting the single kinases.

      (2) The authors do not report any data on whether the global inhibition of WNKs affects insulin levels as such. Since the authors demonstrate the synergistic effect of simultaneous insulin treatment and WNK1-4 inhibition, such data are missing.

      Revisions 1: The authors added Fig. S5A to address this point. It is appreciated that authors performed the needed experiment. Unfortunately, no significant change was found, therefore, the authors still cannot conclude that they demonstrate a synergistic effect of simultaneous insulin treatment and WNT1-4 inhibition. It is a missed opportunity that the authors did not measure insulin in the CSF or tissue lysate to support the data.

      (3) The study discovered that the Sortilin receptor binds to OSR1, leading the authors to speculate that Sortilin may be involved in the insulin-dependent GLUT4 surface trafficking. The authors conclude in the result section that "WNK/OSR1/SPAK influences insulin-sensitive GLUT4 trafficking by balancing GLUT4 sequestration in the TGN via regulation of Sortilin with GLUT4 release from these vesicles upon insulin stimulation via regulation of AS160." However, the authors do not provide any evidence supporting Sortilin's involvement in such regulation, thus, this conclusion should be removed from the section. Accordingly, the first paragraph of the discussion should be also rephrased or removed.

      Revisions 1: The authors added Fig. 5M-N to address this point. The new experiment is appreciated. However, the authors still do not show that sortilin is involved in insulin or WNK-dependent GLUT4 trafficking in their set up since the authors do not demonstrate any changes in GLUT4 sorting or binding. The conclusions should therefore be rephrased or included purely in the discussion. Moreover, the discussion was not adjusted either, leading to over interpretation based on the available data.

      (4) The background relevant to Figure 5, as well as the results and conclusions presented in Figure 5 are quite challenging to follow due to the lack of a clear introduction to the signaling pathways. Consequently, understanding the conclusions drawn from the data is also difficult. It would be beneficial if the authors addressed this issue with either reformulations or additional sections in the introduction. Furthermore, the pulldown experiments in this figure lack some of the necessary controls.

      Revisions 1: The Authors insufficiently addressed this point during the revisions and did not rewrite the introduction as suggested.

      (5) The authors lack proper independent loading controls (e.g. GAPDH levels) in their immunoblots throughout the paper, and thus their quantifications lack this important normalization step. The authors also did not add knock-out or knock-down controls in their co-IPs. This is disappointing since these improvements were central and suggested during the revision process.

      (6) The schemes that represent only hypotheses (Fig. 1K, 4A) are unnecessary and confusing and thus should be omitted or placed at the end of each figure if the conclusions align.

      (7) Low-quality images, such as Fig. 5H should be replaced with high-resolution photos, moved to the supplementary, or omitted.

    3. Reviewer #2 (Public review):

      This study by Jaykumar and colleagues seeks to expand the field's appreciation of insulin responses in the brain, specifically by implicating WNK kinase function in various neuronal responses, ranging from behavioral / memory changes to GLUT4 trafficking to the cell surface with subsequent glucose uptake. This revised study is now comprehensive and presents a logical and reasonably documented cascade of molecular interactions responsible in part for GLUT4 trafficking under the regulation of WKK and insulin. Additional data allow the authors to dissect a plausible WNK/OSR1/SPAK-sortilin pathway for the modulation of GLUT4 trafficking, in part by capitalizing on an overlay of various techniques and systems. The data - much of it in vivo or ex vivo - showing a potential role for WNK function in brain glucose utilization remains a compelling part of the story, with the dissection of the signaling cascade and a potential role for sortilin in mediating WNK function via effects on GLUT4 cellular localization now more convincing.

      Initially, the group shows that oral WNK463 treatment - an inhibitor of WNKs broadly - in mice augments a number of memory readouts. These findings fit within the context of the overall story the authors present: that WNK function is critical to brain glucose utilization, which impacts learning. Multiple approaches are used to show that WNK463 treatment, i.e. inhibition of WNKs, increases glucose uptake, including labeled 2-deoxyglucose uptake in vivo in the brain and in isolated synaptosome, and uptake in ex vivo hippocampal slices. These findings are solid and consistent. With the exception of some relatively minor comments regarding the data presentation made to the authors and now fully addressed, the findings showing that WNK463 treatment increases GLUT4-mediated glucose uptake and surface localization of GLUT4 are reasonable, with the hippocampal slice data being particularly relevant.

      While the details of the WNK signaling cascade is dense, in the revised application one clearly appreciates the molecular interrogation and interactions the group is dissecting, supported by the use of multiple models. With the additional findings, these systems and the data now reinforce each other, presenting a strongly documented overall story.

      A limitation of the study with the initial submission was the authors' reliance upon a single pharmacological tool (WNK463) to inhibit WNK kinases. WNK463 apparently has substantial specificity for WNKs and WNK463 treatment lessened OSR1 phosphorylation (a WNK substrate). Nevertheless, the cohesiveness of the findings in terms of the broader pathway engagement (GLUT4 trafficking, glucose uptake) is consistent with the author's proposed mechanisms and conclusions. The authors have additionally addressed this concern in the revised manuscript with more information supporting the specificity of WNK463 as well as the multiple approaches to confirm the effect of WNK463 on the WNK signaling pathway of interest.

      The final few paragraphs of the discussion that weave the author's findings into the field more broadly, including Sortilin function and neurological disorders, are appreciated. Additional clarity in the Methods section is also helpful.

    4. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      (1) The study used a WNK643 inhibitor as the only tool to manipulate WNK1-4 activity. This inhibitor seems selective; however, it has been reported that it exhibits different efficiency in inhibiting the individual WNK kinases among each other (e.g. PMID: 31017050, PMID: 36712947). Additionally, the authors do not analyze nor report the expression profiles or activity levels of WNK1, WNK2, WNK3, and WNK4 within the relevant brain regions (i.e. hippocampus, cortex, amygdala). Combined, these weaknesses raise concerns about the direct involvement of WNK kinases within the selected brain regions and behavior circuits. It would be beneficial if the authors provided gene profiling for WNK1, 2, 3, and -4 (e.g. using Allen brain atlas). To confirm the observations, the authors should either add results from using other WNK inhibitors or, preferentially, analyze knock-down or knock-out animals/tissue targeting the single kinases.

      Revisions 1: The authors added Fig. S1A during the revisions to show expression of Wnk1-4. While the expression data from humans is interesting, the experimental part of the study is performed in mice. It would be more informative for the authors to add expression profiles from mice or overview the expression pattern with suitable references in the introduction to address this point. The authors did not add data from knock down or knockout tissue targeting the single kinases.

      Thank you for the excellent suggestion. We have added mouse in situ hybridization data curated from Allen Brain Atlas and found mRNA encoding WNK1 and WNK2 highly expressed in the hippocampus compared to WNK3 and WNK4. We also have included WNK1 knockdown data from cell lines (Figure S7A-F).

      Whole body WNK1 knockout is embryonically lethal, and we do not have access to brain tissue specific WNK knockout animal models. In addition, knockout of WNKs from brain tissue samples from animals is not very efficient from our experience and therefore, we included data from cell lines. In other non-neuronal cell lines, WNK1 knockdown replicates the effect of WNK463 (Figure S7A-D). However, in SHSY5Y cells, WNK1 knockdown did not replicate the effects of WNK463 on pAKT levels (Figure S7EF). This suggests tissue-specific effects of WNKs, and it also supports our suggestion that cooperativity among WNK family members is required in neuronal cells. This further supports our conclusion that WNK463 is an ideal tool to test our hypothesis in this study as it targets all 4 WNKs (WNK1-4) and furthermore, WNK463 was reported in the literature to inhibit only the four WNKs out of more than 400 kinases tested, indicating more selectivity than many small molecules used to target other enzymes.

      (2) The authors do not report any data on whether the global inhibition of WNKs affects insulin levels as such. Since the authors demonstrate the synergistic effect of simultaneous insulin treatment and WNK1-4 inhibition, such data are missing.

      Revisions 1: The authors added Fig. S5A to address this point. It is appreciated that authors performed the needed experiment. Unfortunately, no significant change was found, therefore, the authors still cannot conclude that they demonstrate a synergistic effect of simultaneous insulin treatment and WNT1-4 inhibition. It is a missed opportunity that the authors did not measure insulin in the CSF or tissue lysate to support the data.

      Thank you for the comment. As suggested, we tried to measure insulin in mouse hippocampal tissue lysate, and the levels fell way below the detectable range (78 - 5000 pg/mL) of the mouse insulin detection kit (Abcam: AB285341) used.

      (3) The study discovered that the Sortilin receptor binds to OSR1, leading the authors to speculate that Sortilin may be involved in the insulin-dependent GLUT4 surface trafficking. The authors conclude in the result section that "WNK/OSR1/SPAK influences insulin-sensitive GLUT4 trafficking by balancing GLUT4 sequestration in the TGN via regulation of Sortilin with GLUT4 release from these vesicles upon insulin stimulation via regulation of AS160." However, the authors do not provide any evidence supporting Sortilin's involvement in such regulation, thus, this conclusion should be removed from the section. Accordingly, the first paragraph of the discussion should be also rephrased or removed.

      Revisions 1: The authors added Fig. 5M-N to address this point. The new experiment is appreciated. However, the authors still do not show that sortilin is involved in insulin or WNK-dependent GLUT4 trafficking in their set up since the authors do not demonstrate any changes in GLUT4 sorting or binding. The conclusions should therefore be rephrased or included purely in the discussion. Moreover, the discussion was not adjusted either, leading to over interpretation based on the available data.

      Thank you for the suggestion. The conclusion has been rephrased as suggested.

      (4) The background relevant to Figure 5, as well as the results and conclusions presented in Figure 5 are quite challenging to follow due to the lack of a clear introduction to the signaling pathways. Consequently, understanding the conclusions drawn from the data is also difficult. It would be beneficial if the authors addressed this issue with either reformulations or additional sections in the introduction. Furthermore, the pulldown experiments in this figure lack some of the necessary controls.

      Revisions 1: The Authors insufficiently addressed this point during the revisions and did not rewrite the introduction as suggested.

      The background information related to figure 5 has been simplified as suggested. Response regarding the controls used is provided in the response to critique 5 as below.

      (5) The authors lack proper independent loading controls (e.g. GAPDH levels) in their immunoblots throughout the paper, and thus their quantifications lack this important normalization step. The authors also did not add knock-out or knock-down controls in their co-IPs. This is disappointing since these improvements were central and suggested during the revision process.

      GAPDH has been used as a loading control wherever applicable for Western blots on lysates (see Figures: 5D, 3F, 4E, 4C) In other cases, such as in Figure 5E, the analysis of pAS160 is normalized to total AS160 as this is more appropriate compared to GAPDH. For IP experiments such as Figure 5G, GAPDH is not an applicable control as we are using purified protein fragments in this case. For IP experiments (Figure 5K, 5L, 5C), IP proteins have been normalized to the input protein levels serving as a loading control for the IP because GAPDH is an intracellular protein which necessarily is not pulled down along with the proteins being IP’ed. Therefore, in this case, GAPDH is not a valid loading control. The choice of our loading controls used are very well supported by previous publications from our lab and other labs working on WNK pathways.

      (6) The schemes that represent only hypotheses (Fig. 1K, 4A) are unnecessary and confusing and thus should be omitted or placed at the end of each figure if the conclusions align.

      Thank you for the suggestion. Figure 1K is already at the end of the figure 1 and it shows the conclusion of that figure. Figure 4A have been placed at the end of the figures as suggested. Other schemes are only added at the end of the figures as suggested.

      (7) Low-quality images, such as Fig. 5H should be replaced with high-resolution photos, moved to the supplementary, or omitted.

      Thank you for your comment. The suggested images have been replaced with higher resolution ones.

      Reviewer #2 (Public review):

      This study by Jaykumar and colleagues seeks to expand the field's appreciation of insulin responses in the brain, specifically by implicating WNK kinase function in various neuronal responses, ranging from behavioral / memory changes to GLUT4 trafficking to the cell surface with subsequent glucose uptake. This revised study is now comprehensive and presents a logical and reasonably documented cascade of molecular interactions responsible in part for GLUT4 trafficking under the regulation of WKK and insulin. Additional data allow the authors to dissect a plausible WNK/OSR1/SPAK-sortilin pathway for the modulation of GLUT4 trafficking, in part by capitalizing on a overlay of various techniques and systems. The data - much of it in vivo or ex vivo - showing a potential role for WNK function in brain glucose utilization remains a compelling part of the story, with the dissection of the signaling cascade and a potential role for sortilin in mediating WNK function via effects on GLUT4 cellular localization now more convincing.

      Initially, the group shows that oral WNK463 treatment - an inhibitor of WNKs broadly - in mice augments a number of memory readouts. These findings fit within the context of the overall story the authors present: that WNK function is critical to brain glucose utilization, which impacts learning. Multiple approaches are used to show that WNK463 treatment, i.e. inhibition of WNKs, increases glucose uptake, including labeled 2deoxyglucose uptake in vivo in the brain and in isolated synaptosome, and uptake in ex vivo hippocampal slices. These findings are solid and consistent. With the exception of some relatively minor comments regarding the data presentation made to the authors and now fully addressed, the findings showing that WNK463 treatment increases GLUT4-mediated glucose uptake and surface localization of GLUT4 are reasonable, with the hippocampal slice data being particularly relevant.

      While the details of the WNK signaling cascade is dense, in the revised application one clearly appreciates the molecular interrogation and interactions the group is dissecting, supported by the use of multiple models. With the additional findings, these systems and the data now reinforce each other, presenting a strongly documented overall story.

      A limitation of the study with the initial submission was the authors' reliance upon a single pharmacological tool (WNK463) to inhibit WNK kinases. WNK463 apparently has substantial specificity for WNKs and WNK463 treatment lessened OSR1 phosphorylation (a WNK substrate). Nevertheless, the cohesiveness of the findings in terms of the broader pathway engagement (GLUT4 trafficking, glucose uptake) is consistent with the author's proposed mechanisms and conclusions. The authors have additionally addressed this concern in the revised manuscript with more information supporting the specificity of WNK463 as well as the multiple approaches to confirm the effect of WNK463 on the WNK signaling pathway of interest.

      The final few paragraphs of the discussion that weave the author's findings into the field more broadly, including Sortilin function and neurological disorders, are appreciated. Additional clarity in the Methods section is also helpful.

      Thank you for the positive response and acknowledging that we have satisfactorily addressed all of your critiques.

    1. eLife Assessment

      This manuscript presents a valuable computational tool for identifying 3-5 gene regulatory network topologies capable of generating oscillatory dynamics. The application of Monte Carlo Tree Search to circuit design is novel and effectively expands the scale at which non-linear behaviours can be explored in silico. The efficiency of the proposed algorithm is convincing, and the work will be of interest to the systems and synthetic biology communities. While the generality of the identified circuit properties is constrained by the simplifying modelling assumptions and parameter choices, the methodological contribution represents a significant advance in the field.

    2. Joint Public Reviews:

      This manuscript presents an algorithm for identifying network topologies that exhibit a desired qualitative behaviour, with a particular focus on oscillations. The approach is first demonstrated on 3-node networks-where results can be validated through exhaustive search-and then extended to 5-node networks, where the search space becomes intractable. Network topologies are represented as directed graphs, and their dynamical behaviour is classified using stochastic simulations based on the Gillespie algorithm. To efficiently explore the large design space, the authors employ reinforcement learning via Monte Carlo Tree Search (MCTS), framing circuit design as a sequential decision-making process.

      This work meaningfully extends the range of systems that can be explored in silico to uncover non-linear dynamics and represents a valuable methodological advance for the fields of systems and synthetic biology.

      Strengths:

      The evidence presented is strong and compelling. The authors validate their results for 3-node networks through exhaustive search, and the findings for 5-node networks are consistent with previously reported motifs, lending credibility to the approach. The use of reinforcement learning to navigate the vast space of possible topologies is both original and effective and represents a novel contribution to the field. The algorithm demonstrates convincing efficiency, and the ability to identify robust oscillatory topologies is particularly valuable. Expanding the scale of systems that can be systematically explored in silico marks a significant advance for the study of complex gene regulatory networks.

      Weaknesses:

      Although the proposed approach substantially expands the scale of tractable searches, the systems explored remain relatively small, being limited to five-node networks. The authors now discuss possible avenues for improving scalability, but extending the framework to substantially larger networks remains an important future challenge.

      Another important limitation concerns the assumption of identical reaction rates for all circuit connections. As the authors' own analysis shows, relaxing this assumption leads to significant qualitative and quantitative changes in oscillatory dynamics. Consequently, it remains unclear how the properties of the identified fault-tolerant oscillators translate to more biologically realistic regulatory circuits, where kinetic parameters vary across interactions.

      The conclusions should also be interpreted within the chosen modelling framework and parameter space. In particular, the sampled parameter ranges and restriction to relatively low Hill coefficients define the subset of regulatory architectures explored. Whether broader parameter regimes, including higher Hill coefficients, would reveal additional oscillatory architectures remains unclear.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      (1) The principal weakness of the manuscript lies in the interpretation of biological robustness. The authors identify network topologies that sustain oscillatory behaviour despite perturbations to the system or parameters. However, in many cases, this persistence is due to the presence of partially redundant oscillatory motifs within the network. While this observation is interesting and of clear value for circuit design, framing it as evidence of evolutionary robustness may be misleading. The “mutant” systems frequently exhibit altered oscillatory properties, such as changes in frequency or amplitude. From a functional cellular perspective, mere oscillation is insufficient — preservation of specific oscillation characteristics is often essential. This is particularly true in systems like circadian clocks, where misalignment with environmental cycles can have deleterious effects. Robustness, from an evolutionary standpoint, should therefore be framed as the capacity to maintain the functional phenotype, not merely the qualitative behaviour. 

      We agree with the reviewers that our framing conflated qualitative robustness (continued oscillation) with functional robustness (oscillation with preservation of properties such as frequency), and that the latter is a more meaningful definition of robustness in an evolutionary context. We have edited the manuscript to remove any suggestions that our results can explain the multiple-oscillator architecture of circadian clocks. We have also added a paragraph in the discussion that highlights the distinction between qualitative and functional robustness, as well as the other ways in which our training environment differs from an evolutionary context.

      Locations of changes:

      Abstract, second-to-last sentence

      Results, final section, first paragraph

      Discussion, first paragraph

      Discussion, new paragraph

      (2) A secondary limitation is that, despite the methodological advances, the scale of the systems explored remains modest. While moving from 3- to 5-node systems is non-trivial, five elements still represent a relatively small network. It is somewhat surprising that the algorithm does not scale further, particularly when considering the performance of MCTS in other domains — for instance, modern chess engines routinely explore far larger decision trees. A discussion on current performance bottlenecks and potential avenues for improving scalability would be valuable.

      We thank the reviewers for raising this important point. We have edited the manuscript to specify that we faced two distinct bottlenecks in scaling our experiments. The first is the runtime and scaling of the underlying Gillespie simulations, which become much more expensive as circuit size increases. The second is our use of the original (“vanilla”) MCTS algorithm, without the deep-learning value and policy networks that have driven the dramatic gains in domains such as Go and chess. We have added a new Discussion paragraph that identifies these two bottlenecks, followed by a paragraph on future methodological enhancements that goes into more detail about potential improvements to the algorithm. Using a power-law extrapolation of the current scaling, we estimate that without further methodological improvements the largest tractable network is approximately 7 nodes, while AlphaZero-style scaling could plausibly extend the approach to roughly 19-node circuits. We have also added an order-of-magnitude estimate of the 5-node search space (≈2×10<sup>9</sup> topologies) to give the reader a more concrete picture of the current scale.

      Locations of changes:

      Results, final section, end of paragraph 1 (Gillespie runtime as the practical bottleneck)

      Discussion, new paragraph 4 (bottlenecks and possible improvements)

      Discussion, new paragraph 5 (projected scaling under deep-learning-based extensions)

      Introduction, paragraph 4, and Discussion, paragraph 1 (search-space size)

      Methods, new section “Estimation of search space size”

      (3) It is worth noting that the emergence of oscillations in a model often depends not only on the topology but also critically on parameter choices and the nature of the nonlinearities. The use of Hill functions and high Hill coefficients is a common strategy to induce oscillatory dynamics. Thus, the reported results should be interpreted within the context of the modelling assumptions and parameter regimes employed in the simulations.

      We agree that the modeling assumptions substantially impact the interpretation of the results, and we have expanded the description of our modeling framework to make these assumptions explicit. To clarify, our model does not use Hill equations directly. Instead, cooperative binding is represented as a sequential, mass-action binding process. In addition to a new Methods section explaining our model in more detail, we have added a Methods section to show analytically that the effective Hill coefficient in our system is always ≤2. It also mentions an important practical benefit of using sequential binding rather than explicit TF dimerization, which is that it improves the size and scaling behavior of the reaction system. 

      Locations of changes:

      Results, section 2, paragraph 1 (clarification that no Hill function is imposed; effective Hill coefficient ≤2)

      Methods, section 1, new subsection “Sequential binding yields Hill coefficients ≤2”

      Recommendations for the authors:

      (4) It would be helpful to include the explicit reaction equations and corresponding reaction rates used in the simulations, to facilitate reproducibility and better understanding of the modelling assumptions.

      We have added the explicit reaction equations, the corresponding rate parameters, and the bounds used during random parameter sampling. The Results section describing our model now reports the rate parameters used in the simulations shown in the paper. Additionally the new subsection at the beginning of the Methods presents the full stochastic model. Finally, the bounds used for parameter sampling are now included in-line in the corresponding Methods subsection. The original tables of parameter values and sampling bounds (Tables 1 and 2) have been retained.

      Locations of changes:

      Results, section 2, paragraph 1 (rate parameters in main text; Table 1 retained)

      Methods, section 1, new subsection “Stochastic model of a transcription factor network”

      Methods, section “Random sampling” (explicit parameter bounds; Table 2 retained)

      (5) Sustained oscillations are notoriously difficult to observe in Gillespie simulations due to stochastic noise. Could the authors comment on whether simulation times were sufficiently long to distinguish sustained oscillations from transient dynamics?

      We agree this is an important methodological point. We have clarified that simulations were run for 11.1 hours. Because nearly all the oscillators we found had periods below 100 minutes and most oscillators had periods below 12 minutes, almost all were observed for at least 6.6 cycles, which we believe is sufficient to distinguish sustained oscillations from transient dynamics.

      Locations of changes:

      Results, section 2, paragraph 2

      (6) While the manuscript alludes to broader applications of the proposed method, it would be beneficial to elaborate on these possibilities. Clarifying how this tool could be extended to other types of dynamical behaviours or biological questions would strengthen the impact.

      We have rewritten the final paragraph of the Discussion to give concrete illustrative examples of how the framework can be extended beyond oscillator design, including the discovery of more complex design principles and the design of synthetic multicellular circuits such as multi-cell type cancer therapies and morphogenetic patterning circuits.

      Locations of changes:

      Discussion, final paragraph

      (7) It remains somewhat unclear whether the contribution lies primarily in the novel application of reinforcement learning or whether there are methodological innovations within the algorithm itself. Are there existing tools with similar objectives, and how does this work improve upon them in terms of performance or capabilities?

      We have edited the Discussion to state more directly that the principal contribution of this work is the novel application of reinforcement learning to the problem of network topology design problem, rather than a fundamentally new RL algorithm. We also contextualize CircuiTree relative to other topology-search approaches and articulate where the use of RL provides a concrete advantage in navigating large combinatorial search spaces.

      Locations of changes:

      Discussion, end of paragraph 3

      (8) Further details on the training of the reinforcement learning algorithm would be appreciated. Was it trained a priori, and if so, how much data was required?

      CircuiTree is not pre-trained. Like AlphaZero, it learns exclusively from simulations performed during the search itself, with no externally curated dataset or a priori domain knowledge. We have clarified this in two places in our manuscript.

      Locations of changes:

      Introduction, beginning of first paragraph

      Results, section 1, paragraph 4

      (9) Some discussion on how computational time scales with increasing network size would be valuable.

      As mentioned above, we have added (i) an order-of-magnitude estimate of the size of the 5-node search space (≈2×10⁹ topologies) to make the current scale concrete; (ii) a power-law extrapolation of the current algorithm’s scaling that suggests ≈7 nodes is the largest tractable network without further improvements; and (iii) a Discussion paragraph projecting that AlphaZero-style scaling, combined with faster simulators, could plausibly extend the approach to ≈19-node circuits. The methodology for estimating the search-space size is described in a new Methods section.

      Locations of changes:

      Introduction, paragraph 4 (search-space size)

      Discussion, paragraph 1 (search-space size)

      Discussion, new paragraph 4 (extrapolated scaling and 7-node bound)

      Discussion, new paragraph 5 (projected scaling under deep-learning extensions)

      Methods, new section “Estimation of search space size”

      (10) The discussion of knockouts that dampen or restore oscillations is interesting. Was this analysis performed systematically, or were the examples selected through observation? Clarifying the methodology here would add rigour to the interpretation.

      We have clarified that the topologies highlighted in this analysis were selected by observing the fault-tolerant properties of a few topologies that exhibited high mutational robustness in our screen, rather than via a systematic analysis of the screen results. The text now states this explicitly so the reader can interpret the examples accordingly.

      Locations of changes:

      Results, final section, paragraph 3

      Other changes made by the authors

      In addition to the changes above, we have made several minor edits to improve clarity and presentation. We corrected spelling and formatting errors throughout, and improved the clarity of language in the Discussion section. No changes were made to the Figures, Tables, Algorithms, or Supplementary Information.

      We are grateful to the reviewers for their input and believe the manuscript has been substantially improved as a result. We hope the revised version meets with the editors’ and reviewers’ approval.

    1. eLife Assessment

      This study reports the discovery of a new circuit mechanism for light-avoidance behavior in the marine annelid, Platynereis dumerilii. Using calcium imaging, molecular perturbations, behavioral measurements, and modeling, the authors provide compelling evidence that nitric oxide is released by postsynaptic neurons onto ciliary photoreceptors to prolong and enhance their response to ultraviolet light. The fundamental new role of nitric oxide described in this study may be conserved across animal phyla and thus will be of broad interests to neuroscientists and neuroendocrinologists.

    2. Reviewer #1 (Public review):

      Summary:

      The ciliary photoreceptor cells and its downstream neurons of larval annelid must be orchestrated in a specific pattern to promote downward swimming in response to long duration of UV exposure. The authors first conducted neuroanatomical examination of the circuit to identify NOS-expression neurons (INNOS) that are immediately downstream to the ciliary photoreceptor cells. The INNOS is activated by UV and produce NO. The NOS is required for UV avoidance by Platynereis larvae and neural dynamics of the photoreceptor cells and their downstream circuit. Following up the RNA-seq data with in-situ hybridization experiments, the authors found that two unconventional guanylate cyclases, NIT-GC1 and NIT-GC2, are expressed and localized in different subcellular domain of the photoreceptor cells. Experiments using the culture cells ang genetically encoded sensors demonstrated that NIT-GC1 can generate cGMP in response to nitric oxide. Finally, authors build mathematical model that fit the live imaging data and used it to predict how the magnitude of the photoreceptor activation varied by intensity and duration of UV light.

      Strengths:

      The authors conducted comprehensive interrogations of the UV avoidance pathway at the molecular and circuit levels and constructed mathematical model. The main conclusions are supported with layers of evidence from different assays.

      Weaknesses:

      The authors addressed these weaknesses in the previous version of the manuscript. Statistics are missing in both figure legends and methods. The perturbations of genes and molecules were not cell-type-specific and therefore the observed behavioral defect could be attributed to the malfunction of the circuit elsewhere not examined in this study. I suggest adding more explanation about the functions of other NOS-expressing cells and conducting a control experiment to test behavioral response to a non-visual stimulus.

    3. Reviewer #2 (Public review):

      Summary:

      This study is quite thorough, tackling this NO-dependent UV avoidance circuit with both breadth and depth. There are several novel discoveries throughout, but the whole package represents perhaps even more than the sum of these parts.

      Strengths:

      The presentation of the work is compelling. The introduction sets up the question and the state of the field very nicely. The discovery of the non-canonical NO receptor pathway in the ciliary photoreceptors is fascinating and will likely open up new avenues for future research into NO-pathways in different species. The use of genetic and pharmacological manipulations of circuit components was well thought-out. The authors applied different experimental techniques expertly throughout the study so that they could develop a comprehensive view from the molecular to the behavioral levels.

      Weaknesses:

      The authors have done an excellent job revising and explaining their model. No important weaknesses remain, in my opinion.

    4. Reviewer #3 (Public review):

      The transition from planktonic to benthic depends upon several physical and chemical cues. Nitric oxide (NO) is known as a critical player in the induction of larval metamorphosis in several invertebrates. Although NO is a widespread signalling molecule in a broad range of organisms regulating key physiological processes, internal regulatory mechanisms studies are scarce. While the UV sensing in larvae of the annelid Platynereis dumerilii using ciliary photoreceptors has been studied, the neuronal signalling mechanism remains unknown. In this study, Kei Jokura et al. investigated how annelid Platynereis dumerilii larvae detect UV sensing and modulate swimming behaviour through nitric oxide feedback. Using existing resources of Platynereis larval connectome/volume EM data, they identified NOS-expressing interneurons within the ciliary photoreceptors circuit (cPRCs). They demonstrated that NO is produced in cPRCs during UV/violet stimulation by using a fluorescent NO-reporter line. Further, they demonstrated that Nitric oxide signalling mediates UV-avoidance behaviour by using NOS-mutant larvae. Finally, they mapped out the signalled mechanisms of the cPRC circuit using published spatially mapped single-cell transcriptome data of Platynereis larvae, the Ca sensor lines, in situ HCR, and immunostaining. Additionally, by using their findings from Ca imagining data of cPRC, INNOS and INRGWa cells collected in wild-type, NOS knockout and NIT-GC2 morphant larvae, Kei Jokura et al. developed a mixed cellular-circuit-level mathematical model. However, my expertise in mathematical modelling is limited, so I cannot comment on this section.

      Comments on revised version.

      Thank you for the opportunity to re-evaluate this manuscript. I have reviewed the authors' responses and the revised manuscript. The authors have carefully and satisfactorily addressed all of my previous comments and concerns. The revisions have strengthened the paper, and I have no further suggestions.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      The ciliary photoreceptor cells and its downstream neurons of larval annelid must be orchestrated in a specific pattern to promote downward swimming in response to long duration of UV exposure. The authors first conducted neuroanatomical examination of the circuit to identify NOS expression neurons (INNOS) that are immediately downstream to the ciliary photoreceptor cells. The INNOS is activated by UV and produces NO. The NOS is required for UV avoidance by Platynereis larvae and neural dynamics of the photoreceptor cells and their downstream circuit. Following up the RNA-seq data with in situ hybridization experiments, the authors found that two unconventional guanylate cyclases, NIT-GC1 and NIT-GC2, are expressed and localized in different subcellular domain of the photoreceptor cells. Experiments using the culture cells and genetically encoded sensors demonstrated that NIT-GC1 can generate cGMP in response to nitric oxide. Finally, authors build a mathematical model that fit the live imaging data and used it to predict how the magnitude of the photoreceptor activation varied by intensity and duration of UV light.

      Strengths:

      The authors conducted comprehensive interrogations of the UV avoidance pathway at the molecular and circuit levels, and constructed a mathematical model. The main conclusions are supported by layers of evidence from different assays.

      Weaknesses:

      Statistics are missing in both figure legends and methods. The perturbations of genes and molecules were not cell-type-specific and therefore the observed behavioral defect could be attributed to the malfunction of the circuit elsewhere not examined in this study. I suggest adding more explanation about the functions of other NOS-expressing cells and conducting a control experiment to test behavioral response to a non-visual stimulus.

      Thank you for this assessment of our work. We have now added additional panels with statistical tests to the figures and included explanatory text in the figure legends.

      Regarding the cell-type-specific effects, we would like to offer a more nuanced view of this. Some of the genes we studied (NIT-GC1 and NIT-GC2) are only expressed in 4 cells in an organism of ~10,000 cells (as we demonstrated by HCR, immunostainings and the analysis of single-cell RNAseq data) and we knocked-down these genes (validated by antibody staining) with two independent morpholinos (to be able to rule out off-target effects). This is as cell-type specific as it gets. NOS is also expressed in a very limited number of cells. In the larval stages, we detected NOS expression only in the four INNOS cells and the pigmented eyes. We could previously show that UV avoidance is only mediated by the cPRC circuit (including INNOS) whereas phototaxis is exclusively mediated by the pigmented eyes Verasztó et al. (2018). Due to this behavioural specificity, we are therefore confident that the effects of the NOS mutations on UV avoidance are due to the lack of NOS from the INNOS cells and not the pigmented eyes. Besides UV avoidance, we have characterised the speed of ciliary swimming without a light stimulus, as well as phototaxis []. In addition, we also tested phototaxis and detected a reduced phototactic reaction in NOS mutants, but not after chemically inhibition of NO production (Figure 3B and 3F). The additional effects of NO are thus well documented in the paper and independent of the function of NO in the UV reaction. In addition, we have now did further quantifications and added new data (Figure 3 – figure supplement 4) to show that NOS mutants show normal lunar periodicity of sexual maturation, similar to wild-type animals. NOS mutants are viable and fertile, are feeding, building tubes and are mating as wild-type animals (these behaviours were not quantified here, we only show the data for lunar periodicity).

      Reviewer #2 (Public Review):

      Summary:

      This study is quite thorough, tackling this NO-dependent UV avoidance circuit with both breadth and depth. There are several novel discoveries throughout, but the whole package represents perhaps even more than the sum of these parts.

      Strengths:

      The presentation of the work is compelling. The introduction sets up the question and the state of the field very nicely. The discovery of the non-canonical NO receptor pathway in the ciliary photoreceptors is fascinating and will likely open up new avenues for future research into NO pathways in different species. The use of genetic and pharmacological manipulations of circuit components was well thought-out. The authors applied different experimental techniques expertly throughout the study so that they could develop a comprehensive view from the molecular to the behavioral levels.

      Weaknesses:

      While I appreciate the intent of bringing together a large set of measurements from connectomics and calcium imaging in the framework of a model, the model seemed rather poorly constrained. How many parameters are in the model shown in Figure 6A? How many of them are well constrained by experimental measurements? The authors also don’t perform sensitivity analysis on the parameters of the model. And ultimately, the conclusion over the model in Figure 7 is somewhat trivial within the unitless construction: larger amplitude and longer duration stimuli lead to increased activation of the downstream neuron thought to lead to the downward swim behavior. I could imagine that a large family of models would arrive at this same result, and without units, there is no way to really test it with new behavioral experiments.

      We thank the reviewer for these comments. We have now thoroughly revised the model based on new experiments and carried out a sensitivity analysis. We are also more positive about the usefulness of the model though, for the following reasons.

      General usefulness of the model: With the model we can now reproduce all the qualitative dynamics of the circuit. The modelling also completely changed the way we were thinking about the system. For example, we needed to include a time-limited step in the cPRC transduction cascade leading to NOS activation to capture the time-invariance of the peak. In the future, we can also use this model to generate prediction e.g. about what the response to repeated stimulations can be. In the revised version, we included a new series of measurements of calcium dynamics in the cPRCs under varying duration and intensity of UV light. These important new results gave us further insights and necessitated a revision of the model.

      Unitless model: The model is indeed unitless, since already our input data from calcium imaging represent normalised data and the model was fit to these data. We would need a lot more information to build up a proper ground-up biophysical model (e.g. including capacitance, ionic concentrations etc.). This does not mean that the model is not useful.

      Sensitivity analysis: We have created a pipeline to carry out local and global sensitivity analysis and carried out a variance-based sensitivity and identifiability analysis. These data and code are included in the revised version. In the model, most parameters are identifiable. If we fix only two of the parameters all other parameters can be identified.

      Constrains and lots of possible models: In terms of the constrains, since we don’t have dimensioned quantitites, our constraints are bounds on the parameters. However, the structural identifibiality analysis tells us where we can identify parameters and can guard against sloppiness and overfitting. We only fitted the model on a subset of the data and can reproduce dynamics on a larger set of data. For important cellular interactions, the signs of the interactions are constrained. The model thus also allows us to rule out a large number of models - e.g. we could rule out a very simple model of progression from step to step as it was not possible to fit the data to such a model.

      We have updated the text to reflect these changes, e.g.: “Many of the couplings in our model are constrained (e.g. UV leads to INNOS activation) and e.g. reversing the sign of some of these couplings would not arrive at the same result. Several earlier variants of the model could not be fit to the data, The model is thus well constrained by our physiological experiments and the 106 circuit map.”

      Reviewer #3 (Public Review):

      The transition from planktonic to benthic depends upon several physical and chemical cues. Nitric oxide (NO) is known as a critical player in the induction of larval metamorphosis in several invertebrates. Although NO is a widespread signalling molecule in a broad range of organisms regulating key physiological processes, internal regulatory mechanisms studies are scarce. While the UV sensing in larvae of the annelid Platynereis dumerilii using ciliary photoreceptors has been studied, the neuronal signalling mechanism remains unknown. In this study, Kei Jokura et al. investigated how annelid Platynereis dumerilii larvae detect UV sensing and modulate swimming behaviour through nitric oxide feedback. Using existing resources of Platynereis larval connectome/volume EM data, they identified NOS-expressing interneurons within the ciliary photoreceptors circuit (cPRCs). They demonstrated that NO is produced in cPRCs during UV/violet stimulation by using a fluorescent NO-reporter line. Further, they demonstrated that Nitric oxide signalling mediates UV-avoidance behaviour by using NOS-mutant larvae. Finally, they mapped out the signalled mechanisms of the cPRC circuit using published spatially mapped single-cell transcriptome data of Platynereis larvae, the Ca sensor lines, in situ HCR, and immunostaining. Additionally, by using their findings from Ca imagining data of cPRC, INNOS and INRGWa cells collected in wild-type, NOS knockout and NIT-GC2 morphant larvae, Kei Jokura et al. developed a mixed cellular-circuit-level mathematical model. However, my expertise in mathematical modelling is limited, so I cannot comment on this section.

      No doubt, the study has been conducted extensively. However, I have a few comments, please see below.

      Page 4: “In contrast, both two- and three-day-old homozygous NOS-mutant larvae showed a strongly diminished UV avoidance response (Figure 3A, B and Figure 3-figure supplement 1B, C).” Instead of using subjective terms like “strongly,” it would be more relevant to provide statistical values. However, I could not locate any means of statistical analysis on larval behaviour. Can the authors indicate the statistical values for all behaviour studies?

      We thank the reviewer for these comments. We have changed the wording and also added the results of statistical analyses to the figures and explanations to the figure legends.

      Page 5: “(D) Vertical displacement in 30 sec bins of wild type and mutant (NOSΔ11/Δ11 and NOSΔ23/Δ23) three-day-old larvae stimulated with 395 nm light from the side, 488 nm light from the top and 395 nm light from the top.” The error bars for WT are too long at the end of the experiment. It is not clear how the authors decided to use this time frame. Did the authors try carrying this out for an extended time period? How did the authors decide on 120 seconds as the time frame for exposure? Authors should provide data on larval behaviour for an extended time.

      The 120 seconds time frame of exposure takes into account the reaction time and swimming speed of the larvae as well as the size of the assay chamber (160 mm water height).

      By the end of a 120 stimulation many larvae tend to accumulate at the bottom of the chamber due to downward swimming and cannot be further tracked. This effect leads to higher variability in the data towards the end of the experiment in the wild-type batches. We have showed both continuous vertical data as well as the binned data. The 30 sec bin was chosen for convenience and for better comparison with our previous paper on UV avoidance behaviour (Verasztó et al. 2018)

      Page 13: “During the UV response, prototroch cilia beat slower than trunk cilia, resulting in a head down stable state (‘rear-wheel drive’). In contrast, during the pressure response prototroch cilia beat faster than trunk cilia, leading to a head-up orientation (‘front-wheel drive’). Testing this hypothesis will require biophysical experiments and mathematical modelling.” Authors should carry out ciliary beating analysis under UV light in the current study with NOS mutant larvae. Since the pressure and UV detection systems are closely related, comparing the difference in ciliary beating is important to 155 demonstrate this hypothesis. Further, did the authors check the Ca sensor GCaMP6s under pressure conditions?

      We thank the reviewer for this suggestion. We have carried out further experiments and analysed the ciliary beat frequency (CBF) of larvae exposed to UV stimulation. We added these data to Figure 3—figure supplement 3. The results (increase of CBF under UV in wild-type but not NOS mutant larvae) were quite surprising to us and falsified our initial hypothesis. We have rewritten the discussion to reflect this important new finding.

      The response of cPRC cilia to changes in hydrostatic pressure has been extensively documented in our recent paper on the mechanism of barotaxis in the Platynereis larva (see Bezares-Calderón et al., https://doi.org/10.1101/2023.02.28.530398).

      Page 18: “strips. One strip contained UV (395 nm) LEDs (SMB1W-395, Roithner Lasertechnik) and the other infrared (810 nm) LEDs (SMB1W-810NR-I, Roithner Lasertechnik).” Authors should test larval swimming behaviour at different wavelengths. Even though they are performed in previous work, the experiment with different wavelengths is necessary to be conducted in NOS mutant larvae in parallel with a control. This will confirm that NOS is principally associated with UV. Further, to demonstrate that this mechanism is associated with ciliary movement, authors need to provide this evidence.

      The diving reaction by non-directional light can only be induced by UV/cyan light, as we have shown previously, and it is mediated by a single UV-opsin photopigment (c-opsin1). The avoidance experiments can thus only be done with UV/cyan light. We also measured swimming 174 behaviour with 480 nm directional light to test phototaxis (Figure 3D and Figure 3—figure 175 supplement 1F).

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      (1) The current introduction focuses on the nitric oxide signaling. It would be helpful for readers to have an introductory section about Platynereis larvae (a total number of neurons, etc) and rationales to use its nervous system as a model to study mechanisms of NO signaling and gating of visual response.

      We added an extra paragraph to give more detail on the number of cells in the larva and why Platynereis larvae are interesting to study to understand how synaptic and volume signalling interact. “The 3-day-old larva has over 9,000 cells classified into 202 neuronal and 92 non neuronal cell types (Verasztó et al., 2025). The synapticly connected subset of the cells in the body form a connectome of over 2,000 cells. Besides synapses, neurons in the larva also signal by volume transmission mediated by a rich repertoire of neuropeptides (Williams et al., 2017) and other modulators (Bauknecht and Jékely, 2017). The transparent and experimentally accessible Platynereis larvae could therefore inform how synaptic and volume signalling interact to mediate behaviour (Jékely and Yuste, 2024).”

      (2) The readers would want to know why these larvae swim downward in response to UV and upward to 480nm light. Is that for maintaining the certain depth from the surface of the water?

      It has been suggested that the ciliary photoreceptor circuit, which senses UV light, and the rhabdomeric photoreceptor circuit, which senses blue light, exchange messages with each other and that the two work together to form a depth gauge. By allowing larvae to swim at their preferred depth, the depth gauge influences where they end up when they become adults.

      (3) Are there splicing isoforms of NOS in the Platynereis dumerilii? If so, do antibodies and probes for in situ distinguish them? In Drosophila, truncated isoforms can inhibit the function of full length isoform, and therefore it was important to use methods to distinguish spicing isoforms.

      We did not identify any alternatively spliced forms of Platynereis NOS in our published transcriptome resources.

      (4) “INRGWs” acronym appears without explanation in the first paragraph of the results.

      Corrected: “the INRGWa neurons (cholinergic interneurons that express an RGW neuropeptide)”

      (5) In Figure 1C, it is difficult to see the projection patterns of individual cell types. Figure supplement 3 can be combined with the current Figure 1.

      We have moved one panel from Figure supplement 3 to the main Figure 1 (panel D) to show the INNOS projections more clearly.

      (6) Add more explanation about NOSp::palmi-3xHA reporter. Is it membrane-targeted reporter with the upstream promotor sequence of NOS? Or is it endogenous NOS that was tagged with palmi-3xHA?

      It is a membrane-targeted reporter driven by the upstream promoter sequence of the NOS gene. The construct is delivered by plasmid injection. We have clarified the description in the text and the figure legend (“The four apical organ cells, but not the eyes, were also labelled with a transiently expressed NOS-reporter transgene. This transgene contains the upstream promoter sequence of the NOS gene that drives a membrane-targeted palmitoylated tdTomato reporter (Figure 1F).” and “Expression of a membrane-targeted reporter driven by the NOS regulatory region (NOSp::palmi-3xHA-Tomato; magenta”). More details are in the Methods section under Transient transgenesis.

      (7) Show lack of anti-NOS immunostaining in NOS mutant to warrant specificity of the antibody. The subcellular localization of NOS in the dendritic arbors of INNOS is essential for the proposed model. The immunostaining image in Figure 4 -figure supplement 3D can be in the main figure. Is NOS also in the axons of INNOS?

      We further optimised the immunostaining with the NOS antibodies in WT and NOS mutant and added the staining data to the main Figure 1G and Figure 3—supplement 1. NOS was observed to be clearly localised in a region in the neuropil corresponding to the dendritic site of the INNOS cells. The staining was completely absent in larvae of both NOS knockout alleles. We have also added these explanations to the text.

      (8) “that was defective in NOS mutant (Figure 5I)” should be corrected as “Figure 5H”.

      We have restructured this part of the text and figure, the data from the Ser-h1 cells are now in Figure 5 - figure supplement 2.

      (9) Related to Figure 7, how do the larvae respond to a sequence of UV light (e.g. ten times 0.5s ON and 0.5 OFF) or ramping up/down UV light? Can the model make any predictions?

      We carried out new experiments and generated new model predictions. In Figure 6 – figure supplement 7 we show how changing the amplitude and duration of a single UV stimulation influences the response and the model output. In Figure 6 – figure supplement 8 we show how the model behaves when we provide repeated stimulations.

      (10) What are the knock down efficiency of NOS and NIT-GC morpholnio?

      We have now quantified the fluorescence intensity by immunostaining in wild-type and morphant larvae. We show the data for NIT-GC1 and NIT-GC2 in Figure 4—supplement 3F. The knock downs are very efficient. For NOS, we did not do morpholino experiments since we have two null alleles (Figure 3 – figure supplement 1).

      Reviewer #2 (Recommendations For The Authors):

      Why is the behavior of the control animals so different in panels B and C of Figure 3? One group reaches only 15 mm vertical position and the other reaches ~60 mm. Is this just batch variation? Am I missing an experimental variable here? If this is indeed batch variation, then some additional text and statistical analyses might help the reader interpret the behavioral data.

      Thank you for pointing this out. This was a mistake in the original Figure 3C of the unit on the Y axis. We have now corrected this. Additionally we have added statistical tests.

      Really, that’s the only, relatively minor issue I could find. This was a pleasure to read, and I learned a lot. Congratulations on an excellent study.

      Thanks a lot for these comments.

      Reviewer #3 (Recommendations For The Authors):

      Page 3, Figure 1 A: The authors reconstructed the cPRC circuit in 3-day-old larvae and detected NOS gene expression in 2-day-old larvae in Figure 1D&E. Can the authors provide a cPRC circuit reconstruction for 2-day-old larvae?

      We do not have a full connectome of the 2-day-old larva. We also show NOS gene expression in three-day-old larvae (Figure 1 – figure supplement 2).

      Page 4, Figure 2: NO produced by UV/violet stimulation to cPRCs: Including a diagram of the larvae would enhance reader understanding.

      We have added a schematic diagram to Figure 2.

      Page 4: Two Platynereis NOS knockout lines (NOSΔ11/Δ11 and Δ23/Δ23) using the CRISPR/Cas9: Could you direct me to the knockout conformation results for NOS knockout lines NOSΔ11/Δ11 and Δ23/Δ23?

      This is shown in Figure 3 – figure supplement 1 (genetic deletion and loss of antibody signal).

      Page 4: “Three day-old but not two-day-old NOS-mutant larvae also showed reduced phototactic behaviour, suggesting a function for NOS in the visual eyes that mediate three-day-old phototaxis” This sentence is unclear. Why is it only three days old but not two days old?

      2-day-old and 3-day-old larvae have very different type of phototaxis. 2-day-old larvae use their eyspots and show non-visual helical phototaxis. 3-day-old larvae use their visual (‘adult’) eyes for visual phototaxis. We have clarified this sentence: “Given that phototaxis in 1 and 2-day-old trochophore larvae is mediated by their non-visual eyespots (Jékely et al., 2008) and in 3-day-old nectochaete larvae by the visual eyes (Randel et al., 2014), these data suggest a function for NOS in the visual eyes (Figure 3D and Figure 3—figure supplement 1E).”

      Page 5: “Figure 3. NOS is required for UV avoidance in Platynereis larvae.” Detailed experiment setup, if possible, schematic or real setup images would help to replicate the experiments.

      We have added a schematic diagram to Figure 3. 

      Page 5: “All trajectories start at 0 x and y position and time 0 corresponding to 10 sec after the onset of 395 nm stimulation from the side.” How did the authors determine the 10-second duration? Are there any reasons for this choice?

      In our experimental setup we had a limit of tracking individual larvae of approximately 40 seconds. We therefore restricted our analysis to a 10 sec pre and 30 sec post-stimulus interval.

      “(A) Swimming trajectories of wild type (WT, n=32) and NOS mutant (NOSΔ11/Δ11, n=26 and NOSΔ23/ Δ23, n=47) three-day-old larvae.” Authors keep shifting between 2-day and 3-day-old larval data in Figures 1, 2, and 3, causing inconsistency.

      Due to their elongated shape and active muscular contractions it is difficult to carry out calcium imaging experiments with 3-day-old larvae. These experiments were therefore done with 2-day-old larvae. However, we did our behavioural experiments with both 2- and 3-day-old larvae. We observed similar patterns of UV-avoidance behaviour, NOS gene expression, ciliary activity, and mutant phenotype. We are therefore confident that the cPRC responses and circuit activity are similar across these two stages.

      Page 5: “Figure 3. NOS is required for UV avoidance in Platynereis larvae.” Why didn’t the authors present any statistics on the plots? A statistical test is required to prove that the difference is significant.

      We have added statistical tests to the data in Figure 3.

      Page 5: “Analysis of sGCs in Platynereis indicated that these genes are not expressed in any of the cells of the cPRC circuit (not shown and (Verasztó et al., 2017)).” Hence the data is relevant; please provide this data in supplementary.

      We re-analysed previous single-cell data (Achim et al., 2018) and found no sGC homologues detected in cPRC and INNOS. However, we detected an sGCβ subunit in the INRGWa cells. We added these data to the source data and amended the figure and the text. “Analysis of sGCs in Platynereis showed that these genes were not expressed in cPRC or INNOS cells, we only detected expression of an sGCβ subunit in the INRGWa cells (Figure 6B).”

      Page 10: “Diagram of the mathematical model with the components, interactions, parameters and equations used to model Ca dynamics.” It would be helpful to include a legend explaining the meaning of each term (e.g. what is K, Co, Cp etc) and arrow color. Additionally, the main text should provide detailed descriptions to aid understanding.

      We have added a new diagram of the model and its parameters to Figure 6. In addition, in the Methods section we have an extensive description of the model and all its parameters.

      Page 12: “This activated state is maintained for several tens of second.” Can authors point out the data?

      We point to the relevant figure now. “The high-Ca2+-state is then maintained for several tens of second (Figure 4A, B).”

      Page 13: “In the Platynereis circuit, our mathematical model indicates that the magnitude of the NO-dependent signal depends on the intensity and duration of the UV/violet stimulus.” Did the authors conduct experiments of various intensity and duration apart from modelling? If not, such experiments should be carried out to validate the mathematical model.

      We have carried out new calcium imaging experiment with varying duration and intensity of the stimulation. These experiments were key to the revision of the model because they clearly pointed to the time-invariance of the NO-dependent peak in the cPRCs. These data and their model fits are summarised in Figure 6 – figure supplement 7.

      References

      Bauknecht P, Jékely G. 2017. Ancient coexistence of norepinephrine, tyramine, and octopamine signaling in bilaterians. BMC Biology 15. doi:10.1186/s12915-016-0341-7

      Jékely G, Colombelli J, Hausen H, Guy K, Stelzer E, Nédélec F, Arendt D. 2008. Mechanism of phototaxis in marine zooplankton. Nature 456:395–399. doi:10.1038/nature07590

      Jékely G, Yuste R. 2024. Nonsynaptic encoding of behavior by neuropeptides. Current Opinion in Behavioral Sciences 60:101456. doi:10.1016/j.cobeha.2024.101456

      Randel N, Asadulina A, Bezares-Calderón LA, Verasztó C, Williams EA, Conzelmann M, Shahidi R, Jékely G. 2014. Neuronal connectome of a sensory-motor circuit for visual navigation. eLife 3. doi:10.7554/elife.02730

      Verasztó C, Gühmann M, Jia H, Rajan VBV, Bezares-Calderón LA, Piñeiro-Lopez C, Randel N, Shahidi R, Michiels NK, Yokoyama S, Tessmar-Raible K, Jékely G. 2018. Ciliary and rhabdomeric photoreceptor-cell circuits form a spectral depth gauge in marine zooplankton. eLife 7. doi:10.7554/elife.36440

      Verasztó C, Jasek S, Gühmann M, Bezares-Calderón LA, Williams EA, Shahidi R, Jékely G. 2025. Whole-body connectome of a segmented annelid larva. eLife 13. doi:10.7554/elife.97964.3

      Williams EA, Verasztó C, Jasek S, Conzelmann M, Shahidi R, Bauknecht P, Mirabeau O, Jékely G. 2017. Synaptic and peptidergic connectome of a neurosecretory center in the annelid brain. eLife 6. doi:10.7554/elife.26349

    1. eLife Assessment

      This manuscript applies a theoretical analysis to two published datasets on yeast and bacterial evolution to compare different ways of quantifying fitness. It makes an important advance by clarifying how discrepancies can arise by using different approaches and provides recommendations for best practices. Overall, this is an impressive and highly beneficial study that is based on convincing evidence and has the potential of setting standards in this rapidly growing field.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      The authors point out that the fitness estimates obtained from different experimental assays (monoculture, pairwise competition or bulk competition) are not generally equivalent, not even with regard to the fitness ranking of different genotypes. Using a computational model based on experimentally measured growth phenotypes for knockout strains in yeast, as well as data from Lenski's Long Term Evolution Experiment (LTEE), they derive a set of best practice rules aimed at extracting the optimal amount of information from such experiments.

      The study is very complete on a technical level, and the conceptual weaknesses raised in the first round of reviews have been fully addressed in the revision.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript "Quantifying microbial fitness in high-throughput experiments" provides a comprehensive analysis of the various approaches to quantifying fitness in microbial evolution, focusing on three primary factors: encoding of relative abundance, time scale of measurement, and the choice of reference subpopulation. The authors systematically explore how these choices impact fitness statistics and provide recommendations aimed at standardizing practices in the field. This manuscript aims to highlight the impact of differing fitness definitions and the methodologies utilized for analysis and how that can significantly alter interpretations of mutant fitness, affecting evolutionary predictions and the overall understanding of genetic interactions in the experiments.

      Strengths:

      The choices for quantifying fitness in evolution experiments are critical and highly relevant given the increasing prevalence of high-throughput experiments in evolutionary biology. The authors methodically categorize fitness statistics and their implications, providing clarity on a complex subject. This structured approach aids in understanding the nuances of fitness measurement. The manuscript effectively highlights how different choices in fitness measurement can influence fitness rankings and the understanding of epistasis, which is important for modeling evolutionary dynamics.

      Comments on revisions:

      The authors have comprehensively addressed all previous comments and suggestions. In particular, the addition of the new methods section: 'A guide to calculate pairwise relative fitness under the logit encoding from bulk competition data' - significantly improves the clarity of the implementation and helps in the overall interpretation of the framework.

    4. Reviewer #3 (Public review):

      Summary:

      The authors present analyses of different fitness measures derived from empirical data from yeast knock-out mutants and the long-term evolution experiment (LTEE) with Escherichia coli to explore discrepancies and identify preferred methods to estimate relative fitness in high-throughput experiments. Their work has three components. They first discuss the different "encodings" of relative abundance data and conclude that logit-transformations are preferred, because they transform nonlinear abundance trajectories into linear trajectories with greater predictive power. Next, they compare per-generation with per-growth cycle relative fitness estimates inferred from simulations of pairwise competitions based on published growth traits for the yeast strains and on published pairwise competition measurements for the LTEE data. Both data sets show quantitative and qualitative (i.e. rank order) discrepancies of estimates across different time scales, which are highlighted by considering possible underlying causes (i.e. trade-offs between growth traits) and consequences (i.e. epistasis among mutations affecting different growth traits). Finally, the authors compare simulated pairwise and bulk (i.e. where many mutants compete during a growth cycle in a single environment) competition assays based on the yeast knock-out mutants and demonstrate an optimal ratio of collective mutants to wild-type strains that minimizes both sampling error and overestimation of fitness estimates when compared with pairwise competitions.

      Strengths:

      The study deals with a highly relevant topic. Fitness is central to general evolutionary theory, but also poorly defined and implies different traits for different organisms and conditions. For microbes, which are often used in evolution experiments, high-throughput experiments may yield different measures to quantify abundance over time, from individual growth traits to bulk competition experiments. Hence, it is relevant to consider discrepancies among those measures and identify preferred measures with respect to predicting population dynamic and evolutionary processes. The present study contributes to this aim by (i) making readers aware of differences among commonly used fitness estimates, (ii) showing that simulated (yeast) and calculated (E. coli) competitive fitness may differ across time scales, and (iii) showing that bulk competitions may yield relative fitness estimates that are systematically higher than pairwise competitions. The study is rather thorough on the theory side, with extensive derivations and analyses of various fitness measures using their resource competition model in the Supplementary Information. The study ends with a few practical recommendations for preferred methods to infer relative fitness estimates, that may be useful for experimentalists and stimulate further investigations.

      Comments on revisions:

      I appreciate the thorough and effective response to all recommendations and have no further comments.

    5. Author response:

      The following is the authors’ response to the previous reviews

      We thank both editors and the three reviewers for their positive feedback on the revised manuscript. Following suggestions from Reviewer #1, we cite additional literature to clarify the use of ’fitness potential’ and we have revised the caption of Figure 3 to better explain how we generate the panel of mutants. We have also added a sentence to emphasize that the selection coefficient should match the time-scale of bottleneck effects in the evolution environment.

    1. eLife Assessment

      The authors ask whether a simple whole-head spectral power analysis of human magnetoencephalography data recorded at rest in a large cohort of adults shows robust effects of age, and their results provide compelling evidence that it does. The relative simplicity of the analysis is a major strength of the paper, and the authors are careful to control for many different confounds - although perhaps highly correlated factors like brain anatomy still pose a slight issue. The paper provides a valuable power analysis framework that should inform researchers across the broader neuroimaging community.

    2. Reviewer #1 (Public review):

      Summary:

      This is a careful, well-powered treatment of age effects in resting-state MEG. Rather than extracting (say) complex connectivity measures, the authors look at the 'simplest possible thing' : changes in the overall power spectrum across age.

      Strengths:

      They find significant age-related changes at different frequency bands: broadly: attenuation at low-frequency (alpha) and increased beta. These patterns are identified in a large dataset (CamCAN) and then verified in other public data.

      Weakness:

      Some secondary interpretations (what is "unique" to age vs global anatomy) maybe go beyond what the statistics strictly warrant in the current form, but these can be tightened with (I think pretty quick) additions already foreshadowed by the authors' own analyses.

      Aims:

      The authors set out to replace piecemeal, band-by-band ageing claims with t-maps, and Cohen's f2 over sensors×frequency ("GLM-Spectrum").

      On CamCAN, six spatio-spectral peaks survive relatively strict statistical controls. The larger effects are in low-frequency and upper-alpha/beta ranges (f2 approx. 0.2-0.3), while lower-alpha and gamma reach significance but with small practical impact (f2 < 0.075). A nice finding is that the same qualitative profile appears in three additional independent datasets.

      Two analyses are especially interesting. First, the authors show a difference between absolute and relative spectral magnitude (basically within-subject normalization). Relative scaling sharpens spectral specificity of the spatial maps while absolute magnitude is dominated by a broad spatial mode that correlates positively across frequencies, likely reflecting head-position/field-spread factors. The replication of the main age profile is robust to preprocessing decisions (e.g. SSS movement compensation choices) - the bigger determinant of the effect is whether they apply sensor normalization (relative vs absolute).

      Second, lots of brain-related things might be related to age and the authors spend some time trying to back out confounds / covariates. This section is handled transparently (in general I found the writing style very clear throughout) - they examine single covariates (sex, BP, GGMV, etc.) and compare simple vs partial age effects. For example, aging is correlated with reductions in global grey-matter volume (GGMV) but it would be nice to find a measure that is independent of this : Controlling for GGMV (via a linear model) reduces age-related effect sizes heterogeneously across space/frequency but does not eliminate them, a nuance the authors treat carefully.

      This is a nice paper and I have only a few concrete suggestions:

      (1) High-gamma<br /> There can be a lot of EMG / eye movement contamination (I know these were RS eyes closed data but still...) above 30-40 Hz and these effects are the weakest anyway. Could you add an analysis (e.g. ICA/label-based muscle component removal) and show the gamma band's sensitivity to that step. Or just note this point more clearly?

      (2) GGMV confound control<br /> Controlling for GGMV reduces, but does not eliminate, age effects. I have a few questions about this: a) Could we see the residuals as a function of age? I wonder if there are non-linear effects or something else that the regression is not accounting for. Also, b) GGMV and age are highly colinear - is this an issue? Can regression really split them apart robustly? I think by some cunning orthogonalisation you can compute the effect of age independent of GGVM. I don't think this is the same as the effect 'adjusted' for GGMV (which is what is shown here if I'm reading it correctly). Finally, of course, GGMV might actually be the thing you want to look at (because it might more accurately reflect clinical issues) - so strong correlations are not really a problem: I think really the focus might even be on using MEG to predict GGMV and controlling for age.

      Minor presentation edits:

      It would be handy to see a single table listing each tested "analysis family" (e.g., sensors×frequency, source parcels×frequency), the multiple control used, and the permutation count. I kept wanting to see this as I was reading to compare back and fore.

      I loved the power-planning content (section 3.2, the table with peak f2, CIs, contour plot). I think you could somehow make this even more explicit because people will use it a lot - both for this age/MEG domain and more generally as a template for other types of power planning in the field. Perhaps a "How to use this paper to plan N" guide in a paragraph? Power analysis is surely both "important and difficult" - but also not impossible. A flowchart?

      Comments on the latest version:

      The authors address all my initial points in their revisions and I have no further comments.

    3. Reviewer #2 (Public review):

      This paper describes application of the "GLM-Spectrum" mass univariate approach to examine the effects of age on M/EEG power spectra. Its strengths include promotion of the unbiased approach, suitable for future meta/mega-analyses, and the provision of effect sizes for powering future studies. These are useful contributions to the literature. What is perhaps lacking is discussion of limitations of this approach, in comparison to other methods.

      An analogy is the mass univariate approach to spatial localisation of effects in fMRI/PET images. This approach is unbiased by prior assumptions about the organisation of the brain, but potentially also less sensitive, by ignoring that prior knowledge. For example, a voxelwise univariate approach is less sensitive to detecting effects in functionally homogeneous brain regions, where SNR can be increased by averaging over voxels. In the context of power spectra, the authors' approach deliberately ignores knowledge about the dominant frequency bands / oscillations in human power spectra. This is in contrast to approaches like FOOOF and IRASA, that explicitly parametrise frequency components. I am not saying these methods are better; I just think that the authors should acknowledge that these approaches have advantages over their mass univariate approach (in sensitivity and interpretation; see below). I guess it is a type of bias-sensitivity trade-off: the authors want to avoid bias, but they should acknowledge the corresponding loss of sensitivity, as well as loss of interpretation compared to model-based approaches (i.e., models that parameterise frequency; I don't mean the statistical models for each frequency separately).

      An example of the interpretational loss can be seen in the authors' observation of opposite-signed effects of age around the alpha peak. While the authors acknowledge that this pattern can arise from a reduction in alpha frequency with age, this is an indirect inference, and a direct (and likely much more sensitive) approach would be to parametrise and estimate the peak alpha frequency directly for each participant, as done with FOOOF for example (possibly with group priors, as in Medrano et al, 2025, EJN). The authors emphasise the nonlinear effects of age in Fig 2A, but their approach cannot test this directly (e.g. in terms of plotting effects of age on frequency, magnitude, width for each participant), so for me, this figure illustrates a weakness of their approach, not a strength.

      Then I think the section "Two dissociable and opposite effects in the alpha range" in the Discussion section is confusing, because if there is a single reduction in alpha peak frequency and magnitude with age, then there is only one "effect", not "two dissociable" ones. If the authors do want to claim that there are two dissociable age effects within the alpha range, then they need to do a statistical test, e.g., that the topographies of low and high alpha are significantly different. This then reveals another limitation of the mass univariate approach - that space (channel) is not parametrised either - so one cannot test for significant channel x effect interactions within this framework, as necessary to really claim a dissociation (e.g., in underlying neural generators).

      While the authors show that normalisation of each person's power spectra by the sum across frequencies helps improve some statistics, they might want to say more about disadvantages of this approach, e.g., loss of sensitivity to any effects (e.g. of age) that are broadly distributed across majority of frequencies, loss of real SI units (absolute effect sizes) (as well as problems if normalisation were used for techniques like FOOOF, where the 1/f exponent would be affected).

      Please give more information how artifactual ICs were defined. This may be important for cardiac artefacts, since Schmidt et al (2004, eLife) have pointed out how "standard" ICA thresholds can fail to remove all cardiac effects. This is very important for effects of age, given that age affects cardiac dynamics (even though the focus of Schmidt et al is the 1/f exponent, could residual cardiac effects cause artifactual age effects in current results, even above ~1Hz?).

      Please could the authors clarify the precise maxfilter arguments, and explain what "reference" was used for the "trans" option - e.g., did the authors consider transforming the data to match a sphere at the centre of the helmet, which might not only remove some of the global power differences due to different head positions, but also be best for generalisation of the effect sizes they report to future studies (assuming the centre of the helmet is the most likely location on average)? And on that matter, did head positions actually differ by age at all?

      Comments on the latest version:

      I am happy with their revised version.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors ask whether a simple whole-head spectral power analysis of human magnetoencephalography data recorded at rest in a large cohort of adults shows robust effects of age, and their results provide compelling evidence that it does. The relative simplicity of the analysis is a major strength of the paper, and the authors are careful to control for many different confounds - although perhaps highly correlated factors like brain anatomy still pose a slight issue. The paper provides a valuable power analysis framework that should inform researchers across the broader neuroimaging community

      Many thanks to the reviewers and editorial team. This is an insightful and engaging set of reviews with a range of productive suggestions. We’re pleased that the strengths of this approach show through and that this can be a positive contribution to the community.

      We have implemented the large majority of suggestions and believe that the paper is greatly improved with them in place.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is a careful, well-powered treatment of age effects in resting-state MEG. Rather than extracting (say) complex connectivity measures, the authors look at the 'simplest possible thing': changes in the overall power spectrum across age.

      Strengths:

      They find significant age-related changes at different frequency bands: broadly, attenuation at low-frequency (alpha) and increased beta. These patterns are identified in a large dataset (CamCAN) and then verified in other public data.

      Weaknesses:

      Some secondary interpretations (what is "unique" to age vs global anatomy) may go beyond what the statistics strictly warrant in the current form, but these can be tightened with (I think, fairly quick) additions already foreshadowed by the authors' own analyses.

      Aims:

      The authors set out to replace piecemeal, band-by-band ageing claims with t-maps, and Cohen's f2 over sensors×frequency ("GLM-Spectrum").

      On CamCAN, six spatio-spectral peaks survive relatively strict statistical controls. The larger effects are in low-frequency and upper-alpha/beta ranges (f2 approx 0.2-0.3), while lower-alpha and gamma reach significance but with small practical impact (f2 < 0.075). A nice finding is that the same qualitative profile appears in three additional independent datasets.

      Two analyses are especially interesting. First, the authors show a difference between absolute and relative spectral magnitude (basically, within-subject normalization). Relative scaling sharpens the spectral specificity of the spatial maps, while absolute magnitude is dominated by a broad spatial mode that correlates positively across frequencies, likely reflecting head-position/field-spread factors. The replication of the main age profile is robust to preprocessing decisions (e.g., SSS movement compensation choices) - the bigger determinant of the effect is whether they apply sensor normalization (relative vs absolute).

      Second, lots of brain-related things might be related to age, and the authors spend some time trying to back out confounds/covariates. This section is handled transparently (in general, I found the writing style very clear throughout) - they examine single covariates (sex, BP, GGMV, etc.) and compare simple vs partial age effects. For example, aging is correlated with reductions in global grey-matter volume (GGMV), but it would be nice to find a measure that is independent of this: controlling for GGMV (via a linear model) reduces age-related effect sizes heterogeneously across space/frequency but does not eliminate them, a nuance the authors treat carefully.

      Thank you for this concise summary of the work. We’re glad that the strengths of the approach come through clearly.

      This is a nice paper, and I have only a few concrete suggestions:

      (1) High-gamma 

      There can be a lot of EMG / eye movement contamination (I know these were RS eyes closed data, but still..) above 30-40 Hz, and these effects are the weakest anyway. Could you add an analysis (e.g., ICA/label-based muscle component removal) and show the gamma band's sensitivity to that step? Or just note this point more clearly?

      Thanks for this suggestion. We agree that there is concern about eye movements for these gamma band analyses. The ICA preprocessing we conducted was relatively thorough, and we were able to remove EOG-related components from the majority of datasets, even though the experimental protocol involved eyes closed at rest. It is possible that some components were missed, but they would require more advanced labelling tools or manual intervention to identify.

      There are alternative automated ICA labelling tools that could be used, but these are either not suitable for MEG (ICLabel; https://labeling.ucsd.edu/tutorial, https://mne.tools/mne-icalabel/dev/api/iclabel.html) or optimised for data from CTF/4D systems (MEGNet; https://doi.org/10.1016/j.neuroimage.2021.118402). It is beyond our capacity to modify one of these tools for CamCAN for the current analyses.

      It was much more straightforward to rerun the analysis without the ICA step to see the impact of reintroducing all ocular artefacts removed from the v1 analysis. These results are shown in Supplemental section XXX and replicated below.

      We have added the following Figure 11 and text to the paper.

      Main Text

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged and low-gamma effects are reduced (see supplemental section A.3). “

      Supplemental materials

      “The age results may be contaminated in some way by residual cardiac or ocular artefacts that are not removed during preprocessing. Though ICA denoising was applied, it is possible that some artefactual components were not identified and removed from the dataset. The results at low frequencies and in the gamma range are most likely to be directly impacted by this contamination.

      To explore the impact this has on our analysis, we reran the core GLM effect of age on the data with no ICA artefact rejection at all, allowing all eye movements and heart rate components to remain in the data (Figure 10).

      This no-ICA analysis has three differences to the original in the publication. The low-frequency decrease with age is stronger and more widespread in the analysis that removes ocular artefacts with ICA. A large negative effect is visible in both analyses, though without ICA several frontal and temporal sensors no longer show significant effects. Similarly, the effect size of the low-frequency effect of age is substantially larger when ICA denoising is applied.

      The low-alpha effect was strongly reduced in the analyses that do not remove artefacts with ICA. A large central-occipital group of sensors shows an effect between 7 and 8.5 Hz in the ICA analysis, but this is reduced to a single sensor at 8 Hz when ICA is not computed.

      At high frequencies, the age effect in frontal sensors is larger with ICA, and the age effect in posterior sensors is larger without ICA, though the position and frequencies of significant effects are largely unchanged. The remaining effects in the high alpha and beta ranges are unchanged by application of ICA.

      Overall, ICA either improves the estimation of age effects (low-frequency, low-alpha, high-gamma) or has negligible effects (high-alpha, beta). Only the posterior low-gamma effect is reduced by ICA. Together, we take this as evidence that our core results are robust to interference by eye movements and that the ICA denoising is working effectively to reduce noise in the analysis.”

      (2) GGMV confound control 

      Controlling for GGMV reduces, but does not eliminate, age effects. I have a few questions about this: a) Could we see the residuals as a function of age? I wonder if there are non-linear effects or something else that the regression is not accounting for. Also, b) GGMV and age are highly colinear - is this an issue? Can regression really split them apart robustly? I think by some cunning orthogonalisation, you can compute the effect of age independent of GGVM. I don't think this is the same as the effect 'adjusted' for GGMV (which is what is shown here if I'm reading it correctly). Finally, of course, GGMV might actually be the thing you want to look at (because it might more accurately reflect clinical issues) - so strong correlations are not really a problem: I think really the focus might even be on using MEG to predict GGMV and controlling for age.

      This is an interesting area with some tricky interpretation. Thanks for the nudge to help us clarify further.

      We have added the following text and Figures 17 & 18 to the paper to clarify these points.

      Main Text Section 2.8

      “It is important to note that the correlation between age and GGMV does impact the interpretation of the GLM results, but does not prevent the model fit. We explore the model validation and diagnostics in detail in Supplemental Section A.6. In brief, the model is able to separate the unique contributions of age and GGMV. However, the correlation between factors leads to an inflation in the standard error of the estimates. A hypothesis test on these estimates is valid. However, the inflated variance reduces our ability to detect statistically significant partial effects.”

      Supplemental Section A.6

      “There is a strong correlation between age and Global Grey Matter Volume (Pearson’s r=-0.75). The shared variance arising from this collinearity adds nuance to the interpretation of the results, which we explore in more detail in this section.

      Firstly, the sum-square residuals for the group-level model fit including both Age and GGMV are shown as a function of frequency in Figure 17. We see that the residuals broadly follow the overall distribution of variance in the data, peaking at low frequencies and in the alpha range. These frequency ranges are where the strongest signal is visible, but also the highest variability between participants. We would expect that the group model would not perform so well in the points of greatest variability. Importantly, though the residuals are relatively high in the alpha, this is still in the context of a very well-performing model with R2 values of around 80%. “

      “Secondly, the correlation between age and GGMV is not inherently problematic for the GLM, but it does add complexity and nuance to the interpretation of the results. Some additional model validation statistics are shown in Figure 18. “

      “The singular value spectrum of the design matrix indicates whether a design is low-rank, the smallest singular value in this case in 0.37 which indicates that there isn’t a rank deficiency which would prevent us from estimating the model. Though we can estimate the model, correlated regressors can reduce its efficiency. The variance inflation factors for the joint AGE-GGMV model are above 1 for both parametric regressors, indicating that the standard errors of their estimates are inflated. The VIF of 2.35 indicates that the standard errors of this joint model are around sqrt(2.35) = 1.533 times greater than they would be in a separate or uncorrelated model. Though there is no hard rule for this, the literature generally suggests that a VIF above 5 (or sometimes 10) indicates severe multicollinearity.”

      Including additional regressors in the model can change the age estimate by ‘partialling’ out the variance that can be attributed to the other variables and by inflating standard errors. In this specific case of GGMV, the partialled estimates are reduced heterogeneously across space and frequency, and the amount of inflation is at a tolerable level. Overall, the regression is able to separate the unique effects of age and GGMV, at the cost of this inflation in the associated standard errors.”

      Reviewer #2 (Public review):

      This paper describes the application of the "GLM-Spectrum" mass univariate approach to examine the effects of age on M/EEG power spectra. Its strengths include promotion of the unbiased approach, suitable for future meta/mega-analyses, and the provision of effect sizes for powering future studies. These are useful contributions to the literature. What is perhaps lacking is a discussion of the limitations of this approach, in comparison to other methods.

      Thank you for this summary and the thoughtful review. The emphasis for this paper is exactly on the points you highlighted, and we’re glad that this has come across well. We agree that a broader comparison to other methods would be a useful addition and have included a series of additional discussion points to address this.

      We will take the opportunity to reclarify that our intention for this method is not to replace other, more complex or targeted approaches, but to establish a more generalisable foundation for their development. We’re fully supportive of other approaches and are working on their application ourselves. On reflection, this was not clear enough in the first submission, and we have added the following text to the introduction to clarify.

      “Reporting of whole-head and full-frequency spectra of effect estimates would make it straightforward to aggregate across studies and eventually enable identification of sub-threshold effects that may be missed in single analyses but are consistent across studies. We argue that this approach provides a generalisable foundation that can support more complex analyses with a frequency component (such as aperiodic slopes, burst detection, and dynamic functional networks) that require more researcher degrees of freedom.”

      An analogy is the mass univariate approach to spatial localisation of effects in fMRI/PET images. This approach is unbiased by prior assumptions about the organisation of the brain, but potentially also less sensitive, by ignoring that prior knowledge. For example, a voxelwise univariate approach is less sensitive to detecting effects in functionally homogeneous brain regions, where SNR can be increased by averaging over voxels.

      In the context of power spectra, the authors' approach deliberately ignores knowledge about the dominant frequency bands/oscillations in human power spectra. This is in contrast to approaches like FOOOF and IRASA, which explicitly parametrise frequency components. I am not saying these methods are better; I just think that the authors should acknowledge that these approaches have advantages over their mass univariate approach (in sensitivity and interpretation; see below). I guess it is a type of bias-sensitivity trade-off: the authors want to avoid bias, but they should acknowledge the corresponding loss of sensitivity, as well as loss of interpretation compared to model-based approaches (i.e, models that parameterise frequency; I don't mean the statistical models for each frequency separately).

      This is an important point, and we are in complete agreement about the importance of giving a balanced description of how this approach fits within the broader literature. We have added the following paragraph to the discussion on limitations to lay this out more clearly.

      “Our approach promotes an exploratory and unbiased approach to quantifying the age effect on neuronal power spectra, which is intended to complement more focused analyses. This has the benefit of reducing researchers’ degrees of freedom and of being broadly generalisable. These come at the cost of a loss in specificity and in sensitivity. Our approach does not specifically quantify features derived from the power spectrum, such as alpha-peak frequency or the aperiodic component of the spectrum. These features are mixed into our full-spectrum estimates but not directly quantified. Thus, they can be challenging to interpret from our approach. Secondly, the mass-univariate approach suffers from a potential loss in sensitivity compared to results that aggregate across spatial or spectral regions that contain consistent results. Where a region or frequency band of interest can be supported from the literature, an approach focusing on a single region has the benefit of reduced noise by averaging estimates from a larger range of observations. Finally, models that consider the whole shape of the spectrum [Donoghue et al. 2021] would also be able to combine information across a range of frequencies rather than depending on a single frequency bin for each estimate. These models have the additional benefit that their parameters are often directly interpretable as features of interest, such as spectral slope or peak frequency.”

      An example of the interpretational loss can be seen in the authors' observation of opposite-signed effects of age around the alpha peak. While the authors acknowledge that this pattern can arise from a reduction in alpha frequency with age, this is an indirect inference, and a direct (and likely much more sensitive) approach would be to parametrise and estimate the peak alpha frequency directly for each participant, as done with FOOOF for example (possibly with group priors, as in Medrano et al, 2025, EJN). The authors emphasise the nonlinear effects of age in Figure 2A, but their approach cannot test this directly (e.g., in terms of plotting effects of age on frequency, magnitude, and width for each participant), so for me, this figure illustrates a weakness of their approach, not a strength.

      We agree that this point might be misleading within its own figure and have moved the result to the supplemental material with a more lightly phrased wording in the main text. Figure 3 on effect sizes has moved to Figure 2, and a new Figure 3 shows the quadratic effect of age as suggested later in the review.

      “We have completed several control analyses to support these findings. Firstly, we have explored the correspondence between alpha peak frequency and the two effects we identified within the canonical alpha range (see supplemental section A.1). Secondly, the overall pattern of findings is consistent in an equivalent source space analysis using LCMV beamforming and parcellation (see supplemental section A.2). Finally, we have repeated the analyses with and without ICA denoising and find that the overall spectral profile is very similar. Effects in low-frequencies, low-alpha and high-gamma are increased with application of ICA, whilst high-alpha and beta remain unchanged, and low-gamma effects are reduced (see supplemental section A.3).“

      This supplemental section contains some additional content relevant to the next comment.

      Then I think the section "Two dissociable and opposite effects in the alpha range" in the Discussion section is confusing, because if there is a single reduction in alpha peak frequency and magnitude with age, then there is only one "effect", not "two dissociable" ones. If the authors do want to claim that there are two dissociable age effects within the alpha range, then they need to do a statistical test, e.g., that the topographies of low and high alpha are significantly different. This then reveals another limitation of the mass univariate approach - that space (channel) is not parametrised either - so one cannot test for significant channel x effect interactions within this framework, as necessary to really claim a dissociation (e.g., in underlying neural generators).

      As above, we agree that this can be misleading and that our intention got muddled in the heading and writing of this subsection. We are not intending to argue that there are definitely two distinct and different effects. We clarify this in the main text of our first version, but not clearly enough:

      “The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak”.

      We choose to present the first paragraph of this section as a discussion of two effects because this is what is already reported in the literature on changes in alpha power with age. Whilst the decrease in alpha frequency with age is well replicated, the decrease in alpha power is less consistently reported, and the literature shows a highly variable picture of the spatial pattern of this effect. We do not claim a strong dissociation based on these results, but both are reported within the literature.

      Our intention is to provide some clarity to this literature by taking a step back and looking at the ‘lay of the land’ in an unbiased way. With this approach, we see different effects of age on magnitude within a canonical alpha range that are completely separated in frequency. With this perspective, it is feasible that different publications taking different regions of interest, frequency band definitions, and processing options could lead to mixed reports of increases and/or decreases in alpha power with age.

      When single papers that take focused but inconsistent approaches report inconsistent results, we would argue that our approach leads to a substantial interpretational gain on the level of collections of papers in the literature.

      We have added the following content to supplemental section A.1 to clarify and Table 3 illustrates the issue. We have retained the paragraph discussing the possibility that these two effects combine to represent a shift in frequency of a single peak.

      Main text section 3.1, replacing the section on ‘Two dissociable effects…’

      “Reconciling conflicting reports of the ageing effect on alpha power.

      The literature exploring how ageing changes alpha power is heterogeneous. Papers that report results in resting alpha power, either from a canonical band or from an individual peak frequency, include reports of a variety of contrasting age effects. This includes positive correlations with age [Rempe et al., 2023, Stier et al., 2023], negative correlations [Thuwal et al., 2021, Lodder and van Putten, 2011, Medrano et al., 2025, Park et al., 2024], both positive and negative effects separated by space [Hoshi and Shigihara, 2020, Pathak et al., 2022], or null results when correcting for individual frequencies and aperiodic slopes [Scally et al., 2018, Merkin et al., 2023] (see supplemental section A.1 more detailed summary). There is broad variability in methodological approaches, which could account for the variety of results. Stier et al. [2023] suggest that analyses in sensor or source space may lead to different effects. Critically for our work, this variability also prevents formal aggregation of results and meta-analyses that could clarify the picture.

      We have proposed that, by taking a step back and tolerating a reduction in sensitivity, we can map out the whole-head whole-frequency structure of the age effect with minimal researcher degrees of freedom and bias. Investigating age effects as a complete spectrum shows that two contrasting age effects on alpha magnitude coexist in close proximity in space and frequency: an increase with a small effect size in central occipital sensors around 7-8.5 Hz and a decrease with a large effect size across a broad set of occipital, temporal and frontal sensors between 9.5-12.5 Hz. Different data samples and different data analysis choices, particularly the selection of regions of interest, source reconstruction, sensor normalisation, or correction of aperiodic components, might emphasise one effect or the other in each analysis. Whilst we have not conclusively explored all possible variants of these analyses, we have provided a framework that would allow future studies to perform formal comparisons and meta-analyses to resolve this bottleneck.”

      “Relationship between effects on alpha power and alpha individual frequency.

      The two effects identified in the present analysis may combine to represent a decrease in power and frequency of a single alpha peak (Seen qualitatively in Figure 1A). This change in alpha peak frequency is highly replicable [Cesnaite et al., 2023, Dustman et al., 1993, Sahoo et al., 2020, Scally et al., 2018, Pathak et al., 2022, Zibrandtsen and Kjaer, 2021] and is a highly predictive spectral marker of ageing [Stier et al., 2024]. Decreases in alpha peak frequency have been linked to a decline in cognitive performance in healthy ageing [Cesnaite et al., 2023, Finley et al., 2024] and MCI [Garc´es et al., 2013, L´opez-Sanz et al., 2016, Puttaert et al., 2021].

      This compelling possibility that the age effect on alpha is a shift in a single peak is complicated by strong evidence for presence of multiple alpha peaks within individuals [Lodder and van Putten, 2011, Chiang et al., 2011, 2008, Klimesch, 1999], with distinct generators and functional relevance [Sokoliuk et al., 2019]. A complex pattern of changes in power, frequency, and spatial distribution likely underlies age-related change in alpha oscillations. Future work will need to explore all three features at the individual level to clearly illuminate the change.”

      Supplemental section A.1

      “Part of our motivation for this method is that variability in the methodological choices in different publications makes it difficult to aggregate varying results across the literature. For example, though a decrease in alpha peak frequency with increasing age is reported highly consistently, the effect of age on alpha power is much more variable. Table 3 shows a representative sample of publications over the last 20 years that report a change in alpha power with age (note that this is intended to be a representative rather than an exhaustive list). Over half of publications (9/15) report a decrease in alpha power, whilst the remaining publications report an increase (2/15), both increases and decreases (1/15), a decrease but only without correcting for aperiodic slope (1/15), a quadratic effect (1/15), and no effect (1/15). Stier et al. [2023] suggest that the choice of analysis space (sensor space or source reconstruction) is likely a source of discrepancies between studies.

      Critically, it is difficult to reconcile these findings with the information reported in the publications. For example, even within the 9 publications that report a decrease, there is little correspondence in the spatial location of the effect. We argue that focused approaches cannot resolve this issue alone, as targeted analyses are more specific to each dataset and less generalisable.

      In the specific case of the mixed literature on change in alpha power with age, our results show that both effects are present and separated in frequency. It is feasible that the different methodological choices and datasets used by each study in our survey means that one or other of these two effects were emphasised. As a result, the literature may not be mixed in scientific terms, but that a rich pattern of results is obscured by methodological variability.”

      While the authors show that normalisation of each person's power spectra by the sum across frequencies helps improve some statistics, they might want to say more about disadvantages of this approach, e.g., loss of sensitivity to any effects (eg of age) that are broadly distributed across majority of frequencies, loss of real SI units (absolute effect sizes) (as well as problems if normalisation were used for techniques like FOOOF, where the 1/f exponent would be affected).

      This is an important point, and we have added the following text to clarify, with one exception. Firstly, the normalisation we applied scales the whole spectrum linearly and would change the 1/f intercept but not the 1/f^alpha exponent.

      Main text section 3.3

      “Both absolute and relative power measures are used throughout the literature, but there is little consensus about their interpretation [Sandre and Troller-Renfree, 2026]. We focus on relative power for most of our results. By normalising each participant’s power spectrum by the sum across frequencies, we found that the results gained specificity in frequency band and reduced concern about wide intersubject differences in overall variance. Though this improved some analyses, relative power has important drawbacks. In particular, it can reduce sensitivity to effects that are broadly distributed across the spectrum and uses arbitrary scaling rather than meaningful physical units. We support calls in the literature to report both relative and absolute power measures [Rempe et al., 2023, Sandre and Troller-Renfree, 2026].”

      The authors should give more information on how artifactual ICs were defined. This may be important for cardiac artefacts, since Schmidt et al (2004, eLife) have pointed out how "standard" ICA thresholds can fail to remove all cardiac effects. This is very important for the effects of age, given that age affects cardiac dynamics (even though the focus of Schmidt et al is the 1/f exponent, could residual cardiac effects cause artifactual age effects in current results, even above ~1Hz?).

      Artefactual components were estimated using standard tools in MNE python. Specifically:

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_ecg

      https://mne.tools/stable/generated/mne.preprocessing.ICA.html#mne.preprocessing.ICA.find_bads_eog

      We have clarified the text in the methods to make the overall process clearer.

      We believe that this process has been broadly effective, and we have rejected an average of 2.25 ECG components within each dataset. There remains a strong possibility that residual ECG artefact is present in the data.

      We have added the following text to methods section 4.2

      “Artefactual components relating to eye movements or the heart rate were automatically identified by correlation with the simultaneous EOG and ECG channels. ECG artefacts were identified using cross-trial phase statistics [Dammers et al., 2008] and an automatic threshold based on the sample rate of the data, as implemented in the mne.preprocessing.ICA.find_bads_ecg function in MNE Python. Between 0 and 3 EOG components were rejected in each dataset, with an average of 0.99 (standard deviation: 0.79) across all datasets. EOG artefacts were identified by correlation with the HEOG and VEOG channels, with a threshold set to r = 0.35, as implemented in the mne.preprocessing.ICA.find_bads_eog function in MNE Python. Between 0 and 5 ECG components were rejected in each dataset, with an average of 2.25 (standard deviation: 0.84) across all datasets. The continuous sensor data were then reconstructed without the influence of the components labelled as artefacts.”

      Similar to the response to Reviewer 1, we have not been able to rerun a more advanced ICA algorithm on the data but have repeated the analysis without any ICA to see if including all ocular and cardiac artefacts influences the results. This change does not introduce any new signal components to the results but does attenuate the low-frequency and low-alpha effects. With the assumption that our initial ICA analysis captured the majority of the largest ECG components, we are confident that our core findings are not compromised by cardiac artefacts.

      Please see the response to comments to Reviewer 1 for additional text in the manuscript on this point.

      The authors should clarify the precise maxfilter arguments, and explain what "reference" was used for the "trans" option - e.g., did the authors consider transforming the data to match a sphere at the centre of the helmet, which might not only remove some of the global power differences due to different head positions, but also be best for generalisation of the effect sizes they report to future studies (assuming the centre of the helmet is the most likely location on average)? And on that matter, did head positions actually differ by age at all?

      We have used the maxfilter files as provided by the CamCAN team and an equivalent implementation defined in OSL-ephys for the Oxford and Cambridge MEGUK data. We have clarified the text in section 4.2

      “All MEG data pre-processing was carried out using MNE-Python [Gramfort, 2013] and OSL-ephys [Quinn et al., 2022, van Es et al., 2025] using the OSL batch pre-processing tools. For the CamCAN data, we proceeded with analysis on post-maxfilter processed data provided by the CamCAN team. Briefly, the data were processed using AA [Cusack et al. 2015] with automatic bad channel detection (limited to 7 channels) with the origin set to the centre of a sphere fitted to the individual’s Polhemus head shape points. Maxfilter signal-space separation was performed with the temporal extension enabled (temporal window of 10 second and correlation threshold of r = 0.98). Head position was continuously estimated and compensated for during periods where the HPI coils were on. After maxfilter processing, head positions were translated into a default head position defined as a point relative to each individual’s origin in a head coordinate frame. Data with the full maxfilter processing and with the head position translation or movement compensation were extracted from the CamCAN database. An equivalent pipeline was implemented in OSL-Ephys and applied to the data from the MEG-UK datasets from Oxford and Cambridge. Files from the MEG-UK Nottingham dataset were processed with third-order gradiometry applied.”

      We found that head position in CamCAN does change with age. We have included the following in Supplemental section A.5

      “The head position of participants within the MEG sensor dewar is an important consideration that has the potential to change the signal-to-noise level of each individual data recording. There are significant differences in head position as a function of age in the CamCAN dataset in Y (front-back) direction indicating that older participants are seated further forward in the dewar than younger participants.”

      As a point of reference, this pattern is consistent with the largest change with age estimated using the un-normalised raw power spectra in Figure 5. A 1-95Hz region shows a change with age that is consistent with older participants sitting further forward in the dewar. This effect is absent in the relative power contrasts. 

      We have not fully explored this final point so have not included it in the main paper, but believe it is a useful addition to the discussion of the reviews.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, both reviewers are enthusiastic about your paper, indicating that it provides compelling empirical support for its key claims and that it represents a valuable theoretical advance for the field. They provide a number of comments that you may want to consider prior to finalizing the manuscript for publication. All the best, Redmond O'Connell

      Reviewer #1 (Recommendations for the authors):

      (1) It would be handy to see a single table listing each tested "analysis family" (e.g., sensors×frequency, source parcels×frequency), the multiple controls used, and the permutation count. I kept wanting to see this as I was reading to compare back and forth.

      This is a helpful suggestion, we have included the table in a new supplemental section which is referenced from the main text at the start of the results

      Main text section

      “A summary of all GLM analyses carried out in this work is available in supplemental section A.7.”

      With the following table included in supplemental section A.7

      (2) I loved the power-planning content (section 3.2, the table with peak f2, CIs, contour plot). I think you could somehow make this even more explicit because people will use it a lot - both for this age/MEG domain and more generally as a template for other types of power planning in the field. Perhaps a "How to use this paper to plan N" guide in a paragraph? Power analysis is surely both "important and difficult" - but also not impossible. A flowchart?

      We’re very glad that this section is working well and agree that the practical planning content should have been more constructive! We have added the following text as a guide

      Main text section 3.2

      “We propose the following steps as a practical guide for researchers looking to plan a data sample with a reasonable chance of correctly identifying a particular effect.

      (1) Define research question and identify previous results: Your research question must be defined in advance and well specified. There should be relevant data or literature that can be used to guide your decision.

      (2) Define the smallest effect size of interest and the decision criterion for the sample decision: It is critical to define the parameters of how you will make your sample size decision ahead of time. We recommend considering what the ’smallest effect size of interest’ [Anvari and Lakens, 2021] would be for your question. This is specific to your question and is about more than statistical significance. What effect size would indicate that there is a practically meaningful effect for the future literature to consider?

      (3) Estimate effect sizes to inform your decision: Either by aggregating reported statistics from the literature, or by dedicated processing of previous data, compute an estimate of the effect size. Effect sizes are only estimates, so it is important to compute confidence intervals around your estimate to get a measure of variability.

      (4) Compute power/precision curves assuming these results: Using the estimated effect sizes, compute power curves [Baker et al., 2021] that visualise the relationship between effect size, statistical power, and sample size.

      (6) Select the sample size that meets your pre-defined criteria: Using your definitions from step 2, and when considering the whole power curve, make a decision about what sample size would give you a reasonable chance of replicating the effect of interest.

      (7) Make note of any differences or deviations from this plan during your data collection: There are many practical reasons why your planned sample might not match the data acquired in practice. Such deviations from a plan are ok but should be acknowledged, and any mitigating steps explained [Lakens, 2024].

      (8) Document your process for inclusion in a preregistration or publication: Include details on how a sample size decision was made in your research outputs including preregistrations, preprints, and publications. These details will be useful for future researchers to understand your process and to implement their own.”

      Reviewer #2 (Recommendations for the authors):

      (1) Though they mention in the Discussion, the authors could have noted earlier in Section 2.1 that one does not need to assume a linear effect of age - one could use a polynomial expansion or even local splines within the same GLM framework. Indeed, it seems unlikely a priori that effects of age across the wide range of ages in the CamCAN data are linear for all frequencies.

      This is an important point, and one that was straightforward to implement in our model. Based on feedback on other sections, we have removed the previous Figure 2, moved Figure 3 on effect sizes to the second position, and added a new Figure 3 containing a model of the quadratic age effects. We believe that this is a substantial improvement in the paper.

      Main text section 2.3

      “(2.3) Quadratic effects of age

      The linear effect of age is a convenient and simple regression model. However, it makes a strong assumption that change with age is uniform across the whole age range. A second group-level GLM was computed with an additional regressor to quantify quadratic effects of age, which have been reported in the ageing literature [G´omez et al., 2013, Rempe et al., 2023, Stier et al., 2023]. The spectrum of t-values for the quadratic age predictor (Figure 3) shows significant effects for a U-shaped change with increasing age in the low frequency (1-5 Hz) range in central sensors. Significant inverted-U shaped effects are present in two frequency ranges in the beta band. A low-frequency beta effect is present in occipital and temporal sensors between 16 Hz and 20 Hz, whilst a second high-beta effect is present in central sensors between 24 Hz and 30.5 Hz. The low-frequency and low-beta effects overlap in space and frequency with linear effects, suggesting that the low-frequency change with age has both a linear increase with age and a U-shaped component, whilst the low-beta effect has a decrease with age and an inverted-U shaped component. The high-beta effect does not overlap in frequency with any of the reported linear effects. The effect sizes for quadratic effects range between Cohen’s F 2 values of 0.033 for low frequency to 0.076 for high beta and are generally lower than the effect sizes for linear change.”

      Main text section 2.4

      “(2.4) Sample size planning for effects of age on the neuronal power spectrum

      We use the 95% confidence intervals around the effect sizes to make recommendations for future sample sizes for future samples that plan to replicate these results. The observed power calculations have no bearing on the interpretation of the present results. Instead, they should be used as a general guideline for planning future studies. Table 1 gives a full summary of the peak statistics and future sample size range for the six age effects identified in Figure 1B. These results have implications for future sample planning for resting-state electrophysiology studies of ageing. Using the upper bound of the sample size estimates as a conservative estimate, the linear changes with age that have relatively large effect sizes would have well-powered replications with sample sizes of around 50-60 participants. However, the smaller linear effects and all the quadratic effects would require samples of 200 or more participants to have the same probability of detecting the effect if it is indeed present (Table 1). This indicates that study samples should be planned with the smallest effect of interest in mind and that ageing effects in different frequency bands may not all be well powered within the same sample.”

      Discussion section 3.1

      “Linear increase and inverted-U effects in the beta band. The literature consistently reports an increase in low-beta power with age [Gomez et al., 2013, Heinrichs-Graham and Wilson, 2016, Heinrichs-Graham et al., 2018, Hubner et al., 2018, Koyama et al., 1997, Rempe et al., 2023, Stier et al., 2023, Veldhuizen et al., 1993, Xifra-Porxas et al., 2019]. We observed significant inverted-U-shaped quadratic effects of age in two frequency ranges in the beta band, a posterior low-beta component (centred around 18 Hz) and an anterior high-beta component (centred around 25 Hz). This is broadly consistent with reports of quadratic ageing effects in the beta band in the literature [Rempe et al., 2023], though, to our knowledge, our results are the first to suggest a separation of effects into different parts of the beta range. These spectrum changes may be associated with age-related changes in underlying bursting dynamics [Brady et al., 2020, Power et al., 2023].”

      “Methods section 4.7

      Age plus quadratic age models: To move beyond linear change and explore U and inverted-U shaped effects of age, we fitted a group model with three regressors. One constant term, one z-transformed age regressor, and one z-transformed quadratic age regressor (age − mean(age))<sup>2</sup>.”

      (2) The authors could point out an obvious extension of their approach to mixed-effects models, which could properly model within- and between-participant effects (repeated measures), e.g., for longitudinal effects of ageing, GAMMs, etc, including hierarchical linear models that combine trials and participants in the same model, and potentially model trial-specific/stimulus effects, etc.

      This is an important point; we have added the following paragraph to the discussion section.

      Discussion section 3.5

      “Future extensions

      A clear future extension for this work is to formally incorporate estimates of within-subject variability into a mixed-effect model. These powerful models would enable modelling of both fixed effects and random effects, allowing researchers to account for variation within individuals over time and between individuals. LMMs also provide improved approaches for handling missing observations and unbalanced designs, making them especially useful for longitudinal and hierarchical data. A second expansion of this work could use Generalised Additive Mixed Models (GAMMs) to model non-linear relationships using smooth functions while also accounting for random effects. This may allow for more realistic representations of complex patterns of change across age. Linear Mixed Modelling comes with a substantial increase in researcher degrees of freedom and can be challenging to implement and report accurately [Meteyard and Davies, 2020]. We have used fixed-effects modelling in this work in line with our objectives of maintaining generalisability and minimising researcher degrees of freedom.”

      (3) Do the authors want to comment on why effect sizes in Figure 4Ci are so much higher for the Oxford sample? I know the authors' main point is that smaller samples lead to more variable estimates of the true effect size, but there are also other interpretations for these sample differences, e.g., recruitment differences, scanner differences, etc. Have the authors considered implementing empirically Bayesian approaches (like COMBAT, a type of mixed-effects model) to adjust for site differences in both offset and scaling, which I think should be a fairly simple extension of GLM-spectrum?

      We agree that our explanation of here is somewhat lacking. To be transparent, we have thought long and hard about this difference and cannot identify a clear reason why this difference is so striking. There is no apparent difference in the other covariates, such as head position, age, or gender, which could explain why the Oxford dataset has larger effect sizes in this specific frequency range (the results in the low-frequency and alpha ranges are consistent).

      We have added the following to the main text to highlight dedicated data harmonisation strategies that could be applied in this situation.

      Main text section 2.5

      “This may arise from relatively poor estimates of the population level variability from smaller data samples. It is possible that more systematic differences in the participant sampling, recruitment, and data acquisition process could contribute to between-site differences. At present, our analyses can show that the core effects of ageing are replicable across several datasets, though we have not formally combined these datasets into a single analysis. Formal methods for data harmonisation, such as ComBat [Johnson et al., 2006], could correct for additive and multiplicative differences in data scaling across sites to improve site comparisons.

      (4) There are a number of formatting problems - at least in the PDF provided to reviewers - for some symbols, e.g., "below ¡7 Hz" (I presume "j" was originally "~" or something?). Also, sometimes a figure is cited by a single, hyperlinked number, without the word "Figure".

      (5) Line 402: "influence" = "influenced"

      (6) Line 431: date for Gelman & Loken?

      Thank you for highlighting these, we have done a thorough proofread and fixed a large number of spelling and formatting issues.

    1. eLife Assessment

      This valuable study demonstrates how rhythmic inputs can shape sequential working memory. The study puts forward plausible dynamic mechanisms, but some of the central behavioral effects are small, so the evidence - at present - remains incomplete.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates how rhythmically presented stimuli support working memory by using task-trained recurrent neural networks (RNNs) endowed with short-term synaptic plasticity. RNNs trained with rhythmic sequences have a marginal performance increase (0.4%) over models trained with jittered input and show increased phase-locking and oscillatory organisation during the sample period. While the question addressed in this paper is highly relevant, the core conclusion that regular temporal structures provide a functional scaffold for sequence working memory lacks evidence. The extensive post-hoc filtering pipeline obscures whether there is phase coding or not, and whether or not the found oscillatory phenomena are truly emergent or a mathematical artefact of the analytical selection criteria.

      Strengths:

      (1) The manuscript addresses a highly relevant question.

      (2) The introduction is nicely written and presents relevant background work.

      (3) The authors' results are robust in the sense that they analysed and trained an ensemble of models instead of single networks.

      Weaknesses:

      (1) Misalignment between analysis epoch and core claims. The manuscript argues that temporal regularity supports sequence working memory. However, the majority of analyses focus on the sample/encoding period rather than the delay period during which memory maintenance occurs.

      (2) Ambiguity in the neural code (rate vs. phase). The decoding accuracies suggest that the memory can be well decoded from the instantaneous activity, implying a rate (not a phase) code. This raises two questions:<br /> a) Can memory-related information be decoded directly from the oscillatory phase, particularly during the delay period?<br /> b) What would be the mechanism with which the increase in phase organisation improves a representation that seems otherwise decoded/represented from activity levels?

      (3) Absence of any RNN activity plots. The manuscript would benefit from showing, e.g., single neuron response plots, raster plots, phase histograms of units, etc. Are there actually spontaneous oscillatory dynamics as the paper writes (line 243)? Can the authors show baseline activity (which is also supposed to be oscillatory, line 219)?

      (4) Potential concerns in the analysis pipeline: The data undergo an intensive, selective pipeline that might be susceptible to introducing circularity and selection bias. I highlighted some points here:<br /> a) Many analyses are performed on (summed) data projected on demixed PCs (extracted from time-warped data). Crucially, dPCAs are not unsupervised; they already explicitly maximise the variance of interest.<br /> b) For the phase extraction during sample encoding: after dPCA percentile clipping, z-scoring, and z-score clipping are applied (lines 762-764), low-amplitude trials are excluded (lines 779-781), and there is further selection based on a valid-point criterion and r2 thresholding (lines 804-805). Do all of these selection criteria risk introducing bias?<br /> c) Some statistical assumptions are not explicitly evaluated. E.g., the sign-flip permutation test (lines 735-742) relies on sign-exchangeability.<br /> d) For selectivity analysis of oscillatory organisation (Figure 4C, lines 893-896): Units are first selected by ANOVA, and then on the selected units further stats (Power and PLV) are computed. Unless the further stats are completely independent of the ANOVA, this may introduce selection bias.<br /> e) The finding of stronger power around f0 given rhythmic inputs of that exact frequency seems somewhat circular (Figure 3A)?<br /> f) The dPCA description seems a little odd, e.g., line 674, for the ordinal component you would normally actually average (i.e., marginalise out) everything except the ordinal axis.

      (5) Conflation of RNN learning dynamics with working memory mechanisms. The authors show that rhythmic input makes learning marginally easier, but in principle, both RNNs reach full performance (so working memory can be done as well with either case). To avoid the findings depending on learning dynamics, it could be of interest to test the RNNs trained with jittered input on fixed input (or retrain RNNs with both jittered and non-jittered input). It is also unclear if the small increase in performance (0.4%) can be expected to hold across different initialisations and/or learning rate /regularisation strengths.

      (6) STSP. It is unclear if the findings rely on STSP being present or not (or what the role of STSP is in the model at all currently). Note that in Liebe et al. 2025, RNNs were trained on an almost identical task without STSP, and phase-coding was demonstrated in the models.

      (7) Writing redundancy. The methods subsections "Population signal construction for oscillatory analysis" and "Oscillatory phase organization during sample encoding" seem to define exactly the same quantity with different characters (activity projected in dPCA space), which leads to confusion (in one section, z is the PC component, in another, it's the complex signal). There are also slightly different definitions of the wavelets in either section, for which the reasoning is unclear.

    3. Reviewer #2 (Public review):

      Summary:

      The authors train E-I recurrent networks with short-term synaptic plasticity on a sequential delayed match-to-sample task, comparing regular versus jittered sample timing. They report a small accuracy gain under rhythmic input, a more separable population geometry during encoding, organization of internal oscillations around the dominant input frequency, a preference for temporal order over feature encoding, and improved decodability and persistence of stimulus information in both activity and synaptic efficacy. A delay-period perturbation shows synaptic efficacy contributes more than activity to maintenance.

      Strengths:

      The model is well-specified. Dale's law, the STSP formulation, the training objective, and the hyperparameters are all reported clearly enough to reproduce, and code is shared. The statistical machinery is appropriate, with cluster-permutation tests for the spectral analyses and across-network sign-flip tests rather than naive pooling. The temporal-order versus stimulus-direction dissociation in Figure 4C is the most interesting result. The negative association between phase locking and direction selectivity is non-trivial and argues against a simple global-gain reading, and it connects to Liebe et al. 2025. The serial-position decoding curves and the synaptic-versus-neuronal perturbation are well-motivated tests of the maintenance claim.

      Weaknesses:

      The behavioral effect is very small. Match accuracy is 0.991 versus 0.987, and non-match is 0.973 versus 0.969, on networks already at the ceiling. The entire mechanistic analysis is built to explain a roughly 0.4 percentage point difference, and the paper does not establish that this difference is functionally meaningful rather than a marginal byproduct of the timing manipulation. The IOI-dependence result meant to support it is weak, with an R-squared of 0.071 at a p-value of 0.029 on n of 67.

      The core spectral and phase results are close to definitional and should be framed that way. The regularity index R is computed from IOI variability, the dominant frequency f0 is computed from the same IOIs, and the oscillatory metrics in Figures 3 and 4 are then measured relative to f0 and correlated against R. This shows that more regular input produces internal phase progression closer to the input-derived reference frequency, partly restating the input statistics rather than uncovering an independent network mechanism. The phase-locking-increases-with-regularity finding is the clearest case. This does not invalidate the analyses, but the manuscript currently reads them as a mechanism when much of the signal is built into the measurement.

      The only genuinely causal manipulation is the delay-period shuffle, and it is underpowered at n of 15. Its main conclusion, that synaptic efficacy matters more than activity for maintenance, largely recovers prior STSP results (Mongillo et al. 2008, Masse et al. 2019) rather than establishing something specific to rhythm. The result the authors most want, that disrupting synaptic state removes the rhythmic advantage, is predicted in the Discussion but not tested.

      The authors should add a control that breaks the circularity (a held-out f0/phase reference, or shuffling R against the metric) and run the causal STSP-disruption test that is mentioned in the Discussion.

      The oscillatory framing is stronger than the model supports. Phase locking to a periodic input can reflect temporal predictability or repeated preparation without self-sustained entrainment, and the authors acknowledge this once but then use entrainment-style language throughout. The signals are extracted from firing-rate units and should not be read as LFP or EEG oscillations.

      The authors should show raw single-unit and population activity so readers can verify the oscillations before the filtered pipeline. The delay perturbation largely recovers Mongillo 2008 / Masse 2019 rather than anything rhythm-specific, and the relationship to Liebe et al. 2025 should be addressed in the Results.

      Appraisal and impact:

      The authors largely achieve their stated aim of describing how temporal regularity constrains recurrent dynamics in this model, and the temporal-order preference is a useful prediction. The reach of the conclusions exceeds the evidence in two places: the functional importance of the behavioral effect and the degree to which the phase results are independent of the input construction. With the framing corrected and one causal test added, this would be a useful contribution to the modeling literature on timing and working memory rather than a definitive account.

    4. Author response:

      We thank the editors and the two reviewers for their careful evaluation of our study, as well as for their positive assessment of the research question, model reproducibility, and the results concerning temporal-order representations. We also agree with the core issues raised in the reviews: the current manuscript has not yet sufficiently distinguished descriptive changes in recurrent dynamics from the functional mechanisms underlying the behavioral advantage; the behavioral effect itself is small and close to the performance ceiling; the phase analyses require more stringent controls; and the specific role of short-term synaptic plasticity in the rhythmic advantage has not yet been directly tested. We plan to add the corresponding analyses in the revised manuscript and to temper several mechanistic claims.

      First, we would like to clarify three aspects of the study design and analysis pipeline. First, the rhythmic and arrhythmic conditions were not performed by two separately trained groups of networks. Each network was jointly trained using balanced batches containing rhythmic-match, rhythmic-non-match, arrhythmic-match, and arrhythmic-non-match trials. Therefore, the behavioral differences were compared within the same independently initialized network. We will revise the relevant descriptions in the Abstract, Results, and Methods to make this joint-training procedure more explicit. Second, Figures 2 and 3–4 used different dimensionality-reduction approaches because they addressed different analytical questions. In Figure 2, dPCA was applied to time-aligned population activity to separate task-related variance into temporal, ordinal-position, and stimulus-related components, allowing us to examine how rhythmicity affected each representational component. In contrast, Figures 3 and 4 used standard PCA to construct a low-dimensional population signal for spectral and phase analyses without explicitly demixing task variables. We will clarify this distinction and the rationale for the two analysis pipelines in the revised Methods. Third, STSP was not introduced as an auxiliary module. Previous theoretical and computational studies have suggested that STSP can contribute to working-memory maintenance (Mongillo et al., 2008; Masse et al., 2019). Based on previous studies, we incorporated STSP alongside persistent neural activity to examine how synaptic and consistent neuronal activities jointly support sequential working memory. Under the current architecture and training settings, our ablation experiments showed that RNNs without STSP had difficulty successfully learning the task. We will emphasize this result and further distinguish the roles of STSP in task learning, delay-period maintenance, and the rhythmic advantage. We will also explore whether vanilla RNNs can successfully learn the same task under alternative hyperparameter settings.

      Our core hypothesis is that temporal regularity improves the encoding of sequential information by organizing recurrent population dynamics and phase structure during the encoding period, and that this organization subsequently influences information maintenance during the delay period and behavioral performance. The current results establish a relationship between temporal regularity and encoding-period dynamics, but direct validation of how this organization influences subsequent working-memory maintenance remains insufficient. In the revision, we will focus on strengthening the link between the encoding and delay periods and test whether population dynamics and phase organization during encoding are associated with subsequent information maintenance and behavioral performance.

      To provide a more direct view of the network dynamics, we plan to add intuitive visualizations of network activity and phase structure, allowing readers to evaluate the reported temporal organization with less dependence on dimensionality reduction, filtering, and complex statistical processing.

      The current behavioral advantage of approximately 0.4 percentage points is small, and network performance in both conditions is close to ceiling. Following the reviewers’ suggestions, we plan to evaluate the stability of this effect across learning, random initializations, and representative hyperparameter settings, and to examine whether the rhythmic advantage becomes more pronounced under higher memory load when ceiling effects are reduced. We will also avoid equating statistical significance directly with functional importance.

      Although the manuscript already includes decoding and perturbation analyses of neural activity and synaptic efficacy during the delay period, we will perform additional delay-period analyses to more directly examine whether the temporal organization established during encoding is associated with subsequent working-memory maintenance. These analyses will help distinguish the contributions of rhythmic input to stimulus encoding and memory maintenance.

      We will also test the role of STSP more directly. The current delay-period shuffle results show that synaptic efficacy makes a functional contribution to delay-period maintenance, but this does not yet demonstrate that STSP specifically supports the rhythmic advantage. We will further illustrate the roles of STSP during encoding and delay-period maintenance. In the revision, we plan to increase the number of independent networks in the perturbation analysis and test the interaction between temporal regularity and STSP disruption.

      Finally, we will further tighten the conceptual framing of the manuscript by more clearly distinguishing stimulus-driven phase organization from self-sustained oscillations, and by clarifying that the analyzed signals are model population signals derived from firing-rate population activity rather than LFP or EEG field potentials. Where the current evidence is insufficient to support interpretations in terms of entrainment or self-sustained oscillations, we will adopt more cautious terminology and revise the corresponding conclusions accordingly. We will also clarify the relationship between our phase-organization results and the findings of Liebe et al. (2025) in the revised Results.

      Overall, the revised manuscript will more precisely frame the study around how temporal regularity improves sequential working memory and how this behavioral advantage is associated with the organization of recurrent population dynamics and synaptic-state representations during encoding. After completing the additional training and analyses, we will also make the corresponding model weights, training configurations, and necessary analysis files publicly available to improve reproducibility.

      We again thank the editors and the two reviewers for their detailed and constructive comments. We will revise the manuscript carefully on the basis of these suggestions.

      References

      Mongillo, G., Barak, O., & Tsodyks, M. (2008). Synaptic theory of working memory. Science, 319(5869), 1543–1546. https://doi.org/10.1126/science.1150769

      Masse, N. Y., Yang, G. R., Song, H. F., Wang, X.-J., & Freedman, D. J. (2019). Circuit mechanisms for the maintenance and manipulation of information in working memory. Nature Neuroscience, 22(7), 1159–1167. https://doi.org/10.1038/s41593-019-0414-3

      Liebe, S., Niediek, J., Pals, M., Reber, T. P., Faber, J., Boström, J., Elger, C. E., Macke, J. H., & Mormann, F. (2025). Phase of firing does not reflect temporal order in sequence memory of humans and recurrent neural networks. Nature Neuroscience, 28, 873–882. https://doi.org/10.1038/s41593-025-01893-7

    1. eLife Assessment

      This important study provides new insight into how frontostriatal circuits encode elapsed time and exhibit decision-related dynamics during an auditory change-detection task. Analyses of simultaneously recorded neurons from the frontal orienting field (FOF) and anterior dorsal striatum (ADS) provide solid evidence that FOF and ADS show similar dynamics during evidence accumulation, but that FOF shows stronger movement-aligned reorganization near the time of the decision report. However, the mechanistic interpretation would be strengthened by clearer links between the population-level analyses, single-neuron activity, and relevant circuit anatomy. The work will be of broad interest to systems neuroscientists studying timing, decision-making, and frontostriatal dynamics.

    2. Reviewer #1 (Public review):

      Summary:

      The authors explore how temporal information and decision-related dynamics are represented across FOF and ADS in rats. The authors used Neuropixels to record neurons simultaneously from FOF and ADS during a free-response auditory change-detection task. They then applied single-trial temporal decoding to estimate both the time elapsed since stimulus onset and the time remaining until movement initiation. When neurons in both FOF and ADS were sorted based on decoder weights, they showed ramping and transient bump-like dynamics aligned to stimulus onset. However, around the decision report, FOF showed a clearer ramping signal and stronger movement-aligned population reorganization than ADS. These results suggest that FOF and ADS share similar temporal dynamics during evidence evaluation, but that FOF undergoes a stronger reorganization near decision commitment.

      Strengths:

      (1) The authors recorded large-scale neural populations simultaneously from FOF and ADS, allowing direct and fair comparison between them in the same sessions.

      (2) The free-response auditory change-detection task, which requires rats to evaluate sensory evidence over time and initiate a decision report, is suited to address the question. The behavioral results support that rats used sensory evidence to guide their choices.

      (3) The authors used multiple approaches, including single-trial temporal decoding, decoder-weight PCA, PC loading trajectory, and population-geometry analyses, to explore the FOF and ADS dynamics. These methods provide converging evidence supporting that FOF and ADS share similar temporal dynamics during evidence evaluation but diverge around movement/decision commitment.

      (4) The population-geometry analysis is quite strong and interesting because it compares epoch-specific neural subspaces and quantifies dimensionality and subspace alignment, showing stable subspaces during evidence evaluation and stronger subspace reorganization in FOF near movement initiation.

      Weaknesses:

      (1) The manuscript failed to include histological confirmation of probe placement.

      (2) The direct FOF-ADS decoding comparison in fig 3f and 4f includes only 16 of 61 sessions because of imbalanced unit counts. While controlling for unit number is important, excluding most sessions may waste data. Restricting analyses to only 16 sessions questions the generalizability of the result.

      (3) Fitted regression curves, and ideally confidence intervals, were missing from Figures 3c and 4c. Also, the confusion matrices in Figures 3a/b and 4a/b show a strong preference for predictions in the first and last time bins. The authors did not explain whether this reflects meaningful event-aligned neural activity or an endpoint artifact from decoding time as bounded discrete classes.

      (4) The interpretation of the neuron groups defined by PCA on the decoder-weight matrix was confusing. The authors perform PCA on an N units by T time-bin matrix of LDA decoder weights, then group neurons according to their PC1 and PC2 scores. This is an interesting approach, but the current wording could make readers think that neurons at the extremes of PC1 or PC2 are necessarily the most important neurons for temporal decoding. In fact, these groups appear to represent neurons whose decoder-weight profiles project strongly onto the dominant weight-space patterns. They are not necessarily the neurons that contribute most strongly to decoding accuracy, nor are they necessarily the most common firing-rate dynamics in the raw neural population.

      (5) Discussion is missing some needed context. First, given the causal role of ADS in evidence-accumulation-based choices (Yartsev et al., 2018), and its position as a key node that may integrate input from FOF (Brody & Hanks, 2016), the weaker decision-aligned transition in ADS compared with FOF should have been further discussed. If ADS contributes causally to the decision process, why does it show a much weaker population-state transition near decision commitment in the present data? Second, in DePasquale et al. (2024), more extensive choice vacillation was found in ADS, while greater choice certainty was found in FOF. Does this follow the same principle as the current manuscript, where FOF shows stronger reorganization near decision commitment compared to ADS?

      (6) Current analyses do not fully exploit the simultaneous nature of the recordings. Apart from the comparison of decoding accuracy, most analyses could have been performed and compared based on the data collected independently from two regions.

      (7) Figures 3-10 are hard to read and unpolished. Fonts are too small, and legends/labels are redundant.

    3. Reviewer #2 (Public review):

      Summary:

      This work investigated differences in the temporal dynamics of neural populations in frontal orienting fields (FOF) and anterior dorsal striatum (ADS) in rodents during an auditory change detection task. The relative roles of these two regions have been studied previously and have been shown to play a role in the accumulation of evidence, with FOF converting this evidence into a categorical decision. By focusing on the temporal dynamics of neurons in these regions, the authors identified a subpopulation of neurons within FOF that displayed an abrupt ramping of activity near the time of decision commitment. Both FOF and ADS contained subpopulations exhibiting ramping activity aligned to stimulus onset. This is an interesting finding, suggesting that FOF contains a subpopulation of neurons that transforms accumulating evidence from other subpopulations in ADS and FOF into an action.

      Strengths:

      The conclusions of this paper are mostly well supported by data.

      Weaknesses:

      (1) In the neural analysis, the authors use a technique in which the weights of a linear decoder are used to define a feature vector for each neuron. These weights are used to measure the overall contribution of a neuron in decoding time (from stimulus or decision commitment). Interpreting decoding weights in this way is technically not correct (Kriegeskorte and Douglas, "Interpreting encoding and decoding models"), as a large weight in a decoder is not necessarily indicative of a large effect. Weights in decoding models can become large in order to cancel noise. Alternative analyses, for instance, treating the time series of each neuron as a feature vector, could have supported the conclusions from this technique.

      (2) In this same analysis, it appears that the abrupt change in response in FOF at the time of decision commitment is coming from a single subpopulation of about 130 neurons. In the example session (Figure 8J), there is a clear outlier (the neuron in the top right corner). A closer inspection of the single neuron responses in this group would strengthen the results to confirm the abrupt change in mean population response is not coming from a relatively small number of neurons and sessions.

      (3) The significance of the dynamical motif corresponding to transient bumps was unclear. For example, when looking at Figure 6K-L, I do not see any neuron groups that exhibit a clear transient bump. I would characterize all groups as ramping, with some groups showing steeper ramps. It would be helpful if the figure displayed the fraction of variance explained by PC2 so that it would be clear how much variance the bump motif is contributing. Given that there was no discussion of the functional relevance of this second motif, interpretation of this result is unclear.

      (4) The finding that FOF contains subpopulations which slowly ramp during the trial as well as a subpopulation which acts like a switch that abruptly turns on at the time of decision commitment is interesting and significant and presents several computational questions. For example, is this subpopulation a non-linear readout of the more slowly ramping populations? The approach based on constructing a feature vector for each neuron, projecting these vectors into a low-dimensional subspace, and partitioning into subpopulations is insightful and allowed distinguishing these different computational functions within a single region (FOF). However, I found this particular result to not be clearly stated and obscured by other seemingly less significant results (e.g., existence of the transient bump motif) and other less interpretable analyses (e.g., subspace re-alignment).

    4. Reviewer #3 (Public review):

      Summary:

      This study investigates how frontostriatal circuits encode elapsed time and exhibit decision-related dynamics during an auditory change-detection task. Using population-level temporal decoding and analyses of low-dimensional neural dynamics, the authors compare activity in the frontal orienting field (FOF) and anterior dorsal striatum (ADS). The manuscript addresses an important question in systems neuroscience: how cortical and striatal circuits represent elapsed time and signal action initiation during decision-making.

      The results suggest that FOF and ADS differ in how they represent decision-related information near decision commitment or behavioral report. In particular, FOF shows greater movement-aligned changes in temporal decoding and population geometry than ADS. These findings are potentially important because they may help clarify how cortical and striatal circuits contribute to timing, decision formation, and action initiation.

      Strengths:

      A major strength of the study is its use of population-level analyses to identify temporal structure and movement-aligned changes in neural dynamics. The analyses provide evidence that neural dynamics and low-dimensional population geometry change around the time of behavioral report, especially in FOF. This provides a useful population-level description of decision-related dynamics beyond what could be inferred from average firing rates alone.

      Another strength is that FOF and ADS activity were recorded simultaneously during the same auditory change-detection task. This design strengthens the regional comparison by minimizing confounds related to session-to-session variability, including differences in task engagement, decision accuracy, or other behavioral variables across recordings. The simultaneous recordings therefore provide a strong basis for comparing temporal decoding and population dynamics between cortical and striatal circuits.

      Weaknesses:

      One limitation is that the physiological interpretation of the population-geometry analyses remains somewhat abstract. Concepts such as low-dimensional subspaces, subspace alignment, and subspace rotation are potentially powerful, but it is not always clear what specific changes in neural activity give rise to these effects. For example, it is difficult to tell whether changes in population geometry primarily reflect recruitment of different neurons, or changes in the dominant temporal profiles of the same neurons. This limits the physiological interpretability of the population-level findings.

      A second limitation is that the mechanistic interpretation of the FOF-ADS difference remains underdeveloped. The observed differences could reflect an internally generated transition in frontostriatal dynamics, similar to the dynamical-regime and neural-mode transition described by Luo et al. (2025). Alternatively, they could reflect a circuit-readout process, analogous to the framework proposed by Stine et al. (2023), in which cortical activity drives threshold crossing in a downstream circuit, triggering orienting or motor signals that terminate the decision process. The current manuscript describes the regional differences clearly, but it does not fully discuss these mechanistic interpretations.

      Finally, the strength of the evidence would be easier to evaluate if the manuscript more clearly reported the number of animals contributing to each major analysis and the consistency of the main effects across animals. Because many analyses are performed across sessions, the absence of this information makes it difficult to assess whether the key findings are robust across animals or could be influenced by one or a small number of animals.

    1. eLife Assessment

      This important study combines behavioral testing, fiber photometry recordings, and optogenetic manipulations to understand the role of CRH+ neurons in the PVN of the hypothalamus in social and non-social settings, while varying the degree of familiarity. The approaches used and the finding that these neurons respond to various social (and non-social) stimuli are strong; however, the conclusion that the activity of these neurons is driven solely by unfamiliar situations and the lack of dynamic analysis of fiber photometry data leaves the manuscript incomplete. This elegant study will likely be of interest to the field of social neuroscience and, if outstanding issues are addressed, would expand our understanding of the role of the PVN in behavior.

    2. Reviewer #1 (Public review):

      Summary:

      Here, the authors examine how CRH neurons in the PVN track social behaviours. They use fiber photometry to record the bulk activity of PVN CRH neurons during the resident-intruder test. They find that PVN CRH activity increases when the intruder enters, and also when mice make movements to approach the intruder. They further show that the magnitude of this response differs depending on the familiarity of the mouse. Specifically, if the intruding mouse is unfamiliar, there is a greater PVN CRH response relative to a familiar mouse. The authors argue that this is specific to social familiarity, as they do not see the same differentiation in the PVN CRH response when mice approach a familiar or unfamiliar object. Finally, the authors conduct optogenetic experiments and show that inhibition of PVN CRH neurons reduces social investigative behaviour. The authors then conclude that PVN CRH neurons are a part of a decision-making circuit to influence behaviour in ambiguous settings, specifically that they are a "key component of the neural circuitry underlying rapid social appraisal, linking endocrine regulation to real-time behavioural decision making".

      The data are interesting and novel. They help us understand the dynamics and range of situations in which PVN CRH neurons are activated. There is some overinterpretation of the data and restriction of what this signal means (i.e., specifically driven by unfamiliar social situations), which doesn't seem to be supported by the data. Indeed, PVN CRH neurons are robustly activated by scenarios outside unfamiliar social ones.

      Strengths:

      The experiments are run and presented very beautifully in a sophisticated way. The data are novel and interesting. They help us understand the time course of PVN CRH responding in social and object settings, and how this differs with the familiarity of a social stimulus.

      The optogenetic manipulation is also very nice. The authors optically inhibit during just the first 20 seconds of the resident-intruder test. They find that this inhibition results in a long-term reduction in social behaviours. To me, this supports an idea that the PVN CRH signal triggers a cascade of behaviours, but is not necessarily driving these behaviours per se.

      Weaknesses:

      It would be great to see more sophisticated analysis of the fiber photometry data, which may reveal interesting effects that are currently being occluded by static AUC analysis. One pipeline that is freely available that could be used is found in Jean-Richard-dit-Bressel, Clifford, and McNally (2020) Frontiers in Molecular Neuroscience. Referred to as waveform analysis, this would allow the authors to examine the significance of their data across time. There are multiple points at which this would be interesting. For example, in Figure 3F, it is possible that differences between the familiar and unfamiliar objects emerge. Also, there seems to be one outlier in this figure in the familiar object group. What happens if it is removed (Figure 3H)?

      Similarly, what do these signals look like when aligned with making contact with the social or object stimuli? It is possible that the objects do not elicit a difference depending on familiarity when approaching because: (1) they are not moving, and (2) it is unclear whether they are familiar or not until contact is made, consistent with the object recognition literature. What would inhibition of the PVN CRH signal do to investigative behaviours directed towards objects?

      Finally, given the robust nature of the response to the approach to the objects and familiar mouse, why is this signal being argued to predominantly act in unfamiliar social settings? The lack of difference between the familiar and unfamiliar objects doesn't negate the importance of this signal. To me, this is the most interesting finding: PVN CRH neurons that are usually activated in stressful situations can also be robustly activated by familiar objects. Relatedly, while the authors argue that inhibition of PVN CRH neurons only reduces social behaviours in the unfamiliar case, there is likely a floor effect in the behaviours that they are looking at, which occludes observation of a reduction via optical inhibition.

    3. Reviewer #2 (Public review):

      Summary:

      The authors investigated the role of hypothalamic CRH neurons in social behavior. They performed fiber photometry recordings in mice from CRH neurons and showed that novel conspecifics trigger stronger and more prolonged responses compared to familiar conspecifics and objects. The activity of CRH neurons appears to be related to risk assessment, as interactions with juvenile unfamiliar mice (lower-risk conspecifics) trigger responses similar to those of familiar adult mice. Behaviorally, CRH neurons were linked to increased anogenital investigation of unfamiliar compared to familiar mice. Optogenetic suppression of CRH neurons decreased anogenital sniffing of unfamiliar conspecifics.

      Strengths:

      The manuscript is elegant, and the results are compelling. The approaches are well justified, and the methods are validated (eg: Arch inhibition).

      The findings substantiate the role of CRH neurons in responses to stress and uncover the involvement of these neurons in the assessment of social risk.

      Weaknesses:

      These are not weaknesses, just some observations: It is somewhat surprising that CRH neurons respond similarly to familiar and unfamiliar objects; it would be good to have more insights into that aspect.

      Similarly, the novel context by itself is expected to lead to increased activity of CRH neurons (based on data from the last author's lab as well as other labs in the field). It is somewhat surprising (and interesting) that the novel environment did not affect the magnitude of CRH responses to unfamiliar conspecifics.

    1. eLife Assessment

      This important work provides new insights into the role of lysine acetylation of alpha-synuclein, the protein involved in Parkinson's Disease. The evidence is convincing and the work will be of interest to researchers in the fields of protein biophysics and post-translational modifications.

    2. Reviewer #1 (Public review):

      [Editors' note: the authors have revised the work in response to the original reviews.]

      Summary:

      This paper describes experiments with alpha-synuclein (aS) with acetylated lysines (acK) at various positions. Their findings on how to use non-canonical amino acid (ncAA) mutagenesis to generate aS with acetylated lysines are valuable. The paper then continues with a range of experiments to characterise the acetylated alpha-synuclein constructs at different positions, with the aim of providing insights into which sites are relevant to disease or their function inside cells. The paper concludes these experiments with the suggestion that inhibiting the Zn2+-dependent histone deacetylase HDAC8 to potentially increase acetylation at lysine 80 may have therapeutic benefit. However, the relevance of most of these experiments is unclear, mainly as the filaments that form from these constructs are different from those observed in human disease (but see below for more details). Moreover, using the recombinantly produced acetylated versions of alpha-synuclein to normalise mass-spectrometry data, the authors themselves report that acetylation of alpha-synuclein does not differ between individuals with Parkinson's disease or healthy controls.

      Strengths:

      The authors report difficulties with chemical synthesis and then decide to make these constructs using non-canonical amino acid (ncAA) mutagenesis, which seems to work reasonably well (yields vary somewhat). In the Conclusion section, the authors report that they used these recombinant proteins to obtain quantitative insights into the levels of acetylation of lysines in individuals with PD versus healthy controls, for which they find no significant differences. This part of the work is valuable.

      Weaknesses:

      The authors then use circular dichroism to show that aSyn with acK at position 43 has less alpha-helical content. From this result, they deduce that "only this site could potentially perturb aS function in neurotransmitter trafficking", but no experiments on neurotransmitter trafficking were performed.

    3. Reviewer #2 (Public review):

      Summary:

      Shimogawa et al. studied the effect of lysine acetylation at different sites in the alpha-synuclein (aS) sequence on the protein-membrane affinity, seeding capacity in the test tube and in cells, and on the structure of fibrils, using a range of biophysical methods. They use non-canonical amino acid (ncAA) mutagenesis to prepare aS lysine acetylated variant at different sites.

      Strengths:

      The major strength of this paper is the approach used for the production of site-specific lysine acetylated variants of aS using ncAA mutagenesis, as well as the combination of a range of biophysical methods together with cellular assays and structure biology to decipher the effect of lysine acetylation on aS-membrane binding, seeding propensity, and fibril structure. This approach allowed the author to find that lysine acetylation at positions 12, 43, and 80 led to lower seeding capacity of aS in the test tube and in cells, but only acetylation at lysine 80 did not affect aS-membrane interaction. These results suggest that lysine acetylation at position 80 may be protective against aggregation without perturbing the proposed functional role of aS in synaptic plasticity.

      Weaknesses:

      SDS is not a good membrane model to investigate the effect of lysine acetylation on aS membrane-binding because it is a harsh detergent and solubilizes membranes. Negatively charged vesicles or vesicles made of a mixture of lipids mimicking the lipid composition of synaptic vesicles are more accepted in the field to study aS-membrane interactions. The authors used such vesicles for the FCS experiments, and they could be used for the initial screening of the 12 lysine acetylated variants of aS.

    4. Reviewer #3 (Public review):

      Shimogawa et al. describe the generation of acetylated aSyn variants by genetic code expansion to elucidate effects on vesicle binding, aggregation, and seeding effects. The authors compared a semi-synthetic approach to obtain acetylated aSyn variants with genetic code expansion and concluded that the latter was more efficient in generating all 12 variants studied here, despite the low yields for some of them. Selected acetylated variants were used in advanced NMR, FCS, and cryo-EM experiments to elucidate structural and functional changes caused by acetylation of aSyn. Finally, site-specific differences in deacetylation by HDAC 8 were identified.

      The study is of high scientific quality, and the results are convincingly supported by the experimental data provided. The challenges the authors report regarding semi-synthetic access to aSyn are somewhat surprising, as this protein has been made by a variety of different semi-synthesis strategies in satisfactory yields and without similar problems being reported.

      The role of PTMs such as acetylation in neurodegenerative diseases is of high relevance for the field, and a particular strength of this study is the use of authentic acetylated aSyn instead of acetylation-mimicking mutations. The finding that certain lysine acetylations can slow down aggregation even when present only at 10-25% of total aSyn is exciting and bears some potential for diagnostics and therapeutic intervention.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      The authors then use circular dichroism to show that aSyn with acK at position 43 has less alpha-helical content. From this result, they deduce that "only this site could potentially perturb aS function in neurotransmitter trafficking", but no experiments on neurotransmitter trafficking were performed.

      We agree with the reviewer that neurotransmitter trafficking studies would be interesting, but they would presumably require the use of acetylation mimic mutants (Lys-to-Gln mutations), which we would want to validate by comparison to our semi-synthetic proteins with authentic AcK. Such experiments are planned for a follow-up manuscript, and we will investigate the reviewer’s suggested experiment at that time. Thus, we have not modified the manuscript to address this issue.

      Subsequently, they measure the aggregation speed of the variants in seeded aggregation experiments with preformed fibrils (PFFs) from WT aSyn, and conclude that acK at positions 12, 43, and 80 yields slower aggregation. They reach similar conclusions when measuring seeded aggregation in primary cultures. As far as I understand it, the seeding experiments in cells use seeds that are assembled from partially acetylated alpha-synuclein, but that are made of non-acetylated wildtype alpha-synuclein, and the alpha-synuclein that is endogenous in the cells is also non-acetylated (or at least not beyond what happens in these cells at endogenous levels). It is therefore unclear how the cellular seeding experiments relate to the in vitro aggregation assays with (partially) acetylated substrates.

      We understand the reviewer’s concerns and have modified the manuscript to clarify that the method of in vitro seeding really reports on the impact of acetylation on the elongation phase of aggregation. We have also clarified that this is different than the role that acetylation plays in seeding cellular aggregation with pre-acetylated fibrils. We note that having the monomer population acetylated in cells presents technical challenges that might also be addressed with Gln mutant mimics, and we plan to pursue such experiments in the follow-up manuscript described above.

      Anyway, both aggregation experiments ignore that the structures of aSyn filaments in Parkinson's disease (PD) or multiple system atrophy (MSA) are different from those formed in these experiments, and that, therefore, the observed aggregation kinetics are likely irrelevant for the speed with which disease-relevant filaments form in the brain.

      Finally, the authors describe the cryo-EM structure of mixtures of acK80:WT aSyn filaments, which are predominantly made of WT aSyn, with a previously described structure. Filaments made of only acK80 aSyn have a modified arrangement of this structure, where the now neutral side chain of residue 80 packs inside a hydrophobic pocket. The authors discuss differences between the acK80 structures and those of other structures from in vitro assembled aSyn filaments, none of which are the same as those observed from PD or MSA brains, nor are any attempts made to transfer observations from the in vitro experiments to the structures of disease. The relevance of the cryo-EM structures for human disease, therefore, remains unclear.

      The Conclusion on p.20 mentions an interesting and valuable result: the authors used the acetylated recombinant proteins to determine the extent of acetylation within human protein samples by quantitative liquid chromatography MS (SI, Figures S41-S49). Their conclusion is that "The level of acetylation was variable - no clear trend was observed between healthy control and patients - nor between patients of different diseases (SI, Table S4, Supplementary Data 1)" This result implies that acetylation of aS is not directly related to its pathogenicity, which again adds doubts on the disease-relevance of the results described in the rest of the paper.

      We acknowledge the concerns raised in the above paragraphs and believe that they can all be addressed by clarifying our purpose. The different fibril polymorphs adopted in PD and MSA are likely the result of an interplay of many PTMs and non-proteinaceous cofactors. Therefore, we are not necessarily trying to claim that our AcK80 fold is populated in health or disease, but that by driving Lys80 acetylation, one could push fibrils to adopt this conformation, which is less aggregation-prone. A similar argument has been made in investigations of alpha-synuclein glycosylation and phosphorylation. Our results in Figure 9 imply that Lys80 acetylation could be increased with HDAC8 inhibition. We have revised the manuscript to make these ideas clearer, while being sure to acknowledge the limitations noted by Reviewer #1.

      Reviewer #2 (Public review):

      Weaknesses:

      SDS is not a good membrane model to investigate the effect of lysine acetylation on aS membrane binding because it is a harsh detergent and solubilizes membranes. Negatively charged vesicles or vesicles made of a mixture of lipids mimicking the lipid composition of synaptic vesicles are more accepted in the field to study aS-membrane interactions. The authors used such vesicles for the FCS experiments, and they could be used for the initial screening of the 12 lysine acetylated variants of aS.

      We have noted this shortcoming in revisions of our manuscript, but have not performed new experiments as we do not believe that using vesicles instead would change the conclusions of these experiments (that only AcK43 produces an effect, and a modest one at that).

      It would help the reader to have the experimental details (e.g., buffer, protein/lipid concentrations) for the different assays written in the figure legend.

      We have added additional detail to the figure captions.

      The authors use an assay consisting of mixing 10% fibrils + 90% monomer to investigate the effect of lysine acetylation on aS. However, the assay only probes fibril elongation and/or secondary processes. The current wording can be misleading, and the term aggregation could be replaced by seeding capacity for clarity. For example, the authors state that lysine acetylation at sites 12, 43, and 80 each inhibits aggregation, but this statement is not supported by the data. Instead, the data show that the acetylation at these sites slows down the fibril elongation and thus decreases the seeding capacity of aS fibrils. In order to state that lysine acetylation has an effect on aS aggregation, fibril formation, the author should use an assay where the de novo formation of fibrils is assessed, such as in the presence of lipid vesicles or under shaking conditions.

      As noted in our response to Reviewer #1, we have clarified which phase of aggregation we were investigating in our in vitro experiments.

      It is not clear from the EM data that the structures of the different lysine acetylated variants are different, unlike what is stated in the text.

      We feel that it is clear from structures in Figure 8 and the EM density maps in Figure S38 that the AcK80 fold is indeed different. Although the overall polymorphs are somewhat similar to WT, the position of K80 clearly changes upon acetylation, altering the local fold significantly and the global fold more moderately. We have added backbone RMSD calculations to quantify the differences in WT and AcK80 folds in Figure SX. They differ by ~5 Å in the fibril core region.

      Reviewer #3 (Public review):

      Weaknesses:

      The challenges the authors report regarding semi-synthetic access to aSyn are somewhat surprising, as this protein has been made by a variety of different semi-synthesis strategies in satisfactory yields and without similar problems being reported.

      We understand the reviewer’s surprise and have edited the manuscript to clarify that the NCL yields were not unusually low, but were comparable to ncAA yields, and since it is significantly easier to scan AcK positions using ncAAs, we felt that ncAAs are the method of choice in this case.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Cryo-EM data processing: particles were pre-processed in cryosparc and ChatGPT-generated scripts. Are priors on tilt and psi angles defined through this procedure? And are segments from individual filaments kept strictly in the same half-sets for gold-standard estimation of resolution? Absence of psi and tilt priors will lead to worse refinements than otherwise possible. Worse, a mixture of segments from different filaments into the two half-sets may lead to overestimated resolution estimates.

      We thank the reviewer for their thoughtful analysis of our cryo-EM data processing approach. In response, we have modified our approach and added additional commentary on processing to the Materials and Methods section.

      “CryoSPARC (.cs) files were then converted to RELION STAR files using the PyEM csparc2star.py script. Following this data conversion, a custom Python script developed with ChatGPT precisely determined the start and end coordinates of each fibril. This script processed the csparc2star.py star file output to create coordinate pairs based on the cryoSPARC Fibril ID, notably without transferring the original tilt and psi angles from CryoSPARC to RELION. The output of this custom script served as the input for the autopick RELION extraction step.

      To handle curved fibrils, a special segmentation strategy was implemented: the script traced the coordinates along the fibril, generating a new start and end coordinate pair for individal segments, with the segment's end coordinate assigned either after spanning 10 particles (around 50 nm in length) or when the end of the fibril was reached. This resulted in shorter, straighter segments for processing. Finally, the createAutopick function from cryolo_boxmanager_tools.py in crYOLO was utilized to generate a STAR file linking the final particle coordinates to their movie files, which was then used to perform particle extraction in RELION. The standard RELION image processing pipeline then followed, including 2D classification, refinement, and 3D classification.

      This approach resolves the issue with curved fibrils, but also results in all picked fibrils being 50 nm or shorter in RELION. Consequently, segments from the same fibril receive unique fibril IDs in RELION and may be split into different half-maps. Although this avoids random splitting of individual particles without regard to their origin from the same fibril, it could still cause an overestimation of resolution.

      To assess any potential overestimation of resolution, another script was created to reassign the fibril ID of each particle in the RELION STAR file to that of the closest particles in the cryoSPRAC data, effectively ensuring that all particles from an individual fibril have the same fibril ID and are not assigned to different half-maps. This revised STAR file was then used as the input images STAR file for 3D auto-refinement in RELION. A comparison of maps before and after reassigning the fibril IDs shows negligible effects on the resulting structures and on their estimated resolution, indicating that the original procedure did not results in any significant overestimation of the resolution.”

      Author response image 1.

      RELION Cryo-EM maps before (gray) and after (yellow) fibril ID reassignment for (A) WT-A (B) WT-B (C) <sup>Ac</sup>K<sub>80</sub>-A (D) <sup>Ac</sup>K<sub>80</sub>-B, showing that the new processing approaches did not significantly change the fibril structures.

      (2) The 25% acK80 structure in S52 is understood to be wildtype-only. The authors mention in the main text that this is because of the strong density for the K80 side chain. An additional argument would be that an acK80 would leave an unshielded negative charge on the neighbouring E46, as K80 and E46 form a salt bridge in this structure.

      We appreciate the reviewer’s idea and have included a comment on the salt bridge impact.

      (3) It would be valuable to include side views of the density for all reported reconstructions to assess to what extent the beta-rungs are separated.

      The requested side views have been included in Figure SX.

      (4) Methods sections should be moved into the main text of the paper.

      This Materials and Methods portion of Supporting Information has been moved to the main text.

      Reviewer #3 (Recommendations for the authors):

      (1) We suggest removing "all" from the manuscript title as this claim might not hold up in the future.

      We understand the reviewer’s concern. Our title was meant to imply “all currently known” rather than “all” forever, but as this wording is awkward, we have deleted “all” as suggested.

      (2) The white font in Figure 1A is sometimes hard to read, especially on the yellow-green background between amino acids 70-80.

      We have changed this to black font.

      Figure 1 panels C-E are not referenced in the manuscript text?

      References to the Figure 1 panels have been added.

      In panel C, a structure is predicted, but based on what data? Why is there both a small and a big structure in panel C?

      Explanations of the Figure 1C images have been added to the caption.

      (3) Typo ε-acetyllysine -> Nε-acetyllysine

      This has been corrected throughout.

      (4) You report solubility issues during NCL, and these are typically alleviated by the use of chaotropes during ligation. Please specify what you mean by "standard NCL conditions" and include parameters such as guanidine concentration, pH, concentrations, volumes, and temperature.

      These details have now been added to the methods section.

      (5) According to Figure 2D, the desulfurization was incomplete. Please explain.

      The small peak observed next to the product peak is an adduct with sinapic acid (+206 Da), the matrix mixed in for MALDI acquisition. We do not see a sign of +32/64Da peak which would correspond to incomplete desulfurization

      Author response image 2.

      (6) Please include sequences of your constructs for ncAA mutagenesis. Especially, which intein was attached at what position. This is important for other groups to fully understand the production of acetylated aSyn variants.

      The full DNA sequence has been added to Supporting Information. This plasmid has also been reported previously in the referenced publications.

      (7) You state ncAA mutagenesis "yielded 0.11-1.5 mg" aSyn. I guess this is per-liter culture expression medium?

      This has been corrected.

      (8) The resolution/DPI of Figures 4, 5, S14-16, and S18 should be increased. They look blurred compared to Figures 6 or S19.

      These figures have been updated.

      (9) You provided only summaries in Figures 4 and 5 because of space restrictions. However, it would be nice to have at least some selected individual experiments (WT, acetylation K12/43/80) right next to it without the need to switch to S14, S15, or S16.

      Select experimental data has been added to main text Figures 4 and 5 as requested.

      (10) Please increase the size of the microscopy images in Figure 6 (in Figure S19, it looks much better).

      This size of the images in Figure 6 has been increased.

      (11) The labelling A1-A3 in Figure S33 was somehow unclear to me.

      These three panels show different sections of the TEM grid, illustrating heterogeneity in the 25% <sup>Ac</sup>K<sub>12</sub> fibrils that was not observed for fibrils of the other acetylation variants. We have clarified this in the figure caption.

      (12) In Figure 9, you present preliminary results regarding site-specific deacetylation by HDAC8. These results are interesting, but considering the exploratory nature of this in vitro experiment, I suggest toning down the highly enthusiastic discussion of potential in vivo effects.

      We understand the reviewer’s concern and have mitigated the claims of potential impact from these in vitro results.

    1. eLife Assessment

      The authors propose the existence of an AQP4-positive tanycyte-associated canal system in the hippocampus based on histological, ultrastructural, immunohistochemical, and RNA-based approaches. There are serious concerns regarding the evidence, and the identification of the tanycyte-related structures can be questioned. At this stage, the evidence for the central claims of the manuscript has to be regarded as inadequate.

    2. Reviewer #1 (Public review):

      In the manuscript by Fabian-Fine et al., the authors employ neuroanatomy to investigate aquaporin-4 expression in cells they consider tanycytes and their supposed involvement in tau tangles and amyloid-beta plaques in the hippocampus. This study includes samples from three mice and two Alzheimer's disease (AD) patients.

      My key concern and question is whether the cells presented in the manuscript are tanycytes. Tanycytes are specialized ependymoglial cells located in the circumventricular organs and are known to express specific markers. Importantly, they are not myelinated cells, which is a crucial distinction that the authors do not address.

      Additionally, the methodologies described in the manuscript lack clarity and controls. For instance, the use of Cdh5-GCaMP882 mice is not adequately justified. It is unclear what these mice contribute to the study's objectives, particularly concerning the aim of investigating waste removal processes in the brain. Moreover, the rationale behind the purported "fluorophore uptake experiments" is unclear and appears to involve the uptake of fluorophore-labeled goat anti-rabbit secondary antibody, which seems implausible to me.

      The hypotheses and claims presented in this manuscript are not sufficiently substantiated and are conceptually unclear. The notion that amyloid beta and tau proteins play structural roles in a hypothesized "tanycytes"-derived canal network is not sufficiently supported by the evidence. Furthermore, the study lacks rigorous data to convincingly establish the proposed interactions between these proteins and the processes of waste internalization.

      In conclusion, due to conceptual and methodological issues, I consider the current evidence as inadequate to support the primary claims.

    3. Reviewer #2 (Public review):

      Summary:

      In this study, the authors propose the existence of an AQP4-positive tanycyte-associated canal system in the hippocampus and suggest that this system participates in waste clearance and contributes to Alzheimer's disease pathology. Using histological, ultrastructural, immunohistochemical, and RNA-based approaches, the manuscript attempts to reinterpret amyloid-β plaques and tau-associated structures as components of a tanycyte-derived waste-internalization system. The work is conceptually ambitious and raises observations that may stimulate discussion regarding glial organization and waste clearance in the diseased brain.

      Strengths:

      A strength of the manuscript is the combination of imaging modalities and anatomical observations across mouse and human tissue. Some of the reported morphological features are intriguing and may warrant additional investigation. The study also attempts to integrate structural observations with broader hypotheses regarding neurodegeneration and Alzheimer's disease.

      Weaknesses:

      The central interpretation depends almost entirely on identifying the observed hippocampal structures as tanycytes, and the evidence supporting this conclusion remains insufficient. Tanycytes are classically associated with ventricular regions in circumventricular organs, particularly in the third ventricle and median eminence region, yet the manuscript does not provide sufficiently specific anatomical or molecular evidence to convincingly distinguish the described structures from astrocytic, ependymal, radial glial-like, oligodendroglial, myelin-associated, vascular-associated, or degenerative elements. The marker profile used throughout the study, particularly the reliance on AQP4 labeling and Luxol-positive structures, is not sufficiently selective to establish tanycyte identity, especially in pathological tissue where reactive glial changes may occur.

      This becomes particularly important because the manuscript repeatedly interprets Luxol-positive and myelin-associated structures as tanycytic processes or "myelin-derived tanycyte protrusions," despite tanycytes not being known to produce myelin. Alternative explanations are not sufficiently explored. Some of the canal-like structures shown in Figure 4 also resemble vascular profiles, and additional vessel markers would be necessary to exclude this possibility.

      Several of the proposed structures and mechanisms are also difficult to reconcile with established cell biology and neuroanatomy. The introduction of new terminology such as "tanysomes," "waste receptacles," and "toroids" further extends the interpretation beyond what is currently demonstrated experimentally.

      The discussion and integration of the existing literature on tanycytes are also insufficient. Tanycytes themselves are not clearly introduced; the manuscript does not adequately discuss what is currently established regarding tanycyte anatomy, ventricular localization, morphology, and function. Foundational literature defining tanycyte biology, including work from the Prévot group or others, is largely absent despite its central importance to the field. Because the manuscript proposes a substantial departure from established neurobiological concepts, it is particularly important that previous literature be discussed comprehensively and critically. The current version does not sufficiently contextualize the proposed model within the existing literature on tanycyte, AQP4, glymphatic, and Alzheimer's disease, making it difficult to evaluate what is genuinely novel versus what is merely being reinterpreted. It is also not entirely clear what is genuinely new here compared with the authors' previous work, particularly reference 11, which appears to present a highly similar conceptual framework.

      More broadly, several of the manuscript's mechanistic conclusions extend well beyond the available evidence. The proposal that amyloid-β plaques and tau pathology represent hypertrophic tanycyte-derived waste structures is provocative and potentially interesting, but currently remains largely correlative and speculative. At several points, it becomes difficult to distinguish direct observations from broader mechanistic interpretation. The manuscript itself acknowledges that the proposed glial-canal hypothesis contradicts the current understanding of nervous system organization and states that ultrastructural serial-section analysis would be required to unambiguously determine the origin of the myelinated profiles described. This point is critical because the study's central conclusions depend on the assumption that these structures are tanycyte-derived. At present, this interpretation remains insufficiently demonstrated, which substantially limits the strength of the broader pathological and mechanistic conclusions proposed throughout the manuscript.

      Although access to human material is understandably limited, the study appears to include only one male and one female AD patient, making it difficult to assess the reproducibility or frequent these structures are across individuals and pathological conditions. The manuscript would benefit from clearer characterization of prevalence, reproducibility, and variability across samples.

      Overall, the manuscript presents an unconventional and thought-provoking model that may stimulate discussion. However, the evidence currently provided does not convincingly establish tanycyte identity for the described hippocampal structures, and several of the broader disease-related interpretations would require substantially stronger anatomical and molecular evidence before the proposed model can be convincingly supported.

    4. Author response:

      Reviewer #1 (Public review):

      My key concern and question is whether the cells presented in the manuscript are tanycytes. Tanycytes are specialized ependymoglial cells located in the circumventricular organs and are known to express specific markers. Importantly, they are not myelinated cells, which is a crucial distinction that the authors do not address.   

      We agree with the reviewer that tanycytes that have been described in the third ventricle have not been reported to be myelinated. However, our study was conducted on the hippocampal formation that borders the ventral horn of the lateral ventricles. We will include images of the myelin-forming ependymal cells that we refer to as tanycytes  

      Additionally, the methodologies described in the manuscript lack clarity and controls.

      We will expand on our methods section and include controls.

      For instance, the use of Cdh5-GCaMP882 mice is not adequately justified. It is unclear what these mice contribute to the study's objectives, particularly concerning the aim of investigating waste removal processes in the brain. Moreover, the rationale behind the purported "fluorophore uptake experiments" is unclear and appears to involve the uptake of fluorophore-labeled goat anti-rabbit secondary antibody, which seems implausible to me.

      Most experiments described in this manuscript were carried out on human brain. However, functional studies will have to be carried out on rodent brain. We will thus process rodent tissue as well to test whether our observed findings are consistent between human and rodent.  The animals used for these experiments were raised for bladder research and are wild type regarding neuronal and glial cells, particularly using the Cy3 channel. Utilizing these brains for our experiments has allowed us to test rodent tissue at both light- and electron-microscopic levels without having to sacrifice additional animals. The consistency of our findings between human and rodent brain further supports that the calcium indicator in the vascular system of these mice did not affect neurons or glial cells.

      Regarding the uptake experiment: When we initially discovered that myelin-forming macroglia form waste-internalizing glial canals within neuronal in spider brain it was unclear where the AQP4-immunoreactive cells were located. The somata of the myelin-forming cells lacked AQP4 immunoreactivity. Suspecting a synergistic interaction between the myelin-forming and AQP4 expressing cells whereby the myelinforming cells create the canal structure that sequesters waste from the neuron and the AQP4-expressing cells create a convective flow toward the waste-internalizing structures. However, unable to locate the somata of these cells, we submerged a freshly dissected spider brain with the attached surrounding tissue intact in physiological spider saline and slowly added blue vital dye solution to test which cells would internalize the dye. We then identified the (blue) cells in the lining of the dorsally located tubular system that we routinely detached from our brain preparations explaining why we were unable to locate these cells.  Immunolabeling of this tubular system revealed the cells that reside in the lining of this tubular system (see Author response images 1 and 2) the original (Figure 8) shows their long slender processes. Interestingly, this system is continuous with the stomatogastric system. 

      Author response image 1.

      Shows the proposed canal system in spiders

      Author response image 2.

      AQP4-immunoreactive cells in the spider primitive ventricular system that we localized due to similar uptake experiments we have conducted in mouse brain.

      We have utilized this method in mouse brain to test the validity of our postulation, that ependymal tanycytes internalize the presented fluorochrome from extracellular spaces and test which areas and structures may be involved in this uptake. As demonstrated in this experiment, the alveus, and a fine network of cell processes within the brain parenchyma show fluorescence, indicative that they internalize the fluorochrome from extracellular spaces. We used goat-coupled fluorochrome to further test with a FITCcoupled secondary antibody that the observed fluorescence is indeed due to uptake of the goat-coupled secondary antibody and not due to intrinsic autofluorescence. Control preparations lacked this fluorescence.  To further test our postulation that the uptake is indeed AQP4-mediated we have applied an AQP4-blocker, which showed a significantly reduced fluorochrome uptake compared to the controls without this blocker.

      As we state in the text, we are aware of the limitations of this experimental design, however, like in our spider experiments we consider these findings helpful as they likely show an overview of the cellular network in the hippocampus that governs waste-uptake and may help identify suitable target areas for similar studies on organotypic tissue cultures utilizing two-photon microscopy.

      The hypotheses and claims presented in this manuscript are not sufficiently substantiated and are conceptually unclear. The notion that amyloid beta and tau proteins play structural roles in a hypothesized "tanycytes"-derived canal network is not sufficiently supported by the evidence. Furthermore, the study lacks rigorous data to convincingly establish the proposed interactions between these proteins and the processes of waste internalization. In conclusion, due to conceptual and methodological issues, I consider the current evidence as inadequate to support the primary claims.

      We respectfully disagree with this comment and hope that the inclusion of additional evidence together with the clear visibility of this canal system in the spider brain will encourage the reviewer to investigate this possibility themselves. We cannot ignore large amounts of amyloid beta-immunolabeled receptacles emanating from tanysomes in swell-bodies and declare them fixation artifacts, particularly when we demonstrate the expression of Presenilin 1 and APP in swell bodies. We furthermore encourage the reviewer to revisit myelinated cells in the brain in both depictions in the available literature and actual brain preparations. We have not been able to locate actual electronmicrographs of longitudinal sections through neurons that show myelination past the axon hillock at the EM-level consistent with our current understanding of myelination. The only depiction of this form of myelination we found were schematic drawings. We will include several new images that show such longitudinal sections through neurons that are easily obtained and we have numerous additional images that we are happy to share. In all our preparations (mouse, rat and human) the myelination pattern is consistent with the images we will depict in new figures 1 and 2.  Not to bring attention to this inconsistency would be dishonest scientific conduct.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors propose the existence of an AQP4-positive tanycyte-associated canal system in the hippocampus and suggest that this system participates in waste clearance and contributes to Alzheimer's disease pathology. Using histological, ultrastructural, immunohistochemical, and RNA-based approaches, the manuscript attempts to reinterpret amyloid-β plaques and tau-associated structures as components of a tanycyte-derived waste-internalization system. The work is conceptually ambitious and raises observations that may stimulate discussion regarding glial organization and waste clearance in the diseased brain.

      Strengths:

      A strength of the manuscript is the combination of imaging modalities and anatomical observations across mouse and human tissue. Some of the reported morphological features are intriguing and may warrant additional investigation. The study also attempts to integrate structural observations with broader hypotheses regarding neurodegeneration and Alzheimer's disease.

      Weaknesses:

      The central interpretation depends almost entirely on identifying the observed hippocampal structures as tanycytes, and the evidence supporting this conclusion remains insufficient. Tanycytes are classically associated with ventricular regions in circumventricular organs, particularly in the third ventricle and median eminence region, yet the manuscript does not provide sufficiently specific anatomical or molecular evidence to convincingly distinguish the described structures from astrocytic, ependymal, radial glial-like, oligodendroglial, myelin-associated, vascular-associated, or degenerative elements. The marker profile used throughout the study, particularly the reliance on AQP4 labeling and Luxol-positive structures, is not sufficiently selective to establish tanycyte identity, especially in pathological tissue where reactive glial changes may occur.

      As mentioned in our response to reviewer 1 we have now included additional experimental evidence that demonstrates the myelinated ependymal cells and additional gene expression experiments. 

      This becomes particularly important because the manuscript repeatedly interprets Luxolpositive and myelin-associated structures as tanycytic processes or "myelin-derived tanycyte protrusions," despite tanycytes not being known to produce myelin. Alternative explanations are not sufficiently explored. Some of the canal-like structures shown in Figure 4 also resemble vascular profiles, and additional vessel markers would be necessary to exclude this possibility.

      Several of the proposed structures and mechanisms are also difficult to reconcile with established cell biology and neuroanatomy. The introduction of new terminology such as "tanysomes," "waste receptacles," and "toroids" further extends the interpretation beyond what is currently demonstrated experimentally.

      The discussion and integration of the existing literature on tanycytes are also insufficient. Tanycytes themselves are not clearly introduced; the manuscript does not adequately discuss what is currently established regarding tanycyte anatomy, ventricular localization, morphology, and function. Foundational literature defining tanycyte biology, including work from the Prévot group or others, is largely absent despite its central importance to the field. Because the manuscript proposes a substantial departure from established neurobiological concepts, it is particularly important that previous literature be discussed comprehensively and critically. The current version does not sufficiently contextualize the proposed model within the existing literature on tanycyte, AQP4, glymphatic, and Alzheimer's disease, making it difficult to evaluate what is genuinely novel versus what is merely being reinterpreted. It is also not entirely clear what is genuinely new here compared with the authors' previous work, particularly reference 11, which appears to present a highly similar conceptual framework.

      More broadly, several of the manuscript's mechanistic conclusions extend well beyond the available evidence. The proposal that amyloid-β plaques and tau pathology represent hypertrophic tanycyte-derived waste structures is provocative and potentially interesting, but currently remains largely correlative and speculative. At several points, it becomes difficult to distinguish direct observations from broader mechanistic interpretation. The manuscript itself acknowledges that the proposed glial-canal hypothesis contradicts the current understanding of nervous system organization and states that ultrastructural serialsection analysis would be required to unambiguously determine the origin of the myelinated profiles described. This point is critical because the study's central conclusions depend on the assumption that these structures are tanycyte-derived. At present, this interpretation remains insufficiently demonstrated, which substantially limits the strength of the broader pathological and mechanistic conclusions proposed throughout the manuscript.

      Although access to human material is understandably limited, the study appears to include only one male and one female AD patient, making it difficult to assess the reproducibility or frequent these structures are across individuals and pathological conditions. The manuscript would benefit from clearer characterization of prevalence, reproducibility, and variability across samples.

      Overall, the manuscript presents an unconventional and thought-provoking model that may stimulate discussion. However, the evidence currently provided does not convincingly establish tanycyte identity for the described hippocampal structures, and several of the broader disease-related interpretations would require substantially stronger anatomical and molecular evidence before the proposed model can be convincingly supported.

      We agree with the reviewer that it is important to correctly investigate and describe cellular structure. The first author of this manuscript is a 30-year veteran of published cellular ultrastructure and the three-dimensional reconstruction of cells and entire cell networks. We have spent the last five years to try and confirm our current understanding of cellular structure, and we are unable to reproduce our current identification of cell structure. One example is mentioned in response to reviewer 1. We are unable to find electonmicrographs in publications that show myelination of neurons consistent with the countless schematic depictions available online and in the literature. This includes publications about myelination. Our findings are all consistent with the new images we will include in our revised manuscript. Longitudinal sections through neurons are easily obtained and neurons can be followed well beyond the axon hillock.

      A second example that is inconsistent with our current understanding of cells and biochemical processes in cells are ‘astrocytes’ and ‘reactive astrocytes’. We will show in the revised manuscript swell-bodies have no defined cytoplasm that every cell requires to fulfil basic cellular functions required for survival. It is very apparent to a structural expert that swell-bodies lack cytoplasm. The ‘consistency’ of this lacking cytoplasm is ‘inconsistent’ with fixation artifacts. In Author response image 3 we demonstrate that immunolabeling for AQP4 shows two types of structures. (1) Immunolabeled structures that are void of immunolabeling in their lumina and show receptacle-like immunoreactivity along the outside (panels I,J,L,M image below) consistent with swellbodies as indicated by the provided amyloid beta-stained and Luxol H&E-stained swellbodies that are associated with immunolabeled receptacle-shaped structures on the outside (panels K,N). A structurally trained eye quickly recognizes that these are not immunolabeled cells but are consistent with swell-bodies and referred to as ‘reactive astrocytes’ in the literature. We show in panel O what an immunolabeled cell looks like, with slender processes and immunolabeled cytoplasm. In this image in panels G1-4 we demonstrate that swell-bodies contain AQP4 mRNA explaining why they are immunoreactive for this protein. It is important that we bring attention to these details and do not randomly describe structure just based on a signal. This is exactly the point we make and I hope that the reviewer recognizes our expertise in cellular neuroscience and in particular in recognizing cellular structure. 

      Again, I urge the scientific community not to dismiss our findings but to actually study the validity of our findings. 

      Author response image 3.

      In the revised manuscript we will include additional experiments that we have carried out, that have also increased the sample number of tested human brain. We have in total so far investigated 13 different human brain samples, six of which are AD-affected. 

      We will revise our discussion to explain our observations in context with the literature better and include so much evidence in support of our hypothesis that it would be unreasonable to dismiss all this compelling and logical evidence.

      This research was not planned; it resulted from our accidental discovery in the spider system when our animals struggled with early onset neurodegeneration that we needed to address. This is when we recognized the waste-internalizing role of myelin in giant spider neurons. As we will discuss in the revised manuscript, such systems are highly conserved throughout evolution, and this is what made us realize that the only images of myelinated neurons we could find were either schematic drawings, single cross sections through myelinated cell profiles or very high magnification insets that also did not show the actual neuron that is myelinated.

      One can argue that there are both myelinated and unmyelinated axons. However, the varicose projections clearly originate in the ependymal lining and double-label for AQP4. This is consistent with our postulation and inconsistent with our current understanding of myelination. 

      I sincerely urge the neuroscience community to re-visit myelination in the brain, we have done this for the past five years with extensive experience in this field and the only hypothesis that is supported by our findings is presented in this manuscript. Please do not dismiss these findings, the spiders have uncovered a waste canal system in the brain and putting this system in place will help us to gain a better understanding regarding neurodegenerative diseases. Lastly, and maybe most importantly, understanding how this system works in spiders and how the myelin is structurally anchored to microtubule that were missing in our degenerating spiders has allowed us to identify the cause for this sudden neurodegeneration and rescue our tropical, cold-blooded spider colony by installing a new heating system and raising the room temperature so that the coldsensitive microtubules no not dissociate anymore.  

      We would like to thank the reviewers to strengthen the content of this manuscript with their critical comments, we hope that our revision will help clarify some doubts.

      (1) Pasquettaz R, Kolotuev I, Rohrbach A, Gouelle C, Pellerin L, Langlet F. Peculiar protrusions along tanycyte processes face diverse neural and nonneural cell types in the hypothalamic parenchyma. Journal of Comparative Neurology. 2021;529(3):553. doi: 10.1002/cne.24965. PubMed PMID: edsgcl.646848213.

    1. eLife Assessment

      This work describes a valuable method for monitoring pathogens using an innovative, field-deployable device that attracts animals and collects saliva on filter paper in a non-invasive manner. While the strategy is effective at collecting samples and ensuring their preservation, support for some claims remains incomplete. This study will interest scientists in various fields, ranging from biodiversity, ecology, and conservation to infectious diseases and public health.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes the development and validation of a low-cost device to identify viruses from saliva samples of animals non-invasively. This device was tested under laboratory conditions to assess whether viruses could be recovered in different environmental conditions and after different durations of time. The devices were then used to sample mice and cats in shelters to assess utility.

      Strengths:

      Sampling animals is cost-effective and highly labour-intensive, and this device has the potential to substantially improve surveillance. The device is relatively low-cost, and the authors demonstrate that the virus can be obtained from these filter papers after different durations of time and in different environmental conditions.

      Weaknesses:

      The authors do not discuss if different volumes were obtained from different animals (for example, due to different behaviours or attractiveness of the odour baits). Additionally, it appears the virus results were cross-validated using the serological status of the animals. While I am not an expert on FIV, there seems that there could be potential for different levels of viral shedding, and it would be more prudent to cross-validate against blood or another gold standard sample. Finally, the statistical analysis could be improved as there appear to be relatively few replicates and limited analysis conducted.

    3. Reviewer #2 (Public review):

      Summary:

      The study introduces an innovative device designed to collect non-invasive saliva samples from animals using disposable cassettes with odor attractants and filter paper. The authors aimed to validate this tool for pathogen monitoring, specifically by detecting pathogen RNA in animal models. While the concept is compelling and the problem statement well-framed, the validation of the device for pathogen detection was not achieved. For example, the rabies virus was not detected in the chosen model, and results were limited primarily to FeLV. The work highlights the potential of saliva-based sampling for microbiota analysis, but the rationale for virus selection and the experimental design require further clarification. Overall, the study presents a novel approach with promise, though its current scope is better suited to microbiota monitoring rather than pathogen surveillance.

      Strengths:

      The innovative design of the device, which enables non-invasive saliva collection through disposable cassettes with odor attractants, represents a creative and practical advance in sampling methodology. The authors undertook an extensive experimental effort, generating a substantial amount of data that highlights the feasibility of saliva-based monitoring. The rationale for exploring saliva as a medium is valid, and the work successfully shows that the device can be applied to microbiota profiling, where the strongest results were obtained. This methodological innovation could be valuable for expanding non-invasive approaches to animal health monitoring.

      The authors acknowledge that metabarcoding sequencing has limitations; however, the study could be refocused on the microbiota in general rather than on pathogen detection. They could give greater prominence to the taxonomic composition of microorganisms in saliva using high-throughput sequencing. That is where they obtained the most results.

      Weaknesses:

      Despite the enormous experimental effort undertaken, the results fall short of the expected success of the proposed test. The rationale and criteria for virus selection are not clearly explained, leaving the experimental design insufficiently justified.

      The central aim of validating the device for pathogen detection was not achieved, particularly in the case of the rabies virus. The mouse infection model used for the rabies virus does not seem to adequately replicate the natural course of the disease. This could explain, at least in part, the negative results obtained.

      Of the three viruses evaluated, satisfactory results were obtained only for FeLV, and the sample size remains limited. According to the literature reviewed, this virus is not common in wild cats, so the applicability of the results would appear to be limited primarily to domestic cats.

      The collected samples were stored at −80 {degree sign}C for later analysis, which likely contributed to the high Ct values observed with the device. The need to store samples at low temperatures may be a limitation to applying this technique in wildlife sampling scenarios where access to dry ice or liquid nitrogen tanks may be difficult.

      Stating that the device can be used for pathogen monitoring in wild animals is not desirable, since the viruses for which results were obtained are not relevant in wild animals. On the other hand, claiming that this is a tool for monitoring diseases in endangered species is also misleading. Endangered species are typically scarce and therefore would not be the reservoirs that these surveillance efforts should target. In fact, groups such as wild rodents would be a better target for monitoring zoonotic pathogens.

    1. eLife Assessment

      This study provides an important and biologically plausible account of how human perceptual judgments of heading direction are influenced by a specific pattern of motion in optic flow fields known as retinal curl. By combining psychophysical experiments and neural modeling, the authors demonstrate that what was previously considered an incidental, 'nuisance' signal actually serves as a functional control signal for estimating heading and steering toward a fixated target. The evidence supporting the role of curl signals is convincing and advances our understanding of vision-based navigation at the behavioral level. A particular strength of the work is the direct manipulation of curl within flow fields, demonstrating that it effectively cancels or reverses heading biases. This work provides an invaluable framework for future exploration into the neural mechanisms of steering control based on retinal curl.

    2. Reviewer #1 (Public review):

      Summary:

      This carefully executed study uncovers the functional relevance of curl signals that impinge on the retina every time an observer's gaze direction and movement direction are not aligned. This finding is important, highlighting the functional role of an abundant incidental signal (curl in retinal motion) that has thus far believed to be a nuisance that needs to be filtered out of the retinal motion stream. As such, the study forms an important contribution to the emerging recognition that incidental sensory signals are not a challenge to the sensorimotor system, but contain functionally relevant and effectively used visual signals. The study's evidence is compelling: A combination of psychophysical experiments and critical manipulations, control theory and neural modeling makes an internally consistent and biologically plausible case for the role of curl signals in estimating heading direction. The experimental and modeling results clearly go beyond previous studies and significantly advance our understanding of vision-based navigation.

      Strengths:

      The study has its strengths in the combination of psychophysical experiments and critical manipulations, control theory and neural modeling, which together make an internally consistent and biologically plausible case for the role of curl signals in estimating heading direction.

      This study uncovers the functional relevance of curl signals that occur on the retina when an observer is moving and gaze is not straight ahead. The experimental and modeling results clearly go beyond previous studies and significantly advance our understanding of vision-based navigation.

      Another clear strength is that the study uses tightly controlled experimental manipulation to provide strong test cases for the hypothesis that curl is used for visual navigation. These conditions are important to constrain the proposed model (and future models) of heading control.

      The modeling is very clearly described and the modeling and analysis code is published and freely available. The authors go beyond a back-of-the-envelope control model and show how it might be implemented at the neural-circuit level. The model is biologically plausible.

      Weaknesses:

      I see no major weaknesses of the study. I expect it to inspire future research that extends these findings to a wider range of visual environments (including walking in natural scenes), motion speeds and kinds of movements.

      Comments on revised version.

      I have no additional comments for the authors.

    3. Reviewer #2 (Public review):

      This study examines how curl in the retinal flow field can be used as a control variable for estimating and controlling the heading of a moving observer. The basic idea (which is not entirely new, see Matthis et al. 2022) is that translation along a path with eccentric gaze (meaning that the subject is not heading toward the point they are looking at) produces a pattern of optic flow on the retina with a rotational component around the point of fixation (which can be captured by the mathematical "curl" operator). The sign and magnitude of retinal curl varies with heading relative to the point of fixation, such that curl can be used as a control variable to steer rightward or leftward to move toward the fixated target. The authors perform behavioral experiments and show that there are biases in perceived heading that seem to be largely governed by retinal curl. They also show that a simple controller model can use curl to steer toward a target, and they provide a neural network model that provides a biologically-plausible implementation of the controller (although there are some questions about that).

      There is a core of interesting work here that I think can be important to the field. However, there is a lack of clarity on several important fronts, including design of the behavioral experiments, presentation of the behavioral data, conceptual framing of what curl can and cannot do, etc. Equally importantly, the manuscript is not written in a manner that will make it accessible to most vision scientists. I consider myself to be pretty knowledgeable about optic flow, and I had to read most of the manuscript 3 or 4 times to be able to understand the bulk of it. And my experience is that most vision scientists do not understand optic flow well, so I fear that most of the readers that the authors should want to reach would struggle to understand the work. As written, this is mainly going to make an impact on a handful of optic flow gurus. Thus, this manuscript is going to need a major overhaul to clarify important issues and make this more accessible.

      Major issues:

      (1) The manuscript contains inconsistent, if not misleading, messaging about what information retinal curl does, and does not, provide regarding heading estimation. In the Abstract, the authors state: "We propose an alternative: the visual system utilizes retinal curl directly to estimate heading, rendering the explicit recovery of the FOE unnecessary." Based on my understanding of the rest of the manuscript, I find this statement to be a misrepresentation for two main reasons:<br /> a. To "directly estimate heading" relative to what? When not qualified, most people interpret "heading" to mean an observer's heading relative to the world (or some allocentric reference frame). But retinal curl only gives information about an observer's heading relative to the point on which their eyes are fixated. Moreover, that point of fixation will change every few hundred milliseconds in natural viewing, so the retinal curl will change with each new fixation even as heading relative to the world remains unchanged. So, I think most readers would grossly misinterpret the claim that retinal curl can be used "directly to estimate heading". Indeed, in the authors' controller model, the initial heading needs to be given and then the controller can work. But from where does the visual system get the initial heading, since it does not come from curl? These issues are left hanging. Thus, while curl can provide a very useful input for steering toward a fixated target, other signals are needed to estimate heading relative to the world. This has to be made much clearer early on, and a conceptual schematic diagram might help. Also, the authors generally do not specify the reference frame of the variables they are talking about, leaving lots of room for misinterpretations. It should be clear each time they are talking about a variable, such as heading, whether it is relative to the fixation target, body, world, etc.<br /> b. It seems to me that retinal curl will depend on other variables, in addition to heading relative to the fixation target. For example, it seems to me that the magnitude of retinal curl will depend on self-motion speed, the depth structure of the scene, the angle of elevation of the fixated target, and perhaps others. This is not discussed at all, and many readers would get the misguided impression that there is a 1:1 mapping from curl to heading (relative to fixation). If I am right that this is not correct, it means that retinal curl can tell the observer whether to steer right or left to move toward the fixated target, but it cannot tell them how much to steer. Indeed, in the authors' controller model, there is a free parameter that calibrates curl to angle. It makes sense that this works to fit trajectory data that are given from a fixed environment, but it is unclear how the brain would use retinal curl to control steering when these other variables are uncertain or changing unpredictably. Moreover, how does the system change the mapping from curl to steering command as the location of fixation changes relative to the current heading? These are issues that need to be brought up in framing the problem and discussed at some length. If the authors can show mathematically that retinal curl is only dependent on heading (relative to fixation) and not any of these other variables, it would be very valuable to show the equations for this relationship.

      (2) The description of the behavioral experiment and presentation of behavioral data leaves a lot to be desired.<br /> a. First, it is stated (line 158) that "Participants continuously reported their perceived direction of self-motion while maintaining fixation on the yellow dot." Again, reference frame is completely unspecified. Participants were reporting their perceived heading relative to what? The fixation target? The world? What exactly were the instructions given to the subjects to perform the task? Based on the description of how perceived paths are computed (line 166-), it seems to be presumed that subjects are reporting their heading relative to the world because those angles are then converted into x and z coordinates in what I presume is a world-centered reference frame. But how do we know that subjects are accurately reporting their heading relative to the world? What if they are biased in their reports by the location of the fixation target relative to the scene, or by some other reference signal? Is it possible for the authors to rule out the possibility that perceptual biases seen in the unaltered curl condition result from observers not fully adopting the assumed reference frame of the task? If this cannot be firmly excluded, it seems to create problems for the rest of the study.<br /> b. I also feel that there is a mismatch between what the behavioral task requires and what the controller model does. Subjects are apparently asked to report their heading relative to the world, but the controller model only controls their heading relative to the point that they are fixating. I understand how this is resolved in the model, but I think this type of distinction is buried and will not be apparent to most readers. Again, the reference frames of what is being measured and controlled need to be specified explicitly in all parts of the paper, and the authors needs to explain how the system would combine curl-based control with some other measures of (at least initial) heading for world-centered heading to be computed. All of the assumptions need to be clearly specified.<br /> c. Second, I found it frustrating that the authors never present raw perceptual data from the observers. Rather, in Figure 2, we see reconstructed trajectories that are perfectly smooth with no indications of noise whatsoever. Since these paths are computed from the perceptual reports, there must be some noise inherent in them. The figures should represent this uncertainty somehow, and it should be explained how these perfectly smooth trajectories are obtained.

      (3) "...the magnitude of retinal curl in the fovea can specify the body trajectory relative to gaze (Matthis et al., 2022)." The main idea put forward by the authors here seems to overlap heavily with this statement that they attribute to Matthis et al. 2022. While I think this paper still adds importantly to the topic, the authors do not discuss how their findings are different from those of Matthis et al. 2022, why they are an important extension, etc. Readers should not have to go read this other paper to have any idea how the present findings are placed in importance relative to the literature.

      (4) The analysis and treatment of eye movements is extremely weak. The authors discarded trials for which gaze deviated from the fixation point by more than 3 degrees (which is a LOT given that the eye speeds are generally in the neighborhood of 0.5 deg/sec), and they provide basic stats on the distribution of positions. But this largely misses the point: it is not small position errors that are likely to matter, but rather velocity errors. Even a small amount of retinal slip of the target while it is being pursued will cause image motion that is going to alter the optic flow field around the fixation target. So, for example, the retinal curl field may no longer be centered on the fixation target. How do we know that some of the perceptual biases are not influenced by image motion resulting from imperfect tracking of the fixation target? This needs to be analyzed and discussed.

      (5) I found the sections of text comparing the separate and joined fits (starting line 287) to be a bit too rosy. The authors show the separate fits in the main text, and it is not very surprising that these fits are good given that the model has 30 parameters, and these data are pretty low dimensional. The authors only show the joined fits in the supplement, and they say that they are almost as good as the separate fits (indeed they are better in a model comparison sense, but this is 30 parameters vs. 2 parameters). However, when I look at the fits of the joined model in the supplement, I don't find them to be very impressive. In particular, the model grossly misses the data for the straight paths for several subjects (e.g., id5, id6, id8, id10). And fitting the straight paths would presumably be easiest. This implies that the joined model is really missing something and that fitting the curved paths interacts strongly with fitting the data for different fixation target locations on the straight path. I think that the authors should discuss the results a bit more soberly and tone down their conclusions here.

      (6) The section of the paper on neural simulations (starting line 387) has a few weaknesses. First, why are only straight paths simulated here? This does not seem to provide a very rigorous test of the model. Second, it is awkward that the simulation results are presented in units of pixels, rather than degrees. Third, the authors seem to downplay the fact that the neural estimates of heading seem to oscillate rather wildly (over a range of hundreds of pixels, whatever that means, see especially Fig. S16). It was far from clear to me how an estimate of heading with these large oscillations is useful. It would seem to require that heading estimates are integrated over substantial lengths of time to be reliable. It was therefore unclear how the model produces such smooth paths from these oscillating estimates.

      Comments on revised version.

      Overall, the authors have done a responsible job of responding to the comments of my previous review, and the manuscript is substantially improved. There are a few points on which I still do not completely agree with the authors, and I think these are important to document for the record:

      (1) Introduction: "Pure visual decomposition should function regardless of 3D depth or whether the rotation stems from an active eccentric fixation." Perhaps in a world of noiseless perfect computation, this might be true. But I generally disagree. When there is more depth structure in an environment, then translation of the observer is generally going to create greater motion parallax. That is a fact that I don't think can be disputed. And greater motion parallax should help to decompose optic flow into components related to translation and rotation (the latter of which is not depth dependent), especially when there is noise in estimating location motion vectors.

      (2) Related to point #9 of my previous review: I had asked why the authors believed that retinal curl was computed in area MSTd. Their response is that previous studies (i.e., Graziano et al. 1994) show selectivity to spiral motion stimuli in MSTd. That is true, but those studies typically placed the spiral stimulus centered on the MSTd receptive field, hence they were not presenting something like retinal curl as defined here. So, I think it is still an open question as to where in the brain retinal curl is encoded, and from which areas it would be possible to decode retinal curl from population responses.

      (3) Related to point #10 of my previous review: I had asked about biological plausibility of the gaze-centered inhibition signal in the model. The authors' response is that parietal neurons show gain fields in which response depends (usually monotonically) on eye position. This is true, but it is not a trivial jump from gain fields in individual neural responses to a gaze-centered inhibition signal, and I think the authors should have been more forthcoming about the lack of an established neural signal that directly signals what they want in their model.

      (4) The authors point out that the perceptual biases they measure take a few seconds to emerge and they attribute this to temporal integration. But in their curl manipulations, they temporally average over a 2.4 second window in computing the curl signals that they use to cancel or over-cancel curl. So, it is not clear whether some of the delay in the behavioral effects might result from their computations.

      (5) Related to point #13 of my previous review: I had asked about empirical evidence for the assumption of a relationship between the heading preferences of MSTd neurons and their receptive field locations. In response, the authors state that such a relationship is built into the Layton and Browning (2014) model. While that is a precedent, citing another model as a response to a question about empirical evidence is not a convincing response. If there is no empirical evidence to support such a relationship, it would have been better for the authors to acknowledge this.<br /> Given the way that the eLife review model works, it is not necessary for the authors to address these comments, but I think they should be included in the public review record.

    4. Reviewer #3 (Public review):

      Major strengths include the use of realistic retinal motion recorded during virtual walking, an elegant manipulation of curl, converging behavioral and modeling evidence, and grounding in control theory. This provides a novel and important contribution to our understanding of how the brain processes motion information and intuition about how that information might be used to guide steering. In addition, they provide a computational mechanism by which retinal flow curl can be used as a control signal.

      The revised ms has been strengthened by more explicit discussion of the literature where there has been mixed evidence for the use of the Focus of Expansion. Since the ms is a strong test of the use of curl as a heading signal, this allows a deeper understanding of the importance of the finding and historical context. The ms has also been strengthened by a more explicit discussion of integration of the time-varying signal over periods of several seconds, which is an important demonstration. The implications of the ms are still a little unclear, as the results involve visual judgements in seated subjects. The use of different sources of information when humans walk from one place to another in real life may be complex and involve a variety of different sources of information.

    5. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editors for their positive assessment of our manuscript, and all the referees for their constructive comments, which have substantially improved this work. We welcome the opportunity to address referee #2's points for the public record, as they highlight key theoretical nuances and valuable future research directions.

      (1) We appreciate the reviewer's point that, in a noisy biological system, the increased motion parallax provided by a rich 3D depth structure naturally aids in separating translation from rotation. We fully agree on this point. Our argument aimed at highlighting a fundamental theoretical distinction. Pure algebraic decomposition algorithms are mathematically capable of solving heading on flat planes. The fact that human perception often shows biases in these zero-depth conditions, unless extra-retinal cues are present, suggests that the visual system does not rely on a generalized, global de-rotation algorithm. Instead, it relies on heuristic, depth-dependent structural signals (like motion parallax and retinal curl). We maintain that while depth certainly reduces noise, its strict necessity points toward an ecologically grounded control strategy rather than a noisy global decomposition process.

      (2) We think the reviewer raises a valid point regarding the exact neural locus of retinal curl encoding. It is true that Graziano et al. (1994) utilized centred spiral stimuli rather than the spatially offset curl geometries defined in our task. We view the spiral tuning of MSTd not as a direct, one-to-one mapping of full-field retinal curl, but rather as the foundational computational building block required to extract such a signal. We fully agree with the reviewer that identifying exactly where and how this population response is decoded into a unified, gaze-relative retinal curl signal remains an exciting and open empirical question for future neurophysiological research.

      (3) We acknowledge the reviewer's call for transparency here. The transition from well-documented multiplicative gain fields (which modulate response amplitude based on eye position) to a direct, localized gaze-centered inhibitory drive is indeed a theoretical abstraction in our model. We utilized this localized inhibition as a functional mechanism to demonstrate how sensory evidence and spatial priors might competitively interact within a standard Mexican-hat recurrent architecture. While gain fields clearly establish that parietal networks integrate gaze position, we agree that the exact local-circuit implementation mapping these gain fields to the specific inhibitory dynamics we modeled has yet to be empirically established.

      (4) The reviewer smartly questions whether the 3-5 second behavioral biases delay emerges from the 2.4-second smoothing window used in our computational flow manipulation. It is important to clarify that this 2.4-second window was used solely to stabilize the computed curl signal against high-frequency gait oscillations. This smoothing was restricted strictly to the modeling phase of the controller and neural model and was not applied to the participants' responses. The gradual build-up of their perceptual bias over 3-5 seconds represents their own intrinsic temporal integration of this trajectory, independent of the smoothing parameters used to smooth the curl used in the fitting of the controller and neural modelling.

      (5) We concede the reviewer's point that citing a computational model (Layton & Browning, 2014) does not constitute direct empirical evidence for a relationship between MSTd heading preferences and their receptive field locations. Our intention was to highlight a successful theoretical framework that elegantly organizes known properties of MSTd into a system capable of bypassing global de-rotation. We readily acknowledge that direct, single-cell empirical validation of this specific topographic relationship is currently lacking in the literature, and we appreciate the reviewer ensuring this distinction is clearly noted for the record.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides an important and biologically plausible account of how human perceptual judgments of heading direction are influenced by a specific pattern of motion in optic flow fields known as retinal curl. By combining psychophysical experiments and neural modeling, the authors demonstrate that what was previously considered an incidental "nuisance" signal actually serves as a functional control signal for estimating heading and steering toward a fixated target. While the evidence for the role of curl signals is convincing and advances our understanding of vision-based navigation, the work's impact would be strengthened by situating these findings among other cues that contribute to heading estimation, and by clarifying both the time course of these computations and their generalizability across different navigational contexts.

      We thank the editors and reviewers for their insightful feedback and positive assessment of our study. In this revised version, we have made substantial modifications to better situate our findings within the broader landscape of cues contributing to heading estimation, while also clarifying the time course of these computations and their generalizability across different navigational contexts. These points are included in new sections in the revised discussion.

      In addition, and following eLife guidelines, we have moved the methods to the end and make sure that the manuscript reads well without needing to go through methods first.

      Next, we address all the concerns raised by the reviewers.

      Reviewer #1 (Public review):

      We appreciate Reviewer #1’s very positive feedback. Incorporating the perspective of ‘incidental’ sensory signals is a valuable suggestion that aligns perfectly with our findings. We agree that this perspective significantly strengthens the impact of our paper.

      In the revised version we have added a last section in the Discussion (Generalizability and Testable Predictions) to comment on the functional utility of 'incidental' signals and incorporated the suggested references. In addition, in the same heading, we briefly elaborate on the predictions and generalizability of the model and possible manipulations that might affect the integration between sensory evidence (curl signal) and straight-ahead prior.

      Reviewer #1 (Recommendations for the authors):

      It would be great if the authors could discuss the implications and predictions of their model.

      First, from a broader perspective, the study forms an important piece in the emerging recognition that incidental sensory signals are not a nuisance to the sensorimotor system, but contain functionally relevant and effectively used visual signals (Rolfs & Schweitzer, 2022). The authors may want to appreciate their contribution to this perspective in the discussion of the impact of their results. Indeed, a similar shift in perspective has been realized in the recognition that saccade-induced motion signals are not entirely suppressed but play a functional role for gaze correction (Schweitzer & Rolfs, 2021).

      Second, the authors could spell out additional predictions of their proposal: What are experimental manipulations that could shift the balance between relying on a straight-ahead prior and sensory estimation of curl? What would happen in extreme cases of such sensory evidence? When would priors become overwhelmingly influential?

      As commented in the public review we have now included these two aspects in the discussion.

      Minor point: After equation 11, the authors may want to specify that, like position, gaze g is also coded as {x,y}, just like position p.

      While the neural model details have been moved to Appendix 2 (including this equation), we added text before Equation 22 (previously Eq. 11) clarifying that gaze is encoded in image coordinates. We omitted point index i because the equation applies to all image points relative to a given gaze g.

      Reviewer #2 (Public review):

      We appreciate the reviewer’s feedback regarding the formalization of our reference frames. We agree that certain definitions were implicitly assumed rather than explicitly stated. We have revised the manuscript to provide all necessary self-contained information, ensuring that the geometry of the task response and the definition of heading are unambiguous. In the last paragraph of the revised introduction, we make clear the response frame of reference which is also included in the caption of fig 1. Also, we have addressed the gap between the task response (in world coordinates) and the functional role of the controller. This is particularly discussed in the discussion (section: reference frames) in which we also provide (and rule out) potential alternatives to our response biases. We also address all the other points raised by the reviewer.

      Major issues:

      (1) The manuscript contains inconsistent, if not misleading, messaging about what information retinal curl does, and does not, provide regarding heading estimation. In the Abstract, the authors state: "We propose an alternative: the visual system utilizes retinal curl directly to estimate heading, rendering the explicit recovery of the FOE unnecessary." Based on my understanding of the rest of the manuscript, I find this statement to be a misrepresentation for two main reasons:

      (a) To "directly estimate heading" relative to what? When not qualified, most people interpret "heading" to mean an observer's heading relative to the world (or some allocentric reference frame). But retinal curl only gives information about an observer's heading relative to the point on which their eyes are fixated. Moreover, that point of fixation will change every few hundred milliseconds in natural viewing, so the retinal curl will change with each new fixation even as heading relative to the world remains unchanged. So I think most readers would grossly misinterpret the claim that retinal curl can be used "directly to estimate heading". Indeed, in the authors' controller model, the initial heading needs to be given, and then the controller can work. But from where does the visual system get the initial heading, since it does not come from curl? These issues are left hanging. Thus, while curl can provide a very useful input for steering toward a fixated target, other signals are needed to estimate heading relative to the world. This has to be made much clearer early on, and a conceptual schematic diagram might help. Also, the authors generally do not specify the reference frame of the variables they are talking about, leaving lots of room for misinterpretations. It should be clear each time they are talking about a variable, such as heading, whether it is relative to the fixation target, body, world, etc.

      In our study, participants were instructed to report their “perceived direction of self-motion” by aligning a rotational encoder (steering wheel) with the direction they felt they were moving within the 3D simulated scene. Consequently, participants reported their instantaneous heading in a world-centered reference frame, from which the 3D trajectories were reconstructed. Since the reviewer had to infer this information, we have clarified this point at the end of the introduction, legend of figure 1, methods (now at the end of the ms.) and discussion to ensure it is immediately evident.

      Participants were informed that the initial heading (i.e. θ<sub>0</sub> in our controller nomenclature) was oriented “straight ahead” relative to their body which was aligned longitudinally with the experimental room. We have modified Figure 1B and revised the Methods section to explicitly clarify this initial alignment and the instructions provided to participants.

      In the revised manuscript, we have clarified that while the participant’s report is world-centered, the retinal curl provides a gaze-relative heading signal. Although this was already mentioned, we emphasize this point. In natural navigation toward a fixated target, a world-centered vector is often unnecessary; an error signal indicating heading relative to fixation is sufficient (as the reviewer also notes). However, the initial alignment of the heading within the 3D scene allows the brain to “calibrate” this internal controller, mapping the retinal curl signal onto the 3D world coordinates required for the task. Ad commented above, a new section in the discussion addresses and hopefully clarifies the relation with the controller.

      The reviewer also asks how we can be certain that participants were reporting in world coordinates rather than an alternative frame, such as “heading relative to the fixation target.” We believe our “Cancelled Curl” (and over-cancelled) conditions provide the most compelling evidence to rule out this alternative. In these conditions, the physical position of the fixation target in the scene remained identical to the unaltered flow condition. If participants were simply reporting heading relative to the fixation target’s spatial location, the observed biases should have persisted regardless of the flow manipulation. Instead, the bias vanished when the curl was removed. This causal evidence proves that the bias is driven by the retinal motion signal (curl) rather than the spatial orientation of the eyes or the target’s position in the scene. Furthermore, the temporal evolution of the response supports a world-centered integration (in agreement with Warren et 2001 Nat Neuro.). For simulated straight paths, the perceived heading remains straight for the first few seconds (consistent with the initial world-centred alignment), with biases only emerging after approximately 3 seconds of integration (a point we elaborate on in our response to Reviewer #3). Had participants been responding based on a simple gaze-relative reference frame from the onset, these biases would have manifested significantly earlier. We have incorporated these points into the revised Discussion to better frame our findings alongside other cues, such as the Focus of Expansion (FOE) and egocentric visual direction that contribute to heading estimation.

      Finally, we have rephrased the abstract sentence for clarity. However, we maintain that the original premise remains valid once the world-centered initial heading is aligned with the gaze-centered reference frame.

      (b) It seems to me that retinal curl will depend on other variables, in addition to heading relative to the fixation target. For example, it seems to me that the magnitude of retinal curl will depend on self-motion speed, the depth structure of the scene, the angle of elevation of the fixated target, and perhaps others. This is not discussed at all, and many readers would get the misguided impression that there is a 1:1 mapping from curl to heading (relative to fixation). If I am right that this is not correct, it means that retinal curl can tell the observer whether to steer right or left to move toward the fixated target, but it cannot tell them how much to steer. Indeed, in the authors' controller model, there is a free parameter that calibrates curl to angle. It makes sense that this works to fit trajectory data that are given from a fixed environment, but it is unclear how the brain would use retinal curl to control steering when these other variables are uncertain or changing unpredictably. Moreover, how does the system change the mapping from curl to steering command as the location of fixation changes relative to the current heading? These are issues that need to be brought up in framing the problem and discussed at some length. If the authors can show mathematically that retinal curl is only dependent on heading (relative to fixation) and not any of these other variables, it would be very valuable to show the equations for this relationship.

      The reviewer notes that we must be clear about the relationship between curl and heading (relative to fixation) and the variables that affect curl. We also thank the reviewer for encouraging to add the equations that show the relation of curl with additional variables. We have now included these equations in appendix 1.

      Beyond the discrepancy between heading (θ) and gaze (ψ), curl is geometrically determined by translational self-motion speed (v), eye height (h), and pitch (α). More specifically, curl = (v.sinψcosα)/h. The derivation is now included in appendix 1. Since h = dsinα, where d is the 3D distance to the fixation point, we could express cos α as a function of distance. Certainly, there is not a 1:1 map from curl signal to heading relative to gaze (e.g. θ-ψ). Participant would need to know v and eye height plus extra-retinal information. Frenz et al (2003, Vis Res.) showed that people can estimate self-motion directly from optic flow, across different simulated eye height and gaze angle; extra-retinal information can, in addition, provide knowledge to ψ and α. It is then plausible that the visual system can use and transform the curl signal from a qualitative directional cue (i.e. steering left or right of fixation) into a quantitative steering command. By combining curl with knowledge of gaze orientation and eye height, the visual system can resolve ambiguities in the flow field and utilize curl as a more precise error signal for locomotor control. These aspects are now included in the new version of the discussion.

      (2B) I also feel that there is a mismatch between what the behavioral task requires and what the controller model does. Subjects are apparently asked to report their heading relative to the world, but the controller model only controls their heading relative to the point that they are fixating. I understand how this is resolved in the model, but I think this type of distinction is buried and will not be apparent to most readers. Again, the reference frames of what is being measured and controlled need to be specified explicitly in all parts of the paper, and the authors need to explain how the system would combine curl-based control with some other measures of (at least initial) heading for world-centered heading to be computed. All of the assumptions need to be clearly specified.

      We thank the reviewer for this point. We have addressed the alignment of the reference frames in our response to Issues 1a and 2a. Once the initial orientation (θ<sub>0</sub>) is established in the world frame, the controller model generates steering adjustments that directly translate into heading predictions within that same world reference frame. By treating the perceptual report as an output of the locomotor controller, we resolve the discrepancy between the steering task and the reported heading.

      (2c) In addition, I found it frustrating that the authors never present raw perceptual data from the observers. Rather, in Figure 2, we see reconstructed trajectories that are perfectly smooth with no indications of noise whatsoever. Since these paths are computed from the perceptual reports, there must be some noise inherent in them. The figures should represent this uncertainty somehow, and it should be explained how these perfectly smooth trajectories are obtained.

      We respectfully disagree with the reviewer’s interpretation regarding data smoothing. The thin lines in Figure 2 represent the mean 3D paths derived directly from the response variable (θ<sub>t</sub>) across trials of identical conditions for each participant (as detailed in the ‘Computation of Perceived Path’ section). No smoothing or filtering has been applied to these plotted trajectories other than computing the mean across trials. We also wish to remind the reviewer that the raw data and analysis code remain publicly accessible for further inspection. Having said that, we include now a supplementary figure showing an example of raw data responses as a function of time. This figure will be a supplemental figure of main Figure 2 (now provisionally included in the Suppl Information).

      Regarding the visual representation: in earlier versions of the manuscript, we included shaded 95% Confidence Intervals (CIs) in Figure 2. However, this addition rendered the plot overly cluttered and obscured the individual trajectories. We therefore chose to present individual participant means (thin lines) alongside group averages (thick lines) to emphasize inter-subject variability. For clarity, the 95% CIs are explicitly displayed in Figure 3, where the data density is more conducive to shaded areas.

      (3) “...the magnitude of retinal curl in the fovea can specify the body trajectory relative to gaze (Matthis et al., 2022)." The main idea put forward by the authors here seems to overlap heavily with this statement that they attribute to Matthis et al. 2022. While I think this paper still adds importantly to the topic, the authors do not discuss how their findings are different from those of Matthis et al. 2022, why they are an important extension, etc. Readers should not have to go read this other paper to have any idea how the present findings are placed in importance relative to the literature.

      We have updated the Discussion to more specifically align our findings with Matthis et al. (2022). We emphasize that our study provides the perceptual validation for their ecological observation that the FOE is often too unstable for reliable use, whereas foveal curl remains a robust signal for path estimation. Our paper provides the causal link, since we manipulate curl in real-time (the ‘cancelled & over cancelled curl’ condition) providing the critical evidence that perceived heading is affected by this signal. The relation with this previous study is made clear in the revised discussion.

      (4) The analysis and treatment of eye movements is extremely weak. The authors discarded trials for which gaze deviated from the fixation point by more than 3 degrees (which is a LOT given that the eye speeds are generally in the neighborhood of 0.5 deg/sec), and they provide basic stats on the distribution of positions. But this largely misses the point: it is not small position errors that are likely to matter, but rather velocity errors. Even a small amount of retinal slip of the target while it is being pursued will cause image motion that is going to alter the optic flow field around the fixation target. So, for example, the retinal curl field may no longer be centered on the fixation target. How do we know that some of the perceptual biases are not influenced by image motion resulting from imperfect tracking of the fixation target? This needs to be analyzed and discussed.

      We thank the reviewer for noting that retinal slip (velocity error) is a more critical metric than positional gaze error. We agree that tracking inaccuracies can introduce translational noise into the flow field. The 3° threshold was established based on the eye tracker’s specifications and the naturalistic setup (1-meter viewing distance without head stabilization). Across all participants, the mean positional error ranged from 1.016° to 1.5° (1 deg is 2.08 cm in our setup). We also calculated retinal slip values, which ranged from 0.12 to 0.27 deg/s (X dimension) and 0.12 to 0.23 deg/s (Y dimension). These values are comparable to natural oculomotor drift (Kowler et al., 1979) and are understandably small given the low velocity of the fixation target. We have added this information about retinal sleep at the beginning of the results section.

      Consequently, it is highly unlikely that retinal slip influenced the results. Furthermore, assuming that tracking error remained consistent across fixation conditions, any present retinal slip cannot explain why the bias followed the retinal curl manipulation as predicted by the controller. We therefore consider retinal slip to be an unlikely confounding factor.

      (5) I found the sections of text comparing the separate and joined fits (starting line 287) to be a bit too rosy. The authors show the separate fits in the main text, and it is not very surprising that these fits are good, given that the model has 30 parameters, and these data are pretty low-dimensional. The authors only show the joined fits in the supplement, and they say that they are almost as good as the separate fits (indeed, they are better in a model comparison sense, but this is 30 parameters vs. 2 parameters). However, when I look at the fits of the joined model in the supplement, I don't find them to be very impressive. In particular, the model grossly misses the data for the straight paths for several subjects (e.g., id5, id6, id8, id10). And fitting the straight paths would presumably be easiest. This implies that the joined model is really missing something and that fitting the curved paths interacts strongly with fitting the data for different fixation target locations on the straight path. I think that the authors should discuss the results a bit more soberly and tone down their conclusions here.

      We thank the reviewer for the opportunity to clarify the logic behind our modeling choices. We acknowledge that the “separate fits” are inherently less informative due to the high number of free parameters relative to the data. Our primary scientific goal was not to achieve perfect descriptive accuracy via 30 parameters, but to test a specific functional hypothesis through the “joint fit.”

      The Logic of the Joint Fit:

      We agree with the reviewer that the joint fit misses some paths in some conditions. Of course, the joint fit reflects a significant compromise. The “Gain” (the weighting of the curl signal) is likely not a static constant but is dynamically tuned based on task demands, confidence in the visual signal, simulated speed, and so on. By using a single Gain parameter, we intentionally ignore this contextual variability to see how much of the behavior can be explained by a “minimalist” controller. In this sense, the 2-parameter joint model is a deliberate attempt to test this limit. By forcing a single Gain parameter to account for all conditions across both straight and curved paths within one flow manipulation (e.g. unaltered flow) we are asking if a single, fixed linear relationship between retinal curl and steering effort/gain can explain the results. We view the joint fit not as a “perfect” model, but as a stronger test of the curl-based control theory. The fact that a 2-parameter model can capture the direction and scale of biases across such a diverse set of conditions (straight/curved paths, five fixation eccentricities) suggests that retinal curl is a robust signal. Upon closer analysis, these discrepancies between the joint model and the data are most pronounced in the over-cancelled condition which is the one when sensory evidence becomes more ecologically inconsistent with the extra-retinal information (gaze direction). While the joint fit successfully demonstrates that a single parameter can capture the general functional role of curl, it fails to account for the complex sensory re-weighting that occurs in ecologically inconsistent conditions (like ‘over-cancelled’ flow). We have updated the manuscript to discuss these limitations in the “fitting the controller” section, framing the model as a parsimonious first-order approximation rather than a complete description of human heading perception based on a minimal set of parameters.

      (6) The section of the paper on neural simulations (starting line 387) has a few weaknesses. First, why are only straight paths simulated here? This does not seem to provide a very rigorous test of the model. Second, it is awkward that the simulation results are presented in units of pixels, rather than degrees. Third, the authors seem to downplay the fact that the neural estimates of heading seem to oscillate rather wildly (over a range of hundreds of pixels, whatever that means, see especially Figure S16). It was far from clear to me how an estimate of heading with these large oscillations is useful. It would seem to require that heading estimates are integrated over substantial lengths of time to be reliable. It was therefore unclear how the model produces such smooth paths from these oscillating estimates.

      We acknowledge that the presentation of the neural model requires more clarity regarding its objectives and its relationship to the behavioral data.

      We first wish to clarify the intended scope of the neural ring-attractor model. Our primary goal was not to provide a comprehensive account of behavioral performance across all conditions (which is the role of the controller model), but rather to demonstrate a biologically plausible mechanism that explains the emergence of the “Opposite-to-Gaze” bias. While the controller demonstrates that the bias follows a specific control law, the neural model shows how such a law can emerge from known primate neurophysiology, specifically, spiral-tuned MSTd neurons, gaze-contingent inhibition, and an egocentric “straight-ahead” prior.

      Why Straight Paths are Sufficient for this Objective. The reviewer asks why only straight paths were simulated. In our study, the straight-path condition with eccentric gaze is the purest test of the bias mechanism. Simulating the straight paths allowed us to isolate the interaction between foveal inhibition and the straight-ahead prior without the confounding variable of path-curvature flow. Given the complexity of the neural network’s parameter space, we focused on these conditions to provide a clear neuro-plausible explanation. We have added text when introducing the model (Neural simulations in the Results section) to make clear why we model straight paths only.

      Units: Pixels vs. Degrees. We acknowledge that the use of “pixels” in the plots of internal neural dynamics may appear awkward. The neural network operates on input stimuli that are defined by the pixel resolution of the videos used in the simulations, we used pixels as the native coordinate system to describe the movement of activity peaks within the network’s internal “map.” We have decided to keep the pixel units in these figures.

      Behavioral Output (Meters): Importantly, the final heading estimates produced by the network are not left in pixels. We use a pinhole camera model to reconstruct the 3D trajectories from the neural activity. These results are expressed in meters, allowing for a direct comparison with the human behavioral data.

      Addressing Wild Oscillations and Smooth Paths. The oscillations observed in the instantaneous heading estimates reflect the stochastic nature of the population peak when tracking high-frequency sensory inputs. In our model, the synaptic time constant (τ) was kept relatively small to ensure a fast, low-latency response to changes in self-motion. While increasing τ would have produced smoother internal dynamics, it would also have introduced delays into the control loop. Instead, we chose to maintain this high sensory responsiveness and applied a temporal moving average later to the network’s decoding to reconstruct the 3D trajectories. This is explicitly stated in the section “Heading Estimation and 3D path reconstruction” in the new appendix 2.

      In addition, the neural activity over time is shown in two ways: the heatmap shows the neuron with preferred heading (one can see more oscillations, specially when the fixation point is closer to the centre (eccentricities -2 and 2), due to larger competition between the sensory evidence and the straight-ahead prior. The other way is the decoded heading. In the ring-attractor model, the decoded heading (φ̂) is not determined by a single neuron but is calculated using a population vector average (equation 19). By summing across the entire population, the decoder effectively integrates sensory evidence from many neurons simultaneously. One can appreciate (see e.g. Fig. 5B) that averaged decoding, leads to a smoother resulting estimate (the white dashed line, whose visibility had been improved in the revised version). Behavioral work by Burr and Santoro (2001) suggests that global motion signals (divergence and rotation in optic flow) are integrated over much longer timescales—roughly 1000ms to 3000ms—compared to local motion units (~200 ms).

      In the previous manuscript, we discussed this aspect in lines 424-426. In the new version, we have added text in the Heading estimation and 3D path reconstruction section (now in appendix 2) stating that we smoothed the decoded signal in agreement with this psychophysical evidence before applying the camera model.

      See also our comment on temporal integration in the responses to reviewer #3

      Reviewer #2 (Recommendations for the authors):

      (7) Line 51: "...a functional role of rotational flow components has been largely neglected in both theoretical and experimental work on heading perception." I feel like this statement is too strong and that the authors try too hard to "sell" their findings by underrepresenting previous work. There are numerous studies (many not cited), both behavioral and electrophysiological, that have examined how heading perception depends on pursuit eye movements, either physical movements or visually simulated ones. These studies directly involve rotational flow components, and several of them have concluded that rotational flow components contribute to estimating heading in the presence of eye movements (just one example is Grigo and Lappe 1999). Because these studies generally involved horizontal pursuit of a target on the horizon, rather than tracking a point in the ground plane (like the authors' work), these studies generally did not involve flow fields with retinal curl around the fixation point. But I consider these older studies just a special case of the more general geometry, and they still involve rotational flow components. Moreover, various previous studies have used stimuli for which there was no FOE present in the visible display (either due to simulated rotation or masking out the FOE), and the authors do not seem to give credit to these works either. In addition, several studies have implicated a role of extraretinal signals in perceiving heading during eye movements, so retinal curl cannot explain everything. Rather than overemphasizing the limitations of previous work, the authors would be better served to explain how their findings extend and generalize from these previous studies.

      We thank the reviewer for pointing out this oversight; it was not our intention to overlook previous work. While our original version cited studies considering rotation-related cue, we have now substantially revised the introduction to include previous work and better acknowledge the informative role of rotation. Our central aim remains to distinguish between models that compensate for rotation to recover a heading vector and our proposal that the visual system exploits retinal curl directly as a primary, functional signal for locomotor control.

      We have now updated the Introduction and Discussion to better situate our work within the context of studies (including Grigo & Lappe, 1999 and some additional ones we have included in the new version) that have investigated the informative role of rotational flow. We now clarify that our study extends these findings by investigating the non-uniform rotational patterns (curl) that emerge during ground-plane fixation, representing a more general and biologically ubiquitous case of locomotor control, while acknowledging previous studies that have also considered the potential role of curl generated by gaze fixation.

      (8) Figure 3: I did not understand why there are two purple and two blue curves in the graphs of the middle column. And the caption does not explain this.

      This a very good observation. This was explained in lines 268-273 (previous version). When the gaze is straight-ahead (same direction as heading), there is no curl. However, we introduced positive or negative curl in the altered conditions. These purple and blue lines refer to these trials and show that the bias re-appears in the expected direction when curl is (unexpectedly) added. We think this adds additional evidence to the curl contributing to heading. Even though this was extensively explained we have added text in the caption of figure 3.

      (9) Line 331: What makes the authors think that retinal curl is computed in area MSTd? They should cite studies to support this idea if it has been shown in physiology.

      Evidence was cited in the introduction (Graziano et al 1994) of sensitivity to spiral motion in addition to neuro-computational models that implement this activity also cited (e.g. work of Leyton et al.)

      (10) The neural network model for computing heading from curl requires a "gaze-centered inhibitory drive" that inhibits activity around where the eyes are looking. This is probably a biologically plausible thing, but is there any evidence to support the idea that this signal exists in the parts of the brain where the authors believe these computations to be happening? They simply posit the existence of this gaze-centered inhibition as though it is common knowledge, but they provide no citations nor discuss any previous evidence for its existence.

      While neurophysiological evidence primarily describes this as gain-field modulation, this process frequently involves localized suppression of neural activity to facilitate coordinate transformations. In parietal areas such as LIP and 7a, eye-position signals do not just enhance responses but can also suppress them, effectively shifting the 'center of gravity' of a population response (Read et al 1997; Born et al 2005, cited in the discussion in the revised section re-evaluating the Focus of Expansion). In the context of our ring-attractor model, this functional modulation is most parsimoniously implemented as a gaze-centered inhibitory drive.

      (11) Lines 482-483: Why should perceptual biases related to retinal curl take seconds to show up?? The curl information itself must be present very quickly, perhaps requiring just a few video frames. So what does this imply about mechanisms? The authors throw out this assertion, but it is left hanging without any further analysis or support.

      The time course reflects the integration requirements of complex motion processing. While local flow is processed rapidly, global patterns like retinal curl require longer temporal windows to reach a stable estimate (Burr et al 2001). In our study, this integration is functionally necessary to filter the higher-frequency 'wobble' induced by gait-cycle oscillations. We now discuss the temporal integration aspects under a new heading in the discussion.

      (12) Line 508: "This suggests that the "bias" observed in our perceived headings may reflect the operation of a control law optimized for action rather than a failure of a perceptual system designed for passive estimation." The authors make this statement to justify why perceptual biases are present with unaltered curl. But I don't fully understand the logic. Are they saying that it is not possible to have a set of computations that can do both things accurately? Is it possible to show this theoretically? Moreover, if it is not possible to rule out other possible sources of the biases, such as those described above (reference frame of judgments, eye movements, etc), then is it necessary to invoke this logic?

      Our logic is that the observed 'bias' is not a representational failure, but a functional byproduct of a control law optimized for active steering. In a closed-loop system, the objective is to null the error signal (retinal curl) to maintain a stable path. When observers are asked to make an open-loop heading report, they likely utilize this same control signal, which manifests as a systematic bias toward the 'null' point of the controller as a result of a sustained fixation in discrepancy with the simulated translation/heading.

      We do not suggest that accurate perception and control are theoretically incompatible; rather, we suggest that perception and action rely in the same underlying information (e.g. work of Brenner & Smeets). While other factors, such as coordinate transformations between reference frames, certainly can contribute to the reporting process, our interpretation provides a parsimonious link between the psychophysical data and the underlying steering mechanism. By framing the bias as a consequence of a 'nulling' strategy, we explain not just the existence of the error, but its specific direction and magnitude relative to the fixated target.

      (13) Line 546: "...MSTd would simultaneously code curvature for trajectory estimation and heading across the neural population, with curvature encoded through the spirality of the most active cell and heading through the visuotopic location of its receptive field center." The latter part of this argument seems to imply a relationship between the heading preferences of MSTd neurons and the locations of their receptive fields. I am not aware of any evidence for such a relationship, so the authors should indicate whether this is based on some experimental data or just a speculation.

      We thank the reviewer for this observation. The proposal that heading is signaled by the visuotopic location of active MSTd populations is a core architectural feature of our model and is supported by several lines of evidence.In the Layton and Browning (2014) framework, MSTd is modeled as a visuotopic map of functional 'hypercolumns'. Each hypercolumn contains neurons tuned to a continuum of spiral patterns, but all neurons in a given hypercolumn share a receptive field center at a specific location in visual space. Consequently, the visuotopic location ($x, y$ coordinates) of the maximally active hypercolumn represents the center of motion (heading), while the spirality (the tuning dimension within that hypercolumn) represents path curvature. We have clarified this in the revised discussion (re-evaluating the FoE) to emphasize that this dual-coding scheme arises from the simultaneous representation of 'where' (population map location) and 'what' (spiral tuning) in MSTd.

      Reviewer #3 (Public review):

      The primary limitation of the paper is that it avoids discussion of some of the inevitable complexities of heading perception. The main issue is what exactly is meant by heading. Different behaviors evolve over different timescales. The geometry of retinal motion defines instantaneous heading, which varies widely through the gait cycle. Time-varying information like this is known to be important in the momentary control of balance. Heading can also be thought of as steering the body toward a distant goal, which evolves over longer timescales. The current manuscript appears to be concerned with heading information integrated over a few seconds and seems to provide evidence that heading is indeed integrated over the gait cycle. The issue of the time scale of the computation is touched on, but it is not related to how it might be used in normal walking or what situations it might apply to. Steering toward a distant goal during walking is not a very difficult problem and may not require evaluation of retinal motion, but control of balance is more challenging and may depend critically on curl. Consequently, the timescale of the computation needs to be considered in order to understand what is meant by heading.

      We thank Reviewer #3 the comments regarding the definition of heading at different time scales, the role of the gait cycle, and the temporal integration of the curl signal. These comments have helped us refine the manuscript’s core arguments.

      We agree that “heading” must be precisely defined within the context of the differing temporal demands of balance and steering. While instantaneous heading provides the high-frequency feedback necessary for momentary postural adjustments and balance, our study is concerned with heading as a gaze-relative signal used for the continuous control of a locomotor trajectory. As such, we have revised the manuscript to specify that the perceived heading measured in our task reflects a signal integrated over the gait cycle to filter out the oscillatory noise induced by head bob and sway (mainly in the Discussion section).

      The reviewer correctly notes that gait-induced head bob and sway produce high-frequency oscillations in the curl signal, yet our behavioral results show smooth, slowly evolving biases. The visual system does not react to “instantaneous” curl, which would lead to jittery, unstable heading estimates. Instead, it integrates flow over a timescale roughly commensurate with a full gait cycle (~500–1000ms). This implies a significant temporal integration process. This temporal integration is consistent with evidence (Burr and Santoro,2001, Vis Res) indicating that optic flow signals (radial and rotational components) are integrated over windows of approximately up to 3 seconds to ensure perceptual stability. Neurally, this likely involves the projection from area MSTd to the Ventral Intraparietal area (VIP), a pathway where fast, eye-centered sensory inputs are transformed into stable, body-centered representations suitable for guiding long-term steering behavior (Chen et al. 2011, JNeurosci.). By grounding our definition of heading in these specific temporal and neural constraints, we tried to clarify how the visual system exploits retinal curl for goal-directed action in natural, dynamic environments and relate our findings to recent studies addressing the role of retinal motion on balance (Powell et al. 2026 Bioarx).

      In our implementation, we explicitly address the high-frequency noise introduced by gait dynamics by smoothing the retinal curl signals computed from the stimulus videos before they are fed into the controller. This temporal filtering allows the fit of the controller’s prediction to the response data while remaining robust to the rapid fluctuations of head bob and sway. In contrast, the neural ring-attractor model would not require an external smoothing step; instead, the integration is an emergent property of the system’s architecture that can be controlled with different parameters, as commented above in a response to Reviewer #2. The dynamics of the synaptic weights and the characteristic “leak” in the population activity naturally implement a leaky integration of sensory evidence, ensuring that the decoded heading reflects a sustained estimate rather than an instantaneous response to visual noise.

      We also agree that we avoided discussing some complexities of the heading perception. In the new version, we also include and integrate the distinction between instant heading and future path in different parts of the ms (introduction) and mainly discussion (temporal integration and steering section) which have been revised substantially.

      Reviewer #3 (Recommendations for the authors):

      There are a number of points that require clarification.

      (1) Head bob and sway were included in the stimulus and need to be addressed in both the analysis of the data and the interpretation. The curl signal in the stimulus varied over time, commensurate with normal gait. However, the results don't reflect the same level of variability that would be produced from curl over the gait cycle. This means that the information must be integrated over some longer timescale. It is not clear from the data analysis what this integration is. Is there an implicit integration with the manipulation of the steering wheel? If subjects indeed appear to be able to use curl to evaluate heading over timescales of seconds, this needs to be explicitly addressed, as it is a novel result. This would require parts of the discussion to be changed/expanded to maintain consistency. For example, line 482 talks about the buildup of biases over time.

      As commented above in the public response, the curl estimated from the optic flow algorithm was smoothed before being input into the controller (path fitting and predictions). The smoothing was only applied to the curl signal, not to the participants responses. Also, as mentioned before, the time course of the bias is consistent with integration times of optic flow reported in the literature. All these aspects are now explicitly included in the new display and conditions section (Flow manipulation conditions).

      (2) Since the experiment included curl variability resulting from the gait cycle, some discussion is needed about the role of retinal motion in the control of balance and posture. There is a large literature about the role of flow in controlling gait and momentary adjustments of the body while walking. Additionally, it should be noted that in the task, head bob and sway from 1 prerecorded individual was shown to all subjects. It is known that gait varies significantly between different individuals, and it should be acknowledged that this may lead to differences at the individual subject level for perceiving heading.

      We have included a point in the discussion addressing the different time scales for different use of optic flow signals (postural control vs locomotion).

      We agree with the reviewer that utilizing a single gait profile for all participants may introduce individual differences in perceived heading, as the simulated head motion might not perfectly match each participant’s unique biological gait signature. However, we prioritized stimulus consistency over idiosyncratic accuracy. By ensuring that every participant viewed the exact same motion profile, we could be certain that the systematic 'opposite-gaze' biases observed across the population were driven by our experimental manipulations of gaze and retinal curl, rather than being confounded by variability in head-motion kinematics. We have added an acknowledgement of this point at the first paragraph of the displays and conditions section in the Methods.

      (3) More information is required about the use of the rotating wheel for the measurement of heading. How easy was it to use? What about time delay, and how does this deal with the bob and sway? Does the wheel impose a de facto integration on the perceptual measurement?

      The rotary encoder provided an intuitive steering-wheel interface that participants found easy to operate. To ensure minimal latency (1–5 ms), the device was interfaced via an Arduino Uno and sampled by a dedicated background Python thread, isolated from the visual rendering loop. We have incorporated these technical details into the Methods (Procedure) section.

      (4) Restructuring the description of the models It remains unclear why the dynamics of the neural network are a necessary inclusion in this paper. It seems interesting, but there is no comparison to actual neural data or other related work. Instead, this appears to be a description of what the network is doing, which is already defined by the equations. This needs to be clarified for its exact interpretation with respect to real neural data, and its importance here for understanding the biases that emerge in heading judgments. The paper would flow better if this section were included as supplementary material or omitted from the paper entirely, as it seems to detract from the other points. If this is a description of why the biases are seen in the controller, then the supplementary material is a good place for it.

      We thank the reviewer for this suggestion. We clarify that the neural model is not intended to simulate specific empirical neural data, but rather to provide a biologically plausible implementation of the controller. This allows us to demonstrate how the observed biases emerge from the dynamics of standard cortical architectures, such as ring attractors. This is now mentioned when introducing the neural simulation results.

      Following the reviewer's suggestion, we have moved the neural model equations to Appendix 2 while retaining the simulation results in the main text (Results). We believe it is essential to present not just the abstract controller, but also its functional implementation, as this provides a mechanistic bridge between retinal signals and locomotor behavior.

      (5) For the modeling approaches, the math would be more appropriate for supplementary materials.

      To ensure a better flow of the paper, we have moved the neural model equations to Appendix 2, while Appendix 1 now details the relationship between measured curl and other variables (speed, yaw, pitch, etc.). We have retained the controller model in the Methods section, consistent with the eLife layout where Methods follows the Discussion.

      (6) How do the models and data analysis deal with the influence of the gait cycle in the input? Do they integrate the information over that timescale? If so, the integration time needs to be specified.

      As commented above, to mitigate gait-cycle fluctuations, we smoothed the computed curl signal before inputting it into the controller and applied a similar smoothing process to the neural model’s readout. Using a LOESS filter, the effective integration window was 2.4 seconds. Close to the integration time reported in Burr et al. 2001. These parameters have now been explicitly specified in the Methods section. For the empirical data, we just utilized trial-averaging.

      (7) What does biologically plausible mean in terms of the neural network model? Especially when control wasn't explicitly a variable in the measured behavior of the subjects.

      By biologically plausible, we mean that our model is constrained by neural architectures documented in the primate brain—specifically ring-attractor dynamics, population coding, and gaze-centered gain-fields. Crucially, the network utilizes recurrent connectivity with a 'Mexican-hat' profile (local excitation combined with lateral inhibition). This is a standard and widely accepted motif in computational neuroscience, representing the consensus on how cortical circuits maintain a stable "bump" of activity to represent spatial variables. Rather than introducing ad-hoc mechanisms, we demonstrate that the observed behavioral biases emerge naturally from these established neural components when they are tasked with maintaining locomotor stability. Since we think this aspect was already emphasized, we haven’t added any additional detail.

      (8) There should be more extensive acknowledgement of the body of literature that has challenged the use of the focus of expansion. That section should also include references to work that has investigated extraretinal signals, as they may also be important.

      We have expanded the Introduction and Discussion to more thoroughly acknowledge research challenging FOE-based models and the critical role of extraretinal signals. These updates, which also align with our response to Reviewer #2, provide a more comprehensive context for our model within the existing body of heading and self-motion literature.

      Minor points:

      (1) Line 69 - In self-generated motion, spiral patterns are almost always centered on the fovea, but many physiological experiments present spirals in the peripheral retina. This is incompatible with the motion generated during self-motion. Therefore, clarify whether the type of spiral motion Graziano investigated was centered on the fovea.

      In the experiments conducted by Graziano et al. (1994), spiral stimuli were centered on the receptive field (RF) of the individual neuron being recorded to accurately characterize its tuning. While the reviewer correctly notes that spiral centers often align with the fovea during active steering (due to fixation on a goal), MSTd neurons possess large RFs that provide a comprehensive 'template' system across the visual field. This population-level representation allows the brain to recover trajectory information even when the focus of motion shifts relative to the fovea—for example, during pursuit eye movements or when fixating on landmarks off the direct path of travel. To keep this part of the text brief, we haven’t add more details concerning this study.

      (2) Line 72 - "Magnitude" instead of "amount".

      This has been changed.

      (3) Line 115 - State explicitly whether the scale of the visual stimulus was matched to the scale of the actual natural images shown in VR.

      We have updated the Methods (Displays and conditions) section to explicitly state that the visual scale was veridical. The virtual camera’s parameters were calibrated such that its field of view (91°) matched the physical dimensions of the projection screen (2.03 m × 1.16 m) at the 1.0 m viewing distance. This ensures that the angular size of the objects and motion gradients in the stimulus were 1:1 with the scale of the simulated natural environment.

      (4) Line 144 - While Farneback is a good dense flow estimation algorithm, it is noisy and may impose biases/variability in the calculation of curl. This should be acknowledged.

      We acknowledge that the Farneback algorithm can introduce variability in curl estimation. To mitigate this, we utilized 10 independent renderings of each experimental trial to compute the flow fields. Although this methodology was reflected in the data previously uploaded to our OSF repository, we have now explicitly added this detail to the manuscript (Flow (curl) manipulation conditions). The computed curl used for the modeling was derived from the aggregate of these different runs, ensuring a robust and stable signal that accounts for potential algorithmic noise.

      (5) Line 206 - "Join fits". Is this a technical term? It sounds awkward. Would "Joint fits" make more sense?

      The referee is right. We have corrected this.

      (6) Line 280 - "Consistent with" (typo).

      This has been corrected.

      (7) Line 286 - Typo in title.

      Also corrected to Fitting the controller.

      (8) Lines 482-493 - There should be more discussion on the time course of integrating the stimulus, and the relationship/generalizability to more natural stimuli.

      This part of the discussion (related to integration time) has been changed considerably to include discussion of postural control in addition to locomotion.

      (9) Lines 516-517 - It is mentioned that retinal flow dynamics override the visual direction cue. This may not be generally true, as the reweighting of the cues might depend on things like task demands or actual stimulus context. In the present experiment, the subject only has access to a large moving textured ground plane, and the body is stationary.

      We agree with the reviewer that cue reweighting is highly context-dependent. However, as noted in the original manuscript (Lines 516-517), we specifically stated that retinal flow dynamics 'can' override the visual direction cue, rather than asserting a universal rule. This phrasing was intentional to acknowledge that while flow is a potent signal—especially in the presence of a large, textured ground plane as used in our paradigm—the relative weighting of these cues remains contingent on the specific sensory and task conditions. We believe this remains a fair and cautious interpretation of our findings.

      (10) Lines 524-256 - It is unclear why Matthis et. al. 2022 is cited for this point. Some of the steering literature, like Wilkie Wann & Allison 2006 or Lappi & Mole 2018 (and some of their other work), would be more relevant for the definition of a control law under these circumstances.

      We agree and this part has been changed substantially.

      (11) Lines 528-532 - Warren et. al. 2001 should be cited in this section because their results were interpreted in terms of focus of expansion, but may result from the curl signal (Powell et. al., 2026. The Role of Retinal Flow in Walking. bioRxiv, 2026-02.).

      We agree with this suggestion and the citation has been added.

      (12) Lines 544-546 - Layton and Browning are focused more on path perception from the implemented spiral tuned cells. Because of this, it wouldn't be appropriate to say it is a shift away from FOE-based heading based on this citation alone. There are more models/psychophysical results that would strengthen this claim.

      We agree with the reviewer that Layton and Browning focus specifically on path perception. Our original intention in citing this work was to emphasize the neurophysiological continuum from radial to circular motion (spiral tuning) rather than discrete expansion/rotation channels. However, we have revised this section to clarify that the reliance on retinal curl represents a mechanism for determining the future path (locomotor trajectory) rather than merely instantaneous heading. This distinction acknowledges that while heading is a momentary vector, the integration of curl signals allows the system to anticipate and control the intended path over time—a framing that better aligns with both the cited literature and our proposed controller model. As commented above this is now extensively discussed in the revised version.

      (13) Figures - There are minor visibility issues for some of the figures. In Figure 2, the thick line is unreadable, and in Figure 1, the x-axis labels are crowded.

      The axis in Fig 1 has been modified to avoid crowdedness. In Fig. 2, the thick (average line) has been modified. We hope they are more visible now.

    1. eLife Assessment

      This manuscript describes a series of studies using four different Go/No Go task variants in combination with fast-scan cyclic voltammetry to determine the role of dopamine release in the ventromedial striatum in action selection, controllability of reward pursuit, effort, and reward approach. The authors conclude that dopamine signals in the ventromedial striatum integrate the invigoration of action initiation with continuous estimation of spatial, but not temporal, proximity to rewards. There is solid support for the conclusions, and the findings are valuable, with theoretical implications for the role of dopamine in reward-driven behavior.

    2. Reviewer #1 (Public review):

      Summary:

      Poh and colleagues investigate dopamine signaling in the nucleus accumbens (ventromedial striatum) in rats engaged in several forms of go/no-go tasks, that differed in reward controllability (self-initiated reward seeking or cue-evoked/quasi-pavlovian), and in the specific timing of the action-reward contingencies. They analysis dopamine recordings made with fast scan cyclic voltammetry and find that dopamine signals vary most consistently to cues that signal a required action (go cues) vs cue signaling action withholding (no go cues). Through various analysis they report that dopamine signals align most clearly with action initiation and with the approach to the reward-delivery location. Collectively these data support aspects of a variety of frameworks related to accumbens dopamine signaling in movement, action vigor, approach, etc.

      Strengths:

      These studies use several task variants that consolidate a few different components of dopamine signal functions and allow for a broad comparison of many psychological and behavioral aspects. The behavioral analysis is detailed. These results touch on many previous findings, larger showing consistent results with past studies.

      Weaknesses:

      The paper is dense and could benefit from some revision to increase clarity of the figures, the methods and analysis. The inclusion of many tasks is a strength but also somewhat overshadows specific points in the data, which could be improved with some revision to focus. There is a lack of strong connection between some of the findings, which if revised would help to emphasize the impact of the work.

    3. Reviewer #2 (Public review):

      Here, the authors record dopamine release using fast-scan cyclic voltammetry in the nucleus accumbens/ ventromedial striatum (VMS) while rats perform variants of a go/no-go task. Two versions are self-paced, in that the rat can initiate a trial by nosepoking at the odor port at any time once the ITI had elapsed, whereas the other two require the rat to wait for a cue-light before responding. Two "long" variants also require either more lever-presses on go trials, or a longer nosepoke time for no-go trials, and also incorporate "free" trials in which the rat is rewarded for just heading straight to the food tray. The authors find that dopamine levels increase more during the response requirement for go than no-go trials, indicating a role for invigorating to-be-rewarded actions. Dopamine levels also steadily increased as rats approached the site of reward delivery, and the authors demonstrate quite elegantly that this was not due to orientation to the food tray, or time-to-reward, or action initiation, but instead reflects spatial proximity to the rewarded location. Contrary to previous reports, the authors did not discern any differences in dopamine dynamics depending on whether the trials were cue- or self-paced, and dopamine release did not scale with effort requirements.<br /> The manuscript is well-written and the authors use figures to great effect to explain what could otherwise be a hard-to-parse set of data. The authors make good use of the richness of their behavioral data to justify or negate potential conclusions.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Poh and colleagues investigate dopamine signaling in the nucleus accumbens (ventromedial striatum) in rats engaged in several forms of Go/No Go tasks, which differed in reward controllability (self-initiated reward seeking or cue-evoked/quasi-pavlovian), and in the specific timing of the action-reward contingencies. They analyze dopamine recordings made with fast scan cyclic voltammetry, and find that dopamine signals vary most consistently to cues that signal a required action (Go cues) vs cues signaling action withholding (No Go cues). Through various analyses, they report that dopamine signals align most clearly with action initiation and with the approach to the reward-delivery location. Collectively, these data support aspects of a variety of frameworks related to accumbens dopamine signaling in movement, action vigor, approach, etc.

      Strengths:

      These studies use several task variants that consolidate a few different components of dopamine signal functions and allow for a broad comparison of many psychological and behavioral aspects. The behavioral analysis is detailed. These results touch on many previous findings, largely showing consistent results with past studies.

      Weaknesses:

      The paper could heavily benefit from some revision to 1) increase clarity of the figures, the methods, and the analysis. 2) The inclusion of many tasks is a strength, but also somewhat overshadows specific points in the data, which could be improved with some revision/reworking. 3) Some conclusions are not fully justified. As shown, support for the conclusion "dopamine reflects action initiation but not controllability or effort" is lacking without more analyses and additional context. 4) Further, the notion that the dopamine signals reported here reflect spatial information could be justified more strongly.

      We thank the reviewer for their detailed evaluation and constructive feedback. We have made substantial revisions to address each concern raised:

      (1) Clarity of figures, methods and analyses

      We have revised the organization of the panels in Figure 1 for clarity.

      We have revised Figure 2 and its caption: we have labelled all comparisons depicted in the figure, and now included a line, “Subject-wise comparisons of dopamine data were made for all alignments”, to clearly show that all statistical tests shown in Figure 2c-e were performed between subjects.

      In the caption of Figure 3, we have now added, “... and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” to clearly show that statistical tests were performed on the trial-wise level for Figure 3c.

      We have now added a table to the Methods section (Table 1), detailing the sample size in each task variant and the number of trials within each No-go classification.

      We have added more information in the Methods section: we now include Videos to show the classified No-go behaviors and other trial types (Go and Free); and provide a schematic of the DLC workflow in the supplementary materials (Supplementary Figure S9).

      (2) Strengthening specific points in the data

      To improve clarity of our main findings, we have revised the layout of the Results section such as including more descriptive headers:

      “Behavioral performance in Go/No-go (“short task”) was unaffected by controllability of reward pursuit”

      “Behavioral performance in Go/No-go/Free (“long task”) was unaffected by controllability of reward pursuit”

      “Motivation to approach the reward magazine was similar between Go and No-go trials”

      “VMS dopamine release encodes reward-related action initiation”

      “VMS dopamine release does not only reflect reward-related action initiation”

      “Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards”

      “Motivational state reflected by No-go behavioral strategy correlates with dopamine signal size during reward approach”

      (3) (4) Conclusions drawn from data

      We have carefully revised our conclusions to accurately reflect our experimental design and analyses, emphasising our core finding that VMS dopamine was consistently increased in Go versus No-go trials, throughout manipulation of the type of trial start (self- and cue-initiated) and effort manipulation (short and long task variants).

      We thank the reviewer for the feedback regarding our VMS dopamine signals reflecting spatial proximity to reward. We performed additional analyses, which we present in Supplementary Figure S10, S11 and S12, to support our interpretation that VMS dopamine encodes spatial proximity to reward.

      We appreciate the reviewer comment relating to the statement, “Dopamine reflects action initiation but not controllability or effort". We have revised the wording of our conclusion to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). This subheading is now changed in the Discussion to “Reward-related dopamine depends on action initiation irrespective of controllability and effort”. The additional analyses that we have performed are shown in Author response image 1entitled: “Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long).

      Additional details on subjects used in each study, analysis details on trialwise vs subjects-wise data, and other context would be helpful for improving the paper.

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures in the original submission of this manuscript.

      To improve the paper, we now include a table in Methods Section 6: Statistical Analysis (Table 1), detailing the sample size in each task variant (with FSCV recordings) and the number of trials within each No-go classification.

      To give more context, we made Author response image 1 to illustrate the number of subjects in each task variant (and their overlap):

      Author response image 1.

      Number of subjects included in each Go/No-go task variant (total n = 21). Values (n) depict overlap of each subject between task variants. One animal was recorded in self-initiated Go/No-go and Cue-initiated Go/No-go/Free (dotted line with arrowheads).

      Reviewer #2 (Public review):

      Here, the authors record dopamine release using fast-scan cyclic voltammetry in the nucleus accumbens/ ventromedial striatum (VMS) while rats perform variants of a Go/No Go task. Two versions are self-paced, in that the rat can initiate a trial by nosepoking at the odor port at any time once the ITI has elapsed, whereas the other two require the rat to wait for a cue-light before responding. Two "long" variants also require either more lever-presses on Go trials, or a longer nosepoke time for No Go trials, and also incorporate "free" trials in which the rat is rewarded for just heading straight to the food tray. The authors find that dopamine levels increase more during the response requirement for Go than No Go trials, indicating a role for invigorating to-be-rewarded actions. Dopamine levels also steadily increased as rats approached the site of reward delivery, and the authors demonstrate quite elegantly that this was not due to orientation to the food tray, or time-to-reward, or action initiation, but instead reflects spatial proximity to the rewarded location. Contrary to previous reports, the authors did not discern any differences in dopamine dynamics depending on whether the trials were cue- or self-paced, and dopamine release did not scale with effort requirements.

      The manuscript is well-written, and the authors use figures to great effect to explain what could otherwise be a hard-to-parse set of data. The authors make good use of the richness of their behavioral data to justify or negate potential conclusions. I have the following comments.

      Re: The lack of relationship between effort to acquire reward in the current study and the magnitude of dopamine release, 1) can the authors unpack this a bit more? 2) Why the difference between the Walton and Bouret studies? Were the shifts in effort requirements comparable across the behavioral tasks? 3) What else could be different between the methodologies?

      We thank the reviewer for the feedback and have responded to each of the three questions below (see points 1-3).

      Firstly, we tried to improve the clarity of our research aims. Our primary comparison throughout the manuscript is between Go versus No-go within each task variant. We ask whether the Go/No-go difference in dopamine signaling persists across different response demands. Thus, testing effort was not central to this main question, but rather a feature of the task that did not affect the Go-No-go dopamine difference.

      (1) Consistent with this aim, we show that VMS dopamine differs between Go and No-go actions persistently across all task variants despite differences in response requirements (action was always accompanied by greater dopamine release compared to action suppression). Our behavioral-training data suggest that the ability to perform short and long tasks differed: rats were first trained to criterion on either a short (∼2 s) or long (∼3 s) Go/No-go variant, with the longer variant requiring substantially more training sessions (short: 18.5 ± 7.6 sessions vs long: 41.1 ± 7.9 sessions; see Author response image 2), indicating behavioral demands were higher for the long-task.

      Author response image 2.

      (2) Regarding the apparent discrepancy with Walton and Bouret (2019), we acknowledge that our original description was imprecise (We wrote: “Previous studies have shown that dopamine signals are influenced by the effort required to obtain rewards”). Our intent was not to suggest a direct contradiction, but rather to emphasize that our findings are consistent with the paper’s broader conclusion that effort encoding by dopamine is limited and highly context-dependent. We have now adjusted the manuscript to better reflect our intent by changing the sentence to, “Previous studies have shown that dopamine signals may be influenced by the effort required to obtain reward but only for particular task conditions (Cousins et al., 1996; Gan et al., 2010; see for reviews, Salamone and Correa, 2024; Walton and Bouret, 2019).”

      For added clarity, these were the main results highlighted in the Walton and Bouret review: Gan et al. (2010) demonstrated that VMS dopamine sensitivity to low-effort costs is prominent early in training (≤ 2 training sessions) and diminishes after extended experience (> 9 sessions). Similarly, Hollon et al. (2014) reported that cue-evoked VMS dopamine primarily tracks reward magnitude with minimal modulation by effort. In line with this literature, our rats were highly trained (≥ 9 sessions until the first recording), and exhibited no difference of average Go minus No-go dopamine between short and long task variants within controllability type (see Author response image 3), supporting the idea that extended training exhibits minimal effort-related modulation of VMS dopamine.

      Author response image 3.

      Average Go minus No-go dopamine reveals no effect of controllability (self vs. cue-initiated) or effort (short vs. long). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of controllability and effort showed moderate evidence of absence (controllability: BF<sub>10</sub> = 0.324; effort: BF<sub>10</sub> = 0.309). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.101), and the full model including a controllability × effort interaction was the least supported of all models examined (BF<sub>10</sub> = 0.043). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither controllability, effort, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (3) With respect to task comparability and methodological differences, our behavioral paradigm differs in important ways from those highlighted by Walton and Bouret, where effort was often manipulated by training animals to associate cues with different numbers of lever presses within the same session, and typically involved only action initiation. In contrast, our task required both action initiation and action suppression, and changes in response contingencies occurred across separate recording sessions rather than within-session cue-based manipulations. Although these paradigms are not directly comparable, and only had the same dopamine recording technique in common (FSCV), a key takeaway of our results is that regardless of effort differences, VMS dopamine during action initiation is consistently higher than during action suppression.

      I would argue that the cue- vs self-initiated distinction was pretty minor, given that there was a fixed ITI of 5s. How does this task modification compare to those used previously to show that dopamine release corresponds to behavioral controllability? It would help the reader if the authors could spend more time discussing these disparate findings and looking for points of methodological divergence/commonality.

      We agree that clarifying how our manipulation of controllability compares to prior work improves the manuscript, and we have made the necessary adjustments. However, we would first like to correct an incomplete characterization of the task design.

      While the short-task variant used a fixed 5 s inter-trial interval (ITI), the long-task variant employed a variable ITI ranging from 15–25 s. In the long-task variant, the timing of trial onset was less predictable, and we believe this manipulation reduced animals’ ability to precisely estimate when reward pursuit could begin. Under these conditions, whether trials were Self-initiated or Cue-initiated had a substantial impact on animals’ control over the initiation of reward pursuit. That said, we agree that the Self- versus Cue-initiated distinction overall represents a moderate manipulation of controllability compared to those used in studies that focus on controllability.

      A key source of divergence across studies lies in the definition of controllability. We defined controllability as the animals’ ability to choose the time point of beginning the reward pursuit, rather than whether an action was required, and have now added the following sentence in the:

      - Introduction section: “... controllability of reward seeking, defined as the ability to determine when to initiate reward pursuit (Self- vs Cue-initiated trials)...”;

      - Results section: “We defined controllability as the rats’ ability to choose the time point of reward pursuit. In Cue-initiated trials, the time point at which trials could be started was dictated by a cue light, whereas in Self-initiated trials rats were able to choose intrinsically (control) when to attempt a trial start.”;

      - Discussion section

      Importantly, the action requirements for Go, No-go, and Free trials were identical across these trial-start conditions. We found that the degree to which controllability was manipulated in our task was insufficient to modulate the action-specific VMS dopamine signal (Go vs No-go difference), which remained robust across conditions.

      In contrast, controllability has been defined by others as the presence versus absence of an operant action requirement for reward. For example, Goedhoop et al. (2023) directly contrasted operant (lever press required) and Pavlovian (no action required) conditions, removing action execution as a prerequisite for reward. In that context, cues signaling operant control elicited sustained VMS dopamine release, which was interpreted as reflecting anticipation or preparation for executing a learned action. Similarly, Hamid et al. (2021) demonstrated that dopamine “wave” directionality across striatal regions depends on controllability defined by operant versus Pavlovian conditioning.

      Taken together, these comparisons (results from the present study and in the literature) suggest that dopamine sensitivity to controllability may depend on how it is manipulated. We have clarified these methodological distinctions in the revised Introduction, Results and Discussion, and emphasized that more extreme manipulations (such as removing action requirements entirely or increasing uncertainty over trial timing) may be necessary to reveal controllability-dependent changes in VMS dopamine signaling. Alternatively, the apparent discrepancies across studies may primarily reflect differences in the underlying definitions of controllability rather than conflicting results.

      Reviewer #3 (Public review):

      Summary:

      The manuscript by Poh et al. investigated whether dopamine release in the ventral medial striatum integrates information about action selection, controllability of reward pursuit, effort, and reward approach. Rats were implanted with FSCV probes and trained in four Go/No Go task variants:

      (1) trials were self-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (2) trials were cue-initiated and had two trial types (Go vs. No Go) that were auditorily cued,

      (3) trials were self-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued, and effort was increased,

      (4) trials were cue-initiated and had three trial types (Go vs. No Go vs. free reward) that were auditorily cued.

      The authors report that dopamine levels rose during Go trials and slowly rose in No Go trials, but this pattern did not differ across task variants that modified effort and whether trials were cued or initiated. They also report that dopamine levels rose as rats approached the reward location and were greater in rats that bit the noseport while holding during the No Go response.

      Strengths:

      (1) Interesting task and variants within the task paradigm that would allow the authors to isolate specific behavioral metrics.

      (2) The goal of determining precisely what VMS dopamine signals do is highly significant and would be of interest to many researchers.

      Weaknesses:

      (1) This Go/No-Go procedure is different from the traditional tasks, and this leads to several problems with interpreting the results:

      (a) Go/No Go tasks typically require subjects to refrain from doing any action. In this task, a response is still required for the No Go trials (e.g., continue holding the nosepoke). The problem with this modified design is that failure to withhold a response on No Go trials could be because i) rats could not continue holding the response, as holding responses are difficult for rodents, or ii) rats could not suppress the prepotent go response. This makes interpreting the behavior and the dopamine signal in No Go trials very difficult.

      We appreciate the reviewer raising this important methodological consideration. We acknowledge that our Go/No-go task differs from traditional paradigms used in humans and primates (e.g. Raud et al. 2020, 10.1016/j.neuroimage.2020.11658; Eagle, Bari & Robbins 2008, 10.1007/s00213-008-1127-6; Roitman & Loriaux 2013, 10.1152/jn.00350.2013).

      However, our design addresses the specific constraints of studying dynamics in freely moving rodents while maintaining the core feature of Go/No-go tasks: requiring suppression of a prepotent response. Our task accomplishes the primary aim of our study, which is to compare VMS dopamine dynamics during action initiation and action suppression, and below we list the reasons why. Therefore, we do not believe that this difference compromises the validity and interpretation of our results.

      It has been suggested for decades that the two main processes governed by mesolimbic dopamine are reward learning and motivated action, and our study aimed to better understand how VMS dopamine integrates reward-related information and motivated action, rather than studying them in isolation. To do so, we trained rats in a modified Go/No-go task.

      More recent work (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019) demonstrates that VMS dopamine signaling incorporates both action and reward-related information, rather than either of the two alone. Importantly, in freely-moving rodents, examining this relationship requires preventing the approach response that occurs when reward delivery is anticipated (Pavlovian bias, go for rewards). Traditional Go/No-go designs that simply require "doing nothing" would not achieve this control in freely-moving rats, as animals immediately approach the reward magazine as soon as reward is inferred (as seen in our Free trials). Thus, we require a No-go condition, as we and others have defined (Syed et al. 2016), whereby animals have to actively suppress the ‘initiation’ response. Action initiation is defined at the beginning of the Discussion section: “... at two distinct points after trial start: 1) when rats began lever pressing (Go), and 2) when rats walked to the reward magazine, either without action requirement (Free) or after successful trial completion (Go and No-go)”.

      To further strengthen our interpretation that we compare action initiation and suppression, and to facilitate cross-species translation of our results (i.e., rodent to human), we also include Free trials, where reward delivery requires no specific action (which are essentially like “doing nothing” trials in traditional tasks). This addition allowed us to directly compare No-go and Free trials, where animals must actively suppress responding while maintaining task engagement, to a condition where no overt action is required for a reward, respectively. The dramatic difference in VMS dopamine between No-go and Free trials demonstrates that VMS dopamine reflects active action suppression during No-go trials, rather than merely the absence of action requirements. This has now been discussed.

      Finally, we only report correct Go, No-go, and Free trials, which differs from that of human go/no-go studies that focus on the failure of appetitive no-go trials (i.e., inhibiting the pre-potent response). In the present study, the dopamine signals that we interpret are restricted to successful trials only: action initiation (moving the lever press), action suppression (i.e., suppressing the prepotent Go response while maintaining their position in the nose-poke port), or no action (no lever press, not staying in the port). While this design differs from human Go/No-go paradigms, we believe our study of correctly performed Go, No-go, and Free trials are necessary for isolating action-dependent components of dopamine signaling in freely moving rats (action initiation vs action suppression vs action free).

      (b) Most Go/No Go tasks bias or overrepresent Go trials so that the Go response is prepotent, and consequently, successful suppression of the Go response is challenging. 1) I didn't see any information in the manuscript about how often each trial type was presented or 2) how the authors ensured that No Go responses (or lack thereof) were reflecting a suppression of the Go response.

      We appreciate the reviewer's attention to this important methodological consideration. The originally submitted version of the manuscript already addressed both concerns raised.

      Trial type presentation frequencies

      The Methods section describes our trial presentation approach: "On recording days, the trial types were counterbalanced. Within a session, Go left, Go right, and No-go trials were presented with 33% probability each, without replacement. For sessions with Free trials, trials were presented with a 25% chance without replacement."

      This design results in overrepresentation of Go trials overall (66% in the short-task; 50% in the long-task), which establishes the prepotent Go response as intended in standard Go/No-go paradigms. To improve clarity, this detail has now been included in the Methods section.

      Ensuring No-go responses reflect suppression of Go response

      Our paradigm incorporates multiple features that ensure successful No-go performance reflects suppression of the prepotent Go response:

      First, the overrepresentation of Go trials (addressed above) establishes response prepotency. Second, during No-go trials, rats must maintain their snout in the nose-poke port for the duration of the action cue, which creates the requirement to suppress the natural tendency to approach rewards (i.e., Pavlovian bias; Jones et al. 2017, 10.1016/j.bbr.2017.05.044; Guitart-Masip et al. 2014, 10.1007/s00213-013-3313-4; Dayan et al. 2006; 10.1016/j.neunet.2006.03.002). In Go trials, such natural bias does not require suppression as the lever can be approached and pressed during the action-cue period. This conflict between the instrumental No-go requirement and the Pavlovian-instrumental bias toward action makes action suppression particularly challenging (consistent with computational accounts of similar paradigms; Lloyd & Dayan 2023, 10.1371/journal.pcbi.1011569; Jones et al. 2017, Guitart-Masip et al. 2014, Dayan et al. 2006).

      Figure 3 provides behavioral evidence of this challenge: animals frequently left the nose-poke port and developed spontaneous motor strategies (such as biting and digging) to stay in the port, suggesting Pavlovian bias interfering with response suppression for rewards. Importantly, all reported No-go data include only correct trials (i.e., those without lever presses), ensuring that the dopamine signal reflects successful response suppression rather than failed Go attempts.

      (2) The authors observe relatively consistent differences in the DA signal between Go and No Go trials after the action-cue onset. However, the response type was not randomized between trial type, so there is a confound between trial type (Go/No Go) and response (lever/nosepoke). The difference in DA signal may have nothing to do with the cue type, but reflects differences in DA signal elicited by levers vs. nosepokes.

      As stated in the Introduction section and discussed in our rebuttal to point 1a, the focus of our investigation is how VMS dopamine signals differ during action initiation versus action suppression for rewards, as this is a central unanswered question in the dopamine field. More recent work demonstrates that dopamine incorporates not only RPE but also action initiation (Syed et al. 2016; Hamid et al. 2016; Mohebi et al. 2019), and our goal is to further our understanding of action-dependent VMS signals during reward pursuit.

      The reviewer suggests that dopamine differences may reflect differences in lever vs. nosepoke rather than cue type (Go vs No-go). We respectfully suggest this concern reflects a misunderstanding by the reviewer of our experimental question. The cue-action relationship is the experimental manipulation itself. It is not possible to study how dopamine encodes instructed action initiation versus suppression without linking specific cues to specific actions. The suggestion to 'randomize' action type across cue types would eliminate the very phenomenon we are investigating: how dopamine signals differ when cues instruct different action requirements.

      Our experimental design specifically compares reward pursuit with action requirements (Go trials: lever press; No-go trials: sustained hold) to reward pursuit without action requirements (Free trials: direct magazine approach). This design allows us to isolate how action initiation and action suppression influence reward-related dopamine signaling, which can reveal how the timing of action initiation influences RPE-dopamine. And which is the point of the study: to show how actions influence RPE dopamine signaling.

      Supporting this interpretation:

      Firstly, trial types were randomly interleaved, and each auditory cue explicitly instructed a specific behavioral response. Our design directly follows established methods demonstrating that VMS dopamine encodes whether actions are initiated or suppressed following action cues (Syed et al. 2016). That study, like ours, intentionally linked cue identity to a specific action requirement to assess how dopamine reflects instructed behavioral control. Thus, the fact that Go and No-go cues map onto different actions is inherent to the question being addressed, not an unintended confound.

      Second, as discussed in our response to point 1b, the asymmetry between Go and No-go trials is theoretically essential. Go trials align with Pavlovian approach tendencies (action initiation to reward), while No-go trials create conflict with this bias by requiring action suppression despite the cue being associated with a reward. This Pavlovian-instrumental conflict makes suppression particularly challenging (Lloyd & Dayan 2023, PLoS Comput Biol 10.1371/journal.pcbi.1011569) and allows us to examine the role dopamine in overriding prepotent responses.

      Third, the inclusion of Free trials (discussed in point 1a) demonstrates that our findings reflect instructed action control rather than simply motor execution. Free trials require neither lever pressing nor nose poke maintenance, yet show dopamine dynamics distinct from both Go and No-go trials, confirming that dopamine signals encode action requirements beyond motor output per se.

      Finally, we demonstrate that VMS dopamine differs in the same trial type (No-go) and can be classified based on different movement patterns (Biting, Digging, Calm). Importantly, the difference in VMS dopamine only appeared after the action was completed, particularly during reward approach (Figure 3). This data argues against the idea that VMS dopamine is particularly tied to the specific operant manipulanda as suggested by the reviewer, but rather, may reflect an internal motivational state for reward.

      Together, the aim of the present study is not to redefine Go/No-go paradigms for rodents, but to utilize this task structure to investigate action-dependent dopamine signalling for rewards, which cannot be answered without the cue-action mapping that we have used.

      (3) Both Go and No Go trials start with the rat having their nose in the noseport. One cue (Go cue) signals the rat to remove their nose from the noseport and make two lever responses in 5 seconds, whereas the other cue (No Go cue) signals the rat to keep their nose in the noseport for an additional 1.7-1.9 s. The authors state that the time between cue onset and reward delivery was kept the same for all trial types, and Figure 1 suggests this is 2 s, so was reward delivered before rats completed the two lever presses? I would imagine reward was only delivered if rats completed the FR requirement, but again, the descriptions in the text and figures are incongruent.

      The reviewer asks whether reward was delivered before rats completed the two lever presses and notes incongruence between text and figures. We respectfully note that these details were stated in the originally submitted version of the manuscript (see below).

      Reward delivery timing

      The reviewer asks whether reward was delivered before rats completed the two lever presses, which refers to the short-task variant. No - reward was always delivered immediately after the second lever press for all Go trials. This is described in Methods Section 3: Behavioral procedures - Self-initiated task variant. For added clarity, we have now added the term “immediately”: “... food pellet dispensed into the reward-magazine immediately.”

      In the Results section, we report that the average latency to complete two lever presses was 1.8s, which closely matches the 1.7-1.9s nose-poke hold maintenance required for No-go trials in the “short” variant. Thus, the time point of reward delivery was matched between Go and No-go trial types.

      Representation of reward delivery timing in figure and text

      The reviewer's confusion appears to stem from the schematic representation in Figure 1 and the task variant structure. There were two overarching task variants with different trial requirements:

      “Short-task” variants: Go trials required two lever presses (completed on average in 1.8s);

      No-go trials required 1.7-1.9s nosepoke maintenance

      “Long-task” variants: Go trials required a ‘rewarded’ press to occur 2.7-3.2s after cue onset (completed within ~3s); No-go trials required 2.7-3s nosepoke maintenance.

      For simplicity in depicting action-cue onset in Figure 2c, we used grey shading with a speaker icon at approximately 0-2s and 0-3s to represent these two variants. This schematic representation was not intended to indicate the precise reward delivery time, which (as stated in the Methods) occurred only upon successful completion of trial requirements in the short-task variant, or 2s after successful completion of trial requirements in the long-task variant. To improve clarity, we have adjusted the legend of Fig. 2 for more clarity, adding “Shaded gray area depicts approximate duration of action-cue onset for “short” and “long” task variants.”, and included more information under Results: “In the short-task variants, action-cues switched off after trial completion and a reward was delivered immediately” and “n the long-task variants, action-cues switched off after trial completion, or in the case of Free trials after 3s, and reward was delivered 2s later (Figure 1a).”

      (4) The manuscript is difficult to understand because key details are not in the main text or are not mentioned at all. I've outlined several points below:

      (a) The author's description in the manuscript makes it appear as a discrimination task versus a Go/No Go task. I suggest including more details in the main text that clarify what is required at each step in the task. Additionally, providing clarity regarding what task events the voltammetry traces are aligned to would be very useful.

      We respectfully note that the requested details were already present in the originally submitted version of the manuscript (see below). However, we acknowledge that the task design is complex and may benefit from additional clarity in the main text to aid reader comprehension.

      Behavioral task

      The reviewer suggests our task appears more like a discrimination task than a Go/No-go task. We acknowledge that our paradigm differs from traditional Go/No-go tasks used in humans and primates, as discussed in our responses to points 1a and 2. However, we classify this as a Go/No-go task because it shares the defining feature: requiring action initiation (Go) and suppression (No-go). This classification is consistent with established rodent literature examining action initiation versus suppression (Syed et al. 2016).

      Moreover, as discussed in our response to point 1a, we included Free trials specifically to demonstrate that No-go trials require active suppression rather than discrimination alone. The distinct dopamine dynamics across Go, No-go, and Free trials confirm that our task captures action initiation, action suppression, and action-free states, which is an important contrast that we needed to address our research question about action-dependent dopamine signaling.

      The key requirements for each trial type and task variant are described at the beginning of the Results section. A full description of each step required in the task was provided in the Methods Section 3: Behavioral procedures, to avoid repetition in the Results. Specifically:

      Self-initiated Go/No-go task variant (“short”)

      Self-initiated Go/No-go/Free task variant (“long”)

      Cue-initiated Go/No-go and Go/No-go/Free task variant

      To improve clarity, we have added a reference in the Results section to the Methods: Behavioral procedures for additional procedural details.

      Voltammetry trace alignment

      The events to which voltammetry traces are aligned were stated in the legend:

      “... when traces were aligned to action-cue onset…”

      “... aligned to the time when animals departed the nose-poke port…”

      “... we realigned traces to the moment animals arrived at the reward magazine… “

      Figure 2c legend: "c) Traces aligned to action-cue onset"

      Figure 2d-e legend: “d) Traces aligned to nose-poke exit and e) magazine arrival….”

      However, to improve clarity, we have now added “aligned to action-cue onset.. “ to make the trace alignment more immediately apparent when results are first presented, and added: “Dopamine data were aligned to events of interest: action-cue onset, nose-poke exit, and magazine arrival.”

      (b) How many subjects were included in each task variant? The text makes it seem like all rats complete each task variant, but the behavioral data suggest otherwise. Moreover, it appears that some rats did more than one version. Was the order counterbalanced? If not, might this influence the DA signal?

      The number of subjects for each task variant was reported in the Methods Section 3: Behavioral procedures, where each task variant description includes the corresponding sample size (“Self-initiated Go/No-go task (‘short’; n =9)”, “Self-initiated Go/No-go/Free task (“long”; n = 5)”, “A total of n = 11 and n = 15 were included in the Cue-initiated Go/No-go and Cue-initiated Go/No-go/Free tasks, respectively.”

      For added clarity, we have also added a table for the separation of No-go trials and the number of subjects that it has come from in the Methods.

      Task variant completion

      Not all animals completed all task variants. As stated in Methods Section 4: Real-time dopamine recordings and analysis, animals had to achieve >60% success rate for each trial type on at least two consecutive training sessions to proceed to recording. Other reasons include electrode degradation before all recordings could be completed (See Author response image 1).

      Training order

      We trained four cohorts of animals. One cohort was trained first in the short-task variant

      (Cue-initiated Go/No-go), and the remaining three cohorts were trained first in Cue-initiated Go/No-go (3s) long-task variant (i.e., without Free trials, behavioral and FSCV data not presented in the manuscript). Training order was not fully counterbalanced due to constraints described above.

      Following additional analyses, our data suggest that training order did not influence our core comparison of Go minus No-go dopamine. To directly address whether training order influenced dopamine signals, we separated animals based on whether they were first trained in the Cue-initiated Go/No-go (2s) or the Cue-initiated Go/No-go/Free (3s). We calculated the average dopamine of each rat during the action-cue period and then calculated the difference between them (Author response image 4). We observed absence of evidence of a difference between the groups.

      Author response image 4.

      Average Go minus No-go dopamine reveals no effect of the initial training variant (2s-first vs 3s-first) or test task version (Go/No-go vs. Go/No-go/Free). A 2 × 2 Bayesian ANOVA (Cauchy prior: fixed effects r = 0.5; random effects r = 1) consistently favoured the null model over all alternatives. The main effect of the initial training variant and test task version both showed moderate evidence of absence (initial training variant: BF<sub>10</sub> = 0.378; test task version: BF<sub>10</sub> = 0.367). The model including both main effects performed more poorly (BF<sub>10</sub> = 0.134), and the full model including an initial training variant × test task version interaction was the least supported of all models examined (BF<sub>10</sub> = 0.069). These results provide moderate evidence in favour of H<sub>0</sub>, suggesting that neither the variant animals were first trained on, the task version administered at test, nor their interaction meaningfully predicted average Go minus No-go dopamine responses.

      (5) There is a major challenge in their design and interpretation of the dopamine signal. Both trial types (Go and No Go) start with the rat having their nose in the noseport. An auditory cue is presented for 2-3 s signaling to the rat to either leave the noseport and make a lever response (Go trial) or to stay in the noseport (No Go trial). The timing of these actions and/or decisions is entirely independent, so it is not clear to me how the authors would ever align these traces to the exact decision point for each trial type. They attempt to do this with the nose-port exit analysis, but exiting the noseport for a Go trial (a rat needs to make 2 lever presses and then get a reward) versus a No Go trial (a rat needs to go retrieve the reward) is very different and not comparable.

      We respectfully disagree with the reviewer’s assertion that our data alignment approach is problematic. Aligning neural activity to specific behavioral epochs that occur at different times and across conditions is a widely used method to investigate the relationship between neural activity and behavior. Just to mention some examples: data collected with fiber photometry (e.g., Tan et al. 2026, doi: 10.1038/s41386-026-02368-4; Hart et al. 2024, doi: 10.1016/j.celrep.2024.113828) and voltammetry (Hamid et al. 2016, doi: 10.1038/nn.4173; Syed et al. 2016, doi:10.1038/nn.4187).

      Alignment method

      We intentionally designed the task so that overall action timing is matched between trial types (as described in our response to point 3), while specific behavioral epochs occur at different times. This allows us to compare dopamine dynamics during comparable behavioral events across Go and No-go trials (e.g., nose-poke exit).

      We align data to three critical behavioral epochs, stated in the Methods Section 4: Real-time dopamine recordings and analysis - FSCV measurement and analysis: action-cue onset, nose-poke exit, and magazine arrival. Each alignment addresses a specific aspect of our research question:

      Action-cue onset alignment captures VMS dopamine dynamics when animals have explicit knowledge of trial type and the required action. This allows us to characterize how dopamine evolves following correct action selection, which is central to our research question about how dopamine differs during successful Go versus No-go action execution, as well as no overt action (Free) in the long-task variant.

      Nose-poke exit alignment captures dopamine dynamics at the moment animals initiate movement. The reviewer suggests that exiting for Go versus No-go trials is "very different and not comparable" because subsequent actions differ (two lever presses vs. direct reward retrieval). However, this is precisely our experimental manipulation: we compare dopamine signals when animals exit the nose-poke port to perform different actions. This comparison is both valid and necessary to address our research question (does dopamine encode action initiation?).

      Magazine arrival alignment captures dopamine dynamics at reward approach. This allows us to differentiate between spatial proximity to reward from other concepts including temporal proximity (how soon is reward) and action requirements.

      We acknowledge that we cannot identify the precise moment of decision formation. In fact, the precise moment of decision formation is irrelevant for our question. However, our aim is to characterize VMS dopamine dynamics during successful action execution for rewards, following cues associated with specific actions.

      (6) The voltammetry analysis did not appear to test the hypotheses the authors outlined in the intro. All comparisons were done within task variants (DA dynamics in Go vs. No Go trials, aligned to different task events), but there were no comparisons across task variants to determine if the DA signal differed in cued vs self-initiated trials.

      Our aim was to investigate whether VMS dopamine signals consistently differed between action initiation and action suppression during reward pursuit. To test this within-variant contrast (Go > No-go), we manipulated how reward pursuit is initiated (self- vs cue-initiated) and the “effort” requirements (short vs long task variants), and our results show that they did not affect the differential between Go and No-go.

      The consistent Go > No-go dopamine that we observed across all task variants, together with the consistent increase during magazine approach, supports our conclusion that VMS dopamine integrates motivated action and reward.

      The reviewer suggests that we should have compared dopamine signals across self- vs cue-initiated task variants. We acknowledge this is an interesting, but entirely different, question and have addressed it in the Discussion section. We note that differences in controllability altered the time course of increased VMS dopamine, presumably by triggering earlier positive RPEs in Cue-initiated tasks as compared to Self-initiated tasks (illumination of the nose-poke light being the earliest predictor of reward). However, since our primary research question relates to the difference in VMS dopamine between action initiation and suppression, our results show that this difference was unaffected in two variants (short and long), strengthening our conclusions about the relationship between action and reward-related dopamine signaling.

      (7) Classification of No Go behaviors was interesting, but was not well integrated with the rest of the paper and was underdeveloped. It also raised more questions for me than answers. For example:

      (a) Was the behavior classification consistent across rats for all No Go trials? If not, did the DA signal change within subjects between biting vs digging vs calm?

      (b) If "biting rats" were not always biting rats on every No Go trial, then is it fair to collapse animals into a single measure (Figure 3C).

      (c) Some of the classification groups only had 2 or fewer rats in them, making any statistical comparison and inference difficult.

      Behavioral classification for each rat was consistent across trials (i.e., 100%, see Author response image 5). Upon reviewing the consistency of classifications within individual animals, we found that “Biting” animals exhibited biting behavior across the majority of their No-go trials. Only one animal (in the Self-initiated Go/No-go "short" variant) showed mixed classifications across trials, occasionally exhibiting digging or calm behavior. For all other animals, the predominant behavioral classification was highly consistent within subjects across sessions. The occasional trials where “biting” animals did not bite were too infrequent to permit meaningful within-animal comparisons. Therefore, we believe collapsing animals by their predominant behavioral phenotype in Figure 3C is appropriate and accurately represents stable individual differences in No-go response strategies.

      As stated in the Methods Section 6: Statistical Analysis - Clustering No-go behaviors and regrouping animals (last-line), we specifically avoided between-subjects statistical comparisons for groups with n≤2, as this would be inappropriate (see Author response image 5), and reported qualitative observations only. These exploratory findings at the individual-trial level suggest behavioral heterogeneity during No-go trials, that others may use for future investigation, but do not form primary conclusions.

      The behavioral classification is integrated with our central findings on VMS dopamine encoding spatial proximity. Our results demonstrate that individual variation in action suppression strategy, in particular Biting behaviors, consistently manipulates the timing of max dopamine release during subsequent reward approach, but not during the action itself (Figure 3b-c). This links our observations of action-dependent dopamine (Figure 2c) with spatial reward approach (Figure 2e). We believe that our findings shed new light into the understanding of how action modulates reward-related dopamine dynamics at the individual level. This has been discussed in Discussion section: Dopamine dynamics are linked to motivated action.

      Author response image 5.

      Rats predominantly stick to a particular strategy to perform No-go trials. Each bar represents an individual animal, and colours represent the % of each classification type.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figures: It would be helpful to have more panel labels on the figures - 1C, for example, labels 4 different dataset panels, similar to a few other cases. This would really help to improve readability, as there are a lot of task types, trial types, behavioral measures, and labels to sift through. The figures overall are very busy, and it is a bit of a challenge to process the tasks and trial comparisons, as well as interpret what the quantification insets mean.

      We thank the reviewer for their suggestion. We have added more panel labels to Figure 1 to improve the ability to follow with the main text. We hope that the reviewer finds it acceptable.

      (2) Figures: A bit more specifically, in Figure 2, the boxplot insets are pretty hard to see, and it's not clear what scale they are on or what data they reflect. Similarly, it's unclear what the horizontal bars reflect in terms of which conditions are being compared. Why are box plots used for some comparisons, and why are some comparisons based on time series bootstrapping, but others are not clear? I would consider broadly reworking this figure and its description for clarity. Figure 3, by comparison, is easier to understand - the quantifications are clearer and labeled.

      We have now labeled the comparisons being depicted by the horizontal bars in Fig 2c. We have also clarified the boxplot analysis in Fig 2d and Fig 2e by adding 'Max dopamine’ labels to the figure, and have made these analysis methods more explicit in the figure legend.

      We used time-series bootstrap analysis to identify when dopamine signals diverged between trial types, which depicts dopamine differences during distinct action requirements. The latency-to-max quantification provides a summary measure to test specific encoding hypotheses (temporal vs. spatial proximity to reward).

      (3) Broadly, more clarity on the FSCV analysis is warranted.

      (a) Targeting: It looks like the dataset contains a mix of medial shell and mostly core accumbens placements. The paper treats VMS as a uniform dopamine region, but it is more standard to separate core and shell (and also other parts of the shell) into subregions. Many of the reported encoding profiles here are known to differ across the accumbens. So, some consideration of this seems appropriate - a minimal signal-behavior comparison for the shell vs core subgroups, for example.

      We thank the reviewer for raising this point. Upon careful re-examination of our histological analysis, we identified an error that occurred when we made the overlay of electrode placements across rostrocaudal planes (to project placements onto a single plane for the sake of simplicity; Fig 2a): We incorrectly assigned some recordings to the nucleus accumbens shell. We corrected this error, which shows that the vast majority of recordings were in the core, with only 2 animals in the shell (black stars). We have adjusted Fig 2a to reflect this, and added an anatomically more complete illustration of electrode placements across rostrocaudal planes as Supplementary Figure S8.

      While we acknowledge reported differences between core and shell dopamine in some contexts, the small number of shell placements precludes meaningful statistical comparison. However, to determine whether average dopamine concentrations differed between core and shell during the action-cue period, we have plotted the values in Author response image 6. The fact that shell data mostly falls centrally into the overall core-data distribution suggests no consistent difference in dopamine release.

      Furthermore, in our experience (and that of colleagues (personal communication)) with appetitive operant tasks, core and shell FSCV dopamine signals do not substantially differ for action-selective encoding and reward approach. Given the sample distribution and our focus on general VMS function in Go/No-go behavior, we believe pooling these regions is appropriate.

      Author response image 6.

      Average dopamine release during the action cue in nucleus accumbens core (circles) and shell (stars) animals showed no distinct separation between regions. Each symbol represents an animal.

      (b) Design: In my understanding of the design, the main distinction between the short and long task variants is a 2-second versus a 3-second required nose poke hold. 3 seconds here is "long" and more "difficult". I'm not sure I agree that a 1-sec distinction really reflects a difference in task difficulty or effort. Can the authors point to a past paper that demonstrates this variation is sufficient to engage a behavioral difference and/or a neural encoding difference? Broadly, some justification of the validity of this manipulation is needed, I think.

      While we lack direct citations for this specific manipulation, we believe that the 2s and 3s hold requirements represent meaningful differences in difficulty based on our extensive rat-behavior experience and behavioral evidence in Author response image 2.

      Importantly, the difficulty of No-go trials does not stem merely from the required time to hold their snouts in the nose-poke port, but from suppressing the motivational/Pavlovian bias to approach reward-associated cues. In No-go trials, subjects must suppress the prepotent tendency to immediately approach reward-related stimuli and instead maintain active suppression of this approach behavior. Even the 2s hold is challenging as animals tend to perform better on Go trials compared to No-go trials (Fig 1b and 1c), demonstrating the inherent difficulty of response suppression even at the shorter duration. The additional 1-second substantially increases this demand, as it represents a 50% increase in hold duration.

      Our training data clearly demonstrate the difficulty in reaching task criterion when increasing the required action (for Go and No-go trials) from 2s to 3s. Across four cohorts of animals trained in Go/No-go task variants, one cohort that was trained first in the 2s task variant, and the remaining three cohorts were trained in the 3s task variant of Go vs No-go. Animals required 41.1 ± 7.9 (n = 27) sessions to learn the 3s hold (approximately 8 weeks), versus

      18.5 ± 7.6 sessions for the 2s hold (approximately 4 weeks; mean ± SEM). This indicates that despite only a 1-second difference in required action performance, animals needed more than double the number of training days to reach criterion, clearly indicating differential effort demands.

      (c) Figure 2 results: the authors state that because there is a greater DA signal to Go vs No Go cues in all the task variants, this means that controllability of reward pursuit and increased task effort do not affect VMS dopamine. But the magnitude of the signals looks different across the task variants - it looks clearly stronger overall in the self-initiated tasks, for example. Given that dopamine signals are not compared across task variants (I think the tasks are all between-subjects?), I don't think the above conclusion is justified.

      We respectfully clarify that our conclusion does not claim controllability and effort have no effect on dopamine magnitude, but rather that these manipulations do not affect the action-selective difference in dopamine (Go > No-go). Our central finding is that the relative difference between Go and No-go remains consistent across all task variants (within-subjects comparison).

      We did not perform across-variant comparisons of absolute dopamine magnitudes because that was not our primary research question. Our focus was to understand whether dopamine differs between action initiation and suppression, and whether this difference can be modulated by controllability or effort.

      We acknowledge the reviewer’s observation that absolute magnitudes appear larger in self-initiated vs cue-initiated task variants. We believe that this likely reflects differences in RPE timing rather than controllability per se: in cue-initiated tasks, the nose-poke light provides an early trial-start signal, distributing RPE temporally across the trial. In self-initiated tasks, trial-initiation and action requirements are temporally integrated. Though understanding how controllability affects absolute dopamine magnitude is an interesting question for future work (e.g., using sophisticated regression-based encoding models), it was beyond the scope of our current investigation, which focuses on action-selective encoding.

      (4) Broadly, I don't think these data, as shown, support the conclusion "dopamine reflects action initiation but not controllability or effort" without more analysis and additional context.

      We have revised the wording of our conclusions throughout the manuscript to better reflect our intent, which is to compare the action-selective encoding of dopamine (i.e., action initiation vs action suppression). It is now “Reward-related dopamine depends on action initiation irrespective of controllability and effort”.

      (a) Figure 3 - more description of the classified behaviors would be helpful for interpreting this part of the data. When are the behaviors occurring - during the hold cue? Or is the classification related to what they do immediately after holding? Or something in between> I guess I'm not sure what digging and biting are in the context of a nose poke hold. As described, it's not clear what the signal differences relate to - movement differences? Generally, it's not clear what to make of the behaviors. They seem to emerge spontaneously, but it's not clear whether the specific actions mean anything, so it's a bit difficult to know what to glean from the dopamine is greater during "biting". It's a very different movement pattern, so perhaps this result relates to that, rather than task engagement or motivational drive per se?

      We thank the reviewer for the comment. We have added relevant information in the figure caption and in the Results section to clarify that classified behaviors occurred during the action-cue period (for No-go trials, the hold cue; Figure 3a caption). In addition, we have included Videos to better depict the classified No-go behaviors during the action-cue period.

      We agree with the comment that these classified behaviors, such as biting, seem to emerge spontaneously. Our interpretation of these behaviors is that they may represent the motivational state of each subject. Most importantly, whereas the behavioral differences occurred during the action-cue period (while animals had to suppress actions and stay within the nose-poke port), the difference in VMS dopamine was only observable after this behavior was completed. Thus, the movement pattern per se is likely not relevant to the dopamine release occurring after its completion. This has been discussed in the Discussion section: Dopamine dynamics are linked to motivation action.

      Based on our videos, it appears as though Digging could be perceived as more vigorous (i.e., more general movement in the nose-poke port). However, we did not observe more dopamine during the action-cue period of Digging trials as compared to Biting trials. Furthermore, more vigor during the action-cue period (e.g. Digging trials) did not result in more dopamine during the reward approach period. Together, the data suggest that another process may underlie the large increase in VMS dopamine in Biting trials during reward approach, such as varying attribution of incentive salience.

      (b) In some cases, but not all, dopamine measurement comparisons are done on a total trial basis, and in others, it seems to be subject averages. It's not clear why different approaches are used for different parts of the data. But also, for the trialwise analysis, what statistical steps were taken to incorporate the subject as a random factor in the analysis? If that is not done, then a trial-wise analysis artificially increases the power for the stat (n=trial#).

      We used different analytical approaches depending on sample size and data structure. To compute differences in Go vs No-go dopamine within each animal, as intended by our experimental design, we performed subject-level comparisons (Figure 2).

      For the behavioral classification analysis (No-go, Figure 3), we performed trial-level analyses to increase the statistical power and better characterize this unexpected and interesting phenomenon. We explicitly chose not to perform subject-level group comparisons because

      (1) some groups had only n=2-3 animals, making subject-level statistics underpowered, and (2) behavioral classifications were highly stable within individual animals (see Author response image 5). We acknowledge that formal between-group comparisons (across subjects) are underpowered due to small n, but the stability of within-subject No-go behavioral strategy and qualitatively distinct VMS dopamine profile suggest that these differences may be biologically meaningful and worthy of future investigation in larger samples. We have made these limitations more explicit in the Results.

      This relates to Figure 3, where all trial data are shown next to individual subjects - the subject-wise group comparisons are between 2-5 or so rats, which is quite low. In Figure 2, a subject n of 27 is listed, so it's not clear why this analysis is on such a small set of rats. Generally, it's not clear how many rats/subjects are in each data bit. The methods say only 5 rats are in the long self-initiated task, but 15 in the cue-initiated task. Clarity in all this is needed, including in the figures/captions.

      We have now added detail about the statistical test performed in Figure 3’s caption:

      “After action-cue offset: No-go (trials)’: Individual No-go trials classified by No-go behavior, and trial-wise statistical Kruskal-Wallis tests were performed for each task-variant.” We have also added in a table in Methods (Table 1), showing the number of trials in each No-go classification, and animals regrouped based on their predominant No-go strategy (see Methods section: Statistical analysis - Clustering No-go behaviors and regrouping animals).

      For the small sample sizes based on the regrouping of animals based on their predominant No-go strategy, we have now added in the caption, “... Rats classified based on their predominant No-go strategy, with no statistical tests performed.”

      (5) I'm also a little confused about the paper's narrative that the dopamine data reflect spatial (but not temporal) proximity to reward - it seems that this conclusion is based on the dopamine signal peaking at magazine entry, but that is different, I think, than a spatial signal per se (space is not manipulated in this study). I think more analysis of the signals during the magazine approach behaviors would be helpful, and possibly comparing rewarded vs unrewarded approaches. The emphasis, including in the title, that a major take-home of the data is that dopamine encodes reward proximity, is not really borne out by the current analyses. Reward expectation is not manipulated independently of the approach action, so it's hard to pin this on "space" vs "reward is soon". This is admittedly a general complexity in characterizing dopamine ramps.

      As the reviewer notes, we acknowledge that 'spatial proximity' and 'reward is soon' are challenging to fully dissociate in appetitive approach paradigms. However, we believe that our data and new additional analyses, which is now included in Results: Maximum VMS dopamine release encodes spatial but not temporal proximity to rewards and Supplementary Figures S10-12, provide compelling evidence that VMS dopamine primarily reflects spatial proximity to the expected reward location, rather than temporal proximity to reward delivery.

      Evidence against full temporal encoding:

      (1) Max VMS dopamine occurred up to 3s before reward delivery in Free trials (Fig 2e, open circles vs triangles). Furthermore, when we calculated the max values of individual trials of realigned traces, maximum dopamine does not consistently coincide with reward delivery across trial types (new Supplementary Figure S11).

      (2) If dopamine encoded temporal proximity from the earliest reward-predictive cue, we would expect consistent accumulation from cue onset (action cue for self-initiated, nose-poke light for cue-initiated). However, realigned trials also did not show a consistent accumulation around these events (new Supplementary Figure S12).

      Evidence for spatial encoding:

      (1) Max dopamine consistently occurred when animals arrived at the magazine across all trial types and task variants, regardless of when reward was actually delivered (Fig. 2e).

      (2) Across individual trials, max dopamine values concentrated around the magazine-panel, with a striking accumulation when animals were in close proximity to it (new Supplementary Figure S10) showing a distance-dependent distribution.

      (3) Assessment of unrewarded magazine approaches during the intertrial interval (ITI) revealed no increase in dopamine release (Fig 2f), indicating that VMS dopamine requires task-relevant reward expectation.

      (6) Examples of the DLC workflow in a supplement would be appropriate. Also, video examples of the 3 kinds of behaviors from the clustering analysis could be useful for understanding what they are/what they mean.

      We have added supplementary figures showing the DeepLabCut workflow (Supplementary Figure S9) and Videos 1-3 demonstrating the three behavioral classifications (biting, digging, calm) during No-go trials.

      (7) Referencing/scholarship: I would suggest broadening the citation pool for the paper to include more older work that has established the notion that dopamine signaling and the accumbens act as a motivation-action interface, as this has been a longstanding notion since at least the 1980s. There is also a sizable literature on dopamine signaling of effort, some of which would be appropriate to cite here.

      We thank the reviewer for the recommendation and have now expanded our citations to include foundational literature to work from the 1980s-90s. These can be found in the Introduction, Results, and Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) Lines 353-357- This came as a surprise, as the relevant results are only featured in supplementary information. These should be moved to the main manuscript. As both the biting behavior and faster lever-press completion lead to larger peak dopamine, does this represent response vigor?

      We appreciate the reviewer’s interest in these data. However, we believe that the data that we report in “Supplementary Fig 7: Quartile analysis of Go trials shows coordinated changes in last lever-press timing and VMS dopamine, does not warrant movement to the main manuscript.”

      The purpose of the lever-press timing analysis was to demonstrate that maximum dopamine release coincides with the moment animals arrived at the magazine, rather than with the action period itself. By sorting Go trials based on last lever-press latency, we show temporal coordination between action completion and dopamine timing but critically, dopamine peaks after action completion, not during it.

      This temporal dissociation argues against a 'response vigour' interpretation. If dopamine encoded motor vigour, we would expect the signal to coincide with or precede the vigorous action. Instead, both the lever-press data (Supplementary Figure S7) and the classified No-go trial data (Figure 3, biting behavior) show that changes in max dopamine occur after the actions themselves, during the subsequent approach to reward.

      Together, these findings demonstrate that VMS dopamine reflects spatial approach to the reward location following action completion, not the vigour of the required actions per se. The lever-press analysis serves as supporting evidence for this temporal relationship but does not introduce a novel finding that warrants main figure emphasis. We mentioned this temporal coordination in Results to ensure readers are aware of the converging evidence while maintaining focus on our central findings regarding action initiation versus suppression.

      (2) Line 13- the experimental work cited refers to midbrain dopamine neurons, rather than dopamine release within the VMS. Please correct.

      We thank the reviewer for pointing out this error. We have now corrected the citation to refer to studies measuring striatal dopamine rather than midbrain dopamine neurons.

      (3) Line 166 - Shouldn't this say "consistently delayed for No Go trials"? Looks like peak dopamine occurs later for these trial types.

      This statement refers to traces aligned to nose-poke exit (not action-cue onset). The peak Go dopamine occurred later than other trial types following exit (Green arrows in Figure 2d).

      (4) Lines 391-392 - however however

      Thank you for the comment, we have adjusted the text.

    1. eLife Assessment

      In their important manuscript, Gangadharan, Kober and Rice focus on how Stu2/XMAP215-family microtubule polymerases use their TOG domains to catalytically promote microtubule growth, testing whether their mechanism follows an enzyme-like kinetic model similar to that of actin polymerases. The authors integrate measurements including microtubule polymerization rates and TOG-tubulin binding kinetics to convincingly show that Stu2 follows an enzyme-like model where tight tubulin binding enables efficient polymerization, revealing a shared mechanism with actin polymerases despite their evolutionary divergence. This work will be of general interest to the cell biology and biophysics communities.

    2. Reviewer #1 (Public review):

      This study by Gangadharan and colleagues provides significant progress towards a quantitative biochemical mechanism for Stu2 polymerase activity. A key conceptual advance is the novel application of an enzyme-like model, initially developed for the actin polymerase Ena/VASP, to Stu2.

      Strength:

      New refined affinity measurements for a Stu2 TOG domain using Bio-layer interferometry show more than an order of magnitude higher affinity of TOG domains to tubulin compared to previously published reports.

      The findings reinforce the "concentrating reactants" or, more specifically, for TOG-domain proteins, the "tubulin-shuttling antenna" model, compared to the "polarized unfurling" model, a more speculative structural hypothesis.

      The manuscript builds upon a series of previous manuscripts that showcase the profound intellectual engagement with microtubule polymerization mechanisms by TOG-domain proteins from the Rice lab, a thought leader in microtubule polymerization for over a decade.

      Minor weakness:

      The affinity discrepancy is not fully resolved by side-by-side measurements, which seem to be not feasible as not all buffer conditions are compatible with all assays.

    3. Reviewer #2 (Public review):

      Summary:

      The manuscript from the Rice lab by Gangadharan et al., submitted to eLife, investigates the polymerization mechanism of the yeast microtubule polymerase Stu2. The lab has published a number of articles demonstrating the structural basis by which the two TOG domains of Stu2 each bind free tubulin heterodimers and has developed a tethered polymerization model by which the TOG domains drive polymerization by shuttling those tubulin subunits onto the microtubule plus end. A second model was proposed by Nithianantham et al. (eLife, 2018) based on a closed - to - open transitional state in which Stu2 unfurls and loads two longitudinal associated tubulin heterodimers onto the microtubule plus end. While the second model is not directly tested, the current work aims to further characterize/model the tethered polymerization model using a kinetic framework developed by developed by Breitsprecher et al. for Ena/VASP actin polymerization activity, using a model that is enzymatic (EMBO J., 2011). The general architecture and function of Ena/VASP on actin polymerization versus Stu2 on microtubule polymerization is a reasonable relation and hits upon, as the authors note, potential convergent mechanistic evolution across distinct cytoskeletal networks. The model effectively treats tubulin as the substrate, and the polymerized microtubule plus end as the product. If Stu2 is "enzymatic" in this framework, the model predicts it would behave with Michaelis-Menten kinetics, that there would a Vmax, and polymerase activity would either be "affinity limited" by TOG:tubulin affinity (KD) and/or "kinetically limited" by TOG:tubulin association (Kon) and transfer of tubulin to the microtubule plus end (Kt). The authors find that the Brietsprecher model works well for Stu2 activity, and that Stu2 best aligns with a "kinetically limited" model. The work is interesting and adds to the growing elucidation of the Stu2 microtubule polymerase model. While yeast microtubule polymerases are somewhat distinct in their architecture, there is significant overlap that findings from the manuscript can be utilized to inform the mechanisms of larger, more complex microtubule polymerases such as human ch-TOG.

      Strengths:

      The manuscript invokes the enzymatic model of Breitsprecher et al. used for Ena/VASP and conducts an elegant series of (mostly established) experiments to determine whether Stu2 microtubule polymerase activity aligns with the model - which they conclude does align, supported by the data/results obtained.

      Weaknesses:

      The authors used biolayer interferometry to measure TOG:tubulin affinity. The affinities obtained were significantly higher affinities than the lab obtained in an earlier publication using analytical ultracentrifugation. While differences in buffer and salt conditions may underlie these differences, additional runs using comparable buffer systems, or use of a third independent assay to measure affinities would have added rigor.

      The discussion could be expanded to better compare and contrast the results with both existing polymerase models introduced in the introduction, as well as expanded to look at reversible enzymatic activity (microtubule depolymerization at low to zero tubulin concentrations) and microtubule plus versus minus end activity.

      Comments on revised version.

      The revised submission has addressed these comments adequately.

    4. Reviewer #3 (Public review):

      Summary:

      This study by Gangadharan and colleagues seeks to establish a quantitative biochemical model for the microtubule polymerase activity of Stu2. Stu2 is the budding yeast member of the XMAP215 protein family, which is broadly conserved across eukaryotes. XMAP215 proteins play a wide variety of important roles in cells, and these are attributes to effects on microtubule dynamics. Many studies over the last ~20 years have shown that XMA215 proteins selectively associate with microtubule ends where they increase rates of microtubule assembly and disassembly. More recently, structural biology and biochemical studies by the authors and other groups have shown that the multiple TOG domains on XMAP215 proteins are tubulin-binding domains that selectively bind to curved tubulin, which is present in solution and at microtubule ends, but not to straight tubulin which is present in the walls of the microtubule lattice. This has led to the general model that XMAP215 proteins promote polymerization by delivering soluble tubulin to the growing plus end, and two distinct models have been proposed to explain the mechanism. The 'concentrating reactants' model proposed previously by the authors suggests that TOG domains grab hold of tubulin in solution and concentrate at the microtubule end. The 'polarized unfurling' model proposed by the Al Bassam lab suggests that XMAP215 delivers multiple tubulins to the end, using a stepwise mechanism involving different roles for each TOG domain. The current study seeks to improve our understanding of the mechanism by developing a quantitative model to explain the binding and release of tubulins, the number of Stu2 molecules at the end, and the overall rate of tubulin addition. The authors accomplish this goal using new experimental data. The final model fills in new details of the mechanism. The authors draw a comparison between Stu2 and the actin polymerase which bears similarity to the Ena/VASP and suggest a convergent strategy for cytoskeletal polymerases.

      Strengths:

      This is a focused and clearly written study that incorporates prior knowledge of XMAP215 and draws inspiration from the actin field. The data are clear and convincing, and the study accomplishes its goal of generating a new, quantitative model for Stu2. The model will be important for microtubule researchers to predict and test key points for altering XMAP215 activity across different organisms and potentially for different tubulin substrates. The comparison to Ena/VASP may also inspire similar comparisons across other microtubule and actin regulators, which could lead to new insights across cytoskeletal fields.

      Weaknesses:

      The study is without major weaknesses.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In their important manuscript, Gangadharan, Kober and Rice focus on how Stu2/XMAP215-family microtubule polymerases use their TOG domains to catalytically promote microtubule growth, testing whether their mechanism follows an enzyme-like kinetic model similar to that of actin polymerases. The authors integrate measurements including microtubule polymerization rates and TOG-tubulin binding kinetics to convincingly show that Stu2 follows an enzyme-like model where tight tubulin binding enables efficient polymerization, revealing a shared mechanism with actin polymerases despite their evolutionary divergence. This work will be of general interest to the cell biology and biophysics communities.

      Thank you for the favorable assessment of our manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Gangadharan and colleagues provides significant progress towards a quantitative biochemical mechanism for Stu2 polymerase activity. A key conceptual advance is the novel application of an enzyme-like model, initially developed for the actin polymerase Ena/VASP, to Stu2.

      New refined affinity measurements for a Stu2 TOG domain using Bio-layer interferometry show more than an order of magnitude higher affinity of TOG domains to tubulin compared to previously published reports.

      The findings reinforce the "concentrating reactants" or, more specifically, for TOGdomain proteins, the "tubulin-shuttling antenna" model, compared to the "polarized unfurling" model, a more speculative structural hypothesis.

      The manuscript builds upon a series of previous manuscripts that showcase the profound intellectual engagement with microtubule polymerization mechanisms by TOG-domain proteins from the Rice lab, a thought leader in microtubule polymerization for over a decade.

      Minor remarks:

      (1) A major new experimental finding of this paper is the affinity of TOG domains, which is more than an order of magnitude lower (10 nM) than previous measurements from the same lab (~200 nM). The authors attribute this change to ionic strength differences between buffer conditions, citing the lab's previous work (Ayaz et al., 2014). This argument left me contemplating what the buffer conditions are in both experiments, and I wonder if other readers would feel the same. After going down the rabbit hole, I believe the difference in ionic strength is ~2.3 fold, and at least on the back of my envelope, this works out beautifully with the measured differences in affinities. A short version of this argument may strengthen the manuscript.

      This is a good comment. We should have been clearer about the different buffer conditions. The revised manuscript now explicitly states how the two buffers in question differ in pH and ionic strength. (Page 8, ‘Tubulin binds rapidly …’ section). We tried to perform comparative measurements of TOG:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is stated in the revised manuscript (Page 9, final paragraph before the ‘Unifying measurements …’ section. Along the lines of the reviewer’s ionic strength calculation, and consistent with the increase in affinity we observed with lower ionic strength, we now also state that prior measurements from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with higher ionic strength (100 mM KCl vs 200 mM KCl).

      (2) I am wondering if there may be an alternative explanation to tubulin binding by TOG being the kinetically rate-limiting step for polymerase function:

      TOG + Tubulin ⇌ TOG:Tubulin (fast binding rate, high-affinity binding)

      TOG:Tubulin + MT_end → TOG:MT (tubulin is incorporated into MT, fast transfer rate)

      The binding rate is 3/s, and the transfer rate is 5/s.

      I was wondering if the following step should be considered, which involves a conformational change of tubulin (e.g., straightening) TOG:MT → TOG + MT (ratelimiting straightening and unbinding of TOG from the lattice).

      This is an interesting thought that highlights a gap in the understanding of microtubule dynamics.

      Presumably, the affinity of TOGs for straight tubulin is practically zero for the purpose of this discussion, as there is no lattice binding, which means unbinding is likely very rapid; however, straightening may be the rate-limiting factor here.

      In theory, straightening should also be rapid; however, we lack measurements of how fast or slow this step occurs within the context of a TOG domain, which presumably skews the process towards curved tubulin.

      We agree (based on prior observations) that the affinity of these TOGs for straight tubulin is negligeable in this context. There is much less data about the timescale of tubulin straightening, with or without a TOG domain bound, or even about how tubulin interactions with the microtubule end affect the balance of preferred conformations and/or the rate of conformational change. It’s an extremely interesting topic. Because the straightening process the reviewer envisions is zero-order, the transfer rate in our model could in principle reflect slow straightening (in this view the ‘delivery’ step would need to be very fast, i.e. not rate-limiting). Because there is so little data about this, and because there are not yet methods to study or perturb the timescale of straightening on the microtubule, we prefer not to engage too deeply. We added a sentence to acknowledge this alternative possibility in the revised manuscript (bottom of Page 4 and top of Page 5).

      A hypothetical Stu2, when bound to the microtubule end and with the TOG domain not disengaged from tubulin, would not permit the processivity of that molecule or the binding of a new molecule.

      To emphasize the importance of unbinding, when it is not efficient, as reported for the T238 mutant that results in Stu2 lattice binding (Geyer et al., 2018), the polymerase becomes inefficient.

      The mechanism of polymerase processivity has not been conclusively determined (the Geyer et al. 2018 eLife paper took a step in that direction, though). The model used in this paper is only concerned with how many polymerases are at the microtubule end at steady-state (as opposed to how long a particular polymerase acts before dissociating), so while we appreciate and are interested in these questions, we think it would be better to leave them for future work.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from the Rice lab by Gangadharan et al. investigates the polymerization mechanism of the yeast microtubule polymerase Stu2. The lab has published a number of articles demonstrating the structural basis by which the two TOG domains of Stu2 each bind free tubulin heterodimers, and has developed a tethered polymerization model by which the TOG domains drive polymerization by shuttling those tubulin subunits onto the microtubule plus end. A second model was proposed by Nithianantham et al. (eLife, 2018) based on a closed-to-open transitional state in which Stu2 unfurls and loads two longitudinally associated tubulin heterodimers onto the microtubule plus end. While the second model is not directly tested, the current work aims to further characterize/model the tethered polymerization model using a kinetic framework developed by Breitsprecher et al. for Ena/VASP actin polymerization activity, using a model that is enzymatic (EMBO J., 2011). The general architecture and function of Ena/VASP on actin polymerization versus Stu2 on microtubule polymerization is a reasonable relation and hits upon, as the authors note, potential convergent mechanistic evolution across distinct cytoskeletal networks. The model effectively treats tubulin as the substrate, and the polymerized microtubule plus end as the product. If Stu2 is "enzymatic" in this framework, the model predicts it would behave with Michaelis-Menten kinetics, that there would a Vmax, and polymerase activity would either be "affinity limited" by TOG:tubulin affinity (KD) and/or "kinetically limited" by TOG:tubulin association (Kon) and transfer of tubulin to the microtubule plus end (Kt). The authors find that the Brietsprecher model works well for Stu2 activity, and that Stu2 best aligns with a "kinetically limited" model. The work is interesting and adds to the growing elucidation of the Stu2 microtubule polymerase model. While yeast microtubule polymerases are somewhat distinct in their architecture, there is significant overlap that findings from the manuscript can be utilized to inform the mechanisms of larger, more complex microtubule polymerases such as human ch-TOG.

      Thank you for the nice summary and favorable comments.

      Strengths:

      The manuscript invokes the enzymatic model of Breitsprecher et al. used for Ena/VASP and conducts an elegant series of (mostly established) experiments to determine whether Stu2 microtubule polymerase activity aligns with the model, which they conclude does align, supported by the data/results obtained.

      Weaknesses:

      The authors used biolayer interferometry to measure TOG:tubulin affinity. The affinities obtained were significantly higher than the lab obtained in an earlier publication using analytical ultracentrifugation. While differences in buffer and salt conditions may underlie these differences, additional runs using comparable buffer systems, or the use of a third independent assay to measure affinities, would have added rigor.

      This is a good question that was also raised by reviewer #1. We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems using biolayer interferometry, then analytical ultracentrifugation and isothermal titration calorimetry, but in each technique one or the other buffer caused aggregation, nonspecific binding, or some other artifact that prevented such an analysis. This is now stated in the revised manuscript (page 9, final paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been independently shown to depend on ionic strength in a way that seems consistent with what we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, final paragraph before the ‘Unifying measurements …’ section).

      The discussion could be expanded to better compare and contrast the results with both existing polymerase models introduced in the introduction, as well as expanded to look at reversible enzymatic activity (microtubule depolymerization at low to zero tubulin concentrations) and microtubule plus versus minus end activity.

      Thank you for the push to be more explicit about the two contrasting models. We made small changes to the introduction (top paragraph on page 3) and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (page 12, penultimate paragraph of the main text).

      The ‘transfer’ reaction is treated as irreversible (analogous to catalysis by an enzyme), so the biochemical model we use for the polymerase cannot account for polymeraseinduced microtubule depolymerization at low to zero tubulin concentration. We added text to state that the model is limited to the growth reaction (page 4, last paragraph) but otherwise prefer to not engage too deeply in questions about the reverse reaction.

      These polymerases are thought to be plus-end specific because of the domain organization of the protein: TOGs bind tubulin such that the N- to C-terminal polarity of the TOG corresponds to the plus- to minus-end polarity of the tubulin, and the basic region used to make a ‘slippery’ connection to the microtubule is located C-terminal to the TOGs. These two factors mean that it is only at the plus-end that TOGs can engage αβ-tubulins with the basic region contacting surfaces ‘deeper’ in the polymer. We added text about these issues, citing prior work, to the legend of Figure 5 (page 11). We chose to not elaborate much since it is not the primary focus of the paper.

      Reviewer #3 (Public review):

      Summary:

      This study by Gangadharan and colleagues seeks to establish a quantitative biochemical model for the microtubule polymerase activity of Stu2. Stu2 is the budding yeast member of the XMAP215 protein family, which is broadly conserved across eukaryotes. XMAP215 proteins play a wide variety of important roles in cells, and these are attributed to effects on microtubule dynamics. Many studies over the last ~20 years have shown that XMA215 proteins selectively associate with microtubule ends, where they increase rates of microtubule assembly and disassembly. More recently, structural biology and biochemical studies by the authors and other groups have shown that the multiple TOG domains on XMAP215 proteins are tubulin-binding domains that selectively bind to curved tubulin, which is present in solution and at microtubule ends, but not to straight tubulin which is present in the walls of the microtubule lattice. This has led to the general model that XMAP215 proteins promote polymerization by delivering soluble tubulin to the growing plus end, and two distinct models have been proposed to explain the mechanism. The 'concentrating reactants' model proposed previously by the authors suggests that TOG domains grab hold of tubulin in solution and concentrate at the microtubule end. The 'polarized unfurling' model proposed by the Al Bassam lab suggests that XMAP215 delivers multiple tubulins to the end, using a step-wise mechanism involving different roles for each TOG domain. The current study seeks to improve our understanding of the mechanism by developing a quantitative model to explain the binding and release of tubulins, the number of Stu2 molecules at the end, and the overall rate of tubulin addition. The authors accomplish this goal using new experimental data. The final model fills in new details of the mechanism. The authors draw a comparison between Stu2 and the actin polymerase, which bears similarity to Ena/VASP, and suggest a convergent strategy for cytoskeletal polymerases.

      Thank you for the good summary and favorable comments.

      Strengths:

      This is a focused and clearly written study that incorporates prior knowledge of XMAP215 and draws inspiration from the actin field. The data are clear and convincing, and the study accomplishes its goal of generating a new, quantitative model for Stu2. The model will be important for microtubule researchers to predict and test key points for altering XMAP215 activity across different organisms and potentially for different tubulin substrates. The comparison to Ena/VASP may also inspire similar comparisons across other microtubule and actin regulators, which could lead to new insights across the cytoskeletal fields.

      Thank you for these comments.

      Weaknesses:

      The study is without major weaknesses, but there are several minor weaknesses worth noting. One is that the final model provides new details regarding the Stu2 mechanism, but does not provide a major new advance in our understanding of how the polymerase works. For example, the discussion does not clearly argue for whether the new results and model rule out either of the prior models. This appears consistent with the 'concentrating reactants' model, but does it clearly rule out the 'polarized unfurling' model?

      Thank you for pointing out what in retrospect was an obvious ‘loose end’ in our discussion. The other referees raised the same point. We made small changes to the introduction and added a paragraph to the discussion to be clearer about how the existing models are or are not consistent with the present results (top paragraph of page 3 and new penultimate paragraph of the manuscript on page 12).

      A second minor weakness is that the comparison to Ena/VASP is not developed at a deep level based on the final model. I found these ideas exciting and want more critical consideration here, but perhaps it is better suited for a commentary piece to follow.

      We appreciate the enthusiasm and understand where this comment is coming from. Because there has been a fair amount of recent movement in the understanding of TOG domains and what they can do, and because some of the mechanistically interesting parallels entail speculation, we agree with the referee’s suggestion that a future commentary will provide a better venue.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Other minor remarks:

      (1) Figure 2 C is missing the label for what should probably be TOG1*-TOG2.

      Fixed

      (2) Figure 5, lower left, is oddly cropped, showing residuals that are slightly distracting from the beauty of the model.

      Apologies that the figure did not look good in the initial submission. We adjusted it and it looks much better in the revised submission.

      (3) The dynamics assay buffer composition stated in the protein purification Method section is not the same as the PEM buffer used for the dynamics assay. And both are different from BRB80, which, with the chambers, are rinsed. This may very well be accurate, but it raises the question of why not stick to one version, as they are virtually the same.

      Thanks for asking these questions, and sorry for the confusion. First, we should have used different names for the (barely) different buffers. This has now been corrected. Second, the PIPES concentration was not 90 mM, it was 100 mM as in our prior work and this discrepancy failed to get caught in proofreading. Why the other small differences? It’s a good question. The differences reflect an arbitrary decision made at the beginning of the work, there is not a deeper rationale.

      (4) Out of curiosity: Why 90 mM PIPES and not 80?

      Why not 80 mM PIPES? This is just a historical difference. The early measurements of yeast microtubule dynamics (e.g. Gupta … Himes MBoC 2002 and Bode … Himes, EMBO Rep 2003) that partly inspired us to use yeast as a model system used 100 mM PIPES as the working concentration, and we never deviated from that.

      (5) Please state the source of PIPES.

      Sorry for the oversight, we have added the source of PIPES (it is Millipore Sigma P6757).

      Reviewer #2 (Recommendations for the authors):

      (1) Page 4, last paragraph, the authors call out "Fig 1C" which I believe should be "Fig 1D".

      We fixed this, thank for catching this error

      (2) Figure 2C: The authors subtract basal tubulin polymerization (growth rate) from the rates measured in the presence of Stu2 constructs. One assumption in doing this is that Stu2 polymerization activity does not compete for the ability of tubulin (not bound to Stu2) to polymerize on the plus end. I think this is a logical assumption, but it would be beneficial for the authors to state this assumption.

      We said this explicitly in the revised submission (first full paragraph on page 7), borrowing from the reviewer’s phrasing.

      (3) Figure 2C: the label for the last bar is missing - presumably: " Stu2 (TOG1*-TOG2)".

      Fixed.

      (4) Figure 2D: Many of the KM values determined are at the border for points measured, or in one case, beyond the concentration of tubulin sampled. In this regard, the authors should discuss how well the fitted curves correlated with their data. Also, as Vmax and KM are calculated, it would be beneficial if another panel were produced (e.g., Figure 2E) in which the data were presented as a Lineweaver-Burk plot. Doing so, the authors would be able to test their enzymatic model by doping the system with their nonpolymerizable tubulin mutants, which should yield competitive inhibitor behavior but not change Vmax.

      Thanks for pointing this out, we should have been more explicit about this point. We incorporated into the results section an explicit statement about this limitation (first full paragraph on page 7). The suggestion to use blocked mutants and Lineweaver-Burke plots is an interesting one that we hope to pursue in future work using blocked or other mechanism-specific mutants. But we think to do so is complicated enough to be beyond the scope of the present study.

      (5) Figure 3: The authors quantitate the amount of Stu2-GFP fluorescence at microtubule plus ends using line scan analysis of the kymograph. Since the kymographs are processed images, it is more appropriate to integrate intensity from the original frames collected using a circular area. i.e., ID points on the kymograph, and return to the respective position in the corresponding frame to calculate background-subtracted GFP intensity at the plus end.

      This is a fair point. We chose to stick with the kymographbased analysis because the symmetry of the point-spread function and lack of rapid variation in GFP and/or background intensity means that the kymograph analysis is adequate for the intensity-based comparison we were doing.

      (6) Page 7, last line, the authors call out "Fig 2D" which I believe should be "Fig 1D".

      Sorry for the error, we have corrected it.

      (7) Figure 4A: The authors discuss the "sortase epitope," but technically, an epitope is the binding site specifically for an antibody, not to be used in general terms for proteinprotein interaction sites. As such, the authors should describe this as the "sortase recognition sequence" or something similar.

      Thank you for noticing this; we had indeed used ‘sortase recognition sequence’ elsewhere in the paper but we did not catch this instance during proofreading. We have now used that same language in the legend for Fig. 4.

      (8) Page 8, the authors state "KM is approximately equal to Kt/Kon and KM negligeable," but I think they mean "...and KD negligeable".

      Thank you for noticing this typo. We corrected it.

      (9) Page 9, first paragraph last sentence: the readership would be aided by modifying the sentence as follows (adding "Kon" and "Kt"): " ... must be kinetically limited by either the rate of TOG:tubulin binding (Kon) or by the rate of TOG-mediated transfer to tubulin to the microtubule end (kt)."

      Very good suggestion, we implemented it.

      (10) Page 9, second paragraph, the authors call out "Fig 2D", but perhaps they intended to call out "Fig 1D"?

      Sorry for the error, the reviewer is correct and we fixed this.

      (11) Page 9, second paragraph: The authors mention that the transfer rate of tubulin to the plus end via a TOG domain is close to the transfer rate of free tubulin to the growing plus end. Can the authors expand on why they are mentioning this comparison?

      Thanks for the push to be clearer about this. The basic idea is that each ‘delivering’ TOG contributes 50% of the background (uncatalyzed) polymerization rate. So the presence of multiple TOGs (in a single polymerase or from multiple end-resident polymerases) can substantially increase the rate of polymerization. We added brief text to try to make this clearer (first paragraph on page 10).

      (12) Page 9, second paragraph: "TOG-TOG2 polymerases" would be better phrased as "TOG2-TOG2 dimeric polymerases". Noting as well that "2" is missing from the first "TOG".

      This is indeed better phrasing and we have adopted it (also corrected the missing ‘2’) (first paragraph on page 10).

      (13) Page 13, BLI methods: The authors should list the final pH for the PIPES buffer (was it pH 6.9 as in the polymerization assay?).

      Sorry for the oversight, we have added the pH and it was indeed 6.9.

      (14) Page 13, BLI methods: What is "LR1-457"?

      LR1-457 is lab-notebook-speak that did not get purged in editing; it refers to the polymerization-blocked tubulin mutant that also carried a sortase recognition sequence. We replaced ‘LR1-457’ with more evocative phrasing and corrected another typo we found there.

      (15) In Ayaz et al. (eLife, 2014) Stu2 TOG1 and TOG2 affinities for tubulin were measured using AUC, for which the fitted curves appeared to correlate with the data quite well. As the authors note, the values were KD = 70 nM and 160 nM, respectively. This contrasts with the BLI measurement for TOG2-tubulin (~10 nM), which suggests that at least one of the experiments was off the mark - or, as the authors do note, that different buffer and salt condition was used could account for the differences, but that the BLI conditions align with the polymerization conditions (though not exactly) and thus are more appropriate to use. In a supplemental discussion, the authors should run the AUC values through the equation for their model and state what types of differences these values could imply for Stu2 mechanism. If the differences are significant for the Stu2 model derived, the authors should give thought as to whether a third assay should be employed to determine TOG-tubulin affinity. Based on the BLI reagents, it appears the authors would be well-positioned to conduct an assay using SPR. As a potential alternative, the authors could repeat the BLI experiment using the buffer conditions from the Ayaz et al., AUC work (25 mM Tris pH 7.5, 1 mM MgCl2, 1 mM EGTA, 100 mM NaCl, 20 μM GTP) - noting that BSA and Triton X-100 may need to be added as well. If the authors are able to replicate the ~160 nM affinity for TOG2:tubulin, this would be a reasonable way to bootstrap to the conclusion that the BLI is measuring affinity correctly and that the current PIPES-based BLI experiments yielded accurate data.

      We tried hard to perform comparative measurements of TOG2:tubulin affinity in the two buffer systems. Unexpected challenges and personnel turnover made this slower than anticipated. The reviewer’s suggestions are completely reasonable, but ultimately it was not possible for us to get side-by-side results for TOG:tubulin affinity using the same measurement technique, whether it was biolayer interferometry, analytical ultracentrifugation, or isothermal titration calorimetry. For each technique one or the other buffer caused aggregation, non-specific binding, or some other artifact that prevented analysis. The fact that we were unable to compare the buffer conditions in this way is stated in the revised manuscript (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added text to state that the affinity of TOG:tubulin interactions have been shown to depend on ionic strength in a way that is consistent with the changes we observed: prior data from the Al-Bassam lab (Nithianantham, 2018) showed that TOG:tubulin affinity decreases ~20-fold with increased ionic strength (100 mM KCl vs 200 mM KCl) (page 9, last paragraph before the ‘Unifying measurements …’ section). We also added some text to address the comment about affinity and whether/when the shuttle model would hold (first paragraph on page 11).

      (16) A sentence or two in the discussion, relating how their data aligns (or not) with the Nithiantham model would be beneficial, especially as discussing the two models in the introduction was a central point.

      We completely agree and have now added a paragraph to the discussion to explicitly address the two models and how are or are not supported by the new model and observations (page 12, penultimate paragraph of the main text).

      (17) Discussion: The model in Figure 5 depicts Stu2 engaged with the microtubule, perhaps using its basic linker region (?). The authors could note this in the figure caption for 5A. The authors do not discuss the basis for plus-end polymerization activity versus polymerization activity at both the plus and minus ends. Do the authors propose that this is due to differential Kt values for the two ends and/or differential localization via the basic region to the two ends?

      Thanks for bring this up. The plus-end selectivity of these polymerases is thought to result from the polarity of TOG:tubulin engagement and the positioning of TOG domains relative to the basic region that provides ‘slippery’ binding to the microtubule lattice. We have partially addressed these issues in the legend to Figure 5 (page 11).

      (18) Brouhard (Cell, 2008) demonstrated that XMAP215 can catalyze the depolymerization of GMPCPP microtubules when no free tubulin, or very low levels of free tubulin, are present. This is interesting in that it indicates that the enzymatic activity is reversible. Can the authors comment on how their model would behave in the low-tozero free tubulin concentration regime? Would a different model have to be invoked?

      This is an interesting comment. Because the enzyme-like model treats the transfer step as irreversible, the model cannot account for the kind of ‘depolymerase’ activity Brouhard and others have noted. A more general model that could also encompass the depolymerase activity at low-to-no free tubulin would need to explicitly model the step(s) involved in microtubule association and dissociation. These steps remain a major open question in the field and while it would be quite interesting, trying to address this in a model is beyond the scope of what we can confidently do given the data we have. To be more explicit about this assumption/limitation, we now point this out in the results section where the model is introduced (page 4, last paragraph).

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1D, legend. "...and a transfer rate constant kf that describes how fast...". Should kf be replaced with kt?

      We made this correction, thanks for pointing the problem out

      (2) Figure 2C. The x-axis label under the blue bar is missing. Also, I find the arrows to the left of the bars confusing and unnecessary.

      We fixed the legend problem. We sympathize with the dislike of the arrows but respectfully prefer to keep them in the hopes of avoiding confusion about the fact we are fitting ‘growth rate attributable to Stu2’, not simply growth rate. The figure legend has been expanded to hopefully smooth this over.

      (3) Figure 3. The kymographs are convincing, but it may be helpful for future studies to state here what the polymerization rates are for 0.6 µM and 1.4 µM yeast tubulin. These values are probably different than what one might expect for mammalian tubulin at those concentrations, and the authors could simply determine them from the slopes in the kymographs.

      Good suggestion, we added the growth rates to the legend (as the reviewer expected, they differ from expectations based on mammalian tubulin).

      (4) Results, page 9, line 11: "...yielded a value of 9.6 nM...". Should this be 8.9 nM, which is that value stated in Figure 4C?

      Actually, these different values are correct. We just wanted to point out that whether we used response amplitudes or measured on- and off-rates, we get very similar values for K<sub>D</sub>. We changed wording to hopefully make this clearer: “Calculating the dissociation constant K<sub>D</sub> from the measured rate constants (K<sub>D</sub> = k<sub>off</sub>/k<sub>on</sub>) instead of from the amplitudes yielded a value of 9.6 nM, in good agreement with the amplitude-based determination of 8.9 nM.” (page 9).

    1. eLife Assessment

      This valuable study explores changes in the Drosophila microbiome in response to environmental temperature over more than ten years. The evidence that temperature leads to diversification of bacterial clades is solid, despite the need for greater clarity in defining and tracking strain competition. The work will interest researchers working with microbiomes, microbial ecology, and evolutionary biology.

    2. Reviewer #1 (Public review):

      Summary:

      The factors that create and maintain diversity in host-associated microbiomes remain poorly understood. A better understanding of these factors will help in the efforts to leverage the adaptive potential of the microbiome to help solve pressing problems in health and agriculture.

      Experimental evolution provides a promising path forward as we can track the causes and consequences in the emergence of novel variants, but experimental evolution remains underutilized in host-microbiome interactions. Here, Gracia-Alvira utilizes a long-term experimental evolution study in Drosophila simulans under hot and cold temperature regimes to identify strain-level variation in an important fly bacterium, Lactiplantibacillus plantarum. They identify three strains of L. plantarum, which are most prevalent in their respective three temperature regimes, suggesting that these are locally adapted bacteria. Then, using a combination of genomics, in vitro, and in vivo, Gracia-Alvira et al attempt to understand the factors that led to the differentiation of the hot and cold L. plantarum and their impacts on the fly host.

      Strengths:

      This is an excellent use of experimental evolution to track the emergence of novelty in the microbiome. The genomic analyses are all solid and appropriate for the data sets. It is especially striking that the comparisons with the other, independent experimental evolution studies in different labs (and across continents between Portugal and South Africa) show a consistent response to temperature. Many have disregarded the microbiome as it is something that is too sensitive to seemingly innocuous variables (particularly in the fly microbiome), such that we cannot find generalities. However, this finding highlights the potential for experimental evolution to uncover these dynamics. The question of how strains emerge and are maintained is timely and is one of the key open questions in host-microbiome evolution currently.

      Comments on revised version:

      I thank the authors for their thoughtful responses to my concerns, and I appreciate the additional experiments to help resolve the questions about subspecies competition. The manuscript remains strongest in the genomic assessment of changes in the L. plantarum genomes, and it is striking and noteworthy that the isolates across multiple countries but same temperature conditions group together phylogenetically.

      I appreciate the additional clarity also incorporated in this revision, but there are still a few key concerns that are unresolved about the microbial ecology described here. I will also note that I apologize if I missed something in the text as no line numbers were provided to point me to where the changes were incorporated in the revised manuscript.

      (1) Competition has many different meanings and many different measurements (see Hart 2018 https://doi.org/10.1111/1365-2745.12954) -and incorporating the effects of competition in shaping an ecological community is, has been, and will continue to drive much research in community ecology. Measuring strain level competition is one of the major questions in host-associated microbiomes, and it is difficult-though there have been significant advances in doing so (see isogenic barcodes, e.g., Daniel 2024 doi: https://doi.org/10.1038/s41564-024-01634-9b, Ordon 2024 https://doi.org/10.1038/s41564-024-01619-8, as well as my previous suggestion to track the outcomes of competition). The inability to directly track and measure competition of the isolates remains a limitation of this manuscript. The authors' explanation of measuring competition is unusual, simplistic, and at times inconsistent.

      They need to be crystal clear about their definitions, logic for making these inferences, and weaknesses in their approach. I think what the authors mean is that competition between the unevolved and C or H in their respective regimes leads to the decrease of the U clade over experimental evolution. But it is not clear how the authors are thinking about competition between C and H clades in the different temperatures.

      The authors state that competition is inferred because changes in relative abundance across the time series-and this is unusual because there are alternative explanations that require no ecological interactions among sub-strains, as I described in my comments on the prior version. This is then combined with in vitro work that shows that the H and C clades can both grow in their mismatched temperature regimes-and thus I think it is to be inferred that because they can grow alone in vitro (and C isolates show lower growth than H isolates in hot temperature), then changes in the relative abundance over fly generations can be attributed to competitive interactions among C and H clades. But then the logic is inconsistent because then the authors just say that in vitro growth curves don't support the differences in relative abundance observed in the flies (lines 224-225). Then the authors argue is it about a combination of diet/sugar metabolism and temperature (line 373), which doesn't make any sense because temperature previously didn't matter (lines 224-225).

      All of this is to say is that the authors need to make clear their logic to the readers-and explain these inconsistencies appropriately. To me, it suggests that there are clear methodological weaknesses that inhibit the ability to track competitive microbial dynamics. Because you can't really assess the microbial dynamics in vivo, it remains further unresolved why clade C isolates have such strong negative fitness effects on the fly but reach such high relative abundances in the C evolving flies. I find that this series of logical inconsistencies (and see my point #2) distracts from the important finding that the C and H clades evolved to utilize sugars differently from the U clade, which is an interesting finding!

      (2) There are also inconsistencies in the patterns observed between the text and the figures. Some of this arises because the authors are not clear what comparisons they are making. For example, line 450 says that clade C outcompeted the other clades, which I presume means only in the cold temperature. Line 456 says that C and H isolates grow faster in the sugar-rich lab diet, but that is not really true because U and C have similar growth rates in Fig. 5, and U and H have similar growth rates in Fig. S4. The text about microbial load is a bit misleading (lines 271-273), as it is confusing that clade C is significantly higher load in both hot and cold temperatures (Fig. S6), which is counterintuitive given Fig. 4, 5, S4. But it is also overly speculative to say that these results suggest that fitness effects depend on microbial load of clade C without connecting the load to the fly fitness measures (and also given the inconsistency with the time series data from evolving lines). Please take care to more carefully phrase these statements to ensure the inference is supported by the experiment design (e.g., clarifying comparison) and statistics (e.g., ensuring agreement with what the figure shows).

      (3) I understand the concern about focusing the reader on the L. plantarum strains. However, it should be clear to the readers that you did not examine the other parts of the microbiome, and that L. plantarum is often very rare in lab and wild fly populations. The data presented on Table S4 (cited line 552, I think citation at line 176 is incorrect) is confusing. If these were colonies picked and then identified, this should be explicit. If it is based off on colonies, then please clarify if this was sampled randomly or occurred when trying to enrich/focus on L. plantarum isolates. If the data was computational (e.g., Kraken to classify), then only taxa richness is not necessarily relevant, but please also include to the relative abundance of each taxa.

      To me, this is relevant information to contextualize these results, particularly because you test this in both D. mel and D. simulans (apologies for the confusion over Mazzucco & Schlotterer 2021), and we have insight into how combinations of Lactobacillus and other taxa impact fitness (Gould PNAS 2018). If the results from D. melanogaster are not applicable to D. simulans, then the authors need to explain this. I understand if incorporating analysis of the broader microbiome is beyond the scope of this manuscript, but at least acknowledging the general rarity in Lactobacillus frequency in Drosophila microbiome and variation in fitness effects will more accurately contextualization these results.

      One small point is that line 452 the citations are OK, but there are fly-specific examples to support this statement, like Gould PNAS 2018, Henry Proceedings B 2025.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gracia-Alvira et al. investigated how environmental temperature affects competition among members of the microbiome, with a focus on intraspecific diversity, using the Drosophila model.

      Notably, the authors identified three clades of Lactiplantibacillus plantarum from a natural population of Drosophila simulans collected in Florida. They tracked the dynamics of these three bacterial clades under two temperature conditions over the course of more than ten years. Using comparative genomics and phylogeny, they showed that these three bacterial clades likely adapted to their host independently in a temperature-specific manner. Further, by combining in vitro culture and in vivo mono-association assays, they demonstrated the functional divergence of these three bacterial clades phenotypically, including their growth dynamics and effects on host fitness. Lastly, they performed pathway analysis and speculated on key genomic variance supporting such functional divergence.

      Strengths:

      The laboratory evolutionary experiment in response to cold or hot environmental temperature is impressive, given its more than ten years of experimental time period. This collection of achieved microbiome samples paired with the fly host data can be a valuable resource for the field.

      Comments on revised version:

      The revised version has addressed my major points raised in the original review.

    4. Reviewer #3 (Public review):

      Summary:

      The study presents an analysis of 297 pangenomes derived from 20 populations of Drosophila simulans, at 19 time points for fast-reproducing individuals in a hot environment, or at 10 time points for slow-reproducing individuals in a cold environment, over a period of more than 10 years. The authors select a particular microbial component of the pangenomes and study the dynamics of Lactiplantibacillus plantarum strains in two environments. They discover that the revealed operational taxonomic units could be divided into three phylogenetic clades, which have their own genomic and genetic features, different adaptive capabilities that depend on the environment, and have a distinct impact on the fitness of the host.

      Strengths:

      The authors prove that bacterial microbiome components are sensitive to the environment and could rapidly (years) be fixed in eukaryotic populations. This study establishes a tractable model that potentially enables the study of variability of the physiological influence of distinct strains of an important commensal species, Lactiplantibacillus plantarum, on the Drosophila host. It is clearly shown that this single species consists of several phylogenetically and functionally diverse strains. The authors did not limit their interest to their own model, but rather they have integrated a comparative approach by analysing phylogenetic relationships among 92 described L. plantarum strains.

      Overall, the study is novel and delivers important discoveries of a longitudinal, well-replicated experiment, generating a substantial amount of genomic data. It highlights an important dimension of research that environmental selection operates at the subspecies level.

      Weaknesses:

      Even though the authors show only one particular example by conducting their longitudinal experiment, they honestly acknowledge failures important for interpretation of the biological significance of the results (gnotobiotic mono-association experiments was done with D. melanogaster, but not D. simulans) and therefore they state limitations of their conclusions (weaker effects in the non-axenic flies are due to the presence of other taxa or to higher-order interactions with other members of the microbiome). These interactions could significantly affect bacterial growth, metabolism, and physiological influence on the host.

      The authors exploit the results of their experiment to speculate about a wide range of evolutionary phenomena, like within-species competition, ecological adaptation and evolution of the host, fitness advantage of bacteria to the host, the benefits of parasitism or mutualism, the domestication of the microbiome, etc. At the end, they conclude that their study "highlights that even subspecies diversity plays a key role in adaptation to environmental temperature". However, the potential mechanisms of such adaptation are barely discussed, so that the focus of the study shifts from the temperature-induced changes in microbial population structures toward metabolism-related adaptations of clade representatives that enable them to diversify their carbon and nitrogen sources. The role of the temperature factor remains elusive.

      In addition to that, the paper has a clearly minimalistic experimental approach to address functional properties of the revealed L. plantarum strains, so that their own fitness, or their relationship with the Drosophila host, is characterised superficially. Therefore, the authors' discourse can be speculative rather than factual (especially when the authors use the expression "likely" to share their guesses in the "Results" section). Nevertheless, these minor drawbacks do not underscore the novelty of the discovered phenotypes and the importance of their further investigation.

      Comments on revised version:

      I have read the authors revisions and find them compelling and they address fully the minor points raised in my review.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The factors that create and maintain diversity in host-associated microbiomes remain poorly understood. A better understanding of these factors will help in the efforts to leverage the adaptive potential of the microbiome to help solve pressing problems in health and agriculture.

      Experimental evolution provides a promising path forward as we can track the causes and consequences in the emergence of novel variants, but experimental evolution remains underutilized in host-microbiome interactions. Here, Gracia-Alvira utilizes a long-term experimental evolution study in Drosophila simulans under hot and cold temperature regimes to identify strain-level variation in an important fly bacterium, Lactiplantibacillus plantarum. They identify three strains of L. plantarum, which are most prevalent in their respective three temperature regimes, suggesting that these are locally adapted bacteria. Then, using a combination of genomics, in vitro, and in vivo, Gracia-Alvira et al attempt to understand the factors that led to the differentiation of the hot and cold L. plantarum and their impacts on the fly host.

      Strengths:

      This is an excellent use of experimental evolution to track the emergence of novelty in the microbiome. The genomic analyses are all solid and appropriate for the data sets. It is especially striking that the comparisons with the other, independent experimental evolution studies in different labs (and across continents between Portugal and South Africa) show a consistent response to temperature. Many have disregarded the microbiome as it is something that is too sensitive to seemingly innocuous variables (particularly in the fly microbiome), such that we cannot find generalities. However, this finding highlights the potential for experimental evolution to uncover these dynamics. The question of how strains emerge and are maintained is timely and is one of the key open questions in host-microbiome evolution currently.

      Weaknesses:

      (1) The framing in the title and throughout the discussion about "subspecies competition" does not match the data that was collected. The subspecies competition requires actually tracking the competitive outcomes between the hot, cold, and unevolved L. plantarum. In the in vivo work, I can see that mixes of the strains were made, but they did not track whether the cold strain outcompeted the hot strain in vivo under cold conditions, for example.

      We thank the reviewer for the honest concern and take this opportunity to defend our claim of "subspecies competition used across the manuscript. As the reviewer states, subspecies competition requires tracking the competitive outcomes between the three clades, and this is what we did by sampling and sequencing across ten years of experimental evolution (Figures 4 and S3). For this reason, we point that the subspecies competition assessment comes from the direct observation of changes in relative abundance across the time series, and not from the follow-up experiments in vivo or in vitro.

      While Figure 4 is suggestive that there is ongoing competition in the hot temperature regime, this is not necessarily shown in the cold, which is dominated by the C clade. It could also be that the bacteria cannot survive in the flies at the different temperatures. The growth curve assays hint that the bacteria can grow, but the plate reader couldn't actually maintain the 18 {degree sign}C temperature (line 455). So all of this evidence is very indirect and insufficient to say that strain competition is driving these patterns.

      We thank the reviewer for the alternative hypothesis that could explain the observed subspecies dynamic. We rule out that dominance of clade C in the cold occurs because the other two clades cannot grow in this regime based on three pieces of evidence:

      (1) In the time series, clades H and U decrease, but never disappear (Figures 4 and S3), even showing some peaks of abundance in specific replicate populations (Figure S3).

      (2) We isolated individuals belonging to clade H in the cold-evolved populations, as shown in figure 2. This is a direct evidence that clade H prevails in the cold-evolved populations, although in low abundance.

      (3) We did grow the three taxa in fly food Petri dishes incubated at both temperature regimes, observing growth in all cases.

      We will include the food growth experiment in the revised manuscript as further supporting evidence for growth in both regimes.

      (2) The in vivo results are interesting in that there appears to be a fitness cost of clade C, but the explanation is underdeveloped. I say under-developed because in Figure 4, the cold L. plantarum remains much higher throughout adaptation to the hot temperature regime than the hot L. plantarum in the cold regime. The hot L. plantarum is low abundance throughout the cold regime. I felt like this observation was not explained, but it seems relevant to understanding the strain dynamics.

      We acknowledge that a strong fitness cost of clade C is observed in axenic D. melanogaster. In the native host, D. simulans, with reduced microbiome, we observed delayed development that could even be an advantage depending on the situation, as pointed out by reviewer 3 in the recommendations.

      Even if we assume that flies colonized with clade C are less fit in the experimental evolution, another caveat is whether the flies can actively select for the L. plantarum clade. Under this assumption, a clade that imposes a fitness cost to the fly (clade C) should be selected against over time because the flies colonized by this clade will have less offspring or develop later than the rest. Alternatively, as the microbiome is shared among all the individuals in the population, the host might not be able to “purge” the pernicious clade, and L. plantarum dynamics might be controlled solely by the relative fitness between clades in the given experimental treatment. We will discuss this hypothesis in the revision as a way to explain the relationship between the abundance of each clade and the effect on the host.

      I will also note that this is not the first time that L. plantarum or other Lactobacillus have been shown to exert fitness costs to Drosophila. Gould, PNAS, 2018, shows that both Lactobacillus plantarum and Lactobacillus brevis in mono-association have lower fitness (measured through Leslie matrix projections using lifespan and fecundity) than axenic flies. Many studies of wild Drosophila fail to find Lactobacillus, or it is low abundance (e.g., Chandler, PLoS Genetics, 2014; Wang, Environmental Microbiology Reports, 2018; Henry & Ayroles, Molecular Ecology, 2022; Gale, AEM, 2025). This might help provide useful context for the in vivo results.

      We thank the reviewer for the references. These observations are compared to our phenotypic results and discussed in the revised version of the manuscript.

      (3) The data in Figure 4 are compelling to focus on the L. plantarum variants. However, I can see from the methods that the competitive mapping included only other strains of Wolbachia.

      We appreciate the thorough reading of the methods by the reviewer. The competitive mapping comprised two steps: first we discarded the reads that mapped to Drosophila, Wolbachia and additional potential contaminants from sequencing facitilies (human, dog...). This step leaves the reads originated from whole the external microbiome of the flies, including L. plantarum. The second competitive mapping step recruits the reads that map any clade of L. plantarum.

      It is not clear how other members of the microbiome changed in response to the temperature regimes. As I note in point #2, given that Lactobacillus is often rare, it is not clear what the rest of the microbiome looks like over the course of adaptation. Indeed, it seems like Mazzucco & Schlotterer, PRSB, 2021 did a broader analysis of the microbiome and found that Acetobacter is by far the most common bacterium (I think this data is also part of the data shown here?). Expanding on why or why not in this context is important and will improve this study, particularly if the focus is on connecting these evolutionary dynamics to ecological competition to explain the emergence of strain diversity.

      We acknowledge that the rest of the Drosophila microbiome is not addressed in this study, as we wanted to focus the storyline around the intraspecific dynamics found in L. plantarum. We consider that a complete characterization of the whole Drosophila microbiome would unnecessarily elongate the paper and thus we treat it as a constant biotic factor.

      We must point out that our dataset is not the one reported by Mazzucco & Schlötterer, which was done in D. melanogaster, rather than D. simulans. Nevertheless, both experiments share the same infrastructure, temperature regimes and fly maintenance.

      We have included a list of taxa that were isolated from the populations, as well as to report L. plantarum prevalence and abundance across the experiment in order to provide context of the microbiome, beyond L. plantarum, to the readership.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gracia-Alvira et al. investigated how environmental temperature affects competition among members of the microbiome, with a focus on intraspecific diversity, using the Drosophila model.

      Notably, the authors identified three clades of Lactiplantibacillus plantarum from a natural population of Drosophila simulans collected in Florida. They tracked the dynamics of these three bacterial clades under two temperature conditions over the course of more than ten years. Using comparative genomics and phylogeny, they showed that these three bacterial clades likely adapted to their host independently in a temperature-specific manner. Further, by combining in vitro culture and in vivo mono-association assays, they demonstrated the functional divergence of these three bacterial clades phenotypically, including their growth dynamics and effects on host fitness. Lastly, they performed pathway analysis and speculated on key genomic variance supporting such functional divergence.

      Strengths:

      The laboratory evolutionary experiment in response to cold or hot environmental temperature is impressive, given its more than ten years of experimental time period. This collection of achieved microbiome samples paired with the fly host data can be a valuable resource for the field.

      Weaknesses:

      The laboratory evolutionary experiment can be limited due to its artificial experimental setup. For example, wild flies rely on a more diverse set of food sources and are constantly exposed to new bacterial inoculations, whereas under laboratory conditions, flies live in a more restricted ecosystem. In addition, environmental temperatures differ among different locations, but they also involve seasonal changes within the same region. This manuscript can be strengthened with further discussions that elaborate on these limitations.

      As the reviewer has correctly noted, our experimental setting is not exempt from limitations. Lab-reared flies are fed with a defined standard diet. Furthermore, although the system is not completely closed to bacterial migration, this is limited as replicate populations are not allowed to mix during the maintenance of the flies. For this reason, we consider our laboratory setting as a compromise between observing wild populations, which undergo all biotic and abiotic stresses but cannot be manipulated, and evolving the bacteria in absence of the host, or in gnobiotic hosts, in which biotic interactions are not fully considered. We will extend on this in the new version of the manuscript.

      Moreover, the extent of host effects involved in these experiments remains ambiguous, because it is unclear whether these Lactiplantibacillus plantarum mostly reside within fly guts or on Drosophila medium. The laboratory evolutionary experiment possibly favored better colonizers on Drosophila medium under either cold or hot temperatures, which subsequently can saturate fly guts. As fully dissociating these variables can be experimentally tedious, the authors may want to comment more on these aspects in the discussion. Or they may want to consider some measurements. For example, measuring the growth rate of these bacteria on Drosophila medium under different temperatures, in addition to the current MRS culture experiments, or measuring the portion of the Lactiplantibacillus on Drosophila medium versus these stably colonizing fly guts.

      The reviewer's point was briefly addressed in the Results chapter: "Phenotypic differences in liquid culture".

      Reviewer #3 (Public review):

      Summary:

      The study presents an analysis of 297 pangenomes derived from 20 populations of Drosophila simulans, at 19 time points for fast-reproducing individuals in a hot environment, or at 10 time points for slow-reproducing individuals in a cold environment, over a period of more than 10 years. The authors select a particular microbial component of the pangenomes and study the dynamics of Lactiplantibacillus plantarum strains in two environments. They discover that the revealed operational taxonomic units could be divided into three phylogenetic clades, which have their own genomic and genetic features, different adaptive capabilities that depend on the environment, and have a distinct impact on the fitness of the host.

      Strengths:

      The authors prove that bacterial microbiome components are sensitive to the environment and could rapidly (years) be fixed in eukaryotic populations. This study establishes a tractable model that potentially enables the study of variability of the physiological influence of distinct strains of an important commensal species, Lactiplantibacillus plantarum, on the Drosophila host. It is clearly shown that this single species consists of several phylogenetically and functionally diverse strains. The authors did not limit their interest to their own model, but rather they have integrated a comparative approach by analysing phylogenetic relationships among 92 described L. plantarum strains.

      Overall, the study is novel and delivers important discoveries of a longitudinal, well replicated experiment, generating a substantial amount of genomic data. It highlights an important dimension of research that environmental selection operates at the subspecies level.

      Weaknesses:

      Even though the authors show only one particular example by conducting their longitudinal experiment, they honestly acknowledge failures important for interpretation of the biological significance of the results (gnotobiotic mono-association experiments was done with D. melanogaster, but not D. simulans) and therefore they state limitations of their conclusions (weaker effects in the non-axenic flies are due to the presence of other taxa or to higher-order interactions with other members of the microbiome). These interactions could significantly affect bacterial growth, metabolism, and physiological influence on the host.

      We agree with the reviewer in that the use gnobiotic animals is a limitation, as by "tuning" the flies' microbiome we are modifying the interactions between members, which can potentially change the phenotypic outcome. Nevertheless, we use it as a complementary approach, rather than the only inference in our study.

      The authors exploit the results of their experiment to speculate about a wide range of evolutionary phenomena, like within-species competition, ecological adaptation and evolution of the host, fitness advantage of bacteria to the host, the benefits of parasitism or mutualism, the domestication of the microbiome, etc. At the end, they conclude that their study "highlights that even subspecies diversity plays a key role in adaptation to environmental temperature". However, the potential mechanisms of such adaptation are barely discussed, so that the focus of the study shifts from the temperature-induced changes in microbial population structures toward metabolism-related adaptations of clade representatives that enable them to diversify their carbon and nitrogen sources. The role of the temperature factor remains elusive.

      We acknowledge that our study does not fully resolve the mechanism by which a different clade ends up dominating each temperature regime. The MRS liquid experiment was an attempt to answer whether differences in optimal growth temperature could explain the temperature-specific abundance of the two clades. Our experiments showed, however, that this was not the case. Beyond this point, it is hard to disentangle the role of the temperature, as it could also act indirectly on the bacteria, for example, through the host or the food.

      A second observation in our time series was that a third clade, U, was unfit in both regimes despite starting the experiment in high abundance. For this reason we also studied what made this clade less fit. Based on our analyses, we propose that the decrease of clade U was driven by the shift to a laboratory diet, shared by all experimental populations.

      In addition to that, the paper has a clearly minimalistic experimental approach to address functional properties of the revealed L. plantarum strains, so that their own fitness, or their relationship with the Drosophila host, is characterised superficially. Therefore, the authors' discourse can be speculative rather than factual (especially when the authors use the expression "likely" to share their guesses in the "Results" section). Nevertheless, these minor drawbacks do not underscore the novelty of the discovered phenotypes and the importance of their further investigation.

      We consider the reviewer's concern and toned down the phrasing when reporting our findings in the revised version of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) One solution to resolve the "competition" issue would be to check that the L. plantarum strains are established at similar or different titers in the in vivo work. Fly phenotypes can be sensitive to microbial load (Keebaugh, iScience, 2018), which might explain some of the counterintuitive in vivo results. In line 227, the authors mention that "bacterial load" is contributing to the magnitude of the effect, but I don't see the data reported anywhere. If this is from the in vitro assays, then the authors need to show that in vitro predicts in vivo L. plantarum abundance.

      Bacterial load inoculated in the in vivo experiment was normalized to OD=0.05 (~5*10<sup>6</sup> CFUs/ml) for the three clades at the beginning of the experiment. Thus, all vials were inoculated with the same titer of L. plantarum. Only the genotype varied between treatments. However, it is possible that, once inoculated, each clade grew to a different titer (as they have different growth rate and carrying capacity).

      Statement in line 227 comes from the differing results in transfers 1 and 2. In transfer 1 we inoculated a fixed load of ~2.5*10<sup>5</sup> CFUs. In transfer 2, however, we inoculated no bacteria to the food, and the flies seeded the vial. Our statement comes from the assumption that bacterial load in transfer 2 has to be lower than in transfer 1 as bacteria seed the vial solely by defecation of the parents.

      Following the reviewer's suggestion, in the revised version of the manuscript we have included a new experiment in which we quantified the bacterial load of each clade in individual flies.

      (2) Tracking the competitive outcomes is tricky, though it could be done with whole genome sequencing. An alternative would be to label the strains with fluorescent proteins (e.g., Obadia Current Biology 2018 has done this in Lactobacillus) and track fluorescence to better understand the results of the "mix" treatment in Figure 6.

      We appreciate the feedback of the reviewer, but consider this rather labour-intensive approach as an interesting option for future follow-up experiments.

      (3) That being said, my main concern with this is the "competition" claim. If the paper were reframed appropriately, this paper could still make an important contribution to the evolution of host-microbiome interactions, but the authors would need to consider what they can and cannot do with this interesting dataset.

      The "competition" claim comes from the changes in relative abundance observed in the time-series data, not from any of the follow-up experiments. Thus, we consider the use of the term "competition" appropriate.

      (4) The text on the figures is very small and hard to read.

      We increased the size of the text in all figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Have you conducted the in vitro culture experiments following the "cold" conditions?

      We have conducted the experiment in "cold" conditions, but with some modifications to the experimental settings, as the plate reader did not have cooling capacity. Instead, we grew a subset of the isolates (four per clade) in glass vials at constant 20 °C, and measured their OD twice a day. We have included the results in the revised version of the manuscript.

      (2) How many technical and biological replicates were measured for the in vitro culture experiments (Figure 5)? Please add this information to the figure legend and method.

      We measured the growth of four isolates from clade U, nine isolates from clade C and sixteen isolates from clade H. Each isolate was grown three times.

      We have included this information, as requested by the reviewer.

      (3) Making the labels in Figures 2, 3, 5, and 6 bigger would be helpful.

      We have increased the font size of all figures.

      Reviewer #3 (Recommendations for the authors):

      (1) Line 268: "Based on our results in experimentally evolved fruit flies, we propose that within-species competition, thus far largely overlooked, could contribute to ecological adaptation and evolution of the host". Overstatement should be avoided, since the evolution of the host was not directly studied here.

      Our results show that reproductive traits of the host differ upon colonization with each clade. Although we don't test the host's evolution, we speculate that flies differing in their offspring number and developmental time might differ in their overall fitness. Finally, we consider the Discussion section as the right place for speculation and development of hypotheses that can be tested in future work.

      (2) Line 258: "These differences do not explain the clade-specific selection, but reflect the different evolutionary histories of the clades". The temperature factor and its possible role in clade selection would be better discussed at least a little bit.

      In this paragraph we described potential metabolic differences between clades using comparative genomics. We did not find enrichment in a function or group of functions that could explain the different dynamics between clades H and C in the temperature regime.

      In the revised version of the manuscript we highlight that we did not find temperature-specific differences from this analysis.

      (3) Line 252: "...This could explain why clade U, which displayed a high growth rate and carrying capacity in liquid culture". The statement could be further developed with a caution. Even if the isolates that belong to the clade U are outcompeted by H or C, it should be noted that the strain U cannot be used as a true reference for fitness, since it could possess its hidden adaptive properties, not being simply "a loser". Such a hypothesis could explain the maintenance of this strain in the wild.

      We agree with the reviewer in that fitness is relative to the selective environment. Clade U is less fit than H and C in our specific experimental conditions, but it could outcompete them in other conditions, such as wild flies or MRS liquid medium. In the revised version of the manuscript we have rephrased this statement to clarify that we specifically refer to clade U's fitness under the new laboratory conditions.

      (4) In a cold environment, association with the clade C induces developmental delay and produces less progeny, which potentially allows the host to survive in case of harsh conditions and potential food limitation. Could the authors speculate and not exclude that this phenotype could be potentially adaptive? It would be curious to check in further studies whether flies associated with C strains are more stress-resistant, for example.

      We thank the reviewer for this alternative hypothesis. In our manuscript we used the Darwinian definition of fitness; reproductive success of an organism in the focal environment. And thus, both higher progeny per female and shorter developmental time would be beneficial in direct competition with other individuals. It is true that delayed developmental time, or less progeny, could be advantageous in specific cases. This could be the case for D. simulans inoculated with clade C. However, we consider that the fecundity levels observed in D. melanogaster upon inoculation with Clade C (average of 0.06 offspring/female*day in the cold) are too low to sustain a population.

      We have included this hypothesis in the Results section.

      (5) It would be highly recommended to add an experiment to complete the story by measuring the quantity of bacteria in the medium and in the flies. This will resolve the hypothesis (Line 785): "Thus, the ability to exploit this ubiquitous source of carbon and nitrogen could be very advantageous in the fly microbiome context, but would not affect the fitness in liquid culture".

      Following the reviewer's recommendation we included two additional experiments. We measured the bacterial load per fly in the native host, D. simulans, inoculated with the three clades. We also compared the clades' growth speed in solid fly food (without host). In the former experiment, we found similar bacterial loads upon inoculation with clades U and clade H. In contrast, in the latter we found delayed growth of clade U relative to H and C in the food. Thus, chitobiose consumption does not seem to provide an advantage in the fly gut to clade H. We attribute the fitness advantage of H and C to their advantage growing on the laboratory fly food, regardless of the host.

      Both experimental results have been included in the revised version of the manuscript, and the comparative genomics paragraph and discussion have been modified in consequence.

      (6) The chapter "Extended clade-specific differences in KEGG metabolic pathways" could be presented in the main text as it contains important results. These results are mentioned in the chapter "Functional divergence on the genomic level", which looks rather humble when it stands alone as it currently does.

      We appreciate the interest of the reviewer in this supplementary chapter. To keep the length of the manuscript digestible for a broad set of readers, we decided to only include in the main text the functional differences that could play a role in adaptation to the new laboratory environment.

      We consider that a full description of the metabolic differences between the three clades has to be published, as it might be relevant for researchers interested in L. plantarum metabolism. However, it does not fully follow the storyline, as the differences reported in the supplementary, such as nitrate respiration or synthesis of molybdenum cofactors, might not be involved in the clade-specific selection observed in the time series.

      (7) Line 773: "Clades C and H encode a shared genetic repertoire related to sugar/riboflavin metabolism that is lacking in clade U". This indeed allows us to hypothesise that the fixation of these clades in fly populations was due to their improved metabolic capabilities. However, the analysis of fitness shows similarity in flies associated with clades H and U, meaning that sugar/riboflavin metabolism in H does not provide an obvious adaptive trait to flies. Moreover, one could say that sugar metabolism in clade C is maladaptive not only for flies, but also for bacteria in liquid cultures. It is recommended to more clearly state the respective limitations of the study.

      Here we have to make a distinction between bacterial fitness and host fitness. The three clades differ in their (bacterial) relative fitness, as evidenced by the time-series dynamics (Figure 4). In the cited statement we hypothesize that a more versatile sugar metabolism repertoire could increase the bacterial fitness of clades H and C (relative to clade U) in the sugar-rich laboratory diet.

      This is independent of the fitness effect that L. plantarum could have in the host. Finally, as it was discussed in the recommendation 3, fitness is specific to the environment. Clade C is the least fit in liquid MRS in hot conditions, but the fittest in cold experimental conditions.

      (8) The authors should better explain why growth in MRS was not performed in a cold temperature regime to further support or refute the hypothesis that capacity and inflection time could partially explain the higher fitness of bacterial strains from clade U.

      We did not perform this experiment in cold conditions due to technical limitations of the plate reader, that does not have cooling capacity. Nevertheless, following the reviewers' suggestion, we have included in the revised version of the manuscript a new MRS growth experiment in cold-like conditions (constant 20 °C).

      (9) When mentioning that L. plantarum can "increase larval fitness of Drosophila melanogaster relative to germ-free flies" (line 196), the authors should specify in which specific conditions this phenotype was observed, and how relevant the mentioned phenotypes are to the current study.

      Following the reviewer's recommendation, we have modified the paragraph in order to clarify the conditions used in other papers and those used in our work. The references cited in this section (PMID: 21907145, 29290388 and 28062579) report that L. plantarum increases the host fitness in protein-poor diets (12 g/l of dried yeast or less), but not in high-protein diet (50 g/l of yeast or higher). Since our experimental diet contains an intermediate amount of protein (24.3 g/l of dried yeast) we were agnostic of whether L. plantarum would benefit the host or not in our conditions. Regarding the phenotypes, we chose two reproductive traits that are affected by changes in the microbiome according to the literature. Developmental time is directly affected by L. plantarum in the aforementioned papers. Offspring number is another fitness component affected by Drosophila microbiome (PMID: 30510004).

      (10) Provide a reference for line 205: "In axenic D. melanogaster none of the L. plantarum clades provided a fitness advantage to the host relative to germ-free controls, contrary to the effects reported in the literature". If the conditions were different from those in the studies referred to, then it would be of no use to compare the fitness advantage (for example, in Reference 24 another type of diet was used).

      Already covered in recommendation 9.

      (11) Please provide more context to this statement (Line 210): "The high content of dried yeast 24.3 g/l in the fly food used in our experiment likely provided already sufficient amounts of essential amino acids, which negated the growth-promoting effects of L. plantarum". It is not clear why amino acids are taken into account, and what the evidence is for the fact that the amount of essential amino acids was sufficient to abolish growth-promoting effects.

      The whole paragraph was modified in order to clarify the relationship between protein input and nutritional fitness benefit of L. plantarum.

      (12) Please provide measurements of bacterial quantity which would support the statement (Line 215): "The fitness reduction was stronger in the first transfer of flies, likely due to a higher bacterial load".

      Upon request of the reviewer, we have estimated the bacterial load per individual fly in D. simulans. Additionally, we have specified the CFUs inoculated in the vials in transfer 1.

      (13) Correct the typo (line 220): "However, the developmental time was significantly extended after inoculation with clade C at cold temperature (Dunn's test, p < 0.05 05 for all significant comparisons)".

      Done.

      (14) Specify more precisely the temperature conditions referred to in line 226: "In summary, we observed that clade C, which is dominant in the cold-evolved populations, decreases host fitness when axenic flies are inoculated". Does it decrease fitness both in hot and cold environments?

      For the axenic flies, we did find a decrease in fitness in both regimes, yes. We specified it in the revised version of the manuscript.

      (15) Please provide evidence for line 227, or otherwise rephrase it: "The magnitude of this effect varies depending on the environmental temperature, the bacterial load, and the presence of other microbial taxa".

      Novel evidence was provided regarding the role of bacterial load on host fitness.

      (16) Correct the following statement, so that it reproduces the results of the original work (reference 19, line 229): "In a low-protein diet, strains that were not isolated from Drosophila enhanced larval growth relative to germ-free individuals, whereas another Drosophila-associated strain did not have any effect".

      This statement was removed from the revised version. This reference was cited in the discussion to state that: " the nutritional symbiosis in L. plantarum is strain-specific".

      (17) Please provide a rationale for using KEGG Orthologs. Why was this database chosen as an appropriate one, even though it is known to be a non-exhaustive metabolomic resource?

      KEGG is a well-known metabolic database that is widely used in comparative genomics (PMID: 40177264) and built in state-of-the-art software for microbial ecology such as Anvi'o (PMID: 33349678). Other similar gene-to-function databases are less focused on metabolic pathways, such as COG or GO, or limited to specific enzymatic activities, like CAZy. Furthermore, the hierarchical organization of KEGG Orthologs in modules and pathways allowed us to map clade-specific orthologs to the broad metabolic context. For these reasons, we considered KEGG to be the best option for this analysis.

      (18) Line 250: "Therefore, we speculate that the ability to exploit this ubiquitous source of carbon and nitrogen in the lab-maintained fruit flies, could be a strong target of selection in the lab environment". This statement concludes the "Results" section but would be more appropriate for the Discussion section, since the authors do not provide any experimental evidence that could support this statement.

      We have modified this chapter, as covered in recommendation 5.

      (19) Line 276: "However, the intraspecific richness of L.plantarum in our flies was three times higher than that estimated in human gut microbiomes". Note that there are other recent studies which show the presence of several OTUs within L.plantarum isolates (for example PMID: 41484402).

      We thank the reviewer for the reference. We comment on it in the revised manuscript.

      (20) Line 287: "Our finding shows that the well-characterized nutritional symbiosis between Drosophila and L. plantarum depends on the bacterial genotype and cannot be generalized to the entire species". Note that such a conclusion has already been previously stated (for example, PMID: 30008290 and 28993620).

      We thank the reviewer for the references. Indeed, these papers show that some L. plantarum strains are beneficial for the host while others are neutral. Furthermore, as commented by Reviewer #1 in the public review, L. plantarum has been shown to reduce the host's fitness relative to axenic flies (Gould, PNAS, 2018).

      Our observations are novel in two ways. (1) The fecundity observed in D. melanogaster, 0.06 offspring/female/day in average, is lethal (in Gould et al. 2018 fecundity never decreased below 1 offspring/female/day). (2) Clade C outcompetes the other clades in the cold, despite being detrimental for the host.

      We have modified the Discussion to account for the previous work.

      (21) Line 343: "In addition, we obtained L. plantarum genomes from two other experimental evolution studies. Two genomes from the South African experiment and seven genomes from the Portugal experiment". Merge two sentences into one.

      Done.

      (22) Line 360: "At sampling, the age of the flies varied between four and eight days for the hot environment and between nine and 16 days for the cold environment". Please comment on the fact that different age of flies (different physiology) is not the reason for bacterial community differences.

      During maintenance, flies are sampled at different ages because the temperature affects their developmental time. We cannot rule out the hypothesis that age difference drives microbiome differences. Temperature could affect clade competition directly (differences in optimal temperature between clades) or indirectly, by affecting either the host (e.g. changes in Drosophila developmental time alters L. plantarum fitness), the surounding microbiome, or the food (e.g. increased metabolic activity in the hot regime changes nutrients profile). We ruled out the direct effect of temperature with growth experiments in liquid MRS medium and solid fly food, but disentangling the indirect effects is not feasible.

      (23) The majority of figures have low-quality labels that are not legible due to the small size of the font. Please improve.

      Done.

      (24) Figure 1 - Correct the legend: There is no "10" label on the picture. Probably by 10, the authors mean "Generation", while by x10 - number of isogenic replicates.

      Done.

      (25) Figure 2 - No numbers at nodes are indicated, whereas it is announced in the legend that they represent bootstrap support values. In addition, it is recommended to show a reference pangenome in the middle panel to clearly refer to the total size of the possible black bar.

      We added high bootstrap support as coloured nodes in figures 2, 3 and S2.

      We do not understand the reference pangenome request. In the middle panel, each black/white bar corresponds to an orthologous gene that can be either present or absent in each of the genomes. These orthologs were sorted based on hierarchical clustering of the their patterns of abundance (columns present in the same set of genomes, together), not by synteny. Thus, a reference pangenome would be simply a black bar.

      (26) Figure 2: It would be advantageous to add a figure that represents the frequency of each strain in each replicate (at the last time point, for instance). It would explain why some “blue” strains appear to be within the “red” cluster. Otherwise, it is confusing to find cold-evolved bacterial strains in hot-evolved fly populations.

      The frequency of each clade in each replicate is shown in figure 4. We think that it would be more confusing to follow the suggestion of the reviewer, as the isolates were sampled at different time points of the experiment. We would not like to call it a confusion that "blue" strains appear in the "red" cluster, but rather the logical consequence of the color code used in figure 2, which corresponds to the temperature regime in which the isolate was sampled (regardless of its clade). Whereas in the following figures colour represents the clade. It is thus possible to find clade H isolates in the cold temperature regime, as this clade is in low frequency but not completely absent in this regime.

      (27) Figure 3 – Add a label for the X-axis.

      We rotated the tree to be able to increase the genome IDs. We have added the label to the Y-axis.

      (28) Figure 4 - Please indicate how the clade relative abundance was assessed.

      Clade relative abundance was inferred by mapping competitively the short reads against the three clades’ reference sequences. It is specified in the legend now.

      (29) Figure 6 - Total number of F1 flies eclosed normalised by day (during which period?). What do T1 and T2 correspond to?

      During the respective number of days that females were allowed to lay eggs: one day in the hot settings and two days in the cold settings in transfer 1. One day and three days, respectively, in transfer two.

      T1 and T2 correspond to the first and second transfers, as described in the Materials and Methods. In first transfer, flies laid eggs in vials pre-inoculated with a set load of L. plantarum. After egg laying, same adults were then transferred to a sterile set of vials and allowed to lay eggs again (second transfer). Bacterial load in these vials was solely seeded by the parents.

      In order to avoid any confusion, in the revised version of the manuscript we have modified figure 6 to show transfer 1 for both Drosophila species, and moved transfer 2 dataset to supplementary figure S5.

      (30) Figure S2 - label the X-axis.

      We guess the reviewer means Y-axis. Done.

      (31) Figure S3 demonstrates the real data and its variability, so it would be better used instead of Figure 4 (which seems to be just a derivative from Figure S3, not a separate dataset and separate type of analysis).

      As the reviewer suggested, we have replaced figure 4 with figure S3.

      (32) Figure S4: Improve plot title: (e.g., C:H:U = 3:3:3).

      Done.

      (33) Figure S6: It is stated that N = 10; however, some datasets do not have 10 points represented. Please specify why. Also, please specify the meaning of "T1/T2".

      For the inoculation experiment in Drosophila simulans, we had nine replicates per treatment, not ten. This has been corrected in the figure and in the Materials and Methods section.

      T1 and T2 correspond to the first and second transfers, already covered in recommendation 29.

      (34) Table S3: provide legend for values (1 - present in all strains, but 0.04 - what does it mean?).

      It means that 4% of the genomes from this clade harbour the specific gene. We have specified it in the legend of the revised table.

    1. eLife Assessment

      This manuscript focuses on developing a structural model of how the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, resulting in downstream receptor phosphorylation and signaling. This is an important study, based on solid data. The results will be of interest to scientists working in vascular biology and RTK signaling.

    2. Reviewer #1 (Public review):

      Summary:

      In this manuscript, Uphoff et al. propose a structural and mechanistic model in which the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, thereby promoting downstream receptor phosphorylation and signaling. Using AlphaFold-based modeling, the authors predict that the CCP20 domain of SVEP1 binds to TIE1, creating a composite surface that facilitates Angiopoietin association and TIE1 activation. The resulting ternary model (SVEP1-TIE1-ANG) offers a structural rationale for how SVEP1 converts TIE1 into a functional, ligand-responsive receptor. Additional models and biological assays suggest roles for other domains of SVEP1, such as CCP5-EGF-L7, although these interactions are predicted with low confidence. The authors interpret these findings as the first structural framework for how SVEP1 enables ANG-TIE1 signaling.

      Strengths:

      (1) The central hypothesis - that SVEP1 enables ANG binding to the orphan receptor TIE1 - is biologically compelling and addresses an important question in vascular biology.

      (2) The AlphaFold-predicted ternary complex (SVEP1-TIE1-ANG) is plausible, high-confidence, and structurally consistent with prior functional data (e.g., poly-Ala scanning from Sato-Nishiuchi et al.).

      (3) The authors' model offers a potential explanation for the previously observed role of SVEP1 in enhancing ANG signaling through TIE1 and may represent the first structural insight into TIE1's transition from orphan to ligand-activated receptor.

      (4) The potential clinical implication - that a combinatorial ligand (ANG+SVEP1) can activate TIE1- could have translational relevance for vascular leak and inflammatory disease.

      Comments on revised version:

      The authors have adequately addressed my concerns.

    3. Reviewer #2 (Public review):

      Uphoff and colleagues present the results of a study focused on characterizing the binding of SVEP1 to TIE1 along with Angiopoietin-2. Starting with computational prediction of SVEP1 binding to TIE1, the authors identify the region of SVEP1 that serves as a high-affinity ligand for TIE1. Advanced studies identify a weak secondary binding site within SVEP1 that appears to be sufficient but not necessary for its interaction with TIE1 based on in vivo rescue experiments. The most novel contribution of the manuscript seems to be the identification of angiopoietin-1 and -2 as co-factors that seem to enhance the binding of SVEP1 with TIE1 and impact downstream AKT signaling. They propose a complex in which SVEP1 binds to TIE1 and ANG2.

      Although the first set of results is essentially confirmatory, the identification of ANG-2 as a "co-factor" enhancing the binding of SVEP1 to TIE1 and associated downstream signaling (i.e., Figures 3 and 4) is novel and is of interest. However, the manuscript and its conclusions would greatly benefit from some clarifying details and additional experiments to ensure rigor and support specific claims.

      Comments on revised version:

      I have no further comments. The authors have addressed my concerns.

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Uphoff et al. propose a structural and mechanistic model in which the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, thereby promoting downstream receptor phosphorylation and signaling. Using AlphaFold-based modeling, the authors predict that the CCP20 domain of SVEP1 binds to TIE1, creating a composite surface that facilitates Angiopoietin association and TIE1 activation. The resulting ternary model (SVEP1-TIE1-ANG) offers a structural rationale for how SVEP1 converts TIE1 into a functional, ligand responsive receptor. Additional models and biological assays suggest roles for other domains of SVEP1, such as CCP5-EGF-L7, although these interactions are predicted with low confidence. The authors interpret these findings as the first structural framework for how SVEP1 enables ANG-TIE1 signaling.

      Strengths:

      (1) The central hypothesis - that SVEP1 enables ANG binding to the orphan receptor TIE1 - is biologically compelling and addresses an important question in vascular biology.

      (2) The AlphaFold-predicted ternary complex (SVEP1-TIE1-ANG) is plausible, high-confidence, and structurally consistent with prior functional data (e.g., poly-Ala scanning from Sato-Nishiuchi et al.).

      (3) The authors' model offers a potential explanation for the previously observed role of SVEP1 in enhancing ANG signaling through TIE1, and may represent the first structural insight into TIE1's transition from orphan to ligand-activated receptor.

      (4) The potential clinical implication - that a combinatorial ligand (ANG+SVEP1) can activate TIE1- could have translational relevance for vascular leak and inflammatory disease.

      Weaknesses:

      (1) Lack of structural validation and mechanistic follow-up: Despite the promising AlphaFold model, there are no figures of the predicted interface, no residue-level interactions shown, no ipTM values reported, and no experimental follow-up to test the interface. PAE plots are incorrectly used as confidence justifications, which is not appropriate for complex predictions.

      We have appended the data showing AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity. We also added ipTM scores and confidence plots for the predicted complexes.

      (2) Biophysical validation is missing: No surface plasmon resonance (SPR), ITC, or biochemical assays are included to confirm ternary complex formation or quantify binding kinetics. Given the manuscript's structural focus, this is a major gap. For instance, an SPR experiment where ANG is immobilized, and TIE1 binding is measured {plus minus} SVEP1, would directly test the model. And allow direct comparison to ANG-TIE2.

      We have addressed this question and performed ELISA assays to measure binding affinities between SVEP1 and TIE1 in presence or absence of ANG1 or ANG2, thus confirming that the affinity is increased in the presence of ANG1 or ANG2.

      (3) Missed opportunity for mutagenesis-driven validation: The manuscript does not include any interface-targeted mutations, despite clear opportunities. For example, mutating T2595 in SVEP1 (to R) or mutating the TIE1-specific residues (residues PL 202-203 to LF) could strongly test the model and potentially reveal dominant-negative behaviors. E.g. A T2595 mutant should block ANG binding but not TIE1 binding.

      We have depicted figures of the interfaces including P202-L203 and included the TIE1 P202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study, for the following reason: Modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2, thus reducing its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both TIE1 and ANG2, the data is not included in the manuscript.

      (4) Overinterpretation of weak models: The additional AlphaFold model involving the CCP5-EGFL7 domains binding TIE1 has extremely low confidence (ipTM < 0.15) when reexamined by this reader and should not be emphasized. There is no biophysical evidence or binding data (SPR) to support this interaction, and its inclusion detracts from the much stronger CCP20 model.

      We agree with this point made by both reviewers and have removed the data on CCP5-EGFL7 from the manuscript.

      (5) Language around modeling is overstated and potentially misleading: Terms like "unequivocal," "high-affinity," or "affirms strong binding" in reference to AlphaFold predictions are inappropriate. These are hypotheses -not confirmations - and must be tested at the biochemical level. This should be clarified throughout the manuscript to ensure non-experts do not misinterpret modeling confidence as binding affinity.

      We agree with the reviewer, and have adjusted the wording.

      (6) Negative stain EM data is not informative due to low resolution and lack of defined interfaces; unless replaced by higher-resolution Cryo-EM, this should be omitted. Better would be co-gel filtration, AUC, or SEC-MALLs with ANG-SVEP1-TIE1.

      We have now appended the data by adding gold-labelled TIE/ANG proteins, thus enhancing clarity.

      (7) Disjointed narrative: The manuscript presents a compelling mechanism involving CCP20-driven ANG binding to TIE1, but then becomes fragmented by introducing the low-confidence CCP5-EGFL7 model and speculative higher-order polymerization models that are not experimentally supported.

      We agree with this point and have have centered the manuscript around CCP20. We removed data concerning CCP5-EGFL7 as suggested by both reviewers.

      Reviewer #2 (Public review):

      Uphoff and colleagues present the results of a study focused on characterizing the binding of SVEP1 to TIE1 along with Angiopoietin-2. Starting with computational prediction of SVEP1 binding to TIE1, the authors identify the region of SVEP1 that serves as a high-affinity ligand for TIE1. Advanced studies identify a weak secondary binding site within SVEP1 that appears to be sufficient but not necessary for its interaction with TIE1 based on in vivo rescue experiments. The most novel contribution of the manuscript seems to be the identification of angiopoietin-1 and -2 as co-factors that seem to enhance the binding of SVEP1 with TIE1 and impact downstream AKT signaling. They propose a complex in which SVEP1 binds to TIE1 and ANG2.

      Although the first set of results is essentially confirmatory, the identification of ANG-2 as a "cofactor" enhancing the binding of SVEP1 to TIE1 and associated downstream signaling (i.e., Figures 3 and 4) is novel and is of interest. However, the manuscript and its conclusions would greatly benefit from some clarifying details and additional experiments to ensure rigor and support specific claims.

      We have addressed the reviewers concerns and significantly appended the manuscript. Most importantly, we provide structural validation of AlphaFold models reporting interfaces, residue-level interactions and ipTM values. We have included new biophysical validation of binding kinetics of SVEP1 and TIE1 in the presence or absence of ANG1 or ANG2. Furthermore, we have removed the data on CCP5-EGFL7 from the manuscript in order to retain focus on the CCP20 domain.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The AlphaFold modeling for the CCP20-based interactions is strong (as determined by this reader rerunning and getting ipTM values and visually inspecting the interactions “because this is not in the manuscript”). As presented, the manuscript stops at the hypothesis-generation stage. Validation is needed to fulfill the paper's title and claims. The structure-function link is not demonstrated, despite an obvious and achievable experimental path (mutagenesis, SPR, kinetics).

      Additional Context and Suggestions

      (1) Show and label AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity.

      We have appended the data in the new supplementary figures 1.1, 1.2, 2.1 and 2.2

      (2) Provide ipTM scores and confidence plots for each predicted complex.

      We have added the values and plots in the new supplementary figures.

      (3) Perform SPR assays with ANG-coated surfaces and measure binding of TIE1 {plus minus} SVEP1. Compare to TIE2 binding for context.

      We performed the proposed experiment using an ELISA assay to measure binding affinities between SVEP-1 and TIE1 in presence or absence of ANG2 and included these data in the manuscript in figure 2.

      (4) Test interface mutants: e.g., T2595R in SVEP1 (should impair ANG recruitment but not TIE1 binding), or PL→LF muta on in TIE1 (should disrupt SVEP1 binding).

      We have depicted figures of the interfaces including P202-L203 and included the TIEP202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study. Our modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2 thus reduce its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both, TIE1 and ANG2, the data is not included in the manuscript.

      (5) Clarify in the Introduction that SVEP1 is a large, multidomain ECM protein to help readers contextualize the domain names early on. Do not use terms like CCP before defining them.

      We added: “Svep1 encodes a 3571 amino acid long extracellular matrix protein containing different domains such as Willebrand factor type A domain (vWF), ephrin-receptor like domains, complex control protein (CCP) domains, and Hyalin repeats at the N-terminus. The C-terminus mainly consists of CCP and EGF domains. Svep1 is expressed in mesenchymal cells, but not in endothelial cells, and functions non-cell-autonomously (Karpanen et al. 2017; Morooka et al. 2017)”.

      (6) Replace or remove negative-stain EM unless higher-resolution cryo-EM data are available.

      We would like to retain the EM data, but have now replaced the negative-stain EM by adding gold-labelled TIE/ANG Proteins to verify that the proteins we show are the ones we expect. The reason we would like to retain the data is that TIE1 has been considered for so many years as an orphan receptor, and thus we consider it appropriate to demonstrate SVEP1/TIE1 binding using multiple methods.

      (7) Reframe claims around modeling to avoid overstatement. For example: "The model suggests a plausible mechanism consistent with prior biochemical data" is more appropriate than "unequivocal".

      We agree with the reviewer, and have adjusted the wording.

      (8) Consider narrowing the focus: the CCP20-TIE1-ANG model is a strong story on its own. The CCP5EGFL7 model and polymerization hypotheses are not essential and may dilute the impact.

      Since this point was made by more than one reviewer, we have removed the data on CCP5EGFL7 from the manuscript.

      (9) Properly define TIE1: Tyrosine kinase with Ig and EGF domains.

      We have corrected the full protein name for TIE1 and added: “TIE1 and Tie2 exhibit a high degree of homology with a globular head domain consisting of three immunoglobulin-like (Ig) domains and three epidermal growth factor-like (EGF) modules and a short stalk formed by three fibronectin type III repeats, while the N-terminal two Ig domains of Tie2 harbor the angiopoietin binding site (Macdonald et al. 2006).” We also added: “D1 and D2 refer to the two N-terminal Ig domains, D3 refers to the three EGF domains and D4 to the third Ig domain of TIE1 or TIE2.”

      (10) Use RTK, not tyrosine kinase receptors TKR.

      Tyrosine Kinase receptor was replaced by receptor tyrosine kinases (RTKs)

      (11) PDBs (.cif and .json files) of the models must be supplied for the readers so they don't need to rerun the AlphaFold jobs.

      We are providing all PDBs with the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) In several locations, the authors state that alphafold detects an "unequivocal" and/or "high affinity" interaction. Experts on computational structure prediction can weigh in, but I am not sure it is accurate to say that alphafold predicts affinity. Quantitative estimates of the prediction confidence or other parameters of the alphafold output are not provided.

      Thank you for the comment, we agree with the reviewer and have adjusted the wording. We have also appended the data showing interfaces, ipTM scores and confidence plots for the predicted complexes.

      (2) In Figure 1C, what concentrations of TIE1 and TIE2 are being used in the SPR experiments shown?

      We have added the concentrations of the proteins to the methods section

      (3) In Figure 1C, what is the affinity constant (KD) of the interaction between SVEP1 and TIE1 and SVEP1 and TIE2?

      We have added the values of affinity constants to the manuscript text.

      (4) In Figure 1C, the authors immobilize a "70kD" fragment of human SVEP1 to determine the interaction between SVEP1 and TIE1. How does the affinity they measure between the SVEP1 fragment and TIE1 compare to the affinity between immobilized full-length SVEP1 with TIE1?

      The largest molecule we used in any assay is not full-length SVEP1, but consisted of a C-terminal SVEP1 protein spanning from the first EGF domain to the C-terminus (approximately 295 kDa) as previous studies have shown that SVEP1 is proteolytically cleaved N-terminal to the first EGF domain, generating a protein of this size. In our ELISA assays, this molecule has a higher affinity to TIE1 than the 70 kD fragment. We have included data for the 70 kDa as well as for the 295 kDa fragment in the manuscript. ELISA assays have used larger SVEP1 fragments (as indicated in the figure legends), the SPR assay was performed with the 70 kDa fragment.

      (5) How does the alphafold structure prediction for the "70kD" fragment of hSVEP1 compare to the prediction of the same 70kD fragment from the full-length protein prediction?

      All modellings using different SVEP1 fragments including the 70 kD and full-length version identify the putative binding site at position CCP20. While AlphaFold3 can predict the correct domain folds in both the 70kDa fragment and full-length SVEP1, the orientation of these domains is highly variable due to flexible linkers between each domain module. This flexibility effects the output confidence metrics, thereby hampering our interpretation of the models. Therefore, we conducted the structural prediction of the complexes with smaller fragments and not the full-length SVEP1.

      (6) What are the amino acids for the 70kD fragment?

      The relevant amino acids are 2261-2890. The accession number and amino acids of each protein have been listed in the key resources table.

      (7) In Figure 1D, what is being depicted by the red stars? This is not explained in the text or figure legend.

      We have replaced figure 1D.

      (8) By itself, Figure 1D is not terribly informative and in my opinion does not support the statement that the authors "were able to directly visulalize the attachment of SVEP1 and TIE1." As a minimum, the authors should repeat the same set of images with SVEP1 and TIE2, but other approaches, such as labeling, could be performed.

      We replaced figure 1D with new data and gold-coated protein enhancing clarity. We think it beneficial to demonstrate SVEP1/TIE1 binding using multiple methods as TIE1 has been considered as an orphan receptor for so many years. We have not performed these experiments with TIE2 proteins as we were not able to show binding of TIE2 to SVEP1 with other assays.

      (9) What is being stained in Figure 1D? Full-length SVEP1/TIE1? Or fragments of these proteins?

      We replaced figure 1D by a new assay with labeled proteins using the 150 kDa version of SVEP1 and the ectodomain of TIE1 as well as ANG1 or ANG2 (new figure 1D and new supplementary figure 2.3). TIE1/ANG proteins were gold-labelled. The protein fragments used in this assay are described in detail in the methods section.

      (10) The authors discover CCP6-EGFL7 as a low-affinity binding region of SVEP1 for TIE1. Is this region in physical proximity to CCP20 (the high-affinity binding region for TIE1) based on alphafold prediction? How would the authors think this is binding TIE1?

      We have removed this data set (see comment to reviewer 1’s request).

      (11) What is the affinity constant (KD) for CCP6-EGFL7 with TIE1?

      We have removed this data set (see comment to reviewer’s 1 request).

      (12) The authors claim that ANG1/ANG2 increase affinity between SVEP1 and TIE1 based on immunoblotting. Immunoblots are semi-quantitative at best. If the claim is higher affinity, I think the authors should measure this by SPR and determine the KD between immobilized SVEP1 with TIE1 in the absence and presence of ANG1 and/or ANG2.

      We conducted ELISA assays (figure 2) showing that the affinity is increased in the presence of ANG1 or ANG2 and agree with this reviewer that this strengthens the data.

      (13) In Figure 3, can the authors explain why ANG1/2 does not pull down with SVEP1/TIE1?

      We noticed that upon transfection of TIE1 into HEK cells, ANG1/2 is almost undetectable anymore in the total lysate. Thus, we believe that after the pull down we are below the detection limit.

      (14) In Figure 3, what is "TL"? I assume total lysate, but this is not specified.

      Thank you, we now specify TL as total lysate.

      (15) In Figure 3 "TL" panel (again, I assume this is total lysate), why are the ANG1/2 immunoblots so weak when co-transfected with TIE1?

      We consider it likely that in the presence of TIE1, ANG1/2 proteins are internalized and digested. Another reason for low signals could be that upon transfection of two plasmids, the amount of ANG1/2 protein is reduced as the cell has limited capacity for transcription and translation.

      (16) In Figure 3B, why is the SVEP1 fragment now 150kD when 70kD fragment was previously used? What domains are contained in this 150kD fragment?

      We now better define the domains of the 150kD SVEP1 fragment. The 150 kD fragment was the one produced first and available in high amounts in our laboratory and thus used for functional assays. The 70 kDa fragment together with ANG2 also induces phosphorylation of AKT, but it was not used in as many conditions/replicates as the amounts we had available were lower.

      (17) In Figure 4A, signaling with SVEP1 by itself and ANG2 by itself should be shown to support the claims being made.

      We added the lines for SVEP1 and ANG2, and also the quantification. SVEP1 itself already affects the phosphorylation of TIE1, most likely because hdLECs produce ANG2 by themselves. We show this with the ANG2 blocking antibody for pAKT.

      (18) For pAKT, what are the concentations of proteins being used and the times of incubation?

      This information is provided in the Materials and Methods section. We added the concentration of the anti-ANG2 antibody, which was missing.

      (19) It seems that p-AKT and AKT are being blotted on different gels. If this is correct, loading controls need to be shown for p-AKT blot.

      We added HSC70 as a loading control for both blots.

      (20) It is interesting that anti-ANG2 antibody inhibits SVEP1-induced p-AKT signaling. As the authors may know, SVEP1 has been identified as a receptor for PEAR1, which also leads to downstream p-AKT signaling, which seems to be independent of ANG2. Do the LECs being used here express PEAR1? If these cells express PEAR1, how do the authors think ANG2 silencing will eliminate SVEP1-associated p-AKT signaling?

      hdLECs express PEAR1. However, we show that p-AKT signaling is attenuated after siRNA KO of TIE1. Thus, the downstream signaling is dependent on TIE1 (Figure3).

      (21) Again, experts on computational structure prediction can weigh in, but I am not sure how to interpret the prediction of the 2:2:2 stoichiometry for the theoretical SVEP1/TIE1/ANG1-2 complex. Are there quantitative estimates of the confidence that can be provided? Did the authors attempt to model this complex with different stoichiometries? It is difficult to know how relevant this model is without any experimental results supporting this result.

      Since 1:1:1 is the smallest possible triple complex, it is our starting point. We can model a 2:2:2 version, but anything larger than this AlphaFold will not run. Furthermore, we now provide quantitative estimates of confidence with the pLDDT, PAE, pTM, ipTM scores for all models including the 2:2:2 complexes.

      Minor comments:

      (1) The authors could consider including a reference to alphafold on line 102.

      We have added a reference for AlphaFold2 and 3

      (2) The authors should refer to surface plasmon resonance (SPR) assays by this term as opposed to using the brand name Biacore.

      We agree with the reviewer and have changed the term Biacore to SPR.

    1. eLife Assessment

      This useful study examines whether microsaccade direction primarily indexes shifts rather than the maintenance of covert spatial attention, offering a potentially informative account of inconsistencies in the prior literature. However, the evidence remains incomplete because the key effects are based on relatively sparse microsaccadic events, while concerns about event detection, fixation control, subject-level robustness, and the interpretation of gaze-density analyses remain unresolved. The correlational design and absence of a neutral condition or independent measure of attentional shifting further limit the central claim. The work will be of interest to researchers studying attention, eye movements, and visuomotor mechanisms.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining the relationship between microsaccades and covert attention. This question has been widely investigated, with numerous studies showing that during sustained fixation, when subjects covertly attend to a peripheral stimulus, microsaccades tend to be biased toward the attended location. Here, the authors ask whether this microsaccade bias reflects a shift of covert attention or the maintenance of covert attention. They conclude that the bias is primarily driven by attention shifts, a finding that also helps reconcile the seemingly conflicting results of prior research, where the bias was questioned in paradigms that largely involved attention maintenance rather than shifting.

      Strengths:

      A large sample size was used.

      Weaknesses:

      The main weakness is that the authors' response does not adequately resolve concerns about the robustness of the microsaccade analyses. The newly reported event counts reveal that the number of microsaccades per participant is very low, especially in Experiment 2, and highly variable across subjects. Because the key analyses rely on proportions of microsaccades toward versus away from the attended location, estimates based on so few events are likely unstable and may not provide reliable subject-level measures.

      A second major concern is that several additional analyses introduced in the revision appear to suffer from the same limitation. The permutation analyses and angle-partition analyses may give the impression of statistical rigor, but if the underlying averages are based on very few microsaccadic events, the resulting probabilities are difficult to interpret. Further subdividing already sparse data into narrower angular bins likely makes the estimates even less reliable.

      A third concern is that the authors have not fully addressed issues related to microsaccade detection and fixation control. The presence of very small-amplitude events with relatively high velocities raises the possibility that some detected microsaccades may be artifacts. The authors also did not implement the requested exclusion of microsaccades smaller than 5 arcmin or the suggested reanalysis using stricter fixation criteria. These omissions leave open the possibility that the reported effects are influenced by detection errors.

      A fourth weakness is that some of the requested analyses or clarifications were addressed only superficially. The comparison with Brandolani et al. remains minimal, despite being highly relevant to interpreting whether the observed microsaccade-direction effect is transient or sustained. Similarly, the gaze-density plots do not show the raw gaze-position distributions that were requested and may therefore be misleading, because difference maps cannot determine whether subjects were actually fixating centrally.

      Overall, the revision raises additional concerns rather than resolving the original ones. The main conclusions remain insufficiently supported unless the authors can demonstrate that the effects are robust at the individual-subject level, based on adequate numbers of microsaccadic events, reliable detection criteria, and appropriate controls for fixation behavior.

    3. Reviewer #2 (Public review):

      Summary:

      This study aims to test the hypothesis that microsaccades are linked to the shifting of spatial attention, rather than the maintenance of attention at the cued location. In two experiments, participants were required to judge an orientation change at either a validly cued location (80% of the time) or an invalidly cued location (20% of the time). This change was presented at varying intervals (ranging from 500 to 3,200 ms) after cue onset. Accuracy and reaction times both showed attentional benefits at the valid versus invalid location across the different cue-target intervals. In contrast, microsaccade biases were time-dependent. The authors report a directional bias primarily observed around 400 ms after the cue, with later intervals (particularly in Experiment 2) exhibiting no biases in microsaccade direction towards the cued location. Noteworthy, it would have been interesting to observe whether directional biases in microsaccades are also evident when compared to a neutral condition. The authors argue that this finding supports their initial hypothesis that microsaccade biases reflect shifts in attention, but that maintaining attention at the cued location after an attention shift is not correlated with microsaccade direction.

      Strengths:

      The results are straightforward given the chosen experimental design. The manuscript is clearly written, and the presentation of the study and its visualisations are of a high standard.

      Weaknesses:

      The link between attention and microsaccades has been the subject of extensive research over the past two decades. The authors present a potential solution to the conflicting past findings, arguing that attention should be considered a dynamic process that can be broken down into an attention shift and a sustained attention phase. To differentiate between the two components, the authors varied the interval between the onset of the attention cue and the test stimulus. It would have been nice to use a theory-driven criterion (or an independent measure), in addition to their data-driven approach, to distinguish between these components of a dynamic attention concept. Moreover, it is important to note that the current experiments take a purely correlational approach.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We thank the reviewers for their time and for their valuable inputs throughout the review process. We wish to clarify, one final time, the primary scope and empirical grounding of our work for prospective readers.

      Our study was designed to evaluate whether microsaccades track (in a correlative manner) covert visual-spatial attentional shifting, maintenance, or both. We did so within a single dedicated paradigm, across a large sample (N = 48 human participants). Despite remaining criticisms concerning per-participant event counts and microsaccade classification criteria, the key observation remains a striking dissociation (of the link between microsaccades and covert attention) during the initial shifting and the subsequent maintenance of visual-spatial attention. Moreover, we note how the robust effect observed during shifting (but not maintaining) attention, mitigates residual concerns regarding microsaccade sparsity or signal-to-noise ratio.

      We thus remain confident in the empirical foundation of our work and we invite readers to examine the full paper, supplementary materials, and open-access data to evaluate these findings independently.


      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the editors for their careful evaluation of our article and for their valuable input. Building on these suggestions, we were able to further corroborate our main conclusions, make our article more comprehensive, and thereby substantially strengthen the manuscript.

      We have one additional point of our own: we noticed that in our original submission, we had smoothed the data more than intended. Having caught this, we have now reduced the smoothing employed by 2.5 times compared to the original amount of smoothing (the exact smoothing values have also been added to the methods section). Importantly, however, while this has affected how the results look, this has not affected any of our original conclusions.

      General summary

      We would like to first respond to the major points brought forward by both the editorial summary and the public reviews. As we understand, the two main points that were raised regard: (1) the novelty and theoretical importance of our work and (2) the (in)completeness of our results. We start by providing our response to both of these main points below.

      Novelty and theoretical relevance of the work

      Regarding the novelty of our work, we believe the reviews and, by extension, the editorial summary underappreciated the main theoretical value of the question we addressed. Our work set out to investigate whether microsaccades track covert attentional shifting, attentional maintenance, or both. We fully recognise that there are ample prior studies that investigated and reported a link between microsaccades and covert attention, but also underscore how other studies report seemingly contradicting evidence by reporting that there is no such link. One such example is a recent paper by Willett & Mayo in PNAS (2023). Prompted by the recent hypothesis that this seemingly conflicting evidence may be due to prior work investigating attention ‘in different stages’ (van Ede, PNAS, 2023), we set out to address precisely this using a dedicated task that we designed for this purpose. As acknowledged by the summary and public reviews, this helps to reconcile seemingly opposing views in the literature. In our view, such reconciliation has substantial theoretical value.

      While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’). To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. 

      Finally, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Completeness of results

      Regarding the completeness of our results, the editorial summary points to “the absence of independent measures, single-trial analyses, and neutral-condition controls needed to substantiate the central claims”. While the raised points are valuable, they pertain to issues that are tangential to our primary question and stem from unfortunate misunderstandings of key analytical choices, as we now better clarify. We consider our results complete and comprehensive with regards to the main question our studies set out to answer.

      First, regarding the portrayed “need” for independent measures to define the ‘shift window’ of interest, we clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research. 

      Second, regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the single-trial level in our discussion, and prompted by the valuable feedback, we have expanded on this important contextualisation as part of our revision.

      Finally, regarding the portrayed “need” for a neural-attention control condition, we agree that inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ also at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that challenges our main conclusions. Having clarified this, we embraced this valuable reflection and revised the article to ensure that we do not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation, but refer to ‘the presence of an attentional modulation’ instead.

      Taken together, the explicit design and analysis choices that we made align with the theoretical aims of our study, and the central question we set out to address. The raised points are valuable and we are grateful to have been able to leverage them to improve our article, but we hope to have also clarified how they do not render our findings “incomplete” (as currently portrayed) with regards to the key goal of our article.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining the relationship between microsaccades and covert attention. This question has been widely investigated, with numerous studies showing that during sustained fixation, when subjects covertly attend to a peripheral stimulus, microsaccades tend to be biased toward the attended location. Here, the authors ask whether this microsaccade bias reflects a shift of covert attention or the maintenance of covert attention. They conclude that the bias is primarily driven by attention shifts, a finding that also helps reconcile the seemingly conflicting results of prior research, where the bias was questioned in paradigms that largely involved attention maintenance rather than shifting.

      Strengths:

      The paradigm and conclusions appear sound and supported by the results. A large sample size was used.

      We thank the reviewer for this clear and supportive summary of our work.

      Weaknesses:

      Weaknesses are mostly related to how the authors enforced fixation in the task, and clarifications are needed regarding some methodological details. A more direct comparison of the effect in the two experimental conditions is missing.

      We thank the reviewer for raising these valuable points. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”). Regarding the fixation points, we will address them in our point-by-point replies below.

      Reviewer #2 (Public review):

      Summary:

      This study aims to test the hypothesis that microsaccades are linked to the shifting of spatial attention, rather than the maintenance of attention at the cued location. In two experiments, participants were required to judge an orientation change at either a validly cued location (80% of the time) or an invalidly cued location (20% of the time). This change was presented at varying intervals (ranging from 500 to 3,200 ms) after cue onset. Accuracy and reaction times both showed attentional benefits at the valid versus invalid location across the different cue-target intervals. In contrast, microsaccade biases were time-dependent. The authors report a directional bias primarily observed around 400 ms after the cue, with later intervals (particularly in Experiment 2) exhibiting no biases in microsaccade direction towards the cued location. The authors argue that this finding supports their initial hypothesis that microsaccade biases reflect shifts in attention, but that maintaining attention at the cued location after an attention shift is not correlated with microsaccade direction.

      Strengths:

      The results are straightforward given the chosen experimental design. The manuscript is clearly written, and the presentation of the study and its visualisations are both of a high standard.

      We thank the reviewer for this clear summary of our work.

      Weaknesses:

      The major weakness of this paper is its incremental contribution to a widely studied phenomenon. The link between attention and microsaccades has been the subject of extensive research over the past two decades. This study merely provides a limited overview of the key insights gained from these papers and discussions. In fact, it attempts to summarise previous work by stating that many experiments found a link, while others did not, and provides only a relatively small number of references. To make a significant contribution, I believe the authors should evaluate the field more thoroughly, rather than merely scratching the surface.

      We thank the reviewer for this valuable reflection. For an elaborate response to the perceived novelty, please see our general summary reply above. In addition, we have added a more thorough evaluation of the field to the introduction (page 2, find relevant paragraph below). We hope that this will provide more context for the manuscript and strengthen its contribution to the field.

      Revised paragraph from introduction:

      “This link between microsaccades and covert visual-spatial attention has been demonstrated repeatedly. Early studies linked the direction of microsaccades to the deployment of covert attention [14, 15] and these findings were later replicated and extended. For example, it has been demonstrated in both humans [14–25] and non-human primates [26, 27]; during both externally directed perceptual attention [14, 15, 17, 20, 21, 24–27] and internally directed attention within visual working memory [16, 18, 19, 22, 23]; and in both perception and action tasks following directional cues [28]. Several studies have further linked the directional microsaccade bias to task performance [18, 21, 25, 28, 29]. For example, following spontaneous microsaccades, perception of visual targets presented in the same direction is better [25], and visual discrimination benefits may start already prior to microsaccade execution [21]. Recent evidence further suggests that microsaccades may even play a causal role in shaping the perception of peripheral stimuli [30].”

      The authors then present a potential solution to the conflicting past findings, arguing that attention should be considered a dynamic process that can be broken down into an attention shift and a sustained attention phase. Although the authors present this as a novel concept, I cannot think of anyone in the field who considers spatial attention to be a static entity. Nevertheless, I was curious to see how the authors would attempt to determine the precise timing of the attention shift and manipulate the different stages individually. However, the authors only varied the interval between the onset of the attention cue and the test stimulus, failing to further pinpoint their dynamic attention concept.

      The current version of the experiment, therefore, takes a correlational approach, similar to initial studies by Engbert and Kliegl (2003) and Hafed and Clark (2002). Meanwhile, we have learned a great deal about the link between microsaccades and attention. Below, I will list just a few of these findings to demonstrate how much we already know. It is important to note that, while the present study cites some of these papers, it does not provide a clear overview of how the current study goes beyond previous research.

      (1) Yuval-Greenberg and colleagues (2014) presented stimuli contingent on online-detected microsaccades. A postcue indicated the target for a visual task, and the target could be congruent or incongruent with the microsaccade direction. The authors showed higher visual accuracy in congruent trials. The authors cited that paper, but it is still important to emphasize how this study already tried to go beyond purely correlational links on a single trial level.

      (2) The Desimone lab (Lower et al., 2018) showed that firing rates in monkey V4 and IT were increased when a microsaccade was generated in the direction of the attended target.

      (3) However, attention can modulate responses in the superior colliculus even in the absence of microsaccades (Yu et al., 2022)

      (4) Similarly, Poletti, Rucci & Carrasco (2017) observed attentional modulations in the absence of microsaccades, or comparable attention effects irrespective of whether a microsaccade occurred or not (Roberts & Carrasco, 2019).

      Thus, in light of these insights, I believe the current study only adds incrementally to our understanding of the link between microsaccades and spatial attention.

      We thank the reviewer for this insightful comment, and for pointing out several important studies on this topic. While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’).

      To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. Moreover, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly and appropriately address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the singletrial level in our discussion. Prompted by the reviewer’s valuable feedback, we have expanded on this important contextualisation in our discussion section (page 8: “Therefore, even if microsaccades may more reliably track shifting than maintaining attention, as our current findings show, our findings should not be taken as evidence that microsaccades reliably track attentional shifts at the single-trial level.”). We also incorporated the valuable reference suggestions in our revised manuscript, including in the revised paragraph in our introduction where we provide a more extensive overview of the prior literature, as shown in response to the preceding comment and in the discussion where we discuss the relationship between microsaccades and attention on a single-trial level.

      In general, it is important to have an independent measure of the dynamics of an attention shift. I think a shift of 200-600 ms is quite long, and defining this interval is rather arbitrary. Why consider such a long delay as the shift? Rather than taking a data-driven approach to defining an interval for an attention shift, it would be more convincing to derive an interval of interest based on past research or an independent measure.

      We thank the reviewer for their question. We wish to clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research.

      The present analyses report microsaccade statistics across all trials, but do not directly link single-trial microsaccades to accuracy. Similarly, reaction times and accuracy were analyzed only with respect to valid vs. invalid trials. Here, it would be important to link the findings between microsaccades and performance on a single-trial level. For instance, can the authors report reaction times and accuracy also separately for trials with vs. without microsaccades, and for trials with congruent vs. incongruent microsaccades?

      We thank the reviewer for their sincere interest in our findings and for the great suggestion of an additional analysis. We have now investigated whether trials with a congruent, incongruent or no microsaccade in the shift window (where congruent or incongruent was determined as based on the first microsaccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to stress that our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we have decided not to include these analyses.

      The study would benefit greatly from including a neutral condition to substantiate claims of attentional benefits and costs. It is highly probable that invalid trials would also demonstrate costs in terms of reaction times and accuracy. It would be interesting to observe whether directional biases in microsaccades are also evident when compared to a neutral condition.

      We thank the reviewer for this valuable reflection. We agree that the inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial-numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that in any way challenges our main conclusions.

      Prompted by this reflection, we have ensured to not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation throughout the article, but refer to this only as ‘the presence of an attentional modulation’ instead (such changes were made on pages 3 and 9, and we kept this phrasing consistent in our additions on pages 5 and 22).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The results resemble recent findings by Brandolani et al. (2025), who also showed that the microsaccade bias-in a similar task and using a comparable analysis approach-was restricted to a narrow time window, primarily around the time of the attention shift. The authors should reference this work, discuss similarities and differences, and, given their larger sample size, consider whether they observe a similar correlation between response times and microsaccade rate.

      We thank the reviewer for pointing out this useful reference to us. We have included this article in our introduction, when sketching the current state of the field (page 2) and also mentioned this related article in the discussion (page 8).

      In addition, prompted by this comment, we have investigated whether we observe a similar correlation between response times and microsaccade rate in valid trials. We have investigated this for both possible timeframes: (1) the ‘shift’ timeframe and (2) the ‘maintain’ timeframe. For both experiments, there was no consistent correlation between response times and overall microsaccade rate, as shown in Author response image 1.

      Author response image 1.

      Relationship between the reaction time and the overall saccade rate. This figure shows the relationship between the average reaction time and the average overall saccade rate for both the ‘shift’ period (from 200 to 600 ms after cue onset) and the ‘maintain’ period (from 600 to 1400 ms after cue onset). Each dot represents one participant. Throughout the entire figure, the following significance levels were used: *: p< 0.05, **: p < 0.01, ***: p < 0.001, ****: p < 0.0001.

      (2) I could not find information on the average number of trials per condition and the average number of microsaccades per subject per condition. Ideally, these numbers should be reported (e.g., in a supplementary table). Since the analysis is based on microsaccade direction, knowing how many microsaccades each subject contributed per condition is critical. Microsaccade rates vary substantially across individuals, and subjects with very few events may add noise to the analysis, as proportions of toward/away microsaccades become unreliable.

      We thank the reviewer for pointing out that this relevant information was missing. We have now included these numbers in a supplementary table as suggested (page 17).

      (3) Relatedly, it was unclear whether the time-course analyses were based on collapsing all microsaccade events across subjects or on subject-level averages. In the Methods, the authors state that "the permutation distribution of the largest cluster size was acquired by randomly permuting the trial-average data at the group level 10,000 times," but this is ambiguous. Please clarify.

      We thank the reviewer for pointing out this ambiguity. We have changed the methods section to reflect more clearly that we first obtain time-courses of the microsaccade rate per participant, and subject these time courses to second level statistics using cluster-based permutation analysis (page 13). We have also reworded the sentence you quoted to remove any ambiguity (page 13): “A permutation distribution of the largest cluster size was acquired by randomly permuting the condition labels of each participant’s trial-averaged time course data (i.e. randomly flipping the sign of the difference in rate of toward vs. away saccades) 10,000 times and identifying the size of the largest clusters observed in these randomised data after each permutation.”

      (4) The authors analyze only downward microsaccades, but the cutoff definition is not specified. Presumably, this includes all directions between 180{degree sign} and 360{degree sign}, which may also include nearly horizontal events. This should be clearly stated. In addition, the rate of upward microsaccades should still be shown, divided into up-left and up-right quadrants to parallel the toward/away analysis. This would provide informative context on whether upward microsaccade rates change systematically over time.

      We thank the reviewer for pointing out that this information was missing, and for suggesting this valuable additional analysis. In the methods section, we now explicitly state the angular cutoffs used for the main analysis (page 12). Additionally, we have added a supplementary figure that shows the time course of upwards microsaccades over time (page 21), please see Supplementary Figure S4.

      (5) If the dataset contains enough microsaccadic events per subject, it would be useful to test more conservative angular cutoffs for defining "toward" versus "away."

      We thank the reviewer for this insightful suggestion. We have included an additional analysis, where only microsaccades were included with a direction within a 45° angle around the exacttoward and exact-away directions. This replicated our main finding. The results from this analysis are now included in the supplementary materials (page 24), please see Supplementary Figure S8.

      (6) Figure 2C: It is unclear what the "Center" and "Border" lines represent. The figure is also potentially confusing because it shows microsaccade amplitudes rather than landing positions. Small amplitudes may still bring gaze close to the target; this distinction should be clarified.

      We thank the reviewer for pointing this out. We have changed the “centre” and “border” labels to include more information (pages 6 and 19). We have also added an in-text clarification of the distinction between saccade amplitude and landing position (page 5: “Note that Figure 2C does not show saccade landing positions. While it is theoretically possible for multiple small unidirectional saccades to lead to a larger change in gaze position, a complementary analysis shows that fixation was maintained during the period of peak microsaccade rate in both experiments (see Supplementary Figure S6).”). In addition, in response to the related comment below, we have now also added heatmaps of gaze showing that gaze overall remained close to fixation in our tasks.

      (7) From the Methods, it appears that in Experiment 1, there was no automatic criterion for discarding trials in which gaze deviated from fixation. In Experiment 2, trials were terminated if gaze left a 2{degree sign} window, but given that the target was only 5{degree sign} from fixation, this seems a relatively loose criterion. It would be important to show the distribution of gaze positions during the task to assess whether fixation control was adequate.

      We thank the reviewer for this great suggestion. We have now added a figure to the supplementary materials (page 23) that shows the probability density of gaze position throughout the ‘shift’ period, for left cued trials and right cued trials separately, please see Supplementary Figure S6. We hope that this will further show that even in Experiment 1, fixational control was successful. We also show the difference between left cued and right cued trials, which again shows a gaze bias towards the cued item.

      (8) Did the authors examine whether there was a response time benefit (e.g., RT in congruent microsaccade trials minus RT in incongruent microsaccade trials, as in Brandolani et al., 2025) or an accuracy benefit when microsaccades were directed toward the target?

      We thank the reviewer for their interest in our findings and for the great suggestion of an additional analysis. As we discussed also in response to the general summary from reviewer #2 above, we have now investigated whether trials with a congruent, incongruent or no saccade in the shift window (where congruent or incongruent was determined as based on the first saccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to again stress how our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we decided not to include these analyses. However, please note that we did now include the outcomes of another analysis that more directly targeted the relation between the spatial modulations in microsaccades and task performance, as we turn to below. 

      (9) Was there a relationship between the size of the attentional effect and the magnitude of the microsaccade bias?

      We thank the reviewer also for this insightful question. We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (10) The criteria for minimum microsaccade amplitude and duration are not specified. This should be clarified. I recommend excluding events smaller than ~5 arcmin, as these are likely noise-especially since eye tracking was monocular. Monocular "microsaccades" can be spurious, but this can be determined only with binocular tracking. It is also unclear whether subjects used chin/head rests. A main-sequence plot in the supplementary material would be helpful.

      We thank the reviewer for pointing this out. We have included a main-sequence plot in the supplementary materials (page 23). The main-sequence plot can also be found in Supplementary Figure S7 and suggests that our saccade-detection algorithm worked well with detected saccades following the main sequence. We have also stated more clearly in the methods section that subjects used a chinrest (page 11).

      (11) Please specify the asterisk convention in figure captions (i.e., what * vs. ** vs. *** correspond to in terms of p-values).

      We thank the reviewer for pointing out that these significance levels were not mentioned in every figure caption, so we have added this information to every figure caption where they were missing (page 4, 6, 7 and 19).

      (12) The fact that stricter fixation criteria reduced the size of the effect suggests the possibility that gaze drift toward the target might have conferred an eccentricity advantage in this discrimination task. A direct comparison of the effect in the two experiments would be valuable. The authors should comment on this. It would be informative to plot the average gaze position around the time of peak microsaccade rate in both experiments. Reanalyzing the data post hoc with a stricter trial-selection criterion (e.g., excluding trials where gaze deviated more than 1{degree sign} from fixation) could also be very valuable, as it would systematically test how fixation control influences the observed microsaccade-attention relationship. This would be informative for the community studying this topic.

      We thank the reviewer for these valuable reflections. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

      Regarding the average gaze position around the time of peak microsaccade rate: in response to reviewer #1, under point (7), we have included Supplementary Figure S6 that shows the probability distribution of gaze position throughout the ‘shift’ period (the same figure is found on page 23 in the article), which shows that fixational control was successful in both experiments. This period is also the period of peak microsaccade rate in both experiments.

      We wholeheartedly agree that systematically investigating the effect of fixational control is important for the field as a whole, and this is also precisely why we set out to perform the same experiment in two different experimental settings with regards to fixational control, and why we decided to include the results from both experiment variants side-by-side in our article.

      Reviewer #2 (Recommendations for the authors):

      In addition to my general concerns in the public review, I have the following recommendations.

      (1) Did the authors distinguish between the initial and subsequent microsaccades during their analysis? Is it possible to produce multiple microsaccades when shifting attention, or do the authors only consider the first microsaccade to be linked to an attention shift?

      We thank the reviewer for pointing out this ambiguity. We have now stated more clearly in the methods section that we consider all microsaccades for our analyses (page 12: “Crucially, we did not restrict our analyses to initial saccades; rather, all detected saccades were included. This allowed us to examine oculomotor behaviour during later trial phases, where initial saccades are unlikely to occur.”). We also believe this methodological choice is important, as otherwise it would be conceivable that no microsaccade bias can be found during the ‘sustain’ period, simply because no ‘first’ microsaccades occur anymore.

      (2) Two microsaccades had to be separated by at least 100 ms. This is an unusually long delay.

      Could the authors please specify how many microsaccades were discarded using this criterion?

      This inter-saccade-interval is quite large on purpose, as we want to minimise the probability of counting the same microsaccade twice. We have re-analysed the data with a minimum ISI of 50 ms, and this led to an increase of found saccades of a, respectively, 5.1% and 1.9% increase for Experiments 1 and 2. However, two participants in Experiment 1 led to a much higher increase in saccades than all other participants (these participants had z-scores of 3.9 and 2.4 for the number of additionally found saccades with an ISI of 50 ms; all other z-scores for Experiment 1 were between -0.5 and 0.5). When those two participants were removed, in Experiment 1 only 1.6% more saccades were found.

      (3) If I understand correctly, the authors did not use staircase procedures to eliminate differences in task difficulty between participants. Could the authors demonstrate how task difficulty relates to the link between microsaccades and performance? For example, is the time course of the microsaccade direction bias correlated with performance?

      We thank the reviewer for this suggestion (that overlaps with a comment of Reviewer 1). We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (4) The authors reported using equiluminant stimuli. Could the authors please specify the exact luminance?

      We thank the reviewer for pointing out this missing information. We have now included this information in the methods section (page 11: “, with a luminance of 88.5 cd/m<sup>2</sup>.”). We have also included the luminance of the background (page 11: “luminance: 29.0 cd/m<sup>2</sup>”).

      (5) Could the authors please provide a full polar plot showing all microsaccade directions, and colour-code those included in the analysis?

      We thank the reviewer for this great suggestion on how to present our results even more clearly and comprehensively. We have included a supplementary figure showing the full polar histograms (with 20 radial bins), for all three timeframes of interest: the whole trial, the ‘shift’ period and the ‘maintain’ period (page 20). As requested, the saccades included in the main analyses are colour-coded. See Supplementary Figure S3

      (6) Can the authors please directly compare the main effects between experiment 1 and experiment 2 (Figure 2B)?

      We thank the reviewer for this great suggestion (that was also made by reviewer 1). We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

    1. eLife Assessment

      The authors describe a cell-specific mechanism by which glutamate transporters regulate the fidelity with which T-stellate cells in the mouse ventral cochlear nucleus relay information from auditory nerve inputs. The study is supported by solid electrophysiological data. It provides valuable insights into how the rapid binding of glutamate to transporters shapes auditory information processing at specific synapses.

    2. Reviewer #1 (Public review):

      In this article, the authors investigate how glutamate transporter function regulates excitability and synaptic coding in T-stellate cells in the mouse ventral cochlear nucleus. They test this in acute brain slices using whole-cell electrophysiology and artificially raise the relative local concentration of glutamate via pharmacological inhibition of transporter proteins. The main finding is that when sub-saturating doses of DL-TBOA are applied, cells become much more sensitive to synaptic input, diminishing the normally high fidelity of EPSP-spike coupling in these neurons. Notably, high-frequency stimulation in the presence of DL-TBOA reveals a large and slowly decaying AMPA receptor component that underlies persistent/rebound firing in earlier recordings. These effects are not seen in other ventral cochlear neurons, suggesting that rapid glutamate clearance in T-stellate cells, particularly, is important for auditory intensity coding. Overall, these experiments are well-performed, and the findings are robust, though there are some aspects that could be expanded to make the work more impactful. These include a better understanding of the relative contribution of neuronal vs glial transporters and an ability to separate the relative contributions of tonic glutamate concentrations in the cleft vs changes in membrane potential in action potential output. Additionally, there were some minor issues of clarity in both the figure presentation and the main text language that should be addressed.

      Major Points:

      (1) Given the dramatic effect of saturating DL-TBOA on tonic leak/RMP and that the sub-maximal concentration used in most of the experiments still varied between 25-50 uM, Figure 1 would be strengthened substantially by a dose-response curve. Ideally, 5 or 6 concentrations, plotting the effect on tonic current or RMP increase.

      (2) Examining the contribution of glial (EAAT1/2) vs. neuronal (EAAT3) transporters (Fig 8) is intriguing but comes across as incomplete here, especially given the small number of recordings. Using a different non-selective EAAT inhibitor (TFB-TBOA) to chase the EAAT1/2 blocker combo seems like an odd choice, given that you have already characterized the effects of DL-TBOA well. One could also try a lower concentration (~50-100 nM) of TFB-TBOA since it is somewhat selective itself for glial EAAT1/2. Given the data presented, neuronal transporters (presumably EAAT3) appear to dominate the rapid clearance of glutamate at this synapse, but this point isn't emphasized or explored sufficiently.

      (3) Separating the effects of depolarization vs. glutamate clearance was never explored. What effect does depolarizing the cell ~10 mV in control conditions (i.e., without TBOA) have on AP number/fidelity during synaptic stimulation experiments? The authors state that submaximal DL-TBOA generally causes no more than a 5 mV change in RMP, but tonic depolarization could also influence spike fidelity. This experiment could demonstrate that the increase in excitability during/after stimulation is not due to increased engagement of voltage-gated channels.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and mechanistically interesting question: whether plasma membrane glutamate transporters contribute only to slow clearance of ambient glutamate or whether they can rapidly shape synaptic signaling during high-frequency auditory activity. This manuscript provides important evidence that EAAT-mediated glutamate uptake is not merely a slow background clearance mechanism but is essential for maintaining reliable synaptic transmission and linear stimulus-intensity coding in ventral cochlear nucleus T-stellate cells during sustained auditory nerve activity.

      Strengths:

      The finding that EAATs may be required for rapid, local control of glutamate during high-frequency auditory nerve activity is interesting and could have broad relevance to auditory processing. The electrophysiological evidence is generally strong, particularly the use of patch-clamp recordings, stimulus trains, partial versus complete EAAT blockade, and comparison with bushy cell/endbulb synapses. The comparison between T-stellate cells and bushy cells/endbulb synapses strengthens the manuscript. The authors demonstrate that EAAT blockade disrupts coding in T-stellate cells but has little effect on bushy cell spike transmission, supporting a cell-type- and synapse-specific role of glutamate uptake.

      Weaknesses:

      However, some mechanistic conclusions, especially the specific contribution of neuronal versus glial EAATs and the absence of glutamate crosstalk between auditory nerve inputs, rely mainly on pharmacological and indirect electrophysiological inference and would be strengthened by additional anatomical, genetic, or direct glutamate-sensing evidence.

      (1) Clarification of DL-TBOA concentration.

      The authors used bath application of 200 µM TBOA and 25-50 µM in the other experiments, stating that "sub-maximal concentrations (25-50 µM)". The authors should provide a clearer rationale for why different concentrations were used across experiments rather than a fixed concentration.

      The reversibility of DL-TBOA effects should be demonstrated by washout experiments. In addition, potential off-target effects of DL-TBOA on postsynaptic receptors, intrinsic membrane excitability, or presynaptic release (e.g., PPR measurement) should be carefully considered. It would also be useful to test the effects of the submaximal DL-TBOA concentrations (25-50 µM) on membrane potential and inward currents, shown in Figure 1, to determine whether these concentrations depolarize the membrane potential in current-clamp mode or induce inward currents under voltage-clamp conditions.

      (2) Potential contribution of altered intrinsic excitability.

      In Figures 3B and 3C, DL-TBOA appears to induce additional action potentials even immediately after the first stimulation, whereas Figures 6 and 7 suggest that the first EPSC is not substantially altered. This raises the possibility that the enhanced firing may partly result from a modest depolarization caused by background glutamate accumulation or from other changes in intrinsic membrane properties after drug treatment. To address this, the authors should provide a quantitative analysis of physiological parameters under submaximal DL-TBOA conditions, including spontaneous action potential frequency, resting membrane potential, input resistance, and spike threshold.

      (3) Spillover/ crosstalk between AN-fiber-synpases.

      The authors should provide more explanation of how altering the number of active auditory nerve fibers demonstrates the absence of glutamate spillover/crosstalk between bouton synapses. Strong stimulation likely recruits more AN fibers, but it may also change release probability, axonal synchrony, or stimulation spread. The authors should more clearly justify the interpretation that strong stimulation recruits additional independent AN fibers rather than altering release probability or activating fibers with different intrinsic properties.

      (4) Interpretation of glial versus neuronal EAAT contributions.

      The authors claim that both neuronal and glial transporters contribute to rapid uptake using pharmacological approaches. The pharmacological data demonstrate that glial EAATs play a major role in glutamate clearance at T-stellate cell synapses. The strong increase in EPSC decay time and synaptic charge after UCPH-101/DHK application supports the conclusion that glial transporters contribute substantially to limiting glutamate accumulation during sustained auditory nerve activity. However, the conclusion that neuronal EAATs contribute directly should be stated with some caution. The evidence for neuronal EAAT involvement is indirect and depends on the pharmacological specificity and completeness of glial EAAT blockade. The conclusion would be strengthened by additional evidence, such as EAAT subtype expression/localization in T-stellate cells or auditory nerve terminals, transporter current recordings, immunohistochemistry, or genetic manipulation of neuronal EAATs. In addition, fitting the decay phase with a double-exponential model may help determine whether glial and neuronal EAATs contribute over distinct temporal windows.

    4. Author Response:

      We are grateful for the careful and extensive reviews, and are pleased that the reviewers found the work of broad interest to sensory processing. Please find our proposal for revision based on public reviews:

      Reviewer 1

      1) Request for dose-response curve for DL-TBOA and leak current or RMP. We can provide this, at least for the initial phase of the curve relevant to the concentrations used for synaptic experiments. Prolonged exposure to higher concentrations leads to very large cationic currents (through AMPAR) which appear to be damaging to membrane integrity.

      2) We will increase the N for glial vs neuronal block with the blockers we already used; this seems more practical than doing new experiments with different concentrations of TFB-TBOA. 

      3) We can include data to test the effect of blockers or small depolarizations on excitability.

      Reviewer 2

      1) We differentiated experiments with “25-50 uM” from 200 uM DL-TBOA because the higher concentration clearly led to massive AMPAR activation and depolarization block, as shown in Fig 1. We then chose lower concentrations to minimize background current while allowing glutamate build-up during exocytosis.  We felt we were clear on this point. 

      As to reversibility and “off target effects” like synaptic changes, we will provide this information. See also response to Reviewer 1, comment 1.

      2) See response to Reviewer 1, comment 3.

      3) We are certain that increasing stimulus strength increases the number of stimulated fibers, and this is well accepted. The stimulus electrode is placed in the auditory nerve root, well away from recorded cell and synapses, minimizing current spread to synapses. We can compare PPR for weak and strong stimuli in our current dataset to confirm no effects on release probability. As to variations in the intrinsic properties of myelinated auditory nerve fibers and their sensitivity to stimulation, there is no information about this, and do not understand how it would be relevant, particularly in as much as we report a negative result: no difference in blocker effect with small or large numbers of fibers active. The 3 main types of myelinated auditory nerve fiber, Type 1a,b,c, are known to respond to different sound thresholds, but that is a synaptic issue in the inner ear, and apparently not related to the myelinated axon bundle.

      4) We appreciate the reviewer's caution about a role for neuronal transporters and will revise accordingly.  We cited molecular evidence for expression of subtypes in the pre and postsynaptic neurons and in glial cells. Of course, given how ubiquitous such expression is across the brain, we suspect the kinds of experiments we provided offer more direct evidence for function.

    1. eLife Assessment

      This work provides a valuable contribution by leveraging simulation-based inference to investigate candidate compensatory mechanisms in neuronal network models and offering new insights into how distinct pathological perturbations may require different interventions. The computational evidence is convincing and has been strengthened by additional reproducibility analyses, although some aspects of inference validation and the biological interpretation of posterior dependencies warrant further investigation. The study establishes a helpful computational framework for exploring disease-specific compensatory mechanisms and will be of broad interest to the computational and systems neuroscience communities.

    2. Joint Public Review:

      Summary:

      This manuscript couples a 32-parameter model with simulation-based inference (SBI) to identify parameter changes that can compensate for three canonical hyperexcitability perturbations (interneuron loss, recurrent-excitatory sprouting, and intrinsic depolarisation). The study demonstrates a careful implementation of SBI and offers a practical ranking of "compensatory levers" that could, in principle, guide therapeutic strategies for epilepsy and related network disorders.

      Strengths:

      (1) By analysing three mechanistically distinct hyper-excitable regimes within the same modelling and inference framework, the work reveals how different perturbations require different compensatory interventions.

      (2) The authors adopt posterior estimation to systematically rank the efficiency of different mechanisms in balancing hyperexcitability.

      (3) Code and data are available.

      Comments on revised version:

      I appreciate the authors' extensive efforts in revising the manuscript and responding to the previous review. The revised version is substantially improved in clarity, organization, and presentation. In particular, the addition of schematic figures, the reorganization of the Methods section, the improved explanation of the model, and the inclusion of replication analyses all strengthen the manuscript.

      The manuscript remains entirely computational, and therefore its conclusions should be interpreted as predictions generated by a specific model rather than validated biological mechanisms. I believe the work has the potential to make a useful methodological contribution. However, several concerns remain regarding validation, interpretation of inferred posteriors, organization of the manuscript, and presentation.

      Major comments:

      (1) The manuscript states that simulation-based calibration showed the amortized posterior estimator was unreliable (85-88), but these results are not shown. The manuscript explicitly states that simulation-based calibration demonstrated substantial failures of the amortized posterior estimator, yet the corresponding analyses are not presented. Since these results motivate the transition to sequential NPE and are central to assessing inference reliability, they should be reported quantitatively, either in the main text or supplementary material.

      (2) The authors present two independently trained estimators and show strong agreement between them. This is a useful robustness analysis. However, the rebuttal occasionally presents this as addressing concerns regarding cross-validation and generalization. The new analysis does not constitute cross-validation in the usual sense and does not directly assess generalization to held-out targets or posterior accuracy.<br /> I recommend that the authors explicitly describe Figure 4 as a reproducibility analysis and avoid presenting it as a substitute for validation.

      (3) Posterior correlations are useful for generating hypotheses about compensatory mechanisms, but they should not be interpreted as direct evidence of compensation. The compensatory interpretation should instead be supported by the perturbation analyses (e.g., Figure 6), which provide mechanistic validation.

      The manuscript consistently treats posterior correlations and conditional posterior shifts as direct evidence of compensatory mechanisms. These are consistent with compensatory mechanisms, but they do not by themselves establish that the corresponding biological parameters causally compensate for the perturbation. I recommend clarifying this distinction and emphasizing that the conditional posterior analyses generate hypotheses regarding compensation, which are then partially supported by the perturbation experiments shown later in the manuscript.

      The language throughout the manuscript should therefore be softened.

      (4) The manuscript repeatedly suggests that the inferred conditional distributions may be useful for identifying precise interventions or guiding personalized treatments (examples include lines 24-29, lines 217-223, lines 242-246, lines 277-282, lines 283-286). These claims go beyond what is directly demonstrated.

      The study does not evaluate treatment outcomes, patient-specific inference, intervention efficacy, or clinical decision-making. Rather, it demonstrates differences in inferred parameter distributions within a computational model. While these results are valuable and may generate clinically relevant hypotheses, they do not yet establish predictive utility for treatment selection or precision medicine. I therefore recommend substantially softening these translational claims and emphasizing that the current findings generate hypotheses that could be tested experimentally in future work.

      (5) The revised manuscript still mixes presentation of findings with interpretation.

      For example, lines 217-226 largely continue to describe findings from Figure 6 and would fit better in the Results section. The Discussion would be strengthened by focusing more exclusively on biological implications, limitations, and future directions.

      A similar issue appears later in the discussion comparing posterior correlations and conditional distributions. Much of this section effectively reinterprets Figures 2 and 3 rather than discussing broader implications.

      (6) The discussion around lines 271-282 overstates what can be concluded from the inferred posteriors.<br /> The statement that correlations "discover broadly applicable mechanisms" whereas conditionals "identify specific mechanisms" is stronger than the presented evidence supports. Likewise, the conclusion that conditional distributions are more useful for precision treatments is speculative and not directly demonstrated.

      I recommend reformulating these statements as interpretations or hypotheses rather than conclusions.

      (7) Around line 84, the manuscript introduces q(theta|x) without clearly defining θ, x, or q. Readers unfamiliar with SBI may struggle to follow the notation. All quantities should be defined when first introduced.

      (8) The manuscript equates larger KS distances between conditional posteriors with greater compensatory potential. While KS distance provides a useful measure of posterior redistribution, it is not obvious that it should be interpreted as a measure of biological efficacy.

      (9) The manuscript would benefit from a discussion of parameter identifiability. The inference problem maps 32 model parameters to 7 summary statistics, implying substantial non-identifiability. While complete identifiability analysis is likely beyond the scope of the current work, this limitation should be discussed explicitly.

      All in all, the revised manuscript is significantly improved and addresses several concerns raised in the previous review. However, important issues remain as discussed above.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Summary:

      This manuscript couples a 32-parameter model with simulation-based inference (SBI) to identify parameter changes that can compensate for three canonical hyperexcitability perturbations (interneuron loss, recurrent-excitatory sprouting, and intrinsic depolarisation). The study demonstrates a careful implementation of SBI and offers a practical ranking of "compensatory levers" that could, in principle, guide therapeutic strategies for epilepsy and related network disorders.

      Strengths:

      (1) By analysing three mechanistically distinct hyper-excitable regimes within the same modelling and inference framework, the work reveals how different perturbations require different compensatory interventions.

      (2) The authors adopt posterior estimation to systematically rank the efficiency of different mechanisms in balancing hyperexcitability.

      (3) Code and data are available.

      We thank the reviewers for their positive comments on our manuscript.

      Weaknesses:

      (1) A highly dense presentation of the simulated models and undefined symbols makes it hard for readers outside the modelling community to follow the biological message. An illustration of the models, accompanied by some explanations and references to the main equations and parameters discussed in this paper, would make the first section much more straightforward.

      Thank you for this feedback. To clarify our methods, we have added Figure 7, which illustrates the dynamics of the point neurons and their synapses. We have also added explanations and definitions of variables right where they appear. These variables were previously defined only in a table on a different page.

      We also moved the methods section to the back of the paper, as is common in many modern manuscripts. We hope that relegating method details to the end makes the manuscript more accessible.

      (2) This methodology appears to be a brute-force approach, requiring millions of simulations to tune 32 parameters in a network of 500-700 cells. It isn't scalable. Moreover, the authors did not use cross-validation, which, with a relatively low increase in computational cost, would provide a quantitative measure as to how well it generalizes; this combination raises doubts about both scalability and reliability.

      Scalability is indeed a key challenge of SBI methods. Amortized neural posterior estimation (NPE) is a brute-force approach in that it samples solely from the prior distribution, which is extremely wide. Many of these samples are therefore not very informative for the biologically plausible dynamics we are interested in, which is a downside of amortized NPE. However, amortized NPE is extremely scalable because once the estimator is trained, it can estimate the parameter distribution of any given output dynamic. We tried to build an amortized NPE for our simulator, but simulation-based calibration (a method to validate posterior estimates using additional simulations) showed that the estimators were unreliable.

      Sequential NPE is not a brute-force approach because it samples from posterior estimates, which are narrower than the prior. Because the amortized NPE failed, we use sequential NPE to create the two estimators for the baseline and the hyperexcitable condition described in the paper. While millions of prior samples are used to generate the initial posterior estimate, which is then sequentially refined, the sequential refinement requires only 80,000 additional simulations. This requires a significant amount of computational resources, which is why we consider the results worth reporting, but we make the simulator, the simulation results, and the trained estimators available, so other researchers can use or train their own estimators without running millions of simulations. We hope our rewrites make the advantages and disadvantages of the approach clearer.

      Regarding reliability and cross-validation, we agree that our initial submission has fallen short. We presented results from only one density estimator per condition, which we considered sufficient given the large number of samples. In the revised version, we present the key results from two additional density estimators trained on partially new training data (Figure 4).

      (3) Several parameters remain so broadly distributed after fitting that the model cannot say with confidence which specific changes matter. Therefore, presenting them as "compensatory levers" is somewhat questionable.

      It is indeed difficult to determine which changes matter because of the simulator’s complexity. Especially the marginal correlation coefficient (Figure 2 C) are small and the pairwise histograms are broad (Figure 2A), because all other parameters are unconstrained. But the conditional correlation coefficients are larger (Figure 2D) and narrower (Figure 2B). We have added the histograms in Figure 2B in the revised version to highlight the difference. We cannot provide a definitive threshold for correlation coefficients to discriminate between important and unimportant mechanisms. Therefore, compensatory mechanisms discovered with SBI should be validated mechanistically, as we do in Figure 5.

      (4) Every conclusion is drawn from simulated data; without testing the predictions on recordings, we have no evidence that the proposed interventions would work in real neural tissue. Because today we cannot diagnose which of the three modelled pathological regimes is actually present in vivo, the paper's recommendations cannot yet be used to guide therapy.

      This is indeed an unfortunate drawback of our current work. We are working to apply this approach to constrain microcircuit simulators with data from epilepsy patients. But that work is currently ongoing and will not fit into the present manuscript.

      Recommendations for the authors:

      Beyond the issues I wrote above, which are methodological, I would like to raise my concern about the way this manuscript is written:

      We highly appreciate this editorial feedback on clarity and style. Such feedback is rare and we have worked to address each point to improve the manuscript.

      (1) Paragraphs - several paragraphs start with: "To quantify/identify/find specific compensatory mechanisms of hyperexcitability with simulation-based inference". It is a good idea to orient the reader with the specific goal of each section, but it is not helpful to repeat the overall message of the paper in every paragraph. Several paragraphs open with "However," or "Additionally,". Please restructure sentences so that connectors appear after a clear topic sentence.

      We have done major rewrites to improve the readability of our manuscript. We start paragraphs with more specific context sentences, rather than the broad research goal, and also made paragraphs much shorter with clearer main messages.

      (2) Section 1: The way you present NPE, it would seem like it's specific to neuroscience (and it's not). The paragraph starting at line 26 is not clear. Please revise it. Line 29 - missing a "." before the next sentence begins. Avoid phrases like "for the longest time".

      We now stress that NPE, like SBI, is used across scientific domains.

      (3) Section 2 was tough to read. Please present each equation in its own numbered display, followed immediately by a plain-language explanation of every symbol and parameter. Provide an illustrative diagram: a small schematic of the AdEx neuron, synaptic connections, and the three perturbations. Even a simple block figure will orient nonexperts. Keep critical methodological decisions (priors, summary statistics, simulation length) in the main text, but move voluminous tables of parameter bounds, learning rates, and hardware specs to the supplement. Remove mentions of which Python functions you used. Readers care about algorithmic choices, not function names. Please reserve specific code references for the GitHub README.

      We have added schematic panels at the beginning of Figures 3 & 4 and added Figure 7, which illustrates the neuron and synapse models of the simulator. We also made major rewrites to the methods section to remove programmatic implementation details and define variables where they appear.

      In general, I think it would be a good idea to have an editor to polish syntax, verb tense consistency, and punctuation. A thorough language edit will improve the readability and impact of this manuscript.

      We have attempted to improve the points raised by the reviewer. In particular, we have carefully rewritten verb tense and punctuation throughout the revised manuscript.

    1. eLife Assessment

      This important study reports a novel phenomenon of maternal growth during pregnancy that is independent of growth hormone (GH), adding a new dimension to maternal biology of reproduction. The evidence is convincing and supported by state-of-the-art methodologies conducted in mice and persuasive observations in humans with hereditary isolated GH deficiency. Revised discussion should focus on possible mechanisms, including the role of IGF2, and on directions for future research.

    2. Reviewer #1 (Public review):

      This work evaluates the impact of reproductive history on growth, body weight and body composition in mammals. In mice, somatic growth is stimulated by the first pregnancy while the second pregnancy increases body weight mainly by increasing adiposity. To probe the role of pituitary growth hormone (GH), the key regulator of somatic growth in these processes, was addressed by comparing the impact of reproduction on growth in normal ("wild type") and genetically GH-deficient females and by detailed characterization of the profile of fluctuations in circulating GH levels in both types of animals. Additional studies addressed the possible role of other endocrine pathways (ghrelin and estrogen) in the pregnancy-related growth. Surprisingly, reproduction-related growth was independent of GH, ghrelin and estrogen. To determine whether these results may apply ("translate") to human physiology, data on various parameters of somatic growth were collected from women with hereditary GH deficiency. The findings indicate that GH-independent stimulation of growth by reproductive events also occurs in women.

      Use of multiple animal models, rigorous characterization of GH levels in normal and GH-deficient females, and inclusion of data derived from a unique and well-characterised population of people with hereditary isolated GH deficiency and no GH replacement therapy are important strengths of these elegant and innovative studies. The results address a broader and clinically significant issue of permanent changes in body size, composition and function that result from pregnancy and lactation. This work also provides important background for further studies aimed at the identification of the mechanism involved and the role of specific reproductive events in the regulation of growth.

    3. Reviewer #2 (Public review):

      This manuscript describes the fascinating phenomenon of growth hormone (GH)-independent growth occurring in the mother during pregnancy. This growth was most pronounced in dwarf mice that are lacking the receptor for growth hormone-releasing hormone (GHRH) and therefore showing isolated GH deficiency. However, the pregnancy-induced growth could also be observed in wild-type mice, suggesting that it is a normal part of the maternal adaptation to pregnancy. The study falls short of identifying the mechanism(s) driving this pregnancy-induced growth response, but it certainly reveals a novel insight into maternal physiology. The authors have completed a range of experiments in mice to prove that, as well as being GH independent, the pregnancy-induced growth also did not require GH signaling in the liver (i.e. not another pregnancy-specific ligand operating through the GHR to promote IGF). They also provided complementary data from a population of humans with untreated isolated GH deficiency that are broadly consistent with the hypothesis. While it is important to consider the significant species differences between rodents and humans, both in terms of growth physiology and also in terms of evolution of placental somato-mammotrophic hormones, this unique population are a valuable resource and adds credence to the study. Overall, I find this a compelling research story, but disappointingly unfinished. There are some areas where additional information could improve the ability to interpret the data, and some additional concepts that could be considered in the discussion. There are also areas where additional experiments might provide important insights. However, I think that such suggestions can be considered as appropriate for future research, rather than delaying consideration of the current manuscript.

      Main comments:

      (1) Data in Figure 1 are remarkable - not so much the growth in pregnancy in the wildtype mice, because while elevated GH is well known in pregnancy, but growth in the dwarf mice is indicative of GH-independent growth. From these data, it seems that there is good evidence that growth in pregnancy is an adaptive function. However, it is possible that growth is achieved in dwarf mice and that in wildtype mice may have been mediated through different mechanisms. The dwarf mice showed an increase in liver and plasma IGF1, suggestive of an additional ligand driving IGF in pregnancy. One could hypothesize that such an effect could be mediated by an additional pregnancy-specific ligand activating the GH receptor. In humans, placental growth hormone could be such a ligand, but as far as we know, there is no placental GH in mice. In contrast, the wildtype animals showed suppression of liver and circulating IGF1, and low levels of pSTAT5 in the liver during pregnancy. These data (in Figure 5) are very surprising. Given the high circulating GH in pregnancy, as well as high placental lactogen (which would be expected to activate STAT5 in the liver through the Prlr), the low levels of pSTAT5 are unexpected and would seem to indicate some sort of acquired insensitivity to GH. Is this entirely driven by down-regulation of STAT5b protein, or could there be activation of other, negative regulators of STAT signalling, such as SOCS? What is causing such a profound suppression of STAT5? Regardless of the mechanism, this suggests that pregnancy-induced growth in wildtype mice is independent of circulating IGF1 (potentially a different mechanism or in addition to that seen in IGHD mice).

      The data shown in Figure 6 are a major strength of the study, showing that the pregnancy-induced changes are not specific to one particular transgenic model, but still occur in a variety of models affecting GH through different approaches. Given the pregnancy-specific nature of the changes, however, it seems an oversight not to have evaluated the role of placental lactogens. Prlr is highly expressed in the liver, but the function of this hormone in the liver is not well established. Could the extremely high levels of PL be mediating this growth response? Given the low expression of STAT5 in the liver and the fact that plasma IGF1 is not markedly elevated, it seems more likely that this growth response may be mediated by locally produced IGF1 in target tissues.

      I think these possibilities could be addressed by an expanded discussion of species variation in placental hormones, to highlight that humans have expansion of the GH locus, but rodents have expansion of the prolactin axis (see Soares, M. J. The prolactin and growth hormone families: pregnancy-specific hormones/cytokines at the maternal-fetal interface. Reprod Biol Endocrinol 2, 51, 2004). Importantly, placental GH and chorionic somatomammotropins (CSM) in humans are all variants of the GH gene, but CSM have preferential activity at Prlr. This seems to be a fundamental species difference in pregnancy biology, but has been interpreted as an example of convergent evolution, with conservation of prolactin and GH-like functions at the maternal-fetal interface, mediated by different mechanisms, likely contributing to the metabolic adaptations of the mother (see Newbern D, Freemark M. Placental hormones and the control of maternal metabolism and fetal growth. Curr Opin Endocrinol Diabetes Obes. 2011; 18: 409-416). While the preceding function has focused on explaining the evolution of placental lactogens (either prolactin or GH variants), the present data suggest that there are also mechanisms to maintain growth in pregnancy, independent of GH (even in the absence of a placental GH).

      (2) The human data are very interesting, and my initial impression was that it seemed unlikely to be the same phenomenon. Was there any real evidence for "growth" in pregnancy? Pubertal maturation of long bone growth might be expected to prevent further growth in adulthood. However, these issues were appropriately discussed, and it seems well justified to evaluate this unique population of women with IGHD who underwent pregnancy. It would be very interesting to know if these women experienced elevated IGF1 during pregnancy, indicative of placental GH contributing to growth. Mechanistically, this might be more like the dwarf mouse situation of IGHD, that the situation in wildtype mice (associated with liver insensitivity to GH and low IGF1).

      (3) It would be useful to include investigations that isolate the effects of pregnancy and the placental hormones. Such studies could include evaluating growth in pseudopregnant mice with IGHD (pregnancy-like changes in hormones but lacking the placental contribution) and in IGHD animals that experience pregnancy but not lactation (pups removed at birth). I accept that this might be too large an additional study to add for the present manuscript.

      (4) It is an important and translationally relevant observation that pregnancy increased the risk of long-term weight gain, and that after the first pregnancy, the pregnancy-induced growth response was more directed to promoting fat deposition. Does this provide any mechanistic insight? Could a metabolic adaptation result in growth?

    4. Reviewer #3 (Public review):

      Summary:

      The study describes an increase in body growth and body composition in both mice and women. In mice, the impact on growth is mainly seen during the first pregnancy, and the changes postpartum on body composition are also different during the first and second pregnancies. The study has used various knock-out models in the growth hormone axis to understand these changes as well as some gene expression analysis related to GH, IGF-1 and estrogen signalling pathways.

      Strengths:

      (1) The inclusion of various knock-out mouse models that allow for exploration of mechanisms related to the above-mentioned changes.

      (2) The investigation of gene expression of GHR, IGF-1R and ER pathways.

      Weaknesses:

      The human findings are dependent on the patient's recollection of bodily changes after their pregnancies.

      Conclusion:

      The authors have partly achieved their aim of describing changes in growth and body composition that remain after pregnancy and the mechanisms behind these changes. This study may have importance for a wide variety of research areas as well as in the clinical setting. The study is also unique in its attempt to bridge findings in mice to a unique human model of congenital GH deficiency.

    1. eLife Assessment

      In this valuable study, Zhang et al. investigated EEG neurofeedback as a method to modulate brain activity prior to painful stimulation and examined its effect on pain perception in a well-powered, double-blind study. Results showed that real, but not sham, feedback enabled learning-dependent enhancement of pre-stimulus α oscillations. However, the evidence for a neurofeedback-specific reduction in pain remains incomplete, as the current paradigm cannot distinguish between genuine neurofeedback effects and placebo effects. Nonetheless, this work is likely to be of interest to researchers in the fields of neurofeedback and pain.

    2. Reviewer #1 (Public review):

      Summary:

      Zhang et al. investigated EEG neurofeedback as a method to modulate brain activity prior to painful stimulation and its effect on pain perception. Neurofeedback was designed to train participants to upregulate alpha power contralateral to the site of painful stimulation. Real or sham neurofeedback was administered to two independent groups. Each group performed two tasks: one in which participants were asked to modulate their brain signals (training task) and another in which they were asked to passively watch the feedback (non-training task). The authors reported an increase in alpha power during real neurofeedback training compared with sham training and non-training conditions. The authors also reported a decrease in pain perception during the training task, both in the real and sham neurofeedback groups. Additionally, in an offline analysis, the authors investigated brain dynamics with microstate analysis during the neurofeedback training. Also, they implemented a mediation analysis to infer which brain responses to neurofeedback training mediated changes in pain perception.

      Strengths:

      (1) The research question is licit and sound. EEG neurofeedback is a promising non-invasive technique with the potential to alleviate at least the sensory component of pain. The rationale for applying neurofeedback at the alpha band in the somatosensory cortex is well justified by the alpha-gating theory in pain modulation.

      (2) The sample size is adequate to capture neurofeedback effects. The effort to conduct a double-blind study with a complex design paradigm and an adequate sample size is valuable and appreciated.

      Weaknesses:

      (1) Reported behavioral effects on pain reduction might be due to the placebo effect rather than neurofeedback, as pain ratings were reduced both in the real and sham neurofeedback groups during training. It is important that authors report this effect appropriately and disclose which information was given to the participants when they enrolled in the study, i.e., whether the paradigm was designed to reduce pain perception.

      (2) The utility of training effects, especially in the sham group, is unclear. I understand that including the non-training condition allows the distinction between neurofeedback effects and arousal effects. However, interpreting training effects should not be the point of this study. What does it tell us that participants who received sham stimulation increased or decreased alpha power in the training session vs the non-training session?

      (3) There might be hidden time effects (habituation/sensitization) on pain responses and/or on brain responses to neurofeedback. A within-session analysis comparing the first half of the training with the second half should be conducted to discard them.

      (4) Connectivity analysis reflects spurious effects. In EEG, deriving phase-based functional connectivity at the sensor level is problematic due to volume conduction effects. EEG functional connectivity should be performed after source reconstruction, and measures discarding instantaneous phase lags should be preferred, which is not the case with magnitude-squared coherence. See (Bastos and Schoffelen, 2015).

      Although neurofeedback is a promising technique for modulating pain perception, the current study adds limited novelty to the field, as its design could not disentangle whether behavioral effects (reductions in pain intensity and unpleasantness) were specific to neurofeedback training or due to non-specific effects (e.g., placebo). Nevertheless, the authors corroborated that brain states before painful stimuli could be modulated with neurofeedback (enhancement of alpha power).

    3. Reviewer #2 (Public review):

      Summary:

      This study uses neurofeedback to modulate alpha-band activity and examines how this influences pain-related processing. The question is timely and methodologically elegant, because it addresses whether noninvasive modulation of ongoing oscillatory activity can causally shape pain perception and/or expectation-related processes.

      Strengths:

      The use of neurofeedback as a tool to modulate alpha activity is a major strength, because it provides a noninvasive and conceptually clean approach to probing the functional role of oscillatory brain activity. The design is also attractive because it links neurophysiological regulation to a psychologically meaningful outcome, namely pain processing. Further, the induced changes were also related to different EEG microstates and ERP components during the processing of the pain stimulus, and therefore the authors demonstrate a clear relation between preparatory prestimulus states and stimulus processing.

      The manuscript appears to address an important and clinically relevant question, and the idea of testing whether alpha regulation can alter pain-related responses is of high interest for systems neuroscience and pain research.

      Weaknesses:

      Methodologically, it is unclear what alpha values were used in the analyses. It is stated that alpha was extracted within 2s windows of the 16s long feedback period. However, the values change across this period. Which value is used for the correlation with the pain ratings and all other analyses? Using the average across the 16s could reflect large values in the first half and low values in the final half, but for the relationship between alpha and pain, the last segments should be more relevant. If the initially elevated alpha activity subsides several seconds before the onset of the pain stimulus, it is difficult to see how it could influence subsequent pain processing.

      Related, after the 16s feedback period, a fixation period is used with a 3-5s length. If alpha band activity is relevant for the consecutive pain processing, the amount of alpha in this period should be relevant. The authors should demonstrate that the induced alpha change during the feedback period remains stable during the fixation period and that the activity in this period is related to pain processing.

      Further, it should be noted that the alpha band modulations related to alpha band training were accompanied by significant effects in other frequencies. Therefore, a clear relationship between alpha and behavioral pain ratings is not the only interpretation. Correlations with other frequencies or combinations of frequency band modulations should be incorporated to allow a more precise interpretation. Furthermore, in the sham feedback group, an increase in alpha band activity was observed (p=0.06), and the small difference in the pain intensity rating may be related to a clear outlier in the Sham group (Figure 4a).

      In both groups, a main effect of training, regardless of sham or real feedback, was reported with a small difference between groups. But the main modulator seems to be related to the instruction to modulate the neural activity, and this large effect should be discussed in more detail regarding, for example, possible attentional processes.

      A further central concern is that the visual feedback signal (the ball movement) may generate expectations that are not specific to alpha activity and that these expectation processes modulate the pain processing (ball down may indicate more pain). It is well known that intensity cues can generate expectations about upcoming perceptions, and the used feedback signal with an increasing or decreasing visual curve clearly signals what intensity should be expected. Therefore, it is important to show that the amount of positive (ball up) and negative visual displays is matched between the sham and real feedback group. Further, the authors should report whether the final ball position can predict the latter pain rating in both groups or differentially. Following this interpretation, alpha band activity is not directly related to pain processing but only serves as a signal that is transformed to a visual stimulus that then generates expectations.

      Finally, the manuscript would benefit from a more explicit analysis of whether individual alpha changes are related to pain ratings within each subject. If higher alpha is truly linked to reduced pain perception, this should be visible at the participant level during learning of the neurofeedback procedure. Relatedly, there is no learning period incorporated, and usually participants are not able to regulate their alpha activity from the first trial on. The authors should include an analysis of the development of alpha band activity over learning and a relation of these individual alpha values and the corresponding pain ratings.

      I cannot find a link to the preregistration in the current manuscript.

      In summary, a "causal" relation of alpha activity with pain perception -that is mentioned several times in the manuscript- is not fully supported by the present results

    1. eLife Assessment

      This important study investigates whether perceived gender is represented in the brain in a category-invariant manner across faces, bodies, and objects, identifying the right middle temporal gyrus (rMTG) as a potential locus. The evidence is incomplete due to major conceptual concerns, weak statistical methods, and unaddressed low-level confounds like stimulus size and motion. This work will be of interest to psychology and social neuroscience researchers in face and person perception literature, provided the authors temper their claims regarding abstract representation.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript investigates whether the human brain contains a shared category-general representation of gender across faces, bodies, and gender-associated objects. The authors acquired fMRI data while participants viewed male and female stimuli from three categories in a one-back task. They then used searchlight MVPA, cross-category decoding, regression-based RSA, CNN vs. brain representational comparisons, and PPI analyses. Their main finding is that gender information could be decoded from distributed occipitotemporal regions within each category, whereas a cluster in the rMTG showed convergence across cross-category decoding and RSA. The authors concluded that this rMTG representation resembles intermediate layers of fine-tuned CNNs and that face and body gender processing share similar functional connectivity patterns.

      Strengths:

      The question is potentially important, particularly for social cognition, object recognition, and the use of neural network models to interpret high-level visual representations. Previous behavioral studies have shown cross-category adaptation between bodies and faces, and even between gender-associated objects and faces, so the attempt to test for a neural counterpart using fMRI is well motivated. The use of multiple complementary analyses including within-category decoding, cross-category decoding, regression RSA, CNN comparisons, and effective connectivity analyses is also a strength. The convergence of cross-category MVPA and RSA in a right MTG cluster is potentially interesting and deserves attention.

      Weaknesses:

      The largest problem is conceptual. The term gender is used as if it refers to the same construct across faces, bodies, and objects. This is not self-evident. In faces and bodies, the stimuli seem to contain visual cues from which observers infer binary gender categories. In objects, however, the relevant information is almost gender stereotype, cultural association, or learned semantic association. These are not equivalent constructs. The manuscript therefore needs to distinguish much more carefully between perceived gender, biological sex cues, gender-associated visual features, and gender stereotypes. Without this distinction, the title and main conclusion are too broad. The object condition is particularly problematic. Javadi & Wee (2012) showed that gender-associated objects can bias subsequent judgments of ambiguous face gender, and they discussed two possible mechanisms, including shared neural substrates or top-down modulation induced by the gender concept. However, their behavioral adaptation study does not directly demonstrate that objects, faces, and bodies are encoded in the same neural representational format. The present manuscript treats these object stimuli as if they provide evidence about the same kind of gender representation as faces and bodies, but that step requires additional empirical support. Independent ratings of object gender association, cultural familiarity, visual similarity, and semantic category are essential here.

      A second major concern is stimulus control. The face images were taken from Chinese male and female actors, the body images were headless bodies in underwear, and the object images were selected because of prior gender associations. This design introduces many possible confounds: hairstyle, makeup, skin texture, body shape, clothing, color, luminance, object category, object function, curvature, spatial frequency, and cultural familiarity. Cross-category decoding can be significant even when a classifier relies on shared visual statistics rather than an abstract gender code. For example, female-associated stimuli may differ from male-associated stimuli in color, shape, brightness, texture, or semantic category in ways that are consistent across faces, bodies, and objects. The present analyses do not adequately rule out these alternatives. Foster et al. (2019) are especially relevant in this respect. They reported that body sex could be decoded from both body- and face-responsive regions. However, the sex of well-controlled faces, for example faces excluding hairstyle cues, could not be decoded from face- or body-responsive regions. This finding should make the authors more cautious. The fact that the present study used more ecological face stimuli may increase sensitivity to gender-related cues, but it also increases the possibilities that decoding is driven by uncontrolled external features rather than by an abstract gender representation. Accordingly, because no additional visual, semantic, or stereotype-based model RDMs were included in the RSA analysis, this result alone cannot establish an abstract, category-independent gender representation. Any systematic difference between male- and female-associated images will load onto the gender RDM. At least, the authors should include additional model RDMs for low-level visual features. In addition, the current RSA analysis has another limitation. The neural RDMs are based on only six condition-level patterns, producing a 6 × 6 matrix. The theoretical model includes only binary gender and category RDMs. This is too coarse to support the claim of category-independent gender representation. Ideally, all the RSA analysis should be performed at the item level rather than at the condition level.

      The cross-category decoding result in rMTG is promising but not yet conclusive. The authors identify a right MTG cluster by overlapping thresholded maps from three cross-category decoding analyses. This is useful descriptively, but it does not by itself establish a common representational code. The overlap of thresholded maps depends on the chosen threshold. If the authors want to make a formal conjunction claim, they should use a valid conjunction-null approach such as a minimum-statistic conjunction evaluated under the appropriate conjunction null, rather than simply displaying the intersection of thresholded maps. Even if this approach cannot be adopted in this study, the issue should be included as a limitation.

      In the PPI analysis, the reported similarity between face and body connectivity matrices is a little bit small (r = 0.08). The claim of a shared functional network should therefore be softened unless the authors test whether this correlation is significantly larger than the face-object and body-object correlations, correct for multiple comparisons, account for the non-independence of matrix elements, and report participant-level distributions and confidence intervals.

    3. Reviewer #2 (Public review):

      Summary:

      The study tests whether male/female-related information is represented in a form that generalizes across faces, bodies, and gender-associated objects. Using within- and cross-category MVPA, regression RSA, comparisons with fine-tuned CNNs, and connectivity analyses, the authors identify a right middle temporal gyrus region whose patterns generalize across the three stimulus classes. They conclude that this region provides a category-general, mid-level representation of gender and acts as a neural hub.

      Strengths:

      The question is novel and important, while the logic of the study is straightforward. Examining faces, bodies, and objects within the same participants provides a useful extension beyond the predominantly face-based literature. Cross-category decoding is also a stronger test of shared information than simple anatomical overlap between within-category maps. The combination of MVPA, RSA, computational modelling, and connectivity analysis is ambitious, and the replication of the CNN layer profile with both AlexNet and VGG16 is a useful characterization of relevant information.

      Weaknesses:

      (1) The construct labelled "gender" is not equivalent across stimulus classes. For faces and bodies, the male/female label is intended to track a property of the depicted person, albeit one inferred imperfectly from appearance; for objects, masculinity or femininity is not an intrinsic property of the object but a culturally contingent association that may vary across observers and contexts. Treating both as levels of a single binary factor risks conflating person-category information with gender-stereotypic object associations and interpreting their common neural discriminability as evidence for one abstract concept of gender. The term "object gender" could also be confused with grammatical gender in some languages (e.g., French or German).

      (2) The CNN analysis does not isolate the shared male/female component. The authors correlate the complete six-condition neural RDM with the complete CNN RDM. However, rMTG also carries substantial information about whether an image is a face, body, or object. Consequently, the peak correspondence with Conv4 may reflect category structure rather than the representation that supports cross-category male/female decoding. The current analysis does not establish that shared gender-related information specifically depends on mid-level features.

      (3) The connectivity interpretation is overstated. PPI measures task-dependent covariance; it does not establish information transmission, directionality, or an upstream-to-downstream processing sequence. The reported face-body connectivity similarity is also small (r=.08). Also, describing rMTG as a "hub" is not justified without network-centrality measures, lesion evidence, or causal perturbation.

      The authors partly achieve their aims. The results provide credible evidence that patterns in rMTG contain information that generalizes across binary male/female-labelled faces and bodies and masculine/feminine-associated objects. They do not yet establish a genuinely abstract representation of gender, a specifically gender-related correspondence with intermediate CNN layers, or a neural hub that transmits information through a directed network. With more precise framing and targeted reanalysis, the study could make a useful contribution to research on social vision and cross-category representation.

    4. Reviewer #3 (Public review):

      Summary:

      In this work, the authors investigate whether gender information is encoded in the brain in a way that is invariant to the object being perceived. They design an fMRI experiment in which 22 participants perform a one-back repetition detection task in a block design. Images shown are of three types (faces, objects, and bodies) and of two perceived genders, male and female. They perform MVPA, RSA, and functional connectivity analyses to determine whether gender information is invariant to the type of image being perceived. They report an area in the posterior right middle temporal gyrus (rMTG) that is found in their gender decoding analysis across categories. To confirm that this area encodes gender information, they perform a regression-based RSA with category and gender model RDMs, and report that the gender model RDM is significantly correlated with brain representations in that area. Finally, to further investigate the representations in this area, they perform a model-based RSA in which they first fine-tune a deep neural network for gender classification, and then study the correlation between model RDMs and brain RDMs. Consistent with a previous report in face processing (Jiahui et al., 2023), they find that gender information is more consistent with representations in middle-to-late layers of the networks. Additional functional connectivity and PPI analyses are reported to reveal differences in co-fluctuation of brain activity within occipital and parietal nodes when perceiving different types of male/female images. Based on these results, the authors conclude that rMTG represents gender information invariant of the category perceived, although rMTG also afforded decoding of category information.

      Strengths:

      Whether perceived gender is represented in a manner invariant to the category of the stimulus is a legitimate and interesting question, and one of relevance particularly to the face and person perception literature.

      The model-based RSA, in which RDMs from networks fine-tuned for gender classification are compared against brain RDMs, is an interesting approach, and the layer-wise profile the authors obtain converges with a previous report in the face processing literature (Jiahui et al., 2023).

      Weaknesses:

      A substantial number of inferences are drawn on the basis of weak statistical methods and a suboptimal design. My concerns are set out below, ordered by severity.

      (1) The statistical tests are not appropriate for classification and RSA, and are prone to false positives. Classification accuracies and RSA correlations may be positively biased, and the true null distribution may therefore be centered above the nominal chance level, or above zero in the case of RSA. Testing against a theoretical value with a one-sample t-test under these conditions inflates the false positive rate, especially with few test samples per classification, and does not afford valid population inference for information-like measures (Combrisson & Jerbi, 2015; Allefeld et al., 2016). The concern applies to every inferential claim in the manuscript, including the identification of the rMTG cluster on which the paper's central conclusion rests. The established remedy is permutation testing, in which the labels are randomly permuted and the full analysis, including cross-validation, is re-computed so that any bias is captured in the empirical null distribution (Stelzer et al., 2013; Etzel & Braver, 2013). This approach has been applied in comparable face-decoding studies using both classification and RSA (Guntupalli et al., 2017). I raise this methodological concern here because it is the clearest way to convey why the reported statistics cannot be safely interpreted at face value.

      (2) The decoding analyses do not appear to test generalization to left-out stimuli. From my reading of the design, each run contained all six conditions presented three times in random order, with each block containing 12 images (10 unique plus two repetitions serving as catch trials). If all images were presented in every run, the same images would be present in both the training and test sets of the cross-validation. Under these conditions, the interpretation of a general "gender" code is difficult to justify: the classifier may be exploiting low-level image features specific to the particular exemplars rather than gender per se. This bears directly on the paper's central claim, which concerns an abstract, category-invariant representation of gender, a claim that requires decoding to generalize to stimuli the classifier has not encountered.

      (3) There is no evidence that participants perceived the stimuli's gender as the authors assumed. Perceived gender may be subject-specific, yet no norming is reported establishing that participants actually rated or processed the stimuli according to the gender the authors assigned to each image. Some images are likely to be more ambiguous than others. This is a construct validity issue rather than an analysis issue: the class labels used throughout the decoding analyses, and the gender model RDM used in the RSA, both rest on an assumption about the participants' percepts that is never tested against the participants themselves.

      (4) The rMTG ROI reported in Figure 2c appears to overlap almost perfectly with the motion-sensitive area hMT+. The reported effects may therefore be driven, at least in part, by low-level motion signals arising from the rapid on/off changes of the stimuli and the associated optic flow. I am not claiming that the results are fully driven by this, but no control reported in the manuscript rules it out, and this region is the centerpiece of the paper's conclusion.

      (5) Stimulus size is confounded with category in the functional connectivity analyses. The authors report that functional connectivity differed between faces and objects, and between bodies and objects. However, faces and bodies were shown with the same visual extent, while objects were larger. Given that the nodes being investigated are in visual areas, it is unclear how these differences can be attributed to category rather than to the low-level difference in stimulus size. The same confound bears on the behavioral task performed within the scanner: participants can perform the one-back task more easily, simply by detecting size differences, since two images of different sizes are clearly not the same image, rather than by processing the image content. This affects what can be assumed about participants' attention to the stimulus category or gender.

      (6) No motion quality control is reported for the functional connectivity analyses. Functional connectivity is well known to be highly susceptible to head motion, yet the manuscript reports no summary of how much subject motion there was, no indication of whether volumes with excessive motion were removed or censored, and no account of quality control on the measured data more generally.

      (7) The use of famous faces introduces an avoidable confound. The face stimuli were famous faces. Famous and familiar faces are known to recruit substantially more widespread activity than unfamiliar faces, extending well beyond the core visual system (Gobbini & Haxby, 2007; Natu & O'Toole, 2011; Visconti di Oleggio Castello et al., 2017; Kovacs, 2020). For a study focused specifically on gender, this introduces a source of variance that unfamiliar faces would have avoided, and it complicates the comparison of the face conditions against the body and object conditions.

      (8) The rationale and benefit of fine-tuning the deep neural networks are not established. The manuscript does not report the original, non-fine-tuned accuracy of the models that required fine-tuning, so the benefit of the procedure cannot be assessed; given that the final validation accuracy is low, it is unclear that fine-tuning actually helped. AlexNet and VGG are trained for object classification on large datasets, and fine-tuning with 2,000 training images may not be sufficient to genuinely shift the objective. Whether the activation patterns and RDMs changed in any significant manner after fine-tuning is not reported, and the rationale for selecting the specific layers used is not stated.

      (9) Taken together, the analyses as presented do not establish the paper's central claim. My concern is not that the reported effects are necessarily absent, but that the combination of statistical tests that do not account for possible positive bias, a cross-validation scheme that may not guarantee generalization across stimuli, a key region that coincides with a motion-sensitive area, and gender labels that were never validated against participants' own perception leaves too many open questions for the results to be evaluated as they stand.

      (10) I would add one broader consideration. Perceived gender is likely to depend on culture and to vary across individuals. A binary male/female contrast in 22 participants, without evidence that those participants perceived the stimuli as the authors intended, is a narrow operationalization of a construct that is unlikely to be so simple. Even if the analyses were fully sound, caution would be warranted in generalizing from this design to claims about how the brain universally represents gender.

    1. eLife Assessment

      This important study extends a model of cortical normalization (ORGaNICs) to interacting cortical areas and shows that communication through coherence and communication subspaces can arise from a single set of dynamics. The evidence is solid, showing analytically that contrast-dependent gamma dynamics and a low-dimensional inter-areal communication subspace arise from one parameter set, though the comparisons to data remain qualitative and the analytics rest on a linearization that is not checked against numerical simulation. The work will interest neuroscientists and theorists concerned with inter-areal communication, cortical oscillations, and divisive normalization.

    2. Reviewer #1 (Public review):

      In this paper, Pal and colleagues propose a mechanistic unification of two influential accounts of inter-areal communication: communication through coherence and communication subspaces. A major strength of the paper is that it does not treat coherence and communication subspaces as independent phenomena, as typically done, but instead derives both from the same circuit with divisive normalization. In this framework, noise-driven fluctuations around the normalized fixed point determine covariance and cross-power structure (which, in retrospect, makes so much sense to be related). Then, they show how these determine linear prediction performance and the effective dimensionality of the communication subspace. They also show (however not very visually, see recommendation below for a figure) how divisive normalization is crucial to shape inter-areal coherence and the dimensionality of communication.

      I found this conceptual contribution potentially very influential, but somewhat obscured by the technical complexity of the model. The central intuition (I think) is that recurrent normalization can organize cross-area fluctuations, both frequency-specific correlations and cross-covariances. Took me a while to grasp this insight, mostly because I was stuck with the model details. Note that I have some experience with network dynamics, but not with this particular model.

    3. Reviewer #2 (Public review):

      Summary:

      The authors extend the ORGaNICs framework (a recurrent circuit that dynamically implements divisive normalization) to connected cortical areas with explicit top-down feedback. Because the network has a known analytical fixed point that coincides with (or closely approximates) the normalization equation, the authors can linearize about that fixed point and derive closed-form expressions for the power spectral density, inter-areal coherence, and communication subspaces. Using a two-area instantiation (V1 & V2) with a single fixed parameter set and no data fitting, they show the model reproduces: (i) contrast-response functions with steeper slope V2; (ii) gamma-band power and coherence peaks that shift to higher frequency with contrast; and (iii) a low-dimensional inter-areal communication subspace that is lower-dimensional than the within-area subspace. They derive parallel predictions of what happens by changing model parameters: feedback gain enhances inter-areal and suppresses within-area communication, and normalization is necessary for both the oscillatory dynamics and the reduced subspace dimensionality. A three-area extension (V1&V4, V1&V5/MT) is used to argue that differential top-down feedback can dynamically route functional connectivity.

      Strengths:

      (1) Analytical tractability: Deriving power spectra, coherence, and communication-subspace structure in closed form from a known fixed point is genuinely valuable.

      (2) Conceptual unification: Framing coherence and communication subspaces as arising from the same normalization-driven dynamics is an elegant and useful contribution.

      (3) Breadth from few assumptions: A large range of phenomena (contrast gain, gamma dynamics) emerges from normalization-based model assumptions.

      (4) Biological grounding: The mapping of model variables onto identified cell types connects the abstract computation to known cortical microcircuitry.

      (5) The prediction that input-gain versus feedback-gain modulation produce distinct spectral signatures gives experimentalists a clear way to test the framework.

      Weaknesses:

      (1) Comparisons are qualitative, not quantitative: The theory/experiment panels are visual side-by-side comparisons. There is no quantitative goodness-of-fit for any predictions.

      (2) The simulations use τ ≈ 1 ms for all cell types, which the authors acknowledge is unrealistically short; realistic values would shift the gamma peaks to lower frequencies.

      (3) Divisive normalization is a special case and is recovered exactly only for the identity recurrent matrix (self-normalization). Some statements that the circuit implements divisive normalization exactly need softening.

      (4) The element-wise (multiplicative) interaction in the modulator dynamics is not tied to a specific cellular mechanism.

    4. Reviewer #3 (Public review):

      Summary

      The work of Pal and colleagues considers a hierarchical and multi-population version of the "oscillatory recurrent gated neural integrator circuits" (ORGaNICs) model, showing through analytics that the model captures multiple relevant experimental results: first of all, its oscillatory dynamics produce a profile with high resemblance to experimental results, both in terms of decay of power at high frequency and in terms of shifting peak as a function of stimulus contrast. Second, inter-areal communication subspace dimensionality is lower than within-area dimensionality. The authors then proceed to further characterize the model's response properties as a function of input and feedback gain. In particular, they find that frequencies transmitted with higher strength also carry more information, that changing gain modifies the dimensionality of communication subspaces, and that these properties can be used in a three-layer model, where an upstream area can select which downstream area to communicate to, based on the strength of feedback gain.

      Strengths

      This work demonstrates that a single-circuit model with normalization properties can capture both the oscillatory dynamics and the inter-areal communication properties measured in cortical circuits, matching multiple experimental results. The full analytical tractability of the model is highly advantageous, allowing for easier exploration of parameters, replicability, and effective interpretations of results compared to purely numerical approaches.

      The work also makes a useful conceptual link between normalization, coherence-based communication, and subspace-based communication. In particular, it shows how both phenomena can emerge from the same circuit dynamics, where normalization is a key factor.

      Interestingly, the model is also extended to multiple areas, showing how attention (in the form of changes in feedback gain) can synchronize the activity of a downstream area with one of two upstream areas, thus effectively selecting which area to communicate with.

      In general, this is an interesting computational framework and a useful starting point for future modeling work. A particular strength is that it connects normalization, oscillatory dynamics, coherence, and communication subspaces within one analytically tractable model, making it possible to generate mechanistic hypotheses about when inter-areal communication should be stronger, lower-dimensional, or preferentially routed through feedback.

      Weaknesses

      Although I see the analytic approach as a strength, at the same time I regard the lack of any numerical comparison as a big weakness. Circuit simulations would not only confirm the correctness of the analytics, but also offer further insights on the error margins and on the regimes where the analytics are valid. This is because, to my understanding, the analytics are based on a linear approximation around the operating regime, which means deviations might be expected, especially for high gain levels in the input, or in the feedforward and feedback pathways.

      Another problem is that the analytically tractable model seems to rely on effective connectivity weights that break Dale's law. Numerical simulations with explicitly modeled excitatory and inhibitory units might give insights into effects due, e.g., to the additional transmission delays mentioned in the Discussion.

      Another weakness is the use of the term "predictions" to indicate features of the model dynamics that are purely described in the context of the model parameters. Although the model's response properties may certainly lead to predictions, I think the term requires a better contextualization in terms of neurophysiology and experimental neuroscience. The Discussion draws very interesting and valuable bridges between neuron morphology, interneuron types, and model parameters. But it seems it's left to the reader to backtrack and figure out which biological mechanisms or experimental manipulations should correspond to changes in input or feedback gain, and how these should be distinguished from possible changes in feedforward gain.

      Relatedly, the manuscript places substantial emphasis on modulation of feedback gain, but does not comparably explore modulation of the feedforward gain, β2, which regulates the V1-to-V2 drive. This seems important because changes in feedforward gain could also influence communication subspace dimensionality and oscillatory dynamics. Therefore, predictions related to top-down feedback modulations should be taken with a grain of salt.

      Last but not least, the model dynamics are split among multiple elements and nonlinear interactions, reaching a level of complexity far higher than the other ORGaNICs formulations present in the literature. The authors derive these dynamics in the supplementary material, as a dynamical system that converges to a fixed-point solution that includes "exact divisive normalization". I wonder, however, if there could be simpler solutions that also produce normalization, either approximate or in a different form than the one proposed by the authors. Note also that the designation of "excitatory neurons" is misleading: despite the presence of two explicitly inhibitory populations, the "excitatory" units also interact with negative effective weights both recurrently and in the inter-areal interactions, thus breaking Dale's law.

    1. eLife Assessment

      This valuable descriptive study describes the expression of a developmentally relevant transcription factor in the adult Tribolium brain. The evidence supporting the claims is convincing and based on a very detailed and rigorous analysis of light microscopy data, which, however, lacks single-cell resolution. This neuroanatomical study is of interest to the field of insect neural development and neuroscience.

    2. Reviewer #1 (Public review):

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both.

      Strengths:

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells.

      Weaknesses:

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative.

    3. Reviewer #2 (Public review):

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically.

      Strengths:

      Thorough and meticulous application of state-of-the-art anatomical methods in a non-standard laboratory organism.

      Weaknesses:

      No weaknesses were identified by this reviewer.

      Comments:

      I don't really have any major suggestions at all. Loved the work.

      There is only one tiny nitpicking aspect:

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively."

      MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere.

      https://pubmed.ncbi.nlm.nih.gov/10454381/

      such as, e.g., visual pattern learning in the CX

      https://pubmed.ncbi.nlm.nih.gov/16452971/

      or motor learning in motor neurons

      https://pubmed.ncbi.nlm.nih.gov/38779314/

      or ventral ganglion, antennal lobes, and median bundle for place learning:

      https://pubmed.ncbi.nlm.nih.gov/10706599/

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest.

    4. Author response:

      We are very happy that our work was positively received by the reviewers and editors and we are looking forward sharing our results via eLife. 

      We have added more details on the generation of the Tribolium brainbow-lines and we have submitted the respective plasmids to Addgene and give the respective IDs. Some additional minor changes were done to make the text more clear. 

      Public Reviews: 

      Reviewer #1 (Public review): 

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both. 

      Strengths: 

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain. 

      We would want to add that this work establishes and introduces the brainbow system for the first time in an arthropod outside Drosophila melanogaster and that we are the first (outside flies) to relate the expression of a neural transcription factor with neural projection and neurotransmitter content.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells. 

      Weaknesses: 

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects. 

      We kindly disagree with the first statement: not all cells of the enhancer trap are labelled but a subset. Therefore, we call it “sparse labelling” in our manuscript while we do not reach “single cell labelling”, which admittedly limits both precision and use.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function. 

      Previously, we published that this gene has an important function in neural development during embryogenesis. Actually, we have done extensive RNAi experiments to test for an e ect during postembryonic development. We found surprisingly small defects when looking at alterations in several imaging lines. However, we found some changes in behavior. Given the extensive data presented in the current paper, we decided to publish these functional data (another 12 figures/suppl. figures) separately. 

      We also note that the identity/function of neurons is determined by a mix of transcription factors. Disentangling the individual role of each of those transcription factors indeed is an exciting question and a major endeavor beyond the scope of this paper.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative. 

      Indeed, we do not reach single cell resolution, which is below the standards of fly neurobiology. However, compared with all other arthropods we reach a unique level of precision. Specifically, we are the only ones outside fly research that relate the expression of a developmental transcription factor to neural projection and neurotransmitter content. 

      We also think that combining our transgenic line with dopamine-expression was su icient to compare the labelled cells to fly neurons. From what we saw in that analysis, we feel that most homology assessments of single neurons across such large evolutionary distances will remain hypothetical to some degree.

      Reviewer #2 (Public review): 

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically. 

      Strengths: 

      Thorough and meticulous application of state-of-the-art anatomical methods in a nonstandard laboratory organism. 

      Thank you for this encouraging comment. 

      Weaknesses: 

      No weaknesses were identified by this reviewer. 

      Comments: 

      I don't really have any major suggestions at all. Loved the work. There is only one tiny nitpicking aspect: 

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively." 

      MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere. https://pubmed.ncbi.nlm.nih.gov/10454381/ 

      such as, e.g., visual pattern learning in the CX https://pubmed.ncbi.nlm.nih.gov/16452971/

      or motor learning in motor neurons  https://pubmed.ncbi.nlm.nih.gov/38779314/

      or ventral ganglion, antennal lobes, and median bundle for place learning: https://pubmed.ncbi.nlm.nih.gov/10706599/ 

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest. 

      Thanks for this clarification – we have rephrased:

      "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023). This relates to the mushroom bodies’ function in olfactory memory, and the function of the central complex in visual pattern learning and goal directed navigation, respectively."

    1. eLife Assessment

      This paper describes a valuable tool for the detection and analysis of dentate spikes, network events in the dentate gyrus that are common yet understudied. This tool could help standardize dentate spike detection and analysis across labs and is therefore likely to be of interest to hippocampal neurophysiologists. However, the strength of evidence for its broad usefulness was viewed as incomplete, due to several identified bugs in the program and insufficient explanations of parameter selection and methods.

    2. Reviewer #1 (Public review):

      Summary:

      Esfahany et al. describe a new platform (Toothy) to identify and analyze dentate spikes and sharp wave ripples from silicon probe electrophysiology data. The goal is to facilitate and standardize the extraction of DS1 and DS2 events, which have highly variable properties across recordings from different labs. The manuscript describes the basic workflow of the Toothy pipeline, including loading data, assigning channels along a linear probe, customizing parameters, selecting ideal channels for analysis, and classifying DS1 and DS2 events.

      Strengths:

      The manuscript is clear and easy to follow and does a good job of describing the platform. Overall, this will be a useful analysis pipeline that can help to standardize DS analysis across labs and datasets.

      Weaknesses:

      The current version has several bugs that prevent analysis, and the documentation of analysis parameters needs to be improved.

      (1) In limited testing, the pipeline had several bugs, and I was not able to complete the full analysis of a dataset. Loading data from .mat or .npy files gave errors (it seemed that the metadata was not loaded correctly from the pop-up window). I was able to load a .nwb file, which worked well. The probe configuration tool was a bit difficult to understand, and there was not much documentation to help, although it worked when simply entering the x-y coordinates of the channels. It also crashed several times while trying to make a probe configuration due to it trying to save when a small typo was briefly entered. The initial analysis worked well, and the auto-selected channels matched our recording notes and seemed appropriate. DSs and ripples were extracted. An error came when trying to classify DSs, and the program repeatedly crashed across a variety of parameters. Overall, parts of the pipeline worked well, but others had significant bugs that need to be addressed.

      (2) The authors should provide test data that can be run through the pipeline. Ideally, this could use a variety of data types, probes, and conditions so that it is clear how they differ.

      (3) There are a lot of parameters that can be adjusted, but very little information about how they are chosen and what goes into parameter selection for a dataset. Additional documentation with more information on adjustable parameters, channel selection, and best practices would help improve the utility of the tool. Ideally, this could also integrate citations (either in the manuscript or documentation) to support some of the choices made during parameter selection.

      (4) There is no validation presented against other analysis methods or datasets. While there is no ground truth of when DSs occur, this may limit the ability of this tool to become the standard for DS analysis. A section comparing the analysis used in the pipeline to other published analyses would be helpful.

      (5) In the manuscript, it would be helpful to further describe the rationale for initially detecting DSs and SPW-Rs on all channels, when they are network events that occur across channels.

      (6) A section on what hardware and software are necessary to run the pipeline should be added.

    3. Reviewer #2 (Public review):

      Summary:

      This work provides an open-source, Python-based, graphical user interface for curating the detection and classification of dentate spikes (DSs) from hippocampal local field potential (LFP) recordings. The tool may also be used to detect, but not classify, sharp wave-ripples (SPW-Rs). The tool utilizes previously published Python packages for loading LFP files and creating experiment-specific probe objects. Detection and classification parameters are clearly defined and logged in a parameter file before starting processing. Once LFP data has been mapped to the probe object, event detection occurs across all channels. DSs are detected as qualifying peaks in the filtered DS band LFP, while SPW-Rs are detected as qualifying peaks in the filtered ripple band amplitude envelope. An initial curation step allows visualization of the LFP, instantaneous current source density (CSD), and depth-by-frequency band power plots for determining the approximate channel locations of key anatomical regions (i.e., CA1, the hippocampal fissure, and the hilus of the dentate gyrus). The optimal channel for detection is further refined in the next step by comparing event waveforms and quality metrics across channels. Artifacts and noisy waveforms can also be manually excluded during this step. Finally, DSs detected from the optimal channel are classified by computing the CSD profile around events and then clustering the first two principal components of all CSDs. The authors claim that this customizable tool will standardize DS detection and classification.

      Strengths:

      Toothy's detection and classification algorithms are appropriate and well-validated in the literature. The ability to change many parameters, the CSD calculation method, and clustering algorithm is helpful for precise replication of methodology that has varied previously. Default parameters optimized for mouse recordings provide a standardized starting point for rodent researchers.

      The authors' commitments to transparency and user-friendliness are to be commended (e.g., clear instructions, defined and logged parameters, multiple visualization options, etc.) and are likely to be appreciated by new users. Researchers with little-to-no coding experience should find this tool especially powerful for jumpstarting their own DS analyses.

      While not the focus of the paper, the capability to detect SPW-Rs provides an additional use case for Toothy and streamlines simultaneous analysis of SPW-Rs and DSs.

      Weaknesses:

      I encountered unexpected errors while trying to load LFP data into Toothy for testing, indicating that the "data ingestion" stage of Toothy requires minor code revision.

      Toothy's utility for recordings that do not produce an LFP depth profile is unclear. According to the authors, Toothy allows probe designs with irregular spatial sampling (e.g., tetrodes) to be used. However, recording from a linear probe with electrodes spanning from approximately the hippocampal fissure to the hilus of the dentate gyrus is required for Toothy's full functionality. For example, Toothy uses a DS type classification algorithm that relies on sufficiently sampled CSD depth profiles that tetrode recordings cannot provide. As such, usage is currently restricted to detection only for certain recording setups.

      The documentation on Toothy's output could be improved. Specifically, the work does not state which files different data are saved to or list the properties saved per detected event. Furthermore, the work does not discuss the potential importance of DS properties that are saved besides those related to the timing of the DS and its type.

    4. Reviewer #3 (Public review):

      Summary:

      Esfahany et al present a novel, UI-based tool to detect dentate spikes from hippocampal local field potential recordings, called Toothy. Toothy is easily accessible, compatible with many popular recording formats, and guides users entirely via UI through the dentate spike curation and analysis process. The functional and interactive visualizations enable users to gain a detailed understanding of their data and rigorously analyze dentate spike phenomena. This tool will be broadly useful for anyone who studies hippocampal electrophysiology. Furthermore, by expanding access to dentate spike analysis, it may encourage more scientists to explore this understudied but critical phenomenon.

      Strengths:

      (1) Toothy provides several ways for users to interact directly with parameters, revealing the ramifications of these choices. Most parameters are adjustable and made obvious via a UI panel. Their effects are then visualized across channels and individual events. This will help users think critically when selecting parameters.

      (2) Toothy is fully UI-based and pip-installable, lowering the barrier to entry far below what most electrophysiology analysis tools offer.

      (3) The channel selection tool is broadly useful for identifying DG hilus and CA1 pyramidal locations. Since subregional and laminar localization of electrode sites is critical to correctly interpret hippocampal recordings, this tool could be more generally used to identify site locations across the hippocampus.

      Weaknesses:

      (1) The rationale behind parameter choices is not explained. In order to function "not as a black-box detector", as the authors state, all initial parameter choices should be explained with citations. If possible, these citations would also be available from Toothy directly, alongside citations describing alternative parameter choices. This will help users make informed choices. For instance, a user analyzing data from rats would need to adjust the default ripple frequency band upwards (150-250Hz), and would benefit from guidance to adjust this properly.

      (2) The Results describe the functions of Toothy from the perspective of the user, but there is no Methods section describing what Toothy does between UI displays. This would allow readers to compare the tool directly to analysis pipelines as described in the Methods sections from other papers. Particular attention should be paid to justifying the analysis decisions that cannot be changed by the user, such as detecting events off of a single representative channel instead of across a consensus of multiple channels.

      (3) It's unclear whether or how Toothy evaluates data quality to confirm that its analyses return interpretable results. At a minimum, the tool should confirm adequate sampling rate (e.g. <=1kHz) and inter-site spacing for CSD (e.g. <=50um).

      (4) The paper does not put Toothy into context among the other common open-source electrophysiology analysis toolboxes. Consider Rippl-AI (Navas-Olive & Rubio et al, 2024) or pynapple (Viejo et al, 2023), to give a few examples. The paper would be strengthened by addressing how Toothy extends beyond the capacities of these other tools and how Toothy can be integrated into a workflow that also uses these other tools.

    5. Author response:

      We thank the editors and reviewers for their thoughtful and constructive assessment of Toothy, and for recognizing it as a potentially valuable resource for standardizing dentate spike (DS) analysis across labs. We are especially glad that the reviewers found the manuscript clear and easy to follow, judged the detection and classification algorithms to be appropriate and well-validated, and appreciated the tool's graphical user interface (GUI) based, pip-installable design for lowering the barrier to entry for DS analysis.

      We also understand the concerns raised. Most importantly, we will resolve the data-ingestion and classification errors that reviewers encountered and release an updated version of Toothy that we have verified end-to-end across input formats and datasets. Alongside this, we will provide downloadable demo dataset(s) spanning multiple file formats, probe types, and recording conditions, so that users can confirm a correct installation and see how these cases differ.

      To make the pipeline more transparent, we will add a section describing what Toothy does between user steps, including the rationale for decisions users cannot change, such as detection from a single representative channel. We will also expand the documentation of parameter choices with supporting citations and alternatives, and surface this guidance within Toothy where feasible, consistent with our aim that the tool not function as a black box.

      We will clarify Toothy's scope and current limitations. Recordings with irregular spatial sampling (e.g., tetrodes) are supported for detection but not for CSD-based DS-type classification, which requires a laminar probe spanning approximately the hippocampal fissure to the hilus; we will state this explicitly and evaluate adding an optional waveform-based classification mode (Santiago et al., 2024) to extend type classification to such recordings. We will also add data-quality checks (including sampling rate and inter-electrode spacing) that warn users when a recording may not support reliable results.

      Finally, we will situate Toothy among existing open-source toolboxes, describing how it differs, extends beyond, and interoperates with them, and we will add a comparison of Toothy's outputs to previously published analyses while being explicit about the limits of such comparisons. We will of course also address the remaining technical clarifications and figure edits raised by the reviewers.

      We are confident that addressing these points will make Toothy clearer and more useful to the hippocampal community.

    1. eLife Assessment

      This study reports important and invaluable findings that advance understanding of how attention is distributed between what we look at directly and what lies outside the center of gaze during active visual search. The evidence supporting the main claims is solid, with a large and rich dataset spanning multiple brain areas, although some aspects of the interpretation would benefit from additional controls and clearer separation of attention from eye-movement planning. The work will be of particular interest to researchers studying attention, visual perception, and eye movements behavior.

    2. Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye movement mediated search neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. Detailed association of simultaneously obtained eye movement sequences and neural parameters are well done. These are valuable data which will contribute to our understanding of attentional modulation in visual search.

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from key mid-tier (V4) and higher order (IT, PFC) areas. They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, marked by a high degree of feature and categorical specificity. That is, while attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This provides valuable data for the concept of a foveal-peripheral spatiotemporal attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors and looks towards and away from the target) and statistical rigor make these findings compelling. There will likely be additional future impacts of this study. For example, the eye movement patterns collected in this study may also provide a valuable dataset for future study of understanding search strategies. Goal-directed vs non-goal-directed task comparisons could be designed to test possible circuit models. Although much remains unknown regarding how and where frontal and temporal signals are integrated during active search, these data contribute important guideposts for future models of active visual search.

    3. Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Fig. 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provide a dataset that is well suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Fig. 2). As a result, the reported attentional modulation coincides with preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19) therefore likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Fig. S3) partially mitigate this concern by demonstrating that feature-based modulation persists through saccade execution.

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Fig. 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      [Editors' note: the authors have provided responses to each of these points.]

    4. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to differentiate between foveal and peripheral attentional mechanisms in visual and frontal brain regions in monkeys engaged in a free-gaze visual search task.

      Strengths:

      The manuscript is clearly written, the question is important, and the behavioral task is interesting.

      Weaknesses:

      I have two major concerns.

      (1) The authors interpret divergence in neural responses to target vs nontarget as attention. But it is not. The subject has to attend to both target and nontarget stimuli to determine the stimulus category and thereby decide on the next action. Thus, divergence between target and nontarget responses could reflect categorical discrimination, but I am not sure this can be interpreted as attentional modulation. While it may be tempting to suggest that finding a stimulus of a specific category is "feature attention", analogous to, e.g., attending to the red stimulus, I don't believe this is correct. For the former, the animals have to attend to a stimulus, and examine the stimulus to determine the stimulus category, unlike a simpler discrimination, which may pop out. Given this, I am unconvinced that the interpretations in this manuscript are valid.

      We thank the reviewer for raising this concern. Selective attention is a process of focusing on goal-relevant stimuli (targets) while ignoring irrelevant distractions. Importantly, attentional selection is not limited to simple visual features (e.g., color, shape, or motion); it can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [1, 2], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [3, 4]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [5-8], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [7].

      Similarly, in our study, monkeys were trained to search for images that matched the category of the cue. The neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors. We also included only neural responses occurring prior to fixations associated with target selection, that is, before the monkeys made a behavioral choice, thereby controlling for potential contributions of target detection or decision-related signals to the observed effects.

      We have clarified and addressed this point in the Discussion as follows:

      “Feature-based attention to simple visual features such as color, shape, or motion has been extensively studied [1, 3, 5, 7-9, 11, 12, 64]. Attention can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [65, 66], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [6, 67]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [68-71], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [70]. In this study, the neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors.”

      (2) Regarding the RF classification of foveal and peripheral RFs for IT and PFC, prior work suggests that neurons in IT cortex (especially AIT) and PFC have RFs that largely include the foveal visual field. So, it would be important to include figures that show the RFs of neurons classified as foveal versus peripheral for all three areas.

      We thank the reviewer for raising this important point. We agree with the reviewer that neurons in IT cortex and PFC often have RFs that include the foveal visual field. We did record foveal units with both focal and broad foveal RFs; however, in our analysis we only included neurons with focal foveal RFs to exclude the influence of peripheral stimuli. We defined focal foveal-RF units as those that responded solely to the cue in the foveal region and not to items in the search array presented at least 5° away from the central fixation point, ensuring that their RFs did not extend to these peripheral locations. The items were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their RFs during fixations. By definition, their RFs were restricted to the central point. This is further supported by Fig. S1A-H, which shows no responses to items in the search array at peripheral locations. We have made modifications in the Results and Methods as follows:

      “Notably, the items in the search array were presented at least 5° from the central fixation point and were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their foveal RFs during fixations.”

      And:

      “In this study, our focus was on units with focal foveal RFs and units with localized peripheral RFs. All further analyses were conducted on these units.”

      We modified Fig. 1 to illustrate the RFs of neurons classified as peripheral, which were also characterized in our previous study using the same dataset [9]. The peripheral population exhibits no responses to the central cue (Fig. S1I–T).

      Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at the center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature-selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with a covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye-movement mediated search, neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal, and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, and areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. The detailed association of simultaneously obtained eye movement sequences and neural parameters is well done. These are valuable data that will contribute to our understanding of attentional modulation in visual search.

      Strengths:

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly, the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from multiple areas (V4, IT, PFC). They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, a high degree of feature and categorical specificity. This provides valuable data for the concept of a foveal-peripheral attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors, and looks towards and away from the target) and statistical rigor make these findings quite compelling.

      Weaknesses:

      While the study is generally quite strong, there are a few weaknesses to be addressed.

      (1) Little rationale is provided for recording in the selected areas, V4, IT, and PFC. Given the respective roles in sensory, object recognition, and goal-directed behavior, some rationale for this design should be offered, and commonalities/distinctions between these areas should be discussed.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction as follows:

      “V4 and inferotemporal cortex (IT), as the middle and high-level areas of the ventral visual stream, are important for object recognition and categorization [27-34], and their roles have been extensively studied in central vision. At the neuronal level, however, most investigations have largely neglected their functions during active, free-gaze visual search. The prefrontal cortex, including LPFC, has long been implicated as a source of top-down signals that bias the selection of attended features and modulate visual cortical responses [6, 9, 11, 35-40]. Although target-related visual responses have been reported in IT during visual exploration [41], and target-selective responses have been observed in the human medial temporal lobe (MTL) [42] and medial frontal cortex (MFC) [43] during visual search, these studies did not map the receptive fields (RFs) of recorded neurons.”

      We also added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area—consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      (2) Given the reliance of all analyses on saccadic behavior (towards target/distractor, towards/away from target), additional description and summaries of eye movement behavior during single trials and across trials should be provided.

      We thank the reviewer for this helpful suggestion. We have added a description of saccade behavior to the Results as follows:

      “The mean number of saccades monkeys made to find the target after the onset of the search array was 2.25 ± 1.35 (mean ± SD across trials; Table 1) of correct trials, and the mean saccade amplitude was 7.99° ± 3.58° (mean ± SD across saccades; Table 1). Monkeys could fixate on each distractor or the target freely, provided they did not maintain fixation on the target for longer than 800 ms. Across sessions, 42.44% ± 3.6% of saccades were directed to distractors, 57.56% ± 3.6% to targets, and 12.59% ± 3.46% were saccades away from targets (see our previous studies [44-46] for detailed behavioral analyses).”

      We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search task.

      We have also included Table 1, which summarizes eye movement behaviors.

      (3) The dependency of findings on top-down (categorical & feature-specific) task design should be discussed.

      We thank the reviewer for the suggestion and added a discussion as follows:

      “In this task, attention is strongly guided by top-down goals, which bias processing toward behaviorally relevant features and object categories [2, 50, 51]. Top-down attention, including categorical and feature-specific components, has been shown to modulate neural processing across the visual pathway based on task demands and to originate from distributed frontoparietal control networks [11, 35-38, 40]. Our study provides further insight into the mechanisms of goal-directed visual attention, as it is among the first to demonstrate foveal feature attention effects during free-gaze visual search, as well as the distribution of feature and spatial attention across the entire visual field.”

      Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper, including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Figure 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provides a dataset that is well-suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      We thank the reviewer for the helpful suggestion and apologize for not explicitly providing essential information about the RFs of the units. We added a detailed description of RF properties to the Results as follows:

      “The RFs of these peripheral units were further mapped using a visually guided saccade task and quantified by the number of stimuli that activated each unit (Fig. 1F-K). The eccentricities of the peripheral RFs were 6.22° ± 1.31° (mean ± SD) in V4, 7.04° ± 1.52° in IT, and 6.68° ± 1.56° in LPFC. The sizes of the peripheral RFs were 3.67° ± 1.87° in V4, 6.86° ± 3.11° in IT, and 8.65° ± 3.02° in LPFC. The numbers of items from the search array falling within peripheral RFs were 1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC (also see our previous study [44]).”

      The reviewer is correct that multiple items from the search array did fall within the RFs of peripheral-RF units. However, for focal foveal units, only the fixated stimulus fell within the RF, due to the design of the search array and the definition of these units used in our analyses (see our reply to Reviewer 1, Public Review, Question 2 for details). We agree with the reviewer that attentional modulation is typically stronger when multiple stimuli fall within RFs. In our design, peripheral RFs, on average, contained more stimuli than foveal RFs. Therefore, this difference in RF size would, if anything, be expected to bias toward stronger attentional modulation in peripheral units. This would make our observation conservative, thereby further supporting rather than undermines our main finding of robust feature-based attentional enhancement in foveal units, challenging the prevailing view that such modulation is predominantly peripheral. However, we agree that, when comparing the latency of attentional effects across brain regions in Fig. 3, we cannot rule out the influence of the number of stimuli arising from differences in RF size.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Figure 2). As a result, the reported attentional modulation coincides with the preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19), therefore, likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Figure S3) partially mitigate this concern by demonstrating that featurebased modulation persists through saccade execution.

      We thank the reviewer for raising this important question. We agree that the temporal overlap of visual, motor planning, target recognition, and behavioral relevance signals with attention can result in mixed activity, which needs to be dissociated. Therefore, when calculating feature-based attention, we did implement a series of controls. We added a discussion as follows:

      “A major challenge in interpreting neural activity related to attentional modulation is the inherent temporal overlap of visual processing, motor planning, and target recognition signals in the free-gaze visual search task [73]. To isolate genuine feature-based attention from potential confounds, we applied several stringent analytical constraints, consistent with prior studies [3, 5, 6]. Specifically, by restricting our analysis to fixations where the subsequent saccade was directed away from the RFs, we dissociated attentional modulation from the preparatory motor activity associated with saccade execution. Furthermore, by comparing responses to the same physical stimulus alternating its role as a target or distractor across trials we eliminated any potential bias introduced by stimulus identity or physical category. We restricted our analysis to fixations preceding target selection that is, before the monkeys made a behavioral choice to minimize contributions from target detection or decision-related signals.”

      We thank the reviewer for pointing out the issue of different timescales for target versus distractor fixations. To address this, we conducted a control analysis by computing foveal feature-based attentional modulation using fixations on targets and distractors with matched fixation durations. We obtained similar results. We have updated Fig. S2 to include this control analysis.

      We also clarified this point in the Results as follows:

      “We also obtained similar results when controlling for fixation durations on targets and distractors (i.e., there was no significant difference between fixation durations on targets and distractors; Wilcoxon signed-rank test, P > 0.05; Fig. S2K–P).”

      Lastly, as the reviewer correctly pointed out, the interpretation that foveal feature-based attention facilitates prolonged fixation on the target was not supported. We have revised the Results as follows:

      “On average, target fixations (256.69 ± 197.44 ms [mean ± SD]) were significantly longer than distractor fixations (156.26 ± 45.94 ms; Wilcoxon rank-sum test, P < 0.0001), and during these prolonged target fixation, foveal feature-based attention modulation was consistently observed.”

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Figures 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      We thank the reviewer for this important question. We performed a directional control analysis by computing spatial attentional modulation using paired fixations from the attention-in and attention-out conditions. Only saccades directed in nearly opposite directions—defined as having a saccade direction angle ≥ 170° within the 0–180° range—were included. We obtained similar results (Author response image 1). 

      Author response image 1.

      Peripheral spatial attentional modulation in V4, IT, and LPFC. Population response to stimuli followed by saccades directed into their RFs (attention in) versus directed approximately opposite and outside their RFs (attention out), shown for V4 (A), IT (B), and LPFC (C). Shaded area denotes ±SEM across units.

      We did control for feature-based attention when calculating spatial attentional modulation. We apologize for the lack of clarity and have added a description of this control to the Methods as follows:

      “The saccade-target stimulus in the RF during attention-in fixations was matched to a stimulus in the same location during attention-out fixations; in both conditions, this stimulus always served as a distractor for that trial, except in the “Distractor fixations to T” condition (Fig. 5 and Fig. S4), in which it instead served as the target. This design eliminates differences due to feature-based attention between the attention-in and attention-out conditions.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3C: Unclear how to compare LPFC vs V4 for foveal units since only data from peripheral LPFC is shown?

      We thank the reviewer for pointing out this mistake. In Fig. 3C, we only compared LPFC peripheral units, V4 peripheral units, and V4 foveal units. We have corrected this in the legend of Fig. 3 as follows:

      “Shown are cumulative distributions of feature-attention effect latencies, computed from individual foveal face-, house-, and non-selective units in V4 and IT, and from peripheral non-selective units in V4,

      IT, and LPFC.”

      (2) On page 8, last para: For units with peripheral RFs ... Is this controlled for whether the saccade is to targets or to distractors?

      We thank the reviewer for the question. We indeed addressed this concern by separating fixations based on whether the subsequent saccade was directed to a target or a distractor, and by analyzing attention modulation within each condition. Therefore, attention effects were evaluated while holding the saccade destination constant, effectively controlling for potential confounds related to saccade target selection.

      (3) Page 9: The authors find that target fixations were longer than distractor fixations and conclude that this supports the idea that foveal feature-based attention increases fixation duration, but this interpretation is pure conjecture, and there is no experimental manipulation presented in this paper that helps to establish this interpretation.

      We thank the reviewer for this important comment. We agree that this observation does not, by itself, support our original interpretation, and we have modified it in the Results. Please refer to the last paragraph of our Reply to Question 2 from Reviewer 3 (Public Review).

      (4) Data analysis: receptive field. The authors state that visual response to a cue and the stimulus array was assessed during the 0-200 ms window after stimulus onset. However, after the array onset, the animal could saccade within the 200 ms window. How do the authors ensure uniform stimulation during the 0-200 ms window?

      We thank the reviewer for this question. The activity of units in V4, IT, and LPFC within the 200 ms window after array onset primarily reflected visual stimulation prior to saccades, because typical saccade latencies were approximately 150–200 ms, and the response onset latencies of these units were around 50 ms.

      (5) On page 19, the authors state that to assess feature attention in peripheral RFs, they divided trials into target and distractor fixations. In the former, there was a target in the neuron's RF. This is confusing. I assume target fixations imply fixating on a target, but the authors may mean fixations where a target is in the RF. Please clarify.

      We thank the reviewer for pointing out this confusion. In the original manuscript, we intended to sort fixations by whether a target stimulus was located within the unit’s peripheral RF. To avoid further confusion, we have revised the description in the Methods as follows:

      “we sorted fixations during the search period, following a procedure similar to that in our previous study [5], into two types: “target” – a target stimulus was located within the unit’s peripheral RF; and “distractor” – the same stimulus appeared in the same peripheral RF location but served as a distractor.”

      (6) Figure S1: Are these example units? How many trials? SEM? The sharp rise and no noise are inconsistent; the former suggests minimal smoothing, while the latter suggests lots of smoothing.

      We thank the reviewer for these questions. We showed average responses across all units in Fig. S1. On average, there were 941.79 ± 182.56 trials (mean ± SD across sessions). Shaded areas indicate ±SEM across units. The sharp rise reflects the synchronous response of neurons to the stimulus, while the smooth appearance and low noise result from averaging across a very large number of units and trials.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) One weakness of this manuscript is the lack of a rationale for choosing V4, IT, and PFC. Specifically, what are the predictions of the roles of these respective areas in the integration of current and peripheral (future foveal) views? There is a significant literature linking the pre-saccadic peripheral stimulus and the post-saccadic foveal stimulus, suggesting that both spatial and temporal integration occur. However, whether such integration occurs at high or low cortical levels is unknown. By recording from mid-tier (V4) and high-order areas (IT, PFC), the authors have an opportunity to address this question. However, there is no mention of this topic, either in the introduction, results, or discussion. I find this omission surprising. At the very least, it should contribute to experimental design rationale and some discussion.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction and a discussion about this integration. Please refer to our Reply to Question 1 from Reviewer 2 (Public Review).

      (2) As both behavior and neural recordings are collected, a figure on saccadic patterns would enhance the reader's understanding. Questions that come to mind are: What does a single search trial look like? How many saccades are there per trial? How often is the target identified after 1, 2, 3, etc saccades? What is the average size of a saccade? Although this is not a study of search strategy per se, a modicum of description of the search sequences would provide context on the behavior. I suggest an illustration of one or more sample trials; a summary of saccade behavior would also be helpful for understanding the data in relation to behavioral performance.

      We thank the reviewer for this helpful suggestion. We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search, providing an example of a single search trial. Additionally, we have added a description of saccade behavior to the Results and included Table 1, which summarizes eye movement behavior. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for further details.

      (3) "Consistently, the probability of making a saccade to a peripheral target was higher following distractor fixations (75.22%) than following target fixations (48.44%, or 63.49% after probability calibration; see Methods), indicating the important role of peripheral feature-based attention in guiding eye movements" It should be noted that this target-oriented visual search is fundamentally a top down task. Once the target is found, the reward is obtained; saccades to distractors are not rewarded, so saccades are more likely. So certainly this task design would increase the post-distractor saccades and decrease the number of post-target saccades. Please clarify the behavioral paradigm: once a reward is obtained, does the task continue, or is a new trial initiated?

      We apologize for the confusion regarding the behavioral paradigm. We would like to clarify that when the target was found and fixated for 800 ms, the reward was delivered and no further saccades occurred. However, if the target was not fixated for 800 ms, the search could continue. It is worth noting that the target fixations in our analyses were restricted to those occurring during ongoing search behavior, excluding target fixations associated with trial termination and reward delivery. Moreover, we compared the probability of making a saccade to the target, rather than the absolute number of saccades, following these fixations. We have modified the Results for clarification, as follows:

      “Two monkeys performed a category-based visual search task, where their objective was to fixate on one of the two search targets that matched the category of the cue (Fig. 1A, B). Specifically, the monkeys were presented with a central fixation point for 400 ms, followed by a cue lasting 500-1300 ms. After a 500 ms delay, a search array appeared with 11 items, including two targets, randomly chosen from 20 possible locations (Fig. 1E). The monkeys had 4000 ms to find one target and maintain fixation on it for 800 ms to earn a juice reward. Fixating on either target completed the trial, and the monkeys did not search for the second target. A new trial began after the reward. It is worth noting that the two target stimuli matched the category of the cue but were different images. The monkeys were required to maintain fixation throughout the cue and delay periods. During search, however, eye movements were unconstrained, and monkeys could revisit each search distractor or target as long as they did not fixate on a target for 800 ms.”

      (4) The fact that there are many more peripheral units in LPFC suggests that this is a region of foveal/periph integration. Combined with the finding that the LPFC leads the attentional effects, this should be a discussion point.

      We thank the reviewer for the suggestion and we added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      Minor comments:

      (1) Figures 2A-D. "These face-selective units also showed slightly enhanced responses to house targets in IT (P < 0.05), but not in V4 (P = 0.89)." It does not appear enhanced.

      We agree with the reviewer that the effect is modest and does not appear strongly enhanced. However, the average response in the 150–225 ms time window to the house target was significantly higher than that to the house distractor in IT face-selective units (Wilcoxon signed-rank test, P = 0.042). We modified the description in the Results as follows: 

      “These face-selective units also showed weakly but significantly enhanced responses to house targets in IT (P < 0.05)”

      (2) Figure 3. For population comparison, a bootstrapped null distribution was used, and a 2-sided permutation test was used to determine the latency difference between the target and distractor; please show these results (described in text) in a figure. Figures 3A-C are described as the latency of individual units. So each of these graphs is the mean of multiple units? So this is also a population analysis? What is the difference between these two comparisons? This is somewhat confusing.

      We apologize for the confusion and thank the reviewer for pointing this out. Each panel in Fig. 3 shows the cumulative distribution of latencies across individual units within each brain region, reflecting the variability of response timing across single neurons. For this analysis, we first calculate the latency of each unit separately. In contrast, population-level latency is measured from the averaged responses of all units within each region (Fig. 2), which captures the overall timing of the population response rather than individual variability. Statistical comparisons at the population level are performed using a two-sided permutation test. We modified Fig. 2 to better illustrate the population-level latency results.

      (3) Did peripheral RFs span more than a single stimulus in the array? If so, how does this impact the interpretation of Figure 5?

      We thank the reviewer for pointing this out. The reviewer is correct that, in peripheral RFs, more than one stimulus from the search array could fall within the receptive field (1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC). We controlled for this in our analysis of both feature-based and spatial attention effects for peripheral units in Fig. 5. For feature-based attention, we performed the analysis in a stimulus-by-stimulus manner within each category (house and face), such that when a given stimulus served as the target, it was the only target within the RF, and when it served as a distractor, it was the only distractor of its category within the RF. Although additional distractor could still fall within the RF, their identities were random across conditions and thus would be averaged out. A similar approach was applied to spatial attention, where the stimulus-by-stimulus comparison was extended across all four categories, and attention-out stimuli were paired with the corresponding saccade-target stimuli in the attention-in condition, with the effects of other randomly present distractors averaged out. Therefore, the effects shown in Fig. 5 reflect comparisons at the level of individual stimulus, minimizing confounds from other stimuli within the RF.

      (4) Figure 5G: "during "Target fixations to D", there was no significant feature attentional enhancement in response to the peripheral target (Wilcoxon signed-rank test, P > 0.05; Figure 5G-I left panels). It appears that there is some effect of spatial attention during Target Fix to D trials.

      We thank the reviewer for pointing this out and have revised the Results as follows:

      “We further found that spatial attentional enhancements to the saccade target were reduced during target fixations compared to distractor fixations in V4 and IT when activity was aligned to fixation onset (Wilcoxon rank-sum test, P < 0.05; Fig. 5G, H versus Fig. 5A, B), although this effect was not completely abolished.”

      (5) The specific areas of IT and LPFC that were recorded should, as much as possible, be mentioned.

      We thank the reviewer for the helpful suggestions and have added a description of the specific IT and LPFC recording sites to the Methods as follows:

      “Recordings in IT spanned the central IT cortex, encompassing the area between the anterior middle temporal sulcus (AMTS) and the posterior middle temporal sulcus (PMTS), including TE and TEO. Recordings in LPFC were located anterior to the arcuate sulcus (AS) and lateral to the principal sulcus (PS), mainly covering areas 45 and 44.”

      (6) It is often difficult to distinguish the different lines, e.g., red solid vs red dotted, due to their overlap. Would the removal of the error band make this clearer? If so, could put full figure with error bands in the Supplementary Figure.

      We thank the reviewer for this helpful suggestion. To improve visual clarity, we adjusted Fig. 6, Fig. 7, Fig. S2, Fig. S3, Fig. S4, and Fig. S6 by changing the line styles and placing the shaded error bands beneath the traces, allowing the lines to remain clearly visible despite overlap.

      (7) For easy access, the number of saccades to/from targets/distractors should be put into a table.

      We thank the reviewer for the suggestion. We calculated the probability of saccades to and from targets and distractors for each session and report the mean ± SD across sessions in Table 1, as the mean number of saccades per trial was only 2.3. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for Table 1.

      Reference

      (1) O'Craven, K.M., P.E. Downing, and N. Kanwisher, fMRI evidence for objects as the units of attentional selection. Nature, 1999. 401(6753): p. 584-7.

      (2) Baldauf, D. and R. Desimone, Neural mechanisms of object-based attention. Science, 2014. 344(6182): p. 424-7.

      (3) Hayden, B.Y. and J.L. Gallant, Combined effects of spatial and feature-based attention on responses of V4 neurons. Vision Res, 2009. 49(10): p. 1182-7.

      (4) Bichot, N.P., et al., A Source for Feature-Based Attention in the Prefrontal Cortex. Neuron, 2015. 88(4): p. 832-844.

      (5) Reddy, L. and N. Kanwisher, Category selectivity in the ventral visual pathway confers robustness to clutter and diverted attention. Curr Biol, 2007. 17(23): p. 2067-72.

      (6) Peelen, M.V., L. Fei-Fei, and S. Kastner, Neural mechanisms of rapid natural scene categorization in human visual cortex. Nature, 2009. 460(7251): p. 94-7.

      (7) Cukur, T., et al., Attention during natural vision warps semantic representation across the human brain. Nat Neurosci, 2013. 16(6): p. 763-70.

      (8) Keller, A.S., et al., Attention enhances category representations across the brain with strengthened residual correlations to ventral temporal cortex. Neuroimage, 2022. 249: p. 118900.

      (9) Zhang, J., et al., Behavioral and Neural Mechanisms of Face-Specific Attention during GoalDirected Visual Search. The Journal of Neuroscience, 2024. 44(46): p. e1299242024.

      (10) Bichot, N.P., A.F. Rossi, and R. Desimone, Parallel and serial neural mechanisms for visual search in macaque area V4. Science, 2005. 308(5721): p. 529-534.

      (11) Bichot, N.P., et al., The role of prefrontal cortex in the control of feature attention in area V4. Nat Commun, 2019. 10(1): p. 5727.

      (10) Cohen, M.R. and J.H. Maunsell, Using neuronal populations to study the mechanisms underlying spatial and feature attention. Neuron, 2011. 70(6): p. 1192-204.

      (11) Maunsell, J.H. and S. Treue, Feature-based attention in visual cortex. Trends Neurosci, 2006. 29(6): p. 317-22.

      (12) McAdams, C.J. and J.H. Maunsell, Attention to both space and feature modulates neuronal responses in macaque area V4. J Neurophysiol, 2000. 83(3): p. 1751-5.

      (13) Motter, B.C., Saccadic momentum and attentive control in V4 neurons during visual search. J Vis, 2018. 18(11): p. 16.

      (14) Sapountzis, P., S. Paneri, and G.G. Gregoriou, Distinct roles of prefrontal and parietal areas in the encoding of attentional priority. Proc Natl Acad Sci U S A, 2018. 115(37): p. E8755-E8764.

      (15) Treue, S. and J.C. Martinez Trujillo, Feature-based attention influences motion processing gain in macaque visual cortex. Nature, 1999. 399(6736): p. 575-9.

      (16) Zhou, H. and R. Desimone, Feature-based attention in the frontal eye field and area V4 during visual search. Neuron, 2011. 70(6): p. 1205-17.

    1. eLife Assessment

      This valuable study characterizes how antibody responses converge on similar functional solutions despite diverse genetic backgrounds, providing a resource that is of importance in understanding immune responses and informing vaccine research. The evidence is solid, with extensive and well-executed analyses supporting the primary findings, although broader conclusions regarding vaccine design would benefit from more cautious interpretation and fuller discussion of the study's limitations. The work will be of interest to researchers studying antibody responses, viral evolution, and vaccine development.

    2. Reviewer #1 (Public review):

      Summary:

      Based on previous work showing that viral evolution follows reproducible patterns in diverse animals, the authors sought to examine whether the antibody response operates under similar constraints. By analyzing over 17,000 B cells isolated from 6 monkeys at 3 different time points, the authors convincingly show that the immune response does follow specific patterns of responses to different classes of epitopes based on the infecting virus. Moreover, each of these clusters has characteristic (cross-) binding and neutralization properties. Importantly, these classes are independent of the underlying immunogenetics, which (as expected) vary significantly between monkeys. This last point is particularly relevant for vaccine design, as it means that immunogens may not need to be as narrowly focused on specific germline genes as previously thought.

      Strengths:

      The large number of B cells cultured for this study is a particular strength, as is the fact that they were isolated in an antigen-unbiased fashion. The experiments are well-designed and comprehensive.

      Weaknesses:

      The genetic element is a relatively minor component overall and more qualitative than quantitative. It would be nice to investigate other properties of the repertoire like CDRH3 length and possible public clones, as well.

    3. Reviewer #2 (Public review):

      Summary:

      Song et al. comprehensively analyzed the SHIV-infected macaque B cell repertoires and commonalities among their antibody responses, despite their diverse genetic background. They suggest these studies would inform HIV-1 vaccine design.

      Strengths:

      This study is well-designed and used proper analysis methods, and the figures are clear and effectively presented.

      Weaknesses:

      However, it tends to overstate its novelty and significance, emphasizing points that are relatively obvious (e.g., different classes of antibodies can recognize a common epitope) and appears to have been overwritten and unnecessarily fancy ("conceptually analogous to ecomorph evolution", "epitopic convergence"). Moreover, some limitations of the rhesus macaque model and the differences between bnAbs and nAbs should be discussed. That said, the underlying data are solid and important in their detail, and the manuscript will be a useful resource for HIV-1 vaccine and pathogen studies.

    1. eLife Assessment

      This is an important study reporting a new phenotype for a gene cluster that has previously been associated with the responses of the Gram-negative opportunistic pathogen Pseudomonas aeruginosa to flow fluid. Expression of the froABCD gene cluster is induced by HOCl in vitro and by activated immune cells, which produce these types of reactive chlorine species and the evidence presented by the authors is in many places convincing. Overall, the authors have been responsive to the previous review, although the exact mechanism of fro-induction by HOCl remains unclear. The high cysteine- and methionine content of the anti-sigma factor FroI hints at a direct oxidative modification of this protein during activation of the operon and a corneal infection model shows that fro is upregulated in P. aeruginosa 20 h after infection, but the evidence that HOCl is the causative agent of fro upregulation under these conditions is at present circumstantial. This study is of interest to infection biologists interested in mechanisms of bacterial pathogenicity.

    2. Reviewer #1 (Public review):

      Summary:

      Foik et al. report that hypochlorous acid, a reactive chlorine species generated during host defense, activates the transcription of the froABCD in P. aeruginosa. This gene cluster had previously been associated with a potential role during flow of fluids and appears to be regulated by the sigma factor FroR and its anti-sigma factor FroI. In the present study, the authors show that froABCD is expressed both in neutrophils and macrophages, which they claim is likely a result of HOCl but not H2O2 production. Fro expression is also induced in a murine model of corneal infection, which is characterized by immune cells invasion. Expression of the fro system can be quenched by several antioxidants, such as methionine, cysteine, and others. FroR-deficient cells that lack froABCD expression during HOCl stress, appear more sensitive to the oxidant.

      Strengths:

      The authors provide a number of data supporting their claim that transcription of the froABCD system is induced by reactive chlorine species. This was shown by RNAseq, qRT-PCR, and through microscopy using a transcriptional reporter fusion. Likewise, elevated expression of froABCD was shown in vitro and in vivo, excluding potential in vitro artifacts. The manuscript, while mostly descriptive, is easy to follow and the data were presented clearly and convincingly. The authors have also been responsive to concerns from the previous review.

      Weaknesses:

      (1) Line 10: "HOCl preferentially oxidizes....". Please consider modifying the language to: "the second-order rate constant of HOCl is significantly higher with Met/Cys compared to other aa."

      (2) I am not sure I completely understand Fig 1B. Is the promoter right upstream of yfp or is yfp located downstream of froA? If the latter is the case, wouldn't this be a translational fusion?

      (3) My previous comment regarding why fro expression is higher during phagocytosis in macrophages compared to neutrophils has been somewhat (albeit not convincingly), addressed by the authors in the response to the reviewer, but this discussion should be part of the manuscript as the macrophage data were shown.

      (4) Line 122: The statement "The degree of fro inhibition by 4-ABAH...." is incorrect unless the authors can provide experimental evidence. Fro expression is not upregulated because MPO is inhibited by 4-ABAH, which results in less hOCL production.

      (5) Can Supp Fig. 1 be quantified in a similar way it was done for HOCl to allow for a better comparison if HOCl or flow is the more potent inducer?

      (6) Overall, the fro expression (YFP/mCherry) seems highly variable for treatment with HOCl (Fig. 2C: ~65; 2D: ~20; why is fro expression 3x lower?

      (7) The authors should provide evidence that N-chlorotaurine can activate fro expression also. They said they weren't able to obtain chlorinated taurine, but this is quite simple to produce: PMCID: PMC1219228

      (8) Fig. 4 supplement 1: Please provide concentrations for the oxidants used in these experiments.

      (9) Lines 251/252: change to: upregulation of instead of in

      (10) Chaperones and other heat-shock genes are more upregulated in ∆froR, indicating elevated HOCl-mediated oxidative damage, which supports their findings.

      (11) Complementation of ∆froR is missing

      (12) Line 198: The growth experiment at 4 uM shows differences between WT and mutant, but at 2 uM cells showed already low fro expression due to cell death (which has not been proven by CFU counts). This discrepancy should at least be discussed.

      (13) The critical in vitro experiment is missing: does purified FroI get oxidized by HOCl and dissociated from FroR?

      (14) Lines: 350-355: The claim that the fro system is the first-line defense is unproven.

    3. Reviewer #2 (Public review):

      Summary:

      Foik et al. studied the regulation of the fro operon in response to HOCl, an oxidant derived from immune cells, especially neutrophils. They use a transcriptional fusion of YFP to the froA promotor in an mCherry expressing P. aeruginosa strain to determine fro-induction under the microscope. They use this system to study fro expression in medium, in the presence of neutrophils and macrophages, neutrophil-conditioned medium, and several chemical stimuli, including NaCl, HOCl, hydrogen peroxide, nitric acid, hydrochloric acid, and sodium hydroxide. They also use a corneal infection model to demonstrate that froA is upregulated in P. aeruginosa 20 h post infection and perform transcriptional analyses in WT and a froR mutant in response to HOCl.

      Strengths:

      Their data clearly shows that HOCl is a strong inducer of the fro Operon. Addition of HOCl-quenching chemicals together with HOCl abrogates the response. They also show that a froR mutant is more susceptible to HOCl than WT. Their transcriptomic data reveals genes under control of the FroR/FroI sigma factor/anti sigma factor system.

      Weaknesses:

      Although the presented evidence is mostly solid, some of their findings need to be evaluated more carefully; explaining the rationale behind some of the experiments might enhance the article; and some of the models proposed by the authors seem far-fetched, as outlined below:

      Unexpected outcomes and open questions for future research:

      (1) As outlined above, HOCl seems to be the main inducer of the fro operon. Interestingly, during interaction with immune cells, macrophages and neutrophils seem to induce a reporter gene under fro control in a similar manner, although macrophages are generally thought to produce less HOCl, when compared to neutrophils. May be this view needs to be revised, or another reactive species, produced by macrophages, can activate the fro operon as well.

      (2) HOCl is typically unstable in the presence of biomolecules. Nevertheless, medium conditioned by activated neutrophils is a strong inducer of the fro operon. The medium used by the authors for this experiment contains taurine, and, as the authors acknowledge, this taurine will likely react with HOCl to form the more stable taurine N-chloramine. Similarly, the MinA bacterial medium used to treat P. aeruginosa with HOCl directly also contains ammonium ions at mM concentrations, which could potentially react with HOCl to form monochloramine. It could be speculated that taurine N-chloramine and other chloramines are as effective as HOCl in activating the fro-operon.

      (3) The fro operon was originally described to be activated by shear stress ("flow-regulated operon"). How shear stress and HOCl-stress are related, or if fro activation by both stimuli is a coincidence, remains unclear. The authors propose a model, in which flow transports oxidizing molecules, which ultimately activate the fro operon. However, the initial work by Sanfilippo et al. (2019, Nat Microbiol) used plain LB medium in a fluidic chamber to induce the shear stress, which should be free of oxidants, and certainly of HOCl.

      Comments on revised version:

      The authors have addressed my concerns appropriately.

    4. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the efforts of the reviewers, which have provided insightful and helpful comments to improve the manuscript. The feedback touches upon a number of topics, focusing on clarification or justification of experimental techniques and on understanding the mechanism by which P. aeruginosa detects HOCl. All reviewers raised the issue of how HOCl activates fro expression, including whether free or protein-bound methionine, cysteine, or other HOCl byproducts induce this expression. For the upcoming revision, we plan to perform experiments that address this issue and will discuss potential mechanistic models in light of the new data. In addition, we plan to perform additional experiments to address a reviewer’s concerns regarding the dependence of the fro response on HOCl production by neutrophils. The revision will correct imprecise statements pointed out by reviewers, and address all remaining issues requiring clarification or further discussion, including the range of HOCl sensitivity, relationship between HOCl and flow sensitivity, and justification for testing the fro response to nitric acid.

      We have completed a number of experiments and responded thoroughly to reviewer comments below. We thank the reviewers again for their details comments, which suggested additional experiments and interpretations that have resulted in significant additional insight into the potential mechanism of HOCl sensing and its relevance with neutrophils.

      Reviewer #1 (Public review):

      Summary:

      Foik et al. report that hypochlorous acid, a reactive chlorine species generated during host defense, activates the transcription of the froABCD in P. aeruginosa. This gene cluster had previously been associated with a potential role during the flow of fluids and appears to be regulated by the sigma factor FroR and its antisigma factor FroI. In the present study, the authors show that froABCD is expressed both in neutrophils and macrophages, which they claim is likely a result of HOCl but not H2O2 production. Fro expression is also induced in a murine model of corneal infection, which is characterized by immune cell invasion. Expression of the fro system can be quenched by several antioxidants, such as methionine, cysteine, and others. FroR-deficient cells that lack froABCD expression during HOCl stress appear more sensitive to the oxidant.

      Strengths:

      The authors provide a number of data supporting their claim that transcription of the froABCD system is induced by reactive chlorine species. This was shown by RNAseq, qRT-PCR, and through microscopy using a transcriptional reporter fusion. Likewise, elevated expression of froABCD was shown in vitro and in vivo, excluding potential in vitro artifacts. The manuscript, while mostly descriptive, is easy to follow, and the data were presented clearly.

      We greatly appreciate the efforts of the reviewer and thank them for their succinct summary of the manuscript.

      Weaknesses:

      (1) Lines 60-62: Some of the authors' conclusions are not supported by the data and thus appear unfounded. One example: "we determine that fro upregulation.....These data suggest a novel mechanism..." Their data do not show that MSR upregulation is a direct effect of FroABCD. Instead, it could be possible that the FroR sigma factor also controls the expression of msr genes, which would be independent of froABCD.

      We thank the reviewer for pointing out this important distinction. We have clarified in lines 63-65 in the clean version of the revision that MSR upregulation depends on FroR rather than FroABCD.

      (2) The authors show increased fro transcription both in neutrophils and macrophages; however, the two types of immune cells differ quite dramatically with respect to myeloperoxidase activation and HOCl production.

      Neither has this been discussed nor considered here.

      We agree that the distinction between the cell types is important and have added a brief description of the differences in respiratory bursts and ROS production between the two cell types and our justification for focusing on neutrophils in lines 102-103. We think it’s very interesting that Fro appears to be activated by macrophages, which are not associated with HOCl production on their own. We think that it would be interesting to identify what is inducing Fro in macrophages in future work.

      (3) With respect to the activation of fro expression upon challenge with conditioned media from stimulated neutrophils, does the conditioned media contain detectable amounts of HOCl? Do chloramines, which are byproducts of HOCl oxidation with amines, also stimulate expression?

      This is an excellent question that addresses which molecules Fro is responding to from neutrophils. We have performed additional experiments that confirm that PMA-stimulated neutrophils produce HOCl (Fig. S2) through the use of a commercial hypochlorite sensor assay, which claims high specificity for detecting HOCl. We further confirmed that this production is inhibited by pretreatment with the MPO-specific inhibitor 4ABAH. These data support the interpretation that the Fro response to stimulated neutrophils requires MPO activity, of which the major product is HOCl. We have described this in lines 117-127.

      Our data does not exclude the possibility that other MPO products could activate Fro expression. We were unable to obtain a reliable source of the major secondary MPO product, taurine chloramine, for our experiments, unfortunately. We believe that understanding the potential for secondary products to activate the response is an important and interesting question that can be explored in a future study. We have added a discussion of this in lines 367-376.

      (4) A better control to prove that this fro expression is indeed induced by HOCl in activated neutrophils would be to conduct the experiments in the presence of a myeloperoxidase inhibitor.

      We thank the reviewer for raising this point. We have performed the suggested set of experiments and found that indeed, the pre-treatment of neutrophils with MPO inhibitor 4-ABAH prior to PMA stimulation suppresses the activation of fro (Fig. 2D and Figure 2-figure supplement 1-2). The results are discussed in lines 117-127.

      (5) The work was conducted with two different P. aeruginosa strains (i.e. AL143 and PAO1F). None of the figure legends provides details on which strain was used. For instance, in line 111, the authors refer to Figure S1B for data that I thought were done with PAO1F, while in 154, data were presented in the context of the infection model, which was conducted with the other strain.

      We thank the reviewer for pointing this issue out. We have ensured that strain names appear in all the revised figure legends. To clarify, only mouse experiments and a related RT-qPCR assay used strain PAO1F due to prior IACUC approval of this strain and its use in previous publications.

      (6) It would be good if immune cell recruitment at 2hrs and 20hrs PI could be quantified.

      We previously quantified neutrophil recruitment at the site of corneal abrasion at 24 hours using the same conditions and strains (Ratitong, B. et al., J. Immun, 2022). While we do not have immune cell recruitment data for the 2 hr and 20 hr time points, the previous data show significant neutrophil recruitment near the latter time point, which is consistent with the interpretation that fro expression is activated by stimulated neutrophils. We have discussed this in lines 212-216.

      (7) The conclusions of Figure 4 are, in my opinion, weak (line 187-188; "It is possible that ....."). These antioxidants likely quench the low amounts of NaOCl directly. This would significantly reduce the NaOCl concentrations to a level that no longer activates expression of fro. There is no direct evidence provided that oxidized methionine induces fro expression. Do the authors postulate that this is free methionine, or could methionine and/or cysteine oxidation in FroR increase the binding affinity of the sigma factor to the promoter? Another possibility is that NaOCl deactivates the anti-sigma factor. None of these scenarios has been considered here.

      We acknowledge that our model of HOCl sensing was unclear and thank the reviewer for their insight. This critique is echoed by reviewer #2 in comment 3 as well. We recently found that the FroI anti-sigma factor has the highest concentration of methionine and cysteine residues of all known P. aeruginosa anti-sigma factors (Appendix 2—Table 1). Given that FroR and FroI form an extracytoplasmic function sigma – anti-sigma pair, which are associated with transducing extracellular signals to the cytoplasm, we have proposed an alternative model in which HOCl or secondary RCS molecules are detected through their oxidation of cysteine and methionine residues in FroI. This is discussed in lines 271-279 and lines 363-367.

      (8) Line 184: The reaction constants of HOCl with Cys and Met are similar.

      We thank the reviewer for pointing out this important clarification. We have revised the sentence to accurately reflect this in lines 243-244.

      (9) Treatment with 16 uM NaOCl caused a growth arrest of ~15 hrs in the WT (Figure 5A), whereas no growth at all was recorded with 7.5 uM in Figure 3A.

      We thank the reviewer for catching this. We have determined that the concentration of NaOCl in the reagent used for this particular experiment was lower than expected, thus requiring a higher concentration to achieve growth inhibition. We have repeated the experiment with new reagent and find that the results (now Figure 6A) are similar to the previous experiment but at a lower concentration of 4 micromolar, consistent with the concentration found to be sub-inhibitory in Figure 3A.

      (10) The concentration range of NaOCl causing fro expression is extremely narrow, while oxidative burst rapidly generates HOCl at much higher concentrations. This should be discussed in more detail.

      We appreciate the reviewer’s comment, which is related to reviewer #2’s comment #9. We have clarified the reported production rates of HOCl, which far surpass the bacterial MIC. After greater consideration, we believe secondary HOCl products including taurine chloramine could have a more significant role in vivo. While this molecule is less potent than HOCl, it is longer-lived, retains bactericidal activity, and retains the ability to oxidize methionine. We have discussed this in lines 377-405.

      Reviewer #1 (Recommendations for the authors):

      (1) Some statements in the text don't match the data shown in the Figures. For instance:

      (a) Figure 2B shows ~65-fold fro expression, but the text states: "...increased expression of fro expression by 30fold..."(line 88).

      The YFP/mCherry value of PMA-stimulated is 67.4 in the Figure (now Figure 2C). The fold-change is computed relative to unstimulated conditioned medium (third column, which has a value of 2.2), which is a 30-fold change. We have added a citation in the main text to the Source Data, which provides these raw values, to help clarify the computation for readers, and added in the legend that the value for unstimulated is greater than 1.

      (b) Figure 3B shows ~35-fold fro expression at 1 uM NaOCl, but the text states: "...NaOCl increased fro expression by up to 74-fold..."(line 110).

      We have clarified that the increase is relative to untreated, added that untreated value is below 1 in the caption, and provided a citation in the main text to the Source Data, which contains the raw values. The change is measured relative to untreated, for which the YFP/mCherry value is 0.48. The value at 1 uM is 35.7, giving a 74-fold change.

      (2) Line 229: While the ∆froR strain was sensitive to HOCl, the strain was tolerant". Please revise.

      We have corrected this typo (now lines 328- 329). We meant to convey that growth was not entirely inhibited in the froR strain.

      Reviewer #2 (Public review):

      Summary:

      Foik et al. studied the regulation of the fro operon in response to HOCl, an oxidant derived from immune cells, especially neutrophils. They use a transcriptional fusion of YFP to the froA promoter in an mCherry-expressing P. aeruginosa strain to determine fro-induction under the microscope. They use this system to study fro expression in medium, in the presence of neutrophils and macrophages, neutrophil-conditioned medium, and several chemical stimuli, including NaCl, HOCl, hydrogen peroxide, nitric acid, hydrochloric acid, and sodium hydroxide. They also use a corneal infection model to demonstrate that froA is upregulated in P. aeruginosa 20 h post-infection and perform transcriptional analyses in WT and a froR mutant in response to HOCl.

      Strengths:

      Their data clearly shows that HOCl is a strong inducer of the fro Operon. The addition of HOClquenching chemicals together with HOCl abrogates the response. They also show that a froR mutant is more susceptible to HOCl than WT. Their transcriptomic data reveal genes under control of the FroR/FroI sigma factor/anti sigma factor system.

      Weaknesses:

      Although the presented evidence is mostly solid, some of their findings need to be evaluated more carefully; explaining the rationale behind some of the experiments might enhance the article, and some of the models proposed by the authors seem far-fetched, as outlined below:

      We greatly appreciate the reviewer’s efforts and thank them for highlighting strengths and areas for improvement.

      (1) In line 76 the authors claim "Relative to P. aeruginosa that were incubated in host cell-free media, P. aeruginosa in close proximity to human neutrophils or that were engulfed in mouse macrophages appeared to increase fro expression (Fig. 1C)". Counting bacterial cells in Figure 1C shows that 1 in 17 bacteria (5.8%) induce the froA-promotor in media in the absence of immune cells, while 4 in 72 bacteria (only 5.5%) do the same in the presence of neutrophils. Contrary to the authors' claims, it appears that P. aeruginosa actually decreases fro-expression in close proximity to neutrophils. There is a slight increase in fro-expression in bacteria co-incubated with macrophages (3 in 21, or 14.3%). A more rigorous statistical analysis might substantiate the authors' claim, but, as is, the claim "neutrophils increase fro expression" is untenable.

      We believe the images alone do not give an adequate representation of the data and have quantified a larger portion of the data, which has been added as Figure 1D. The quantification supports the original claim that fro expression is increased during co-incubation with macrophages and neutrophils. Since there was not sufficient statistical sampling to distinguish engulfed P. aeruginosa from free ones, this part of the claim has been removed from the text (updated in lines 80-84).

      (2) The authors should explain the rationale behind some of the chemicals used. Why did they use nitric acid? Especially at these high concentrations, a strong acid such as nitric acid might have a significant influence on the medium pH. I understand that the medium is phosphate-buffered, but 25 mM nitric acid in an unbuffered medium would shift the pH well below 2. Similar considerations apply to hydrochloric acid and sodium hydroxide.

      We thank the reviewers for pointing out the need for this clarification. We have updated Figure 3D with a lower concentration of NaOH at 1 uM, which is the same concentration as NaOCl that activates fro expression. Due to the high buffering capacity of our medium, a high concentration of 6 mM NaOH was needed to induce a discernible change in pH and this high concentration of NaOH had no obvious effect on growth. Neither 1 uM nor 6 mM NaOH produced a change in fro expression, consistent with our previous findings that the effect is not due to sodium ions or higher pH. These updated findings are described in lines 181-190.

      Since the effect of chloride is already controlled for using NaCl and the concentration of HCl used was not sufficient to cause a significant change in pH in the buffered medium, we have removed the HCl group from the data.

      We have clarified that nitric acid was used because it is a strong oxidizer that is not found in neutrophils and that concentrations used were near the minimal inhibitory concentrations (lines 145-147 and lines 191-198). We acknowledge that the growth inhibition from HNO<sub>3</sub> could be due to pH or oxidation. However, since no change in fro expression was observed at concentrations approaching the inhibitory concentration, we did not address the potential effects of low pH from nitric acid on fro expression.

      (3) In line 187, the authors state that "It is possible that oxidized methionine increases fro expression" and they suggest a model to that effect in Figure 5D. It is unclear why the authors singled out methionine sulfoxide, since a number of other things get oxidized by HOCl. In line 184, the authors state, in the same vein, that "HOCl oxidizes methionine residues 100-fold more rapidly than other cellular components". The authors should state which other cellular compounds they are referring to. Certainly not cysteine and other thiols, which react equally fast and are highly abundant in the cell: P. aeruginosa contains 340 µM GSH, 140 µM CoA-SH (https://doi.org/10.1074/jbc.RA119.009934) plus free cysteine and cysteines in proteins (based on codon usage, 1.34% of amino acids in proteins are cysteine, while methionine is only slightly more present at 2.10%, although a number of starting methionines are removed from mature proteins).

      We acknowledge that our HOCl sensing model had been vague and unclear and thank the reviewer for their insight. This critique is echoed by reviewer #1 in comment 7 as well.

      Our initial suggestion that methionine sulfoxide was sensed was motivated by the observation that methionine sulfoxide reductases are upregulated by HOCl. However, we have revised this based on feedback from reviewers and further consideration of chlorine redox chemistry. Interestingly, we found that the FroI anti-sigma factor has the highest concentration of methionine and cysteine residues of all the P. aeruginosa anti-sigma factors (Appendix 2—table 1). Given that FroR and FroI form an extracytoplasmic function sigma – anti-sigma pair, which are associated with transducing extracellular signals to the cytoplasm, we propose a model in which HOCl or secondary RCS molecules are detected by their oxidation of cysteine and methionine residues in FroI. This is discussed in lines 271-278 and lines 362-366.

      (4) Overall (and this is probably not addressable with the authors' data), some very interesting questions remain unanswered: what is the molecular mechanism of fro-induction? How is the FroR/FroI system modulated by HOCl? Does the system sense free or protein-bound methionine-sulfoxide? Are certain methionine residues in these proteins directly oxidized by HOCl? Many "HOCl-sensing" proteins are also modified at cysteine residues or amino groups; could those play a role? And lastly: what is the connection between shear/fluid flow and HOCl, or are these totally separate mechanisms of fro-induction?

      We thank the reviewer for raising these excellent mechanistic questions. Issues relating to HOCl sensing are addressed in the preceding comment.

      Regarding the connection to shear sensing, Padron et al., 2023 found that the detection of flow in P. aeruginosa can be attributed to chemical transport, in particular to H<sub>2</sub>O<sub>2</sub> that was present in growth media. Based on the same principle, we expect Fro to be upregulated in flow at much lower concentrations than those observed in stationary fluids. The activation of Fro and the effects of HOCl would thus be expected to be flow-sensitive. We have commented on this important factor in the discussion in lines 406-417.

      Reviewer #2 (Recommendations for the authors):

      (1) To address 1, the authors could evaluate the microscopic images in the same manner in which they evaluated the other microscopic images, as, for example, presented in Figure 2B or Figure S1B.

      See response to Weakness point (1).

      (2) To address 2, please explain the choice of the chemicals (why nitric acid?), but also provide the pH of the media with those high concentrations of strong acids and bases, and interpret them in light of the permissible pH range for P. aeruginosa growth. More sensible controls might be a lower NaOH concentration in the range that would be reached through the amount of NaOH in the NaOCl stock at the highest NaOCl concentrations used. As for the acids, I don't see a reason to use these acids at these high concentrations. Please explain.

      See response to Weakness point (2).

      (3) To address 3: The authors could specify their methionine-sulfoxide model a bit more, so that testable hypotheses can be developed. If the authors think the FroR/FroI system senses free methionine sulfoxide, they or others could add methionine sulfoxide to the medium and check induction. If they think specific methionine residues in these proteins are oxidized, they could provide evolutionary evidence of conserved methionine residues. Or, based on a structure or structural prediction, they (or others) could mutate methionine residues, e.g., at the protein's surface or potential protein/protein-interaction sites and assess the effect on HOCl-based activation. Or they could consider other amino acids known to be highly reactive towards HOCl and mutate those in a future study.

      See response to Weakness point (3).

      Further comments:

      (4) The headline of the figure legend of Figure S1 seems incomplete. Please mention the flow experiments shown in Figure S1A.

      We have updated the title of this figure, which now appears in the eLife format as Figure 1 – figure supplement 1.

      (5) What is the difference between the data presented in Figure 4C (bars "UTR" and "NaOCl") and the same bars in Figure S2A? Is this redundant or a re-plot?

      In this revision, Figure 4C has become Figure 5B and Figure S2A is now Figure 5—figure supplement 1. Only the 1 uM NaOCl condition is replotted. We have described this in the legend for Figure 5—figure supplement 1.

      (6) Line 228: hpd is more likely a gene of the aromatic amino acid catabolism.

      We thank the reviewer for pointing this out. We have removed the ‘branched’ descriptor in this sentence, now in lines 325-326.

      (7) Line 229: "While the ΔfroR strain was sensitive to HOCl, the strain was tolerant (Figure 5A)". Please clarify. Which strain was tolerant? WT?

      We have corrected this typo (now line 328-329). We meant to convey that growth was not entirely inhibited in the froR strain.

      (8) Line 247: "Activated neutrophils produce HOCl concentrations as high as 50 µM [24]." This "50 µM" number is often quoted; the citation trail typically leads to Weiss et al. 1982 (https://doi.org/10.1172/JCI110652). However, a more factually correct statement based on that paper would be "2 x 10^6 neutrophils, activated with 30 ng/mL PMA at 37C in 1 mL of Dulbecco's buffer can produce around 50 nmol HOCl per hour". In the particular reference 24, Dybpukt et al used a methodology similar to Weiss et al., and here around 50nmol were produced by the same number of cells in 30 min in response to 100 ng/mL PMA. Please clarify accordingly.

      We thank the reviewer for bringing this to our attention. We have altered the language to indicate that the production rate is for this specific set of parameters. Related to this, Reviewer #1 Comment #10 requested a more detailed discussion of the significance of the Fro response, since it is at much lower HOCl concentration than produced by neutrophils. We have discussed this in lines 377-405.

      (9) Line 275: "which is strain PA14 strain". Please clarify.

      We have fixed this error and entered it into the Key Resource Table.

      (10) Line 311: "Cultures containing densities below 10 P. aeruginosa per frame were concentrated using a syringe filter with 0.2 or 0.8 µm pore sizes (Millipore, Burlington, MA)." Isn't that a bit concerning when testing the induction of an operon that is supposedly activated by shear through fluid flow? Did the authors convince themselves that this procedure does not induce fro?

      We do not expect the filtering procedure to cause changes in gene expression because cells are imaged immediately after filtering. Nonetheless, we performed additional experiments (Fig 3F and 2D) entirely without concentrating cells and found that YFP/mCherry levels were consistent with previous data in Figure 3. This rationale has been added to the Methods section under the “Fluorescence and Phase Contrast Microscopy” section.

    1. eLife Assessment

      This study provides important insights into the neural mechanisms linking sleep and long-term memory consolidation. By combining behavioural, genetic, imaging, and connectomic approaches in Drosophila, it identifies a target neural circuit that will be of broad interest to researchers studying sleep, memory, and neural circuits. The evidence supporting the involvement of the identified circuit in the regulation of sleep and memory is solid and represents a substantial advance in the field. Nevertheless, there is limited evidence to support the mechanistic claim that this circuit directly links sleep and memory consolidation within the available data, and some results should therefore be interpreted with appropriate caution.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have considered and discussed the comments raised in the previous round of review.]

      Summary:

      The authors aim to use state-of-the art behaviour, imaging and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      The investigation follows a logical strategy to collect huge dataset of sleep, appetitive memory and live imaging. The authors identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation dependant manner. The author also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the new data provide TRIC-LUC provided better temporal resolution of neural activity correlates for PAMalpha1-DPM inhibition. Importantly, the author demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP at the memory consolidation period after training.

      Weaknesses:

      Although the revised version carries arguments to satisfy the reviewers' concern, the writing is now less cohesive. Crucially an explanation however remains required for the following experimental contradiction: the central observation of the study indicates that PAM alpha1 activation cause DPM inhibition which disrupt sleep and memory consolidation. Therefore, one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found the opposite is true from now enhanced TRIC-LUC dataset. The authors indicate this data reinforce the inhibitory nature of PAM-alph1-DPM, but it does not explain why such a reduced DPM activity is observed after training.

    3. Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain incompletely understood. The authors examined a specific subset of PAM dopaminergic neurons, PAM-α1, and DPM neurons in Drosophila. These neurons have previously been implicated in memory, and DPM neurons have also been linked to sleep. The study explores whether this circuit provides a mechanistic link between sleep and memory consolidation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-α1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly during the night. The authors further show that perturbation of PAM-α1 and DPM neurons impairs sleep and appetitive memory consolidation under starvation conditions, and that pharmacological sleep induction during the night rescues the LTM defects. Together, these findings suggest that PAM-α1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important observations that advance our understanding of the circuits regulating sleep and memory consolidation.

      Weaknesses:

      Some claims require additional evidence or clarification.

      (1) Previous studies linking impaired memory to reduced sleep have primarily examined conditions involving severe sleep deprivation. In contrast, this manuscript argues that relatively modest decreases in total sleep, accompanied by sleep fragmentation, are sufficient to impair memory consolidation. It remains unclear whether sleep fragmentation of this magnitude is itself critical for LTM consolidation. An independent method for inducing comparably mild sleep loss and fragmentation would be needed to directly test this interpretation.

      (2) It is unclear why both activation and inactivation of PAM-α1 neurons produce similar effects on sleep and memory. In addition, MB299B-labeled neurons exert stronger effects on memory than MB043B-labeled neurons, whereas MB043B-labeled neurons have stronger effects on sleep. If sleep disruption is the primary driver of impaired memory consolidation, a stronger correspondence between the sleep and memory phenotypes might be expected. The authors speculate that MB043B may affect sleep through non-PAM neurons, but without identifying the relevant neurons, this remains speculative.

      (3) The complex schematic model (Fig. 12), with parallel circuits and unidentified neuronal groups, underscores the difficulty of interpreting the current data. In the "less activity" arm of the model, distinct circuits are proposed to regulate sleep and LTM, respectively, and DPM neurons are not included. This makes it difficult to reconcile the model with the central claim that the PAM-α1-to-DPM microcircuit links sleep and LTM consolidation.

      (4) The TRIC-LUC reporter system is not ideal for resolving dynamic changes in neuronal activity. Activity-dependent Ca²⁺ signaling must first reconstitute the TRIC transcriptional system, which then drives luciferase transcription, translation, and accumulation. The original characterization of TRIC indicates that TRIC signals accumulate and decay over several hours. Thus, the kinetics of the TRIC-LUC reporter should be interpreted cautiously, particularly when inferring transient or precisely timed changes in neuronal activity.

      (5) Including data from training under fed conditions would provide a more complete understanding of state-dependent neural activity and would help distinguish starvation-specific effects from more general circuit mechanisms.

    4. Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence how dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strength:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuit interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of the role of dopamine in cognitive functions. Additional experiments investigating the contribution of MBON-α1 to the circuit, connectomic analysis together with a dissection of dopamine receptor function further strengthen the proposed circuit motif and its biological relevance.

      Weaknesses:

      While the study is well designed, presents compelling findings and has been further strengthened by additional experiments, some aspects remain unclear. The role of DPM neurons in memory consolidation seems not yet fully resolved, as different genetic approaches yield variable results. Furthermore, some manipulations impair memory without affecting sleep fragmentation - or vice versa, suggesting that the observed memory deficits cannot be explained solely by impaired sleep-dependent consolidation. It would also have been interesting to discuss potential mechanisms by which dopamine receptor-mediated cAMP signaling could lead to a reduction in Ca²⁺ signals. I am confident that these questions can be addressed in future studies.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopaminergic circuits interact to regulate memory consolidation in Drosophila and may reveal general principles underlying the neural regulation of memory.

    5. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This study provides important insights into the neural mechanisms linking sleep and long-term memory consolidation. By combining behavioural, genetic, imaging, and connectomic approaches in Drosophila, it identifies a target neural circuit that will be of broad interest to researchers studying sleep, memory, and neural circuits. The evidence supporting the involvement of the identified circuit in the regulation of sleep and memory is solid and represents a substantial advance in the field. Nevertheless, there is limited evidence to support the mechanistic claim that this circuit directly links sleep and memory consolidation within the available data, and some results should therefore be interpreted with appropriate caution.

      We appreciate the reviewer’s careful evaluation of our manuscript, and we agree that (as with any experimental study) there is still a lot to do to fully understand the mechanisms underlying the linkage between sleep and memory consolidation.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behaviour, imaging and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      The investigation follows a logical strategy to collect huge dataset of sleep, appetitive memory and live imaging. The authors identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation dependent manner. The author also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the new data provide TRIC-LUC provided better temporal resolution of neural activity correlates for PAMalpha1-DPM inhibition. Importantly, the author demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP at the memory consolidation period after training.

      Weaknesses:

      Although the revised version carries arguments to satisfy the reviewers' concern, the writing is now less cohesive. Crucially an explanation however remains required for the following experimental contradiction: the central observation of the study indicates that PAM alpha1 activation cause DPM inhibition which disrupt sleep and memory consolidation. Therefore, one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found the opposite is true from now enhanced TRIC-LUC dataset. The authors indicate this data reinforce the inhibitory nature of PAM-alph1-DPM, but it does not explain why such a reduced DPM activity is observed after training.

      We thank the reviewer for their point of view. We have faithfully reported all experimental observations acquired from our enhanced TRIC-LUC dataset as objectively as possible. We note that the reviewer postulates a particular expected activity shift (reduced PAM-α1 activity and elevated DPM activity after memory training), yet it is not clear why. As our data show, this is a highly connected microcircuit and it is not easy to predict how it might change. Our data show that there is change and dismissing our empirically measured results purely based on a theoretical expectation is not justified. The brain operates as an intricately interconnected network; neural activity dynamics cannot always be simply inferred from static circuit polarity. Progress in deciphering neural circuit function relies on iterative rounds of experimental testing. Like most neuroscience investigations, the present study cannot resolve every open question, and we explicitly acknowledge several unresolved directions worthy of future exploration in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain incompletely understood. The authors examined a specific subset of PAM dopaminergic neurons, PAM-α1, and DPM neurons in Drosophila. These neurons have previously been implicated in memory, and DPM neurons have also been linked to sleep. The study explores whether this circuit provides a mechanistic link between sleep and memory consolidation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-α1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly during the night. The authors further show that perturbation of PAM-α1 and DPM neurons impairs sleep and appetitive memory consolidation under starvation conditions, and that pharmacological sleep induction during the night rescues the LTM defects. Together, these findings suggest that PAM-α1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important observations that advance our understanding of the circuits regulating sleep and memory consolidation.

      Weaknesses:

      Some claims require additional evidence or clarification.

      (1) Previous studies linking impaired memory to reduced sleep have primarily examined conditions involving severe sleep deprivation. In contrast, this manuscript argues that relatively modest decreases in total sleep, accompanied by sleep fragmentation, are sufficient to impair memory consolidation. It remains unclear whether sleep fragmentation of this magnitude is itself critical for LTM consolidation. An independent method for inducing comparably mild sleep loss and fragmentation would be needed to directly test this interpretation.

      We appreciate the reviewer’s suggestion. While alternative assays for inducing sleep loss or sleep fragmentation are indeed available, this line of investigation lies beyond the core scope of the present study. We will certainly take this valuable suggestion into consideration for the future studies.

      Regarding the question of whether sleep fragmentation of this magnitude per se is critical for long-term memory consolidation, we would like to highlight relevant published evidence. Prior work in rodents and human (Bonnet and Arand, 2003; Van Someren et al., 2015; Ramesh et al., 2012; Baud et al., 2014) and our earlier study (Liu et al., 2019) have demonstrated that alterations in sleep architecture, independent of changes in total sleep amount, can affect multiple physiological processes, including memory. Given this existing supporting evidence, we respectfully argue that this does not constitute a weakness of the present manuscript.

      (2) It is unclear why both activation and inactivation of PAM-α1 neurons produce similar effects on sleep and memory. In addition, MB299B-labeled neurons exert stronger effects on memory than MB043B-labeled neurons, whereas MB043B-labeled neurons have stronger effects on sleep. If sleep disruption is the primary driver of impaired memory consolidation, a stronger correspondence between the sleep and memory phenotypes might be expected. The authors speculate that MB043B may affect sleep through non-PAM neurons, but without identifying the relevant neurons, this remains speculative.

      The concern raised by the reviewer that certain interpretations remain speculative represents a common situation in most published research. This interesting direction warrants further investigation in future work, but falls outside the scope of the present study. In the revised manuscript, we have elaborated on the differences observed between these two GAL4 drivers. We have also conducted additional experiments to investigate a well-characterized memory-related recurrent loop of PAM-α1 neurons in sleep regulation. We respectfully note that no single study can comprehensively address all outstanding questions.

      (3) The complex schematic model (Fig. 12), with parallel circuits and unidentified neuronal groups, underscores the difficulty of interpreting the current data. In the "less activity" arm of the model, distinct circuits are proposed to regulate sleep and LTM, respectively, and DPM neurons are not included. This makes it difficult to reconcile the model with the central claim that the PAM-α1-to-DPM microcircuit links sleep and LTM consolidation.

      We appreciate this careful comment on our schematic model in Figure 12. This diagram aims to summarize the key findings obtained in the present study while also explicitly laying out unresolved questions that await future investigation. In our view, including open, outstanding questions in the working model does not undermine the interpretation of our existing experimental results. In the revised manuscript, we have modified the corresponding text to clarify this point and distinguish firmly between conclusions supported by our data and tentative components requiring follow-up validation.

      (4) The TRIC-LUC reporter system is not ideal for resolving dynamic changes in neuronal activity. Activity-dependent Ca<sup>2+</sup> signaling must first reconstitute the TRIC transcriptional system, which then drives luciferase transcription, translation, and accumulation. The original characterization of TRIC indicates that TRIC signals accumulate and decay over several hours. Thus, the kinetics of the TRIC-LUC reporter should be interpreted cautiously, particularly when inferring transient or precisely timed changes in neuronal activity.

      We fully acknowledge the inherent limitations of the TRIC-LUC reporter system, as pointed out by the reviewer. Every experimental tool comes with characteristic strengths and drawbacks. Although TRIC-LUC suffers from temporal delays, our experiment does not aim to capture acute, immediate effects; instead, it examines long-term dynamics of neuronal activity. To date, within Drosophila neurobiology, no superior technique is available for non-invasive long-term monitoring of neuronal activity in freely behaving flies. We share the hope that new tools capable of reporting neuronal activity in real time will be developed and applied, which will facilitate deeper mechanistic understanding of neuronal dynamics.

      (5) Including data from training under fed conditions would provide a more complete understanding of state-dependent neural activity and would help distinguish starvation-specific effects from more general circuit mechanisms.

      We appreciate this suggestion. First, our memory paradigm relies on reward-based associative learning, and starvation is required for flies to express robust memory, so to do this would require a completely new experimental set up. Second, our core findings demonstrate that transient perturbations of this neuronal circuit trigger sleep disturbances and memory deficits specifically under starvation conditions. Therefore, measurements of neural activity under fed conditions are not directly relevant to the central conclusions of the present study. We agree that related experiments on other behaviors under fed states constitute an interesting direction and could be pursued in future investigations.

      Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence how dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strength:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuit interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of the role of dopamine in cognitive functions. Additional experiments investigating the contribution of MBON-α1 to the circuit, connectomic analysis together with a dissection of dopamine receptor function further strengthen the proposed circuit motif and its biological relevance.

      Weaknesses:

      While the study is well designed, presents compelling findings and has been further strengthened by additional experiments, some aspects remain unclear. The role of DPM neurons in memory consolidation seems not yet fully resolved, as different genetic approaches yield variable results. Furthermore, some manipulations impair memory without affecting sleep fragmentation - or vice versa, suggesting that the observed memory deficits cannot be explained solely by impaired sleep-dependent consolidation. It would also have been interesting to discuss potential mechanisms by which dopamine receptor-mediated cAMP signaling could lead to a reduction in Ca<sup>2+</sup> signals. I am confident that these questions can be addressed in future studies.

      We greatly appreciate the reviewer’s positive evaluation of our work and the thoughtful suggestions regarding future directions. As acknowledged in the manuscript, our study centers on identifying a shared circuit that coregulates sleep and memory processes. We have performed preliminary investigations of downstream circuitry, and our results indeed support the idea that sleep and memory can be modulated independently. Importantly, we have avoided drawing definitive conclusions that memory deficits arise purely from impaired sleep-dependent consolidation. Instead, we emphasize the existence of a common circuit mechanism governing both processes.

      We also thank the reviewer for drawing attention to dopamine receptor-mediated cAMP signaling and Ca<sup>2+</sup>dynamics. This observation constitutes an additional finding that requires more extensive mechanistic follow-up. Given the scope of the current work, we have not dedicated a separate discussion section to dissecting this pathway. This promising line of inquiry will be pursued in our future research.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopaminergic circuits interact to regulate memory consolidation in Drosophila and may reveal general principles underlying the neural regulation of memory.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further issue, apart from the reverse figure 13 are not found in the main text as indicated.

      We thank the reviewer for this careful check. We have performed a full-text search for Figure 13 throughout the revised manuscript and found no relevant citation. We have also carefully cross-checked all figure numbering and confirm that all figure labels are accurate in the current version.

      Reviewer #2 (Recommendations for the authors):

      In Fig. 8B, some individual GCaMP measurements show values below −100% ΔF/F₀. Under the stated definition, ΔF/F = (Fn - F0) / F0, values below −100% would require Fn to be negative. Since raw fluorescence intensity cannot be negative, values below −100% require further explanation. The authors should clarify whether Fn represents raw fluorescence or processed fluorescence, and whether the plotted traces underwent any normalization, subtraction, detrending, or transformation beyond the stated formula.

      We greatly appreciate the reviewer’s rigorous scrutiny of our data and the valuable question raised. Our fluorescence signals were calculated using the standard formula ΔF/F = (Fn − F0)/F0, where Fn = F_ROI − F_background, and all calculations were implemented accordingly. We have carefully revisited all raw imaging datasets and identified the source of the issue. During initial data processing, we retained all acquired recordings without excluding samples exhibiting focal plane drift. This drift occasionally yielded negative values for Fn (F_ROI − F_background). Beyond the five DPM cell bodies from four brains in Figure 8B highlighted by the reviewer, we further detected six additional DPM cell bodies from four brains in Figure 2A affected by the same artifact. We have now excluded these drifting preparations, regenerated all corresponding plots, and updated the statistical analyses in the revised manuscript. For transparency, we upload both the raw and processed datasets as supplementary materials to clarify this point.

    1. eLife Assessment

      This important study introduces NoSeMaze, a semi-naturalistic platform for continuous, high-dimensional tracking of social and cognitive behaviors in group-housed mice, and uses it to show that individual social rank is stable across changing social contexts. By integrating automated dominance measures, proactive social behaviors, and reinforcement-learning-based profiles, the authors demonstrate a novel framework for examining how stable individual differences shape social structure. The findings provide compelling evidence that dominance traits are stable across changing social contexts and largely independent of non-social cognitive performance, supporting the view that social rank reflects an intrinsic dimension of individuality however, the broader functional significance of dominance in this paradigm remains somewhat ambiguous, including the extent to which the measured behaviors capture how dominance operates in more naturalistic social settings. This work will be of broad relevance for behavioral neuroscience and social behavior research.

    2. Reviewer #2 (Public review):

      Summary:

      This manuscript presents the "NoSeMaze", a novel automated platform for studying social behavior and cognitive performance in group-housed male mice. The authors report that mice form robust, transitive dominance hierarchies in this environment and that individual social rank remains largely stable across multiple group compositions. They further demonstrate that social dominance and aggressive behaviors, like chasing, are partially dissociable and that dominance traits are independent of non-social cognitive performance. The study includes a genetic manipulation of oxytocin receptor expression in the anterior olfactory nucleus, which showed only transient effects on social rank.

      Strengths:

      (1) Innovative Methodology:<br /> The NoSeMaze platform is a technically elegant and conceptually well-integrated system that enables fully automated, long-term monitoring of both social and cognitive behaviors in large groups of group-housed mice. It combines tube-test-like dominance contests, voluntary chase-escape interactions, and an embedded operant olfactory discrimination task within a single, ethologically relevant environment. This modular design allows for high-throughput, minimally invasive behavioral assessment without the need for repeated handling or artificial isolation.

      (2) Experimental Scale and Rigor:<br /> The study includes 79 male mice and over 4,000 mouse-days of observation across multiple group reshufflings. The use of RFID-based identification, automated data logging, and longitudinal design enables robust quantification of individual trait stability and group-level social structure.

      (3) Multidimensional Behavioral Profiling:<br /> The integration of social (tube dominance, proactive chasing), physical (body weight), and cognitive (olfactory learning task) measures offers a rich, multi-dimensional profile of each individual mouse. The authors' finding that social dominance traits and non-social cognitive performance are largely uncorrelated reinforces emerging models of orthogonal behavioral trait axes or "animal personalities".

      (4) Clarity and Data Analysis:<br /> The analytical framework is well-suited to the study's complexity, with appropriate use of dominance metrics, mixed-effects models, and permutation tests. The analyses are clearly explained, statistically rigorous, and supported by transparent supplementary materials.

      Weaknesses:

      (1) Scope Limitations (Sex):<br /> The study is limited to male mice, which represents a common but problematic bias.

      (2) Ambiguity of Dominance as a Construct:<br /> While the study robustly quantifies social rank and hierarchy structure, the broader functional meaning of "dominance" remains unclear.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The goal of the study was to address the question of the degree to which social position in a group is a stable trait that persists across conditions. Reinwald et al. use a custom-built cage system with automated tracking and continuous testing for social dominance that does not require intervention by the experimenter. Remixing of individuals from different groups revealed that social position was rather stable and not really predictable from other measures that were taken. The authors conclude that social position is multifaceted but dependent on characteristics like personality traits.

      Strengths:

      (1) Reductionistic, highly controlled setting that allows for the control of many confounding variables.

      (2) Very interesting and important question.

      (3) Confirms the emergence of inter-individual behavior-driven differences in inbred mice in a shared environment.

      (4) Innovative paradigm and experimental setup.

      (5) Fresh perspective on an old question that makes the best use of modern technology.

      (6) Intelligent use of behavioral and cognitive covariables to generate a non-social context.

      (7) Bold and almost provocative conclusion, inviting discussion and further elaboration.

      We thank Reviewer #1 for this constructive and balanced evaluation of our work, and for highlighting both the conceptual importance of the question and the strengths of our automated, highly controlled approach.

      Weaknesses:

      (1) Reductionistic, highly controlled setting that blends out much of the complexity of social behavior in a community.

      Our goal in developing the NoSeMaze was to provide an enriched yet standardized environment that allows animals to interact freely in a complex setting where arena geometry and access contingencies are held constant across groups. As the Reviewer also notes as a strength, this highly controlled setting with minimal experimenter interference enables us to minimize confounds and provides reproducibility across rounds and groups. We now clarify this explicitly and frame the design as a trade-off between ecological complexity and experimental controllability.

      The NoSeMaze is designed as an open system that can incorporate additional configurations. In this first study with the system, we intentionally used a design that focused on single-sex groups without mating, no intruders or external threats, stable environmental conditions, and adult mice. This allowed us to establish social behavior under one defined condition. From here, future studies can add certain levels of ecological and social complexity to progressively understand their impact on specific behaviors and group dynamics.

      Accordingly, we have revised the manuscript to clarify the scope and the role of social and environmental conditions to the here observed phenomena. We also describe how future studies with systematically modified conditions can be used to understand how they change behaviors. We further replaced potentially over-broad terms (e.g., “naturalistic,” “real-world”) with more precise wording throughout.

      We modified the following sections:

      Abstract

      We removed “… within naturalistic mouse groups.” (ll. 52-54) and changed the sentence to “The approach thus enables longitudinal modeling of individuality and social position as key resilience factors.”

      We changed “… in naturalistic groups.” (l. 38) to “… in larger male mouse groups.” and “… from naturalistic tube competitions …” (ll. 41-42) to “… from incidental competitions in the integrated tube tests …”. We also deleted “naturalistic” in l. 50.

      Discussion (ll. 691-701)

      “…The NoSeMaze aims to increase environmental complexity and group dynamics while retaining experimental control, enabling longitudinal high-dimensional phenotyping of individuals. At the same time, it is a controlled laboratory group-housing habitat optimized to capture a subset of the determinants of social complexity present in natural communities. The strength of this design lies in the continuous, observer-independent observation of complex behaviors in defined environmental and social contexts. Accordingly, we interpret our findings as applying to the social contexts and environmental conditions tested here. Building on this, its modular design allows for introducing additional environmental and social factors like stressors, mating behavior, or resource competition in the future to understand their respective impact in modifying social behaviors.”

      Conclusion

      We changed “… complexity of real-world behavior …” (ll. 729-730) to “… complexity of group behavior in semi-naturalistic conditions ...”

      (2) The motivation to enter the test tube is not "trait" (or at least not solely a trait) but the basic need to reach food and water; chasing behavior would be less dependent on this stimulus.

      Tube traversals may reflect different motivations, including routine movement between compartments to access food, the water lickport, the open arena, or the housing area. Nevertheless, we do not interpret tube-entry motivation itself as a ‘trait’. Importantly, our hierarchy readout does not quantify which animals enter the tubes, nor is it confounded by tube-entry frequency itself (cf. ll. 527-532). Rather, it captures the consistent outcomes of incidental dyadic competitions once two animals meet in the tube (push vs. retreat), aggregated across many interactions. Thus, the stable signal we report lies in repeatable competition outcomes, not in traversal propensity. Consistent with this interpretation, the overall number of tube competition events was not associated with social rank, arguing against systematic competition avoidance by low-ranking animals. In contrast, chasing is a voluntarily initiated, asymmetric interaction between an initiator and a recipient and therefore adds a social-action component beyond incidental access-linked encounters, providing a complementary readout of social behavior.

      We also noted that the term ‘trait’ may not be the optimal description in this context and replaced it throughout the ms. with more precise wording, including “internalized social rank” and “propensity to chase”.

      Results

      We added the following paragraph (ll. 524-532):

      “Tube crossings in the NoSeMaze are motivated by the intent to eat, drink, sleep, or socialize. Accordingly, competitions within the tube arise by chance when two animals enter from opposite sides at the same time. Social rank therefore captures the consistent outcomes of repeated incidental competitions (push versus retreat). This measure is not confounded by differences in tube engagement, as social rank was neither associated with participation in tube competitions (ρ = 0.11, p = 0.139; Fig. 6C, Supplementary Fig. S12C) nor with the overall number of tube detections when controlling for chasing (Spearman’s partial correlation between detection count and z-scored David’s score, corrected for the fraction of active chases: ρ = 0.022, p = 0.758).”

      Discussion

      We changed the following paragraphs:

      Line 608-611

      “Importantly, participation frequency and differences in tube entry time did not confound the resulting social ranks, underscoring the robustness of the automated incidental rank assessment.”

      Lines 653-662

      “These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure.”

      (3) Dominance is only one aspect of sociality, social structure is reduced to rank. The information that might lie in the chasing behavior is not optimally used to explain social behavior beyond the rank measure.

      In this manuscript, we focused on the relationship between social rank derived from incidental tube competitions and chasing. We agree that social rank derived from tube competitions captures only one aspect of social structure. We have now clarified this point more explicitly in the revised manuscript. Here, our specific aim was to determine how chasing relates to competition-derived social rank and whether it provides information beyond rank in these mouse groups.

      We also refer to another manuscript dedicated to additional aspects of social structure captured in the video data, including approach and interaction behavior as well as social clique formation. There, these measures are again examined in relation to chasing and social rank. We also plan future studies leveraging chasing in the NoSeMaze as a readout for neurophysiological investigations.

      In the present study, social rank is based on dominance and subordination in incidental competitions in the integrated tube test. We treat chasing as a distinct, volitional social dimension rather than redundant “rank information”. Importantly, our chasing analyses add beyond social rank information in three ways: (1) structural asymmetry (initiator vs recipient roles) not captured by symmetric social rank measures; (2) elite-centric reciprocal dynamics rather than broad top-down enforcement; and (3) context dependence, with social rank–chasing coupling strengthening when group transitivity is lower. We revised the Discussion/Conclusion accordingly and additionally note that other aspects of social organization (e.g., affiliative bonding and higher-order network structure such as clique/rich-club organization; Nelias et al., 2025) require complementary measures (e.g., video-derived interaction networks) and are not the scope of the present study.

      To avoid ambiguity, we also added the following operational definitions in the Introduction and Results:

      Introduction

      Lines 90-94

      “In this study, we use social hierarchy to denote the group-level structure inferred from incidental competitions in the integrated tube tests, and social rank for an individual’s level within that hierarchy. Social position serves as an umbrella term for social rank and chasing behaviors.”

      Results

      Line 306-308

      “Here, social hierarchy refers to the group-level structure reconstructed from incidental competitions in the integrated tube tests, social rank to an individual’s position within that structure, and social position to social rank together with chasing.”

      We additionally clarified the distinction between chasing and social rank in the following sections.

      Discussion (ll. 632-664)

      “In more constrained or despotic conditions, chasing can serve as a unidirectional, dominance-related behavior directed at subordinates [9,43,44]. The larger groups observed in the NoSeMaze reveal a more nuanced role for chasing behavior. Chasing levels are individually stable, but the expression and meaning of chasing are context-sensitive. Chasing was neither broadly distributed nor consistently directed down the social hierarchy. Instead, it was initiated by a small subset of individuals – primarily those occupying high competition-based social ranks – and frequently occurred reciprocally within this group, suggesting intra-elite social dynamics rather than broad dominance enforcement. Rather than solely serving to impose social hierarchy, chasing appeared to function as a means through which individuals with high social rank monitor, negotiate, and maintain their relative standing within the top tier. Notably, the identity of frequent chasers remained stable over time and persisted across changing group compositions, indicating that the propensity to initiate chases reflects a consistent individual-level tendency rather than a purely situational response.

      However, its coupling to social rank depended on the group’s hierarchy structure. In groups with less clearly defined social hierarchies (i.e., lower transitivity), active chases aligned more strongly with social rank, suggesting that mice in less structured groups rely more on proactive signaling to clarify social rank. Indeed, the top-ranked mice in these groups exhibited relatively high levels of active chases. This context-sensitivity highlights chasing’s dual role: it serves as a tool for negotiating social rank among mice at the upper end of the hierarchy, and additionally functions to establish or reinforce hierarchical clarity when social structures are ambiguous. These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure. These findings reveal a novel aspect of social dynamics: chasing is not merely a dominance display but a flexible context-dependent mechanism shaped by both individual disposition and group-level social structure.”

      Conclusion (ll. 718-723)

      “Crucially, social position is not fully described by a single behavioral dimension. We focused here on two separable dimensions: competition-based social rank and proactive chasing. Future work should integrate additional dimensions of social organization, such as affiliative bonding and higher-order network measures [36], to capture additional aspects of social organization. Chasing behavior played a dual role, reflecting a stable individual propensity while also adapting to group-level structure.”

      (4) Focus on rank bears the risk of overgeneralization for readers not familiar with the context.

      As already discussed above (see also point 2 and 3), we have sharpened the framing to reduce the risk of overgeneralization. Throughout the manuscript, we now refer to tube-derived social rank explicitly as a dominance-subordination-related axis of social organization rather than a comprehensive measure of “sociality”, and we also avoid the term “traits”, but prefer the use of “internalized social rank” or “propensity to chase” when discussing stability across rounds and contexts. Also, we removed “personality-like” as the scope of the study is to set specific behaviors in relation and not to enter the field of mouse personality classification.

      We clarified the definition of social rank in this manuscript in the Introduction (ll. 90-94) and the beginning of the Results (ll. 306-308, see also point 3).

      Abstract

      Lines 41-46

      “… Across more than 4,000 mouse-days, hierarchies derived from incidental competitions in the integrated tube tests were non-despotic, transitive, and stable even when group compositions changed. This stability supports an internalized component of competition-based social rank. Chasing was also stable across contexts. Notably, chasing was concentrated among high-ranking individuals, consistent with ongoing negotiation of social rank among individuals at the upper end of the hierarchy.”

      Lines 50-54

      “In summary, high-dimensional tracking with the NoSeMaze reveals that social position in mice is multifaceted and shaped by stable dimensions of individual behavior that persist across changing social contexts. The approach thus enables longitudinal modeling of individuality and social position as key resilience factors.”

      Discussion (ll. 622-626)

      “This temporal and contextual stability supports interpreting social rank as a stable, internalized characteristic of individuals. In this sense, ‘internalized’ refers to stability across repeated rounds and reshuffled groups in this paradigm and to relatively stable tube-competition outcomes.”

      Changes in the Conclusion starting l. 709 as highlighted in answer to the above point 3.

      (5) Conclusion only valid for the reductionistic setting, in which environment, social and non-social changes only within narrow limits, and in which the mouse population does not face challenges

      Our conclusions are bounded by the conditions tested, namely an enriched but stable environment, controlled group composition and remixing, and the absence of explicit ecological challenges such as resource scarcity or predators. As described above in the answer to point 1, we now state explicitly that our conclusions apply to the social-context variation tested here (controlled group reshuffling) under otherwise stable conditions, and we frame resource-competition and environmental-stressor manipulations as future variation to test of how this manipulation affects the here described behaviors.

      We accordingly changed the Discussion as already highlighted in point 1 above and added ll. 693-704.

      (6) Animals are not naive at the beginning of the experiment, but are already several weeks old.

      Animals entered the study as adults and were continuously group-housed (3-5 mice/cage) before entering the NoSeMaze. To reduce effects of initial apparatus novelty, mice underwent two NoSeMaze habituation sessions prior to data collection (each several hours). We now clarify these points explicitly in the Methods and scope the inference accordingly. One of our future steps is to extend this framework to earlier developmental stages (e.g., adolescence/weaning) to capture full lifespan trajectories. This however first required establishing the approach in adult mice under controlled conditions as done in the present study.

      Methods (ll. 739-747)

      “A total of 79 adult male homozygous OXTRfl/fl mice (B6.129(SJL)-Oxtrtm1.1Wsy/J, RRID: IMSR_JAX:008471, Jackson Laboratory) backcrossed > F10 to C57BL/6J background (Charles River, Sulzfeld) were used for the experiments. Animals entered the NoSeMaze as adults (see Supplementary Table S2 for ages at NoSeMaze entry across rounds) and had been continuously group-housed after weaning (3–5 mice/cage), i.e., they were not developmentally or socially naïve at study onset. Of the 79 mice, 26 mice were injected six weeks before the start of the experiment with an AAV expressing Cre recombinase (rAAV1/2-CBA-Cre) into the AON pars centralis to induce bilateral OXTR deletion (OXTRΔAON). The remaining 53 animals received an AAV expressing only dTomato (rAAV1/2-CBA-dTomato). …”

      Discussion (ll. 629-631)

      “… A future direction is to extend this framework to earlier developmental stages to understand which early experiences shape later trajectories of social position.”

      In summary, this is a wonderful study, but not one that is easy to interpret. The bold conclusion is valid only within the constraints of the study, but nevertheless points in an important direction. The paradigm is clever and could be used for many interesting follow-ups.

      To define social position as a personality trait will elicit strong opposition and much debate; the nuances of the paper might be lost on many readers and call for the (re)-consideration of many concepts that are touched. I find this attitude a strength of the paper, but the approach bears the risk of misunderstanding.

      We thank the Reviewer for the helpful comments. As detailed above, we tightened terminology and framing throughout the revised manuscript by defining stability explicitly in terms of repeatability across rounds and reshuffled groups within this paradigm, and described tube-derived social rank as one dimension of social behavior. We hope these edits preserve the conceptual message while reducing the risk of misunderstanding.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents the "NoSeMaze", a novel automated platform for studying social behavior and cognitive performance in group-housed male mice. The authors report that mice form robust, transitive dominance hierarchies in this environment and that individual social rank remains largely stable across multiple group compositions. They further demonstrate that social dominance and aggressive behaviors, like chasing, are partially dissociable and that dominance traits are independent of non-social cognitive performance. The study includes a genetic manipulation of oxytocin receptor expression in the anterior olfactory nucleus, which showed only transient effects on social rank.

      Strengths:

      (1) Innovative Methodology:

      The NoSeMaze platform is a technically elegant and conceptually well-integrated system that enables fully automated, long-term monitoring of both social and cognitive behaviors in large groups of group-housed mice. It combines tube-test-like dominance contests, voluntary chase-escape interactions, and an embedded operant olfactory discrimination task within a single, ethologically relevant environment. This modular design allows for high-throughput, minimally invasive behavioral assessment without the need for repeated handling or artificial isolation.

      (2) Experimental Scale and Rigor:

      The study includes 79 male mice and over 4,000 mouse-days of observation across multiple group reshufflings. The use of RFID-based identification, automated data logging, and longitudinal design enables robust quantification of individual trait stability and group-level social structure.

      (3) Multidimensional Behavioral Profiling:

      The integration of social (tube dominance, proactive chasing), physical (body weight), and cognitive (olfactory learning task) measures offers a rich, multi-dimensional profile of each individual mouse. The authors' finding that social dominance traits and non-social cognitive performance are largely uncorrelated reinforces emerging models of orthogonal behavioral trait axes or "animal personalities".

      (4) Clarity and Data Analysis:

      The analytical framework is well-suited to the study's complexity, with appropriate use of dominance metrics, mixed-effects models, and permutation tests. The analyses are clearly explained, statistically rigorous, and supported by transparent supplementary materials.

      We thank the Reviewer for their evaluation of our manuscript. We appreciate their recognition of the NoSeMaze as an innovative and conceptually integrated platform, the scale and longitudinal rigor of the dataset, and the clarity and appropriateness of the analytical framework.

      Weaknesses:

      (1) Conceptual Novelty and Prior Work:

      While the study is carefully executed and methodologically innovative, several of its core findings reaffirm concepts already established in the literature. The emergence of stable, transitive social hierarchies, the persistence of individual differences in social behavior, and the presence of non-despotic social structures have all been previously reported in mice, including under semi-naturalistic conditions (e.g., Fan et al., 2019; Forkosh et al., 2019). Although this work extends those findings with greater behavioral resolution and scale, the manuscript would benefit from a clearer articulation of what is genuinely novel at the conceptual level, beyond the technological advance.

      We agree with the Reviewer that transitive dominance hierarchies and stable inter-individual differences have been demonstrated previously in mice, including in semi-naturalistic settings. Our intent was therefore to address two more specific conceptual questions that prior work typically has not tested in an integrated way:

      (1) Whether an individual’s social position generalizes across distinct social contexts created by systematic changes in group composition rather than reflecting stability that is only observable within a fixed group;

      (2) How the two frequently employed social dominance metrics tube competition-based social rank and chasing relate to each other in unperturbed larger groups.

      Specifically, the conceptual advance is enabled by (A) continuous, handling-free estimation of competition-based social rank from incidental tube contests over weeks, (B) a repeated group-reshuffling (“accelerated longitudinal”) design that explicitly tests whether an individual’s social rank generalizes across distinct social contexts, and (C) parallel, continuous measurement of chasing and non-social reinforcement-learning behaviors in the same individuals and environment.

      This integration also lets us dissociate dominance-related social dimensions (competition-based social rank vs chasing) and test their relation to individual styles in non-social reinforcement-learning. Importantly, these questions are addressed in larger intact societies (9–10 mice), where hierarchy structure is shaped by more complex network-level dynamics than in smaller groups.

      We revised the Introduction and Discussion to foreground these conceptual points and to position them more explicitly in the context of prior works.

      Introduction

      We adapted the following section (ll. 76-98) and integrated the suggested references:

      “… In mice, small groups tend to form highly despotic hierarchies [24-26], whereas larger groups exhibit more complex structures [27]. These patterns suggest a strong influence of emergent group-level dynamics on social structure [12,27]. Yet, animals do not enter social groups as blank slates [28-30]; stable latent factors in the individual may also contribute to hierarchy formation [13]. Thus, it remains unclear to what extent an individual’s social position is internalized and persists across different social contexts [31], or instead is primarily an emergent property of group-level dynamics. Disentangling these possibilities requires experimental conditions that allow unperturbed, continuous tracking of all individuals in sufficiently large groups, together with systematic changes of group composition to modulate social context.

      While stable hierarchies and behavioral identity domains have been described previously in semi-naturalistic settings [9-13], many studies quantify these features either within fixed group compositions or in separate assays. Here, we therefore use systematic group reshuffling in 10-member societies to directly test whether individual differences in social rank and chasing persist across distinct social groups. In this study, we use social hierarchy to denote the group-level structure inferred from incidental competitions in the integrated tube tests, and social rank for an individual’s level within that hierarchy. Social position serves as an umbrella term for social rank and chasing behaviors. By continuously measuring social rank, chasing, and reinforcement-learning behavior in the same individuals, we further test how chasing and social rank are related to each other and how they relate to individual styles in non-social reinforcement learning.

      To enable these tests, we developed the Non-invasive Sensor-rich Maze (NoSeMaze).”

      Discussion

      We added a brief framing statement at the beginning of the Discussion (ll. 577-583) to clarify the study’s conceptual novelty:

      “Ecologically enriched, yet experimentally controlled assessments allow us to study behavioral individuality and social structure in group-living animals over extended timescales. Here, we show that individual mice carry stable, individual-specific, and multi-faceted profiles of social position and cognitive styles across changing group contexts. Our approach goes beyond prior work by testing cross-context stability under repeated, systematic group reshuffling in larger mouse societies, while measuring competition outcomes, chasing, and reinforcement-learning behavior in parallel. …”

      (2) Role of OXTR Deletion:

      The inclusion of the OXTR manipulation feels somewhat disconnected from the manuscript's central aims. The effects were minimal and transient, and the authors defer full interpretation to a separate study.

      We appreciate the Reviewer’s point and agree that the OXTR<sup>ΔAON</sup> manipulation can appear secondary to the manuscript’s central aims. We included OXTR<sup>ΔAON</sup> because oxytocin-dependent social recognition memory was hypothesized originally to impact potentially also learning an individual’s position in social hierarchy networks. Even though the effects were small and transient, we nevertheless believe it is relevant to report them, also in relation to the more profound effects of the genetic manipulation reported in a related manuscript (Nelias et al., bioRxiv 2025, 10.1101/2025.08.26.672298). Therefore, we explicitly account for genotype as a covariate in the analyses such that the main conclusions do not depend on this manipulation.

      To improve coherence, we have added one sentence in the Introduction motivating why OXTR<sup>ΔAON</sup> was included, and finally more clearly signposted that deeper mechanistic interpretation is beyond the scope of the present manuscript.

      Introduction (ll. 115-120)

      We added one short sentence introducing the rationale behind the perturbation:

      “… and (3) determine whether social rank, chasing, and non-social reward-seeking behaviors represent stable individual characteristics or dynamic features across time and changing group composition. As a secondary analysis, motivated by oxytocin’s established role in social recognition memory [32,33], we also tested whether OXTR deletion in the anterior olfactory nucleus produces detectable shifts in rank dynamics.”

      Results

      Lines 161-164

      “… This manipulation was included as a secondary biological perturbation. The primary analyses and conclusions focus on the platform and cross-context stability, and genotype is treated as a covariate unless stated otherwise. …”

      Lines 451-454

      “… As a secondary analysis, we tested whether OXTR<sup>ΔAON</sup>, which impairs de novo social recognition memory required for social clique formation in this cohort [36], also affects the measures reported here. Consistent with largely internalized features, OXTR<sup>ΔAON</sup> produced only transient effects. …”

      Discussion (ll. 594-605)

      “This study focused primarily on the relation of social rank and chasing. We however also considered their relation to additional variables including the loss of oxytocin receptors in the olfactory cortex in the adult (OXTR<sup>ΔAON</sup>), involved in de novo social recognition learning. The propensity to chase was largely unaffected by OXTR<sup>ΔAON</sup>. Mice carrying OXTR<sup>ΔAON</sup> displayed a transient reduction in social rank during the first week that normalized thereafter. This transient effect contrasts to the persistent impairment by OXTR<sup>ΔAON</sup> in forming higher-order social bonds that enable membership in stable cliques, as identified by video tracking of self-paced interactions in the same cohort [36]. Together, these findings suggest that OXT-dependent olfactory learning is critical for the formation of social context-dependent higher-order bonds, but plays a limited role in shaping hierarchy-related behaviors.”

      (3) Scope Limitations (Sex and Age):

      The study is limited to male mice, and although this is acknowledged, the title and overall framing imply broader generalizability. This sex-specific focus represents a common but problematic bias. Additionally, results from the older mouse cohort are under-discussed; if age had no effect, this should be explicitly stated.

      We thank the Reviewer for this relevant point. The study is limited to male mice. We therefore revised the title and the abstract and strengthened the limitations to make the sex-specific scope explicit.

      We additionally note ongoing work extending the same framework to female groups, where we find similar hierarchy structure and cross-context stability in the NoSeMaze in preliminary unpublished data.

      Regarding age, our design included two separate cohorts of different adult age ranges (young and older adult animals), but age was not the primary experimental factor. As detailed in our response to Reviewer #3 (point 1), we now quantify age structure explicitly and test age effects using a decomposition that separates between-group age differences from within-group age variation (mean age per group and each animal’s deviation from that mean). We also recomputed stability estimates with age-adjusted ICC models, and the resulting ICCs are highly similar to the original estimates (cf. new Supplementary Table S4), indicating that the reported metrics’ stability is not driven by age differences across groups.

      Title

      We change the title from “Individual differences drive social hierarchies in mouse societies” to “Individual differences drive social hierarchies in male mouse societies”

      Abstract

      We also added male in the abstract (ll. 37-38):

      “The interaction of these behaviors in the shaping of social position in larger male mouse groups remains largely unknown.”

      Discussion (ll. 613-617)

      “… While this study focused on male mice, in which social hierarchies are best established [9], future work is needed to explore sex-specific expressions of social structure and their neurobiological underpinnings in female groups. We therefore restrict our interpretation to male mice. Critically, male social ranks were robustly maintained within the same group over time. …”

      For a more detailed integration of age in the Methods, Results, and Discussion, we kindly refer to the reply to Reviewer #3, point 1.

      (4) Ambiguity of Dominance as a Construct:

      While the study robustly quantifies social rank and hierarchy structure, the broader functional meaning of "dominance" remains unclear. As in prior work (e.g., Varholick et al., 2019), dominance rank here shows only weak associations with physical attributes (e.g., body weight), cognitive strategy, or neuromodulatory manipulation (OXTR deletion). This recurring pattern, where rank metrics are reliably established yet poorly predictive of other behavioral or biological traits, raises important questions about what such measures actually capture. In particular, it challenges the assumption that outcomes in paradigms like the tube test or chase frequency necessarily reflect dominance per se, rather than other constructs.

      We thank the Reviewer for this clarifying point. We agree that stable social rank metrics do not necessarily imply a complete or unitary measure of “dominance.” In the revised manuscript, we therefore clarified that, in this study, social rank is operationalized as consistent competitive outcomes in incidental tube-test encounters in the NoSeMaze (see also Reviewer #1, points 2-4).

      Our data indicate that this competition-based social rank is related to, but not identical with, other social behaviors such as chasing. Likewise, body weight significantly contributes to social rank, but explains only part of the variance. We therefore do not interpret weak or partial associations with other variables as invalidating the social rank measure. Rather, we interpret them as indicating that social position is multidimensional and only partially captured by any single assay.

      In independent subsequent studies that are currently in preparation or revision, we observed that heterogeneity in the neurobiology and response to challenges was best predicted by the competition-based social rank, also compared to the other behaviors assessed here. While these observations are beyond the scope of the present manuscript, they support our view that incidental tube competition and the resulting dominance-subordination structure may provide a biologically informative measure of one important dimension of social position. At the same time, the observation here and in many previous studies that factors such as body weight explain only a limited portion of the variance remains important for our understanding of these constructs.

      We have revised the manuscript accordingly to make this distinction more explicit. In particular, we now define more precisely the terms social hierarchy, social rank, and social position in the context of this manuscript (see also reply to Reviewer #1, point 3; ll. 90-94 in the Introduction and ll. 306-308 in the Results), and we use these terms more consistently throughout. We also revised the text to avoid overstating the meaning of “dominance” where the data support a more specific interpretation.

      Finally, we now emphasize more clearly both the strength and the limitation of the present approach. The NoSeMaze allows these relationships to be assessed continuously in a minimally perturbed group-housing ecology, laying ground for future incorporation of further variables and dimensions to capture how they shape social organization.

      Specifically, we added text in the Discussion and Conclusion to clarify that competition-based social rank and proactive chasing represent separable dimensions related to dominance and subordination, but do not exhaust sociality or individuality, and that future work should integrate additional measures such as affiliative behavior and higher-order network structure.

      Discussion

      Lines 624-6631:

      “… In this sense, ‘internalized’ refers to stability across repeated rounds and reshuffled groups in this paradigm and to relatively stable tube-competition outcomes. The tube-derived social rank describes here a dimension of individual social behavior. The capacity to quantify stable individual differences across changing social contexts highlights the value of the NoSeMaze for lifespan-oriented studies of behavioral individuality. A future direction is to extend this framework to earlier developmental stages to understand which early experiences shape later trajectories of social position.”

      Conclusion

      Lines 718-722:

      “Crucially, social position is not fully described by a single behavioral dimension. We focused here on two separable dimensions: competition-based social rank and proactive chasing. Future work should integrate additional dimensions of social organization, such as affiliative bonding and higher-order network measures [36], to capture additional aspects of social organization.”

      Reviewer #3 (Public review):

      Reinwald et al. present the NoSeMaze, a semi-natural behavioral system designed to track social behaviors alongside reinforcement-learning in large groups of mice. Accumulating more than 4,000 days of behavioral monitoring, the authors demonstrate that social rank (determined by tube competitions) is a stable trait across shuffled cohorts and correlated with active chasing behaviors. The system also provides a solid platform for long-term measurements of reinforcement learning, including flexibility, response adaptation, and impulsiveness. Yet, the authors show that social ranking and chasing are mostly independent of these cognitive traits, and both seem mostly independent of oxytocin signaling in the AON.

      Strengths:

      (1) The neuroethological approach for automated tracking of several mice under semi-natural conditions is still rare in social behavioral research and should be encouraged.

      (2) The assessment of dominance by two independent measures, i.e., spontaneous tube competitions and proactive chasing, is innovative and valuable.

      (3) The integration of a long-term reinforcement-learning module into the semi-natural system provides novel opportunities to combine cognitive traits into social personality assessments.

      (4) The open-source system provides a valuable resource for the scientific community.

      Limitations:

      (1) Apparent ambiguity and inconsistency in age structure and cohort participation across rounds, raising concerns about uncontrolled confounds.

      (2) Chasing behavior appears more stable than tube-test competitions (Figure 4D vs. Figure 3D), which challenges the authors' decision to treat tube competitions as the primary basis for hierarchy determination.

      We thank the Reviewer for the evaluation of our work. We have addressed the limitations raised below with additional analyses, clarifications, and corresponding manuscript revisions.

      Major concerns:

      (1) Unclear and inconsistent handling of age groups and repeated sampling. The manuscript repeatedly refers to "younger" and "older" adults, but it is unclear whether age was ever controlled for or included in models. Some mice completed only one round, others 2-5 rounds, without explanation of the criteria or balancing.

      We thank the Reviewer for this clarifying comment. We now explicitly quantify participation in the different rounds in new Supplementary Table S3. Importantly, our primary stability analyses are implemented using variance-component mixed models (REML) that naturally handle unbalanced repeated-measures data. This is explained in more detail in the Methods section (ll. 1003-1111), where we added a statement that the LMEs are well suited for unbalanced repetitions.

      To further address unbalanced participation, we additionally performed a conservative sensitivity analysis restricted to a balanced subset, including only sessions 1 and 2 and only mice with observations in both sessions. For the key social measures, ICC estimates were highly similar in the full dataset and in the balanced subset (e.g., z-scored competition David’s score, ICC across cohorts: 0.55 without age adjustment vs. 0.56 in the balanced first-two-session subset; active chasing: 0.74 vs. 0.72; being chased: 0.61 vs. 0.60; Supplementary Table S4), indicating that unbalanced participation did not inflate the stability estimates.

      We also addressed age structure explicitly. Although age was not the primary experimental factor, the inclusion of two age cohorts allowed us to assess whether the observed behaviors and their interrelations were robust across most of the adult lifespan (cf. new Figure 1). We therefore decomposed age into a between-group component (mean age per group) and a within-group component (each animal’s deviation from its group mean) and included these terms in the relevant LME models. Tube-based dominance rank (David’s score, z-scored) showed no age effect (p<sub>age, cond.</sub> = 0.63), whereas chasing metrics showed modest age associations (cf. new Supplementary Table S5). This however only indicates that chasing was associated with age to some degree. More importantly, recomputing all stability estimates using age-adjusted ICC models yielded nearly identical ICCs for the core social measures (new Supplementary Table S4), indicating that the reported stability was not affected by age differences across groups.

      In addition, we revised the study design schematic (new Fig. 1; cf. Reviewer #1, Recommendations for the authors) to depict the separate age cohorts and the reshuffling procedure more clearly.

      We adapted the following sections accordingly.

      Methods

      Lines 980-985

      “… Most mice (n = 68) participated in at least two NoSeMaze rounds with reshuffled group members. Supplementary Table S2 summarizes the number of rounds per mouse and missing data due to technical problems. We examined the stability of social and reward-seeking metrics by correlating values from the first and second round (Spearman’s correlation, cf. Fig. 5). These round-1-to-round-2 correlations use one paired observation per mouse and are therefore not inflated by mice contributing >2 rounds.”

      Lines 991-993

      “The ICC treated mouse identity as a random intercept and NoSeMaze group (i.e., round-specific social group) as a random effect, with repetition included as fixed effect (i.e., stability across groups while holding repetition means constant).”

      Lines 1003-1011

      “Variance components for ICC estimation were obtained from LMEs fit by restricted maximum likelihood (REML), which yields less biased variance-component estimates and is well suited for unbalanced repeated-measures designs (i.e., different numbers of rounds per mouse). To assess potential confounding by age structure, age was decomposed into a between-group component (group-mean age) and a within-group component (each animal’s deviation from its group mean) and included as covariates. ICCs were recomputed in age-adjusted models (see Supplementary Table S4). Finally, we performed a conservative sensitivity analysis restricted to the first two sessions per mouse and to mice with observations in both sessions (“balanced first2”), to additionally account for unbalanced participation structure.”

      Results

      Lines 167-173

      “… The study population comprised two age cohorts: younger (16-30 weeks) and older adults (55-97 weeks) (Fig. 1B). The two age cohorts were run as separate experimental series, and group reshuffling was performed within each cohort (Fig. 1C, see Supplementary Table S2). Mice lived in groups of 9-10 for multiple rounds in the NoSeMaze, with different group members in each round (Fig. 1D, see Supplementary Table S1). This allowed us to test which individual behaviors were stable across different group compositions. …”

      Lines 441-450

      “Because the number of rounds in the NoSeMaze was unbalanced between animals (Supplementary Table S3) and age varied across groups (Supplementary Table S1-2), we also performed robustness checks. Stability estimates changed only minimally when recomputed in age-adjusted ICC models (between-group mean age and within-group age deviation; see Methods) and when restricting analyses to a balanced first-two-session subset (sessions 1–2 only; mice with both sessions) (Supplementary Table S4). Mixed models indicated that some chasing and reinforcement-learning measures showed modest age- and/or session-related shifts in absolute levels (Supplementary Table S5), but importantly, these did not affect the observed stability patterns.”

      (2) Stability of chasing appears stronger than the stability of tube competitions. Figure 4D shows highly consistent chasing behavior across weeks, while Figure 3D shows weaker and more variable correlations for tube-based David scores. This is also evident from Figure 5A-B,D. Thus, it appears that chasing, which serves to quantify dominance in similar semi-natural setups, may be a more reliable and behaviorally meaningful measure of dominance than the incidental tube competitions.

      Indeed, active chasing showed higher cross-round correlations than tube-derived social rank (e.g., R1–R2 Spearman ρ = 0.75 vs 0.57; ICC<sub>across cohort</sub> 0.74 vs 0.55, see Fig. 5 and new Supplementary Table S4). Importantly, this does not contradict the central finding that tube-derived social rank is stable across time and across remixed groups. Rather, it may highlight that chasing and social rank capture different aspects of dominance-subordination-related behavior with different statistical properties. Chasing reflects an individual’s propensity to actively initiate interactions (a strongly expressed, asymmetric behavior), which can be highly consistent across contexts. By contrast, social rank is a relational measure inferred from symmetric dyadic win–loss outcomes based on incidental competitions in the integrated tube tests and can vary with the specific set of competitors and interaction opportunities in each reshuffled NoSeMaze group, while still remaining substantially stable overall.

      Accordingly, we do not interpret the higher repeatability of chasing as evidence that it is the “better” hierarchy measure. Tube competitions yield symmetric dyadic outcomes that directly support formal hierarchy reconstruction (David’s score/Elo, transitivity, steepness) and show convergent validity with traditional tube testing. Chasing, in contrast, is asymmetric and volitional and, in our data, is concentrated in the upper social ranks and modulated by group-level social hierarchy structure, consistent with rank negotiation/signaling rather than a mechanism that assigns a full ordering to all individuals. These different properties lead to very different “win-lose” relationships when comparing tube competition events to chasing events that we specifically illustrated in Supplementary Fig. S13. We therefore revised the Results and Discussion to frame tube-derived social rank and chasing as complementary social dimensions.

      Results (ll. 402-413)

      “… Specifically, stability was high for the tube competition-based David’s score (Fig. 5A, ρ = 0.57, p < 0.001), as well as for the fraction of active chases (Fig. 5B, ρ = 0.75, p < 0.001) and of times being chased (Fig. 5C, ρ = 0.54, p < 0.001). Notably, the fraction of active chases showed slightly higher across-round stability than tube-derived David’s score and the fraction of being chased. This is in line with active chases capturing an individual propensity to initiate this behavior, whereas competition-based social rank is a relational measure that is also influenced by the set of competitors and interaction opportunities in each reshuffled group. Nonetheless also social rank and being chased were overall stable. Across the full series, ICCs for these social measures were also in the good–excellent range, indicating high across-round stability (Fig. 5D, ICC = 0.548 to 0.890).”

      Discussion

      Line 653-664 (see also Reviewer #1, point 2)

      “These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure.”

      (3) Unbalanced participation across rounds compromises stability analyses. Stability analyses (e.g., ICCs, round-to-round correlations) assume comparable sampling across individuals. However, some mice contribute 1 round, others 2, 3, 4, and even 5 rounds. This imbalance may inflate stability estimates or confound group reshuffling effects, and the rationale for variable participation is not explained.

      We thank the Reviewer for raising this point. Indeed, participation was unbalanced across rounds (see the same Reviewer #3, Major concerns 1 and new Supplementary Table S3), mainly due to missing data from occasional technical failures of the RFID detectors or the reinforcement learning water port during acquisition in some groups (for details, see Supplementary Table S2). We added this rationale behind variable participation to our Methods section (for details, see Major concern 1, ll. 981-982, “Supplementary Table S2 summarizes the number of rounds per mouse and missing data due to technical problems.”)

      We now quantify round participation (new Supplementary Table S3) and directly address potential bias from unequal sampling in two ways. First, the round-1-to-round-2 (R1–R2) stability correlations use one paired observation per mouse and therefore are not inflated by mice contributing more than two rounds. Second, ICCs were estimated from REML variance-component mixed models that account for unbalanced repeated-measures by design and use all available observations. To further rule out inflation from unequal sampling or non-random missingness, we additionally report a conservative sensitivity ICC restricted to a balanced subset including only each mouse’s first two observed sessions and only mice with both sessions (“first2-balanced”). Full-sample ICCs and first2-balanced ICCs were highly similar (new Supplementary Table S4), indicating that participation imbalance did not affect the stability estimates. For details on the changes made in the manuscript, see Reviewer #3, Major concern 1.

      Recommendations for the authors: 

      Editor's notes:

      Should you choose to revise your manuscript, if you have not already done so, please include full statistical reporting including exact p-values wherever possible alongside the summary statistics (test statistic and df) and, where appropriate, 95% confidence intervals. These should be reported for all key questions and not only when the p-value is less than 0.05 in the main manuscript.

      Readers would also benefit from noting that the mice were male in the abstract.

      We have revised the manuscript accordingly and now report exact p-values values wherever possible for the key results in the main text. Because full statistical reporting for every analysis in the main text would substantially interrupt readability, we provide the complete statistical details in a Supplementary Excel File (Supplementary Material – Systematic Statistical Reporting), including sample sizes, degrees of freedom, the number and type of permutation tests, exact p-values, and 95% confidence intervals. We also provide Extended Data Sheets for all linear-mixed effects models that account for covariates such as age and round of participation in the NoSeMaze (cf. Reviewer #3, Major Concern (1) for the additional analyses). To guide readers to these resources, we now explicitly refer to the Supplementary Material – Systematic Statistical Reporting at several points in the manuscript:

      Results

      Lines 271-274

      “Detailed statistical reporting for all analyses, including n, degrees of freedom, exact p-values, and 95% confidence intervals, is provided in the Supplementary Material - Systematic Statistical Reporting. Extended Data includes additional linear mixed-effects models controlling for potential confounding variables.”

      Lines 430-431

      “All statistical details are provided in the Supplementary Material – Systematic Statistical Reporting.”

      Additional references to the supplementary statistical reporting were inserted at lines 252-254, 562-563, and 567-568.

      Methods

      Lines 962-965

      “Full details on the statistical tests, including n, degrees of freedom, exact p-values, and 95% confidence intervals, are provided in the Supplementary Material – Systematic Statistical Reporting, as well as in the Extended Data for the LMEs accounting for different covariates.”

      We also revised the abstract and the title to explicitly state that the mice were male.

      Manuscript changes:

      Results, figure legends, and supplementary tables expanded to include full statistical reporting; abstract revised to specify male mice.

      Reviewer #1 (Recommendations for the authors):

      I would recommend being much more explicit about the reductionistic nature of the study and how the limitations are turned here to an advantage, while at the same time acknowledging the challenges of extrapolating beyond these boundaries. Most importantly, dominance should be positioned more clearly and cautiously within a framework of social behavior (and social structure) in general. The authors include cognitive tests, etc., to generate context, but this context is dependent on the same circumstances that possibly contribute to the social structure. The information that lies in the chasing behavior as an additional measured variable might be used better to provide more context.

      We thank the Reviewer for these recommendations. We have better clarified the framing to make explicit that the NoSeMaze is a controlled laboratory group-housing system. The specific conditions are now discussed in more detail in relation the observed social behaviors. We describe the tube-derived social rank as one dimension of social organization, and proactive chasing as another one. This study presents a first necessary step to understand their shared and distinguishing features. We also elaborated the Discussion to better clarify the trade-off between ecological complexity and experimental control, and to highlight that the value of the NoSeMaze for continuous, observer-independent phenotyping under standardized conditions. We now clearly state the importance to vary conditions in the system to see how the social behaviors and also their relation to non-social features changes depending on context conditions. For detailed changes, see our responses to Reviewer #1, points (1), (3), (4), and (5), and Reviewer #2, Weakness (4).

      Manuscript changes:

      Abstract, Introduction, Discussion, and Conclusion revised to clarify scope, construct interpretation, and the complementary roles of competition-based social rank and chasing.

      The precision of the description of the experimental design should be improved. When were the cohorts mixed, or did they stay separate? Which animals were old, which were young? This remained a bit confusing.

      We revised the presentation of the study design accordingly. Specifically, we clarified that the younger and older adult cohorts were run as separate experimental series and that group reshuffling occurred within, but not across, these cohorts. We also revised the study schematic in Fig. 1 and the corresponding description in the manuscript to more explicitly depict cohort structure, timing of NoSeMaze rounds, and between-round reshuffling (details are provided in the Supplementary Tables 1-3). For new age-related analyses and additional robustness checks, see our response to Reviewer #3, Major Concern (1).

      Manuscript changes:

      Figure 1 and its legend were revised to depict cohort structure, NoSeMaze rounds, and between-round reshuffling more explicitly. In addition, the Results subsection “Ecological longitudinal assessment in the NoSeMaze” were updated to clarify that the younger and older adult cohorts were run separately, that reshuffling occurred within but not across cohorts, and which animals belonged to each age-defined cohort. We also now cross-reference Supplementary Table S2 for round-specific age information and cohort composition.

      Results

      Lines 167-170

      “The study population comprised two age cohorts: younger (16-30 weeks) and older adults (55-97 weeks) (Fig. 1B). The two age cohorts were run as separate experimental series, and group reshuffling was performed within each cohort (see Supplementary Table S2).”

      It is also not fully clear across which groups stability measures were obtained: are these across all groups or within the subgroups with a given characteristic?

      We thank the Reviewer for highlighting this point. We now state explicitly that stability measures were computed across all eligible rounds, using mixed-effects models that account for repeated observations of the same mouse and for round-specific group membership. We also added a conservative sensitivity analysis restricted to a balanced first-two-session subset. Full and balanced-subsample stability estimates were highly similar, indicating that the main conclusions are robust to the participation structure. For details, see our response to Reviewer #3, Major Concerns (1) and (3).

      Manuscript changes:

      Methods and Results revised to clarify the level of analysis and model structure, as well as new supplementary robustness tables (Supplementary Tables 3-5) added.

      Minor point: In Figure 3, 18 groups are mentioned, but 19 are shown.

      We thank the Reviewer for this clarification. The apparent discrepancy arose because the group labels in Figure 3 follow the numbering of all experimental groups, whereas only groups with available tube-competition data are shown in this panel. Accordingly, group 16 is absent because no tube data were available for that group (see Supplementary Table S2), and groups 20 and 21 are likewise not included for the same reason. We have revised the figure legend and corresponding text to make this explicit and to avoid the impression of a numbering inconsistency.

      “Fig. 3: Social rank derived from incidental competitions in the integrated tube tests of the NoSeMaze.”

      “C, Box plots of metrics characterizing social hierarchy for 18 groups, including transitivity, steepness, stability, and uncertainty-by-repeatability (for details, see ‘Source Data’). Group labels correspond to original experimental group IDs. Only groups with available tube-competition data are shown. Therefore, numbering is non-consecutive (e.g., groups 16, 20, and 21 are absent; see Supplementary Table S2).”

      Reviewer #2 (Recommendations for the authors):

      (1) To better distinguish this study from previous literature, we recommend incorporating a more focused discussion (or adding to the intro) of how the findings advance our understanding of social hierarchy beyond prior works.

      We have sharpened the conceptual framing in both the Introduction and Discussion. In particular, we now distinguish more explicitly between the well-established observation that hierarchies form and remain stable within fixed semi-naturalistic groups, and the more specific question addressed here: whether an individual’s social position generalizes across changing social contexts created by repeated group reshuffling. We added additional references and also emphasize that the present study integrates continuous measurements of competition-based social rank, chasing, and reinforcement-learning features in the same individuals and environment. For details, see our response to Reviewer #2, Weakness (1).

      Manuscript changes:

      Introduction and Discussion revised to foreground conceptual novelty relative to prior semi-naturalistic work.

      (2) Given the weak associations between dominance rank and other traits such as body weight, cognitive performance, and oxytocin receptor manipulation, we suggest further clarifying what is being captured by these measures.

      We agree and have clarified this point throughout the manuscript. We now define tube-derived social rank explicitly as an operational measure based on repeated competitive outcomes from incidental dyadic tube tests, highlighting it as one dimension of social behavior.

      We also make clearer that proactive chasing is not redundant with rank, but instead captures a distinct, partly dissociable dominance-related interaction mode. For details, see our response to Reviewer #2, Weakness (4).

      We understand the question on the meaning of social rank as the correlations to body weight and cognitive performance are only punctual. We would like to mention here already that the competition-based social rank turns out to be a strong predictor of individual reactivity in a series of challenges. These works are in currently in preparation for publication and will make the relevance of these measures more clear. Related to this, we find in these studies that chasing and competition-based social rank predict different behavioral and neuronal aspects of individual reactivity.

      Manuscript changes:

      Abstract, Results, Discussion, and Conclusion revised to clarify construct interpretation and multidimensionality.

      (3) We recommend modifying the title and abstract to more clearly reflect the male-only design of the study. In addition, please indicate whether any age-related differences were observed. If age had no measurable effect, this should be stated explicitly to justify the combination of age groups.

      We revised the title and abstract to make the male-only design explicit. We also now report age-related analyses directly in the manuscript. Briefly, age was modeled by separating between-group age structure from within-group age variation. Some measures showed modest age associations in mean level, but most importantly, age-adjusted stability estimates for the core social metrics were highly similar to the original estimates, indicating that the reported stability is not driven by age differences across groups. For details, see our response to Reviewer #3, Major Concern (1).

      Manuscript changes:

      Title and abstract revised. Methods, Results, and supplementary robustness analyses expanded to report age effects explicitly.

    1. eLife Assessment

      This important study provides evidence that locus coeruleus activity is coordinated with heart rate during sleep, confirming previous work in mice and humans, with a possible role for sleep-dependent memory consolidation. The claims are supported by convincing evidence. This work will be of interest to neuroscientists focusing on sleep, memory, and autonomic functions.

    2. Reviewer #2 (Public review):

      Summary:

      This convincing study builds on previously published findings in both mice and humans to advance quantitative insights into the coupling between noradrenergic activity fluctuations during mouse NREM sleep and heart rate fluctuations. The work reaffirms the presence of coordinated infraslow fluctuations in sigma power and heart rate during NREM sleep and that this coordination is enabled by noradrenaline-releasing neurons in the locus coeruleus. Also supporting previously published work in mice and humans, the authors describe a link between the strength of these infraslow fluctuations and memory consolidation in mice and humans.

      Strengths:

      A major finding of this study is the mechanistic insight it provides into the regulation of the previously understudied very-low-frequency (0-0.15 Hz) component of heart rate variability, and the demonstration, through elegant optogenetic bidirectional interference, that infraslow noradrenergic fluctuations are an underlying driving force. This finding will promote recognition of heart rate variability in sleeping mice as a read-out of neuronal activity patterns that control autonomic balance.

      Another strength of the study is its translational part, whereby the sigma power-heart rate coupling in mouse is used to identify a previously unrecognized correlation between such coupling and memory consolidation in humans. This widens the applicability of heart rate variability measures, highlighting their use as biomarkers for noradrenergic fluctuations and associated sleep-dependent memory consolidation.

      Weaknesses:

      The study impresses by the thorough parallel analysis of both mouse and human correlational data between electrophysiological and fluorescent activity measures of the sleeping brain. Further work will be needed to disentangle the mechanisms by which heart rate is regulated, notably the contribution of parasympathetic and sympathetic nervous systems, to establish the very low frequency heart rate variability in mice as a novel biomarker for noradrenergic dynamics in the sleeping brain.

    3. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examined whether infraslow fluctuations in noradrenaline and in heart rate are coupled and how they are affected by sleep transitions. The authors used the fluorescent NA biosensor GRAB-NE2m in the medial prefrontal cortex of mice to record extracellular NA while also recording EEG and EMG during sleep-wake episodes. They also analyzed previously published human data to reproduce relationships they found between sigma power and RR intervals in mice.

      Strengths:

      This is an impressive study with significant strengths, as it involves a rich set of data that includes not only observations of associations between heart rate and noradrenergic dynamics but also optogenetic manipulation of the locus coeruleus. Human data is presented to show parallels in the association between sigma power during sleep and phasic heart-rate bursts.

      We thank Reviewer #1 for their thoughtful, detailed, and constructive evaluation of our manuscript. We appreciate their recognition of the strengths of the study, particularly the integration of noradrenergic recordings, optogenetic manipulation, and cross-species analyses. We are especially grateful for the reviewer’s careful attention to clarity, experimental interpretation, and control comparisons. The comments have helped us sharpen the framing of our hypotheses, clarify causal claims, improve statistical reporting, and better explain our closed-loop approach and heart rate analyses. We have addressed each point in detail below and believe that the revisions substantially strengthen the manuscript.

      Weaknesses:

      (1) Language could be clearer and more precise. As detailed below, in both the introduction and the discussion, the way the hypotheses and study objectives are described could use some revision to be more precise and accurate.

      Thank you for this helpful comment. We have sharpened the description of the study objectives, hypotheses, and interpretation of the findings to better distinguish between what was directly tested, what was inferred, and what remains speculative. We revised the language throughout these sections to improve clarity, accuracy, and overall readability

      (1A) In the introduction on p. 4: The overarching question is framed as "could the peripheral autonomous systems be a read-out of the central LC-NE system and thus be a biomarker of memory consolidation and LC dysfunction?" This gives the impression that the LC function would be the main influence on peripheral autonomous systems. There are, of course, many influences on peripheral autonomous systems, so it would be advisable for the authors to be more specific here about what signal(s) in particular would be predicted to be sensitive markers of LC function.

      Thank you for this important point. We agree that heart rate reflects the integrated output of multiple autonomic mechanisms and should not be interpreted as being exclusively driven by LC activity. Cardiac dynamics arise from the balance between sympathetic and parasympathetic influences, which themselves are regulated by several central and peripheral systems. In addition, recent work shows that the infraslow oscillations observed during NREM sleep are not restricted to norepinephrine alone but also involve other neuromodulatory systems, including acetylcholine and serotonin (e.g., Teng et al., PNAS 2025, Kjaerby et al., iScience, 2026). Our intention was therefore not to imply that the LC is the sole driver of peripheral autonomic dynamics. We have revised our overarching questions to make them more specific: Please see new text below:

      (Introduction, page 4/5). “The sympathetic and parasympathetic autonomic nervous system are involved in HRV, which is conventionally analyzed across three primary frequency bands: high frequency (HF) HRV, low frequency (LF) HRV, and very low frequency (VLF) HRV (Berntson et al., 1997). Due to the frequency overlap with VLF HRV, we wondered if central infraslow NE dynamics could be linked to this poorly understood HRV indicator. Furthermore, are infraslow NE fluctuations directly reflected by HRV under different physiological states or does LC–HR coupling scale differently with LC output? Specifically, if infraslow NE oscillations display faster frequencies - as occurs during sleep fragmentation - will cardiac dynamics exhibit corresponding changes? Conversely, given that stronger infraslow NE dynamics correlate with memory consolidation through their regulation of sleep spindles, could the peripheral autonomic signatures provide an accessible cross-species biomarker of spindle-dependent memory consolidation? Addressing these questions could help bridge mechanistic insights into LC-mediated sleep regulation with established HRV metrics used in human physiology.”

      (1B) In the discussion on p. 12: "In this study, we leveraged real-time measurements of mPFC NE levels and HR measurements from EMG recordings in mice to investigate the causal link between the two variables with high temporal resolution in freely moving sleeping mice, with similar inspection in humans." To test the causal link between mPFC NA levels and HR measures, the study would manipulate NA levels just in the mPFC and not elsewhere in the brain. However, in this study, the manipulation occurred in the LC, and so there would be broad cortical changes in NA levels. Thus, it could be that LC activity causes HR changes via a non-PFC pathway.

      We thank the reviewer for this important comment. Indeed, mPFC NE is merely a readout of LC activations and we expect that NE in other brain regions would show the same patterns. Indeed, mPFC NE is not expected to provide any causal link to heart rate. We have revised added a sentence to the results section and also changed the initial summary part of the discussion to reflect this better.

      (Results, page 6). “mPFC was selected as a representative cortical readout of LC-mediated norepinephrine dynamics, as infraslow NE fluctuations are coordinated across widespread brain regions.”

      (Discussion, page 14/15). “Variability in HR is a non-invasive biomarker of autonomic nervous system function and is frequently disrupted in ageing and Alzheimer’s disease. Here, by combining real-time measurements of mPFC NE dynamics with simultaneous HR recordings in freely sleeping mice, we demonstrate that HR closely tracks the infraslow phasic activity of the LC–NE system.”

      (2) Comparisons with the control condition need further development.

      (2A) While the authors did include a key YFP control condition, in the main text no direct statistical comparison between the closed-loop optogenetic stimulation (ChR2) condition and the YFP control condition was reported. (It was reported in Supplementary Figure 2c-d.) Instead, in the main text, the authors only reported that the effects of stimulation were significant in the closed-loop condition and not in the control. However, that is not the same as demonstrating that the two conditions significantly differed from each other, and it is the direct test that is important for the conclusions, so it seems important to include this result in the main presentation.

      We thank the reviewer for this important point and agree that direct statistical comparisons between ChR2 and YFP conditions are important for interpretation. These comparisons were performed and are shown in Supplementary Figure 2c–d, but we acknowledge that this was not sufficiently emphasized in the main text. We are now more clearly referring to this comparison in Result section:

      (Results, page 9). “The magnitude of pre-stimulation NE descent and post-stimulation NE ascent was reduced as the thresholds increased, indicating less pronounced NE dynamics as LC stimulation became more frequent (Fig. 2e, for direct comparison with YFP control, see Suppl. Fig. 2c-d).”

      Our rationale for prioritizing the within-animal threshold comparisons in the main figure was that the central experimental question concerned how progressive shifts in infraslow NE oscillatory frequency influence the NE–HR relationship. Because variability in viral expression levels (both NE sensor expression and LC opsin expression) introduces substantial between-animal variability, we considered within-animal comparisons across threshold conditions to provide the most informative representation of how changes in LC-driven NE dynamics alter cardiac responses. That said, we agree that highlighting the direct ChR2 versus YFP comparison is important for the overall interpretation. As mentioned, the reference to Suppl. Fig. 2c– d, where the between-group analyses are visualized are now clearly referred to.

      (2B) In addition, the authors should address the issue that the pre-stimulation NE was consistently significantly lower in the YFP condition than in the ChR2 condition (see Supplementary Figure 2c), which is a potential confound.

      We thank the reviewer for bringing up this important point. We agree that differences in stimulation timing between ChR2 and YFP animals could complicate interpretation in a closed-loop design and appreciate the opportunity to clarify this aspect of the experiment.

      In the ChR2 condition, animals were exposed to repeated optogenetic LC activation designed to mimic progressively faster infraslow NE dynamics. Such repeated stimulation is expected to produce a gradual elevation in tonic NE levels across the recording session, which explains the higher pre-stimulation baseline relative to YFP controls. We acknowledge that elevated tonic NE levels could introduce additional physiological effects. For example, higher NE tone would be expected to increase α2mediated autoinhibitory feedback on LC neurons and presynaptic NE release. Within the LC itself, we expect that optogenetic stimulation would largely override such effects due to the strong Na+-mediated depolarization induced by ChR2 activation. However, NE release in downstream regions such as the mPFC may be influenced to some extent by elevated tonic noradrenergic tone. Importantly, such feedback mechanisms would likely also occur under physiological conditions characterized by elevated LC activity, such as stress or sleep fragmentation. Because the goal of our stimulation paradigm was to model progressively faster infraslow NE dynamics under physiologically relevant conditions, we believe this feature of the manipulation may in fact increase the translational relevance of the model. We specifically address this elevation in Figure 3, where we show that very rapid stimulation regimes are accompanied by signs of compensatory cardiovascular regulation, likely reflecting baroreceptor-mediated responses to sustained increases in heart rate.

      We agree that the precise contribution of elevated tonic NE to the overall manipulation cannot be fully disentangled in the present study. We therefore avoid overinterpreting these effects and have instead added text to the manuscript acknowledging this consideration without extensive speculation.

      (Results, page 9) “Since pre-stimulation NE baseline levels progressively became higher in the ChR2 condition compared with YFP controls (Fig. 2c+e, Supplementary Fig. 3a-f), it demonstrates that they arise from stimulation-dependent modulation of noradrenergic tone rather than nonspecific signal drift. As a result, the pre-stimulation state at higher thresholds differed between ChR2 and control conditions, which should be considered when interpreting the immediate effects of laser stimulation across thresholds.”

      (2C) Direct comparison of the strengths of correlations shown in Figure 2h vs. Supplementary Figure 2f should be included. Currently, we see relatively weak correlations in both ChR2 and YFP conditions, and it is not clear if the relationships differ in the control. It seems they are still present in the control condition but weaker which would contradict the apparently broad claim on p. 7 that "No such effects were present in the control condition" (it is not entirely clear whether this claim refers to all effects discussed in the figure or just a subset - this language should be clarified).

      We thank the reviewer for this important comment and agree that the original wording could be interpreted as implying a complete absence of an NE–RR relationship in the YFP condition. To address this concern, we directly compared the strength of the NE– RR relationship between ChR2 and YFP animals. Please see Supplementary Figure 3h.

      Using a linear mixed-effects model that accounted for repeated measurements within animals, we found that the slope of the NE–RR relationship was significantly steeper in ChR2 animals than in YFP controls (all thresholds: slope difference = 1.915, p < 0.0001; thresholds −15, −10, and −5 only: slope difference = 2.367, p < 0.0001). Consistent with this result, comparison of Pearson correlations using Fisher's r-to-z transformation also indicated significantly stronger coupling in ChR2 animals than in YFP controls (see Statistics in the Supplementary File).

      These analyses demonstrate that an inverse relationship between NE and RR is present under physiological conditions in YFP animals, but that optogenetic LC activation substantially strengthens this coupling. We have revised the corresponding text to clarify this:

      (Results, Page 9): “An inverse NE–RR relationship was present in both ChR2 (Fig. 2h) and YFP (Suppl. Fig. 3f) groups but was significantly stronger in ChR2 animals than in YFP controls (Suppl. Fig. 3h).”

      (2D) Did the YFP controls vs. ChR2 animals show any differences in the number of NA states that triggered stimulation in the closed-loop system? With ChR2 animals, stimulation changes NA, which could change future triggering. In YFP animals, nothing changes NA (other than natural fluctuations), so the dynamics of stimulation timing could diverge between groups in a way that complicates interpretation. Specifically, if ChR2 stimulation raises NA and prevents future threshold crossings, ChR2 animals may end up receiving fewer subsequent stimulations than YFP animals (or a different temporal clustering). If the number or pattern of stimulation differed in two groups, it would be important to have a yoked control where matched animals get the same stimulation pattern but not triggered by their own NA.

      We thank the reviewer for this important point. We agree that differences in stimulation timing between ChR2 and YFP animals could complicate interpretation in a closed-loop design.

      Importantly, stimulation triggering was based on relative declines in NE fluorescence calculated against a rolling 2-minute baseline, rather than absolute NE levels. Thus, as tonic NE levels gradually increased in ChR2 animals, the threshold adapted accordingly, reducing the likelihood that elevated baseline NE alone would prevent future triggering. Instead, stimulation continued to occur when NE declined relative to the recent baseline, thereby preserving the infraslow closed-loop structure.

      The number of stimulation events across thresholds is already reported in the manuscript (Methods, p. 30 and corresponding figure legends), but we have now indicated more clearly in the result section where to find the information:

      (Results, Page 9). “Mean traces of NE and RR were aligned to LC stimulation onset (Fig. 2c-d, for number of laser stimulations see Fig. 2 legend or Methods).”

      (Methods, Page 31). “For the LC activation-related analysis, 108 events were found for Threshold -15 (11 of these being YFP), 260 events for Threshold -10 (55 of these YFP), 777 events for Threshold -5 (296 of these YFP), 1,444 events for Threshold 0 (510 of these YFP), and 1,148 events for Threshold 5 (377 of these YFP) across ten animals (four being YFP).”

      While the total number of events was lower in YFP animals, this is expected in part due to the smaller group size (4 YFP vs. 6 ChR2 animals included in this analysis). Furthermore, because ChR2 stimulation increased NE levels by design, more pronounced subsequent declines in NE may have modestly facilitated additional threshold crossings.

      Nevertheless, the overall temporal structure of the stimulation paradigm remained comparable across groups, and YFP animals underwent the same closed-loop stimulation protocol. Our primary comparison was mainly based on shifts in within-animal NE oscillatory frequency and how this change would impact the connection to HR. Thus, we believe the present control condition appropriately addresses the central question of whether optogenetic LC activation are able to conduct a continuum of NE oscillations.

      (3) Some more discussion/explanation of the rationale for the closed-loop approach and how it influences how we should interpret the results could be useful. For instance, currently, it is not clear whether LC stimulation needs to be timed after an NA dip to yield the effects seen.

      We thank the reviewer for pointing this out. The rationale for the closed-loop LC stimulation approach was to test whether the relationship between infraslow LC–NE dynamics and heart rate is maintained only under physiological infraslow conditions or whether it breaks down when the rhythm becomes progressively faster, as occurs during sleep fragmentation and other high-arousal states. Specifically, we asked whether heart rate continues to track LC–NE fluctuations as the infraslow rhythm shifts toward higher frequencies, thereby assessing its utility as a potential biomarker of disrupted restorative sleep.

      To address this while preserving the intrinsic temporal structure of infraslow LC activity, we implemented a closed-loop strategy in which stimulations were triggered following defined declines in the NE signal, using a rolling preceding 2-minute window as baseline. This allowed LC activation to occur during the descending phase of the endogenous infraslow cycle, maintaining its physiological phase structure while systematically increasing its effective frequency. By progressively relaxing the decline threshold, stimulations were triggered earlier in the cycle, thereby compressing the infraslow period in a controlled manner.

      Importantly, the intention was not to test whether LC stimulation specifically needs to occur after an NE dip to elicit the observed effects. Rather, triggering stimulation during the decay phase provided a way to accelerate the infraslow rhythm without disrupting sleep through indiscriminate stimulation. This enabled us to examine whether the coupling between LC–NE dynamics and heart rate remains stable under increasingly rapid infraslow regimes. Our results indicate that this relationship weakens at higher infraslow stimulation frequencies, suggesting that heart rate reliably reflects physiological LC–NE oscillations but becomes less tightly coupled when the rhythm is compressed beyond its normal range.

      We have clarified this rationale in the revised manuscript by adding the below section in the result section.

      (Result, page 8/9). “This approach enabled controlled compression of the infraslow NE cycle by triggering LC activation during the descending phase of the endogenous NE signal, thereby increasing the effective oscillatory frequency while preserving the temporal structure of physiological LC–NE dynamics. This strategy allowed us to test whether heart-rate responses continue to track LC-driven NE fluctuations as the infraslow rhythm becomes progressively faster.”

      (4) The section on heart rate decelerations is hard to follow. In particular, I was not sure how to interpret Figure 3f-j. For Figure 3f, what does the middle line represent? The laser onset or the max RR value after laser onset? What is the baseline that is used to correct the values to obtain amplitudes? If it is the whole period before the maximal RR value or the laser onset, wouldn't baseline values differ significantly across conditions and so potentially account for differences seen between conditions in the reported HR decelerations? Larger HR decelerations may be seen in conditions with higher HR simply as a regression to the mean phenomenon.

      We thank the reviewer for this feedback and agree that additional clarification of Figure 3f–j is warranted.

      For Figure 3f, the central line represents the peak RR value (maximal heart-rate deceleration) identified within the 2–7 s window following laser onset, rather than the laser onset itself. We realize this was not sufficiently clear and have revised the figure and corresponding Results text to clarify this point.

      Regarding baseline correction, RR amplitudes were calculated as the difference between the RR at peak deceleration and the mean RR during the 8–10 s period preceding the RR peak, as described in the manuscript. Thus, the baseline was defined locally for each event and was not based on the entire pre-laser period or stimulation onset. We chose this approach to account for shifts in baseline heart rate across conditions and to capture the relative magnitude of the deceleration response rather than absolute RR values.

      We appreciate the reviewer’s point regarding potential regression-to-the-mean effects, particularly in conditions with higher baseline heart rates. This is an important consideration. However, because the amplitude measure was baseline-corrected on an event-by-event basis, we believe the reported differences are unlikely to be explained solely by higher pre-stimulation heart rate. At the same time, our findings clearly show that elevated baseline heart rate influence the dynamic range of deceleration responses. Our findings show that under physiological conditions with elevated HR, larger heart-rate fluctuations would also contribute to increased HRV, which is often interpreted positively, despite potentially reflecting fragmented or dysregulated sleep states in this context.

      To improve readability, we have revised the figure and associated text.

      (Results, page 10). “To quantify the HR decelerations that happened after LC activation, we took the maximal RR value (so slowest HR) 2-7 s after LC stimulation and baseline corrected the value to the mean RR during the 8–10 s period preceding the RR peak to obtain their amplitude.”

      (5) The findings regarding LC suppression could be further clarified.

      (5A) Page 8: "observed a response in NE decline" - please be more precise. Did NE decline more or less?

      We thank the reviewer for this suggestion and agree that the original wording was imprecise. To clarify the direction and nature of the response, we have revised the text to state:

      (Results, page 11). “...we observed a gradual NE decline sustained throughout the laser period that was not observed in the YFP condition…”

      (5B) It would be helpful to also show the correlation between NE and RR in the control (YFP) condition and whether there were any differences between YFP and Arch conditions (Figure 4e).

      We thank the reviewer for this suggestion. We have now added the corresponding YFP correlation to Figure 4e. In the YFP group, the relationship between NE and RR showed a similar negative trend but did not reach statistical significance (p = 0.053). To directly assess whether the NE–RR relationship differed between Arch and YFP animals, we performed both a linear mixed-effects analysis and a Fisher r-to-z comparison.

      Neither analysis revealed a significant difference between groups. The linear mixed-effects model showed no significant Group × NE interaction (slope difference = 0.994, p = 0.51), indicating that the NE–RR coupling was not altered by LC suppression. Similarly, Fisher's r-to-z comparison found no significant difference between the correlations (p = 0.42).

      We believe this result is consistent with the relatively modest nature of the LC suppression paradigm. While Arch stimulation produced a clear reduction in NE levels, it did not induce a large shift in the overall NE–RR relationship. Instead, the data suggest that heart-rate responses remain coupled to noradrenergic fluctuations under both physiological conditions and during mild LC suppression. We have added the YFP data and clarified this interpretation in the revised manuscript.

      (Results, page 12): “A similar negative relationship was observed in YFP controls (Fig. 4e), and the strength of the NE–RR association did not differ significantly between Arch and YFP animals (Suppl. File, Statistics), suggesting that LC suppression did not substantially alter the underlying coupling between these measures.”

      (5C) This sentence took me multiple readings to understand - it would be helpful to rewrite to make it clearer: "indicating that, while HR generally did not respond strongly to LC suppression, the variability in RR responses was dependent on NE changes to the suppression (Figure 4e)."

      We agree with the reviewer that this phrasing is hard to understand and we have optimized for better clarity. Please see new version below:

      (Results, page 11/12). “Notably, despite the absence of a robust group-level HR effect, NE and RR responses remained negatively correlated across trials (Fig. 4e), indicating that HR dynamics continued to track the magnitude of noradrenergic suppression at the individual-response level”.

      (5D) The two colors in Figure 4 are similar and hard to distinguish.

      We agree that the colors are hard to separate and have altered them to make them easier to separate.

      (5E) The correlations shown in Figure 4j seem to be driven by just two of the cases. Are the effects significant when outliers are removed?

      We thank the reviewer for raising this point. To assess whether the observed correlation in Figure 4j was disproportionately driven by a small number of data points, we performed a formal outlier analysis using the ROUT method (Q = 1%). This analysis did not identify any statistical outliers in the dataset. Therefore, we did not have an objective basis for excluding any observations from the analysis.

      (5F) Page 10: Were there any differences in memory performance between the Arch and YFP conditions?

      We thank the reviewer for this question. The memory experiments were based on a previously published dataset (Kjaerby, Andersen et al., 2022), in which the primary objective was to assess the effect of LC suppression on sleep spindle dynamics and memory consolidation. In the present study, we performed an additional analysis by extracting heart-rate (RR) measures from these recordings to evaluate whether cardiac responses could serve as a biomarker of LC-mediated noradrenergic regulation.

      However, due to technical limitations (EMG recording often suffers from noise) in extracting reliable RR signals from all animals in this dataset, the number of subjects available for this secondary analysis was reduced. All these considerations are described in Methods/Mice. As a result, we were not sufficiently powered to perform a direct statistical comparison of memory performance between Arch and YFP groups based on RR measures alone. Instead, we examined whether RR responses to LC suppression predicted behavioral performance across animals. When pooling Arch and YFP conditions, we observed a correlation between the magnitude of the RR response and subsequent memory performance, suggesting that heart-rate dynamics reflect noradrenergic modulation relevant for memory consolidation.

      To avoid overinterpretation, we have therefore limited our conclusions to reporting this association rather than making direct group-level comparisons between Arch and YFP animals, and we have clarified this point in the revised manuscript.

      (Results, page 13): “Interestingly, across pooled Arch and YFP animals, larger RR increases following LC suppression were associated with better subsequent memory performance. Additionally, RR and NE responses to LC suppression were negatively correlated indicating that animals showing stronger NE reductions also exhibited larger RR changes. Together, these findings suggest that heart-rate dynamics covary with noradrenergic responses during sleep and may reflect physiological processes relevant for sleep-dependent memory consolidation.”

      (5G) Page 10: "We found a correlation between RR responses to LC suppression and sigma power, suggesting that a stronger HR reduction response is linked to higher spindle power." It should be noted in the text that the correlation was not specific to sigma (it was also seen for theta and beta, Figure 4i).

      We agree with reviewer that this should be highlighted. We have changed the sentence:

      (Results, page 12). “Furthermore, we found a correlation between RR responses to LC suppression and sigma power, suggesting that a stronger HR reduction response is linked to higher spindle power; similar correlations were also observed in the theta and beta frequency ranges (Fig. 4h-i).”

      (6) It is not clear which of the sigma power and RR interval findings do/do not exactly line up between the mice and humans. It could be helpful to have a table comparing them. For instance, was the finding in humans that pre-HRB sigma power was positively associated with slowing in heart rate after the HRB also seen in mice? Was there evidence in mice (as seen in the human sample) that sleep-dependent memory improvement was associated with pre-HRB sigma power?

      We thank the reviewer for this thoughtful comment and agree that the cross-species comparisons could be communicated more clearly. Our intention was not to imply exact one-to-one correspondence between all mouse and human findings, but rather to examine whether central–autonomic coupling surrounding phasic heart-rate events shows conserved features across species while acknowledging species-specific physiological differences.

      Importantly, the mouse and human analyses were designed to address related but not identical questions. In mice, we leveraged optogenetic LC suppression to probe a more causal relationship between noradrenergic activity, heart-rate slowing, spindle-related dynamics, and memory consolidation. Previous work using this dataset demonstrated that 2-minute LC suppression robustly enhances spindle density and that spindle enhancement correlates with improved memory performance. In the present study, we therefore asked whether heart-rate slowing covaries with this LC-mediated spindle/memory relationship, supporting HR as a potential biomarker of these restorative processes.

      By contrast, causal manipulation of LC activity is not feasible in humans. Instead, we focused on naturally occurring HRBs as putative downstream signatures of phasic LC– NE activity, motivated by our mouse findings that NE increases precede HR accelerations. We observed conserved coupling between sigma activity and HR dynamics across species, although the temporal profile differed, with sigma activity occurring closer to the HRB in humans than in mice. These temporal differences may reflect species-specific differences in cardiac and sleep physiology.

      Regarding the reviewer’s specific questions, the positive relationship between pre-HRB sigma power and post-HRB heart-rate slowing was tested in humans, where this metric showed the strongest relationship to behavioral outcome. We did not directly test the same measure in mice because, unlike humans, HR recovery following HRBs did not show a pronounced baseline shift (Fig. 5b), limiting the interpretability of this comparison. Similarly, we did not directly correlate pre-HRB sigma power with memory performance in mice, as the more causal LC suppression paradigm already demonstrated a spindle–memory relationship in this species and was the focus of our mechanistic analysis.

      To reduce confusion, we have revised the Results conclusion to make a clearer overview:

      (Results, page 14): “In conclusion, mice and humans displayed evidence of conserved autonomic-central coupling, reflected in coordinated HR and sigma power dynamics surrounding phasic cardiac events.

      However, the temporal relationship between these events differed across species, likely reflecting differences in sleep and cardiovascular physiology. In mice, causal manipulation of LC activity demonstrated that heart-rate dynamics covary with LC-mediated noradrenergic and spindle-related processes linked to memory consolidation. In humans, sigma power preceding HR bursts was associated with both post-HRB heart-rate slowing and sleep-dependent memory improvement, suggesting that autonomic–central coupling surrounding HR events may provide a translational marker of restorative sleep processes (Fig. 5n).”

      (7) Page 18: It is not clear if the sex of mice was balanced across controls and optogenetics groups.

      We thank the reviewer for this important comment and agree that the sex distribution should be reported more clearly. These experiments relied on the availability of animals from the heterozygous TH-Cre transgenic line, and given the relatively small cohort sizes, perfect balancing across sex and experimental groups was not always feasible.

      For the LC activation experiments, the sex distribution was: YFP: 2 male / 2 female; ChR2: 4 male / 2 female. For the LC suppression experiments, the distribution was: YFP: 4 female; Arch: 3 female / 1 male.

      Although the groups were not perfectly sex balanced, we had no strong reason to expect robust sex-dependent differences in the physiological effects of these optogenetic manipulations, particularly given the relatively strong and acute nature of the intervention. At the same time, we acknowledge that the present study was not powered to assess sex as a biological variable, and subtle sex-dependent effects therefore cannot be excluded. To improve transparency, we have now clarified the sex distribution in the Methods/Mice section.

      Reviewer #2 (Public review):

      Summary:

      The major part of this study reproduces previously published findings in both mice and humans and provides incremental analyses on these findings. In essence, the work reaffirms the presence of coordinated infraslow fluctuations in sigma power and heart rate during NREM sleep. It further confirms previous findings that coordination depends on noradrenaline-releasing neurons in the locus coeruleus. Also supporting previously published work in mice and humans, the authors describe a link between the strength of these infraslow fluctuations and memory consolidation in mice and humans.

      Strengths:

      The authors successfully replicate key previously reported phenomena across both mice and humans. Confirmatory studies and demonstrations of reproducibility are essential for progress in neuroscience. To maximize their value, such studies should clearly acknowledge their confirmatory nature and carefully situate what, in their view, are novel results, going beyond existing literature.

      Weaknesses:

      The authors' interpretation of their data needs to be revised. Many of their claims regarding the mechanistic basis of their findings and the predictive value of their correlative datasets are not supported by the available evidence.

      In the present manuscript, several citations of literature on the work they reproduce lack precision or completeness, which reduces transparency and obscures how the reported findings relate to previously established results.

      We thank Reviewer 2 for the thoughtful comment regarding positioning of our findings relative to the literature, and caution in mechanistic interpretation. In response, we have revised the Introduction, Results, and Discussion to more clearly acknowledge foundational studies in this area and to better clarify how the present work extends beyond them.

      We agree that prior work has demonstrated infraslow coupling between sigma activity, norepinephrine (NE) dynamics, and heart rate (HR), and has established a role for the locus coeruleus (LC) in coordinating these oscillations. However, cardiac measures in these studies were typically treated as secondary observations rather than as primary experimental targets. A central goal of the present study was therefore to provide a systematic and mechanistically grounded characterization of NE-mediated HR dynamics during sleep across multiple timescales, including infraslow oscillations, sleep–wake transitions, and causal manipulations of LC activity.

      Importantly, we also aimed to relate infraslow HR fluctuations to the very-low-frequency (VLF) component of heart rate variability (HRV), which remains comparatively under-characterized and mechanistically unresolved in the clinical HRV literature. By linking LC activity, NE dynamics, and HR fluctuations across behavioral states, our findings provide a biologically grounded framework that may help explain this component of HRV.

      A second major objective of the study was translational. Because direct LC recordings are not feasible in humans, we asked whether cardiac dynamics alone could reflect the infraslow, memory-consolidating potential of sleep and thus serve as a noninvasive biomarker. By directly manipulating LC activity and demonstrating corresponding changes in HR dynamics, our results strengthen the mechanistic rationale for using HRV—particularly its VLF component—as an accessible proxy of LC-dependent sleep physiology.

      We therefore respectfully disagree with the suggestion that the present study does not provide novel insight. Rather, the revised manuscript now more clearly emphasizes that our contribution lies in (i) systematically characterizing NE-dependent HR dynamics across sleep states, (ii) linking these dynamics to the poorly understood VLF component of HRV, and (iii) establishing a causal and translational framework for using cardiac measures as markers of LC-mediated sleep processes.

      We hope the reviewer finds that the revised Introduction and Discussion better highlight both the existing literature and the specific advances provided by the present work.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I have been convinced by discussions about the replicability crisis that it should be a standard practice to share data upon publication in a publicly accessible online repository such as Open Science Framework or OpenNEURO. Making data available can increase the impact of the research study. Simply stating "data available upon request" as done in the current draft is not sufficient as, unfortunately, when data are not shared upon publication in a public repository, it can be impossible to gain access by request from the researchers - the majority of requests from other researchers to obtain data are not complied with (e.g., Vanpaemel, Vermorgen, Deriemaecker, & Storms, 2015; Wicherts, Bakker, & Molenaar, 2011).

      We fully agree with the reviewer about the importance of data sharing for transparency, reproducibility, and maximizing the impact of research. Consistent with these principles, we will make the human dataset publicly available on the Open Science Framework upon publication and have added the link to this repository in the manuscript under Data Availability (https://osf.io/g6emj/). With respect to the mouse dataset, we respectfully note that this dataset is currently the subject of multiple planned and ongoing analyses that extend beyond the scope of the present manuscript. Releasing these data publicly at this stage could compromise these efforts and lead to potential misinterpretation prior to completion of the full analytic pipeline. For this reason, we believe that it would not be appropriate to share the mouse data in a public repository at this time. However, we remain committed to transparency and will make the mouse data available upon reasonable request during this period, with the intention of publicly releasing the dataset once the planned analyses are complete.

      (1) p. 4: There is some orphan text at the top of the page ("marker of Alzheimer's disease. Furthermore, maintaining LC neural density prevents neurodegeneration (13). Given the reported reduction in HRV in aging and Alzheimer's disease, suppressed").

      Thank you. We have removed the orphan text.

      (2) p. 22: What does NFR refer to?

      Novel-to-familiar ratio. We have removed the abbreviation from the main text. It is now only used in the figure.

      (3) I might have missed this information, but it was not clear to me how the epochs to be analyzed were selected, and for the mice, how much wake vs. sleep they included.

      We thank the reviewer for this comment and apologize that the epoch selection criteria were not sufficiently clear. They were mentioned under Methods. For the optogenetic analyses in mice, epochs were selected based on NREM sleep including microarousals to ensure that physiological responses were evaluated within stable sleep conditions while allowing for natural brief interruptions.

      For the LC activation (ChR2) experiments, stimulation epochs were included only if NREMinclMA began at least 30 s before laser onset and continued for at least 30 s after laser onset. For the LC suppression (Arch) experiments, the criterion was similarly ≥30 s of NREMinclMA prior to laser onset, but extending ≥60 s following laser onset to accommodate the longer suppression response profile.

      Thus, analyses were intentionally restricted to sleep periods, and wakefulness was not included except where it emerged naturally as an outcome of the manipulation or transition under investigation. We have now clarified these inclusion criteria in the Methods/ Event marker selection section to improve transparency.

      (4) In the Supplement, Figure 3 appears before Figure 2.

      Thanks. We have corrected it.

      Signed, Mara Mather

      Reviewer #2 (Recommendations for the authors):

      The authors' interpretation of their data needs to be revised. Many of their claims regarding the mechanistic basis of their findings and the predictive value of their correlative datasets are not supported by the available evidence.

      There are three major directions in which this study would need to be revised.

      Part 1 - Literature citations of both mouse and human literature need to be revised, and citations placed in a manner that accurately reflects what has been previously done. In detail:

      (1.1) The statement on p. 4 regarding "the extent to which phasic infraslow NE fluctuations ...is not well understood" disregards a previous publication in which closed-loop optogenetic stimulation of LC was already shown to regulate HR variations on the infraslow time scale (10.1016/j.cub.2021.09.041).

      We thank the reviewer for this important point. The study by Osorio-Forero et al. (2021) was already cited as ref. 4 in our original manuscript and in the discussion, we specifically highlighted this study: ‘These findings build on Osorio-Forero et al. (35) as well as other studies’; however, we agree that our wording did not sufficiently emphasize its key finding that closed-loop optogenetic manipulation of LC activity can coordinate infraslow heart-rate fluctuations and spindle clustering during NREM sleep.

      We have revised the relevant paragraph in the introduction to explicitly acknowledge that ref. 4 demonstrated causal coordination between LC activity, sleep spindle dynamics, and heart-rate fluctuations. We have also refined our statement of the knowledge gap to clarify which novel questions our study addresses. We believe these revisions more accurately position our work within the existing literature and clearly distinguish our contributions from prior studies.

      We have updated the introduction in several places to address these comments:

      (Introduction, page 4). “It has previously been reported that HR fluctuates at similar infraslow frequencies as NE fluctuations and sigma power in mice (Lecci et al., 2017; Osorio-Forero et al., 2021) and that phasic HR fluctuations correlate with infraslow changes in pupil diameter (Carro-Domínguez et al., 2025), a proxy for changes in NE levels (Murphy et al., 2014; Reimer et al., 2016). Importantly, optogenetic manipulation of the LC has demonstrated that infraslow LC activity coordinates sleep spindle clustering and heart rate fluctuations during NREM sleep (Osorio-Forero et al., 2021). Together, these findings support a functional coupling between the central LC–NE system and peripheral cardiac dynamics, further supported by findings that HR increases accompany MAs during NREM sleep (Carro-Domínguez et al., 2025;

      Osorio-Forero et al., 2025).”

      (Introduction, page 4/5). “Due to the frequency overlap with VLF HRV, we wondered if central infraslow NE dynamics could be linked to this poorly understood HRV indicator. Furthermore, are infraslow NE fluctuations directly reflected by heart-rate variability under different physiological states or does LC–HR coupling scales differently with LC output? Specifically, if infraslow NE oscillations display faster frequencies - as occurs during sleep fragmentation - will cardiac dynamics exhibit corresponding changes? Conversely, given that stronger infraslow NE dynamics correlate with memory consolidation through their regulation of sleep spindles, could the peripheral autonomic signatures provide an accessible cross-species biomarker of spindle-dependent memory consolidation? Addressing these questions could help bridge mechanistic insights into LC-mediated sleep regulation with established HRV metrics used in human physiology.”

      (1.2) In this same published paper, optogenetic stimulation of LC was already used to "determine the causal relationship..." (p.6). However, the authors do not cite these data.

      We thank the reviewer for this comment. As noted in our response to Comment 1.1, ref. 4 was already cited in the Introduction, and we have now revised that section to more explicitly emphasize this paper. In addition, we have modified the wording in the Results section.

      (Results, page 8). “After finding the inverse correlation between NE and RR in natural sleep transitions, we next sought to further characterize the causal influence of LC activity on HR dynamics, using a closed-loop optogenetic approach to modulate NE oscillatory frequency during NREM sleep.”

      (1.3) The authors' speculation about LC-induced sympathetic and parasympathetic actions is premature: this study does not provide pharmacological experiments in this direction. However, two published studies implied a parasympathetic mechanism linked to infraslow fluctuations of LC activity (10.1016/j.cub.2021.09.041, 10.1016/j.cub.2017.12.049). These findings should be appropriately cited. While HR decelerations may reflect compensatory autonomic responses, there is no direct evidence in the present study that these effects are sympathetically mediated. An alternative, and equally plausible, interpretation is enhanced parasympathetic activity. This distinction is particularly important given that the observed increases in mean HR during LC stimulation cannot distinguish between reduced parasympathetic activity and increased sympathetic drive.

      We thank the reviewer for raising this important point and agree that the current study does not provide direct mechanistic evidence to disentangle sympathetic versus parasympathetic contributions to LC-mediated heart rate regulation. This was not the intention of our study; rather, the relevant discussion section was meant to provide mechanistic interpretations and hypotheses based on the observed physiology. To avoid overstating our conclusions, we have revised the wording to more clearly emphasize the speculative nature of these interpretations. Furthermore, we have added the suggested references demonstrating that muscarinic blockade reduces heart rate and pupil fluctuations during sleep, which support the possibility of a parasympathetic contribution to infraslow LC-related dynamics.

      (Discussion, page 16/17). “Prior findings demonstrate that LC is linked to the autonomic nervous system. Stimulation of LC projections decrease parasympathetic cardiac vagal activity (Wang et al., 2014) and also influences sympathetic output through direct projections to the preganglionic cells in the sympathetic nervous system (Karemaker, 2017; Nygren and Olson, 1977; Samuels and Szabadi, 2008). While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014), there is a clear parasympathetic component (Taylor et al., 1998). This combined with the ability of pharmacological blockage of the parasympathetic system to block infraslow oscillations of HR (Osorio-Forero et al., 2021) and pupil diameter during sleep (Yüzgeç et al., 2018), led us to expect a slowing of HR during LC suppression due to parasympathetic disinhibition.”

      (1.4) It is not clear why prefrontal NE signals were associated with HR fluctuations. Literature evidence indicates other brain areas that are functionally more directly linked to autonomous fluctuations.

      We thank the reviewer for raising this point. We do not speculate that mPFC NE activity is causally linked to heart rate fluctuations or that the mPFC directly mediates the observed autonomic dynamics. Instead, mPFC NE signaling was used as an experimentally accessible readout of LC activity. Importantly, accumulating evidence suggests that infraslow NE oscillations are globally coordinated phenomena that are expressed across multiple brain regions during sleep, making mPFC NE a valid proxy for LC-driven neuromodulatory state dynamics. We have clarified this in the result section:

      (Results, page 6). “mPFC was selected as a cortical readout of LC-mediated norepinephrine dynamics, as infraslow NE fluctuations are coordinated across widespread brain regions.”

      (1.5.a) The lack of effect of Arch-inhibition of LC on infraslow NE signals is concerning. Prior work showed that bilateral LC inhibition does affect NE signals and also infraslow sigma power fluctuations (10.1016/j.cub.2021.09.041, 10.1038/s41593-024-01822-0). This discrepancy should be explicitly acknowledged and discussed on p.9.

      We thank the reviewer for highlighting this important point. To clarify, we did observe a robust effect of LC inhibition on noradrenergic signaling, with clear suppression of the NE signal following Arch-mediated LC inhibition (Fig. 4c), aligned to laser onset. In addition, LC suppression increased neuronal synchronization, including enhanced sigma power relative to YFP controls (Fig. 4h), consistent with previous reports showing that reduced LC activity promotes synchronized sleep-related oscillations. Thus, we do not interpret our findings as indicating an absence of LC suppression effects.

      The apparent discrepancy relates specifically to the absence of a statistically significant group-level shift in infraslow NE or HRV power during NREM sleep, rather than the efficacy of the manipulation itself. We note that our inhibition paradigm was intentionally mild, consisting of repeated 2-minute suppression periods separated by 4-minute intervals, and was designed to introduce subtle shifts within physiological ranges rather than globally reorganize infraslow sleep structure. Accordingly, we consider the immediate NE and heart-rate responses to LC inhibition to be the most sensitive physiological readouts of LC-mediated regulation in this context.

      We also note that the studies cited by the reviewer used different suppression paradigms, including more frequent manipulations. Thus, while our suppression scheme did not result in detectable infraslow reorganization, we do not believe they reflect an ineffective LC suppression as we demonstrated a clear NE reduction in response to time-locked LC suppression.

      We have added a sentence to the result section to explain the lack of effect infraslow power:

      (Results, page 12). “This likely reflects the relatively mild and intermittent LC suppression paradigm, which was designed to remain within physiological ranges and therefore did not globally reorganize infraslow sleep dynamics.”

      (1.5.b) Additionally, the authors state in the Discussion that they "find no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection is more strongly associated with sympathetic activity rather than parasympathetic inhibition". However, the results primarily demonstrate an absence of modulation in mean HR, while preserving a significant relationship between RR intervals and stimulation. This indicates a modulation of heart rate variability, even in the absence of changes in average HR. Notably, such variability-related effects may fall outside the VLF range and could instead involve higher-frequency components.

      We thank the reviewer for this important clarification. We agree that our original wording may have conflated the absence of modulation in mean HR with the absence of autonomic modulation more generally. To address this point, we revised the Discussion to emphasize that LC activity may influence VLF HRV through sympathetic activation and/or indirect modulation of cardiac vagal activity, even in the absence of robust changes in average HR. We have softened our previous interpretation that the LC-HR relationship is primarily sympathetic in nature and instead discuss a more nuanced interaction between sympathetic and parasympathetic influences on HRV dynamics.

      (Discussion, page 16/17). “Prior findings demonstrate that LC is linked to the autonomic nervous system. Stimulation of LC projections decrease parasympathetic cardiac vagal activity (Wang et al., 2014) direct projections to the preganglionic cells in the sympathetic nervous system (Karemaker, 2017; Nygren and Olson, 1977; Samuels and Szabadi, 2008). While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014) there is a clear parasympathetic component (Taylor et al., 1998). This combined with the ability of pharmacological blockage of the parasympathetic system to block infraslow oscillations of HR (Osorio-Forero et al., 2021) and pupil diameter during sleep (Yüzgeç et al., 2018) led us to expect a slowing of HR during LC suppression due to parasympathetic disinhibition. Interestingly, we found no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection may also somehow be driven by sympathetic outflow. Previous research had indicated that LF power may also represent sympathetic activity, but this interpretation has been challenged due to the mixed contribution of both autonomic branches (Houle and Billman, 1999; Japundzic et al., 1990; Reyes et al., 2013). LF and HF ratio (LF/HF) were traditionally thought to reflect balance between sympathetic and parasympathetic activities (i.e., the sympatho-vagal balance), though currently considered as an oversimplification of non-linear integration of autonomic signals (Billman, 2013; Pagani et al., 1986). Our findings implicate the VLF may offer a precise marker for central arousal states, given its overlaps with infraslow phasic fluctuations of LC-NE levels. Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014). Within this framework, infraslow LC–NE dynamics may represent one central contributor to these slow cardiac fluctuations during sleep either directly through sympathetic activation or indirectly by inhibition of cardiac vagal activity.”

      (1.6) The relationship between sigma power fluctuations and HR is different in humans than in mice. This has been shown before (10.1126/sciadv.1602026, 10.1038/s41593025-02159-y). This work should be mentioned on p. 11.

      We already acknowledge prior studies demonstrating species differences in the relationship between sigma power fluctuations and heart rate in the discussion section. However, to accommodate the reviewer’s comment and improve clarity for the reader, we have now also added the suggested references to the Results section, where the relationship between sigma power fluctuations and HR is first discussed.

      (Results, page 14). “These temporal differences may reflect species-specific physiology differences in cardiac timescales, which has also been previously reported (Bergel et al., 2025; Carro-Domínguez et al., 2025; Lecci et al., 2017).”

      (1.7) Correlations between the strength of infraslow sigma power fluctuations and memory consolidation have been published and should be discussed (10.1126/sciadv.1602026). It is surprising to see that correlations with learning in humans are done using pre-HRB sigma peaks rather than heart rate. This is a measure that is very close to the one used by Lecci et al.; this similarity should be clearly acknowledged. The way the data are currently presented limits this manuscript's novelty, also in its translational aspect.

      We agree that the work by Lecci et al. (2017) established an important relationship between the association of infraslow sigma power fluctuations and memory consolidation, which is highly relevant to our findings.

      Importantly, the underlying infraslow fluctuations in neuromodulatory tone are increasingly recognized as key regulators of sleep spindle dynamics (sigma power), including from our own previous work demonstrating that direct manipulation of locus coeruleus–norepinephrine infraslow rhythms alters spindle organization and sleep continuity. Thus, our findings are conceptually aligned with prior studies linking sigma fluctuations to memory consolidation.

      However, we would like to clarify an important distinction in our translational approach. While Figure 4 demonstrates that heart rate dynamics during sleep can predict memory performance in mice, Figure 5 was designed to address the translational potential of these findings in humans, where direct neuromodulatory readouts are not readily accessible. Here, we deliberately focused on heart rate bursts and their associated sleep dynamics as a clinically tractable physiological measure.

      We acknowledge that the pre-HRB sigma increase may appear conceptually similar to the measure used by Lecci et al.; however, our approach is not equivalent. Rather than selecting spindle or sigma peaks themselves, we aligned analyses to heart rate accelerations and examined the robust upregulation of sigma activity preceding these events. In this framework, sigma activity serves as a physiological readout linked to autonomic dynamics, rather than being the primary anchor of analysis. We chose this measure because heart rate bursts are influenced by multiple physiological factors, and the associated sigma dynamics provided the clearest and most robust relationship with memory outcomes in the human dataset.

      We have revised the Discussion to more clearly acknowledge the similarity to prior work. We believe our findings extend prior observations by providing evidence that sleep-related heart rate fluctuations may serve as a non-invasive readout of the memory-preserving function of sleep.

      (Discussion, page 18). “Previous work in humans demonstrated that the strength of infraslow sigma oscillations correlates with sleep-dependent memory consolidation in humans (Lecci et al., 2017).”

      (1.8) Conclusions as to whether sigma fluctuations might be slightly slower and less powerful in mice are not justified. More work is required to determine which infraslow manifestations are most useful for cross-species comparisons. Moreover, little is currently known about LC activity in human sleep. A careful look into how pupil diameter correlates with sigma power should provide clues for further discussion (see Carro-Dominguez et al). A detailed study of infraslow fluctuations in human sleep should also be discussed https://doi.org/10.1101/2024.11.06.620875.

      We thank the reviewer for this thoughtful comment. We agree that our original phrasing suggesting that sigma dynamics in mice may be “slower and less powerful” than in humans was overly interpretive. We have removed this sentence from the Results section. In addition, we have expanded the Discussion to more thoroughly integrate recent human literature on infraslow sleep dynamics.

      (Discussion, page 19). “Importantly, infraslow fluctuations of sigma power in human sleep have received growing attention. Recent work demonstrates that the infraslow fluctuation of sigma power segments N2 sleep into functional phases associated with arousal and memory-related sleep markers (Dimitriades et al., 2024). Complementary findings using pupillometry show that pupil diameter fluctuates on similar infraslow timescales during NREM sleep and is inversely related to spindle clustering, providing indirect evidence that arousal-related noradrenergic dynamics shape human sleep microstructure (Carro-Domínguez et al., 2025). However, LC activity during human sleep remains inferred rather than directly measured, and systematic perturbation studies linking LC output to spindle–autonomic coupling in humans are currently lacking. Together, these observations underscore both the promise and the current limitations of cross-species comparisons of infraslow sleep dynamics.”

      (1.9) Citation of literature should be as explicit as possible. Referring to "many studies rely on plasma levels..." while including some that actually did real-time fiber photometric measures is misleading.

      We thank the reviewer for this suggestion and agree with the reviewer about the importance of accurately representing prior studies. We have now updated the discussion to clarify this.

      (Discussion, page 15). “Many studies have linked HR to NE (Fawaz and Simaan, 1963; Sundaram et al., 1991; Tanoue et al., 2022; Watson et al., 1979). Many rely on plasma levels of NE, which, while linked to central NE (Gurguis and Uhde, 1998), has low temporal resolution, making causal interpretations harder. In recent years, the use of biosensors and fibre photometry allows for very reliable estimate of the temporal dynamics of NE changes making association to HR more precise (Osorio-Forero et al., 2021).”

      (1.10) Regarding the discussion on the baroreflex: The emphasis is placed predominantly on sympathetically mediated effects. However, the description of the baroreflex loop is incomplete, as it overlooks the substantial contribution of parasympathetic modulation. In particular, heart rate adjustments within the baroreflex are primarily mediated by parasympathetic mechanisms.

      We agree with reviewer that this important notion should be added. We have rephrased the discussion as below:

      (Discussion, page 20). “These neurons suppress the activity of the rostral ventrolateral medulla, ultimately resulting in reflex parasympathetic activation with sympathetic inhibition lowering the HR (Aicher et al., 2000; Lanfranchi and Somers, 2002).”

      (1.11) Regarding the interpretation of HRV analysis, the discussion places disproportionate weight on sympathetic modulation in the interpretation of HRV metrics. This framing is inconsistent with recent conceptual clarifications, including a recent Nature Reviews Cardiology article by Menuet et al. (10.1038/s41569-02501160-z), which cautions against simplistic low-frequency/high-frequency (LF/HF) interpretations of autonomic balance.

      We thank the reviewer for this thoughtful comment. We agree that HRV frequency bands should not be interpreted as exclusive markers of specific autonomic branches. In the original manuscript, we cited Billman (2013), which challenges the validity of LF/HF as a measure of sympatho-vagal balance, to acknowledge these conceptual limitations. However, we recognize that some of our phrasing, particularly in the section discussing compensatory HR decelerations, may have implied branch-specific dominance.

      We have updated the Discussion section as follows:

      (Discussion, page 17). “Interestingly, we found no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection may also somehow be driven by sympathetic outflow. Previous research had indicated that LF power may also represent sympathetic activity, but this interpretation has been challenged due to the mixed contribution of both autonomic branches (Houle and Billman, 1999; Japundzic et al., 1990; Reyes et al., 2013). LF and HF ratio (LF/HF) were traditionally thought to reflect balance between sympathetic and parasympathetic activities (i.e., the sympatho-vagal balance), though currently considered as an oversimplification of nonlinear integration of autonomic signals (Billman, 2013; Pagani et al., 1986). Our findings implicate the VLF may offer a precise marker for central arousal states, given its overlaps with infraslow phasic fluctuations of LC-NE levels. Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014). Within this framework, infraslow LC–NE dynamics may represent one central contributor to these slow cardiac fluctuations during sleep either directly through sympathetic activation or indirectly by inhibition of cardiac vagal activity.”

      We have also replaced the title of the Discussion section title “Locus-coeruleus-mediated sympathetic control drives compensatory heart rate decelerations” with “Locus-coeruleus activation drives compensatory heart rate decelerations via autonomic feedback mechanisms”

      (1.12) In particular, the manuscript attributes VLF power primarily to sympathetic activity, despite evidence that very-low-frequency RR-interval oscillations are strongly dependent on parasympathetic integrity. Notably, parasympathetic blockade has been shown to nearly abolish VLF oscillations in humans (Taylor et al., Circulation, 1998; doi:10.1161/01.CIR.98.6.547).

      We thank the reviewer for highlighting this important point. It was not our intention to imply that VLF HRV is driven primarily by sympathetic activity or to disregard the important role of parasympathetic integrity in shaping VLF oscillations. Our intention was to present the multifactorial and incompletely resolved nature of the VLF component; however, if our wording can be interpreted otherwise, we agree that clarification is warranted.

      In response, we have revised the manuscript to better reflect the current understanding of VLF physiology. Specifically, we now make clearer that, while VLF HRV remains less mechanistically defined than HF and LF HRV, substantial evidence supports a strong parasympathetic contribution to the expression of VLF oscillations. We have incorporated the reviewer-suggested references within this comment and others and clarified this point throughout the revised manuscript.

      (Discussion, page 16). “While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014) there is a clear parasympathetic component (Taylor et al., 1998).”

      (Discussion, page 17). “Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014).”

      (1.13) Moreover, based on work by Armour (2003) and Kember et al. (2000, 2001), the VLF rhythm is thought to emerge from stimulation of afferent sensory neurons within the heart, further arguing against a purely sympathetic interpretation. Together, these findings indicate that the discussion overemphasizes sympathetic mechanisms and underrepresents the contribution of parasympathetic and afferent cardiac pathways to HRV, particularly in the VLF range.

      We thank the reviewer for this important point. Our response to this comment is largely aligned with our response to comment (1.12). In the revised Discussion, we now more clearly acknowledge that VLF oscillations likely arise from more complex cardioautonomic interactions than simple sympathetic and parasympathetic innervation alone. At the same time, we have intentionally avoided an extensive discussion of these mechanisms, as our experimental design does not directly address these pathways.

      Part 2 - There are a number of conceptual issues that need more careful elaboration:

      (2.1) A highly problematic point throughout this study is the choice of AUCs, notably of NE signals and RR intervals, rather than the signal amplitudes. It confuses the correlations to events of different durations, such as MAs and wakefulness. It is not possible to draw conclusions of the kind "cardiac rhythm are tightly coupled with the infraslow phasic NE ..." because such statements ignore that AUCs conflate amplitude and duration of a signal.

      We thank the reviewer for this comment and agree that the distinction between amplitude- and duration-related signal features is important. We deliberately chose AUC measures because our intention was to capture the overall physiological response over time, including both the magnitude and temporal evolution of the signal, rather than relying solely on a single peak value. In this context, AUC provides information about the shape and sustained nature of NE and RR changes within a defined time window, which we considered particularly relevant for temporally dynamic responses.

      At the same time, we appreciate the reviewer’s concern that AUC may conflate response amplitude and duration, particularly when comparing events of different lengths such as microarousals and wakefulness. To minimize this issue, the AUC windows were intentionally kept relatively short and fixed, thereby limiting the influence of prolonged wake episodes or differences in transition duration on the measure. Thus, our intention was not to quantify the total duration of awakenings, but rather the immediate physiological response profile surrounding the event.

      To address this concern more directly, we compared amplitude- and AUC-based measures across all mice. The results are now shown in Suppl. Figure 1g. We found a strong correspondence between amplitude and AUC measurements for NE signals, indicating that the observed relationships are not dependent on the choice of metric. A similar, albeit weaker, relationship was observed for RR responses. This likely reflects physiological constraints on heart-rate dynamics, where the initial heart-rate acceleration is relatively similar across vigilance-state transitions (Figure 1e), while the duration of the response differs substantially. As a result, amplitude measures may underestimate differences between transitions, whereas AUC better captures the extent of the cardiac response. Importantly, direct comparison of NE and RR amplitudes still revealed a significant relationship, supporting the overall conclusion that cardiac and noradrenergic responses are coupled. However, this relationship was weaker than that observed using AUC measures, suggesting that incorporating temporal aspects of the response captures additional biologically relevant information.

      In addition to the figure, we have revised the Results section to include these considerations:

      (Results, page 7): “These comparisons were performed using area under the curve (AUC) estimates of NE and R-R responses. Importantly, comparing AUC and peak amplitude measures showed a similar overall relationship, and the coupling between NE and RR remained significant when only peak amplitudes were considered (Suppl. Fig. 1g), indicating that the findings are not solely driven by the response duration aspect of AUC. The somewhat stronger relationship observed with AUC-based measures may reflect rapid saturation of heart-rate responses across vigilance-state transitions, making response persistence an informative component of the physiological signal.”

      (2.2) A next problematic aspect of the study is the analysis of noradrenergic signals and HR at transitions (e.g., from NREM sleep to sleep wakefulness or to microarousals). The abstract does not mention these data, leaving open how they fit into the paper's message. There are also two problems with it: a) Noradrenaline levels increase with wakefulness, as do many other neuromodulators. This is not novel. Furthermore, why use an AUC measure for 0-25 s when MAs last only 5 or 15 s? b) biosensor signal comparisons are difficult to make for state transitions, because blood flow changes and modifies the fluorescent signal.

      We thank the reviewer for these comments and appreciate the opportunity to clarify the rationale and interpretation of these analyses.

      Regarding the inclusion of vigilance-state transitions, our intention was not to claim novelty in the observation that NE levels increase during wakefulness. Rather, these analyses were included to provide an additional physiological context in which to examine the coupling between NE and heart-rate dynamics. Specifically, the transition analyses allowed us to determine whether graded changes in NE across sleep-to-wake transitions were mirrored by corresponding RR changes (Fig. 1d–e) and whether these responses covaried (Fig. 1f), thereby strengthening the overall conclusion that cardiac dynamics track noradrenergic signaling across naturally occurring sleep-state fluctuations.

      Regarding the use of AUC measures, we deliberately chose this metric because it captures the overall physiological response over a defined time window, including both magnitude and temporal evolution, rather than relying solely on a peak value. In this context, AUC was intended to reflect differences in the overall response profile, for example that NE and RR responses during microarousals may return more rapidly toward baseline than during sustained wakefulness. To minimize the influence of differing event durations, the analysis window was kept fixed (0–25 s) across all transition types. Importantly, we directly compared AUC- and amplitude-based measures and found that they produced largely similar relationships (Suppl. Fig. 1g). The coupling between NE and RR responses remained significant when only peak amplitudes were considered, indicating that the observed relationship is not solely driven by response duration. However, the relationship was somewhat stronger with AUC-based measures, likely because heart-rate responses rapidly saturate across vigilance-state transitions, making response persistence an informative component of the physiological signal. As mentioned in the previous comment, we have added a Figure and new result text to highlight this.

      Regarding the concern about blood-flow related artifacts in fluorescent biosensor signals, we agree that hemodynamic contamination is an important consideration for neuromodulator recordings and applies broadly to biosensor-based measurements, including analyses of infraslow fluctuations. To minimize this issue, ΔF/F calculations were performed using the isosbestic control channel, which serves to correct for movement- and hemodynamic-related signal fluctuations. Furthermore, hemodynamic artifacts typically occur rapidly at state transitions, whereas GRAB-NE signals display slower dynamics. If uncorrected hemodynamic contamination strongly influenced the signal, we would expect abrupt signal distortions tightly aligned to arousal onset, which was not evident in our recordings. While we cannot completely exclude residual hemodynamic influences, we do not believe they account for the graded NE responses observed across vigilance-state transitions or the corresponding relationship with RR dynamics.

      (2.3) Wordings such as 'extent of arousal' to compare MAs and wakefulness are problematic. Microarousals and wakefulness are qualitatively different behaviorally, physiologically, and in terms of neuromodulatory conditions

      We thank the reviewer for the helpful comment. Our intention was not to imply that microarousals and wakefulness are qualitatively identical states differing only in magnitude. Rather, we used arousal in the broader neurophysiological sense, referring to the degree of activation of central arousal systems, with wakefulness representing part of this continuum. However, we recognize that in the sleep field, arousal is often used more specifically to describe brief EEG desynchronization events and that we furthermore use micro-arousals as a broad term for these sleep arousals, which may make our wording confusing. To avoid ambiguity, we have revised the manuscript to replace this terminology with vigilance state, sleep–wake state, or state transitions, depending on the context.

      (2.4) To interpret linear correlations between datasets, even for the ones with Rsquare values < 0.5, as 'predictive' represents an overextended interpretation that is not supported by available evidence. This concern is aggravated due to the use of AUCs that confound amplitudes and time courses. This is in particular the case for Figure 5d, 5k, or 5m…

      We thank the reviewer for this important comment. We agree that the term predictive may overstate the interpretation of these correlations, particularly given the modest R² values in some analyses. Our intention was to highlight an association between heartrate and NE-related measures rather than imply strong predictive performance or causality. We have therefore revised the wording throughout the manuscript to avoid predictive language and instead refer to these relationships as associations or correlations.

      Regarding the use of AUC, we selected summary metrics based on the physiological characteristics of the signal of interest. In cases where responses were characterized primarily by rapid shifts to a new level, amplitude measures were used. In contrast, AUC was chosen when both the magnitude and duration of the response were considered physiologically relevant.

      Part 3. A substantial number of experimental and analytical points require clarification. Here is a list of a few examples; many observations noted here apply equally to other figure panels.

      (3.1) Many figure panels leave it open about whether averages or representative data are shown, how many animals are included, and what kind of measures are plotted.

      We thank the reviewer for this comment and agree that clarity in figure presentation is important. In response, we have carefully revised the figure legends throughout the manuscript to more explicitly state whether data shown are representative examples or group averages, clarify the number of animals included in each analysis, and specify the measures being plotted. We have also added “data are shown as mean ± SEM” where this information was previously not explicitly stated and clarified when n refers to the number of animals. We believe the revised figure legends now provide clearer guidance for interpretation, and further statistical details are available in the accompanying statistics table. It should also be noted that number of animals used for all experiments can be found in the Method section ‘Mice’.

      (3.2) In case experiments were done in a paired manner (e.g., the LC stimulations), individual data points should be shown connected for the different conditions.

      We thank the reviewer for this suggestion. While we agree that connecting individual data points is valuable for paired experimental designs, this visualization is not appropriate for the LC stimulation analyses presented here. Specifically, each condition reflects pooled stimulation events selected across animals rather than a single summary value per animal that can be directly matched across columns. As such, individual points in one condition do not map one-to-one onto points in the next condition, making connected visualizations potentially misleading.

      Importantly, although the data are presented as pooled event-level measures, the paired structure of the experiment was accounted for in the statistical analyses, such that repeated measurements within animals and the paired nature of the design were included in the relevant comparisons.

      (3.3) Numerous analyses involve heart rate measures from the neck EMG during wakefulness. However, it is not specified how, in this case, RR peaks could be detected within the high-activity EMG.

      We thank the reviewer for this question. The procedure for heart-rate detection during wakefulness is described in detail in the Heart rate detection Methods section. Briefly, RR intervals were extracted from preprocessed neck EMG recordings using a previously validated approach for mouse sleep studies. To minimize contamination from movement-related EMG activity during wakefulness, R-peaks were not detected within periods extending 100 ms before and 250 ms after detected movement, as these segments were considered too noisy for reliable peak detection. Heart-rate estimates during these excluded periods were subsequently interpolated using surrounding valid RR intervals to preserve temporal continuity and enable analysis of HR dynamics before and after movement episodes.

      (3.4) Figure 1b: PSD for RR intervals. The supplementary figure says that 11-minutelong NREMS or 5-minute-long NREMS periods were used. These are very rare events in mice, for which the average bout duration is around 2 min and the cycle length is 10 minutes. The methods do not explain how these bouts were chosen, how many of them were included, and why shorter bouts were not analyzed. Single cases or means?

      We thank the reviewer for this important point and agree that additional clarification was warranted. The selection of long NREM (including microarousals) periods was motivated by methodological considerations related to spectral analysis of very slow oscillations rather than by an assumption that these bout lengths are representative of average NREM duration in mice. Because our analysis focused on very-low-frequency dynamics (~1 cycle every 50 s), sufficiently long continuous recordings are required to reliably estimate power at these frequencies and avoid fragmentation-related edge effects that disproportionately affect shorter bouts.

      For this reason, shorter NREM episodes were excluded from the PSD analysis, as they do not provide sufficient duration to robustly capture slow-frequency components. We initially compared PSD estimates using NREM periods of at least 11 min (allowing ~10 VLF cycles) and 5 min duration to assess whether the shorter 5 min recordings introduced bias (Supplementary Fig. 1f). While longer periods resulted in higher overall power estimates, the frequency distribution remained highly similar between conditions. We therefore selected 300 s (5 min) as the inclusion criterion for the remainder of the study, as periods exceeding 10 min are uncommon in mice and would substantially limit analyses across experimental paradigms.

      To further reduce bias related to bout duration, PSD estimates were weighted by NREM episode length, as longer bouts showed systematic effects on power estimates. For Figure 1, the analysis included 144 NREM-with-MA bouts across 7 animals, and data shown represent group means rather than single examples.

      This description is also included in the Methods section/Data Analysis.

      (3.5) Figure panel 1c,d: In Panel c, what is plotted?

      As stated in the figure legend, it is the cross-correlation between NE and R-R. We have added more description in the figure legend.

      A cross-correlation between two signals? If yes, how were these signals chosen per transition? Looks rather like they plot some time course across a transition. What is time point 0?

      We thank the reviewer for pointing out that this analysis was insufficiently explained. The analysis shown represents a cross-correlation between the NE and RR signals, performed to assess their temporal relationship across different vigilance-state transitions. Specifically, for each transition type, NE and RR signals were extracted within the corresponding time windows and cross-correlated to determine the strength and timing of their interaction.

      In this context, time point 0 (lag = 0) represents perfect temporal alignment between the two signals. Positive or negative lags indicate whether changes in one signal systematically precede or follow changes in the other. Due to methodological differences between the rapid electrical heart signal and the slower fluorescent NE signal, we intentionally avoided overinterpreting fine temporal lead–lag relationships. Rather, the aim of this analysis was to characterize the overall interaction between the two signals across transitions.

      Because the relationship between NE and RR was predominantly inverse, negative cross-correlation values indicate that increases in NE are associated with decreases in RR (i.e., faster heart rate), and vice versa. Thus, this analysis served primarily to confirm and extend our other findings by quantifying the interaction between NE and heart-rate dynamics across vigilance-state transitions.

      We have revised the descriptive sentence in the Results section to improve clarity and have added a more detailed description of the signals included directly in the figure panel.

      (Results, page 7). “Here, cross-correlation analysis revealed a predominantly negative relationship between NE and RR across vigilance-state transitions, indicating that increases in NE were associated with reductions in RR (i.e., faster heart rate; Fig. 1c).”

      (Figure 1 legend). “Cross correlation (how strongly and at what temporal offset the two signals covary) between NE and RR during transitions”.

      In Panel d, what is measured here? Is the time point of the dotted line a NA trough or the moment of a transition? Show the data with connected lines. Heart rate calculation during wakefulness?

      We thank the reviewer for these questions and apologize that this was not sufficiently clear. In panel d, the dotted line indicates the NE trough, not the moment of a vigilance state transition, as specified in both the figure and figure legend. The analysis is aligned to detected NE troughs during sleep and examines the subsequent physiological dynamics, including transitions into wakefulness.

      We have updated the sentence in the result section to make it a bit more clear:

      (Results, page 7): “To explore the NE-RR relationship across sleep-wake transitions, we examined four progressive sleep-to-wake transitions using the preceding NE trough as the time stamp (time 0):…”

      Regarding the suggestion to connect data points, as noted in an earlier response, these analyses are based on event detection, where multiple events contribute from each animal. Thus, each condition represents pooled events across animals rather than a single matched value per animal, making connected-line visualizations inappropriate and potentially misleading. Importantly, the paired structure of the experimental design was accounted for in the statistical analyses.

      Heart-rate detection during wakefulness was usually not possible due to movement artefacts in the EMG. In the detection, we excluded periods with movement. Our EMG-based detection approach is described in the Heart rate detection Methods section. Briefly, movement-contaminated periods were excluded from R-peak detection (100 ms before and 250 ms after detected movement), and RR intervals were subsequently interpolated using surrounding valid values to preserve temporal continuity. Because analyses were aligned to NE troughs occurring during sleep, we limited the temporal window to avoid excessive contamination from movement-related noise associated with subsequent wakefulness. We have slightly revised the text to make these points clearer.

      Panel C is a cross correlation between NE and RR from the same traces included in the mean traces. x=0 is the NE through like the other figures.

      Panel D is the mean traces of NE during transitions with the dotted line at x=0 being the NE though

      Yes, but x = 0 for the cross-correlation is not the NE trough. It says something about how aligned NE and R-R are in time (see response further up in (3.5)).

      (3.6) The sigma band should ideally be chosen between 10-15 Hz for better consistency with the literature.

      We thank the reviewer for this suggestion and agree that consistency in the definition of frequency bands is important for comparison across studies. The sigma range used in the present manuscript was selected to match our previous publications and analyses, thereby allowing direct comparison with our earlier findings on LC-mediated regulation of sleep spindles and infraslow sleep dynamics. Maintaining the same band definition also ensured consistency across the datasets analyzed in this study.

      We acknowledge that having a shared definition of sigma power would improve alignment and facilitate comparisons across laboratories. At present, there remains some variability in the exact frequency boundaries used for spindle and sigma analyses across studies and species, although there are ongoing efforts within the sleep field to improve standardization. Importantly, we do not expect that modest adjustments of the sigma-band boundaries would materially affect the conclusions of the present study, as the spindle-related activity of interest lies well within the selected frequency range and the observed effects are broad rather than restricted to a narrow frequency bin.

      Moving forward, we aim to follow emerging consensus recommendations where appropriate. To clarify this point for readers, we have added a statement in the Methods section explaining that the sigma band was chosen to maintain consistency with our previous publications.

      (Methods: EEG power, page 29). “The sigma band was defined as 8 - 15 Hz to maintain consistency with our previous publications. Modest differences in sigma-band boundaries are not expected to affect the main conclusions.”

      (3.7) What are 'extreme LC stimulation frequencies'. The only information available is that stimulations were done for 2s at 20 Hz.

      We thank the reviewer for pointing out that this wording was unclear. By “extreme LC stimulation frequencies”, we did not refer to the within-stimulation pulse frequency (which remained constant at 20 Hz for 2 s across all conditions). Rather, we referred to the effective frequency of LC activation at the infraslow timescale, which was progressively increased through the closed-loop stimulation paradigm.

      Specifically, stimulations were triggered when NE levels crossed increasingly permissive thresholds during the descending phase of the endogenous NE signal. As thresholds increased over time (from −15 ΔF/F (%) to +5 ΔF/F (%)), stimulations occurred progressively earlier in the infraslow cycle, thereby compressing the oscillatory period and increasing the effective frequency of LC recruitment while preserving the endogenous temporal structure of NE dynamics.

      Thus, “extreme stimulation frequencies” refers to the highest rate of repeated LC activations achieved through the closed-loop paradigm, where stimulations became increasingly frequent at the infraslow level rather than changes in the 20 Hz pulse train itself. To avoid confusion, we have revised the wording throughout the manuscript to refer more explicitly to faster infraslow LC activation frequencies or increased infraslow stimulation frequency.

      We have added more information in the result section to highlight this better.

      (Results, page 8). “We employed a closed-loop paradigm, where LC stimulations (2 s 20 Hz (10 ms) blue laser pulses with a light intensity of 5 mW) were triggered when NE levels fell below increasing thresholds (-15, -10, -5, 0 and 5 ΔF/F (%), Fig. 2a-b, Methods). This approach enabled controlled compression of the infraslow NE cycle by triggering LC activation during the descending phase of the endogenous NE signal, thereby increasing the effective oscillatory frequency while preserving the temporal structure of physiological LC–NE dynamics. This strategy allowed us to test whether heart-rate responses continue to track LC-driven NE fluctuations as the infraslow rhythm becomes progressively faster.”

      (3.8) Could the discrepancy between panels 4c, left and right, be due to limited sample size? The whole figure lacks indications of sample numbers, making interpretation difficult.

      We thank the reviewer for this comment. We assume the reviewer is referring to the apparent discrepancy between the NE response and heart-rate response following LC suppression (Fig. 4c–d), where LC inhibition induced a clear reduction in NE levels, whereas mean heart-rate responses were less pronounced.

      The figure legends state sample sizes and event numbers. Specifically, for these analyses, n = 8 animals (4 Arch, 4 YFP) were included, comprising 48 Arch events and 38 YFP events, our interpretation is that this discrepancy reflects a biological observation rather than a failed manipulation. Specifically, LC suppression robustly reduced NE levels, confirming the effectiveness of the optogenetic intervention, whereas HR did not exhibit a similarly consistent group-level response. However, as highlighted by the correlation analyses, variability in RR responses remained associated with the magnitude of NE suppression, suggesting that heart-rate dynamics still reflected noradrenergic modulation at the individual-response level despite the absence of a strong mean effect.

      (3.9) It would be great if Figure 4 j could be more explicitly illustrated. For example, behavioral traces that lead to higher NFR and corresponding changes in RR AUC should be shown for animals with large and small effect sizes. The sample size seems excessively low. Can this explain the difference in slopes compared to Figure 4e, right panel?

      We thank the reviewer for this suggestion. To improve the interpretation of Figure 4j, we have now added representative examples illustrating animals with high and low memory performance and their respective RR responses following LC suppression.

      We agree that the sample size for this analysis is limited. As noted in the Methods, Figure 4j is based on a secondary analysis of a previously published dataset (Kjaerby, Andersen et al., 2022), where heart-rate measures were retrospectively extracted from EMG recordings. Due to noise-related limitations in RR detection, reliable cardiac measures could not be obtained from all animals, reducing the number of subjects available for this analysis. For this reason, we have deliberately avoided direct statistical comparisons between Arch and YFP animals and instead limited our conclusions to the observed association between RR responses and memory performance across animals.

      Regarding the difference in slope compared with Figure 4e, the two analyses are based on different levels of aggregation and address different questions. Figure 4e examines the relationship between NE and RR responses across individual LC suppression events, resulting in multiple observations per animal. In contrast, Figure 4j uses a single mean RR response and a single behavioral outcome per animal. Furthermore, Arch and YFP animals were pooled in Figure 4j to maximize statistical power and because the dataset was not sufficiently powered for direct group comparisons. Consequently, the slopes are not expected to be directly comparable between the two figures.

      (3.10) It would be important to show anatomical validation of viral expression in THcre animals and optic fiber positioning.

      We thank the reviewer for this comment. Anatomical validation of viral expression in TH-Cre animals and optic fibre placement was performed for these experiments and has been reported previously in Kjaerby et al. Nature Neuroscience paper, from which this dataset was derived. Specifically, viral targeting and fibre positioning were histologically verified as part of the original experimental validation. This is mentioned in the Method section ‘Surgery’: Viral expression and injection sites were validated through immunostaining of perfused brain slices from the experimental animals (see Kjaerby et al. (3) for more information).

    1. eLife Assessment

      This important study reveals distinct representations of task-related information in the dendrites and somata of cortical neurons during sensorimotor learning and behavioral adaptation. The evidence is compelling, combining imaging of dendritic and somatic activity during behavior to demonstrate compartment-specific encoding of sensory cues, motor actions, and corrective signals. The work will be of broad interest to neuroscientists studying dendritic computation, motor learning, and the cellular mechanisms underlying adaptive behavior.

    2. Reviewer #1 (Public review):

      [Editors' note: this version has been assessed by the Reviewing Editor without further input from the original reviewers. The authors have addressed the comments raised in the previous round of review.]

      Summary:

      In this manuscript, Scheib et al. identify distinct calcium dynamics in the somata and tuft dendrites of layer 5 pyramidal cells in mice performing a licking task. Animals are trained to lick water ports on the left or right following an acoustic cue, and can adjust their targeting when the ports are displaced. For tongue premotor cortical neurons projecting to the ventromedial thalamus, calcium transients in tuft dendrites are tightly locked to the direction-instructive cue, while somatic calcium signals are more broadly dispersed and more frequently synchronized with tongue motion and port contact. Finally, when the targets are shifted, tufts exhibit a sparse but large corrective signal on an improperly-targeted first lick, and the changes in population activity in the tufts and somata differ after adaptation to the new port locations.

      Strengths:

      In my opinion, this is a very strong manuscript which reports several novel and significant observations, contains high-quality data and (for the most part) reasonable analyses, and is clear and well-written. Most prior studies of cortical sensorimotor processing have measured the output of neurons using extracellular recording - an approach which obscures potentially important signaling differences between neuronal compartments. This study leverages cutting-edge imaging techniques in mice to document large, time-dependent differences between calcium signals at cortical somata and tuft dendrites. This phenomenon could have major implications at the cellular level for synaptic plasticity, and at the systems and behavioral levels for motor adaptation.

      Weaknesses:

      At a conceptual level, the authors may wish to elaborate a bit on what sensorimotor computation they think the circuit is implementing, and how their results help explain this implementation. Several possibilities are raised: tuft activation could "prime" the pyramidal cells in advance of movement initiation (line 319ff), or could track errors to engage plasticity (line 351ff) and solve the credit assignment problem (line 362ff). It might be helpful to make one of these proposals more concrete with a computational model, but this is not strictly necessary. [The authors explain that they will address this with modeling work in subsequent research.]

    3. Reviewer #2 (Public review):

      Summary:

      The authors set out to compare functional encoding in the tuft dendrites and somata of a specific cortical cell type during motor planning and learning.

      Strengths:

      The investigation of a specific projection type (L5 ET) is a strength that aids reproducibility and interpretation. The elegant approach to increasing the depth of field of dendritic imaging is another strength. The data analyses are largely clear in their methods, scope, and interpretation. The writing is extremely clear and appropriately referenced, with an excellent Introduction, in particular.

      Weaknesses:

      This work is largely observational, describing signals that might reflect computational transformations and/or instruct plasticity, but those possibilities have not yet been deeply investigated. The manuscript does a good job of laying out these as future directions.

    4. Reviewer #3 (Public review):

      Summary:

      This article by Scheib et al. investigates how layer 5 extratelencephalic (ET) neurons in the frontal cortex encode sensorimotor information during motor learning, focusing on differences between their apical tuft dendrites and somas. The authors alternated recordings among these ET neuronal compartments in the mouse anterior lateral motor cortex (ALM) during a cued directional licking task with a target port shift. They found that while tuft dendrites predominantly encode sensory cues, with a subset selectively active during corrective actions, somatic activity was more strongly associated with action timing. Additionally, learning induced divergent plasticity: tuft dendrites increased their selectivity but decreased response gain, maintaining stable net selectivity, whereas somas showed increased net selectivity early in learning. Together, these findings reveal distinct sensorimotor representations and learning-related plasticity in dendritic and somatic compartments, providing insight into how compartment-specific activity in the frontal cortex may contribute to motor skill acquisition.

      Strengths:

      The authors developed an innovative imaging approach and a comprehensive data analysis pipeline to address a knowledge gap in the literature. By alternating imaging of dendritic tufts and somas in the same animals, they compare compartment-specific activity during motor learning and identify distinct encoding of task variables and learning-related plasticity across these compartments. Interestingly, a subset of dendritic tufts shows activity associated with corrective actions. The findings are discussed in the context of current theories of dendritic computation, credit assignment, and motor learning, providing a useful foundation for future mechanistic studies.

      Weaknesses:

      No major weaknesses were identified.

    5. Author Response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study reveals distinct representations of task-related information in the dendrites and somata of cortical neurons during sensorimotor learning and behavioral adaptation. The evidence is compelling, combining simultaneous imaging of dendritic and somatic activity during behavior to demonstrate compartment-specific encoding of sensory cues, motor actions, and corrective signals. The work will be of broad interest to neuroscientists studying dendritic computation, motor learning, and the cellular mechanisms underlying adaptive behavior.

      Thank you for this excellent summary. We recommend one change: removing the word “simultaneous”. It could perhaps be replaced with “concurrent” or simply omitted. Tuft dendrites and somata were imaged on alternating days, and most readers will probably interpret “simultaneous” as implying a faster, interleaved sampling rate.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Scheib et al. identify distinct calcium dynamics in the somata and tuft dendrites of layer 5 pyramidal cells in mice performing a licking task. Animals are trained to lick water ports on the left or right following an acoustic cue, and can adjust their targeting when the ports are displaced. For tongue premotor cortical neurons projecting to the ventromedial thalamus, calcium transients in tuft dendrites are tightly locked to the direction-instructive cue, while somatic calcium signals are more broadly dispersed and more frequently synchronized with tongue motion and port contact. Finally, when the targets are shifted, tufts exhibit a sparse but large corrective signal on an improperly-targeted first lick, and the changes in population activity in the tufts and somata differ after adaptation to the new port locations.

      Strengths:

      (1.1) In my opinion, this is a very strong manuscript which reports several novel and significant observations, contains high-quality data and (for the most part) reasonable analyses, and is clear and well-written. Most prior studies of cortical sensorimotor processing have measured the output of neurons using extracellular recording - an approach which obscures potentially important signaling differences between neuronal compartments. This study leverages cutting-edge imaging techniques in mice to document large, time-dependent differences between calcium signals at cortical somata and tuft dendrites. This phenomenon could have major implications at the cellular level for synaptic plasticity, and at the systems and behavioral levels for motor adaptation. As described below, I have only one major technical concern (which should be addressable with additional analysis), along with several relatively minor suggestions for improving the manuscript.

      We thank the reviewer for their insightful summary of the significance of the differences that we identified in the task-related activity of tuft dendrites and somata.

      Weaknesses:

      (1.2) At a conceptual level, the authors may wish to elaborate a bit on what sensorimotor computation they think the circuit is implementing, and how their results help explain this implementation. Several possibilities are raised: tuft activation could "prime" the pyramidal cells in advance of movement initiation (line 319ff), or could track errors to engage plasticity (line 351ff) and solve the credit assignment problem (line 362ff). It might be helpful to make one of these proposals more concrete with a computational model, but this is not strictly necessary.

      We thank the reviewer for this feedback. We absolutely agree that detailed computational models of each proposed computation will be very valuable and constitute an important follow-up to this work. We hope to collaborate with theorists to take that next step. Each possible computation noted by the reviewer reflects distinct differences that we observed in the task-related activity of tuft dendrites and somata. They are not mutually exclusive hypotheses to explain the same phenomenon. As such, we think they are best addressed independently in future modeling work. By making the data and a concise description of the main findings available immediately, we hope to allow computational experts in each of these areas to take advantage of the results of this study without delay.

      (1.3) My only major technical concern relates to the analyses in Figures 4F-H, 5G-I, and 6H-K (c.f. equations 2-5). Typically, one identifies population-level factors by projecting neural activity onto fixed dimensions of interest; this makes it possible to see how activity evolves over time along interpretable coordinates. Here, however, the coding directions are redefined at each time point, so the "choice" activity at time t is actually a different signal from the "choice" activity at t+1. This procedure is a bit like comparing the activity of one neuron at one time point with the activity of a different neuron at a later time point. It also makes the physiological interpretation more complicated: if the dimensions are fixed, one can see how a downstream neuron could "read out" the signal by computing a weighted sum of the activity of upstream neurons, but it is harder to see how this could happen if the weights are always rotating.

      We thank the reviewer for raising this point. We agree that our use of projections along coding directions (CDs) defined at each time point is a less conventional use of coding directions, although nearly identical calculations have been previously used to assess population-level selectivity and code stability in this task (Chen et al., 2017; Yang et al., 2022). As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task-dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task-variable. Finally, because the primary goal of the study was to identify any differences between tuft dendrite and somatic encoding, we think that calculating the population selectivity at each timepoint gives readers a less biased view of the selectivity of the two compartments, whereas calculating a CD over a single arbitrary time window could conflate differences in dynamics with differences in selectivity.

      We agree that calculating the CD at each timepoint makes it hard to see where the code is stable and where it is rotating, and thus how a downstream neuron might “read out” the signal. To provide this information, we have added new panels to the supplement showing the correlation of CDs across time (Figure 4 - figure supplement 1B,D). We also now provide this information for CR-CA in Figure 5—figure supplement 2A (the plots previously presented in 2A were the correlations of CR with CA, rather than CR-CA; an error that has been fixed). The following changes were also made to the Results section to clarify this issue:

      “From the linear model, we calculated coding directions (CDs) at each timepoint that maximally separated Stimulus, Choice, and Outcome activity (Figure 4F; Figure 4—figure supplement 1A) and estimated the direction and selectivity along each dimension across time (Figure 4—figure supplement 1B-E; see Methods). Allowing CDs to rotate in time (see Figure 4—figure supplement 1B,D), although unconventional, ensured that comparisons of population selectivity across the two compartments were not biased by the selection of an arbitrary CD time window.”

      We also identified a mistake in the description of CD orthogonalization in the Methods, which has been corrected as follows:

      “For each timepoint, each selectivity CD was then orthogonalized with respect to the other two selectivity CDs by a QR decomposition in which that selectivity CD was last in the order.”

      (1.4) A few comments on the behavioral task and results. After the port shift, the error rate is quite high, and doesn't diminish much between the early and late epochs (approximately 42% and 38% error rate, respectively; Figure 1I). That is, mice do not seem to fully master the task. Clearly, animals do alter their aim, but even this does not seem to change much between early and late periods (Figure 1J). I recommend that the authors show the behavioral data at a finer level of granularity (e.g., by plotting the change in exit trajectory on all individual trials across sessions, with a loess fit) to allow an assessment of the adaptation rate and when adaptation saturates. It would also be more conventional to refer to the behavioral changes as "motor adaptation," instead of "skill learning." (The latter would be appropriate if the port offset were randomized across trials, and animals received two separate cues for direction and offset, but I suspect this task would be too difficult for mice to learn.)

      We agree with the reviewer that by the end of the late period, performance on the right side (Figure 1I) has still not returned to pre-shift levels. This may reflect mice not fully mastering the task, as the reviewer suggests, or it may reflect that after the shift, the right port is substantially more difficult to reach than the left port. Unfortunately, because of high animal-to-animal and lick-to-lick variability, plotting the post-shift lick angle at a finer level of granularity is not statistically informative.

      With regard to the nature of the learning in our task, we selected “skill learning” as the best description of the motor learning task based on distinctions between adaptation and the learning of motor skills by Krakauer et al., 2019 and Heald et al., 2021. Conceptually, the difference is whether an existing motor controller memory is simply updated with new parameters, or whether the motor context has changed sufficiently that a distinct motor controller memory (which can still use parts of previous memories) is formed. In our task, after the port shift the left port forms an obstacle to reaching the right port. This obstacle was simply not present before the shift. Before the shift, ports were approximately equidistant from the mouth and easily avoided given the port separation and tongue width. Thus, avoiding an obstacle would presumably not be part of the initial motor controller memory and a distinct memory would need to be constructed.

      We agree with the reviewer, however, that given that we do not have fine-timescale dynamics of behavioral changes in response to the shift, and did not conduct other experiments (such as returning the ports to their original location) that would typically be conducted to identify “adaptation-like” or “skill learning-like” dynamics, we cannot empirically distinguish between the two. We now clarify in “Study limitations” that we call the studied behavior “skill learning” based on the nature of the task, but that our behavioral analysis cannot distinguish between adaptation and skill learning:

      “We refer to the behavioral paradigm as motor “skill learning” strictly based on the nature of the task. After the shift, mice must avoid a new obstacle close to the mouth (i.e., the left port), which we assume requires the formation of a distinct motor controller memory and therefore would be considered skill learning (Krakauer et al., 2019). However, we did not confirm that the mice exhibited specific behavioral characteristics of skill learning and it is possible that other kinds of motor learning (e.g., motor adaptation) were dominant.”

      (1.5) This is perhaps a semantic point, but it might not be entirely accurate to refer to the activity evoked by the directional cue as "sensory." Typically, a "sensory" response should encode some feature of a stimulus - in this case, the frequency of a tone. Here, it seems likely that the cue-aligned activity reflects the instructed lick direction, rather than the auditory information per se. (Presumably, these premotor neurons do not have well-behaved auditory tuning curves.) By comparison, in macaques performing center-out reach tasks, activity in dorsal premotor cortex rapidly ramps up following a visual cue instructing the direction of an upcoming reach, but one usually wouldn't refer to this activity as "visual" or "sensory" (though this is sometimes done). I suggest the authors either use "Instruction" or similar (e.g., in Figure 4F), or clarify in the text whether they think the activity is a genuine auditory response or something else.

      We understand how this could cause confusion. “Sensory” was meant to denote the nature of the differences in external events between the trial types used to calculate selectivity, not to imply that the activity was necessarily selective for detailed features of the cues outside the context of the task. Previous work in ALM cortex has labeled this selectivity direction as “stimulus” (Yang et al., 2022; Chen et al., 2024) to better emphasize that it is simply defined by the external cue. Where appropriate, we have revised the article to use “stimulus” or “instructional cues” in place of “sensory” for clarity and to better conform with convention.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to compare functional encoding in the tuft dendrites and somata of a specific cortical cell type during motor planning and learning.

      Strengths:

      (2.1) The investigation of a specific projection type (L5 ET) is a strength that aids reproducibility and interpretation. The elegant approach to increasing the depth of field of dendritic imaging is another strength. The data analyses are largely clear in their methods, scope, and interpretation. The writing is extremely clear and appropriately referenced, with an excellent Introduction, in particular.

      We thank the reviewer for their appreciation of the study design, imaging methods, and scholarship of the article.

      Weaknesses:

      (2.2) It is not obvious whether the selected labeling strategy avoids labeling Layer 6 CT neurons, which would contaminate dendritic recordings. The images provided suggest enrichment in L5, but a discussion of this important potential caveat is warranted, especially since within-cell comparisons of apical dendrites to somata were not performed.

      We thank the reviewer for emphasizing the need to discuss this potential issue. For the following reasons, it is likely that the vast majority of dendrites we imaged in layer 1 originated from layer 5 ET neurons. First, as the reviewer notes, the provided images suggest enrichment in layer 5. This enrichment likely reflects the fact that most L6 CT neurons in motor and premotor cortex send denser projections to other thalamic nuclei than to VM thalamus (Winnebust et al., 2019, Cell), where we targeted our retrograde-Cre injections. Second, L6 CT neurons are predominantly untufted (Ledergerber and Larkum, 2010, J. Neurosci.), including in motor and premotor cortex (Peng et al., 2021, Nature; Ichikawa, 2025, Front. Neuroanat.). A recently identified subclass of L6 CT neurons in secondary motor cortex has dense projections to VM thalamus, but this class also appears to extend minimal dendrites into L1 (Li et al., 2024, bioRxiv). Nonetheless, we did not label post-hoc tissue collected from imaged mice with markers of precise laminar boundaries, and thus cannot definitively rule out the possibility that dendrites from a subclass of L6 CT neurons with tuft dendrites were also imaged. We have added the following paragraph to the “Study limitations” section to make readers aware of these issues:

      “L5 ET neurons in premotor cortex elaborate extensive tuft dendrites in L1, whereas Layer 6 (L6) corticothalamic (CT) neurons are predominantly untufted (Jiang et al., 2020; Peng et al., 2021). Thus, although we cannot rule out the possibility that dendrites from a subclass of L6 CT neurons were also sampled, it is likely that the vast majority of dendrites we recorded in L1 originated from L5 ET neurons.”

      (2.3) The application of DeepInterpolation to dendritic data appears to be novel, and little detail or vetting is provided. The reader is left guessing: Was the model retrained or fine-tuned on dendritic data? How does the denoising affect the resulting segmentation and activity traces? Is denoising necessary for this workflow?

      We thank the reviewer for requesting this useful additional information.

      In all cases, the model was retrained for each dendritic or somatic imaging session. Denoising improved segmentation consistency, as measured by comparing segmentations of individual sessions from the same animal. This is now specified in the Methods as follows:

      “The DeepInterpolation model was trained on each imaging session prior to denoising of that session. Denoising prior to NMF-based segmentation resulted in more robust and consistent dendrite segmentation than NMF-based segmentation without prior denoising (0.79 +/- 0.01 ⍴ vs. 0.46 +/- 0.01 ⍴; mean of the max Spearman correlation of components across sessions; random subsample of N = 3 mice, 15 sessions, 400 components).”

      With regard to how denoising impacts activity traces, examples were shown in Figure 2I, K. To provide more quantitative information to the reader, we calculated estimates of the power and reliability of the spectral content of dendrite activity traces extracted with or without denoising. These data are now shown in the new panel, Figure 2 - figure supplement 2G. The power spectral density of the denoised activity and the estimated reliable power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C). Some frequencies beyond this point have been suppressed beyond what would be expected due to photon shot noise (as estimated by the replicate coherence-weighted PSD, or “recoverable” PSD). Further characterization of the precise nature of the suppressed high-frequency information – which could be suppressed artifacts (e.g., fast brain motion) or lost signal detail (i.e., GCaMP8m rise kinetics) – is beyond the scope of this paper.

      Details of the PSD calculations have been added to the Methods, and the following statement has been added to the Results: “Power spectral density of the denoised traces and the coherence-weighted power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G; Methods), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C).”

      (2.4) The activity patterns of the recorded cells appear to lack the characteristic ramping during the delay epoch previously reported in both calcium imaging and electrophysiology studies. Given that a major contribution to the significance of the work is to constrain models of ALM function, a discussion of how the data aligns with previous measurements in the same circuit would improve the work.

      Preparatory selectivity and ramping activity can be seen in Figure 3H, Figure 6I, and Figure 5 – figure supplement 1B. We note that in ALM cortex, the ramping mode explains a minority of the total variance (~17%, Yang et al., 2022), but it can appear particularly prominent in projections along certain fixed CDs.

      (2.5) It would be very informative to compare differences in signals between dendrites and somata of the same cells. Consistently tracing dendrites to their respective somata would assuage worries of potential contamination from dendrites of deeper cells and enable more direct comparisons of signal transformations between dendrites and somata. It would be good to understand the relationship between dendritic calcium signals and backpropagating action potentials in this task. The authors detect less frequent calcium events in tufts versus somata; is this due to selective backpropagation of action potentials? The dynamics of this process were recently investigated by Adam Cohen's group in vivo and in vitro, and measurements in the present settings could be compared to such work.

      We agree with the reviewer that being able to compare differences in signals between the dendrites and somata of the same cells would be very valuable. However, reliable tracing of tuft dendrites to somata from in vivo 2P anatomical imaging requires extremely sparse labeling, such that very few neurons are recorded per animal (Kerlin et al., 2019, eLife; Otor et al., 2022, Science). As stated in the “Study limitations” section of the Discussion, we suspected (correctly) that some task-related selectivity (i.e., selectivity for corrective action) would be sparsely represented in the dendrites, and thus adopted a labeling and image processing strategy that allowed us to record from many dendrites per animal. This strategy necessarily comes at the expense of generating a labeling density that precludes reliable tracing of tuft dendrites to their respective somata based on 2P morphology alone. As discussed in our response to reviewer comment 2.2 and a new paragraph of “Study limitations,” substantial contamination of the dendrite recordings by dendrites of L6 CT neurons is highly unlikely. Future studies could use simultaneous functional imaging across large volumes combined with activity-based segmentation or post-hoc high-resolution imaging of tissue sections registered to in vivo 2P imaging to accomplish both high-throughput dendritic imaging and reliable tracing.

      We thank the reviewer for pointing out that we could discuss selective backpropagation as a potential mechanism more explicitly. Our results are consistent with previous studies of L5 tufts in vivo (Francioni et al., 2019, eLife), including in ALM cortex (Maristany de las Casas et al., 2026, Science), that reported that rates of multi-branch calcium transients in the tuft dendrites of L5 neurons are lower than somatic spike rates. As discussed in “Study limitations,” there is not a clear approach in our data to determine the precise nature of the events underlying the calcium transients we measured in the tuft dendrites. Selective backpropagation of action potentials is certainly one possibility and we agree that recent research from Dr. Adam Cohen’s group should be discussed. We have added the following to the Discussion:

      “Based on previous calcium imaging of L5 tufts in ALM cortex of mice engaged in similar tasks (Kerlin et al., 2019; Maristany De Las Casas et al., 2026), we suspect that most of the activity we measured was coincident with global tuft or hemi-tree events, as well as somatic spiking. Recent in vivo voltage imaging in the hippocampus has also indicated that most spikes in distal dendrites start as bAPs that have been selectively amplified (Wu et al., 2026; Lee et al., 2026).”

      (2.6) The Coding Direction analyses presented in this work, while consistent with previous literature on population codes in ALM, are at odds with the nature of the measurements here. The changes in representation that occur between the dendrites and soma of an individual cell are probably best thought of in terms of the dynamics of signals themselves within individual neurons, rather than in the information encoded across a population.

      We thank the reviewer for giving us the opportunity to clarify this issue. As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task variable in the population activity. Thus, the analyses are not at odds with the nature of the measurements in the study.

      Nevertheless, it is true that by recalculating the CD at each timepoint, our selectivity projections do not provide the same information as conventional projections along a fixed CD, which can indicate where the selectivity code is stable and where it is changing. To provide this information we have added new panels to the supplement showing the correlation of selectivity CDs across time (Figure 4 - figure supplement 1B, D).

      (2.7) This work is largely observational, describing signals that might reflect computational transformations and/or instruct plasticity, but those possibilities have not yet been deeply investigated. The manuscript does a good job of laying out these as future directions.

      We agree with the reviewer. As noted by the reviewer in comment (2.1), we combined a number of approaches in an innovative manner to explore how tuft dendrite activity differs from somatic activity at the population level during motor learning. These measurements provide the necessary foundation for future mechanistic studies and we think it is appropriate to share them at this stage of investigation and in the format of this article.

      Reviewer #3 (Public review):

      Summary:

      This article by Scheib et al. investigates how layer 5 extratelencephalic (ET) neurons in the frontal cortex encode sensorimotor information during motor learning, focusing on differences between their apical tuft dendrites and somas. The authors alternated recordings among these ET neuronal compartments in the mouse anterior lateral motor cortex (ALM) during a cued directional licking task with a target port shift. They found that while tuft dendrites predominantly encode sensory cues, with a subset selectively active during corrective actions, somatic activity was more strongly associated with action timing. Additionally, learning induced divergent plasticity: tuft dendrites increased their selectivity but decreased response gain, maintaining stable net selectivity, whereas somas showed increased net selectivity early in learning. Together, these findings reveal distinct sensorimotor representations and learning-related plasticity in dendritic and somatic compartments, providing insight into how compartment-specific activity in the frontal cortex may contribute to motor skill acquisition.

      Strengths:

      The authors developed an innovative imaging approach and a comprehensive data analysis pipeline to address a knowledge gap in the literature. By alternating imaging of dendritic tufts and somas in the same animals, they compare compartment-specific activity during motor learning and identify distinct encoding of task variables and learning-related plasticity across these compartments. Interestingly, a subset of dendritic tufts shows activity associated with corrective actions. The findings are discussed in the context of current theories of dendritic computation, credit assignment, and motor learning, providing a useful foundation for future mechanistic studies.

      We thank the reviewer for highlighting interesting findings in the paper and their assessment that it provides a “useful foundation for future mechanistic studies”.

      Weaknesses:

      No major weaknesses were identified.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A very minor suggestion: it would be useful to mention the model organism in the abstract or title.

      (1.6) Thank you for catching this. We have added the model organism to the abstract as follows:

      “Using longitudinal two-photon calcium imaging, we investigated sensorimotor encoding in the apical tuft dendrites and somata of L5 extratelencephalic (ET) neurons in the frontal cortex of mice during learning of a discrete change to a cued dexterous action.”

      Reviewer #3 (Recommendations for the authors):

      Major:

      (3.1) Lines 197-199: It is unclear why the authors conclude that somas have stronger representations of choice and task outcome. In Figure 4G, there is no significant difference between dendrites and somas for Choice or Outcome coding selectivity. The differences in the Sensory/Choice and Sensory/Outcome ratios shown in Figure 4H,I are likely explained by stronger Sensory selectivity in dendrites (Fig 4g), rather than by stronger Choice or Outcome encoding in somas.

      We agree with the reviewer’s interpretation of the data. The statement at 197 - 199 was meant to reflect relative selectivity, but it was imprecise. We have replaced that sentence with the following, more precise sentence:

      “Somatic activity also encoded these features, but the representation of the stimulus was weaker – and the representation of action timing was stronger – than in the tuft dendrites.”

      (3.2) Figure 4B: The authors realign FL-associated IRFs to GO-cue timing using the mean FL latency for each trial type and animal. Because FL timing is jittered across trials and may differ between CL and CR trials, this could smear the realigned traces and complicate the interpretation of contact-associated activity. The authors should consider using trial-by-trial FL timing for realignment or quantify the impact of FL-timing variability on the resulting traces.

      We aligned average GO- and contact-IRFs in Figure 4B so that comparisons of their magnitudes could be drawn from the same time window.

      With regard to jitter across trials, we think the reviewer may have misinterpreted how the mean IRFs in Figure 4B are calculated. The contact-IRF, by its nature, is calculated once per animal and trial type with respect to FL timing and shifted once based on mean FL latency. There is no smearing due to trial-to-trial FL timing.

      With regard to systematic differences in FL timing across animals and CR vs. CL, the reviewer is correct that this could – in theory – smear the realigned mean contact-IRF shown in Figure 4B. However, differences in mean FL latency across animals and trial-types are small compared with the long-timescale contact-IRFs. Thus, the non-realigned (i.e., always FL-aligned) mean contact-IRF looks nearly identical to Figure 4B just globally offset in time, as shown in Author response image 1:

      Author response image 1.

      Since this is nearly identical to data already presented in Figure 4B, we do not think it is necessary to include it in the revised article. However, we have added the following to the Methods:

      “Population averages of contact-IRFs that were not shifted prior to averaging were nearly identical (excluding the overall temporal shift; data not shown), indicating that pooling of mean IRFs across animals and trial types produces minimal smearing of the final population IRF.”

      (3.3) Figure 5A: Are CA trials specific to motor learning, or do they reflect a corrective lick toward the alternative port after an unrewarded lick? An analysis of the second lick on left-error trials or pre-shift right-error trials could help distinguish whether correction licking reflects a general decision change after failed reward, or a motor-command correction specific to post-shift motor learning. The authors should also report the prevalence of CA versus AP trials and clarify whether these trial types are behaviorally distinct.

      We thank the reviewer for highlighting the need to emphasize that CA trials reflect a distinct behavior related to reaching the displaced port.

      By definition, CA trials started as Motor Error trials and thus reflected a corrective lick toward the same port after an unrewarded lick. Almost all first contact licks on Motor Error trials were well outside the distribution of correct left licks both pre- and post-shift (Figure 1 - figure supplement 1B,D), consistent with the interpretation of this first lick as directed toward the right port. Thus, we see no evidence suggesting that CA trials involve a decision change. CA trials are exceedingly rare pre-shift, because Motor Error trials are rare pre-shift (Figure 1I, only ~5% of all right trials).

      With regard to other error types before the shift, most expert-trained mice did not immediately sample the other port with a second lick after an unrewarded lick. They usually either stopped licking immediately or licked the unrewarded port multiple times before switching ports. When port switches occurred pre-shift, timing was highly variable across mice and trials. Even on rewarded trials, some mice would “check” the unrewarded port after consuming the reward, as can be seen in Figure 3I, J. All of these behaviors are clearly distinct from the stereotyped second lick that occurred on CA trials after the shift. We agree that the prevalence of CA and AP trials, as well as the prevalence of immediate port alternation, should be reported, and we have added that information to the article as follows:

      “On Correction Attempted (CA) trials, the first lick made contact with the incorrect port, and the mouse chose to direct a second lick toward the correct port (Figure 5A; prevalence: 54% of motor error trials). We interpreted these licks as a corrective action, because the tongue exit angle shifted further toward the correct target (Figure 5B). Abandoned Port (AP) trials were the same as CA trials, except the mouse either did not make a second attempt or the second lick was directed toward the incorrect port (Figure 5A; prevalence: 46% of motor error trials).”

      (3.4) The classification of pre-shift errors into motor and decision errors is not clear. If error-trial exit angles follow a unimodal distribution (Figure 1- Figure Supplement 1C), then the distinction between motor and decision errors may not be behaviorally well separated. The authors should explain how these categories are validated and whether conclusions depending on this classification are robust to alternative definitions.

      We do not conclude that motor errors and decision errors are distinguishable pre-shift. Pre-shift licks were classified into motor error and decision error categories only to demonstrate that the boundary we established for classifying post-shift licks classifies extremely few (~5%, Figure 1I) pre-shift licks as motor errors. No conclusions were drawn from comparisons between pre-shift licks classified as decision errors and those classified as motor errors. The categorization is defined by the distribution of exit angles pre-shift and validated by the bimodal distribution of exit angles on error trials post-shift. To improve clarity regarding our classification of pre-shift errors, we have added the following to the Results:

      “Exit angles after the shift exhibited a bimodal distribution across error trials (Figure 1G,H; Figure 1—figure supplement 1C,D), supporting this distinction in error type. The frequency of licks classified as motor errors on right-cued trials increased significantly after the shift (median pre-shift 0.06, median post-shift 0.42, p < 0.001; Figure 1I; Figure 1—figure supplement 1C,D), reflecting the new challenge of avoiding the left lickport. In contrast to after the shift, exit angles on error trials before the shift were unimodal (Figure 1—figure supplement 1C). These errors were classified based on the fixed CB in order to demonstrate that very few pre-shift licks qualify as motor errors (Figure 1I), and not to suggest that tongue trajectories before the shift are behaviorally well-separated.”

      (3.5) Figure 1- Figure Supplement 1D, post-shift decision errors: Are these truly decision errors? The lick angles appear similar to those observed before the shift, suggesting that these trials may reflect execution of a "default" or "uncertain" lick trajectory rather than an incorrect choice under the new contingency.

      The post-shift exit angles on right-cued decision error trials (Figure1 - figure supplement 1D, grey) are similar to the lick angles on correct left-cued trials pre-shift (Figure 1 - figure supplement 1A, red) and clearly different from the correct right-cued trials pre-shift (Figure 1 figure supplement 1A, blue). Thus, to the extent that the animal’s intention can be measured from lick trajectory, it was targeting the incorrect (left) port. It is also true that it may still target the previous location of the left port (a “default” left trajectory), but because the decision error makes precise targeting irrelevant to the task outcome (it is easy to reach the left port after the shift), we do not designate it as a joint decision error and motor error. As to whether the deliberative process leading to this action is somehow cognitively distinct from other behaviors typically labeled as decision errors or incorrect choices, we cannot say.

      Minor:

      (3.6) Vocabulary consistency: soma vs somata.

      When data are shown for, or derived from, multiple somata, we use “somata”. When data are shown for an individual soma (such as in a panel with data from a single example soma), we use “soma.” We could not find any use of “somas,” which would indeed be inconsistent.

      (3.7) Figure 1C: I am not sure why the lick trajectories do not depict the tongue exiting the mouse. What time window is shown? Why does it look like the trajectories are shifted to the left?

      We thank the reviewer for identifying this issue. The definition of the location labeled “mouth” was accidentally omitted. The lick trajectories in Figure 1C do depict the tongue tip once it became visible to the cameras. Jaw opening and shifting partly determined the location where the tongue became visible in the videography. These movements varied from mouse to mouse and trial to trial, so exit angle was measured from the approximate midpoint between the temporomandibular joints, which is the grey point in 1C. We have fixed the captions and Methods to precisely define this location. With regard to the appearance of a slight leftward shift in the trajectories, this reflects how the tongue exits the mouth and how the tongue tip curves downward as the tongue approaches the port.

      (3.8) Figure 1- Figure supplement 1: it could ease the comparisons to report population statistics, such as median, from panel A to panel B and D, population statistics from B to D.

      Thank you. We have added these statistics to the Figure 1 - figure supplement 1 caption.

      (3.9) Choice boundary (CB) should be defined in line 100, not 110.

      Thank you. We have fixed this.

      (3.10) Line 109: claim not supported by referenced figure (Figure 1 - Figure Supplement 1). Lick angle histogram to the right port, pre-shift does not overlap substantially with lick angle to the left port, post-shift.

      We thank the reviewer for the opportunity to clarify this. We agree that Figure 1 - figure supplement 1 is not sufficient to support the claim. First, we want to make clear that Figure 1 - figure supplement 1 does not contradict the claim. The new location of the left port can obstruct the tongue during right-cued licks, regardless of the distributions of left licks pre- or post-shift. Second, to confirm that the new location of the left port would obstruct a substantial fraction of pre-shift right-cued lick trajectories, we measured the minimum distance between tongue trajectories and the post-shift location of the left port. Of pre-shift right-cued exit trajectories, 30 +/- 5% came within 1.25 mm – half of the combined tongue width (1.5 mm) and port width (1 mm) – of the port center.

      To make this claim more precise, we have changed the statement as follows:

      “Thus, on right-cued trials, mice continuing to follow the pre-shift motor plan would be biased to more frequently contact the new left port location (Figure 1E,F; 30 +/- 5% of pre-shift trajectories came within a tongue-width of the new location) and receive punishment (i.e., timeout).

      (3.11) Line 113: claim not supported by referenced figure. Figure 1G does not display error trials.

      We have changed the line to refer to “both correct and error trials”, such that reference to Figure 1G is also appropriate.

      (3.12) Figure 2 - Figure Supplementary 3 & method: how is noise estimated?

      Thank you. The following has been added to the Methods:

      “For Figure 2 - figure supplement 3, noise was estimated as the square-root of the geometric mean of the Welch power spectrum in a high-frequency band (0.25–0.5 times the frame rate; Giovannucci et al., 2019).”

      (3.13) Figure 2D: Was imaging during the shift epoch always performed in dendrites? If so, could the imaging schedule bias comparisons between dendritic and somatic activity during learning, especially given that mice show behavioral learning between early and late post-shift sessions (Figure 1J)?

      No, imaging during the shift was not always performed in the dendrites. The following has been added to the Methods to make clear that the post-shift data reflect dendritic and somatic imaging conducted on the day of the shift with roughly similar frequency:

      “For Figure 5 and Figure 6, which make comparisons between dendritic and somatic activity during the post-shift period, 67% of animals providing somatic data (4 of 6 mice) underwent somatic imaging on the day of the shift and 80% of mice providing dendritic data (8 of 10 mice) underwent dendritic imaging on the day of the shift.”

      (3.14) Lines 163-164, "we observed that the onset of tuft activity was consistently time-locked to the GO cue (vertical green line; Figure 3B). This was in contrast to somatic activity, which had more variable timing (Figure 3E)." The authors cite panels B and E in support of this point, but these appear to be example ROIs. It would be helpful to clarify how representative these examples are, since the corresponding population summaries in panels G and H do not make the effect immediately apparent.

      These examples are representative, as supported by the population summary of activity time-locked to the GO-cue versus port contact in Figure 4B.

      (3.15) Figure 4B: It could be useful to add the lick traces here as well. To allow the reader to have an idea of contact timing with respect to the Go cue and compare the sustain response with the licking pattern.

      We understand how this could be helpful. However, since these exact traces are already present in Figure 3I,J, we think that adding them to Figure 4 is unnecessary and would add complexity to an already very busy figure.

      (3.16) Figure 5E: Why are the imaging sessions labeled 0 and +1 rather than 0 and +2? Are the dendritic and somatic imaging not alternated?

      Yes, imaging was not alternated for 3 of the 22 mice. We have clarified this in the Methods, as follows:

      “Somatic and dendritic imaging sessions alternated every other day (19 of 22 mice), except for 3 mice in which only one compartment was imaged daily (dendrite-only: 2 mice, soma-only: 1 mouse). The exceptions were due to brain curvature or the angle of the coverslip with respect to the brain, such that only one compartment could be imaged and the other compartment was underneath skull regrowth or dural thickening that made high-quality imaging impossible.”

      (3.17) Figure 6C, legend: I suppose the authors meant "remapping", not "Post-shit SI distribution" for the description of the right column.

      Thank you. We have fixed this label.

    1. eLife Assessment

      This study provides a useful analysis of the effects of arginine vasopressin (AVP) on islet cells in pancreatic tissue slices, using sophisticated spatiotemporal calcium recordings to show that AVP affects α and β cells differently depending on glucose concentration. The calcium imaging, analytical approaches, and V1b receptor-targeted peptide ligands are strengths of the work. However, the reviewers were concerned that the proposed mechanistic model is not sufficiently supported by the data. Characterisation of β-cell responses remains incomplete, and potential off-target effects and limited receptor specificity raise alternative explanations, including indirect effects mediated through α cells. The study would have been strengthened by signalling pathway analyses, genetic validation (e.g. β-cell-specific V1bR deletion), or selective V1b receptor silencing. The RNAscope data included in the revision indicate broader expression patterns but do not clearly establish receptor localisation within specific endocrine populations.

    2. Reviewer #1 (Public review):

      Summary:

      The paper investigates how AVP modulates pancreatic alpha and beta cell activity using acute mouse pancreatic tissue slices, calcium imaging, hormone secretion assays, RNAscope, and newly synthesized receptor-selective ligands. The Authors report that AVP regulates islet cell activity in a glucose- and state-dependent manner, with maximal effects occurring within physiological AVP concentrations and a bell-shaped concentration-response profile. They conclude that V1b receptors are the principal mediators of these effects and propose that IP3 receptor-dependent signaling underlies the observed nonlinear responses.

      Strengths:

      The use of fresh pancreatic tissue slices preserves islet architecture and cell-cell interactions, providing a physiologically relevant experimental model compared with isolated islets or immortalized cell lines.

      The combination of live calcium imaging, hormone secretion measurements, RNAscope, and pharmacological characterization of newly synthesized receptor-selective ligands represents a technically comprehensive experimental approach that addresses AVP signaling from multiple complementary perspectives.

      Weaknesses:

      (1) The central mechanistic model of the manuscript is not supported by the experimental data. Although the Authors repeatedly attribute the observed bell-shaped responses to IP3 Receptor activation and inactivation, no direct mechanistic evidence is provided to implicate IP3 receptors. Experiments assessing IP3 receptor function using genetic manipulation and direct measurements of IP3 signaling are necessary before such mechanistic conclusions can be drawn.

      (2) The Authors should directly demonstrate V1b receptor expression in β cells using complementary approaches, since the RNAscope data indicate broader expression but do not convincingly establish receptor localization within specific endocrine populations.

      (3) In my opinion, the central conclusion that V1b receptors are the predominant mediators of the observed effects is insufficiently supported because definitive loss-of-function experiments are lacking. Genetic deletion or selective silencing of V1b receptors should be provided to validate the proposed mechanism.

      (4) The heterogeneous responses observed among islets substantially weaken the proposed mechanistic model. Data should be provided to identify the determinants responsible for activation, absence of response, or inhibition in individual islets.

      (5) Please explain why the marked changes in alpha-cell calcium activity were not accompanied by corresponding alterations in glucagon secretion. This apparent discrepancy requires additional experimental evidence.

      (6) The Authors need to provide stronger evidence linking the observed calcium dynamics with insulin secretion, since calcium measurements alone cannot establish the proposed functional consequences.

      (7) Proper assays should be provided to assess whether the newly synthesized ligands exhibit comparable selectivity and efficacy at murine receptors rather than relying primarily on pharmacological characterization performed using human receptor-expressing cell lines.

      (8) The proposed absence of V1a receptor involvement is based primarily on pharmacological inhibition. Independent experimental approaches should be provided to exclude a contribution of this receptor subtype.

      (9) They must provide additional quantitative analyses demonstrating that the reported bell-shaped concentration-response relationship is robust across individual experiments rather than reflecting substantial biological variability.

      (10) The Authors should include experiments evaluating endogenous AVP signaling under more physiological conditions instead of relying predominantly on exogenous agonist administration.

      (11) I believe the role of forskolin deserves further clarification because many conclusions were obtained under cAMP-permissive conditions that may substantially influence AVP responses. Additional experiments without pharmacological cAMP stimulation should be presented.

      (12) Please clarify how beta cells and alpha cells were identified exclusively from functional activity patterns during calcium imaging and provide independent validation of cell identity within the analyzed recordings.

      (13) In my opinion, the manuscript relies heavily on changes in intracellular calcium activity as a surrogate for endocrine function, whereas the secretion data do not consistently support the proposed functional conclusions. Additional evidence is needed to establish a direct relationship between the observed calcium dynamics and hormone release.

      (14) The Authors should better reconcile their findings with previous reports showing minimal or absent AVP receptor expression in β cells and explain how the current data resolve these discrepancies rather than adding another possible interpretation.

    3. Reviewer #2 (Public review):

      Summary:

      In this paper Drs. Kercmar, Murko and Bombek make a series of observations related to the role of AVP in pancreatic islets. They use the pancreatic slice preparation that their group is well known for. The observations on the slide physiology are technically impressive. However, I am not convinced by the conclusions of this manuscript for a number of reasons. At the core of my concern is perhaps that this manuscript appears to be motivated to resolve 'controversies' surrounding the actions of AVP on insulin and glucagon secretion. This manuscript adds more observations, but these do not move the field forward in improving or solidifying our mechanistic understanding of AVP actions on islets. A major claim in this manuscript is the beta cell expression of the V1b Receptor for AVP, but the evidence presented in this paper fall short of supporting this claim. Observations on the activation of calcium in alpha cells via V1b receptor align with prior observations to this effect and can explain the effects of beta cell calcium and insulin secretion better than an explanation where beta cells express functional V1BR, for which direct evidence is lacking.

      I have focused my main concerns below. I hope the authors will consider these suggestions carefully - please be assured that they were made with the intent to support the authors and increase the impact of this work.

      Strengths:

      The main strength of this paper is the technical sophistication of the approach and the analysis and representation of the calcium traces from alpha and beta cells.

      Weaknesses:

      (1) There are excellent data that indicate that the actions of AVP are mediated via V1bR on alpha cells and that V1bR is 1) not expressed by beta cells and 2) does not activate beta cell calcium at all at 10 nM - which is the same concentration used in this paper (Figure 4G) for peak alpha cell Ca2+ activation (see https://doi.org/10.1016/j.cmet.2017.03.017; cited as ref 30 in the current manuscript). Any published stimulatory actions of AVP on insulin secretion can be explained by the potentiating effects of glucagon, released in response to AVP stimulation of alpha cells.

      (2) The RNAscope data offered in the revision as a second line of evidence for the expression of the V1bR in beta cells do not convince. I applaud the authors for trying as these are hard experiments to do well, as evidenced from the Gcg RNAscope signal that is not at all concentrated in the islet periphery, and in fact both color puncta occur outside of the islet at similar density. Absent a convincing concentration of Gcg signal (which is a very abundant transcript in alpha cells), it is hard to depend on these results. They certainly do not substitute experiments to determine cell autonomous activation of isolated beta cells by AVP. Claim of beta cell expression of V1br, require a more direct demonstration by staining (if appropriate antibodies exist), by beta cell-specific deletion of V1br, or by documenting the direct calcium activation in isolated beta cells in the absence of alpha cells. This should include a demonstration of Gaq-dependence in isolated beta cells.

      (3) We know from bulk RNAseq data on purified alpha, beta, and delta cells from both the Huising and Gribble groups that there is no expression of V2a. I will point you to the data from the Huising lab website published almost a decade ago (http://dx.doi.org/10.1016/j.molmet.2016.04.007) - which is publicly available and can be used to generate figures (https://huisinglab.com/data-ghrelin-ucsc/index.html). They indicate the absence of expression of not only AVP2 receptors anywhere in the islet - but the lack of expression of V1bra, V1brb, and Oxtr in beta cells. These AVP/OXT receptor expression data are largely and helpfully confirmed by the efforts in this paper that involved the generation of the V1aR agonist and V2R antagonist.

      (4) Importantly, the lack of V1br from beta cells does not invalidate observations that AVP affects calcium in beta cells, but it does indicate that these effects are mediated 1) indirectly, downstream of alpha cell V1br or 2) via an unknown off-target mechanism (less likely). The different peak efficacies in Figure 4G would also suggest they are not mediated by the same receptor. The recent work by Huixia Ren and colleagues (PMID: 41916313) that demonstrates that glucagon accelerates the frequency of beta cell calcium is in line with such a scenario.

      (5) The use of forskolin across almost all traces complicates the interpretation of the results. The design does not account for the elevation of cAMP in alpha cells and subsequent release of glucagon - particularly upon co-stimulation with AVP which permits glucagon release by activating a calcium response in alpha cells. This glucagon then could activate beta cells. If resolving the mechanism of action is the goal, often less is more. The activation of Gaq-mediated calcium is not cAMP dependent (although the downstream hormone secretion clearly often is). As was shown, AVP does not activate calcium in beta cells in the absence of cAMP. The experiments should have been completed in the absence of cAMP/forskolin, which would likely have had different outcomes on the beta cell responses and to the hormone secretion.

      (6) It is motivated by a desire to 'study the AVP dependence of both alpha and beta cells at the same time'. As best as I can determine, the design choice to conduct most studies under sustained forskolin stimulation is related to the permissive actions of AVP on hormone secretion in response to cAMP-generating stimuli. The permissive actions by AVP that are cited are on hormone secretion - which in many cell types requires activation of both calcium and cAMP signaling. Whether the activation of V1br and subsequent calcium responsive is permitted by cAMP is unclear. I believe the argument the authors are making here is that the activation of beta cell calcium by AVP is permitted by forskolin. i.e. the cAMP stimulated by it in beta cells.

      (7) Figure 9 suggests a pharmacological activation of beta cell V1bR in the low pM range. How do the authors reconcile this compare with the apparent absence of an effect of AVP stimulation at low pM to low nM doses in beta cells (Figure 5A). I note that there are changes over time with sustained beta cell stimulation with 8 mM glucose, but these changes are relatively subtle, gradual and quite likely represent the progression of calcium behaviors that would have occurred under sustained glucose irrespective of these very low AVP concentrations. I will note that the Kd of the V1bR for AVP is around 1 nM, with tracer displacement starting around 100 pM according to the data in figure 6B, which is hard to reconcile with changes in beta cell calcium by AVP doses that start 10-100-fold lower than this dose at 1 and 10 pM (Figure 9).

    4. Reviewer #3 (Public review):

      Summary:

      This work aims to better understand the role of arginine vasopressin (AVP) in the control of islet hormone secretion. This builds on previous literature in this area reporting on the actions of AVP to stimulate islet hormones. The gap in literature being addressed by these studies is primarily focused on the glucose-dependency of AVP on both insulin and glucagon secretion. A secondary objective is to explore the role of individual receptors with the use of newly generated peptides and existing tools. The methods include the use of Ca2+ imaging in pancreas slices from mice, with additional outcomes including insulin secretion in some areas. The conclusions presented are that AVP acts through V1b receptors in both alpha- and beta-cells, that this activity occurs in the high cAMP environment, and is glucose dependent.

      Strengths:

      The area of research is emerging with plenty of room for new contributions. The concept of AVP stimulating islet hormone secretion is important and deserving of further insight. The use of pancreas tissue to image primary cells makes the experiments physiologically relevant. The advancement of novel tools in this area should be helpful to other groups investigating the actions of AVP.

      Comments on revised version:

      Overall, the authors have modified their conclusions to more accurately capture the results of this manuscript. They also add the significant limitations outlined in the review process to the Discussion. With the addition of new data, however, a few concerns have emerged.

      (1) The rational for showing somatostatin staining in Figure 1 a is unclear. It also does not appear to be the same region of interest as in panel B.

      (2) It is difficult to assess the success of the RNAscope with the representative image used in Figure 1b. It is surprising how low the Gcg signal is in this image, suggesting some optimization is required. Additionally, how many mice were used (biological replicates) and how many V1b receptor+ and Gcg+ cells per mouse were quantified for the RNAscope images? This must be indicated in the methods section and should be sufficiently powered to make a conclusion.

      (3) While the addition of insulin and glucagon secretions with AVP ramp provide a functional output to the calcium imaging, it is unclear why the measurements are sometimes log10 transformed (Figure 5 K and L) but not always (Figure 5E). It is difficult to interpret negative glucagon values. What is the functional output of the dose-dependent calcium response to AVP in alpha cells if it is not glucagon?

      (4) Finally, the highlights section has not been refined to the revised interpretations of the manuscript.

    5. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents a useful finding on the effects of arginine vasopressin (AVP) on islet cells in pancreatic tissue slices, using technically sophisticated spatio-temporal calcium recordings to confirm that AVP influences α and β cells differently depending on glucose concentrations. While the study's methods - particularly the calcium imaging techniques and peptide ligand design targeting V1b receptors - are strong, the reviewers were concerned about several aspects of the experimental design. However, the results on βcell responses are incomplete and insufficient to support the manuscript's claims, especially due to the high variability of islet responses and lack of mechanistic and functional (hormone release) data. There are also concerns about the possibility of off-target effects and incomplete receptor specificity, noting that the study would have been significantly strengthened by inclusion of signaling pathway interrogation, hormone output assays, genetic validation (e.g., β cell-specific deletion of V1br), and receptor localization, although the work will still be of interest to researchers studying islet physiology in the context of health and diabetes.

      We sincerely thank the reviewers and editors for their thorough evaluation of our manuscript and their recognition of its technical strengths, including the advanced spatio-temporal calcium imaging and the rational design of selective V1b receptor ligands. We appreciate their acknowledgement of the study’s relevance for understanding AVP effects in a physiologically intact islet context and their positive assessment of our methodological rigour and innovation. The reviewers’ constructive feedback has helped us clarify the boundaries and intent of our study, which focuses on the glucose- and context-dependent modulation of α- and β-cell activity, rather than exhaustive molecular dissection.

      While the reviewers rightly emphasize the importance of receptor specificity and downstream signaling validation, we respectfully suggest that some of their concerns may reflect a lingering bias toward reductionist frameworks. Our interpretation is rooted in the emerging understanding that β-cell behaviour is largely defined by dynamic intercellular interactions within the islet collective, rather than by static gene expression or receptor localization alone (Jin et al., 2025; Korošak et al., 2021; Rutter et al., 2024). Recent studies have demonstrated that roles such as “leader” or “hub” β cells are transient and emergent, governed more by timing, environment, and local network structure than by fixed molecular identity (Postic et al., 2023; Gosak et al., 2018).

      This has profound implications for how we interpret cell responsiveness to agents like AVP: what appears as biological variability may in fact reflect context-sensitive transitions within a non-linear, self-organizing system (Stožer et al., 2021). Hence, we chose to focus on functional collective dynamics using intact pancreatic slices, rather than isolated cell models which fail to preserve the essential network architecture of islets. Although the addition of genetic models or isolated receptor measurements would strengthen receptor-specific conclusions, we argue that such approaches alone cannot resolve the physiological complexity of a system where function arises from cell–cell communication and spatiotemporal context.

      Indeed, the lack of direct correlation between receptor transcript abundance and functional outcomes has been noted in prior studies, reinforcing the view that function cannot be strictly predicted by molecular presence (Rutter et al., 2024). As articulated in our manuscript, the islet behaves as a sensory collective (Fancher & Mugler, 2017), where emergent patterns— not static cell identity—determine behaviour. This perspective aligns with broader shifts in biology away from strict genetic determinism toward causal emergence and collective agency (Ball, 2023; Levin, 2021).

      We therefore believe our study contributes not only new pharmacological insights but also a conceptual reframing of how AVP responses should be interpreted in a complex organ like the pancreas. We have added new data addressing reviewer suggestions—such as glucagon secretion assays, clarifications on the role of forskolin, and an analysis of event timing—that further support our conclusions. We also expanded the discussion on how islet variability is functionally meaningful, not just noise, and explained why β-cell responses to AVP must be interpreted within this probabilistic framework.

      We agree that future work should include receptor-specific knockouts and more direct signaling pathway assays, but these would need to be designed with careful consideration of the islet’s dynamic topology and the emergent nature of β-cell roles. In this light, we see our study not as the final word, but as a necessary systems-level foundation for more targeted interventions. We thank the reviewers again for their careful critiques and hope that our response clarifies both the rationale and scope of our work. Our revisions aim to enhance the paper’s clarity while maintaining its commitment to an integrative, physiology-rooted approach.

      We thank the reviewers and editors for their thoughtful and constructive assessment of our work. We are especially grateful for their recognition of the study’s technical strengths, including the use of spatio-temporal calcium imaging in intact pancreatic tissue and the strategic development of receptor-selective peptide ligands. We also appreciate their acknowledgement that our study contributes to the understanding of glucose-dependent AVP effects in islet physiology. The reviewers’ concerns regarding variability, receptor specificity, and functional validation helped us further clarify the scope and context of our study.

      We respectfully submit that some reservations stem from a reductionist framing that may not fully account for the collective behaviour of islets. As we and others have shown, β-cell function arises from emergent, self-organizing network dynamics, not just from static gene expression or receptor abundance (Jin et al., 2025; Korošak et al., 2021; Postic et al., 2023). In this view, pharmacological heterogeneity across islets is not simply noise or experimental inconsistency, but a signature of dynamic attractor states within the islet network (Stožer et al., 2021). Because an islet functions as a coupled system, most response variability originates from its emergent collective behavior, which eclipses variability in receptor expression or metabolic state.

      For this reason, even single-islet receptor quantification or ATP measurements would provide limited explanatory power: it is the state of the network—not absolute receptor levels—that determines whether a perturbation elicits activation or inhibition. As we illustrate in our graphical abstract, a single islet tested repeatedly under identical glucose conditions can yield divergent responses, simply because it occupies different dynamic states. These findings are in line with systems biology and network science approaches, which have revealed that cell function, especially in the β-cell collective, cannot be fully understood through reductionist parameters alone (Gosak et al., 2018; Ball, 2023).

      We have included glucagon secretion assays and new analyses to address key reviewer suggestions. Still, we chose not to pursue extensive knockouts or cAMP imaging, as these would require a different experimental scope and could risk disrupting the very dynamics we aim to understand. Likewise, while direct measurements of V1bR or IP3R expression would add molecular detail, they are not definitive without network context. The bell-shaped AVP dose-response curve and its explanation through IP3R inactivation are supported by prior studies; we invoke this mechanism not speculatively, but because it provides the most parsimonious explanation for the glucose-dependent shift in β-cell responsiveness.

      We also clarify that our study does not aim to resolve every mechanistic detail, but rather to offer a systems-level insight into how AVP modulates islet dynamics across varying glucose and cAMP contexts. The implications extend beyond AVP pharmacology, suggesting that perturbations to β-cell function must be understood within a probabilistic, state-dependent framework (Fancher & Mugler, 2017). This resonates with emerging concepts in cell physiology that emphasize causal emergence and local agency over static molecular determinism (Levin, 2021; Rutter et al., 2024).

      In summary, we see our work as part of a necessary shift in perspective—from linear receptor-function models to context-sensitive dynamic systems. We are grateful for the opportunity to revise our manuscript in response to insightful feedback and hope our clarifications and new data will strengthen its impact for the islet research community.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors confirmed earlier findings that AVP influences α and β cells differently, depending on glucose concentrations. At substimulatory glucose levels, AVP combined with forskolin - an activator of cAMP -did not significantly stimulate β cells, although it did activate α cells. Once glucose was raised to stimulatory levels, β cells became active, and α cell activity declined, indicating glucose's suppressive effect on α cells and permissive effect on β cells. Under physiological glucose levels (8-9 mM), forskolin enhanced β-cell calcium oscillations, and AVP further modulated this activity. However, AVP's effect on β cells was variable across islets and did not significantly alter AUC measurements (a combined indicator of oscillation frequency and duration). In α cells, forskolin and AVP led to increased activity even at high glucose levels, suggesting that α cells remain responsive despite expected suppression by insulin and glucose.

      Experiments with physiological concentrations of epinephrine suggest that AVP does not operate via Gs-coupled V2 receptors in β cells, as AVP could not counteract epinephrine's inhibitory effects. Instead, epinephrine reduced β cell activity while increasing α cell activity through different G-proteincoupled mechanisms. These results emphasize that AVP can potentiate αcell activation and has a nuanced, context-dependent effect on β cells.

      The most robust activation of both α and β cells by AVP occurred within its physiological osmo-regulatory range (~10-100 pM), confirming that AVP exerts bell-shaped concentration-dependent effects on β cells. At low concentrations, AVP increased β cell calcium oscillation frequency and reduced "halfwidths"; high concentrations eventually suppressed β cell activity, mimicking the muscarinic signaling. In α cells, higher AVP concentrations were required for peak activation, which was not blunted by receptor inactivation within physiological ranges.

      Attempting to further dissect the role of specific AVP receptors, the authors designed and tested peptide ligands selective for V1b receptors. These included a selective V1b agonist; a V1b agonist with antagonist properties at V1a and oxytocin receptors; and a selective V1a antagonist. In pancreatic slices, these peptides seem to replicate AVP's effects on Ca<sup>2+</sup> signaling, although responses were highly variable, with some islets showing increased activity and others no change or suppression. The variability was partly attributed to islet-specific baseline activity, and the authors conclude that AVP and V1b receptor agonists can modulate β cell activity in a statedependent manner, stimulating insulin secretion in quiescent cells and inhibiting it in already active cells.

      We applaud the reviewer to capture the essence of work in their introduction.

      Strengths:

      Overall, the study is technically advanced and provides useful pharmacological tools. However, the conclusions are limited by a lack of direct mechanistic and functional data. Addressing these gaps through a combination of signaling pathway interrogation, functional hormone output, genetic validation, and receptor localization would strengthen the conclusions and reduce the current (interpretive) ambiguity.

      Thank you!

      Weaknesses:

      (1) The study is entirely based on pharmacological tools. Without genetic models, off-target effects or incomplete specificity of the peptides cannot be fully ruled out.

      We partially agree with this comment and acknowledge that genetic models would provide a valuable complementary approach to address possible off-target effects or incomplete peptide specificity. However, genetic models also have important limitations, particularly when the aim is to resolve subtle, population-level physiological differences in beta cell activity. We therefore used pharmacological tools at different concentrations to test whether the observed effects were concentration-dependent and consistent with the expected receptor-mediated actions. An advantage of the pancreatic slice preparation is that it preserves much of the native tissue environment and allows pharmacological manipulation within concentration ranges closer to in vivo efficacy, thereby reducing the likelihood of nonspecific effects. To compensate for the lack of genetic models, we now emphasize the collective activity analysis as an additional strength of the study and have clarified this limitation in the revised manuscript.

      (2) Despite multiple claims about β cell activation or inhibition, the functional output - insulin secretion - is weakly assessed, and only in limited conditions. This aspect makes it very hard to correlate calcium dynamics with physiological outcomes.

      We agree that the functional output needed stronger support and have therefore expanded the hormone secretion experiments. While the effects of AVP and its analogues were tested during a stable plateau phase in the Ca<sup>2+</sup> imaging experiments, this phase provides only a narrow dynamic range for insulin release measurements in mouse slices. We therefore added a sequence of stimulations on the same slices, using 8 mM glucose and 500 nM forskolin, with glucose lowered to a non-stimulatory range between different AVP concentrations. These new experiments better define how AVP-dependent changes in Ca<sup>2+</sup> dynamics translate into insulin secretion under conditions with a broader secretory dynamic range. The new insulin and glucagon secretion data have now been added to the manuscript as Figure 5, and the text has been revised accordingly.

      (3) Insulin and glucagon secretion assays should be provided; the authors should measure hormone release in parallel with Ca2+ imaging, using perifusion assays, especially during AVP ramp and peptide ligand applications.

      We added insulin and glucagon secretion assays for AVP ramp to Figure 5.

      Additionally, there is no standardization of the metabolic state of islets. The authors should consider measuring islet NAD(P)H autofluorescence or mitochondrial potential (e.g., using TMRE) to control for metabolic variability that may affect responsiveness.

      We agree that standardization of the metabolic state of the islets would further strengthen the interpretation of the responsiveness data. We attempted to address this experimentally, but the results were inconclusive and therefore not included in the manuscript. Based on our previous unpublished observations, NAD(P)H levels appear to be significantly higher and less variable in islets within tissue slices than in isolated islets, suggesting that the slice preparation may better preserve the native metabolic state. However, we acknowledge that this remains an important limitation and we now indicate that in the manuscript. Additional experiments will be required to establish a robust and standardized approach, for example by combining NAD(P)H autofluorescence and/or mitochondrial potential measurements with Ca<sup>2+</sup> imaging.

      (4) There is a high degree of variability in response to AVP and V1b agonists across islets (activation, no effect, inhibition). Surprisingly, the authors do not fully explore the cause of this heterogeneity (whether it is due to receptor expression differences, metabolic state, experimental variability, or other conditions).

      This is a well-taken point and has indeed been one of the major bottlenecks in interpreting the results of this study. We agree that the variability in responses to AVP and V1b agonists may reflect several factors, including receptor expression, metabolic state, experimental conditions, and differences in the functional state of individual islets. However, our data also suggest that the beta cell population within an islet should be considered as a dynamic, non-linear system, in which even small differences in initial conditions or collective state can result in qualitatively different outcomes, including activation, no apparent effect, or inhibition. In this framework, the response to AVP is not determined by receptor expression alone, but by the current physiological context of the islet network. This is also why we believe that pharmacological tests are most informative when interpreted within a defined functional state rather than as isolated receptor-specific readouts. As indicated in the graphical abstract, apparently similar islets may occupy different dynamic states and therefore respond differently to the same Gq/PLC/IP3R stimulus. We have now expanded the discussion to make this interpretation more explicit and to acknowledge that receptor expression, metabolic variability, and experimental factors remain possible contributors that will require further targeted studies.

      The following text has been added to expand the discussion:

      “The heterogeneous responses observed across different islets, where some showed increased activity while others showed no detectable change or inhibition, could intuitively be attributed to variability in V1b receptor expression or signaling capacity among β cells. Such an explanation would be consistent with differences in receptor density, coupling efficiency to Gq proteins, or downstream signaling components such as PLC or IP<sub>3</sub> receptors. However, our data suggest that receptor-level variability alone is unlikely to fully explain the observed response spectrum, and that the current functional state of the islet collective must also be considered. The islet behaves as a non-linear dynamic system in which the same molecular perturbation can produce different functional outcomes depending on the current state of the β-cell collective. In such systems, cells or cell populations do not occupy a single deterministic activity state, but rather move within a landscape of possible states, with perturbations shifting the probability distribution of transitions between them. This concept is well established in dynamical systems approaches to biological cell-state transitions, where attractor landscapes, noise, and signaling inputs determine the probability of moving between alternative functional states rather than enforcing a single fixed output.

      In this framework, AVP and V1b receptor-selective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that βcell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone. Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (5) There is no validation of V1b receptor expression at the protein or mRNA level in α or β cells using in situ hybridization, immunohistochemistry, or spatial transcriptomics.

      We agree with the reviewer that spatial validation of V1b receptor expression is important for interpreting the cellular targets of AVP signaling in the islet. We have therefore added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      (6) AVP effects are described in terms of permissive or antagonistic effects on cAMP (especially in relation to epinephrine), but direct measurements of cAMP in α and β cells are not shown, weakening these conclusions. The authors should use Epac-based cAMP FRET sensors in α and β cells to monitor the interaction between AVP, forskolin, and epinephrine more conclusively.

      We agree that direct measurements of cAMP dynamics in alpha and beta cells would provide a more conclusive assessment of the interaction between AVP, forskolin, and epinephrine signaling. We attempted to address this experimentally; however, within the time domain of the Ca<sup>2+</sup> oscillations analyzed here, the temporal resolution and robustness of currently available cAMP readouts were not sufficient to resolve these interactions reliably. Even at slower time scales, cAMP sensor signals can be difficult to interpret quantitatively and may be overinterpreted if not tightly linked to the functional readout. We have therefore moderated the wording of the manuscript and now describe the proposed permissive or antagonistic interaction between AVP/V1b and cAMP-dependent signaling as an interpretation supported by the pharmacological Ca<sup>2+</sup> response patterns, rather than as a directly demonstrated cAMP mechanism. We now explicitly acknowledge in the limitations that most experiments were performed under cAMP-permissive conditions, which increases sensitivity for detecting AVP-dependent modulation but complicates the separation of direct beta cell effects from intra-islet interactions. Future studies using optimized cell-type-specific Epac-based sensors will be required to resolve this interaction.

      (7) Single-islet transcriptomics or proteomics (also to clarify variability) should be provided to analyze receptor expression variability across islets to correlate with response phenotypes (activation vs inhibition). Alternatively, the authors could perform calcium imaging with simultaneous insulin granule tracking or ATP levels to assess islet functional states.

      We agree that single-islet transcriptomics, proteomics, or simultaneous metabolic readouts could provide useful complementary information, particularly for describing molecular variability across islets. However, we do not think that differences in receptor expression or ATP levels alone are sufficient to explain the diversity of response phenotypes observed here. Our interpretation is that the beta cell population behaves as a collective dynamic system, in which the same input can lead to different outcomes depending on the current state of the network and its local physiological context. In such a system, AVP/V1b signaling does not necessarily impose a single deterministic response, but changes the probability distribution of accessible states, including activation, inhibition, or no detectable response. Theoretically and partially confirmed by the preliminary data, even the same islet exposed repeatedly under apparently identical conditions could be expected to display different responses if it occupies a different position within this dynamic state space at the time of stimulation. This concept is summarized in the graphical abstract and is central to our interpretation of the pharmacological data. We have therefore clarified in the Discussion that receptor expression, ATP levels, and other molecular parameters may modulate the response landscape, but are unlikely to fully define the observed functional phenotype without considering the collective dynamics of the islet.

      Added to Discussion section: “In this framework, AVP and V1b receptorselective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that β cell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone (63). Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (8) While the study implies AVP acts through V1b receptors on β cells, the signaling downstream (e.g., PLC activation, IP3R isoforms involved) is simply inferred but not directly shown.

      We agree that downstream signaling was not directly resolved at the level of PLC activation or specific IP3R isoforms. However, we did not infer Gq/PLC/IP3R involvement solely from AVP pharmacology, but used ACh as an independent Gq-coupled receptor reference stimulus in the same pancreatic slice preparation. With ACh concentration ramps, we could reproduce both activation and inactivation patterns observed with AVP/V1b stimulation, supporting the interpretation that these responses arise from modulation of the Gq-dependent Ca<sup>2+</sup> signaling axis.

      In addition, in prelilminary expriments we could observe that inhibition of Gq activity with YM254890, as well as interference with IP3R-dependent signaling using Xestospongin C, diminished the response, although not completely. This incomplete suppression is important, because it suggests that beta cell Ca<sup>2+</sup> homeostasis and collective islet activity are not controlled by a single linear pathway, but by partially redundant and context-dependent mechanisms. We have therefore revised the manuscript to state more cautiously that our data support the involvement of Gq/PLC/IP3R-dependent signaling, while acknowledging that direct measurements of PLC activity and IP3R isoform-specific contributions remain outside the scope of the present study.

      (9) The interpretation that IP3R inactivation (mentioned in the title!) underlies the bell-shaped AVP effect is just hypothetical, without direct measurements. Assays in β (and/or α)-cell-specific V1b KO mice and IP3R KO mice must be provided to support these speculations.

      We agree that the involvement of IP3R-dependent signaling should be stated with appropriate caution. However, the concept of IP3R inactivation as a mechanism contributing to bell-shaped Gq-dependent Ca<sup>2+</sup> responses is not purely hypothetical, since IP3R inactivation has been directly demonstrated in previous studies and provides a parsimonious explanation for the shift from activation to suppression at higher AVP concentrations. In the present study, this interpretation is further supported by the glucose dependence of the AVP concentration-response relationship, where different stimulatory glucose conditions shift the apparent efficacy peak.

      We also agree that cell-specific V1b receptor and IP3R knockout experiments would be valuable future approaches. In this respect, we have obtained preliminary results from a small sample of IP3R triple-knockout mice, which cannot yet be fully included because they are part of an ongoing collaboration. In these experiments, supraphysiological AVP concentrations did not produce the IP3R-like beta-cell response pattern observed in controls, namely reduced halfwidth and increased frequency, whereas alpha cell stimulation was preserved similarly to WT slices.

      At the same time, we believe that definitive knockout experiments must be carefully designed, because the beta cell population behaves as a dynamic collective system in which the response to AVP depends on the current functional state of the islet, glucose context, and intercellular coupling. We therefore now present IP3R inactivation as a strongly supported mechanistic interpretation rather than as a directly proven mechanism in this study, and we explicitly acknowledge that cell-specific V1b and IP3R genetic models will be required to fully resolve this pathway.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Drs. Kercmar, Murko, and Bombek make a series of observations related to the role of AVP in pancreatic islets. They use the pancreatic slice preparation that their group is well known for. The observations on the slide physiology are technically impressive. However, I am not convinced by the conclusions of this manuscript for a number of reasons. At the core of my concern is perhaps that this manuscript appears to be motivated to resolve 'controversies' surrounding the actions of AVP on insulin and glucagon secretion. This manuscript adds more observations, but these do not move the field forward in improving or solidifying our mechanistic understanding of AVP actions on islets. A major claim in this manuscript is the beta cell expression of the V1b Receptor for AVP, but the evidence presented in this paper falls short of supporting this claim.

      Observations on the activation of calcium in alpha cells via V1b receptor align with prior observations of this effect.

      I have focused my main concerns below. I hope the authors will consider these suggestions carefully - please be assured that they were made with the intent to support the authors and increase the impact of this work.

      We thank the reviewer for their detailed input and support to increase the impact of our work and our understanding of important cellular processes overall. We have considered their suggestions carefully to further expand the strenghts of our approach and analysis.

      Strengths:

      The main strength of this paper is the technical sophistication of the approach and the analysis and representation of the calcium traces from alpha and beta cells.

      Thank you!

      Weaknesses:

      (1) The introduction is long and summarizes a substantive body of literature on AVP actions on insulin secretion in vivo. There are a number of possible explanations for these observations that do not directly target islet cells. If the goal is to resolve the mechanistic basis of AVP action on alpha and beta cells, the more limited number of papers that describe direct islet effects is more helpful. There are excellent data that indicate that the actions of AVP are mediated via V1bR on alpha cells and that V1bR is a) not expressed by beta cells and b) does not activate beta cell calcium at all at 10 nM - which is the same concentration used in this paper (Figure 4G) for peak alpha cell Ca2+ activation (see https://doi.org/10.1016/j.cmet.2017.03.017; cited as ref 30 in the current manuscript).

      We thank the reviewer for this important comment and agree that the literature on AVP actions in vivo is complex, with several possible sites of action outside the islet. We have therefore revised the Introduction to make the rationale more focused and to better separate systemic effects of AVP from studies addressing direct actions on pancreatic islet cells. At the same time, we chose not to restrict the Introduction only to the alpha cell V1bR literature, because one of the aims of the manuscript is precisely to address why AVP effects on insulin secretion have remained difficult to interpret across experimental contexts.

      Our results fully confirm a central aspect of the study cited by the reviewer, namely that V1bR activation robustly stimulates alpha cell Ca<sup>2+</sup> activity under non-stimulatory glucose conditions, and that 10 nM AVP does not produce a uniform activation of beta cell Ca<sup>2+</sup> activity. In fact, in a substantial fraction of beta cell populations, 10 nM AVP failed to activate oscillations, consistent with the view that alpha cells are the more sensitive and more direct cellular target of AVP/V1bR signaling. However, we do not think that the available transcriptomic evidence is sufficient to categorically exclude V1bR expression or functional relevance in beta cells. Re-analysis of the published dataset, together with more recent datasets and our newly added RNAscope data, supports a higher relative expression of V1bR transcripts in alpha than in beta cells, but does not justify treating beta (or non-alpha) cell expression as absent.

      We have therefore revised the manuscript to avoid overstating beta cell V1bR expression as a major isolated claim. Instead, we now present the data as evidence that AVP/V1bR signaling acts most prominently through alpha cells, while beta cell responses emerge in a concentration-, glucose-, and statedependent manner within the intact islet. This interpretation is consistent with the reviewer’s concern that 10 nM AVP preferentially activates alpha cells, but it also accommodates our observation that beta cell collective activity can be modulated under defined pharmacological and metabolic conditions. We believe that this is an important distinction, because the absence of a uniform beta cell Ca<sup>2+</sup> activation at one AVP concentration does not exclude beta cell modulation by AVP/V1bR signaling within the intact islet network. The Introduction and Discussion have been revised accordingly to clarify that our study does not simply challenge the alpha cell V1bR model, but expands it by examining how AVP-dependent alpha cell activation, possibly lower beta-cell receptor expression, and collective beta cell dynamics interact in the native pancreatic slice preparation.

      (2) We know from bulk RNAseq data on purified alpha, beta, and delta cells from both the Huising and Gribble groups that there is no expression of V2a. I will point you to the data from the Huising lab website published almost a decade ago (http://dx.doi.org/10.1016/j.molmet.2016.04.007) - which is publicly available and can be used to generate figures (https://huisinglab.com/dataghrelin-ucsc/index.html). They indicate the absence of expression of not only AVP2 receptors anywhere in the islet, but also the lack of expression of V1bra, V1brb, and Oxtr in beta cells. Instead of the detailed list of expression of these 4 receptors elsewhere in the body, it would be more directly relevant to set up their pancreatic slice experiments to summarize the known expression in pancreatic islets that is publicly available. It would also have helped ground the efforts that involved the generation of the V1aR agonist and V2R antagonist, which confirm these known AVP/OXT receptor expression patterns.

      We thank the reviewer for pointing us more directly to the publicly available islet expression datasets. We agree that the expression of AVP/OXT receptors in purified alpha, beta, and delta cells provides an important reference frame for interpreting our pharmacological data, and we have revised the manuscript to summarize these islet-specific datasets more directly rather than emphasizing receptor expression in other organs. These data support the absence or very low expression of V2 receptors in islet endocrine cells and confirm that V1b receptor expression is substantially enriched in alpha cells compared with beta cells.

      At the same time, as outlined in our response above, we do not think that the currently available transcriptomic datasets are sufficient to categorically exclude low-level V1bR transcript expression or functional relevance in beta cells within the intact islet. For this reason, we added independent RNAscope validation to assess V1bR transcripts in the pancreatic slice preparation. These data confirm stronger V1bR expression in glucagon-positive alpha cells, while also showing a broader expression pattern within the islet and pancreas.

      We have also revised the rationale for the pharmacological experiments using V1aR- and V2R-directed tools. We now present these experiments not as evidence for unexpected receptor expression, but as functional controls that are consistent with the known AVP/OXT receptor expression patterns in pancreatic islets. This better aligns the manuscript with the existing transcriptomic literature while preserving the main physiological question of the study: how AVP/V1bR-dependent signaling reshapes alpha-cell activity and beta-cell collective dynamics in intact pancreatic tissue.

      (3) Importantly, the lack of V1br from beta cells does not invalidate observations that AVP affects calcium in beta cells, but it does indicate that these effects are mediated a) indirectly, downstream of alpha cell V1br or b) via an unknown off-target mechanism (less likely). The different peak efficacies in Figure 4G would also suggest that they are not mediated by the same receptor.

      We agree with the reviewer that the absence or very low abundance of V1bR transcripts in beta cells in published transcriptomic datasets would not invalidate the observation that AVP modulates beta-cell Ca<sup>2+</sup> activity. It does, however, raise the important question of whether this modulation is mediated indirectly through alpha-cell V1bR activation, through V1bR expression in beta cells that is difficult to resolve transcriptomically, or through another mechanism. To address this more directly, we have now added RNAscope data, which confirm relatively stronger V1bR transcript enrichment in glucagon-positive alpha cells, but also show a broader V1bR transcript signal within the islet and pancreas. Thus, while our data support alpha cells as the dominant V1bR-positive endocrine population, they do not support a strict absence of V1bR-associated signaling capacity in the beta cell compartment.

      We also agree that different peak efficacies in alpha and beta cells could be interpreted as evidence for distinct receptors or indirect mechanisms. However, we favor a different interpretation: the apparent efficacy of AVP depends strongly on the physiological state in which the cells are tested. This is particularly evident in beta cells, where the AVP efficacy peak shifts with glucose concentration, suggesting that the beta-cell response is shaped by the metabolic and Ca<sup>2+</sup>-handling context rather than by receptor occupancy alone. In this framework, the same V1bR/Gq-dependent input can generate different downstream Ca<sup>2+</sup> outcomes in alpha and beta cells because the two cell types operate in different dynamic regimes.

      We have therefore revised the manuscript to acknowledge this dilemma more explicitly. We now state that beta cell effects of AVP could include indirect alpha cell-dependent components, but given the magnitude and statedependence of the beta cell Ca<sup>2+</sup> response it is unlikely to be driven by alpha cell activation. Instead, our preferred interpretation is that AVP/V1bR signaling acts within the intact islet as a context-dependent perturbation of the collective beta cell Ca<sup>2+</sup> system, with IP3R-dependent mechanisms being modulated by glucose-dependent changes in beta cell excitability and intracellular Ca<sup>2+</sup> handling.

      (4) The rationale for the use of forskolin across almost all traces is unclear. It is motivated by a desire to 'study the AVP dependence of both alpha and beta cells at the same time'. As best as I can determine, the design choice to conduct all studies under sustained forskolin stimulation is related to the permissive actions of AVP on hormone secretion in response to cAMPgenerating stimuli. The permissive actions by AVP that are cited are on hormone secretion, which in many cell types requires activation of both calcium and cAMP signaling. Whether the activation of V1br and subsequent calcium response is permitted by cAMP is unclear. I believe the argument the authors are making here is that the activation of beta cell calcium by AVP is permitted by forskolin. i.e., the cAMP stimulated by it in beta cells. However, the design does not account for the elevation of cAMP in alpha cells and subsequent release of glucagon, particularly upon co-stimulation with AVP, which permits glucagon release by activating a calcium response in alpha cells. This glucagon could then activate beta cells. If resolving the mechanism of action is the goal, often less is more. The activation of Gaq-mediated calcium is not cAMP dependent (although the downstream hormone secretion clearly often is). As was shown, AVP does not activate calcium in beta cells in the absence of cAMP. The experiments in Figures 1, 2, and 4 should have been completed in the absence of cAMP first.

      We agree with the reviewer that the use of forskolin needs to be explained more clearly, and we have revised the manuscript accordingly. Our rationale was based on the established permissive role of cAMP in AVP-dependent endocrine responses, but we acknowledge that this does not necessarily imply that the upstream V1bR/Gq-mediated Ca<sup>2+</sup> response itself is cAMP-dependent. The reviewer is also correct that forskolin elevates cAMP broadly and therefore may affect both alpha and beta cells, including the possibility that AVP-enhanced alpha cell activation and glucagon release secondarily influence beta cell activity.

      In fact, our initial experiments were performed without forskolin and revealed an important difficulty: stimulatory glucose alone can increase cAMP levels to a variable extent, as also supported by our previous work on epinephrine signaling, thereby shifting the apparent peak efficacy of AVP stimulation. Thus, forskolin was originally used to reduce this variability and create a more defined cAMP-permissive background in which alpha and beta cell responses could be compared in the same slice. However, we agree that this design works against isolation of beta cell-autonomous AVP effects.

      Within a scope of another study we have done an independent series of more focused experiments using GLP-1 receptor stimulation, which preferentially increases cAMP signaling in beta cells compared with the broad cAMP elevation produced by forskolin. We have clarified that the modulation of the AVP-dependent pathway by GLP-1 and related ligands at largely supports beta cell-autonomous AVP effects. It is part of ongoing work and will be reported independently, because a full mechanistic dissection of cAMP–AVP interactions goes far beyond the scope of the present study.

      (5) It is unexpected that epinephrine in Figure 2 does not activate the alpha cell calcium? A recent paper from the same group (Sluga et al) shows robust calcium activation in alpha cells in a similar prep by 1 nM epinephrine, which is similar to the dose used here.

      We thank the reviewer for pointing this out, but we would like to clarify that epinephrine did significantly activate alpha-cell Ca<sup>2+</sup> activity in our experiments, as shown in Fig. 3F. This result is consistent with our previous study by Sluga et al., where low nanomolar epinephrine robustly activated alpha cell Ca<sup>2+</sup> signals in the pancreatic slice preparation. The main point of the present comparison was therefore not that epinephrine is inactive in alpha cells, but that AVP produces a substantially stronger and reproducible alpha cell Ca<sup>2+</sup> response under comparable experimental conditions. This is also consistent with the data of van der Meulen et al., supporting the view that AVP/V1bR signaling is a particularly potent activator of alpha cell activity. We have revised the text to make this comparison clearer and to avoid the impression that epinephrine failed to activate alpha cells in our preparation.

      (6) Figure 8 suggests a pharmacological activation of beta cell V1bR in the low pM range. How do the authors reconcile this comparison with the apparent absence of an effect of AVP stimulation at low pM to low nM doses in beta cells (Figure 4A)? I note that there are changes over time with sustained beta cell stimulation with 8 mM glucose, but these changes are relatively subtle, gradual, and quite likely represent the progression of calcium behaviors that would have occurred under sustained glucose, irrespective of these very low AVP concentrations. I will note that the Kd of the V1bR for AVP is around 1 nM, with tracer displacement starting around 100 pM according to the data in figure 5B, which is hard to reconcile with changes in beta cell calcium by AVP doses that start 10-100-fold lower than this dose at 1 and 10 pM (Figure 8).

      We agree that the interpretation of low-pM AVP effects requires caution, particularly when compared with reported V1bR binding affinities. The apparent discrepancy between Fig. 4A and Fig. 8 most likely reflects differences in experimental design, stimulation context, and readout sensitivity. In Fig. 4A, we assessed acute AVP effects under conditions in which beta cell Ca<sup>2+</sup> responses are relatively threshold-dependent and where low AVP concentrations produced little or no activation. In contrast, Fig. 8 analyzes prolonged beta cell population dynamics during sustained stimulation with 8 mM glucose, a physiological stimulatory context in which even weak modulatory inputs may become detectable at the level of collective Ca<sup>2+</sup> activity.

      Importantly, the strongest and statistically significant effect was observed at 100 pM AVP, while lower pM concentrations showed only a trend. We therefore do not interpret the low-pM range as evidence for robust direct pharmacological activation of beta cell V1bR. Rather, these data suggest that AVP may exert permissive or modulatory effects within an already active beta cell network, where glucose-dependent excitability, receptor-effector coupling, and Ca<sup>2+</sup> amplification mechanisms can enhance the apparent efficacy of weak inputs. This interpretation is consistent with the known permissive role of AVP in endocrine responses, where AVP may not act as a primary activator alone but can increase the efficacy of other physiological stimuli.

      We have also clarified that sustained 8 mM glucose alone does not account for these effects, since Suppl. Fig. 1 shows no comparable time-dependent progression of Ca<sup>2+</sup> behavior under sustained glucose stimulation alone. Thus, we now present the low-concentration AVP effects as subtle, contextdependent modulation within the physiological stimulatory range, rather than as evidence for direct beta cell activation at concentrations below the expected receptor affinity range.

      Reviewer #3 (Public review):

      Summary:

      This work aims to better understand the role of arginine vasopressin (AVP) in the control of islet hormone secretion. This builds on previous literature in this area reporting on the actions of AVP to stimulate islet hormones. The gap in literature being addressed by these studies is primarily focused on the glucose-dependency of AVP on both insulin and glucagon secretion. A secondary objective is to explore the role of individual receptors with the use of newly generated peptides and existing tools. The methods include the use of Ca2+ imaging in pancreas slices from mice, with additional outcomes including insulin secretion in some areas. The conclusions presented are that AVP acts through V1b receptors in both alpha- and beta-cells, that this activity occurs in the high cAMP environment, and is glucose dependent.

      Strengths:

      The area of research is emerging with plenty of room for new contributions. The concept of AVP stimulating islet hormone secretion is important and deserving of further insight. The use of pancreas tissue to image primary cells makes the experiments physiologically relevant. The advancement of novel tools in this area should be helpful to other groups investigating the actions of AVP.

      We would like to thank the reviewer for recognizing the potential of our emerging area of research.

      Weaknesses:

      The conclusions are only modestly supported by the data and lack experimental depth and rigor. The rationale for only conducting studies at high cAMP conditions is not entirely clear and limits the conclusions that can be made. The use of Ca2+ is helpful, but it is a surrogate for hormone secretion. Additional measurements of hormone secretion are needed to enhance the robustness of these conclusions. Consideration of paracrine effects between alpha- and beta-cells is only superficially made and is likely essential in the context of the experimental design. For instance, there is clear literature that alpha-cells secrete several factors that work in paracrine interactions on beta-cells and autocrine actions back on alpha-cells. Conducting these studies in a high cAMP context only completely overlooks these interactions, skewing the interpretations made by the investigators. Finally, the clarity of the experiments and results could be significantly enhanced.

      We thank the reviewer for this balanced assessment and for emphasizing several issues that are central to the interpretation of our study. We agree and now explicitly state in the Limitations, that Ca<sup>2+</sup> oscillations are a surrogate readout for hormone secretion and that currently used stimulation protocols are not optimized to directly quantify the relationship between Ca<sup>2+</sup> dynamics and secretory output. To address this limitation, we have now expanded the functional part of the study by adding complete insulin and glucagon secretion measurements during AVP concentration ramps. These new data provide a stronger functional framework for interpreting the Ca<sup>2+</sup> imaging results, while also clarifying that Ca<sup>2+</sup> activity and secretion cannot be assumed to correlate linearly under all stimulation protocols.

      We have also revised the rationale for the high-cAMP experimental condition. The original aim was to reduce variability arising from glucose-dependent endogenous cAMP signaling and to study alpha and beta cell responses in a common permissive background. However, we agree that broad forskolin stimulation complicates the interpretation of cell-autonomous versus paracrine mechanisms. Independent experiments within a scope of another study demonstrate that using GLP-1 co-stimulation, which provides a more beta-cell-oriented cAMP-permissive condition and supports the interpretation that AVP can modulate beta-cell collective activity in a manner that is not solely secondary to alpha-cell activation.

      We fully agree that paracrine interactions within the islet are physiologically important and must be considered, particularly in intact pancreatic slices. Nevertheless, the rapid onset of the AVP effects observed in beta cell Ca<sup>2+</sup> activity argues against a mechanism mediated predominantly by slower indirect paracrine loops. High AVP concentrations, as shown in Fig. 5, significantly shorten the intervals between Ca<sup>2+</sup> events in both alpha and beta cells, but that the activity of the two cell populations remains largely noncoordinated. This temporal dissociation does not exclude paracrine modulation altogether, but it argues against a simple alpha-cell-driven explanation for the beta-cell response.

      We have revised the manuscript to state these points more clearly and to moderate conclusions where the data support modulation rather than definitive cell-autonomous receptor action. We believe that the added secretion experiments and alpha/beta event-timing analysis substantially strengthen the physiological interpretation of the study, and we thank the reviewer for raising these issues.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The paragraph discussing the benefits of slice physiology over islets is not reflective of how most - if not all of your colleagues who do islet experiments conduct these. Many labs have reported for years high-quality GSIS experiments, synchronous calcium responses, and a plethora of studies detailing the mechanism of hormone and neurotransmitter actions using islet models, and have done so well. Slice physiology is a unique and helpful model that can have advantages over other models. This particular reviewer uses both models in their lab and each has benefits and - inevitably - drawbacks. Many of the possible drawbacks cited for islet studies apply equally to slices, including the possibility of altered gene expression, lack of innervation, and circulation. Added drawbacks are the exposure to higher levels of pancreatic enzymes from the slice, which require co-culture with enzyme inhibitors.

      We agree with the reviewer and have revised these limitations accordingly. Our intention was not to imply that isolated islet preparations are generally inferior, since they have provided a highly productive and rigorous experimental platform for GSIS, synchronized Ca<sup>2+</sup> dynamics, and mechanistic studies of hormonal and neurotransmitter regulation with standardized protocols with all their positive and negative sides. We modified the presentation of pancreatic slices as a complementary model with specific advantages, particularly preservation of local tissue architecture, while also acknowledging their limitations. The revised text therefore avoids a comparative hierarchy between slices and isolated islets and instead emphasizes that both models have distinct strengths and drawbacks depending on the experimental question.

      (2) If you want to demonstrate direct actions on beta cells, deconstructing the islet would be a better way to go. Less complicated, not more. Dissociated beta cells, instead of slices, were used just to prove or disprove the hypothesis of direct beta cell effects of AVP.

      We agree that dissociated beta cells can be a useful reductionist model to test whether AVP is capable of acting directly on individual beta cells. However, this approach would also remove the collective beta-cell activity that is central to the physiological question addressed in the present study. Since our data indicate that AVP effects emerge within the intact islet as rapid and extensive changes in coordinated Ca<sup>2+</sup> dynamics, dissociation would not necessarily provide a more informative model for understanding these responses. The fast onset and magnitude of the beta cell response argue against a predominantly indirect non-autonomous mechanism, and this interpretation is further supported by the GLP-1 co-stimulation experiments, which are more consistent with beta cell-autonomous modulation. We have therefore clarified in the revised manuscript that dissociated-cell experiments would be valuable for a narrowly defined receptor-cell autonomy question, but would not resolve the collective islet dynamics that are the focus of this work.

      (3) If you want to sustain the claim of beta cell expression of V1br, you would have to demonstrate this far more directly by staining (if appropriate antibodies exist), by beta cell-specific deletion of V1br, or by highly selective, well-validated pharmacology. This should include a demonstration of Gaqdependence in isolated beta cells.

      We have added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      Minor

      (1) O'Carroll et al. should be cited in the context of islet permissive actions of AVP/cAMP. PMID: 18434353, although that paper offers no evidence that the AVP-dependent potentiation of insulin release is mediated directly by beta cells. It does confirm dependence on PKC.

      We agree and have added O’Carroll et al. in the revised manuscript in the context of AVP/cAMP-dependent permissive actions on islet hormone secretion. We therefore use it as support for the broader concept of AVP-dependent amplification of secretion in a permissive signaling context, rather than as direct evidence for beta-cell-autonomous V1bR signaling.

      (2) Figure 2 E-H, glucose concentration mislabeled.

      We thank the reviewer for pointing this out. The glucose concentration label has been clarified: panels E–H show pooled data from separate experiments performed at 8 mM glucose, whereas panel D shows a representative experiment performed at 9 mM glucose.

      (3) The insulin secretion in 4E is difficult to interpret without a low-glucose control. If this is hard to do in a slice preparation, a separate static islet secretion experiment would help here. The possibility that the inhibition of insulin secretion traces back to the activation of delta cells by AVP could be considered - I struggle to come up with a plausible mechanistic explanation why AVP (which activates calcium in alpha and in beta cells in the presence of 8 mM G plus forskolin according to your data) would inhibit insulin secretion.

      We agree that the original insulin secretion experiment was difficult to interpret without a clearer low-glucose reference condition. To address this, we have added new insulin release experiments in which glucose was lowered to a non-stimulatory range between AVP concentrations, followed by sequential stimulation with 8 mM glucose and 500 nM forskolin in the same slices. These new data provide a broader dynamic range for assessing insulin secretion and allow a more direct comparison between AVP-dependent Ca<sup>2+</sup> modulation and secretory output.

      We also agree that AVP-dependent inhibition of insulin secretion requires careful interpretation. One possible explanation is not simply activation of delta cells, but a failure of the beta cell collective to maintain coordinated activity at very high AVP concentrations. In the Ca<sup>2+</sup> imaging data, high AVP concentrations increase activity in many beta cells, but numerous cells within the islet fail to keep pace with the collective oscillatory rhythm, leading to fragmented and less synchronized population activity. Thus, despite increased frequency of Ca<sup>2+</sup> oscillations in the islet, the integrated beta cell output may become less efficient for insulin secretion. We have added this interpretation to the revised manuscript and now discuss delta cell activation as a possible contributing mechanism, but not as the primary explanation supported by our current data.

    1. eLife Assessment

      In this important work, the authors develop methods to forecast epidemic growth from viral sequencing data alone. The evidence for the usefulness of the approach is solid, but some justifications and methodological details are incomplete. This study should be of broad interest to the community interested in viral dynamics and epidemiology.

    2. Reviewer #1 (Public review):

      Summary:

      This paper develops a formalism for quantifying epidemic dynamics in terms of relative fitnesses of circulating variants, uses the formalism to elucidate fundamental tradeoffs of epidemics driven by variants with increased transmissibility versus immune escape capability, shows the formalism implies a natural quantity measuring the impact of selection on epidemic growth, and demonstrates that the formalism enables a decomposition of epidemic dynamics into circulation among different immunity groups. The relative fitness formalism enables these analyses to be performed with genetic sequence data only, a major benefit of the model given the relatively high availability of sequence data compared to other data streams such as case counts and titers.

      Strengths:

      Linking epidemic dynamics to pathogen evolution is a fundamental problem in studies of antigenically variable pathogens, with models of epidemic dynamics and immune-driven evolution going back decades in applications to respiratory pathogens such as influenza. The COVID-19 pandemic heightened the urgency for developing methods for quantifying epidemic growth in contexts where novel variants emerge, leading to differential susceptibility among individuals with diverse exposure histories with implications for vaccination strategies. Real-world data streams such as case counts and immunological measurements have a variety of shortcomings that pose major challenges for quantitative models aiming to inform policy. In recent years, genetic sequencing data has become widely available for pathogens including SARS-CoV-2 and influenza, allowing tracking of pathogen evolution at unprecedented detail in real time, yet biases in the collection of sequence data across different populations make connections between absolute epidemic size and variant frequencies from sequence data not immediately transparent.

      This paper's contributions are exciting because they demonstrate new ways to link pathogen evolution and epidemic dynamics using very accessible data. From a theoretical perspective, the model is appealing because of its simple derivation in terms of compartmental models of epidemics, which are standard in the literature, and its clear extension to populations with heterogeneous immune histories. The latter extension leads directly to new methods for inferring immune groups with differential susceptibility to antigenically distinct variants in populations with heterogeneous immune histories without access to immunological data such as titers, an important advance given the wide applicability of quantification of antigenic relationships among variants in real populations.

      Weaknesses:

      While the demonstrated methods for forecasting short-term epidemic growth and for quantifying population immunity using sequence data are exciting as proofs of principle, the validation and statistical support provided in the analyses have drawbacks that are not fully addressed in the manuscript, weakening the evidence for the usefulness of the methods in their current form.

      The analyses forecasting epidemic growth using Gaussian process models are justified using Pearson correlation coefficients whose values are extremely low for the test data period. The explanation given for this is that the case data used to validate the predictions has worse ascertainment over time, but it is not shown directly that the model may be working well despite the low correlations. Whereas, by eye, the predicted epidemic growth curves appear to capture features of the observed epidemic growth curves, the computed metrics don't support the claim of success of the predictions. Additionally, nearly all the model fits lack estimates of uncertainty, so it is not possible to discern the significance of departures between the model and data, or subtle differences in relative fitness calculations across geographies.

      The analysis of latent pseudo-immune components also suffers drawbacks that render it more of an interesting proof of principle than a convincing tool for prediction at this point. In particular, in figures S18 and S19, metrics meant to quantify the statistical significance of the results show no difference from null models computed by permuting variants and their escape vectors, yet no interpretation is given for the lack of significance. Moreover, the model fits relating titer distance to pseudo escape distance seem unsuccessful for JN.1 infection and XBB infection histories, which is not adequately accounted for in the text, which cites just "weaker correlations" in these cohorts.

      In several instances, the evidence for the new data analyses is weakened by a lack of clarity in the presentation of the technical details of the methods. For example, in the discussion of the Gaussian process models, it was not clear what features of the problem inform the choice of kernel (Matern 5/2), which hyperparameters were used, and how novel this use of Gaussian processes is. In the section describing methods for predicting epidemic growth rate from selective pressure, the discussion of the gradient boosting regressor model provided no intuition as to why this method performed better than the others tested or whether this was particularly important to the conclusions, and the lack of discussion of uncertainty or variability in the model predictions makes it difficult to assess the significance of the time series estimates alone. In the discussion of the latent immune factor model, the mismatch between the notation used in Equation 5 compared to that in Equation 18 made the derivations more difficult to follow. Subsequently, the explanation of the fitting of the pseudo-immune model left out details, such as an explicit definition of distance in pseudo-escape space, to what extent the group-level mean aggregated titer measurement captured features of the titer data (despite ignoring interindividual variability), and a thorough discussion of the successes and shortcomings of the fits in different scenarios. More explicit presentation of the mathematical choices going into the methods, sources and quantification of uncertainty, and cases where the model performs well or poorly could significantly bolster the case for the usefulness of sequence data in quantitatively predicting epidemic growth and antigenic relationships among variants in practice, in more general settings than those carried out here.

    3. Reviewer #2 (Public review):

      Summary:

      The authors first introduce a framework to understand how different phenotypic drivers of viral evolution, i.e., changes in transmissibility versus immune escape, complicate epidemic forecasting using only genetic data. To overcome these complications, they advance an evolutionary "selective pressure" metric to predict population-wide epidemic growth from genetic data alone. Separately, they introduce a latent space model to infer a "pseudo" population immune structure from geographic variation in viral lineage dynamics, and find that the inferred pseudo-structure predicts human serological data.

      Strengths:

      This paper begins with a useful pedagogical exposition on the connection between fitness-driven frequency dynamics and underlying mechanisms of viral-immune co-evolution. A major contribution of this paper - a method to infer variant-specific escape properties from geographically non-uniform variant frequency dynamics alone - is an interesting and potentially timely one, given the advance of sequencing-based surveillance.

      Weaknesses:

      The logical flow of the pedagogy part of the text works against the reader, which is problematic since it motivates the rest of the text. Moreover, some important modelling choices and procedures, particularly with respect to the selective pressure metric, are only cursorily described in the methods section. The lack of explanation and detail, especially relative to more simple choices that are seemingly motivated by the authors' own theory, makes it difficult to understand and therefore assess their validity and/or necessity.

    4. Reviewer #3 (Public review):

      Summary:

      This study introduces a new analytical framework to analyze how viral variant frequencies change over time and in different locations. Two examples are given that demonstrate where this approach can be useful and where other approaches can be ambiguous in characterizing novel variants. The authors then demonstrate that the spatiotemporal dynamics of variant frequencies can be used to predict future epidemic growth rates and to investigate how variants differ in immune escape.

      Strengths:

      (1) Examples are provided that make the study accessible for a general audience.

      (2) The authors demonstrate that their approach is predictive both of overall epidemic growth rates and immunological distance between variants.

      (3) The approach introduced in this study can be readily applied to current and future epidemiological challenges that are similar to SARS-CoV-2 with respect to the relative evolutionary timescales wherever there is spatiotemporal heterogeneity in the susceptible population.

      Weaknesses:

      (1) The authors conclude their abstract claiming that their method provides an early signal of epidemic growth. Can this be quantified? Could the authors perform retrospective analyses for sequences available through various cutoff times, identify how early significant new variants are detected, and compare this to other detection methods?

      (2) Analysis depicted in Figure 4 and Figure S9 could be explored further than speculatively attributing weak correlation to declining reporting rates for US states. Exploring how correlation between data and prediction varies over time during the test period might identify periods/events that explain weak correlation overall. The authors could explore predicting growth rates for estimated state prevalences rather than reported cases.

    1. eLife Assessment

      The paper provides a valuable, foundational dataset that will undoubtedly provide substantial value to the cancer research community. The aggregation and harmonization of a broad, rich, multi-omic dataset is convincing and will potentially support further downstream research on GIST, which currently has poor supporting datasets. Support for the accompanying claims is more incomplete, however: the biological demonstration rests on a single gene tested by one approach, several analyses lack sample sizes and multiple-testing correction, the claims made for the LLM assistant are not backed by direct evaluation, and the curated data matrices would ideally be deposited independently of the web.

    2. Reviewer #1 (Public review):<br /> <br /> Summary:

      This tumour type is missing from the big pan-cancer databases, so none of the popular online analysis tools works for it. That's a real gap, and it's the right one to go after. The authors build an online resource that gathers the scattered public molecular datasets for this disease, adds three of their own patient cohorts, ties everything to clinical data, and exposes interactive tools, downloads, and programmatic access so other people can build on it. To show what it does, they take one gene through the whole platform - clinical, gene-expression, protein, single-cell, immune, and drug-response and then test that gene in cell lines. So there are really two things on offer here: a resource and a practical example of using it. They land very differently.

      Strengths:

      The resource is the real contribution, and it's done with care. It covers 37 centres and nearly 2,000 samples across five kinds of molecular data, and the authors are honest about provenance: how they screened datasets in or out, where they recorded the diagnostic codes, and why they dropped ambiguous mixed-tumour collections. The key methodological decision is the right one; every analysis runs inside its own cohort, and the cross-cohort views are explicitly "for looking, not for combining." That's exactly how you should treat heterogeneous public data, and they say so plainly instead of quietly pooling everything. Their three pathologist-confirmed cohorts add genuine independent material, so this isn't a re-skin of data that already existed. And because the code and a public access point are actually available, the reuse claim holds.

      The example is internally consistent, which is what makes it persuasive. The gene reads higher in higher-risk patients across several independent cohorts and in their own protein data, tracks with the disease spreading and recurring, and lines up with worse survival. The single-cell data put it in the dividing cells; the pathway analysis points to proliferation. Three independent data types landing on the same proliferation story are the strongest part of the biology.

      Weaknesses:

      The honest problem is that the entire biological story rests on one gene, tested one way. The lab work is two cell lines with the gene knocked down, showing less growth and migration: there is no rescue to confirm the effect is real, no second gene to show the approach generalises, nothing in a living animal. That earns the modest claim: the resource can point you at a candidate worth testing. It does not earn the headline claim that the platform reliably generates good target hypotheses, because we only ever watch it succeed once. One example illustrates a workflow; it doesn't establish a method.

      Some of the statistics won't survive scrutiny. The clearest case is a perfect separation between treatment-resistant and treatment-sensitive cases from a single immune cell population, reported with no error bars, no check for information leakage, and apparently from very few samples. A perfect result in that setting is almost always overfitting or a small-sample artefact, not a strong classifier. The same pattern shows up elsewhere: small groups, p-values with no effect sizes or error bars, and no correction for the enormous number of features and cohorts being tested across the whole platform. Separately, one drug result is a correlation against a predicted sensitivity score from a model.

      The AI assistant gets far more weight than the evidence supports. Credit where due: the authors are clear and consistent that it only helps interpret and navigate, and never touches the data, the statistics, or the results. That's the correct line to draw, and they hold it. But the assistant itself is never tested, no accuracy numbers, no benchmark, no error analysis, no described way for a human to check what it produces. Calling it something that "fundamentally transforms the user experience" is an assertion, not a finding. And since even the literature feature is admitted not to be a proper systematic review, the prominence of the artificial-intelligence framing runs ahead of what's been shown.

    3. Reviewer #2 (Public review):

      Summary:

      dbGIST appears to be the first dedicated multi-omics resource worldwide that is specifically focused on GIST.

      Strengths:

      The main value of the paper is not simply that the authors collected datasets, but that they built a usable resource around them, with cohort-aware analyses, curated clinical labels, interactive visualizations, downloadable results, selected API access, and an optional LLM-assisted interface. The work is solid, and the database is likely to be useful for GIST researchers interested in target discovery, cross-dataset validation, drug-response hypotheses, and translational follow-up.

      The MCM7 analysis is a reasonable use case. It shows how a user can start from one candidate gene and then move across transcriptomic, proteomic, clinical, single-cell, immune-related, drug-response, and experimental evidence. I do not see this as the main discovery of the paper, but rather as a practical demonstration of what the database can do. That is appropriate for a resource manuscript.

      Weaknesses:

      (1) The authors should make the organization of the platform a little easier to follow. The manuscript refers to five primary omics layers, six omics-focused pages, and eight analytical modules. This structure is understandable after reading the relevant sections, but it may not be immediately obvious to readers. A brief clarification of how the omics layers, web pages, and analytical modules relate to each other would help.

      (2) Since dbGIST is a live web resource, the authors should provide a clear versioning statement. The manuscript should indicate which version of the database corresponds to the analyses and figures reported in the paper, and how future updates will be distinguished from the version evaluated here. This is a small point, but it matters for reproducibility.

      (3) The API function is a strength of the resource, but it is still described rather generally. The authors should give more concrete documentation of what can be accessed through the API, what inputs are required, and what type of output is returned. This could be placed in the supplementary materials. It would make the database more useful for computational users.

      (4) The manuscript should clarify the status of downloadable data. It is clear that figures, source-data tables, and selected derived outputs are available, but it is less clear whether the full processed matrices used internally by the platform are downloadable or only maintained for deployment. This distinction should be stated plainly.

      (5) The statistical reporting in the MCM7 clinical-association analyses needs a little more care. Several p-values are shown across different cohorts and clinical variables. The authors should state whether these are nominal p-values or adjusted p-values. If they are nominal, that is acceptable for a resource demonstration, but the exploratory nature of the analyses should be made clear.

      (6) The ROC analyses for imatinib response should include sample sizes, and confidence intervals for AUC values would be useful if available. Some of the AUC values are high, and without group sizes, it is difficult to judge how stable those estimates are. The authors should avoid implying that these ROC results are validated predictive models.

      (7) The interpretation of MCM7 should be slightly more cautious. MCM7 is a well-known DNA replication and cell-cycle gene, and the single-cell analyses seem to support its association with proliferative cell states. This is biologically consistent, but it also means that MCM7 expression should not be presented as tumour-cell-specific without qualification. The manuscript should frame it mainly as a proliferation-associated signal in the current analysis.

      (8) The drug-response section would benefit from a clearer explanation of the response metric. The authors report correlations between MCM7 expression and predicted response to C6-ceramide, but readers need to know whether the predicted value represents IC50, AUC, sensitivity score, or another metric. The direction of interpretation should also be made explicit, since a negative correlation can mean different things depending on the scoring system.

      (9) The single-cell annotation would be more convincing if the authors provided a compact marker-gene summary for the major cell types in each single-cell cohort. The current description of annotation by source labels, marker inspection, and manual curation is reasonable, but users of the database would benefit from seeing the marker evidence behind the labels.

      (10) The experimental validation section should include a few routine details that are currently not easy to find. The siRNA sequences or target regions, number of biological replicates, statistical tests for the CCK-8 and wound-healing assays, and details of wound-closure quantification should be reported. These additions would make the in vitro part more reproducible.

      (11) The wound-healing result should be interpreted with caution. Since MCM7 knockdown reduces proliferation, reduced wound closure could reflect changes in proliferation, migration, or both. Unless proliferation was controlled during the wound-healing assay, the authors should avoid describing this result as purely migratory.

      (12) The LLM-related claims should remain conservative. The assistant is a useful feature for navigation, plain-language explanation, and user support, especially for clinicians or wet-lab researchers. However, the strongest statements about the LLM transforming interpretation or automating analysis should be toned down. The important point is that the LLM layer helps users interact with the resource, while the numerical analyses come from predefined dbGIST modules.

    4. Reviewer #3 (Public review):

      Summary:

      The dbGist dataset/tool would provide substantial value to the cancer research community.

      Strengths:

      The manuscript presents dbGIST, a dedicated GIST-focused multiomics resource integrating data from 37 centers and ~2k samples across genomics, transcriptomics, proteomics, phosphoproteomics, and single-cell transcriptomics. Given that GIST is virtually absent from major cancer genomics consortia (TCGA, ICGC), this resource fills a genuine gap and represents a valuable contribution to the GIST research community.

      (1) The MCM7 case study effectively demonstrates the platform's utility, linking a resource-derived candidate to survival outcomes.

      (2) The LLM-assisted interface (dbGIST Assistant) is a reasonable addition for accessibility, lowering the barrier for clinicians and wet-lab researchers, who may not always have the skill set required for proper data analysis, especially for a rich and wide dataset like the dataset in question.

      Weaknesses:

      (1) Data deposition (major):

      While the manuscript references public accessions for raw source datasets and provides a GitHub repository for code, it remains unclear where the **curated, harmonized data matrices** - which represent the core value-add of this work - are independently deposited. Access to these processed data appears to depend entirely on the dbGIST web interface and API. The authors should deposit the harmonized matrices in a persistent, general-purpose repository to ensure long-term availability independent of the web platform.

      (2) LLM agent capabilities underspecified:

      The manuscript would benefit from a clearer description of the assistant's capabilities and boundaries. Specifically, what tools or actions are available to the LLM agent? Can it execute code against the underlying data, trigger analytical modules programmatically, or is it limited to natural-language explanation of pre-computed results? Clarifying this would help readers assess the scope of the AI layer and distinguish it from agentic platforms that perform computation on behalf of the user.

    1. eLife Assessment

      This important study presents the development of a model to quantify the dynamics of stress-induced volatile emissions in plants and identify biologically relevant differences in these responses. The model is based on a convincing methodology and represents a starting point for studies of similar responses in other plants or investigations of induced phenotypes in different biological systems.

    2. Reviewer #1 (Public review):

      Summary:

      The manuscript by Waterman et al. describes the development of a mathematical model that quantifies plant volatile emissions dynamics in response to mechanical/biotic stress. Model outputs were based on volatile emission measurements from maize plants using PTR-MS. Modeling revealed differences in emission patterns dependent on the intensity of wounding damage, application of herbivore oral secretions, age of leaf, circadian clock, and genotype. Differences were also observed between different types of volatiles, and the response curves somewhat correlated with expression patterns of biosynthetic genes. Moreover, the model showed priming effects from overlapping response curves upon multiple wounding events.

      Strengths:

      As a non-expert in modeling, this reviewer assesses the work from a broader point of view. Overall, I consider this model to be useful for other researchers to quantify volatile emission dynamics for their plant system. Generating the models does not seem to be overly complicated as long as emissions can be measured with a real-time system such as PTR-MS, which is costly and not available to every lab. The advantage of this approach is that it does not rely on parameters of underlying enzymatic pathways or transport processes. The authors claim that it can be easily applied to other biological responses.

      Weaknesses:

      The manuscript lacks a deeper discussion of how the model can help make predictions of volatile emission dynamics from plants in the greenhouse or field. Can the model be trained and validated with volatile measurements from plants under different environmental conditions? How realistic is this approach given the complexity of a field environment? It would be helpful to provide a better outlook of the application of the model for scientists in the field of plant volatile biology and beyond.

      The authors state that "emissions can be regulated independently of each other" (Line 359). I would assume that regulatory mechanisms in different genotypes are similar but show genotype-specific variation.

    3. Reviewer #2 (Public review):

      This is a study of the dynamics of plant volatile emissions, using a curve-fitting approach to describe salient properties of the dynamics of plant volatile chemicals. The study is interesting and unique in taking this approach. Some of the dynamics uncovered (e.g. lagged emission of many sesquiterpenes) are already well known using less sophisticated approaches, while other properties (diurnal cycles in emission dynamics) are newly uncovered. The approach in general is new for the topic of plant volatile emissions, but curve-fitting is widely used to describe the dynamics or function-valued responses of plants and other organisms. The study thus reads as rather methods-focused, giving tidbits of interesting properties of the dynamics of plant VOCs rather than being structured strongly around clear biological hypotheses. The method seems like a logical and robust way to analyze the dynamics of plant VOCs. I believe the impact of the work will largely depend on whether there are substantial and meaningful outcomes (for herbivores, downstream processes of induction, etc) due to the differences in VOC dynamics described via these methods that would be hard to observe in other ways. If so, there will be a need to adopt robust methods such as this to describe the salient features of those dynamics. At present, I do not believe there is evidence one way or another as to whether the subtle differences in VOC dynamics have large consequences.

      The paper sells itself as describing a new technique for describing response curves generally across biological systems, but it only uses this technique to look at the dynamics of induced plant volatiles. I believe to show general utility of this approach, a wider range of examples of plastic responses to stimuli across organismal groups would be needed. I am, however, convinced that this approach is both novel and useful within the scope in which the examples are shown (i.e. in describing the dynamics of induced plant responses). Some of the text purporting novelty in uncovering shared and divergent responses across the tree of life seems pretty overstated.

      Much of the introductory and discussion text is quite broad, and I wonder if the technique is really meant to be applicable to the specific case that is described (repeated measures of an induced volatile response). Likewise, there has been considerable work in such realms as behavioral science, function-valued traits (e.g. Stinchcombe et al 2012), performance curves (Kingsolver various papers), etc to describe dynamic or variable responses phenomenologically, and there are approaches including GAMs, parametric curve fitting, and other techniques that probably report the same salient features as the approach here. Indeed, there are already statistical techniques to assess the macroevolution of response curves (e.g. Goolsby 2015) and wide discussions as to how to compare function-based responses among organisms (The Functional Phylogenies Group 2012). So in the broad scheme of biology, I am not sure I'm convinced of the novelty of the approach. However, I believe it is novel within the context in which it is used here. The salient part of the methods is that it uses predefined attributes of dynamics (onset, duration, etc) based on a gamma distribution that the researchers (with good reason) believe to be biologically meaningful. This is in contrast to multivariate approaches (e.g. Izem et al 2005) that attempt to find salient dynamic features in a less constrained way.

      I would have liked to see a clear description of model fits (e.g., how much of the variation in the real data is described by the fitted model). This seems important because there are quite a number of constraints placed on model fitting - so presumably when a model blind to those constraints picks unrealistic parameters, that would suggest that the constrained model probably does not fit the data all that well.

      I am curious about the normalization process in the 'normalized emission' that is analyzed throughout the study. Normalization to leaf size makes sense, though I was less clear about L459: "Additionally, values were normalized to the maximum response observed in each experiment, yielding a range of positive values < 1." Why was this needed? Is the 'maximum response observed in each experiment' across all plants/compounds/treatments or within a single plant? In general, is there a way of reporting VOC emission rates in absolute values (e.g. umol / Liter air)? Normalization would presumably not impact most curve properties very much, but it could have effects on 'integral', and the need for within-experiment normalization would suggest a lack of transferability or comparability among datasets from different experiments (at least as regards 'integral'), which is suggested as a major advantage of this approach in the discussion.

    1. eLife Assessment

      This important study presents a cell-based screen for small-molecule activators of GCN2, an eIF2α kinase that regulates the Integrated Stress Response under diverse stress conditions. The authors identify a compound as a potent GCN2 activator with GCN1-independent activity under the tested conditions, providing a new pharmacological tool for probing GCN2 regulation. The revised manuscript provides compelling support for the central conclusions through a clearer description of the screening workflow, extended kinetic analyses, and demonstration that the identified compound causes a GCN2-dependent reduction in protein synthesis.

    2. Reviewer #1 (Public review):

      Summary:

      This manuscript describes a chemical screen for activators of the eIF2 kinase GCN2 (EIF2AK4) in the integrated stress response (ISR). Recently, reported inhibitors of GCN2 and other protein kinases have been shown at certain concentrations to paradoxically activate GCN2. The study uses CHO cells and ISR reporter screens to identify a number of GCN2 activator compounds, including a potent "compound 20." These activators have implications for the development of new therapies for ISR-related diseases. For example, although not directly pursued in this study, these GCN2 activators could be helpful for the treatment of PVOD, which is reported for patients with certain GCN2 loss-of-function mutations. The identified activators are also suggested to engage with the GCN2 directly and can function devoid of GCN1, a co-activator of GCN2.

      Strengths:

      The manuscript appears to be a largely rigorous study that flows in a logical manner. The topic is interesting and significant.

      Weaknesses:

      Portions of the manuscript are not fully clear. There are some experimental presentation and design concerns that should be addressed to support the stated conclusions.

    3. Reviewer #2 (Public review):

      Summary:

      In this manuscript Zhu, Emanuelli and colleagues describe a novel pharmacological activator of the Integrated Stress Response kinase GCN2. The work is conclusive and biochemically solid. This work significantly adds to the pharmacological arsenal targeting the ISR and in particular GCN2.

      Strengths:

      Strong biochemistry, novel molecular activator of GCN2 (GCN1 independent).

      Weaknesses:

      Rationale for the screen not exploited in the results (e.g. pathogenic GCN2 mutants), lots of cell-based read-outs not endogenous.

      Comments on revised version.

      The authors did a great job at addressing my initial critique on their manuscript and consequently I have no further comment.

    4. Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors describe the results of a high throughput screen for small molecule activators of GCN2. Ultimately, they find 3 promising compounds. One of these three, compound 20 (C20) is of the most interest both for its potency and specificity. The major new finding is that this molecule appears to activate GCN2 independent of GCN1, which suggests that it works by a potentially novel mechanism. Biochemical analysis suggests that each bind in the ATP binding pocket of GCN2, and that at least in vitro C20 is a potent agonist. Structural modeling provides insight into how the three compounds might dock in the pocket and generates testable hypotheses as to why C20 perhaps acts through a different mechanism than other molecules.

      Strengths:

      Of the 3 compounds identified by the authors, C20 is of the most interest, not just for its intriguing mechanistic distinction as being GCN1-independent (shown genetically in two distinct cell lines, CHO and 293T, and in contrast to other GCN2 activators) but also for its potency. Ultimately, C20 might be a tool for providing mechanistic insight into the details of GCN2 activation and regulation and could be exploited therapeutically.

      Weaknesses:

      The chief limitation of this work is that the experiments exploring the effects of C20 on ISR output in cells are limited, so how useful these compounds are both experimentally and therapeutically remains to be determined.

      Comments on revised version.

      The authors have satisfactorily addressed my comments. A more extensive analysis of UPR signaling in cells (transcription and cell death in particular) would have further strengthened the paper, but that can be left to future work.

    5. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a chemical screen for activators of the eIF2 kinase GCN2 (EIF2AK4) in the integrated stress response (ISR). Recently, reported inhibitors of GCN2 and other protein kinases have been shown at certain concentrations to paradoxically activate GCN2. The study uses CHO cells and ISR reporter screens to identify a number of GCN2 activator compounds, including a potent "compound 20." These activators have implications for the development of new therapies for ISR-related diseases. For example, although not directly pursued in this study, these GCN2 activators could be helpful for the treatment of PVOD, which is reported for patients with certain GCN2 loss-of-function mutations. The identified activators are also suggested to engage with the GCN2 directly and can function while devoid of GCN1, a co-activator of GCN2.

      Strengths:

      The manuscript appears to be a largely rigorous study that flows in a logical manner. The topic is interesting and significant.

      Weaknesses:

      Portions of the manuscript are not fully clear. Some experimental presentation and design concerns should be addressed to support the stated conclusions.

      We thank the reviewer for their supportive comments. We agree that portions of the manuscript were not fully clear and that some aspects of the experimental presentation and design required clarification.

      To address this, we have revised the manuscript to make the experimental logic more transparent. In particular, we now explain more clearly the rationale for the screening strategy, including the use of histidinol as a canonical GCN2 activator, latrunculin A as a modulator of PPP1R15A-mediated eIF2α dephosphorylation, and tunicamycin as a PERK-dependent ER-stress control. We also clarify why a submaximal concentration of histidinol was used: this was intended to reveal compounds that enhance ISR signalling when GCN2 is partially activated.

      We have clarified the use of the two ATF4 reporter systems. The ATF4–NanoLuc reporter was used for sensitive primary screening, whereas the ATF4–luc2 reporter was used as a more stringent orthogonal assay to prioritise robust ISR activators. We now state explicitly why some initial hits were not retained after testing in the second reporter line, and why the NanoLuc system was subsequently used again for mechanistic experiments.

      Finally, we have revised the presentation of the orthogonal validation steps to make clearer how they support the stated conclusions. These include assays designed to distinguish GCN2-dependent ISR activation from indirect activation through ER stress, additional analysis of GCN2 dependence, and clearer interpretation of biochemical and docking data.

      We hope that these revisions address the reviewer’s concern that the experimental design and data presentation needed to be made clearer in order to support the conclusions.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Zhu, Emanuelli, and colleagues describe a novel pharmacological activator of the Integrated Stress Response kinase GCN2. The work is conclusive and biochemically solid. This work significantly adds to the pharmacological arsenal targeting the ISR and, in particular, GCN2.

      Strengths:

      Strong biochemistry, novel molecular activator of GCN2 (GCN1 independent).

      Weaknesses:

      The rationale for the screen is not exploited in the results (e.g., pathogenic GCN2 mutants), and lots of cell-based read-outs are not endogenous.

      We thank this reviewer for their positive assessment of the work. We address the three major concerns in turn below.

      Major points

      (1) Regarding the justification of the work. Since the authors justify the screen for GCN2 activators with loss-of-function mutants associated with diseases, it would be of interest to evaluate whether the best compounds identified in the study are indeed able to prompt activation of those mutants (or at least of the most prevalent). This approach could actually go in parallel with the docking experiments carried out in the last figure of the manuscript, where mutants could be modelized as well.

      To address this point, we tested whether the lead compounds could activate disease-associated GCN2 variants linked to pulmonary veno-occlusive disease. In contrast to GCN2iB, the new compounds did not activate these variants. We now state this explicitly in the manuscript, thereby clarifying that although the compounds identify a new mode of GCN2 activation, they do not rescue the pathogenic GCN2 variants tested here.

      Results

      “Moreover, in contrast to GCN2iB (17), the current compounds did not activate disease-associated GCN2 variants linked to PVOD [data not shown].”

      (2) The compounds are only tested using « artificial » proximal signaling outputs. It would be interesting to evaluate whether the best identified compounds are capable of prompting endogenous eIF2alpha phosphorylation in cellular models.

      We thank the reviewer for this suggestion. Detecting eIF2α phosphorylation following activation of GCN2 is technically challenging and typically produces weaker signals compared to activation of other ISR kinases, such as PERK (e.g. by thapsigargin). For this reason, many studies rely on downstream reporter assays to monitor GCN2 activity. To address the reviewer’s concern, we have now included an orthogonal readout of ISR activation by assessing global translation using a puromycin incorporation assay. Using this approach, we show that compound 20 significantly reduces translation, and importantly, this effect is attenuated in GCN2-deficient cells, supporting a GCN2-dependent mechanism.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      (3) Other GCN2 activators (other than GCN2iB, e.g., HC-7366) were recently identified. In this context, it would be of interest to carry out a small benchmarking study to evaluate how the compounds identified in the current study perform against the previously identified molecules.

      We thank the reviewer for this suggestion. In response, we obtained HC-7366 and assessed its activity alongside our compounds in the CHO ATF4::NanoLuc reporter assay. In this system, compound 20 demonstrated greater potency than HC-7366 (see reviewer figure below). However, we note that HC-7366 showed relatively limited activity in CHO cells in our hands, despite previously reported strong effects in other cellular systems and in vivo models. This context-dependent activity makes direct benchmarking difficult. Accordingly, we have included this comparison in the revised manuscript and discuss this limitation in the Discussion.

      Discussion

      “We also evaluated the reported GCN2 activator HC-7366 in our CHO ATF4::NanoLuc reporter system. In this context, HC-7366 showed limited activity relative to compound 20, despite its reported efficacy in other cellular systems and in vivo (data not shown). This highlights potential context dependence in small‑molecule activation of GCN2 and limits direct cross-study comparison.”

      Author response image 1.

      ISR activation by compound 20 and GC-7366 in CHO cells

      Normalised fold-change in ATF4 signal in CHO ATF4::NanoLuc reporter cells treated for 19 hours with Compound 20 or HC-7366. DMSO was used as vehicle control. (representative experiment, mean ± SEM, n=3 technical replicates).

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors describe the results of a high-throughput screen for small-molecule activators of GCN2. Ultimately, they find 3 promising compounds. One of these three, compound 20 (C20), is of the most interest both for its potency and specificity. The major new finding is that this molecule appears to activate GCN2 independent of GCN1, which suggests that it works by a potentially novel mechanism. Biochemical analysis suggests that each binds in the ATP-binding pocket of GCN2, and that at least in vitro, C20 is a potent agonist. Structural modeling provides insight into how the three compounds might dock in the pocket and generates testable hypotheses as to why C20 perhaps acts through a different mechanism than other molecules.

      We agree that GCN1-independent activation suggests a potentially distinct mechanism of action. While we are currently unable to define the mechanistic basis underlying the GCN1-independence of compound 20, prior work provides some relevant context. Recent studies have shown that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions, notably in the context of the GCN2 E26A mutant [Carlson, 2023]. This observation raises the possibility that, under certain conditions, engagement of the kinase domain, potentially via the ATP-binding pocket, may bypass the requirement for GCN1. However, in our system, we did not observe GCN1-independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow or context-dependent window for such activity, or differences between wild-type and mutant GCN2. These findings suggest that GCN1-independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds, although further work will be required to define the underlying mechanism for compound 20. We have added the following to the main text:

      Discussion

      “Recent work suggests that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions using a GCN2 E26A mutant (9). In our hands, we did not observe GCN1‑independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow concentration window for GCN1‑independent activation or context‑dependent effects of the E26A mutation. These findings raise the possibility that GCN1‑independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds.”

      Strengths:

      Of the 3 compounds identified by the authors, C20 is the most interesting, not just for its intriguing mechanistic distinction as being GCN1-independent (shown genetically in two distinct cell lines, CHO and 293T in Figure 4, and in contrast to other GCN2 activators) but also for its potency. In in-cellulo assays, compound 21 appears as more of an ISR enhancer than an activator per se, and although compound 18 and compound 21 lead to upregulation of the ISR targets (Figure 2), that degree of upregulation is probably not significantly different from that induced by those compounds in Gcn2-/- cells. For C20, the effect appears stronger (although it is unclear whether the authors performed statistical analysis comparing the two genotypes in Figure 2D). In Figure 3, only C20 activates the ISR robustly in both CHO and 293T. Ultimately, C20 might be a tool for providing mechanistic insight into the details of GCN2 activation and regulation, and could be exploited therapeutically.

      Prompted by this suggestion, we assessed C20 in two additional commonly used cell lines: human colon carcinoma HCT116 cells and African green monkey COS7 cells. C20 showed no activity in these models. In contrast, primary mesothelioma cells (Mesobank T12) exhibited robust PPP1R15A induction in response to the compound. The following text has been added to the manuscript.

      Results

      “We went on to examine downstream cellular consequences of GCN2 activation in multiple models. While compounds did not induce detectable ISR signalling in HCT116 or COS‑7 cells under the conditions tested, induction of PPP1R15A was observed in Mesobank T12 primary mesothelioma cells, indicating context-dependent biological responses… [data not shown].”

      Weaknesses:

      There are some limitations to the existing work. As the authors acknowledge, they do not use any of the compounds in animals; their in vivo efficacy, toxicity, and pharmacokinetics are unknown. But even in the context of the in cellulo experiments, it is puzzling that none of the three compounds, including C20, has any effects in HeLa cells when Neratinib does. It's beyond the scope of this paper to address definitively why that is, but it would at least be reassuring to know that C20 activates the ISR in a wider range of cells, including ideally some primary, non-immortalized cells. In addition, the ISR is a complex, feedback-regulated response whose output varies depending on the time point examined. The in cellulo analysis in this paper is limited to reporter assays at 18 hours and qRT-PCR assays at 4 and 8 hours. A more extensive examination of the behaviour of the relevant ISR mRNAs and proteins (eIF2, ATF4, CHOP, cell viability, etc.) for C20 across a more extensive time course would give the reader a clearer sense of how this molecule affects ISR output.

      We thank the reviewer for this insightful suggestion. To address the need for a more comprehensive assessment of ISR signalling, we have extended our analysis across a broader time course and incorporated additional functional readouts. Using the ATF4-NanoLuc reporter, compounds 20 and 21 exhibit peak ISR activation at approximately 6-8 h in wild-type cells. In parallel, we assessed global mRNA translation using puromycin incorporation and found that compound 20, but not compound 21, induces a progressive reduction in translation over this period, which is dependent on GCN2. While we agree that direct measurement of upstream ISR markers such as eIF2α phosphorylation can be informative, detection of GCN2-mediated eIF2α phosphorylation is technically challenging and often less robust than activation of other ISR kinases (e.g. PERK). For this reason, we have prioritised orthogonal downstream functional readouts, including reporter activity and translational output, to capture ISR pathway engagement. These additional data provide a clearer picture of the kinetics and functional consequences of compound-induced ISR activation and have been incorporated into the revised manuscript.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      I also find it a bit strange that the authors describe C20 as "demonstrat(ing) weak inhibition of ... PKR" - the measured IC50 is ~4 μM, which is right around its EC50 for GCN2 activation. This raises the confounding possibility that C20 would simultaneously activate GCN2 while inhibiting PKR. While perhaps inhibition of PKR is not relevant under the conditions when GCN2 would be activated either experimentally or therapeutically, examining in cells the effects of C20 on GCN2 and PKR across a dose range would shed light on whether this cross-reactivity is likely to be of concern.

      We thank the reviewer for highlighting compound 20 as the most interesting lead compound and for recognising its apparent ability to activate GCN2 independently of GCN1. The reviewer identified several limitations relating to cell-type specificity, the temporal behaviour of ISR activation, and possible PKR cross-reactivity.

      In response to the concern about cell-type specificity, we tested compound 20 in additional cellular models. Compound 20 did not induce detectable ISR signalling in HCT116 or COS-7 cells under the conditions tested, consistent with the reviewer’s observation that activity is not universal across cell types. However, we observed induction of PPP1R15A in primary Mesobank T12 mesothelioma cells. We have therefore revised the manuscript to present compound 20 activity as cell-context dependent rather than broadly generalisable across all cell types.

      To address the reviewer’s concern that the ISR output was examined only at limited time points, we extended the time-course analysis for compounds 20 and 21. Using live-cell ATF4::NanoLuc reporter measurements, both compounds showed maximal reporter activation at approximately 6-8 hours. We also measured translational output by puromycin incorporation and found that compound 20, but not compound 21, caused a progressive GCN2-dependent reduction in translation. These data provide a clearer view of the kinetics and functional consequences of compound 20-mediated ISR activation.

      The reviewer also noted that compound 20 inhibits PKR in vitro at concentrations close to those required for GCN2 activation in cells. We agree that this is an important potential liability. Because PKR signalling was not robustly or reproducibly inducible in our CHO-based reporter system, we were unable to perform a reliable cellular dose-response analysis of PKR engagement in the present study. We have therefore revised the Discussion to acknowledge kinase cross-reactivity, including possible PKR inhibition, as an important limitation and an issue for future development of this chemical series.

      Finally, we have moderated our mechanistic interpretation of compound 20. Although the data support direct engagement of GCN2 and suggest a mechanism distinct from canonical GCN1-dependent activation, we now discuss GCN1-independent activation more cautiously and in the context of prior reports that GCN2iB can display GCN1-independent activity under specific experimental conditions.

      Discussion

      “While the functional relevance of PKR inhibition in our cellular systems is uncertain, these observations highlight the potential for kinase cross-reactivity, which will be important to address in future studies.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):<br /> (1) The description of the chemical screen for Gcn2 activators is not sufficiently clear and detailed. a) Briefly and early on provide the rationales (modes of action) for using histindinol and latruculin A. Explain further the rationale in Figure 2, outlining the purpose for the combined compound + submaximal dose of histindinol.

      The text has been amended.

      Results

      “Histidinol activates GCN2 by inhibiting histidyl‑tRNA synthetase, leading to the accumulation of uncharged tRNAHis. This uncharged tRNA binds to GCN2, relieving its autoinhibition and activating the kinase (35). Latrunculin A sequesters G‑actin, thereby inhibiting PPP1R15A activity (36, 37). Tunicamycin inhibits protein glycosylation in the endoplasmic reticulum (ER), resulting in activation of PERK (38).”

      “This submaximal concentration was used to allow detection of compounds that enhance ISR signalling when GCN2 is partially activated.”

      b) What is unique about the second CHO: ATF4-luc2 reporter line? Why do only 89 out of the original 130 compounds induce the ISR in this line versus the original CHO: ATF4-Nanoluc cell line? This is confusing for the reader about how compounds were triaged for characterization.

      The ATF4‑Nanoluc and ATF4‑luc2 reporter lines differ only in the luciferase used, but this has important practical consequences. The Nanoluc reporter is substantially more sensitive, so it was used for the primary screen to detect even weak ISR activation. The luc2 reporter has lower sensitivity and a narrower dynamic range, making it a more stringent orthogonal assay. As a result, not all hits from the Nanoluc screen (130 compounds) reproduced in the luc2 line; the 89 compounds retained are those that robustly activate the ISR under these more stringent conditions. This step was therefore used to prioritise stronger, more reproducible activators for downstream characterisation.

      Results

      “While primary screening was performed in ATF4‑Nanoluc lines for maximal sensitivity, hits were subsequently re-tested in a second CHO ATF4::luc2 reporter line as a more stringent orthogonal assay to prioritise robust ISR activators. Of the 130 hits identified in the sensitive Nanoluc screen and passing early toxicity assessment, 89 were confirmed in the luc2 assay, consistent with enrichment for higher-amplitude ISR activators under more stringent detection conditions.”

      c) The study uses a second CHO reporter line in the flow scheme (CHO:ATF4-luc2) and then switches back to an ATF4-Nanoluc line to establish GCN2 dependence. What is the rationale for switching back to the original reporter line?

      The luc2 reporter line was used as a more stringent, orthogonal validation step to prioritise robust ISR activators. For subsequent mechanistic studies, including assessment of GCN2 dependence, we returned to the ATF4‑Nanoluc line because its higher sensitivity and simpler single‑reagent assay format are better suited to multi‑point measurements and comparative analyses. In effect, the luc2 reporter was used for triage, whereas the Nanoluc system was retained for mechanistic characterisation and downstream screening.

      Results

      “In subsequent mechanistic studies, the ATF4‑Nanoluc reporter was again used to take advantage of its higher sensitivity and simpler assay format for multi‑condition comparisons.”

      d) The rationale for the first orthogonal screen described in the results section to identify inducers of ER stress is not clearly explained. The compounds were already determined to be dependent on GCN2 prior to this test, and one would have thought that this criterion would have covered ER stress and alternative eIF2 kinase activators.

      We agree with the reviewer that, in principle, establishing GCN2 dependence should reduce the likelihood of capturing compounds acting through alternative eIF2α kinases. However, we performed this orthogonal ER stress screen to address two practical considerations. First, high‑throughput screening is inherently prone to false positives, as it is typically conducted at a single concentration and time point, and compound libraries may contain degraded or chemically inconsistent material. We therefore used a lower‑throughput, more controlled ER stress assay with freshly sourced compounds and additional readouts (e.g. CHOP and XBP1) to improve confidence in the hits. Second, despite prior evidence of GCN2 dependence, ER stress signalling via PERK converges on the same downstream endpoints: eIF2α phosphorylation and ATF4 induction. We therefore wished to explicitly exclude compounds that activate the ISR indirectly via ER stress. In practice, this proved important, as the orthogonal assay did identify compounds that induced ER stress, which we subsequently excluded from the lead set.

      Results

      “Although hits were prioritised for GCN2 dependence, we performed an additional orthogonal screen to exclude compounds that activate the ISR indirectly via ER stress, which converges on the same downstream outputs.”

      (2) A major point of the manuscript is that there is GCN1 independence for the small molecule activation of GCN2, and this has not yet been reported. One report for this GCN1 independence is reference 9 [Carlson … Wek 2023] (Figure 5). In this report, low doses of GCN2iB that can activate GCN2 (although by the present manuscript at much lower levels than the identified new compounds) induce ATF4 expression in cells expressing an E26A mutant of GCN2 that is suggested to negate GCN1 binding and enhancement of GCN2 activity. Halofuginone induction of ATF4 expression was thwarted by the GCN2 E26A mutant.

      We thank the reviewer for highlighting this important point. We agree that Carlson et al. (2023) suggest that, under certain conditions, GCN2iB can activate GCN2 independently of GCN1 using the E26A mutant. In our experiments, however, we did not observe GCN1‑independent activation with GCN2iB under the conditions tested, i.e. similar low doses. One possible explanation is that the GCN1‑independent activity reported by Carlson et al. occurs only within a narrow concentration range; their observations were made at very low compound concentrations, whereas higher concentrations may engage additional regulatory mechanisms. In our study, we used concentrations optimised for robust ISR activation, which may mask such effects. We also note that the E26A mutation (E18A in yeast), originally identified by two‑hybrid analysis, disrupts the GCN2-GCN1 interaction but may not completely eliminate all modes of functional coupling under all conditions. Taken together, these observations raise the possibility that GCN1‑independent activation represents a context‑dependent mechanism that may be unmasked only under specific experimental conditions or by particular classes of compounds.

      We have revised the Discussion to acknowledge this prior report explicitly and to clarify how our findings relate to it.

      Discussion

      “Recent work suggests that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions using a GCN2 E26A mutant (9). In our hands, we did not observe GCN1‑independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow concentration window for GCN1‑independent activation or context‑dependent effects of the E26A mutation. These findings raise the possibility that GCN1‑independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds.”

      (3) The authors state that the ISR was exaggerated in Ppp1r15a KO cells. It would be helpful to include statistical analyses to support this statement.

      Thank you for enabling us to be more precise. New text added:

      Results

      “Activation of the ISR by tunicamycin was exaggerated in the Ppp1r15a<sup>-/-</sup> cells owing to their defective dephosphorylation of eIF2a (wild type vs Ppp1r15a<sup>-/-</sup>, p<0.05).”

      (4) The results state that Chop and Ppp1r15a mRNAs were measured following 4 hours of treatment with compound 18, 20, or 21, but Figure 2D shows treatment from 0 to 8 hours? It appears that the compound still induces these mRNAs in GCN2 KO cells, possibly with delayed kinetics. A lengthened time course study would help determine if this is indeed the case.

      We thank the reviewer for this careful observation. To address the reviewer’s point regarding delayed or GCN2‑independent signalling, we have extended our analysis using compounds 20 and 21, which are more tractable experimentally (solubility). Using the ATF4‑Nanoluc reporter, both compounds show peak ISR activation at ~6–8 h in wild-type cells over an extended time course. In parallel, functional readouts of mRNA translation (puromycin incorporation) demonstrate that compound 20, but not 21, induces a progressive, GCN2‑dependent reduction in translation over this period. These clarify the temporal aspects of signalling by these two compounds.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      Legend

      “Supplementary Figure S1. Kinetics of responses to compounds 20 and 21

      (A-B) Wild-type CHO cells stably expressing the ATF4::nanoLuc-PEST reporter were treated with Nano-Glo and either (A) 13mM compound 20 or (B) 13mM compound 21. Median bioluminescence (fold change normalised to DMSO control) ± 95% confidence. Representative experiment (n=4 technical repeats). (C-F) Representative immunoblot of lysates from wild-type or Eif2ak4<sup>-/-</sup> CHO cells treated with 10μM compound 20 or 7.5μM 21 for the indicated times. Immediately before harvesting, cells were treated with 10μg/mL puromycin to label newly synthesised polypeptides. “-“ indicates cells not incubated with puromycin. “U” cells were treated with puromycin but without test compound. “CHX” represents the cycloheximide control (100μg/mL). Molecular size in kDa. (E-F) Quantification of puromycinylated proteins normalised to GAPDH. Mean ± SEM. CHO WT (black) and Eif2ak4<sup>-/-</sup> cells (turquoise. N = 4 independent experiments. Two-way ANOVA with Šídák's multiple comparisons test; ***: p ≤ 0.001.”

      (5) In the section describing the differences between cell lines in the ability of compounds to induce the ISR, this is difficult for the reader to interpret, as no controls are included. How does histidinol (or other canonical inducers of the ISR) behave in the three reporter assays (CHO, 293T, and HeLa)?

      As requested, we now provide ATF4::Nanoluc reporter activation (3mM, 20 hours because of this drug’s slow kinetics)

      Results

      “To benchmark ISR activation in these models, each cell type was treated with 3mM histidinol (Figure S2). Reporter activation was most robust in CHO cells, followed by 293T cells, then HeLa cells.”

      Discussion

      “Moreover, histidinol-induced ISR activation showed a clear hierarchy across cell lines, with CHO cells being the most responsive and HeLa cells the least.”

      Legend

      “Supplementary Figure S2. Cell-type differences in response to histidinol

      Fold-change of ATF4::NanoLuc reporter signal in HEK293T, HeLa and CHO cells transiently transfected with reporter and treated for 20 hours with 3mM histidinol. Fold-change calculated relative to vehicle control. Mean ± SEM).”

      We thank the reviewer for this important question. However, we respectfully disagree that a direct correspondence between the concentrations required for target engagement in the BRET assay and for ISR activation in functional assays should necessarily be expected. BRET (including NanoBRET) is a target engagement assay that measures compound binding to the protein in intact cells, typically by competition with a labelled tracer, and thus reports on apparent intracellular affinity and occupancy rather than downstream biological effect (Robers 2019, PMID 30519940). By contrast, ISR activation is a functional readout that reflects amplification through signalling networks, and can be influenced by multiple additional variables including pathway non-linearity, feedback, and kinase regulation. Consequently, it is well established that potencies derived from target engagement assays do not always align with those measured in functional assays. For example, intracellular kinase profiling studies using NanoBRET have demonstrated systematic potency offsets between binding/engagement measurements and downstream cellular activity, arising from factors such as intracellular ATP competition and pathway context (PMID Capener 2026, PMID 41495225). More generally, target engagement assays provide a quantitative measure of binding, whereas functional assays measure biological outcome, and these readouts need not coincide because they capture distinct aspects of a compound’s mechanism of action. Accordingly, we interpret our BRET data as evidence of direct interaction with GCN2 in cells, rather than as a predictor of the concentration required to activate the ISR. The observation that higher concentrations are required in the BRET assay is therefore not unexpected and does not argue against a requirement for kinase-domain engagement in ISR activation. Instead, it reflects the different mechanistic endpoints captured by the two assay formats.

      We will clarify this point explicitly in the revised manuscript.

      Results

      “The concentrations required to detect target engagement in NanoBRET assays did not directly mirror those required for ISR activation, reflecting the distinction between ligand binding and downstream pathway output.”

      (7) In Figure 5E, the authors suggest that compounds 18 and 20 are non-competitive inhibitors of GCN2 since the Vmax increases with increasing ATP concentration. What is the Km for ATP in the absence or presence of compound 18 or 20? It would be helpful to include progress curves as supplementary data to support the Vmax plots in Fig. 5E. Consider providing more specific units (currently arbitrary units) for the y-axis.

      We thank the reviewer for this insightful comment and agree that our original wording overstated the mechanistic interpretation of these data. In particular, the use of the term “non‑competitive” is not well supported by the current analysis and may be misleading, especially given that our data are consistent with binding within or proximal to the ATP-binding pocket. We have therefore revised the text to remove this designation and instead describe the data more conservatively in terms of changes in apparent Vmax, without assigning a specific inhibition mechanism. With respect to kinetic analysis, we agree that full determination of K<sup>m</sub> values and inclusion of progress curves would provide a more rigorous mechanistic interpretation. However, given the primary focus of this manuscript on identifying and characterising small‑molecule activators of GCN2 in cells, we believe that a detailed steady‑state kinetic analysis would be beyond the scope of the current study. We have therefore moderated our conclusions accordingly and now present these data as preliminary kinetic observations rather than definitive evidence of inhibition modality.

      Results

      “Compounds 18 and 20 altered the apparent kinetic parameters of GCN2, including an increase in the observed V<sub>max</sub>; however, these data do not allow assignment of a specific inhibition modality.”

      (8) The full-length GCN2 assay presented in Figure 5G appears to be unresponsive to uncharged tRNA, a known regulator of GCN2. The statement that compound 20 induces eIF2 phosphorylation to a greater extent than tRNAs is true for the in vitro assay, but arguably is because the in vitro assays do not recapitulate the in vivo arrangement.

      We thank the reviewer for this comment. We respectfully disagree that the assay is unresponsive to uncharged tRNA. In our hands, GCN2 does exhibit activation in response to tRNA; however, the magnitude of this effect is modest (~4‑fold) compared to the substantially stronger activation observed with compound 20 (~40‑fold). As a result, the tRNA response can appear compressed when both are plotted on the same scale. We agree with the reviewer that the in vitro assay does not fully recapitulate the in vivo regulatory environment, where factors such as GCN1 and ribosome association are known to potentiate GCN2 activation. This limitation likely explains the relatively weaker response to tRNA under our assay conditions and was a key motivation for incorporating cellular assays in our study. Interestingly, the marked difference in activation magnitude between tRNA and compound 20 in vitro raises the possibility that compound-mediated activation may, at least in part, bypass regulatory features that normally constrain GCN2 activity in a GCN1‑dependent manner. While we have not directly tested this hypothesis, we will temper the wording and include this as a speculative point in the Discussion.

      Results

      “Of note, uncharged tRNA produced a modest (~4‑fold) activation of GCN2 under these conditions, whereas compound 20 induced substantially greater (~40‑fold) activation.”

      Discussion

      “The markedly greater activation observed with compound 20 compared with uncharged tRNA in vitro raises the possibility that such compounds may partially bypass regulatory constraints on GCN2 activation, including those normally mediated by GCN1.”

      (9) Using purified GCN2 kinase domain at low ATP concentrations (10 μM), compounds 18 and 20 were shown not to inhibit GCN2 up to concentrations of 3 μM. In previous assays, much higher concentrations of compounds 18 and 20 were used to inhibit GCN2. Why were different concentrations used? This makes this interpretation of this data difficult for the reader to draw conclusions.

      We thank the reviewer for this comment and agree that the use of different concentration ranges across assays may not have been sufficiently clear. The kinase‑domain assay performed at low ATP (10 μM) was specifically designed to assess whether compounds 18 and 20 have a propensity to inhibit GCN2 under conditions that sensitise detection of ATP‑competitive effects and facilitate comparison with related eIF2α kinases. This assay was therefore optimised for detecting inhibition, rather than activation. In contrast, the higher concentrations used in other experiments were selected to robustly measure ISR activation in cellular or full‑length protein contexts, where higher compound exposure is required to observe downstream signalling outputs. These two assay systems therefore address distinct mechanistic questions—targeting inhibition under controlled biochemical conditions versus activation in more complex functional settings—and are not directly comparable in terms of concentration–response relationships. We will revise the manuscript to clarify this distinction and to emphasise that the kinase‑domain assay was not intended to define the activation potency of the compounds.

      Results

      “This assay was performed at low ATP concentrations to sensitise detection of ATP-competitive inhibition and was not optimised to detect compound-mediated activation of GCN2.”

      (10) In silico docking studies support the binding of compound 20 in the ATP-binding pocket of the GCN2 kinase domain. How is this compatible with the stated ATP non-competitive mechanism?

      We agree with the reviewer that our previous description was misleading. The designation of compounds 18 and 20 as “ATP non‑competitive” is not supported by the available data and is inconsistent with the docking results suggesting binding within the ATP‑binding pocket. We have therefore revised the manuscript to remove this terminology and to describe the kinetic behaviour more cautiously, without assigning a specific mode of inhibition.

      (11) The study does not appear to feature biological assays demonstrating the effects of GCN2 activation. For example, does compound 20 reduce translation or growth of cells in a GCN1/GCN2-dependent manner, or do the compounds overcome PVOD mutations akin to the authors' Hum Mol Genet 2024 Aug 18;33(17):1495-1505 article?

      We thank the reviewer for this suggestion. We agree that defining downstream biological consequences of GCN2 activation is an important goal. We did explore this using several cellular systems; however, these effects were context-dependent and not consistently observed across models. Specifically, compounds did not induce a detectable ISR in HCT116 or COS‑7 cells under the conditions tested, despite responsiveness of these systems to canonical activators such as histidinol. In contrast, we did observe induction of PPP1R15A in primary mesothelioma cells, indicating that biological responses can be elicited in certain cellular contexts. We also tested whether these compounds could rescue disease-associated GCN2 variants linked to PVOD, as previously reported for GCN2iB, but did not observe activation of these mutants. These findings suggest that while the compounds robustly activate GCN2 signalling in reporter assays, downstream biological outputs are context-dependent and may require specific cellular conditions or co-factors. Given the variability across systems, we have limited our conclusions to ISR activation and have not generalised broader biological effects. We will clarify this point in the revised manuscript.

      Results

      “We went on to examine downstream cellular consequences of GCN2 activation in multiple models. While compounds did not induce detectable ISR signalling in HCT116 or COS‑7 cells under the conditions tested, induction of PPP1R15A was observed in Mesobank T12 primary mesothelioma cells, indicating context-dependent biological responses. Moreover, in contrast to GCN2iB (17), the current compounds did not activate disease-associated GCN2 variants linked to PVOD [data not shown].”

      (12) For the control of neratinib and other activators linked with ATP binding that are suggested to be dependent on GCN1 in Fig. 4, include reporter induction by drug treatment in Gcn2-/- cells. It would be helpful to be clear in the list of compounds between those suggested to be direct activators versus those that may create stress that leads to GCN2 activation.

      We thank the reviewer for this suggestion. In response, we have performed additional reporter assays in both WT and GCN2<sup>-/-</sup> cells to assess the dependence of drug-induced ISR activation on GCN2. These experiments reveal clear GCN2 dependence for reporter induction upon treatment with sunitinib, NXP800, WEE1-in-4, Debio0123, gefitinib, and erlotinib, as signal is markedly reduced in GCN2⁻/⁻ cells compared to WT.

      In contrast, dovitinib and AZD1775 show less clear dependence, with relatively low reporter signal even in WT cells (notably lower than observed in Fig. 4C), limiting interpretation. Interestingly, dabrafenib induces stronger reporter activity in GCN2<sup>-/-</sup> cells than in WT, indicating that its effects are independent of GCN2 and may reflect activation of alternative stress or signalling pathways.

      Results

      “To further assess the mechanism of compound-induced ISR activation, we evaluated reporter responses in GCN2-deleted cells (Supplementary Figure S3). Several compounds, including sunitinib, NXP800, WEE1-in-4, Debio0123, gefitinib, and erlotinib, showed reduced reporter activity in GCN2-deficient cells, consistent with GCN2-dependent activation. In contrast, dovitinib, AZD1775 and dabrafenib produced weaker or inconclusive responses even in the paired wild-type lines, limiting analysis. These data allow us to distinguish compounds consistent with direct or GCN2-dependent activation from those more likely to induce ISR indirectly through cellular stress upstream of GCN2. These findings support a distinction between compounds that activate the ISR through GCN2-dependent mechanisms and those that likely act indirectly via alternative stress pathways.”

      Legend

      “Supplementary Figure S3. GCN2-dependence of ISR activation by putative GCN2 agonists

      Normalised fold-change in ATF4 signal in CHO WT (purple) and Gcn2-/- (blue) ATF4::NanoLuc reporter cells treated for 19 hours with a panel of ATP-competitive kinase inhibitors reported to activate GCN2 (at 1 and 3µM, sunitinib used at 3 and 10µM, AZD1175 used at 0.3 and 1µM, gefitinib and erlotinib used at 3 and 10µM). DMSO was used as vehicle control. (n=3; mean ± SEM).”

      Minor Comments:

      (1) In the introduction, the authors state that "Several type 1 and 1.5 kinase inhibitors can activate GCN2 at low concentrations while inhibiting at higher concentrations". What is the evidence that type I inhibitors can activate GCN2?

      Thank you for this query. We believe our initial phrasing was open to misinterpretation. The new wording is

      Introduction

      “Several kinase inhibitors classified as type 1 or type 1.5 with respect to their canonical targets can activate GCN2 at low concentrations while inhibiting it at higher concentrations; however, their binding mode to GCN2 remains undefined.”

      (2) A CHO:ATF4-Nanoluc translation reporter screen was used to screen 123K compounds and divide them into a pool of inhibitors and a pool of activators. The criteria used to make this distinction are not sufficiently described in the manuscript, and screen results are not supplied as supplementary data.

      We thank the reviewer for highlighting the need for greater clarity regarding the screening criteria and reproducibility. Compounds from the primary screen (BioAscent library of 123,222 drug-like compounds performed by the ALBORADA Drug Discovery Institute) were classified based on Z-score thresholds, with activators defined as those with Z-score > 3. To assess robustness, the primary screen data were re-analysed independently. In an initial analysis of 121,000 compounds (excluding plates failing quality control), 6,521 compounds met the activator threshold. Of these, 6,461 overlapped with the original hit list. The small number of discrepancies included: (i) compounds absent from the analysed dataset (e.g. originating from failed plates), and (ii) compounds with Z-scores close to the threshold (typically between −3 and −3.02), consistent with minor analytical variation (e.g. rounding). Overall, these analyses show a high degree of concordance in hit identification, with differences restricted to borderline cases near the selection threshold. We have clarified these criteria below.

      Results

      “Compounds were classified based on Z-score thresholds derived from the primary screen, with activators defined as those with Z-score > 3. An initial analysis of 121,000 compounds, excluding plates failing quality control, identified 6,521 activators, of which 6,461 overlapped with the subset selected for follow-up screening. Minor discrepancies were restricted to compounds absent from the analysed dataset (e.g. originating from failed plates) or those with Z-scores close to the threshold (3 to 3.02), consistent with limited analytical variation. Using this approach, we assembled a subset of 6,461 compounds enriched for potential ISR activators and screened these in 384-well format at 10 µM for 16 hours.”

      Reviewer #3 (Recommendations for the authors):

      (1) The legend to Figure 4 should read "Compound 20 displays GCN1 independence", not "dependence".

      Thank you for spotting this error. We have made the correction.

      (2) The text describes Figure 2D as examining 4h, but it also examines 8h.

      We have amended the text to:

      “After validating their effects using the assays described above, we next confirmed activation of the ISR by these compounds at the transcriptional level at 4 and 8 hours.”

      (3) Unless I'm misreading Table S1, compound Z134826202 is listed as activating the ISR, but the authors describe it in the text as inactive.

      Thank you. We have corrected this error.

    1. eLife Assessment

      This important paper describes the role of the Pre-rRNA in meiotic sex chromosome inactivation in mouse spermatocytes. The cytological analyses of nucleolar components and the chemical inhibition of RNA polymerase I for rDNA transcription provided solid evidence, supporting the authors' conclusions. However, the results were not well described or explained in the text, making the logic difficult to follow. This paper will be of interest to researchers in meiotic chromosome structure and the nucleolus.

    2. Reviewer #1 (Public review):

      The authors show that during prophase I of male meiosis, nucleoli disassemble and nucleolar components relocalize to the sex chromosome (XY) body. They further demonstrate that this process is regulated by the ATR-dependent signaling pathway that mediates meiotic sex chromosome inactivation (MSCI). Pharmacological disruption of pre-rRNA synthesis using the RNA polymerase I inhibitor BMH-21 leads to the recruitment of RNA polymerase II to the sex chromosomes and ectopic expression of sex chromosome-linked genes. These findings uncover a previously unrecognized role for pre-rRNAs in maintaining transcriptional silencing during meiosis. The study employs a combination of cell biology, genetics, and genomics approaches, and the conclusions are supported by compelling, well-organized data.

      Comments:

      (1) The current study focuses on transcriptional regulation of the sex chromosomes. It would be interesting to know whether perturbation of pre-rRNA synthesis also affects transcription of autosomal genes.

      (2) Is ribosome biogenesis still active during prophase I of male meiosis? Additional discussion of the timing and extent of rRNA synthesis at this stage would help place the findings in a broader biological context.

      (3) A recent preprint reports active RNA polymerase II-mediated transcription of Y chromosome genes within nucleolus-like bodies (NLBs) during prophase I of meiosis in Drosophila male germ cells (https://doi.org/10.64898/2026.05.20.726666). These findings suggest that the meiotic nucleolus may have species-specific roles in regulating sex chromosome gene expression. It would be valuable for the authors to discuss how their findings compare with these observations and the potential evolutionary implications.

    3. Reviewer #2 (Public review):

      Summary:

      The authors showed the localization pattern of nucleolus components, including Pre-rRNA, a precursor of rRNAs, changes during meiotic prophase I, particularly with the localization of these nucleolar components to the X-Y body, which shows inactivation of RNA polymerase II transcription, during pachynema. The localization of Pre-rRNA depends on ATR kinase and gammaH2AX. The chemical inhibition of rRNA transcription disrupts the binding of pre-rRNA to the X-Y body and suppresses the inhibition of the RNA polymerase II-mediated transcription on the sex chromosomes.

      Strengths:

      The cytological analysis, combined with the chemical inhibition, provided solid evidence to support the idea that, together with the remodeling of the nucleolus structure, pre-rRNA is an essential component of sex chromosome inactivation in male mouse meiosis. The role of pre-rRNA in sex chromosome inactivation in male meiosis helps our understanding of how the X-Y body, which would be a biological condensate, would be formed; e.g. for example, this Pre-rRNA may promote phase separation.

      Weaknesses:

      However, there is limited information on how Pre-rRNA is recruited to only sex chromosomes and how the RNA promotes the inactivation of sex chromosomes. Of course, these will be a target of future study. One major weakness of this paper is a poor description of the results, with fair presentation and interpretation of the data.

    1. eLife Assessment

      The authors addressed a significant biological question, namely the role of glutamine metabolism in humoral responses, and they obtained solid conclusions. The strength of this study is that the authors used state-of-the-art transgenic mouse models together with in vitro analysis, thereby providing important insights into the question posed. The manuscript has been further substantiated by adding more appropriate experimental controls and describing more in-depth functionality/physiological relevance.

    2. Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigate the functional requirements for glutamine and glutaminolysis in antibody responses. The authors first demonstrate that the concentrations of glutamine in lymph nodes are substantially lower than in plasma, and that at these levels, glutamine is limiting for plasma cell differentiation in vitro. The authors go on to use genetic mouse models in which B cells are deficient in glutaminase 1 (Gls), the glucose transporter Slc2a1, and/or mitochondrial pyruvate carrier 2 (Mpc2) to test the importance of these pathways in vivo. Interestingly, deficiency of Gls alone showed clear antibody defects when ovalbumin was used as the immunogen, but not the hapten NP. For the latter response, defects in antibody titers and affinity were observed only when both Gls and either Mpc2 or Slc2a1 were deleted. These latter findings form the basis of the synthetic auxotrophy conclusion. The authors go on to test these conclusions further using in vitro differentiations, Seahorse assays, pharmacological inhibitors, and targeted quantification of specific metabolites and amino acids. Finally, the authors document reduced STAT3 and STAT1 phosphorylation in response to IL-21 and interferon (both type 1 and 2), respectively, when both glutaminolysis and mitochondrial pyruvate metabolism are prevented.

      Strengths:

      (1) The main strength of the manuscript is the overall breadth of experiments performed. Orthogonal experiments are performed using genetic models, pharmacological inhibitors, in vitro assays, and in vivo experiments to support the claims. Multiple antigens are used as test immunogens--this is particularly important given the differing results.

      (2) B cell metabolism is an area of interest but understudied relative to other cell types in the immune system.

      (3) The importance of metabolic flexibility and caution when interpreting negative results is made clear from this study.

      Weaknesses:

      (1) All of the in vivo studies were done in the context of boosters at 3 weeks and recall responses 1 week later. Primary responses, including germinal centers, may still be ongoing at 3 weeks after the initial immunization and defects in GCs may contribute to the findings. Nonetheless, the authors do check antibody levels prior to the boost, and it is likely that most of the observed defects in Gls/Mpc2-deficiency are driven by faulty recall responses.

    3. Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors investigate how glutaminolysis (GLS) and mitochondrial pyruvate import (MPC2) jointly shape B cell fate and the humoral immune response. Using inducible knockout systems and metabolic inhibitors, they uncover a "synthetic auxotrophy": When GLS activity/glutaminolysis is lost together with either GLUT1-mediated glucose uptake or MPC2, B cells fail to upregulate mitochondrial respiration, IL 21/STAT3 and IFN/STAT1 signaling is impaired, and the plasma cell output and antigen-specific antibody titers drop significantly. This work thus demonstrates the promotion of plasma cell differentiation and cytokine signaling through parallel activation of two metabolic pathways. The dataset is technically comprehensive and conceptually novel, but some aspects leave the in vivo and translational significance uncertain.

      Strengths:

      (1) Conceptual novelty: the study goes beyond single-enzyme deletions to reveal conditional metabolic vulnerabilities and fate-deciding mechanisms in B cells.

      (2) Mechanistic depth: the study uncovers a novel "metabolic bottleneck" that impairs mitochondrial respiration and elevates ROS and directly ties these changes to cytokine-receptor signaling. This is both mechanistically compelling and potentially clinically relevant.

      (3) Breadth of models and methods: inducible genetics, pharmacology, metabolomics, seahorse assay, ELISpot/ELISA, RNA-seq, two immunization models.

      (4) Potential clinical angle: the synergy of CB839 with UK5099 and/or hydroxychloroquine hints at a druggable pathway targeting autoantibody-driven diseases.

      Comments on revised version.

      Authors extensively modified the text with great care and provided new data e.g. Fig. 5. Collectively, this is convincing and hence, I have no further comments.

    4. Author response:

      The following is the authors’ response to the current reviews.

      We thank the referees for noting the substantive revisions and for the praise of the work. While we each have somewhat different weightings of likelihood, we feel the appraisals are fair and reasonable.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors addressed an important biological question, namely the role of glutamine metabolism in humoral responses, and they obtained solid conclusions. The strength of this study is that the authors used state-of-the-art transgenic mouse models together with in vitro analysis, thereby providing significant insights into the question posed. The following would strengthen the manuscript: i) adding more in-depth functionality/physiological relevance in the discussion part, and ii) regarding the experiments, the inclusion of more appropriate controls and a clearer and more accurate description of the methods.

      We are grateful for the decision of the Editors to select this submission for in-depth peer review and to the Reviewing Editor and referees for the thoughtful and constructive comments.

      We mostly agree with the specific comments and evaluation of strengths of what the work adds as well as with indications of limitations and caveats that apply to the breadth of conclusions. We have edited the text to be more clear and provide more details about certain aspects of the Methods and Legends. In addition, although we try to avoid Discussion sections that are unduly long or have flights of fancy, we will add to the Discussion as well as edit it for directness about potential relevance, basic explorations of mechanisms, and functionality.

      The revised manuscript also contains new data, some of it dealing with comments of the referees, other additions representing work done while the manuscript was under review. While we would be inclined to do more, the sad practical problem is one of limits placed by both the absence of any grant funds and the institution's terminations (RIFs) of the two experimenters in the lab.

      While we believe the original data interpretable as presented originally, up to a point it nonetheless is good to enhance scope or have even better data and add refinements about some of the technical issues. Ultimately, the question becomes "when is enough enough?"

      In the detailed point-by-point response below, we outline changes prompted by the reviewers. We also comment on a few points more expansively that would be suitable for the paper itself, and offer some skepticism or disagreement, (longer and more detailed explanations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Cho et al. present a comprehensive and multidimensional analysis of glutamine metabolism in the regulation of B cell differentiation and function during immune responses. They further demonstrate how glutamine metabolism interacts with glucose uptake and utilization to modulate key intracellular processes. The manuscript is clearly written, and the experimental approaches are informative and well-executed. The authors provide a detailed mechanistic understanding through the use of both in vivo and in vitro models. The conclusions are well supported by the data, and the findings are novel and impactful. I have only a few, mostly minor, concerns related to data presentation and the rationale for certain experimental choices.

      Detailed Comments:

      (1) In Figure 1b, it is unclear whether total B cells or follicular B cells were used in the assay. Additionally, the in vitro class-switch recombination and plasma cell differentiation experiments were conducted without BCR stimulation, which makes the system appear overly artificial and limits physiological relevance. Although the effects of glutamine concentration on the measured parameters are evident, the results cannot be confidently interpreted as true plasma cell generation or IgG1 class switching under these conditions. The authors should moderate these claims or provide stronger justification for the chosen differentiation strategy. Incorporating a parallel assay with anti-BCR stimulation would improve the rigor and interpretability of these findings.

      We edited the manuscript to be clear that total splenic B cells were used in this set-up figure and the rest of the paper. In addition, we performed new experiments to improve this "set-up figure (Fig. 1)" and moved the older data using alternative experimental conditions to a supplemental figure, Figure 1 - supplement 1. We also used new conditions that included styles of stimulating proliferation and differentiation - to foster an increased sense of generality. The findings in no way change the supported conclusions of the work. Specifically, we used mitogenic stimulation with anti-IgM <sup>+</sup> anti-CD40, all with BAFF, IL-4, and IL-5 in addition to the anti-CD40 stimulation of the original manuscript, bearing in mind excellent work from Aiba et al, Immunity 2006; 24: 259-268, and similar papers. In addition, we added a panel with representative flow cytometric profiles. These new data are presented in Figure 4 - supplement 1 (panels ae).

      To be transparent and add to a more open public discussion (using the virtues of this forum), the senior author and colleagues would caution about whether any in vitro conditions exist that warrant complete confidence. That is the reason for proceeding to immunization experiments in vivo. That is not said to cast doubt on our own in vitro data - there are some experiments (such as those of Fig. 1a-c and associated Fig 1 - supplement 1) that only can be done in vitro or are better done that way (e.g., because of rapid uptake of early apoptotic B cells in vivo).

      For instance: Well-respected papers use the CD40LB and NB21.2D9 systems to activate B cells and generate plasma cells. Those appear to be BCR-independent and yet continue in common use. [We found that these cellular systems (CD40LB; NB21.2D9) cannot be used in experiments with a.a. deprivation or the inhibitors due to effects on the engineered stroma-like cells.] In considering BCR engagement, Reth has published salient points about signaling and concentrations of the Ab, the upshot being that this means of activating mitogenesis and plasma cell differentiation (when the B cells are costimulated via CD40 or TLR (4 or 7/8) is also artificial. Moreover, although Aiba et al, Immunity 2006; 24: 259-268 is a laudable exception, one rarely finds papers using BAFF despite the strong evidence it is an essential part of the equation of B cell regulation in vivo and a cytokine that modulates BCR signaling - in the cultures.

      (2) In Figure 1c, the DMK alone condition is not presented. This hinders readers' ability to properly asses the glutaminolysis dependency of the cells for the measured readouts. Also, CD138<sup>+</sup> in developing PCs goes hand in hand with decreased B220 expression. A representative FACS plot showing the gating strategy for the in vitro PCs should be added as a supplementary figure. Similarly, division number (going all the way to #7) may be tricky to gate and interpret. A representative FACS plot showing the separation of B cells according to their division numbers and a subsequent gating of CD138 or IgG1 in these gates would be ideal for demonstrating the authors' ability to distinguish these populations effectively.

      In the revised manuscript, we have added new experimental data (Figure 1).

      We agree that exact placement of divisions and deconvolution by FlowJow is more fraught than might be thought from presentations in many or most papers. We include the data shown to the right as representative FACS plot(s) with old and new data that illustrate the gating on CTV fluorescence. With the representative examples pasted in here and presented in Fig 1 - supplement 1f, g of the revised manuscript, we will aver that using divisions 0-6, and ≥7 was and is entirely reasonable.

      Ditto for DMK with normal glutamine. However, in the spirit of eLife transparency lacking in many other journals, this comparison is more fraught than the referee comment would make things seem. The concentration tolerated by cells is highly dependent on the medium and glutamine concentration, and perhaps on rates of glutaminolysis (due to its generation of ammonia). In practice, DMK becomes more toxic to B cells unless glutamine is low or glutaminolysis is restricted. Thus, the concentration of DMK that is tolerated and used in Fig. 1b, c can become toxic to the B cells when using the higher levels of glutamine in typical culture media (2 mM or more) - at which point the "normal conditions <sup>+</sup> DMK" "control" involves the surviving cells in conditions with far greater cell death and less population expansion than the "low glutamine <sup>+</sup> DMK". condition.

      (3) A brief explanation should be provided for the exclusive use of IgG1 as the readout in classswitching assays, given that naïve B cells are capable of switching to multiple isotypes. Clarifying why IgG1 was preferentially selected would aid in the interpretation of the results.

      On lines ~112-3 and ~182-5, we edited the text in light of the referee's suggestion that we focus the presentation of serologic data on IgG1 in the immunization experiments. We also rearranged figures and panels to be more explicit and harmonize. That said, and [Brief explanation - IgG1 provides the strongest signal and hence better signal/noise both in vitro and with the alum-based immunizations that are avatars for the adjuvant used in the majority of protein-based vaccines for humans. Perhaps for this reason, the majority of papers on molecular mechanisms seem only to analyze IgG1. Nonetheless, since molecular regulation can differ according to isotype, and the more pro-inflammatory mouse IgG2c is more pertinent to some forms of anti-pathogen immunity and some auto-immune disease models, we believe it valuable to retain these data in supplements to the related Figures.]

      (4) The immunization experiments presented in Figures 1 and 2 are well designed, and the data are comprehensively presented. However, to prevent potential misinterpretation, it should be clarified that the observed differences between NP and OVA immunizations cannot be attributed solely to the chemical nature of the antigens - hapten versus protein. A more significant distinction lies in the route of administration (intraperitoneal vs. intranasal) and the resulting anatomical compartment of the immune response (systemic vs. lung-restricted). This context should be explicitly stated to avoid overinterpretation of the comparative findings.

      We appreciate the positive assessment, and agree with the referee that it is possible the conditions of immune challenge or re-exposure may contribute to the observed differences. We edited the text of the revised manuscript accordingly [lines ~152-153; ~159-160]. Certainly, the difference in how the anti-ova response is elicited compared to the anti-NP response in the same mice or with a bit different an immunization regimen might be another factor - or the major factor - explaining why glutaminolysis was important after ovalbumin inhalations (used because emergence of anti-ova Ab / ASCs is suppressed by the NP hapten after NP-ova immunization) but not needed for the anti-NP response unless Slc2a1 or Mpc2 also was inactivated. Thank you prompting addition of this important caveat!

      Nevertheless, it seems fair to note that in Figures 1 and 2, the ASCs and Ab are being analyzed for NP and ova in the same mice, albeit with the NP-specific components not being driven by the inhalations of ovalbumin. With that in mind, when one compares the IgG1 anti-NP ASC and Ab to those for IgG1 anti-ovalbumin (ASC in bone marrow; Ab), the ovalbumin-specific response was reduced whereas the anti-NP response was not. [lines ~171-172]

      (5) NP immunization is known to be an inducer of an IgG1-dominant Th2-type immune response in mice. IgG2c is not a major player unless a nanoparticle delivery system is used. However, the authors arbitrarily included IgG2c in their assays in Figures 2 and 3. This may be confusing for the readers. The authors should either justify the IgG2c-mediated analyses or remove them from the main figures. (It can be added as supplemental information with proper justification).

      We rearranged the Figure panels to move IgM and IgG2c data to Supplemental Figures (Figure 3 - supplements 1, 2, 4, 5 in the eLife system).

      For purposes of public discourse, we note first that in contrast to the premise about weak IgG2c responses, the data [previously, Figure 3(c, g); now in the supplements] show substantial levels of NP-specific IgG2c. The referee is quite right that the class switching and in vitro ASC generation were done with IL-4 / IgG1-promoting conditions.

      To assist readers, the revised manuscript takes note of the important role of IgG2c (mouse - IgG1 in humans) in controlling or clearing various pathogens as well as in autoimmunity [lines ~182-5]. Moreover, we continue to think that these measurements add substantial value both from the standpoint of providing a better sense of generality to the loss-of-function effects, and in considering potential ways of translating the findings to B cell-dependent autoimmune conditions such as systemic lupus erythematosus.

      [As a scientific aside, we speculate that a greater or lesser IgG2c anti-NP response may arise due to different preparations of NP-carrier obtained from the vendor (Biosearch) having different amounts of TLR (e.g., TLR4) ligand. In any case, the points of presenting the IgG2c (and IgM) data were to push against the limiting boundaries of convention (which risks perpetuating a narrow view of potential outcomes) and make the breadth of results more apparent to readers.

      (6) Similarly, in affinity maturation analyses, including IgM is somewhat uncommon. I do not see any point in showing high affinity (NP2/NP20) IgMs (Figure 3d), since that data probably does not mean much.

      As noted in the reply immediately preceding this one, we appreciate this suggestion from the reviewer and moved the IgM and IgG2c to supplemental status.

      Nonetheless, in collegial discourse we disagree a bit with the referee in light of our data as well as of work that (to our minds) leads one to question why inclusion of affinity maturation of IgM is so uncommon - as the referee accurately notes. Of course a defect in the capacity to class-switch is highly deleterious in patients but that is not the same as concluding that recall IgM or its affinity is of little consequence.

      In some of the pioneering work back in the 1980's, Bothwell showed that NP- carrier immunization generated hybridomas producing IgM Ab with extensive SHM (~11% of the 18 lineages; ~ 1/3 of the IgM hybridomas) [PMID: 8487778], IgM B cells appear to move into GC, and there is at least a reasonable published basis for the view that there are GC-derived IgM (unswitched) memory B cells (MBC) that would be more likely, upon recall activation, to differentiate into ASCs. [As an example, albeit with the Jenkins lab anti-rPE response, Taylor, Pape, and Jenkins generated quantitative estimates of the numbers of Ag-specific IgM<sup>+</sup> vs switched MBC that were GC-derived (or not). [PMID: 22370719]. While they emphasized that ~90% of IgM<sup>+</sup> MBC appeared to be GC-independent, their data also indicated that ~1/2 of all GC-derived MBC were IgM<sup>+</sup> rather than switched (their Fig. 8, B vs C; also 8E, which includes alum-PE). And while we immensely respect the referee, we are perhaps less confident that IgM or high-affinity Ag-specific IgM doesn't mean that much, if only because of evidence that localized Ab compete for Ag and may thus influence selective processes [PMCID: PMC2747358; PMID: 15953185; PMID: 23420879; PMID: 27270306].

      (7) Following on my comment for the PC generation in Figure 1 (see above), in Figure 4, a strategy that relies solely on CD40L stimulation is performed. This is highly artificial for the PC generation and needs to be justified, or more physiologically relevant PC generation strategies involving anti-BCR, CD40L, and various cytokines should be shown.

      In line with our response to point (1), we tested BCR-stimulated B cells (anti-CD40 plus anti-IgM with BAFF, IL-4, and IL-5, parallel to the analyses with anti-CD40 but no BCR engagement). These results align with and reinforce the utility of the data with anti-CD40 as the sole mitogen.

      (8) The effects of CB839 and UK5099 on cell viability are not shown. Including viability data under these treatment conditions would be a valuable addition to the supplementary materials, as it would help readers more accurately interpret the functional outcomes observed in the study.

      We added presentation of data that provide cues as to relative viability / cxmsurvival under the experimental conditions used.

      [FSC X SSC as well as 7AAD or Ghost dye panels; we also generated new data that in[ further experiments scoring annexin V staining (see Fig 4 - supplement 1d, e, and Fig 5 - supplement 1e, f)].

      (9) It is not clear how the RNA seq analysis in Figure 4h was generated. The experimental strategy and the setup need to be better explained.

      Including text added at lines ~291-293 and ~582-585, the revised manuscript provides more information in the Results, Methods and Legend for Fig 4j-l. We agree entirely with the concern and apologize that in this and a few other instances we inadvertently sacrificed sufficiency of detail on the altar of attempting brevity.

      [As a synopsis: In three temporally and biologically independent experiments, cultures were harvested 3.5 days after splenic B cells were purified and cultured as in the experiments of Fig. 4a-e. Total cellular RNA was prepared from the twelve samples (three replicates for each of four conditions - DMSO vehicle control, CB839, UK5099, and CB839 <sup>+</sup> UK5099), then analyzed by RNA-seq. RNA-seq data were initially processed using the pipeline described in the Methods. For panels g & h of Fig 4, DESeq2 was used to quantify and compare read counts in the three CB839 <sup>+</sup> UK5099 samples relative to the three independent vehicle controls and identify all genes for which variances yielded P<0.05. In Fig 4g, all such genes for which the difference was 'statistically significant' (i.e., P<0.05) were entered into the indicated Immgen tool and thereby mapped to the B lineage subsets shown in the figure panels (i.e., g, h). In (g), these are displayed using one format, whereas (h) uses the 'heatmap' tool in MyGeneSet.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigate the functional requirements for glutamine and glutaminolysis in antibody responses. The authors first demonstrate that the concentrations of glutamine in lymph nodes are substantially lower than in plasma, and that at these levels, glutamine is limiting for plasma cell differentiation in vitro. The authors go on to use genetic mouse models in which B cells are deficient in glutaminase 1 (Gls), the glucose transporter Slc2a1, and/or mitochondrial pyruvate carrier 2 (Mpc2) to test the importance of these pathways in vivo.

      Interestingly, deficiency of Gls alone showed clear antibody defects when ovalbumin was used as the immunogen, but not the hapten NP. For the latter response, defects in antibody titers and affinity were observed only when both Gls and either Mpc2 or Slc2a1 were deleted. These latter findings form the basis of the synthetic auxotrophy conclusion. The authors go on to test these conclusions further using in vitro differentiations, Seahorse assays, pharmacological inhibitors, and targeted quantification of specific metabolites and amino acids. Finally, the authors document reduced STAT3 and STAT1 phosphorylation in response to IL-21 and interferon (both type 1 and 2), respectively, when both glutaminolysis and mitochondrial pyruvate metabolism are prevented.

      Strengths:

      (1) The main strength of the manuscript is the overall breadth of experiments performed. Orthogonal experiments are performed using genetic models, pharmacological inhibitors, in vitro assays, and in vivo experiments to support the claims. Multiple antigens are used as test immunogens--this is particularly important given the differing results.

      (2) B cell metabolism is an area of interest but understudied relative to other cell types in the immune system.

      (3) The importance of metabolic flexibility and caution when interpreting negative results is made clear from this study.

      Weaknesses:

      (1) All of the in vivo studies were done in the context of boosters at 3 weeks and recall responses 1 week later. This makes specific results difficult to interpret. Primary responses, including germinal centers, are still ongoing at 3 weeks after the initial immunization. Thus, untangling what proportion of the defects are due to problems in the primary vs. memory response is difficult.

      We performed new experiments and added the data on differences prior to a boost [see below; new Fig 3d, e; etc].

      (2) Along these lines, the defects shown in Figure 3h-i may not be due to the authors' interpretation that Gls and Mpc2 are required for efficient plasma cell differentiation from memory B cells. This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. The more likely interpretation is that ongoing primary germinal centers are negatively impacted by Gls and Mpc2 deficiency, and this, in turn, leads to reduced affinities of serum antibodies.

      We have edited the wording of the conclusion to add a possibility we consider unlikely and downplay a conclusion that MBCs bearing switched BCRs are affected once reactivated. [see lines ~221-230] We also have added citations pertaining to the topic, including work from the Victora lab which seems to put the point succinctly: "Recall GCs in mice consist almost entirely of naïve B cells, whereas recall antibodies derive overwhelmingly from memory B cells." [emphasis added] [PMID: 38838672; new ref #83]. While unclear as to the reasoning - as one looks at the data - and skeptical as to the accuracy of the referee's point (2), it suggests that the matter is open to reasonable doubt. In line with the point and the edits, we also have added citation of a bioRxiv preprint from the Victora lab, which touches on the concept of what one could call boost-induced reinvigoration of a pre-existing GC [new ref #82].

      Beyond the textual changes, we performed a new series of experiments to investigate partially, and present the results in Fig 3d, e as well as Fig 3 - supplement 1d, e. Unfortunately, time before lab closure was an enemy both for the period between primary and recall immunizations in performance and multiple replication of work to extend that presented in Figure 3, panels g & h, and the related Supplemental Data (Fig 3 - supplements 4d, 5a-g). Unfortunately, it was not possible to do a longer-term memory experiment with recall immunization out at 8 weeks.

      The intriguing concerns and questions of points 1 & 2 provide a springboard for consideration of generalizations and simplifications. Germinal center durability is not at all monolithic, and instead is quite variable**. It is true that in the literature (especially with the substantially different approach of transferring BCR-transgenic / knock-in versions of an NP-biased BCR) there may be meaningful pools of IgG1 and IgG2c GC B cells. The premise (cognitive bias, perhaps?) in our interpretation is that in our previous work we measured few if any GC B cells - NP-APC-binding or otherwise - above the background (non-immunized controls) three weeks after immunization with NP-ovalbumin in alum. While recognizing that the immunogen can matter, we note for the readers and referee that Fig. 1 of the Taylor, Pape, & Jenkins paper considered above [PMID: 22370719] reported 10-fold more Ag-specific MBCs than GC B cells at day 29 post-immunization (the point at which the boost/recall challenge was performed in our Figure 3g, h. [That work did not use NP-carrier in alum to immunize, or measure the anti-NP response.]

      Viewing Fig. 3i from that perspective, the surmise of the comment is that a major contribution to the differences in both all-affinity and high-affinity anti-NP IgG1 (whose production requires differentiation into plasma cells) derived from the immunization at 4 wk stimulating persistent GC B cells as opposed to memory B cells.

      The issue and question also relate to rates of output of plasma cells or rises in the serum concentrations of class-switched Ab. To this point, our prior experiences agree with the long-published data of the Kurosaki lab in Figure 3c of the Aiba et al paper noted above (Immunity, 2006) (and other such time courses). Readers can note that the IgG1 anti-NP response (alum adjuvant, as in our work) hits its plateau at 2 wk, and did not increase further from 2 to 3 wk. The most likely interpretation is that GC are on the decline and Ab production has reached its plateau by the time of the 2nd immunization in Fig. 3h.

      Assuming we understand the comment and line of reasoning correctly, we also lean towards disagreeing with the statement " This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. Our evidence shows that both low-affinity as well as high-affinity anti-NP Ab (IgG1) were reduced due to combined gene-inactivation after the peak primary response (Fig. 3h; also, see the new data in Fig 3 and Fig 3 - supplement 1). Recent papers show that affinity maturation is attributable to greater proliferation of plasmablasts with high-affinity BCR. Accordingly, the findings with loss of GLS and MPC function are quite consistent with the interpretation that much of the response after the second immunization draws on MBC differentiation into plasmablasts and then plasma cells, where the proliferative advantage of high-affinity cells is blunted by the impaired metabolism. Notwithstanding these issues, the revised manuscript includes the alternative, if less likely, interpretation proposed by the review [lines ~221-230].

      **In some contexts, of course, especially certain viral infections or vaccination with lipid nanoparticles carrying modified mRNA, germinal centres are far more persistent; also, in humans even the seasonal flu vaccine

      (3) The gating strategies for germinal centers and memory B cells in Supplemental Figure 2 are problematic, especially given that these data are used to claim only modest and/or statistically insignificant differences in these populations when Gls and Mpc2 are ablated. Neither strategy shows distinct flow cytometric populations, and it does not seem that the quantification focuses on antigen-specific cells.

      The revised manuscript improves these aspects of the presentation, using old and new data. See Fig 3 - supplement 3a, c; Fig 3 - supplement 4a. We note for readers that many other papers in the best journals show plots in which the separation of, say, GC-Tfh from overall Tfh is based on cut-off within what essentially is a continuous spectrum of emission as adjusted or compensated by the cytometer (spectral or conventional).

      The revised manuscript presents results from new experiments that deal with the subset of GC B cells whose BCRs bind NP-APC with enough affinity to retain a positive signal after washing. These new data are presented in Fig 3 - supplement 3c & 3e. In practice, the new findings suggest that the metabolic requirement applied more to the NP-binding B cells than the overall GC B cell population.

      (4) Along these lines, the conclusions in Figure 6a-d may need to be tempered if the analysis was done on polyclonal, rather than antigen-specific cells. Alum induces a heavily type 2-biased response and is not known to induce much of an interferon signature. The authors' observations might be explained by the inclusion of other ongoing GCs unrelated to the immunization.

      We apologize for ambiguity or insufficient clarity and, as noted above, have edited the text to be more clear that the in vitro experiments do not represent GC B cells and that the RNA-seq data were from experiments that did not involve alum and were not an Ag (SRBC)-specific subset.

      New text in the Results, an expanded Legend, and tweaking the Methods make it more readily clear that the RNA-seq data (and hence the GSEA) involved immunizations with SRBC (not the alum / NP system. That said, we note that the hapten-carrier experiments in which the immunogen was adjuvantized with alum actually generated a robust IgG2c (type 1-driven) response along with the type 2-enhanced IgG1 response, in line with what has been reported by others with alum-adjuvanted vaccination.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors investigate how glutaminolysis (GLS) and mitochondrial pyruvate import (MPC2) jointly shape B cell fate and the humoral immune response. Using inducible knockout systems and metabolic inhibitors, they uncover a "synthetic auxotrophy": When GLS activity/glutaminolysis is lost together with either GLUT1-mediated glucose uptake or MPC2, B cells fail to upregulate mitochondrial respiration, IL 21/STAT3 and IFN/STAT1 signaling is impaired, and the plasma cell output and antigen-specific antibody titers drop significantly. This work thus demonstrates the promotion of plasma cell differentiation and cytokine signaling through parallel activation of two metabolic pathways. The dataset is technically comprehensive and conceptually novel, but some aspects leave the in vivo and translational significance uncertain.

      Strengths:

      (1) Conceptual novelty: the study goes beyond single-enzyme deletions to reveal conditional metabolic vulnerabilities and fate-deciding mechanisms in B cells.

      (2) Mechanistic depth: the study uncovers a novel "metabolic bottleneck" that impairs mitochondrial respiration and elevates ROS, and directly ties these changes to cytokinereceptor signaling. This is both mechanistically compelling and potentially clinically relevant.

      (3) Breadth of models and methods: inducible genetics, pharmacology, metabolomics, seahorse assay, ELISpot/ELISA, RNA-seq, two immunization models.

      (4) Potential clinical angle: the synergy of CB839 with UK5099 and/or hydroxychloroquine hints at a druggable pathway targeting autoantibody-driven diseases.

      We agree and thank the referee for the positive comments and this succinct summary of what we view as contributions of the paper.

      Weaknesses:

      (1) Physiological relevance of "synthetic auxotrophy"

      The manuscript demonstrates that GLS loss is only crippling when glucose influx or mitochondrial pyruvate import is concurrently reduced, which the authors name "synthetic auxotrophy". I think it would help readers to clarify the terminology more and add a concise definition of "synthetic auxotrophy" versus "synthetic lethality" early in the manuscript and justify its relevance for B cells.

      We edited the Abstract, Introduction, and Discussion to try to do better on this score. Conscious of how expansive the prose and data are even in the original submission, we appear to have taken some shortcuts that we will try to rectify or at least mitigate. Thank you for highlighting this need to improve on key concepts !!

      Specifically, the revised text expands a bit on the notion that synthetic auxotrophy represents effects on differentiation that go beyond additional mechanisms of reducing division efficiency and a modest impact on selective death. [see the 10th - 11th lines in Abstract and lines ~84-85, Introduction] Even though decreased population expansion is observed and new evidence supports a model in which the altered metabolism contributes to enhanced death in vivo, at equal division numbers the frequency of CD138<sup>+</sup> progeny is lower once glutaminolysis and mitochondrial pyruvate are reduced by either genetic or pharmacological means.

      This comment of the review raises interesting semantic questions about what represents "physiological relevance". The fundamental point is to explore a basic science question - what, if any, are limits to metabolic flexibility? In principle, shouldn't B cells be able to use fatty acid metabolism to generate enough ATP and provide the backbones for biosynthesis during growth? Put a different way, the point is that a basic curiosity to understand why decreasing glucose influx did not have an even more profound effect than what was observed, combined with curiosity as to why glutaminolysis was dispensable in relatively standard vaccine-like models of immunize/boost, provided a springboard to identification of new vulnerabilities. The manuscript shows one physiological limitation (and hence vulnerability). Be that as it may, the revised text of the Discussion section more clearly addresses this issue (lines ~531-549 at the end of the Discussion).

      While the overall findings, especially the subset specificity and the clinical implications, are generally interesting, the "synthetic auxotrophy" condition feels a little engineered.

      CAR-T cells are 'a little engineered' (or more than a little) and yet they do seem to have had an impact on understanding the centrality of B cells in various autoimmune conditions as well as in the direction of cancer therapy research. So it is a matter of balancing this perspective of the referee against the strengths they highlight in points 1, 2, and 4. In editing the revision, we try to expand and be more explicit about this in the Discussion of the revised manuscript.

      In brief, even were the money not all gone, we would not believe that expanding the heft of this already rather large manuscript and set of data would be appropriate. As matters stand, a basic new insight about metabolic flexibility and its limits leads to evidence of a way to reduce generation of Ab and a novel impairment of STAT transcription factor induction by several cytokine receptors. The vulnerability that could be tested in later work on B cell-dependent autoimmunity includes the capacity to test a compound that already has been to or through FDA phase II in patients together with an FDA-approved standard-of-care agent.

      Therefore, the findings strongly raise the question of the likelihood of such a "double hit" in vivo and whether there are conditions, disease states, or drug regimens that would realistically generate such a "bottleneck".

      Hence, the authors should document or at least discuss whether GC or inflamed niches naturally show simultaneous downregulation/lack of glutamine and/or pyruvate. The authors should also aim to provide evidence that infections (e.g., influenza), hypoxia, treatments (e.g., rapamycin), or inflammatory diseases like lupus co-limit these pathways.

      Again, we appreciate some 'licensing' to be more expansive and explicit, and will try to balance editing in such points against undue tedium or tendentiously speculative length in the Discussion. In particular, we will note that a clear, simple implication of the work is to highlight an imperative to test CB839 in lupus patients already on hydroxychloroquine as standard-of-care, and to suggest development of UK5099 (already tested many times in mouse models of cancer) to complement glutaminase inhibition.

      As backdrop, we note that the failure to advance imaging mass spectrometry to the capacity to quantify relative or absolute (via nano-DESI) concentrations of nutrients in localized interstitia is a critical gap in the entire field. Techniques that sample the interstitial fluid of tumour masses or in our case LN as a work-around have yielded evidence that there can be meaningful limitations of glucose and glutamine, but it needs to be acknowledged that such findings may be very model-specific and, as can be the case with cutting-edge science, are not without controversy. That said, yes, we had found that hypoxia reduced glutamine uptake but given the norms of focused, tidy packages only reported on leucine in an earlier paper [PMID27501247; PMCID5161594].

      Beyond all that, another impetus to and inspiration for these experiments stems from quite data that we generated in a model of short-term protein-restricted diet (loosely akin to kwashiorkor in humans), based on an excellent publication showing that such a regimen quickly led to lower circulating glutamine and mTORC1 activity (**). In brief, we found that a low-protein diet did, in our experiments, preferentially lower glutamine but - importantly - led to reduced Ab responses (which would match what we have modeled here). The findings were not a well-enough connected evidentiary component to include in the "story" but I'll append slides with the relevant data to this Response to Reviews for the referee's perusal (and anyone else who reads this online discourse).

      It would hence also be beneficial to test the CB839 + UK5099/HCQ combinations in a short, proof-of-concept treatment in vivo, e.g., shortly before and after the booster immunization or in an autoimmune model. Likewise, it may also be insightful to discuss potential effects of existing treatments (especially CB839, HCQ) on human memory B cell or PC pools.

      We certainly agree that the suggestions offered in this comment are important next steps and the right approach to test if the findings reported here translate toward the treatment of autoimmune diseases that involve B cells, interferons, and pathophysiology mediated by auto-Ab. As practical points, performance and replication of such studies would take more time than the year allotted for return of a revised manuscript to eLife and in any case neither funds nor a lab remain to do these important studies.

      Concrete evidence for our concurrence was embodied in a grant application to NIH that was essential for keeping a lab and doing any such studies. [We note, as a suggestion to others, that an essential component of such studies would be to test the effects of these compounds on B cells from patients and mice with autoimmunity]. Perhaps unfortunately for SLE patients, the review panelists did not agree about the importance of such studies. However, it can be hoped that the patent-holder of CB839 (and perhaps other companies developing glutaminase inhibitors) will see this peer-reviewed preprint and the public dialogue, and recognize how positive results might open a valuable contribution to mitigation of diseases such as SLE.

      (2) Cell survival versus differentiation phenotype

      Claims that the phenotypes (e.g., reduced PC numbers) are "independent of death" and are not merely the result of artificial cell stress would benefit from Annexin-V/active-caspase 3 analyses of GC B cells and plasmablasts. Please also show viability curves for inhibitor-treated cells.

      This comment leads us to see that the wording on this point may have been overly terse in the interests of brevity, and thereby open to some odd misunderstanding. The CD138<sup>+</sup> events are scored among VIABLE CELLS, so a decrease in the %CD138<sup>+</sup> at similar division number represents an effect independent from (or beyond) survival and division-counting. Accordingly, we expanded the text of the Abstract and elsewhere in the manuscript, to be more clear. In addition, we added data from new experiments addressing death in vitro and among GC-phenotype B cells in vivo. To clarify in this public context, it is not that an increase in death (along with the reported decrease in cell cycling) can be or is excluded. The point is that beyond any such increase, and taking into account division number (since there is evidence that PC differentiation and output numbers involve a 'division-counting' mechanism), the frequencies of CD138<sup>+</sup> cells and of ASCs among the viable cells are lower, as is the level of Prdm1-encoded mRNA even before the big increase in CD138<sup>+</sup> cells in the population.

      (3) Subset specificity of the metabolic phenotype

      Could the metabolic differences, mitochondrial ROS, and membrane-potential changes shown for activated pan-B cells (Figure 5) also be demonstrated ex vivo for KO mouse-derived GC B cells and plasma cells? This would also be insightful to investigate following NP-immunization (e.g., NP+ GC B cells 10 days after NP-OVA immunization).

      We performed a series of new experiments to have enough biologically independent replications for meaningful and statistical analyses. The new results, added in as Fig 5 - supplement 1, showed that the combined pathway interruption by loss-of-function increased ROS, mtROS, and death (annexin V / 7AAD) upon analyzing GCphenotype B cells immediately upon harvest. The findings align well with the data in Fig 5 (cultured B cells).

      (4) Memory B cell gating strategy

      I am not fully convinced that the memory-B-cell gate in Supplementary Figure 2d is appropriate. The legend implies the population is defined simply as CD19+GL7-CD38+ (or CD19+CD38++?), with no further restriction to NP-binding cells. Such a gate could also capture naïve or recently activated B cells. From the descriptions in the figure and the figure legend, it is hard to verify that the events plotted truly represent memory B cells. Please clarify the full gating hierarchy and, ideally, restrict the MBC gate to NP+CD19+GL7-CD38+ B cells (or add additional markers such as CD80 and CD273). Generally, the manuscript would benefit from a more transparent presentation of gating strategies.

      In considering the referee's viewpoint, we further expanded the supplemental data displays to include more of the gating and analytic schemes, which we believe should mitigate one concern noted here. In addition, we now include flow data from the non-immunized control mice that had been analyzed concurrently in the experiments.

      Third and finally, we performed new experiments and analyses in which the focus was the frequencies of memory-phenotype (IgD<sup>neg</sup> GL7<sup>neg</sup> CD38<sup>+</sup> / CD38<sup>hi</sup> aka CD38<sup>+</sup><sup>+</sup>) NPbinding B cells after immunization. While this time, as opposed to previously, the NP-APC staining met our standard for interpretability, the gist of the findings was that the two independent repeat experiments yielded a split decision and a degree of variability. With time being up due to the funds running out, we have elected to delete the issue and the data panel in question.

      That said, it bears noting that in the previous figure panel, the labeling indicated that the gating included the important criterion that cells be IgD<sup>neg</sup>, which excludes the vast majority of naive B cells but measures memory-phenotype B cells independent from consideration of whether or not they were NP-binding.

      [In principle marginal zone (MZ) B cells might fall within this gate. However, the MZ B population is unlikely to explain the differences shown.

      (5) Deletion efficiency - [The] mRNA data show residual GLS/MPC2 transcripts (Supplementary Figure 8). Please quantify deletion efficiency in GC B cells and plasmablasts.

      Even were there resources to do this, the degree of reduction in target mRNA (Gls; Mpc2) renders this question superfluous. To the best of our understanding, the proteins (for which there might be some phenotypic lag) are translated from RNA. Might there be a small subpopulation of B cells (or their PC progeny) with only one, or even neither, allele converted from fl to D? Yes, but they would be a minor subset in light of the magnitude of mRNA reduction, in contrast to our published observations with Slc2a1. As to plasmablasts and plasma cells, the pre-existing populations make such an analysis misleading, while the scarcity of such cells recoverable with antigen capture techniques is so low as to make both RNA and genomic DNA analyses questionable. We also refer readers to the supplemental figure that presents the results of experiments testing the issue one might infer from the question about extents of deletion in PC (i.e., how much counter-selection might have occurred by the PC stage).

    1. eLife Assessment

      This interesting study adapts machine learning tools to analyze movements of a chromatin locus in living cells in response to serum starvation. The machine learning approach developed is useful, the experiments are well controlled, and the data are solid. The study will benefit from future work testing predictions by perturbation experiments. This work will be of interest to those studying chromosome biology and gene expression patterns.

    2. Reviewer #1 (Public review):

      Summary:

      Redchuk et al. explore the dynamic properties of chromatin upon serum starvation using machine learning approaches. They use CRISPR-tagging to visualize a region on chromosome 1 in human cells and show that in their system, chromosome 1, but not the previously reported chromosomes 10, 13, and X, undergo a change in radial position upon serum starvation. Live cell imaging showed a position change towards the periphery after serum starvation. They then apply a machine learning algorithm for the analysis of the imaging data, which reveals changes in nuclear area during serum starvation and longer displacements of the chromosome 1 locus near the nuclear periphery. Differential behavior of homologues is also reported.

      Strengths:

      (1) The study of chromatin dynamics is an interesting and important area of research.

      (2) The use of machine learning approaches to analyze live cell imaging data is timely.

      (3) With serum starvation, the authors use a simple, well-controllable model system.

      Weaknesses:

      (1) This study provides limited new insight into chromatin dynamics.

      (2) It was not immediately evident what the use of machine learning approaches added to this study. It appears that the main conclusions could have been reached by conventional analysis.

      Comments on revised version:

      The authors have added some technical information, but have not made any major efforts to clarify some of the major points or to strengthen the paper. The degree of advance remains limited and several conclusions are not convincingly supported by the presented data.

    3. Reviewer #2 (Public review):

      Summary:

      The study demonstrates that CRISPR-Sirius provides a powerful approach to investigating chromosome dynamics in living cells during environmental stress. By focusing on serum starvation, the authors show that this process induces global nuclear changes, including a reduction in nuclear area and increased morphological dynamism, while at the same time driving specific reorganization of chromosome 1. Chromosome 1 relocates toward the nuclear periphery and displays distinctive patterns of motion, maintaining overall motility but punctuated by occasional long-distance displacements, particularly near the nuclear envelope. Importantly, the analysis reveals that homologous copies of chromosome 1 do not behave uniformly: peripheral loci become more mobile and responsive to starvation, whereas central homologs remain comparatively stable, often associated with nucleolar subcompartments. By integrating live imaging with machine learning and explainable AI analysis, the study highlights the complexity of nuclear organization and provides valuable insights into how chromosome-specific and locus-specific responses to stress are orchestrated within the three-dimensional nuclear landscape.

      Strengths:

      The study uses live-cell imaging to investigate the dynamics of loci during starvation. Live-cell tracking and data interpretation are carried out using machine learning and AI models, which is a major strength.

      Weaknesses:

      The manuscript is at times difficult to follow, partly because the methodological descriptions are highly specialized, especially for non-expert biologists. In addition, the observations are not tested for a mechanistic basis. Experiments that could provide deeper insights are missing, for example, why chromosome 1 moves, why the peripheral homologue dislocates, or why a "long jump" is observed at the periphery even though the speed of the loci does not change. It is also unclear whether a displacement of 0.5 μm is functionally meaningful.

      Comments on revised version:

      The authors have added some technical information and provided a better discussion of the data, but beyond that, they have not strengthened the conclusions. The observations are not supported by any perturbation assays.

    4. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This interesting study adapts machine learning tools to analyze movements of a chromatin locus in living cells in response to serum starvation. The machine learning approach developed is useful, the experiments are well controlled, and the data are solid. The study would be greatly strengthened by testing key predictions made using perturbation experiments. This work will be of interest to those studying chromosome biology and gene expression patterns.

      We thank eLife for this nice assessment. We indeed believe that the presented machine learning approach will be useful for many types of research questions, and this was the main aim of this manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Redchuk et al. explore the dynamic properties of chromatin upon serum starvation using machine learning approaches. They use CRISPR-tagging to visualize a region on chromosome 1 in human cells and show that in their system, chromosome 1, but not the previously reported chromosomes 10, 13, and X, undergo a change in radial position upon serum starvation. Live cell imaging showed a position change towards the periphery after serum starvation. They then apply a machine learning algorithm for the analysis of the imaging data, which reveals changes in nuclear area during serum starvation and longer displacements of the chromosome 1 locus near the nuclear periphery. Differential behavior of homologues is also reported.

      Strengths:

      (1) The study of chromatin dynamics is an interesting and important area of research.

      (2) The use of machine learning approaches to analyze live cell imaging data is timely.

      (3) With serum starvation, the authors use a simple, well-controllable model system.

      Weaknesses:

      (1) This study only provides limited new insight into chromatin dynamics.

      We respectfully disagree with this conclusion. To the best of our knowledge, our study is the first to provide any insights into chromatin dynamics upon serum starvation. Previous studies are solely based on studies in fixed cells, and the dynamics have remained unexplored. Moreover, for example the notion that homologous chromosomes show differential dynamic behavior is novel and will likely have implications and relevance to many chromatin-based processes beyond the example studied here.

      (2) It was not immediately evident what the use of machine learning approaches added to this study. It appears that the main conclusions could have been reached by conventional analysis.

      First, we would like to point out that the other reviewer found our machine learning analysis pipeline a major strength of our manuscript. Indeed, analyzing single features and assessing their impact on the studied phenomenon could have been achieved relatively easily by conventional analysis. However, this analysis would have ignored the interactions (some of which were not intuitively obvious) between different features and thereby limited the knowledge gain from the experiment.

      Unbiased analysis of the interactions between the different features would have been already very difficult and time-consuming with conventional approaches. We believe that our analysis pipeline, especially with the Shapley values, addresses the key issue of combinatorial explosion prominent to multiparametric data, such as imaging data, and helps the researcher to navigate complex datasets.

      (3) There are several specific technical points:

      (a) It was not clear what the CRISRP-Sirius probes actually labelled. The chromosome 1 sgRNA sequence is provided, but I could not find information as to which region(s) of the chromosome are actually labelled (size, location, etc.).

      We have added a schematic as Supplementary Figure 1A to show the region of the chromosome that is labelled. In addition, the target sequence, together with the relevant references can be found in the Materials and methods (page 16). Please see also below Reviewer #1 (Recommendations for the authors) point 4a.

      (b) The authors visualize a relatively small region of chromosome 1 but make conclusions regarding the entire chromosome. Additional probes on the same chromosome should be used.

      Related to this point, the discussion of why the authors are unable to reproduce the prior findings of relocation of chromosomes 10, 13, and X is not satisfying. It would be worth comparing the FISH-based painting of entire chromosomes, which generated the results suggesting relocation of these chromosomes, with the point-labelling method used here.

      We agree that our approach to labeling chromosome 1 is very different than the FISH-based probes utilized before. However, we also feel that we discuss this aspect, and the difference between our and previous results, which may also stem from the used cell model, in quite a detail in the first paragraph of the results (page 4). Also, we are very careful throughout the manuscript to indicate that here we study the dynamics of a specific chromosome loci, not the entire chromosome, and have further amended the text to emphasize this. In the future, it would be very interesting to study the dynamics of also other loci of chromosome 1. As indicated also below in response to reviewer 2, we have failed to identify further gRNAs that would reliably and reproducibly label further chromosome 1 loci, suggesting that we would need to change the labeling system entirely. Unfortunately, this is not in the scope of this manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 1.

      (c) The study lacks controls. Since in their hands chromosomes 10, 13, and X do not change position, they should be used as a negative control in all experiments demonstrating a shift in the location of chromosome 1.

      We disagree that our study lacks controls, since we use telomeres as controls throughout the manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 2,3.

      (d) I did not find information about the spatial or temporal resolution of the imaging modality. This is important to assess whether the observed changes in position, relative to time, are meaningful.

      To estimate the spatial resolution, we have added new data using fixed cells (Supplementary figure 1E; corresponding text in results on page 5); temporal resolution is indicated in Materials and methods (page 17). Please see also below Reviewer #1 (Recommendations for the authors) point 4d.

      (e) The authors analyze surprisingly early timepoints (up to 40 minutes) of serum starvation. Would these results look different if longer serum starvation timepoints of several hours were analyzed?

      We chose to analyze early time points of serum starvation based on the previous literature reporting the chromosome relocation within the first 15 minutes of starvation. Indeed, the results might look very different later during serum starvation, since we already observe differences between 0-20 min vs 20-40 min into starvation (see for example Figure 5A-D). Analyzing further time points is not in the scope of this manuscript.

      (f) The authors can do a better job of explaining what the biological meaning of the various parameters (DistR, TDist, etc.) they measure is.

      We have amended Table 1 to describe the measured features more clearly. Please see also below Reviewer #1 (Recommendations for the authors) point 4e.

      (g) I did not understand the reasoning for the authors' conclusion of differential behavior of homologues. Please explain this better, or idealy use more direct labeling methods that identify the individual homologues.

      The differential behavior of homologues is best demonstrated in Figure 6H, which shows that in serum-containing media, the peripheral homolog has equal probability of being faster or slower compared to its homolog. However, the distribution changes upon starvation, with the peripheral loci being more frequently the faster homolog. We completely agree that further studies are needed to understand this phenomenon better, but changing the labeling method is not in the scope of this manuscript.

      (h) In many figures, statistical analysis of the data is missing, including, but not limited to, Figures 1B, C, G, Figures 4, 5, 6.

      We have added a Supplementary table to include inferential statistics. See also below Reviewer #1 (Recommendations for the authors) point 4b.

      (i) No information is provided throughout the manuscript as to how many cells were analyzed in each experiment. This should be indicated in every figure legend.

      The number of analyzed loci or nucleus is indicated in every figure. See also below Reviewer #1 (Recommendations for the authors) point 4c.

      Reviewer #2 (Public review):

      Summary:

      The study demonstrates that CRISPR-Sirius provides a powerful approach to investigating chromosome dynamics in living cells during environmental stress. By focusing on serum starvation, the authors show that this process induces global nuclear changes, including a reduction in nuclear area and increased morphological dynamism, while at the same time driving specific reorganization of chromosome 1. Chromosome 1 relocates toward the nuclear periphery and displays distinctive patterns of motion, maintaining overall motility but punctuated by occasional long-distance displacements, particularly near the nuclear envelope. Importantly, the analysis reveals that homologous copies of chromosome 1 do not behave uniformly: peripheral loci become more mobile and responsive to starvation, whereas central homologs remain comparatively stable, often associated with nucleolar subcompartments. By integrating live imaging with machine learning and explainable AI analysis, the study highlights the complexity of nuclear organization and provides valuable insights into how chromosome-specific and locus-specific responses to stress are orchestrated within the three-dimensional nuclear landscape.

      Strengths:

      The study uses live-cell imaging to investigate the dynamics of loci during starvation. Livecell tracking and data interpretation are carried out using machine learning and AI models, which is a major strength.

      Weaknesses:

      The manuscript is at times difficult to follow, partly because the methodological descriptions are highly specialized, especially for non-expert biologists. In addition, the observations are not tested for a mechanistic basis. Experiments that could provide deeper insights are missing, for example, why chromosome 1 moves, why the peripheral homologue dislocates, or why a "long jump" is observed at the periphery even though the speed of the loci does not change. It is also unclear whether a displacement of 0.5 μm is functionally meaningful.

      We appreciate the comment about the readability of our manuscript, and have seriously evaluated this point. We also completely agree that it would be interesting and important to understand the mechanistic and functional basis of the observed changes in chromatin dynamics take place upon serum starvation. However, we feel that it is not in the scope of the present manuscript. See also below Reviewer #2 (Recommendations for the authors) points 3,7-11.

      Recommendations for the authors:

      Reviewing Editor Comments:

      I would like to first offer my congratulations on a very interesting study; second, I would like to encourage you to test a few key predictions using a perturbation experiment. Two reviewers with deep expertise in this area were supportive of the work, and both noted that such an addition would greatly increase the impact and visibility of this work in the field. I welcome a revision that addresses this seminal point. Thank you for sending your work to eLife!

      We thank eLife for the positive assessment. We have aimed to address all of the reviewers comments and suggestions. However, we feel that some of the suggestions are not in the scope of this particular manuscript, since they would require setting up a different chromatin labeling system.

      Reviewer #1 (Recommendations for the authors):

      The following experiments would strengthen the study:

      (1) Please label additional regions on chromosome 1 so as not to rely on a single point to represent the behavior of the entire chromosome.

      This is an excellent suggestion, but unfortunately, despite our extensive efforts, we have failed to identify further gRNAs that would reliably label chromosome loci with the CRISPR-Sirius system. Changing the labeling system is not in the scope of the presented manuscript.

      (2) Please use chromosomes 10, 13, or X as a negative control since these chromosomes do not change position in the authors' hands.

      (3) Please compare the behavior of the homologues to that of either random loci or control loci on 10, 13, or X to assess whether the differential behavior observed for chromosome 10 is a specific effect.

      Related to points 2 and 3, we opted to use telomeres as controls in this study. Throughout the manuscript, the behavior of chromosome 1 loci is compared to telomeres, demonstrating the specific effect of serum starvation on chr 1. For example, Figure 5A and 5B show that when analyzing mean locus displacement, chr1 and telomeres show the opposite behavior.

      (4) In addition:

      (a) Please provide detailed information on the sequence and location of the probes used.

      We have added a schematic showing the location of the probes as Supplementary Figure S1A. In addition, the sequences are indicated in Materials and methods (page 16).

      (b) Please provide a statistical analysis in all graphs.

      To make statistical analysis more comprehensive, we have added supplementary table 1, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. See also below Reviewer #2 (Recommendations for the authors) point 5.

      (c) Please provide throughout the manuscript in each figure legend information as to how many cells were analyzed in each experiment.

      The number of analyzed loci (or nucleus) is indicated in each graph.

      (d) Please provide information on the spatial and temporal resolution of the imaging modality.

      The imaging settings are indicated in Materials and methods, including the temporal resolution of 0.25 frames per second (page 17). To estimate spatial resolution, and especially its relationship with the observed repositioning of the chromosome loci, we performed experiments in fixed cells, using an optically identical set-up as utilized for live imaging. Unfortunately, the microscope utilized for live imaging was taken out of use by the core facility after submission of the original draft of this manuscript, but we used a microscope with essentially a similar set-up. The data from fixed cells is now presented as Supplementary figure 1E and discussed in results on page 5. This analysis indicates that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error.

      (e) Please better explain what the various measured parameters mean in biological terms.

      We have amended Table 1 to provide better explanation of the measured parameters.

      (f) Please add a scale bar to Figure 6I.’

      Scale bar has been added to figure 6I.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) SHAP analysis identified nuclear area (MA) and its change (sA) as the most predictive features of starvation state, while motility features (MD, MaxD, TD) showed strong interactions with nuclear morphology. Discrete features, such as displacement outliers and homolog subclassification by speed/proximity, influenced classification, particularly in MLP models. Could the authors clarify why morphological and motility features act in combinatorial and context-dependent ways? A biological interpretation of this interdependence would strengthen the study.

      Unfortunately, we do not have a good biological interpretation for this. The fact that some interactions are context-dependent indicates that there could be subpopulations of cells/analyzed loci. For example, we found that the predictive value of nuclear area was high in a subgroup of low motility loci (Figure 4F and Supplementary figure 4D-F). We do not believe that adding more speculation would strengthen the study.

      (2) The manuscript shows that chromosome 1 moves toward the periphery within the first 20-40 minutes of serum withdrawal. However, it remains unclear whether the locus eventually "touches" the periphery and whether it subsequently stabilizes or retracts. It would be valuable to compute the time point of minimal nuclear distance and examine whether this is transient or sustained.

      With the experimental set-up utilized here, we imaged the loci for only two minutes at random time point within the first 40 minutes of the starvation. Hence extracting the time point of minimal nuclear distance is not meaningful from this dataset. As we discuss in the manuscript, following the dynamics of the same locus for longer periods of this would be very interesting in the future. However, this is not in the scope of the present manuscript.

      (3) The manuscript is at times difficult to follow, partly because methodological descriptions are highly detailed in the main text. Consider moving more of the methodological content into Supplementary Methods and emphasizing the main results and interpretations in the main text for clarity.

      We have carefully evaluated this point. Most methodological descriptions in the manuscript relate to the machine learning models and their explanation with SHAP. As we feel that this combination is an essential part of the manuscript, and likely the aspect that can have widest impact beyond chromatin dynamics studies, we feel that the background and our reasoning related to the chosen methods are important.

      (4) The distinction between the first 20 minutes and the latter 40-minute window is intriguing. Could these different time scales be paralleled with early versus delayed gene expression responses to serum starvation? A discussion of this temporal connection would add biological depth.

      This is an intriguing idea, and we have added a short note on this in the discussion (page 13). However, as we do not know how the U2OS cells utilized here respond transcriptionally to serum starvation, we are hesitant to speculate too much.

      (5) If the observed interpretations are robust, could this be demonstrated more explicitly through statistical principles or reproducibility tests across independent datasets?

      To provide further evidence of the robustness of our findings, we have 1) added new data to estimate the spatial resolution (Supplementary figure 1E) and 2) expand the statistical analysis as supplementary table 1. Regarding the spatial resolution (see also the response to reviewer 1), our experiments on fixed cells demonstrate that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. To make statistical analysis more comprehensive, we added a supplementary table, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. MW test was used as a non-parametric statistic, with null hypothesis assuming the samples are coming from the same distribution. Null hypothesis was rejected at the p-value below 0.05. In most cases (bold font in table) MW test results were in accordance with the descriptive statistics confirming the conclusions in the study. In case of exceptions (MD, TD for telomeres and TDist), the results were reported, for example, as an “appearing trend” to reflect descriptive statistics while highlighting certainty levels.

      (6) Figure labeling is difficult to follow. Please include abbreviation explanations directly in the figure panels or legends for clarity.

      Abbreviations have been added to figure legends. Adding them to figures themselves would have made the figures too busy.

      (7) The manuscript reports higher displacement at the nuclear periphery. Can the authors explain why displacement amplitudes increase near the periphery and how this relates to nuclear architecture?

      We speculate in the manuscript (results, page 11; discussion, page 14) that actually the lower displacement observed with the central locus may, at least partially, result from anchoring this locus to the nucleolus (Fig 6I). Nevertheless, alternative explanations, such as differences in transcriptional and/or chromatin states may exist (see also the response to point 11), and this is now mentioned in the discussion (page 14).

      (8) How is the movement of chromosome 1 directed specifically toward the periphery, rather than being random fluctuations? This point requires clarification.

      This is an important question, but unfortunately our data does not provide an answer to this, and suggesting any mechanism would be pure speculation. Nevertheless, our results agree with previous studies utilizing fixed cells that also demonstrated movement of chromosome 1 towards nuclear periphery (Mehta et al., 2010), arguing against random fluctuation.

      (9) Only chromosome 1, and not the other tested chromosomes, undergoes this relocalization. Could the authors elaborate on why some chromosomes but not others display this behavior?

      Previous studies (Mehta et al., 2010) utilizing chromosome paints in fixed cells actually show the relocalization of several chromosomes upon serum starvation. The fact that we observed the relocalization of only chr1 loci is likely due to the labeling method and/or the cell model utilized in this study. This is quite explicitly discussed in the first paragraph of results (page 4).

      (10) The magnitude of these movements appears relatively small (0.5 micron). Can the authors discuss whether such small but reproducible displacements are likely to be biologically meaningful in terms of nuclear function or gene regulation?

      At the moment, our experimental set up allows us to analyze the dynamics of only a small portion of chr1, which indeed shows an average 0.5 micron displacement towards the nuclear periphery. Based on the chromosome painting data from fixed cells, the displacement at the level of whole chromosome is significantly larger. As mentioned in the discussion (page 13), the functional implications of radial repositioning of chromosomes upon serum starvation is not known. Therefore further discussion on the relevance of the magnitude reported here would be pure speculation.

      (11) Peripheral homologs of chromosome 1 became faster and more dynamic under starvation. Why might these loci be more prone to movement? Could this be linked to differences in transcriptional activity or chromatin state between central and peripheral homologs?

      At the moment we favour the idea that the central homolog is constrained by its anchorage to the nucleolus (Figure 6I). However, transcriptional activity and/or chromatin state may also play a role, and this possibility is now mentioned in the discussion on page 14.

      Minor points:

      (1) Figure legends use inconsistent capitalization and panel labels. These should be standardized across all figures for better readability.

      We apologize for these inconsistencies, and have aimed to standardize all labeling.

      References

      Mehta, I.S., Amira, M., Harvey, A.J., and Bridger, J.M. (2010). Rapid chromosome territory relocation by nuclear motor activity in response to serum removal in primary human fibroblasts. Genome Biol 11, R5.

    1. eLife Assessment

      This important study reports insights into how the caspase Dcp-1, best known for cell death, can also promote tissue growth in Drosophila, extending the authors' earlier work by identifying regulatory factors that shape this non-lethal activity. The compelling findings identify a physical and functional interaction between Dcp-1 and Bruce, as well as new Dcp-1-interacting proteins that function in autophagy: Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a. This work helps broaden the understanding of the non-lethal roles of Dcp-1.

    2. Reviewer #1 (Public review):

      The authors clearly demonstrate that overexpressed Dcp-1, but not Drice, is activated without canonical apoptosome components.

      Using TurboID-based proximity labeling they revealed distinct proximal proteomes, among which Sirtuin 1, an Atg8a deacetylase, which promotes autophagy, was specifically required for Dcp-1 activation. Additionally, the show that autophagy-related genes, including Bcl-2 family members Debcl and Buffy, are required for Dcp-1 activation. Using structure-based prediction using AlphaFold3 they identified that Bruce, an autophagy-regulated inhibitor of apoptosis, as a Dcp-1-specific regulator acting outside the apoptosome-mediated pathway. Finally, they show that Bruce suppresses wing tissue growth. These findings indicate that non-lethal Dcp-1 activity is governed by the autophagy- Bruce axis, enabling distinct non-lethal functions independent of cell death.

      Comments on revised version.

      No further comments.

    3. Reviewer #2 (Public review):

      Summary:

      The Drosophila executioner caspase Dcp-1 has established roles in cell death, autophagy, and imaginal disc growth. This study reports previously unrecognized factors that work together with Dcp-1. Specifically, the authors performed a turboID-based proximal ligation experiment to identify factors associated Dcp-1 and Drice. Dcp-1-specific interactors were further examined for their genetic interaction. The authors report autophagy-related genes, including Debcl and Buffy, to be required for Dcp-1 activation. In addition, the authors present evidence of an interaction between Bruce and Dcp-1. Bruce expression blocks the Dcp-1 overexpression phenotype. Inhibition of effector caspases or overexpression of Bruce commonly reduced wing growth, suggesting a relationship between the two proteins.

      Strengths:

      The study identifies new Dcp-1-interacting proteins and provides a functional link between Dcp-1 and Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a. During the revision, the authors have also added convincing new data supporting the interaction between Dcp-1 and Bruce. They further make a strong case regarding the quality of the turboID-proteomics data. Overall, this is a strong manuscript supporting an interesting discovery.